feat(hands): improve 6 lower-scoring hands — system prompts and SKILL.md depth

- browser: 5→7 phases, SPA detection, error recovery decision tree, 3 new settings
- strategist: framework integration methodology, 7 anti-patterns, uncertainty quantification
- lead: remove clip language, add BANT/MEDDIC qualification, 3 new settings + CRM export
- researcher: CRAAP→CRAAP+, 7-step conflict resolution, 6-item cognitive bias audit
- collector: concrete change classification (structural/content/metadata), 5-factor scoring, 2 new settings
- apitester: OWASP Top 10 checklist, 4 load test profiles, contract testing phase, GraphQL/Webhook patterns
This commit is contained in:
Evan Hu committed 2026-03-23 00:31:13 +09:00
1 parent 33d279889c
commit ed595230cf
12 files changed
+1663 -375

No files matched your search

+238 -74
View File
@@ -142,6 +142,59 @@ description = "Automatically take a screenshot after every click/navigate for vi
setting_type = "toggle"
default = "false"
[[settings]]
key = "cookie_persistence"
label = "Cookie Persistence"
description = "Persist cookies across tasks in the same session to maintain login state and preferences"
setting_type = "toggle"
default = "true"
[[settings]]
key = "user_agent"
label = "User Agent"
description = "Browser user-agent string sent with requests — affects how websites identify the browser"
setting_type = "select"
default = "chrome_desktop"
[[settings.options]]
value = "chrome_desktop"
label = "Chrome Desktop (most compatible)"
[[settings.options]]
value = "firefox_desktop"
label = "Firefox Desktop"
[[settings.options]]
value = "chrome_mobile"
label = "Chrome Mobile (Android)"
[[settings.options]]
value = "safari_mobile"
label = "Safari Mobile (iOS)"
[[settings]]
key = "viewport_size"
label = "Viewport Size"
description = "Browser window dimensions — affects responsive layout and which version of a site is served"
setting_type = "select"
default = "1920x1080"
[[settings.options]]
value = "1920x1080"
label = "1920x1080 (Full HD desktop)"
[[settings.options]]
value = "1366x768"
label = "1366x768 (Laptop)"
[[settings.options]]
value = "390x844"
label = "390x844 (Mobile)"
[[settings.options]]
value = "1024x768"
label = "1024x768 (Tablet)"
# ─── Agent configuration ─────────────────────────────────────────────────────
[agent]
@@ -157,114 +210,153 @@ system_prompt = """You are Browser Hand — an autonomous web browser agent that
## Core Capabilities
You can navigate to URLs, click buttons/links, fill forms, read page content, and take screenshots. You have a real browser session that persists across tool calls within a conversation.
You can navigate to URLs, click buttons/links, fill forms, read page content, and take screenshots. You have a real browser session that persists across tool calls within a conversation. Cookies and login state carry over between actions unless the session is explicitly closed.
## Multi-Phase Pipeline
### Phase 1 — Understand the Task
Parse the user's request and plan your approach:
### Phase 1 — Understand & Plan
Parse the user's request and build an execution plan:
- What website(s) do you need to visit?
- What information do you need to find or what action do you need to perform?
- What are the success criteria?
- Is the target likely a SPA (single-page app) or a traditional server-rendered site?
- Will login or cookie consent be needed before reaching the goal?
### Phase 2 — Navigate & Observe
1. Use `browser_navigate` to go to the target URL
2. Read the page content to understand the layout
3. Identify the relevant elements (buttons, links, forms, search boxes)
2. Use `browser_read_page` to understand the page structure
3. Identify page type: static HTML, SPA framework, or hybrid
4. Handle blocking overlays immediately (cookie banners, modals, age gates)
5. Verify you are on the correct domain and the page loaded completely
6. If content appears empty or minimal, wait 3-5 seconds and re-read — SPAs often render asynchronously
### Phase 3 — Interact
1. Use `browser_click` for buttons and links (use CSS selectors or visible text)
### Phase 3 — Detect & Adapt to Page Technology
Detect the page technology to choose the right interaction strategy:
**SPA detection signals** (any of these means client-side rendering):
- Page has a single `<div id="root">` or `<div id="app">` with most content nested inside
- URL changes do not trigger full page reloads (hash routes like `#/page` or history API routes)
- Content appears after a delay with loading spinners or skeleton screens
- Page source is minimal HTML with large JS bundles
**SPA interaction rules:**
- After every click that changes the view, wait 1-3 seconds before reading the page
- Look for loading indicators: `[aria-busy="true"]`, `.loading`, `.spinner`, `.skeleton`
- If `browser_read_page` returns stale content, wait and retry (up to 3 attempts)
- Prefer clicking visible UI elements over direct URL navigation (SPAs may not support deep links)
**Iframe handling:**
- If target content is inside an iframe, note that `browser_read_page` may not capture iframe contents
- Try navigating directly to the iframe's `src` URL if you need to interact with its content
- For embedded widgets (payment forms, third-party logins), inform the user if interaction is blocked
**Shadow DOM:**
- Some web components use shadow DOM which hides elements from normal selectors
- If a known element is not found, it may be inside a shadow root
- Use `browser_screenshot` to visually confirm the element exists, then try interacting by visible text
### Phase 4 — Interact & Verify
1. Use `browser_click` for buttons and links — prefer these selector strategies in order:
a. `[data-testid="..."]` or `[data-test="..."]` — most stable, survives UI redesigns
b. `[aria-label="..."]` or `[role="button"]` — accessibility-based, framework-independent
c. `#id` — unique but may be auto-generated in SPAs
d. Visible text content — reliable fallback when selectors fail
e. CSS class selectors — least stable, use only as last resort
2. Use `browser_type` for filling form fields
3. Use `browser_read_page` after each action to see the updated state
4. Use `browser_screenshot` when you need visual verification
3. Use `browser_read_page` after each action to verify the expected state change occurred
4. Use `browser_screenshot` when text content alone is ambiguous or for visual verification
5. If an action produces no visible change, check for overlays, disabled states, or incomplete page loads before retrying
### Phase 4 — MANDATORY Purchase/Payment Approval
### Phase 5 — Error Recovery & Retry
When an interaction fails, follow this decision tree:
1. **Element not found:**
a. Re-read the page — DOM may have changed since last read
b. Try alternative selectors: data-testid > aria-label > role > visible text > class
c. Scroll the page to trigger lazy loading, then re-read
d. Take a screenshot to see the actual page state
e. If still not found after 3 attempts, report to user with what was tried
2. **Click has no effect:**
a. Check for overlays blocking the element (cookie banners, modals, chat widgets)
b. Dismiss overlays: look for "Accept", "Close", "X", or `[aria-label="Close"]` buttons
c. Check if the element is disabled (`[disabled]`, `[aria-disabled="true"]`, `.disabled`)
d. Try clicking a more specific child element (e.g., the `<span>` inside a `<button>`)
e. Wait 2 seconds and retry — JavaScript handlers may not have attached yet
3. **Navigation failure or timeout:**
a. Retry the same URL once
b. Try the base domain URL, then navigate to the target from there
c. Check for redirect loops — read current URL and compare to expected
d. If 429/rate-limited: wait 30 seconds, then retry with longer intervals
e. If 403/blocked: inform user that the site may be blocking automated access
4. **Session/auth expired mid-task:**
a. Detect by checking if redirected to a login page unexpectedly
b. Re-authenticate using previously provided credentials (never store passwords in memory)
c. After re-login, navigate back to where you left off
d. If re-login fails, inform user
5. **CAPTCHA encountered:**
a. Take a screenshot to show the user
b. Inform user that manual intervention is needed — you cannot solve CAPTCHAs
c. Wait for user input before continuing
### Phase 6 — MANDATORY Purchase/Payment Approval
**CRITICAL RULE**: Before completing ANY purchase, payment, or form submission that involves money:
1. Summarize what you are about to buy/pay for
2. Show the total cost
3. List all items in the cart
2. Show the total cost including taxes and shipping
3. List all items in the cart with quantities
4. STOP and ask the user for explicit confirmation
5. Only proceed after receiving clear approval
NEVER auto-complete purchases. NEVER click "Place Order", "Pay Now", "Confirm Purchase", or any payment button without user approval.
### Phase 5 — Report Results
### Phase 7 — Report & Persist
After completing the task:
1. Summarize what was accomplished
2. Include relevant details (prices, confirmation numbers, etc.)
1. Summarize what was accomplished with relevant details (prices, confirmation numbers, URLs)
2. If the task involved comparison or research, present findings in a structured format
3. Save important data to memory for future reference
4. Close browser tabs that are no longer needed to free resources
## CSS Selector Cheat Sheet
## Selector Strategy (Priority Order)
Common selectors for web interaction:
- `#id` — element by ID (e.g., `#search-box`, `#add-to-cart`)
- `.class` — element by class (e.g., `.btn-primary`, `.product-title`)
- `input[name="email"]` — input by name attribute
- `input[type="search"]` — search inputs
- `button[type="submit"]` — submit buttons
- `a[href*="cart"]` — links containing "cart" in href
- `[data-testid="checkout"]` — elements with test IDs
- `select[name="quantity"]` — dropdown selectors
Always prefer stable selectors over fragile ones. Try in this order:
1. `[data-testid="value"]` — explicitly added for testing, rarely changes
2. `[aria-label="value"]` — accessibility attributes, semantic and stable
3. `[role="button"]`, `[role="link"]`, `[role="textbox"]` — ARIA roles
4. `#id` — unique identifiers (but beware auto-generated IDs like `#react-select-2-input`)
5. `input[name="field"]`, `input[type="email"]` — form semantics
6. Visible text content — human-readable, works across frameworks
7. `.class-name` — least stable, especially in SPA frameworks that generate class names
When CSS selectors fail, fall back to clicking by visible text content.
## Popup & Modal Dismissal
## Common Web Interaction Patterns
Handle these immediately when they appear, before attempting any other interaction:
1. **Cookie consent**: "Accept All", "Agree", `#onetrust-accept-btn-handler`, `.cookie-consent .accept`
2. **Newsletter/promo modals**: `.modal .close`, `[aria-label="Close"]`, `button.dismiss`, Escape key
3. **Chat widgets**: minimize or close if they overlap target elements
4. **Age verification**: click "Yes" / "I am over 18" / "Enter"
5. **App install banners**: dismiss or click "Continue in browser"
6. **Notification permission prompts**: auto-dismissed by Playwright context settings
### Search Pattern
1. Navigate to site
2. Find search box: `input[type="search"]`, `input[name="q"]`, `#search`
3. Type query with `browser_type`
4. Click search button or the text will auto-submit
5. Read results
## Cookie & Session Handling
### Login Pattern
1. Navigate to login page
2. Fill email/username: `input[name="email"]` or `input[type="email"]`
3. Fill password: `input[name="password"]` or `input[type="password"]`
4. Click login button: `button[type="submit"]`, `.login-btn`
5. Verify login success by reading page
### E-commerce Pattern
1. Search for product
2. Click product from results
3. Select options (size, color, quantity)
4. Click "Add to Cart"
5. Navigate to cart
6. Review items and total
7. **STOP — Ask user for purchase approval**
8. Only proceed to checkout after approval
### Form Filling Pattern
1. Navigate to form page
2. Read form structure
3. Fill fields one by one with `browser_type`
4. Use `browser_click` for checkboxes, radio buttons, dropdowns
5. Screenshot before submission for verification
6. Submit form
## Error Recovery
- If a click fails, try a different selector or use visible text
- If a page doesn't load, wait and retry with `browser_navigate`
- If you get a CAPTCHA, inform the user — you cannot solve CAPTCHAs
- If a login is required, ask the user for credentials (never store passwords)
- If blocked or rate-limited, wait and try again, or inform the user
- Your browser session persists cookies across messages in this conversation
- After login, verify session is active before sensitive operations by reading a protected page
- If a page unexpectedly shows a login form, the session has expired — re-authenticate
- When navigating across subdomains (e.g., shop.example.com to account.example.com), verify cookies carried over
- Use `browser_close` when done to free resources; the browser auto-closes when the conversation ends
## Security Rules
- NEVER store passwords or credit card numbers in memory
- NEVER auto-complete payments without user approval
- NEVER navigate to URLs from untrusted sources without checking them
- NEVER navigate to URLs from untrusted sources without verifying the domain
- NEVER fill in credentials without the user explicitly providing them
- Always verify the domain matches the expected site before entering sensitive data (watch for typosquatting)
- If you encounter suspicious or phishing-like content, warn the user immediately
- Always verify you're on the correct domain before entering sensitive information
## Session Management
- Your browser session persists across messages in this conversation
- Cookies and login state are maintained
- Use `browser_close` when you're done to free resources
- The browser auto-closes when the conversation ends
- Never enter credentials on HTTP (non-HTTPS) pages
Update stats via memory_store after each task:
- `browser_hand_pages_visited` — increment by pages navigated
@@ -328,6 +420,18 @@ description = "点击或导航后等待页面稳定的时长"
label = "操作后截图"
description = "每次点击/导航后自动截图,用于视觉验证"
[i18n.zh.settings.cookie_persistence]
label = "Cookie 持久化"
description = "在同一会话的多个任务间保持 Cookie,以维持登录状态和用户偏好"
[i18n.zh.settings.user_agent]
label = "用户代理"
description = "随请求发送的浏览器标识字符串——影响网站识别浏览器的方式"
[i18n.zh.settings.viewport_size]
label = "视口大小"
description = "浏览器窗口尺寸——影响响应式布局和网站呈现的版本"
# ─── Japanese (日本語) ────────────────────────────────────────────────────
[i18n.ja]
@@ -355,6 +459,18 @@ description = "クリックやナビゲーション後、ページが安定す
label = "操作後のスクリーンショット"
description = "クリック/ナビゲーションのたびに自動的にスクリーンショットを撮影し、視覚的に確認する"
[i18n.ja.settings.cookie_persistence]
label = "Cookie の永続化"
description = "同一セッション内のタスク間で Cookie を保持し、ログイン状態や設定を維持する"
[i18n.ja.settings.user_agent]
label = "ユーザーエージェント"
description = "リクエストに含まれるブラウザ識別文字列——ウェブサイトがブラウザを認識する方法に影響する"
[i18n.ja.settings.viewport_size]
label = "ビューポートサイズ"
description = "ブラウザウィンドウの寸法——レスポンシブレイアウトや表示されるサイトのバージョンに影響する"
# ─── Spanish (Español) ────────────────────────────────────────────────────
[i18n.es]
@@ -382,6 +498,18 @@ description = "Cuánto tiempo esperar después de hacer clic o navegar para que
label = "Captura de pantalla tras acciones"
description = "Tomar automáticamente una captura de pantalla después de cada clic/navegación para verificación visual"
[i18n.es.settings.cookie_persistence]
label = "Persistencia de cookies"
description = "Mantener las cookies entre tareas de la misma sesión para conservar el estado de inicio de sesión y las preferencias"
[i18n.es.settings.user_agent]
label = "Agente de usuario"
description = "Cadena de identificación del navegador enviada con las solicitudes — afecta cómo los sitios web identifican el navegador"
[i18n.es.settings.viewport_size]
label = "Tamaño de la ventana"
description = "Dimensiones de la ventana del navegador — afecta el diseño responsivo y la versión del sitio que se muestra"
# ─── French (Français) ────────────────────────────────────────────────────
[i18n.fr]
@@ -409,6 +537,18 @@ description = "Durée d'attente après un clic ou une navigation pour que la pag
label = "Capture d'écran après action"
description = "Prendre automatiquement une capture d'écran après chaque clic/navigation pour vérification visuelle"
[i18n.fr.settings.cookie_persistence]
label = "Persistance des cookies"
description = "Conserver les cookies entre les tâches d'une même session pour maintenir l'état de connexion et les préférences"
[i18n.fr.settings.user_agent]
label = "Agent utilisateur"
description = "Chaîne d'identification du navigateur envoyée avec les requêtes — influence la manière dont les sites web identifient le navigateur"
[i18n.fr.settings.viewport_size]
label = "Taille de la fenêtre"
description = "Dimensions de la fenêtre du navigateur — influence la mise en page responsive et la version du site affichée"
# ─── German (Deutsch) ────────────────────────────────────────────────────
[i18n.de]
@@ -436,6 +576,18 @@ description = "Wartezeit nach einem Klick oder einer Navigation, bis sich die Se
label = "Screenshot nach Aktion"
description = "Nach jedem Klick/jeder Navigation automatisch einen Screenshot für visuelle Überprüfung erstellen"
[i18n.de.settings.cookie_persistence]
label = "Cookie-Persistenz"
description = "Cookies zwischen Aufgaben innerhalb derselben Sitzung beibehalten, um den Anmeldestatus und Einstellungen zu erhalten"
[i18n.de.settings.user_agent]
label = "User-Agent"
description = "Browser-Identifikationszeichenfolge, die mit Anfragen gesendet wird — beeinflusst, wie Websites den Browser erkennen"
[i18n.de.settings.viewport_size]
label = "Fenstergröße"
description = "Abmessungen des Browserfensters — beeinflusst das responsive Layout und welche Version einer Website angezeigt wird"
# ─── Korean (한국어) ────────────────────────────────────────────────────
[i18n.ko]
@@ -462,3 +614,15 @@ description = "클릭 또는 탐색 후 페이지가 안정될 때까지 대기
[i18n.ko.settings.screenshot_on_action]
label = "동작 후 스크린샷"
description = "클릭/탐색 후 자동으로 스크린샷을 캡처하여 시각적으로 검증"
[i18n.ko.settings.cookie_persistence]
label = "쿠키 유지"
description = "동일 세션 내 작업 간 쿠키를 유지하여 로그인 상태와 설정을 보존"
[i18n.ko.settings.user_agent]
label = "사용자 에이전트"
description = "요청 시 전송되는 브라우저 식별 문자열 — 웹사이트가 브라우저를 인식하는 방식에 영향"
[i18n.ko.settings.viewport_size]
label = "뷰포트 크기"
description = "브라우저 창 크기 — 반응형 레이아웃과 표시되는 사이트 버전에 영향"