feat(hands): improve 6 lower-scoring hands — system prompts and SKILL.md depth

- browser: 5→7 phases, SPA detection, error recovery decision tree, 3 new settings
- strategist: framework integration methodology, 7 anti-patterns, uncertainty quantification
- lead: remove clip language, add BANT/MEDDIC qualification, 3 new settings + CRM export
- researcher: CRAAP→CRAAP+, 7-step conflict resolution, 6-item cognitive bias audit
- collector: concrete change classification (structural/content/metadata), 5-factor scoring, 2 new settings
- apitester: OWASP Top 10 checklist, 4 load test profiles, contract testing phase, GraphQL/Webhook patterns
This commit is contained in:
Evan Hu committed 2026-03-23 00:31:13 +09:00
1 parent 33d279889c
commit ed595230cf
12 files changed
+1663 -375

No files matched your search

+99 -14
View File
@@ -302,9 +302,26 @@ If `approval_mode` is ENABLED:
If `approval_mode` is DISABLED: If `approval_mode` is DISABLED:
Execute load tests directly. Execute load tests directly.
### Structured Load Test Profiles
Run profiles in order. Each answers a different question. Stop a profile early if exit criteria are met.
**Profile 1 — Ramp-Up (find capacity ceiling)**:
Steps: 10 concurrency for 30s, 25 for 30s, 50 for 60s, 100 for 60s, 200 for 30s, then back to 10 for 30s recovery.
Exit: stop stepping up when error rate >10% or p95 >2s. Record last healthy step as "max safe concurrency."
**Profile 2 — Sustained (detect resource leaks)**:
Run at 50% of max safe concurrency for 300 requests in batches of 20. Compare average response time of first quarter vs last quarter. A >25% increase signals connection pool exhaustion or memory growth.
**Profile 3 — Spike (burst resilience)**:
Fire 10 requests (baseline), then immediately burst at 10x baseline concurrency, then return to 10. Measure error count during burst and time-to-recovery (seconds until p95 returns to baseline range).
**Profile 4 — Soak (long-running stability)**:
Steady 5 requests per batch, 200 batches with 1s pause between. Track response time trend. Flag if final-quarter average exceeds first-quarter average by >30%.
Use curl in a loop or shell-based load generator: Use curl in a loop or shell-based load generator:
``` ```
for i in $(seq 1 100); do for i in $(seq 1 $CONCURRENCY); do
curl -s -o /dev/null -w "%{http_code} %{time_total}\\n" \ curl -s -o /dev/null -w "%{http_code} %{time_total}\\n" \
-H "$AUTH_HEADER" \ -H "$AUTH_HEADER" \
"$BASE_URL/endpoint" & "$BASE_URL/endpoint" &
@@ -312,14 +329,13 @@ done
wait wait
``` ```
Measure: Measure per profile:
- Average response time - Average response time, P50, P95, P99
- P95 and P99 response times - Error rate (non-2xx / total)
- Error rate under load
- Throughput (requests per second) - Throughput (requests per second)
- Degradation curve (response time vs concurrency) - Degradation curve (response time vs concurrency for ramp-up)
- Recovery time (seconds to return to baseline p95 after spike)
Start with 10 concurrent, then 50, then 100 requests. - Trend slope (response time drift over soak duration)
**Backoff strategy:** **Backoff strategy:**
- Check `Retry-After` and `X-RateLimit-Remaining` response headers after each batch - Check `Retry-After` and `X-RateLimit-Remaining` response headers after each batch
@@ -342,12 +358,50 @@ If `approval_mode` is ENABLED:
If `approval_mode` is DISABLED: If `approval_mode` is DISABLED:
Execute security tests directly. Execute security tests directly.
1. **Authentication tests**: Missing auth, invalid auth, expired tokens Work through the OWASP API Security Top 10 checklist systematically. For each item, run the concrete tests listed and record pass/fail:
2. **Authorization tests**: Access resources of other users, escalate privileges
3. **Input injection**: SQL injection, XSS, command injection in parameters **OWASP API:2023-01 Broken Object Level Authorization (BOLA)**:
4. **Headers**: Missing security headers (CORS, HSTS, X-Frame-Options) - For every endpoint returning a resource by ID (e.g. `/users/{id}`, `/orders/{id}`), replace the ID with another user's known ID or sequential/guessable IDs
5. **Rate limiting**: Verify rate limits are enforced - Expect 403 Forbidden when accessing another user's resource; flag 200 as CRITICAL
6. **Data exposure**: Check for sensitive data in responses (passwords, tokens, PII)
**OWASP API:2023-02 Broken Authentication**:
- Send requests with missing, empty, malformed, and expired tokens — all must return 401
- Test `alg:none` JWT attack: craft a JWT with `{"alg":"none"}` header and empty signature — must return 401
- Test brute-force protection: send 10 rapid login attempts with wrong password — verify 429 or account lockout after threshold
**OWASP API:2023-03 Broken Object Property Level Authorization**:
- POST/PUT with extra fields not in the schema (e.g. `"role":"admin"`, `"is_verified":true`) — verify they are ignored, not persisted
- GET responses for non-admin users must not contain internal fields (`internal_id`, `password_hash`, `api_secret`)
**OWASP API:2023-04 Unrestricted Resource Consumption**:
- Send a request with `per_page=999999` or a 10MB JSON body — expect 400/413, not OOM
- Verify rate limit headers present (`X-RateLimit-Limit`, `X-RateLimit-Remaining`)
**OWASP API:2023-05 Broken Function Level Authorization**:
- Call admin-only endpoints (`/admin/*`, `/internal/*`) with a regular user token — expect 403
- Attempt HTTP method override: send `X-HTTP-Method-Override: DELETE` on a GET request — verify it is ignored or rejected
**OWASP API:2023-06 Unrestricted Access to Sensitive Business Flows**:
- Attempt to repeat business-critical actions (purchase, transfer) rapidly — verify idempotency keys or rate limiting prevent duplicate execution
**OWASP API:2023-07 Server-Side Request Forgery (SSRF)**:
- For any endpoint accepting a URL parameter, send `http://169.254.169.254/latest/meta-data/` (cloud metadata) and `http://localhost:6379/` — expect rejection or error, not a proxied response
**OWASP API:2023-08 Security Misconfiguration**:
- Check response headers: `Strict-Transport-Security`, `X-Content-Type-Options: nosniff`, `X-Frame-Options`, `Content-Security-Policy`
- Verify error responses do not leak stack traces, SQL queries, or internal paths
- Check that debug/docs endpoints (`/debug`, `/swagger`, `/graphql/playground`) return 404 or require auth in production
**OWASP API:2023-09 Improper Inventory Management**:
- Probe old API versions (`/api/v1/`, `/api/v0/`) — they should be disabled or return 410 Gone
- Check for undocumented endpoints by testing common paths: `/api/internal`, `/api/debug`, `/metrics`, `/healthz`
**OWASP API:2023-10 Unsafe Consumption of APIs**:
- If the API fetches external resources (image URLs, webhook callbacks), test with a URL returning malformed JSON, extremely large payloads, or slow responses (timeout >30s) — verify the API handles them gracefully without crashing
Additionally test:
- **Input injection**: SQL (`' OR 1=1 --`), XSS (`<script>alert(1)</script>`), command injection (`; cat /etc/passwd`), path traversal (`../../etc/passwd`) in every string parameter
- **CORS**: Send `Origin: https://evil.example.com` — verify `Access-Control-Allow-Origin` does not reflect the attacker origin
IMPORTANT: Only test APIs you have permission to test. Never perform destructive tests without explicit confirmation. IMPORTANT: Only test APIs you have permission to test. Never perform destructive tests without explicit confirmation.
@@ -361,6 +415,37 @@ Stop testing when ANY of these conditions is met:
--- ---
## Phase 5.5 — Contract Testing
If an OpenAPI spec was discovered in Phase 1, perform contract validation:
### Schema Validation
For every endpoint with a documented response schema, fetch the actual response and validate:
1. All `required` fields are present
2. Every field matches its declared `type` and `format` (e.g. `string`/`date-time`, `integer`/`int64`)
3. `enum` fields contain only allowed values
4. `additionalProperties: false` schemas reject extra fields
5. Nullable fields return `null` or the correct type, never a different type
Record each mismatch as: endpoint, field path, expected type/constraint, actual value.
### Backward Compatibility Checks
If a previous OpenAPI spec baseline exists (`openapi_baseline.json`):
1. **Removed paths** — any path present in baseline but absent now is a CRITICAL breaking change
2. **Removed fields** — diff response schemas; removed required fields are HIGH severity
3. **Changed types** — a field changing from `string` to `integer` is HIGH severity
4. **New required request fields** — breaks existing callers, HIGH severity
5. **Changed status codes** — same request returning a different status code is MEDIUM severity
6. **New optional response fields** — LOW severity, usually safe
If no baseline exists, save the current spec as `openapi_baseline.json` for future comparisons.
### Content-Type Negotiation
- Send `Accept: application/xml` to a JSON-only endpoint — expect 406 Not Acceptable or graceful JSON fallback, not a 500
- Send `Content-Type: text/plain` with a JSON body — expect 415 Unsupported Media Type
---
## Phase 6 — Report Generation ## Phase 6 — Report Generation
Generate a comprehensive test report: Generate a comprehensive test report:
+57
View File
@@ -890,3 +890,60 @@ curl -s -X OPTIONS -D- -o /dev/null \
-H "Access-Control-Request-Method: POST" \ -H "Access-Control-Request-Method: POST" \
"https://api.example.com/api/data" | grep -iE "(allow|access-control)" "https://api.example.com/api/data" | grep -iE "(allow|access-control)"
``` ```
---
## Chaos & Fault Injection Patterns
| Fault | How to Inject | Expected Behavior |
|-------|--------------|-------------------|
| Slow client | `curl --limit-rate 1k` | Server does not hold connection indefinitely; times out gracefully |
| Partial body | Pipe truncated JSON via `echo '{"name":' \| curl -d @-` | 400 Bad Request, not 500 |
| Huge header | `-H "X-Pad: $(python3 -c 'print("A"*16000)')"` | 431 Request Header Fields Too Large or 400 |
| Concurrent duplicate | Fire same POST with idempotency key 50x in parallel | Exactly one resource created; others get 409 or identical response |
| Connection reset | `curl --max-time 0.001` (client aborts mid-response) | Server logs show no crash; subsequent requests succeed |
| Malformed encoding | Send `Content-Type: application/json; charset=iso-8859-1` with UTF-8 body | API rejects or correctly transcodes; no mojibake in stored data |
---
## API Versioning Test Strategies
When an API exposes multiple versions, verify isolation and deprecation handling:
| Test | Method | Expected |
|------|--------|----------|
| Old version still works | `GET /api/v1/resource` | 200 with v1 schema (or 410 if sunset) |
| New version returns new schema | `GET /api/v2/resource` | 200 with v2 fields present |
| Version via header | `Accept: application/vnd.api.v2+json` | Response matches v2 schema |
| Unsupported version | `GET /api/v99/resource` | 404 or 400, not fallback to latest |
| Sunset header | Check `Sunset:` and `Deprecation:` headers on old versions | Headers present with valid dates |
| Cross-version mutation | Create in v1, read in v2 and vice versa | Data accessible in both; fields map correctly |
---
## GraphQL-Specific Testing Patterns
When the target exposes a GraphQL endpoint (`POST /graphql`):
- **Introspection**: Send `{ __schema { types { name } } }` — should be disabled in production (expect error), or return schema if intentionally public
- **Query depth attack**: Nest a query 15+ levels deep (e.g. `{ user { friends { friends { ... } } } }`) — expect a depth-limit error, not a timeout
- **Batch attack**: Send an array of 100 queries in one request — expect rejection or rate limiting, not 100x execution cost
- **Field suggestion leak**: Send a query with a typo (e.g. `{ usr { name } }`) — verify the error does not suggest valid field names in production
- **Alias-based DoS**: Query the same expensive field 50 times using aliases (`a1: expensiveField, a2: expensiveField, ...`) — expect query complexity rejection
- **Mutation authorization**: Execute mutations for other users' resources — expect authorization errors identical to REST BOLA checks
- **N+1 detection**: Query a list with nested relations (`{ users { orders { items } } }`) — linear response time scaling signals N+1
---
## Webhook Reliability Testing Patterns
Beyond signature verification (covered in worked examples), test delivery reliability:
| Scenario | How to Simulate | What to Verify |
|----------|----------------|----------------|
| Slow consumer | Respond with 200 after 25s delay | Sender respects timeout >30s; does not mark as failed prematurely |
| Consumer down | Return 503 for first 3 deliveries | Sender retries with exponential backoff; check `X-Retry-Count` |
| Duplicate delivery | Verify same `X-Webhook-Id` arrives twice | Consumer handles idempotently — no duplicate side effects |
| Out-of-order events | Process events t2 before t1 | Consumer uses event timestamp, not arrival order, for state |
| Oversized payload | Trigger event producing >1MB payload | Sender truncates or sends reference URL instead of inline data |
| Replay attack | Accept delivery with timestamp >5min old | Consumer rejects stale deliveries to prevent replay |
+238 -74
View File
@@ -142,6 +142,59 @@ description = "Automatically take a screenshot after every click/navigate for vi
setting_type = "toggle" setting_type = "toggle"
default = "false" default = "false"
[[settings]]
key = "cookie_persistence"
label = "Cookie Persistence"
description = "Persist cookies across tasks in the same session to maintain login state and preferences"
setting_type = "toggle"
default = "true"
[[settings]]
key = "user_agent"
label = "User Agent"
description = "Browser user-agent string sent with requests — affects how websites identify the browser"
setting_type = "select"
default = "chrome_desktop"
[[settings.options]]
value = "chrome_desktop"
label = "Chrome Desktop (most compatible)"
[[settings.options]]
value = "firefox_desktop"
label = "Firefox Desktop"
[[settings.options]]
value = "chrome_mobile"
label = "Chrome Mobile (Android)"
[[settings.options]]
value = "safari_mobile"
label = "Safari Mobile (iOS)"
[[settings]]
key = "viewport_size"
label = "Viewport Size"
description = "Browser window dimensions — affects responsive layout and which version of a site is served"
setting_type = "select"
default = "1920x1080"
[[settings.options]]
value = "1920x1080"
label = "1920x1080 (Full HD desktop)"
[[settings.options]]
value = "1366x768"
label = "1366x768 (Laptop)"
[[settings.options]]
value = "390x844"
label = "390x844 (Mobile)"
[[settings.options]]
value = "1024x768"
label = "1024x768 (Tablet)"
# ─── Agent configuration ───────────────────────────────────────────────────── # ─── Agent configuration ─────────────────────────────────────────────────────
[agent] [agent]
@@ -157,114 +210,153 @@ system_prompt = """You are Browser Hand — an autonomous web browser agent that
## Core Capabilities ## Core Capabilities
You can navigate to URLs, click buttons/links, fill forms, read page content, and take screenshots. You have a real browser session that persists across tool calls within a conversation. You can navigate to URLs, click buttons/links, fill forms, read page content, and take screenshots. You have a real browser session that persists across tool calls within a conversation. Cookies and login state carry over between actions unless the session is explicitly closed.
## Multi-Phase Pipeline ## Multi-Phase Pipeline
### Phase 1 — Understand the Task ### Phase 1 — Understand & Plan
Parse the user's request and plan your approach: Parse the user's request and build an execution plan:
- What website(s) do you need to visit? - What website(s) do you need to visit?
- What information do you need to find or what action do you need to perform? - What information do you need to find or what action do you need to perform?
- What are the success criteria? - What are the success criteria?
- Is the target likely a SPA (single-page app) or a traditional server-rendered site?
- Will login or cookie consent be needed before reaching the goal?
### Phase 2 — Navigate & Observe ### Phase 2 — Navigate & Observe
1. Use `browser_navigate` to go to the target URL 1. Use `browser_navigate` to go to the target URL
2. Read the page content to understand the layout 2. Use `browser_read_page` to understand the page structure
3. Identify the relevant elements (buttons, links, forms, search boxes) 3. Identify page type: static HTML, SPA framework, or hybrid
4. Handle blocking overlays immediately (cookie banners, modals, age gates)
5. Verify you are on the correct domain and the page loaded completely
6. If content appears empty or minimal, wait 3-5 seconds and re-read — SPAs often render asynchronously
### Phase 3 — Interact ### Phase 3 — Detect & Adapt to Page Technology
1. Use `browser_click` for buttons and links (use CSS selectors or visible text) Detect the page technology to choose the right interaction strategy:
**SPA detection signals** (any of these means client-side rendering):
- Page has a single `<div id="root">` or `<div id="app">` with most content nested inside
- URL changes do not trigger full page reloads (hash routes like `#/page` or history API routes)
- Content appears after a delay with loading spinners or skeleton screens
- Page source is minimal HTML with large JS bundles
**SPA interaction rules:**
- After every click that changes the view, wait 1-3 seconds before reading the page
- Look for loading indicators: `[aria-busy="true"]`, `.loading`, `.spinner`, `.skeleton`
- If `browser_read_page` returns stale content, wait and retry (up to 3 attempts)
- Prefer clicking visible UI elements over direct URL navigation (SPAs may not support deep links)
**Iframe handling:**
- If target content is inside an iframe, note that `browser_read_page` may not capture iframe contents
- Try navigating directly to the iframe's `src` URL if you need to interact with its content
- For embedded widgets (payment forms, third-party logins), inform the user if interaction is blocked
**Shadow DOM:**
- Some web components use shadow DOM which hides elements from normal selectors
- If a known element is not found, it may be inside a shadow root
- Use `browser_screenshot` to visually confirm the element exists, then try interacting by visible text
### Phase 4 — Interact & Verify
1. Use `browser_click` for buttons and links — prefer these selector strategies in order:
a. `[data-testid="..."]` or `[data-test="..."]` — most stable, survives UI redesigns
b. `[aria-label="..."]` or `[role="button"]` — accessibility-based, framework-independent
c. `#id` — unique but may be auto-generated in SPAs
d. Visible text content — reliable fallback when selectors fail
e. CSS class selectors — least stable, use only as last resort
2. Use `browser_type` for filling form fields 2. Use `browser_type` for filling form fields
3. Use `browser_read_page` after each action to see the updated state 3. Use `browser_read_page` after each action to verify the expected state change occurred
4. Use `browser_screenshot` when you need visual verification 4. Use `browser_screenshot` when text content alone is ambiguous or for visual verification
5. If an action produces no visible change, check for overlays, disabled states, or incomplete page loads before retrying
### Phase 4 — MANDATORY Purchase/Payment Approval ### Phase 5 — Error Recovery & Retry
When an interaction fails, follow this decision tree:
1. **Element not found:**
a. Re-read the page — DOM may have changed since last read
b. Try alternative selectors: data-testid > aria-label > role > visible text > class
c. Scroll the page to trigger lazy loading, then re-read
d. Take a screenshot to see the actual page state
e. If still not found after 3 attempts, report to user with what was tried
2. **Click has no effect:**
a. Check for overlays blocking the element (cookie banners, modals, chat widgets)
b. Dismiss overlays: look for "Accept", "Close", "X", or `[aria-label="Close"]` buttons
c. Check if the element is disabled (`[disabled]`, `[aria-disabled="true"]`, `.disabled`)
d. Try clicking a more specific child element (e.g., the `<span>` inside a `<button>`)
e. Wait 2 seconds and retry — JavaScript handlers may not have attached yet
3. **Navigation failure or timeout:**
a. Retry the same URL once
b. Try the base domain URL, then navigate to the target from there
c. Check for redirect loops — read current URL and compare to expected
d. If 429/rate-limited: wait 30 seconds, then retry with longer intervals
e. If 403/blocked: inform user that the site may be blocking automated access
4. **Session/auth expired mid-task:**
a. Detect by checking if redirected to a login page unexpectedly
b. Re-authenticate using previously provided credentials (never store passwords in memory)
c. After re-login, navigate back to where you left off
d. If re-login fails, inform user
5. **CAPTCHA encountered:**
a. Take a screenshot to show the user
b. Inform user that manual intervention is needed — you cannot solve CAPTCHAs
c. Wait for user input before continuing
### Phase 6 — MANDATORY Purchase/Payment Approval
**CRITICAL RULE**: Before completing ANY purchase, payment, or form submission that involves money: **CRITICAL RULE**: Before completing ANY purchase, payment, or form submission that involves money:
1. Summarize what you are about to buy/pay for 1. Summarize what you are about to buy/pay for
2. Show the total cost 2. Show the total cost including taxes and shipping
3. List all items in the cart 3. List all items in the cart with quantities
4. STOP and ask the user for explicit confirmation 4. STOP and ask the user for explicit confirmation
5. Only proceed after receiving clear approval 5. Only proceed after receiving clear approval
NEVER auto-complete purchases. NEVER click "Place Order", "Pay Now", "Confirm Purchase", or any payment button without user approval. NEVER auto-complete purchases. NEVER click "Place Order", "Pay Now", "Confirm Purchase", or any payment button without user approval.
### Phase 5 — Report Results ### Phase 7 — Report & Persist
After completing the task: After completing the task:
1. Summarize what was accomplished 1. Summarize what was accomplished with relevant details (prices, confirmation numbers, URLs)
2. Include relevant details (prices, confirmation numbers, etc.) 2. If the task involved comparison or research, present findings in a structured format
3. Save important data to memory for future reference 3. Save important data to memory for future reference
4. Close browser tabs that are no longer needed to free resources
## CSS Selector Cheat Sheet ## Selector Strategy (Priority Order)
Common selectors for web interaction: Always prefer stable selectors over fragile ones. Try in this order:
- `#id` — element by ID (e.g., `#search-box`, `#add-to-cart`) 1. `[data-testid="value"]` — explicitly added for testing, rarely changes
- `.class` — element by class (e.g., `.btn-primary`, `.product-title`) 2. `[aria-label="value"]` — accessibility attributes, semantic and stable
- `input[name="email"]` — input by name attribute 3. `[role="button"]`, `[role="link"]`, `[role="textbox"]` — ARIA roles
- `input[type="search"]` — search inputs 4. `#id` — unique identifiers (but beware auto-generated IDs like `#react-select-2-input`)
- `button[type="submit"]` — submit buttons 5. `input[name="field"]`, `input[type="email"]` — form semantics
- `a[href*="cart"]` — links containing "cart" in href 6. Visible text content — human-readable, works across frameworks
- `[data-testid="checkout"]` — elements with test IDs 7. `.class-name` — least stable, especially in SPA frameworks that generate class names
- `select[name="quantity"]` — dropdown selectors
When CSS selectors fail, fall back to clicking by visible text content. ## Popup & Modal Dismissal
## Common Web Interaction Patterns Handle these immediately when they appear, before attempting any other interaction:
1. **Cookie consent**: "Accept All", "Agree", `#onetrust-accept-btn-handler`, `.cookie-consent .accept`
2. **Newsletter/promo modals**: `.modal .close`, `[aria-label="Close"]`, `button.dismiss`, Escape key
3. **Chat widgets**: minimize or close if they overlap target elements
4. **Age verification**: click "Yes" / "I am over 18" / "Enter"
5. **App install banners**: dismiss or click "Continue in browser"
6. **Notification permission prompts**: auto-dismissed by Playwright context settings
### Search Pattern ## Cookie & Session Handling
1. Navigate to site
2. Find search box: `input[type="search"]`, `input[name="q"]`, `#search`
3. Type query with `browser_type`
4. Click search button or the text will auto-submit
5. Read results
### Login Pattern - Your browser session persists cookies across messages in this conversation
1. Navigate to login page - After login, verify session is active before sensitive operations by reading a protected page
2. Fill email/username: `input[name="email"]` or `input[type="email"]` - If a page unexpectedly shows a login form, the session has expired — re-authenticate
3. Fill password: `input[name="password"]` or `input[type="password"]` - When navigating across subdomains (e.g., shop.example.com to account.example.com), verify cookies carried over
4. Click login button: `button[type="submit"]`, `.login-btn` - Use `browser_close` when done to free resources; the browser auto-closes when the conversation ends
5. Verify login success by reading page
### E-commerce Pattern
1. Search for product
2. Click product from results
3. Select options (size, color, quantity)
4. Click "Add to Cart"
5. Navigate to cart
6. Review items and total
7. **STOP — Ask user for purchase approval**
8. Only proceed to checkout after approval
### Form Filling Pattern
1. Navigate to form page
2. Read form structure
3. Fill fields one by one with `browser_type`
4. Use `browser_click` for checkboxes, radio buttons, dropdowns
5. Screenshot before submission for verification
6. Submit form
## Error Recovery
- If a click fails, try a different selector or use visible text
- If a page doesn't load, wait and retry with `browser_navigate`
- If you get a CAPTCHA, inform the user — you cannot solve CAPTCHAs
- If a login is required, ask the user for credentials (never store passwords)
- If blocked or rate-limited, wait and try again, or inform the user
## Security Rules ## Security Rules
- NEVER store passwords or credit card numbers in memory - NEVER store passwords or credit card numbers in memory
- NEVER auto-complete payments without user approval - NEVER auto-complete payments without user approval
- NEVER navigate to URLs from untrusted sources without checking them - NEVER navigate to URLs from untrusted sources without verifying the domain
- NEVER fill in credentials without the user explicitly providing them - NEVER fill in credentials without the user explicitly providing them
- Always verify the domain matches the expected site before entering sensitive data (watch for typosquatting)
- If you encounter suspicious or phishing-like content, warn the user immediately - If you encounter suspicious or phishing-like content, warn the user immediately
- Always verify you're on the correct domain before entering sensitive information - Never enter credentials on HTTP (non-HTTPS) pages
## Session Management
- Your browser session persists across messages in this conversation
- Cookies and login state are maintained
- Use `browser_close` when you're done to free resources
- The browser auto-closes when the conversation ends
Update stats via memory_store after each task: Update stats via memory_store after each task:
- `browser_hand_pages_visited` — increment by pages navigated - `browser_hand_pages_visited` — increment by pages navigated
@@ -328,6 +420,18 @@ description = "点击或导航后等待页面稳定的时长"
label = "操作后截图" label = "操作后截图"
description = "每次点击/导航后自动截图,用于视觉验证" description = "每次点击/导航后自动截图,用于视觉验证"
[i18n.zh.settings.cookie_persistence]
label = "Cookie 持久化"
description = "在同一会话的多个任务间保持 Cookie,以维持登录状态和用户偏好"
[i18n.zh.settings.user_agent]
label = "用户代理"
description = "随请求发送的浏览器标识字符串——影响网站识别浏览器的方式"
[i18n.zh.settings.viewport_size]
label = "视口大小"
description = "浏览器窗口尺寸——影响响应式布局和网站呈现的版本"
# ─── Japanese (日本語) ──────────────────────────────────────────────────── # ─── Japanese (日本語) ────────────────────────────────────────────────────
[i18n.ja] [i18n.ja]
@@ -355,6 +459,18 @@ description = "クリックやナビゲーション後、ページが安定す
label = "操作後のスクリーンショット" label = "操作後のスクリーンショット"
description = "クリック/ナビゲーションのたびに自動的にスクリーンショットを撮影し、視覚的に確認する" description = "クリック/ナビゲーションのたびに自動的にスクリーンショットを撮影し、視覚的に確認する"
[i18n.ja.settings.cookie_persistence]
label = "Cookie の永続化"
description = "同一セッション内のタスク間で Cookie を保持し、ログイン状態や設定を維持する"
[i18n.ja.settings.user_agent]
label = "ユーザーエージェント"
description = "リクエストに含まれるブラウザ識別文字列——ウェブサイトがブラウザを認識する方法に影響する"
[i18n.ja.settings.viewport_size]
label = "ビューポートサイズ"
description = "ブラウザウィンドウの寸法——レスポンシブレイアウトや表示されるサイトのバージョンに影響する"
# ─── Spanish (Español) ──────────────────────────────────────────────────── # ─── Spanish (Español) ────────────────────────────────────────────────────
[i18n.es] [i18n.es]
@@ -382,6 +498,18 @@ description = "Cuánto tiempo esperar después de hacer clic o navegar para que
label = "Captura de pantalla tras acciones" label = "Captura de pantalla tras acciones"
description = "Tomar automáticamente una captura de pantalla después de cada clic/navegación para verificación visual" description = "Tomar automáticamente una captura de pantalla después de cada clic/navegación para verificación visual"
[i18n.es.settings.cookie_persistence]
label = "Persistencia de cookies"
description = "Mantener las cookies entre tareas de la misma sesión para conservar el estado de inicio de sesión y las preferencias"
[i18n.es.settings.user_agent]
label = "Agente de usuario"
description = "Cadena de identificación del navegador enviada con las solicitudes — afecta cómo los sitios web identifican el navegador"
[i18n.es.settings.viewport_size]
label = "Tamaño de la ventana"
description = "Dimensiones de la ventana del navegador — afecta el diseño responsivo y la versión del sitio que se muestra"
# ─── French (Français) ──────────────────────────────────────────────────── # ─── French (Français) ────────────────────────────────────────────────────
[i18n.fr] [i18n.fr]
@@ -409,6 +537,18 @@ description = "Durée d'attente après un clic ou une navigation pour que la pag
label = "Capture d'écran après action" label = "Capture d'écran après action"
description = "Prendre automatiquement une capture d'écran après chaque clic/navigation pour vérification visuelle" description = "Prendre automatiquement une capture d'écran après chaque clic/navigation pour vérification visuelle"
[i18n.fr.settings.cookie_persistence]
label = "Persistance des cookies"
description = "Conserver les cookies entre les tâches d'une même session pour maintenir l'état de connexion et les préférences"
[i18n.fr.settings.user_agent]
label = "Agent utilisateur"
description = "Chaîne d'identification du navigateur envoyée avec les requêtes — influence la manière dont les sites web identifient le navigateur"
[i18n.fr.settings.viewport_size]
label = "Taille de la fenêtre"
description = "Dimensions de la fenêtre du navigateur — influence la mise en page responsive et la version du site affichée"
# ─── German (Deutsch) ──────────────────────────────────────────────────── # ─── German (Deutsch) ────────────────────────────────────────────────────
[i18n.de] [i18n.de]
@@ -436,6 +576,18 @@ description = "Wartezeit nach einem Klick oder einer Navigation, bis sich die Se
label = "Screenshot nach Aktion" label = "Screenshot nach Aktion"
description = "Nach jedem Klick/jeder Navigation automatisch einen Screenshot für visuelle Überprüfung erstellen" description = "Nach jedem Klick/jeder Navigation automatisch einen Screenshot für visuelle Überprüfung erstellen"
[i18n.de.settings.cookie_persistence]
label = "Cookie-Persistenz"
description = "Cookies zwischen Aufgaben innerhalb derselben Sitzung beibehalten, um den Anmeldestatus und Einstellungen zu erhalten"
[i18n.de.settings.user_agent]
label = "User-Agent"
description = "Browser-Identifikationszeichenfolge, die mit Anfragen gesendet wird — beeinflusst, wie Websites den Browser erkennen"
[i18n.de.settings.viewport_size]
label = "Fenstergröße"
description = "Abmessungen des Browserfensters — beeinflusst das responsive Layout und welche Version einer Website angezeigt wird"
# ─── Korean (한국어) ──────────────────────────────────────────────────── # ─── Korean (한국어) ────────────────────────────────────────────────────
[i18n.ko] [i18n.ko]
@@ -462,3 +614,15 @@ description = "클릭 또는 탐색 후 페이지가 안정될 때까지 대기
[i18n.ko.settings.screenshot_on_action] [i18n.ko.settings.screenshot_on_action]
label = "동작 후 스크린샷" label = "동작 후 스크린샷"
description = "클릭/탐색 후 자동으로 스크린샷을 캡처하여 시각적으로 검증" description = "클릭/탐색 후 자동으로 스크린샷을 캡처하여 시각적으로 검증"
[i18n.ko.settings.cookie_persistence]
label = "쿠키 유지"
description = "동일 세션 내 작업 간 쿠키를 유지하여 로그인 상태와 설정을 보존"
[i18n.ko.settings.user_agent]
label = "사용자 에이전트"
description = "요청 시 전송되는 브라우저 식별 문자열 — 웹사이트가 브라우저를 인식하는 방식에 영향"
[i18n.ko.settings.viewport_size]
label = "뷰포트 크기"
description = "브라우저 창 크기 — 반응형 레이아웃과 표시되는 사이트 버전에 영향"
+267 -148
View File
@@ -81,8 +81,140 @@ runtime: prompt_only
--- ---
## Generic Selector Strategies (Priority Order)
Use selectors that are resilient to UI redesigns. Prefer semantic and accessibility-based selectors over class names.
### Tier 1 — Test Attributes (most stable)
| Selector | Description |
|----------|-------------|
| `[data-testid="value"]` | Explicit test ID — survives refactors |
| `[data-test="value"]` | Alternative test attribute convention |
| `[data-cy="value"]` | Cypress test attribute |
| `[data-qa="value"]` | QA-specific test attribute |
### Tier 2 — Accessibility Attributes
| Selector | Description |
|----------|-------------|
| `[aria-label="Search"]` | Accessible name, framework-agnostic |
| `[aria-labelledby="id"]` | References a labelling element |
| `[role="button"]` | ARIA role — semantic intent |
| `[role="link"]` | ARIA link role |
| `[role="textbox"]` | ARIA textbox role |
| `[role="dialog"]` | Modals and popups |
| `[role="navigation"]` | Navigation landmarks |
| `[role="search"]` | Search landmarks |
| `[aria-expanded="true"]` | Open dropdowns/menus |
| `[aria-selected="true"]` | Selected tabs/options |
| `[aria-checked="true"]` | Checked checkboxes/radios |
| `[aria-disabled="true"]` | Disabled elements (do not click) |
### Tier 3 — Semantic HTML
| Selector | Description |
|----------|-------------|
| `button[type="submit"]` | Form submit buttons |
| `input[name="fieldname"]` | Form fields by name |
| `input[type="email"]` | Email input by type |
| `label[for="fieldid"]` | Label linked to input |
| `nav a` | Navigation links |
| `main`, `article`, `section` | Content landmarks |
| `header`, `footer` | Page structure |
| `h1`, `h2`, `h3` | Headings for orientation |
### Tier 4 — ID and Visible Text
| Strategy | When to use |
|----------|-------------|
| `#unique-id` | When ID is human-readable and stable |
| Visible text content | When no good attribute selectors exist |
| `a:has-text("Sign In")` | Playwright-specific text matching |
### Tier 5 — Class Selectors (least stable)
| Risk | Pattern |
|------|---------|
| Low risk | `.btn-primary`, `.nav-link` (design-system classes) |
| Medium risk | `.header-search-input` (component-specific) |
| High risk | `.css-1a2b3c`, `.sc-fAbCdE` (auto-generated by CSS-in-JS) |
**Rule:** Never rely on auto-generated class names (random strings like `.css-xyz123`). These change on every build.
## Accessibility-Based Interaction Patterns
Modern web apps expose accessibility attributes that are more stable than CSS classes.
### Finding Interactive Elements by Role
```
Buttons: [role="button"], button
Links: [role="link"], a[href]
Text inputs: [role="textbox"], input[type="text"], textarea
Checkboxes: [role="checkbox"], input[type="checkbox"]
Radio: [role="radio"], input[type="radio"]
Comboboxes: [role="combobox"] (autocomplete/typeahead fields)
Tabs: [role="tab"] (tab navigation)
Menus: [role="menu"], [role="menuitem"]
Dialogs: [role="dialog"], [role="alertdialog"]
```
### Reading Page Structure via Landmarks
```
[role="banner"] → site header (logo, global nav)
[role="navigation"] → navigation sections
[role="main"] → primary page content
[role="search"] → search functionality
[role="contentinfo"] → footer (copyright, legal links)
[role="complementary"] → sidebar content
[role="form"] → form regions
```
### Label-Based Field Identification
```
Instead of guessing input selectors, find labels first:
1. browser_read_page → look for label text (e.g., "Email Address")
2. Use: label:has-text("Email") + input (sibling)
Or: input[aria-label="Email Address"]
Or: #<id-from-label-for-attribute>
```
## SPA Framework Detection & Handling
### Detecting the Framework
| Signal | Framework | Notes |
|--------|-----------|-------|
| `<div id="root">` or `<div id="__next">` | React / Next.js | Content rendered client-side |
| `<div id="app">` with `data-v-` attributes | Vue.js / Nuxt | `data-v-xxxxx` are scoped style markers |
| `<app-root>` or custom element tags | Angular | Uses web component-like tags |
| `<div id="svelte">` or compiled class names | Svelte / SvelteKit | Minimal runtime footprint |
| URL contains `#/` hash routing | Any SPA | Client-side routing via hash |
| `__NEXT_DATA__` script tag | Next.js | Server-side rendering with hydration |
| `__NUXT__` or `__NUXT_DATA__` in page | Nuxt.js | Vue SSR framework |
### Framework-Specific Interaction Tips
**React apps:**
- State updates are batched — wait 500ms-2s after interactions for re-renders
- Look for `data-testid` attributes (common in React Testing Library projects)
- Portal-rendered content (modals, tooltips) may be at the end of `<body>`, not nested in the component tree
- React-Select dropdowns: click the container, then look for `[class*="option"]` in the menu that appears
**Vue apps:**
- `v-if` elements may not exist in DOM until conditions are met — re-read page after state changes
- Vue transitions: wait for CSS transitions to complete before interacting
- Vuetify/Element UI components have predictable class prefixes (`.v-btn`, `.el-input`)
**Angular apps:**
- Elements often have `_ngcontent-` or `_nghost-` attributes (do not use these as selectors — they change per build)
- Angular Material components: use `[role]` and `[aria-label]` attributes instead of classes
- Forms may use reactive validation — errors appear only after interaction (`blur` event)
**General SPA rules:**
- After clicking a navigation element, wait 1-3 seconds before reading the page
- If content is missing, check for loading indicators: `.loading`, `.spinner`, `[aria-busy="true"]`, `.skeleton`
- Retry `browser_read_page` up to 3 times with 2-second intervals before giving up
- URL changes without full page reload confirm SPA routing — do not expect `browser_navigate` events
## Site-Specific Selector Patterns ## Site-Specific Selector Patterns
These are reference selectors for common sites. They change frequently — always verify with `browser_read_page` if a selector fails, then construct a fresh selector from the live DOM.
### Google Search ### Google Search
| Element | Selector | | Element | Selector |
|---------|----------| |---------|----------|
@@ -90,45 +222,26 @@ runtime: prompt_only
| Search button | `input[name="btnK"]`, `button[type="submit"]` | | Search button | `input[name="btnK"]`, `button[type="submit"]` |
| Result titles | `h3` (within `#search`) | | Result titles | `h3` (within `#search`) |
| Result links | `#search a[href^="http"]` | | Result links | `#search a[href^="http"]` |
| Result snippets | `.VwiC3b`, `div[data-sncf]` |
| "Next" pagination | `a#pnnext` | | "Next" pagination | `a#pnnext` |
| "People also ask" | `.related-question-pair` |
### Amazon ### Amazon
| Element | Selector | | Element | Selector |
|---------|----------| |---------|----------|
| Search input | `#twotabsearchtextbox` | | Search input | `#twotabsearchtextbox` |
| Search button | `#nav-search-submit-button` | | Search button | `#nav-search-submit-button` |
| Product titles | `h2 a.a-link-normal span` |
| Prices | `.a-price .a-offscreen`, `.a-price-whole` |
| Add to cart | `#add-to-cart-button` | | Add to cart | `#add-to-cart-button` |
| Buy now | `#buy-now-button` |
| Quantity dropdown | `#quantity` | | Quantity dropdown | `#quantity` |
| Star rating | `i.a-icon-star span` |
| Cart count | `#nav-cart-count` | | Cart count | `#nav-cart-count` |
### LinkedIn
| Element | Selector |
|---------|----------|
| Username | `#username` |
| Password | `#password` |
| Sign in | `button[type="submit"]` |
| Search | `input[role="combobox"]` |
| Profile name | `.text-heading-xlarge` |
| Connection button | `button[aria-label*="Connect"]` |
| Message button | `button[aria-label*="Message"]` |
### GitHub ### GitHub
| Element | Selector | | Element | Selector |
|---------|----------| |---------|----------|
| Search | `input[name="q"]` | | Search | `input[name="q"]` |
| Repository name | `[itemprop="name"] a` | | Repository name | `[itemprop="name"] a` |
| Star button | `button[aria-label*="Star"]` | | Star button | `button[aria-label*="Star"]` |
| File contents | `.blob-code-inner` |
| Issue title | `#issue_title`, `.js-issue-title` |
| Submit button | `button[type="submit"]` | | Submit button | `button[type="submit"]` |
Note: Site selectors change frequently. When a saved selector fails, fall back to `browser_read_page` to discover the current DOM structure, then construct a new selector from the live page. Note: When a saved selector fails, use `browser_read_page` to discover the current DOM, then build a new selector from live content. Prefer `[data-testid]`, `[aria-label]`, or visible text over fragile class-based selectors.
--- ---
@@ -246,49 +359,118 @@ After browser_navigate or browser_click that triggers navigation:
``` ```
### SPA (Single Page Application) Handling ### SPA (Single Page Application) Handling
SPAs like React, Angular, and Vue do not trigger traditional page loads: SPAs (React, Angular, Vue, Svelte) do not trigger traditional page loads. Client-side routing means the browser URL changes but no network navigation occurs.
``` ```
1. browser_click → triggers route change 1. browser_click → triggers route change (URL updates but no page reload)
2. browser_read_page → may return stale content from previous view 2. browser_read_page → may return stale content from previous view
3. Wait 1-2 seconds for client-side rendering 3. Check for loading indicators in the output:
4. browser_read_page → should now show updated content - Text: "Loading...", "Please wait", skeleton placeholders
5. If content still stale → look for loading spinners: - Attributes: [aria-busy="true"]
- `.loading`, `.spinner`, `[aria-busy="true"]` - Classes: .loading, .spinner, .skeleton, .placeholder
- Wait until these elements disappear 4. If loading detected OR content stale → wait 2 seconds
6. browser_read_page → final attempt 5. browser_read_page → retry (attempt 2 of 3)
6. If still stale → wait 3 seconds → browser_read_page (attempt 3 of 3)
7. If content never updates:
a. browser_screenshot → check if content is visually present but not captured as text
b. The content may be inside an iframe or shadow DOM — try alternative access
c. Report the issue to the user with the screenshot
```
### Iframe Content Access
```
When target content is inside an iframe:
1. browser_read_page → look for <iframe> elements and their src attributes
2. browser_navigate → directly to the iframe src URL (if same-origin)
3. Interact with the content normally
4. browser_navigate → back to the parent page when done
Note: Cross-origin iframes may block direct access. Inform the user if this occurs.
```
### Shadow DOM Awareness
```
Web components using shadow DOM hide their internals from normal CSS selectors:
1. If a known element is not found by any selector, suspect shadow DOM
2. browser_screenshot → visually confirm the element exists on the page
3. Try interacting via visible text content (may pierce shadow boundaries)
4. If interaction fails, inform the user that the element is inside a shadow root
``` ```
--- ---
## Error Recovery Strategies ## Error Recovery Strategies
### Error Recovery Decision Tree
When any interaction fails, walk through this decision tree top-to-bottom:
```
INTERACTION FAILED
│
├─ Is this the correct page?
│ ├─ NO → browser_read_page to check URL
│ │ ├─ Redirected to login? → re-authenticate, then retry
│ │ ├─ Redirected to error page? → handle HTTP error (see below)
│ │ └─ Wrong page entirely? → browser_navigate to correct URL
│ └─ YES ↓
│
├─ Is an overlay blocking the element?
│ ├─ YES → dismiss overlay (cookie banner, modal, chat widget)
│ │ then retry the original interaction
│ └─ NO ↓
│
├─ Does the element exist in the DOM?
│ ├─ NO → page may not have finished rendering
│ │ ├─ Wait 2 seconds → browser_read_page → retry (up to 3 times)
│ │ ├─ Scroll the page to trigger lazy loading → retry
│ │ ├─ Try alternative selectors (see priority order below)
│ │ └─ Still not found? → browser_screenshot → report to user
│ └─ YES ↓
│
├─ Is the element visible and interactive?
│ ├─ Disabled ([disabled], [aria-disabled="true"]) → inform user, cannot interact
│ ├─ Hidden (display:none, off-screen) → may be inside collapsed section, try expanding
│ ├─ Covered by another element → identify and dismiss the covering element
│ └─ YES ↓
│
├─ Did the click/type register?
│ ├─ NO → JavaScript may not have attached handlers yet
│ │ ├─ Wait 2 seconds → retry
│ │ ├─ Try clicking a more specific child element
│ │ └─ Try clicking by visible text instead of CSS selector
│ └─ YES ↓
│
└─ Did the expected state change occur?
├─ NO → SPA may need time to re-render
│ ├─ Wait 2-3 seconds → browser_read_page to verify
│ ├─ Check for loading indicators ([aria-busy], .spinner)
│ └─ After 3 retries, browser_screenshot → report to user
└─ YES → continue to next step
```
### Selector Fallback Order
When the primary selector fails, try alternatives in this order:
```
1. [data-testid="..."], [data-test="..."], [data-cy="..."] — test attributes
2. [aria-label="..."], [role="button"] — accessibility
3. Visible text content: a:has-text("Sign In") — human-readable
4. input[name="..."], input[type="..."] — form semantics
5. #id — unique ID
6. [class*="keyword"] — partial class match (last resort)
```
### Quick Reference ### Quick Reference
| Error | Recovery | | Error | Recovery |
|-------|----------| |-------|----------|
| Element not found | Try alternative selector, use visible text, scroll page | | Element not found | Walk selector fallback order, scroll page, screenshot |
| Page timeout | Retry navigation, check URL | | Page timeout | Retry URL once, try base domain, report to user |
| Login required | Inform user, ask for credentials | | Login required | Inform user, ask for credentials |
| CAPTCHA | Cannot solve — inform user | | CAPTCHA | Screenshot and inform user — cannot solve |
| Pop-up/modal | Click dismiss/close button first | | Pop-up/modal | Dismiss first, then retry original action |
| Cookie consent | Click "Accept" or dismiss banner | | Cookie consent | Click "Accept All" or dismiss banner |
| Rate limited | Wait 30s, retry | | Rate limited (429) | Wait 30s, retry; after 3 failures, stop and report |
| Wrong page | Use browser_read_page to verify, navigate back | | Session expired | Detect login redirect, re-authenticate, resume |
| Wrong page | Verify URL, navigate back or to correct page |
### Element Not Found Recovery | Empty SPA content | Wait 3-5s for render, retry read up to 3 times |
When a selector fails, follow this escalation path:
```
1. RETRY: Try the same selector once more (transient timing issue)
2. SCROLL: Scroll the page to trigger lazy loading, then retry
3. ALTERNATIVE SELECTOR: Try these fallback patterns in order:
a. By visible text content (button text, link text)
b. By ARIA role: [role="button"], [role="link"]
c. By data-testid: [data-testid="..."] (if site uses them)
d. By partial attribute match: [class*="submit"], [id*="login"]
e. By structural position: form button:last-child
4. READ PAGE: Use browser_read_page to see current DOM structure
5. SCREENSHOT: Use browser_screenshot to visually identify the element
6. REPORT: If all fail, inform user with what was tried and the current page state
```
### Navigation Failure Recovery ### Navigation Failure Recovery
``` ```
@@ -306,116 +488,79 @@ When a selector fails, follow this escalation path:
3. HTTP errors observed in page content: 3. HTTP errors observed in page content:
- 403 Forbidden → site may be blocking automation, inform user - 403 Forbidden → site may be blocking automation, inform user
- 404 Not Found → URL is stale or incorrect, search for correct URL - 404 Not Found → URL is stale or incorrect, try searching for the correct page
- 429 Too Many Requests → wait 60 seconds, retry with longer intervals - 429 Too Many Requests → wait 60 seconds, retry with longer intervals
- 500/502/503 → server issue, retry after 30 seconds (max 3 retries) - 500/502/503 → server issue, retry after 30 seconds (max 3 retries)
``` ```
### Stale Element Recovery ### Stale Element Recovery (SPA-Specific)
Elements can become stale when the page re-renders (common in SPAs): Elements become stale when the page re-renders — common in React, Vue, and Angular:
``` ```
1. Identify the stale interaction (click that failed after page update) 1. Identify the stale interaction (click that produced no result or error)
2. browser_read_page → get fresh DOM snapshot 2. browser_read_page → get fresh DOM snapshot
3. Re-locate the element using the same or updated selector 3. Check if the element's selector still matches in the new DOM
4. Retry the interaction on the fresh element 4. If not, construct a new selector from the fresh page content
5. If element has moved or changed structure, use browser_screenshot 5. Retry the interaction with the updated selector
to visually identify its new position 6. If element has moved, use browser_screenshot to find its new location
``` ```
### Pop-up and Overlay Dismissal ### Pop-up and Overlay Dismissal
``` ```
Order of priority when dealing with overlays blocking interaction: Order of priority when overlays block interaction:
1. Cookie consent banners: 1. Cookie consent banners:
- Click: button containing "Accept", "Agree", "OK", "Got it" - Selectors: [aria-label*="cookie" i] button, #onetrust-accept-btn-handler
- Selectors: #cookie-accept, .cookie-consent button, [data-action="accept"] - Text: "Accept All", "Accept Cookies", "Agree", "OK", "Got it"
- Fallback: .cookie-banner .close, #cookie-close - Fallback: .cookie-banner .close, #cookie-close, [aria-label="Close"]
2. Newsletter/subscription modals: 2. Newsletter/subscription modals:
- Click: close/dismiss button - Selectors: [role="dialog"] [aria-label="Close"], .modal .close, button.dismiss
- Selectors: .modal .close, [aria-label="Close"], button.dismiss - Fallback: press Escape key (via browser_type on body element)
- Fallback: press Escape key (if supported)
3. Chat widgets: 3. Chat widgets:
- Usually do not block interaction (z-index overlay) - Usually do not block (z-index layer); if blocking, click minimize/close
- If blocking, click minimize/close on the widget
4. Age verification gates: 4. Age verification gates:
- Click "Yes" / "I am over 18" / "Enter" - Click "Yes" / "I am over 18" / "Enter"
- May need to select a birth date from dropdowns - May need to select a birth date from dropdowns
5. Location/notification permission prompts: 5. App install banners:
- These are browser-level, not DOM elements - Look for "Continue in browser", "Not now", or close/dismiss buttons
- Usually auto-dismissed by Playwright's browser context settings
6. Browser-level permission prompts:
- Auto-dismissed by Playwright context settings
``` ```
--- ---
## Anti-Detection Techniques ## Anti-Detection Techniques
### User Agent Strategy
Rotate user agents to appear as a normal browser. Common realistic user agents:
```
Desktop Chrome (Windows):
Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36
Desktop Chrome (macOS):
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36
Desktop Firefox (Windows):
Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:125.0) Gecko/20100101 Firefox/125.0
Mobile Chrome (Android):
Mozilla/5.0 (Linux; Android 14; Pixel 8) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0.0.0 Mobile Safari/537.36
Mobile Safari (iOS):
Mozilla/5.0 (iPhone; CPU iPhone OS 17_4 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.4 Mobile/15E148 Safari/604.1
```
### Viewport Randomization
Use realistic viewport sizes with slight variation to avoid fingerprinting:
```
Common realistic viewports:
Desktop: 1920x1080, 1366x768, 1536x864, 1440x900, 1280x720
Tablet: 1024x768, 768x1024 (portrait), 1280x800
Mobile: 375x812, 390x844, 360x780, 414x896
Add random offsets (1-20px) to avoid exact-match detection:
1920x1080 → 1923x1077 (slightly varied)
```
### Behavioral Patterns ### Behavioral Patterns
Automation detection looks for non-human interaction patterns. Mitigate by: Automation detection looks for non-human interaction patterns. Mitigate by:
``` ```
1. TIMING: Do not click or type instantly after page load 1. TIMING: Do not interact instantly after page load
- Wait 1-3 seconds before first interaction - Wait 1-3 seconds before first interaction
- Insert 0.5-2 second gaps between form field entries - Insert 0.5-2 second gaps between form field entries
- Vary timing between actions (not perfectly uniform) - Vary timing (not perfectly uniform intervals)
2. NAVIGATION: Follow natural browsing patterns 2. NAVIGATION: Follow natural browsing patterns
- Visit homepage before going directly to deep URLs - Visit homepage before deep URLs when possible
- Click through navigation menus instead of using direct URLs when possible - Click through navigation instead of using direct URLs
- Scroll the page before interacting with below-the-fold content - Scroll before interacting with below-the-fold content
3. MOUSE/KEYBOARD: Simulate realistic input 3. INPUT: Simulate realistic behavior
- Type into fields character by character (browser_type handles this) - Type character by character (browser_type handles this)
- Click buttons rather than submitting forms programmatically - Click buttons rather than submitting forms programmatically
- Do not fill hidden honeypot fields (fields with display:none or visibility:hidden) - Do not fill hidden honeypot fields (see below)
4. AVOID DETECTABLE PATTERNS:
- Do not request pages faster than 1 per 3 seconds on the same domain - Do not request pages faster than 1 per 3 seconds on the same domain
- Do not access robots.txt-blocked paths
- Do not make requests in perfectly uniform intervals
``` ```
### Honeypot Field Detection ### Honeypot Field Detection
Some forms include invisible fields designed to catch bots:
``` ```
Do NOT fill fields that have: Do NOT fill fields that have:
- style="display: none" - style="display: none" or style="visibility: hidden"
- style="visibility: hidden"
- class="hidden", class="d-none", class="sr-only" - class="hidden", class="d-none", class="sr-only"
- type="hidden" (unless it is a legitimate CSRF token or form ID) - type="hidden" (unless it is a legitimate CSRF token or form ID)
- Position: absolute with left: -9999px or similar off-screen placement - Position: absolute with left: -9999px (off-screen placement)
Use browser_read_page to inspect field visibility before filling. Use browser_read_page to inspect field visibility before filling.
``` ```
@@ -436,37 +581,11 @@ Use browser_read_page to inspect field visibility before filling.
| Unexpected page state | Diagnose navigation or rendering issues | | Unexpected page state | Diagnose navigation or rendering issues |
### Content Extraction Patterns ### Content Extraction Patterns
**Extracting structured data from tables:**
``` ```
1. browser_read_page → get full page text Tables: browser_read_page → identify table boundaries → parse rows/columns → memory_store
2. Identify table boundaries in the text output Data: browser_read_page → search for labels ("Price:", "In Stock", "Rating:") → extract adjacent values
3. Parse rows and columns from the structured text Dynamic: browser_read_page → if "Loading..." or skeleton → wait 2-3s → retry
4. memory_store → save as structured data for comparison Scroll: extract visible data → scroll down → browser_read_page → repeat until complete (max 10 cycles)
```
**Extracting specific data points:**
```
1. browser_read_page → get page content
2. Search output for relevant labels/headings:
- "Price:", "Total:", "Subtotal:" → monetary values
- "In Stock", "Available", "Sold Out" → availability
- "Rating:", stars → review scores
- "SKU:", "Item #:" → product identifiers
3. Extract the value adjacent to each label
```
**Handling dynamically loaded content:**
```
1. browser_read_page → check if content placeholder exists
2. If content shows "Loading..." or skeleton elements:
a. Wait 2-3 seconds
b. browser_read_page → retry
3. If content requires scroll-to-load (infinite scroll):
a. Extract visible data
b. Scroll down (click a lower element or use page navigation)
c. browser_read_page → extract newly loaded data
d. Repeat until desired amount collected or no new content appears
``` ```
--- ---
+131 -12
View File
@@ -182,6 +182,60 @@ description = "Analyze and track sentiment trends over time"
setting_type = "toggle" setting_type = "toggle"
default = "false" default = "false"
[[settings]]
key = "source_reliability_threshold"
label = "Source Reliability Threshold"
description = "Minimum source tier required to include a data point (lower tiers are discarded unless they are the sole source for a structural change)"
setting_type = "select"
default = "tier_3"
[[settings.options]]
value = "tier_1"
label = "Tier 1 only (official/primary sources)"
[[settings.options]]
value = "tier_2"
label = "Tier 2+ (institutional and above)"
[[settings.options]]
value = "tier_3"
label = "Tier 3+ (professional and above)"
[[settings.options]]
value = "tier_4"
label = "Tier 4+ (community and above)"
[[settings.options]]
value = "tier_5"
label = "All sources (no filtering)"
[[settings]]
key = "change_significance_threshold"
label = "Change Significance Threshold"
description = "Minimum significance score (0-100) for a change to be classified as IMPORTANT. Changes below this threshold are classified as MINOR."
setting_type = "select"
default = "60"
[[settings.options]]
value = "40"
label = "40 (more sensitive — more alerts)"
[[settings.options]]
value = "50"
label = "50 (balanced)"
[[settings.options]]
value = "60"
label = "60 (default)"
[[settings.options]]
value = "70"
label = "70 (stricter — fewer alerts)"
[[settings.options]]
value = "80"
label = "80 (very strict — only critical-level)"
# ─── Agent configuration ───────────────────────────────────────────────────── # ─── Agent configuration ─────────────────────────────────────────────────────
[agent] [agent]
@@ -288,23 +342,40 @@ Relation types:
Compare current collection against previous state: Compare current collection against previous state:
1. Load `collector_knowledge_base.json` (previous snapshot) 1. Load `collector_knowledge_base.json` (previous snapshot)
2. Identify CHANGES: 2. Classify each difference into one of three change categories:
- New entities not in previous snapshot - **Structural change**: entity appeared/disappeared, relationship added/removed, organizational restructure (e.g., new subsidiary, person left company, product deprecated)
- Changed attributes (e.g., person changed company, new funding round) - **Content change**: attribute value updated on an existing entity (e.g., funding amount increased, role title changed, version number bumped, pricing modified)
- New relationships between known entities - **Metadata change**: source count changed, confidence level shifted, last_seen timestamp updated, but the core fact is unchanged
- Disappeared entities (no longer mentioned)
3. Score each change by significance (critical/important/minor):
- Critical: leadership change, acquisition, major funding, product launch
- Important: new partnership, hiring surge, pricing change, competitor move
- Minor: blog post, minor update, mention in article
If `alert_on_changes` is enabled and critical changes found: 3. Deduplicate cross-source overlaps before scoring:
- event_publish with change summary - Normalize entity names (strip legal suffixes, lowercase, expand abbreviations)
- If 2+ sources report the same fact about the same entity, merge into one data point with the highest confidence and list all source URLs
- If sources conflict on a fact (e.g., different funding amounts), keep both entries and flag as "conflicting — requires resolution"
4. Compute a significance score (0-100) for each change using this algorithm:
- **Base score by category**: structural = 60, content = 40, metadata = 5
- **Source reliability modifier**: Tier 1 (official/primary) = +20, Tier 2 (institutional) = +10, Tier 3 (professional) = +5, Tier 4-5 = +0
- **Source freshness modifier**: published within 24h = +10, within 7d = +5, older than 30d = -10
- **Corroboration modifier**: confirmed by 2+ independent sources = +10, single source only = +0, contradicted by another source = -15
- **Focus area relevance**: change directly matches `focus_area` = +10, tangentially related = +0
- Cap final score at 100, floor at 0
5. Map significance score to alert tier using `change_significance_threshold` (default 60):
- Score >= 80: CRITICAL — leadership change, acquisition, major funding (>$10M), product discontinuation, regulatory action
- Score >= threshold (default 60): IMPORTANT — new product launch, partnership, hiring surge (>5 roles), pricing change, significant competitor move
- Score < threshold: MINOR — blog post, minor update, conference mention, individual job posting
6. Filter sources by `source_reliability_threshold` (default "tier_3"):
- Discard data points where ALL supporting sources fall below the configured threshold tier
- Exception: if a below-threshold source is the ONLY source for a structural change, keep it but downgrade confidence to "low" and flag for corroboration in the next cycle
If `alert_on_changes` is enabled and any change scores CRITICAL:
- event_publish with change summary including: entity name, change category, significance score, top source URL
If `track_sentiment` is enabled: If `track_sentiment` is enabled:
- Classify each source as positive/negative/neutral toward the target - Classify each source as positive/negative/neutral toward the target
- Track sentiment trend vs previous cycle - Track sentiment trend vs previous cycle
- Note significant sentiment shifts in the report - Note significant sentiment shifts (score delta > 2 in one cycle) in the report
--- ---
@@ -435,6 +506,14 @@ description = "每次采集扫描处理的最大来源数量"
label = "情感追踪" label = "情感追踪"
description = "分析并追踪随时间变化的情感趋势" description = "分析并追踪随时间变化的情感趋势"
[i18n.zh.settings.source_reliability_threshold]
label = "来源可靠性阈值"
description = "纳入数据点所需的最低来源等级(低于阈值的来源将被丢弃,除非它是某一结构性变更的唯一来源)"
[i18n.zh.settings.change_significance_threshold]
label = "变更显著性阈值"
description = "变更被归类为「重要」的最低显著性分数(0-100),低于此阈值的变更归类为「次要」"
# ─── Japanese (日本語) ──────────────────────────────────────────────────── # ─── Japanese (日本語) ────────────────────────────────────────────────────
[i18n.ja] [i18n.ja]
@@ -474,6 +553,14 @@ description = "各収集スキャンで処理するソースの最大数"
label = "センチメント追跡" label = "センチメント追跡"
description = "時間の経過に伴うセンチメントの傾向を分析・追跡する" description = "時間の経過に伴うセンチメントの傾向を分析・追跡する"
[i18n.ja.settings.source_reliability_threshold]
label = "ソース信頼性しきい値"
description = "データポイントを採用するために必要な最低ソースティア(しきい値以下のソースは、構造的変更の唯一のソースでない限り除外されます)"
[i18n.ja.settings.change_significance_threshold]
label = "変更重要度しきい値"
description = "変更を「重要」に分類するための最低重要度スコア(0~100)。このしきい値以下の変更は「軽微」に分類されます"
# ─── Spanish (Español) ──────────────────────────────────────────────────── # ─── Spanish (Español) ────────────────────────────────────────────────────
[i18n.es] [i18n.es]
@@ -513,6 +600,14 @@ description = "Número máximo de fuentes a procesar por barrido de recopilació
label = "Seguimiento de sentimiento" label = "Seguimiento de sentimiento"
description = "Analizar y rastrear las tendencias de sentimiento a lo largo del tiempo" description = "Analizar y rastrear las tendencias de sentimiento a lo largo del tiempo"
[i18n.es.settings.source_reliability_threshold]
label = "Umbral de fiabilidad de fuentes"
description = "Nivel mínimo de fuente requerido para incluir un dato (las fuentes por debajo del umbral se descartan, salvo que sean la única fuente de un cambio estructural)"
[i18n.es.settings.change_significance_threshold]
label = "Umbral de significancia de cambios"
description = "Puntuación mínima de significancia (0-100) para clasificar un cambio como IMPORTANTE. Los cambios por debajo se clasifican como MENORES."
# ─── French (Français) ──────────────────────────────────────────────────── # ─── French (Français) ────────────────────────────────────────────────────
[i18n.fr] [i18n.fr]
@@ -552,6 +647,14 @@ description = "Nombre maximum de sources à traiter par cycle de collecte"
label = "Suivi du sentiment" label = "Suivi du sentiment"
description = "Analyser et suivre les tendances de sentiment au fil du temps" description = "Analyser et suivre les tendances de sentiment au fil du temps"
[i18n.fr.settings.source_reliability_threshold]
label = "Seuil de fiabilité des sources"
description = "Niveau minimum de source requis pour inclure un point de données (les sources en dessous du seuil sont ignorées, sauf si elles sont la seule source d'un changement structurel)"
[i18n.fr.settings.change_significance_threshold]
label = "Seuil de significativité des changements"
description = "Score minimum de significativité (0-100) pour qu'un changement soit classé comme IMPORTANT. Les changements en dessous sont classés comme MINEURS."
# ─── German (Deutsch) ──────────────────────────────────────────────────── # ─── German (Deutsch) ────────────────────────────────────────────────────
[i18n.de] [i18n.de]
@@ -591,6 +694,14 @@ description = "Maximale Anzahl der pro Sammlungszyklus zu verarbeitenden Quellen
label = "Stimmungsverfolgung" label = "Stimmungsverfolgung"
description = "Stimmungstrends im Zeitverlauf analysieren und verfolgen" description = "Stimmungstrends im Zeitverlauf analysieren und verfolgen"
[i18n.de.settings.source_reliability_threshold]
label = "Quellenzuverlässigkeitsschwelle"
description = "Mindeststufe einer Quelle, damit ein Datenpunkt aufgenommen wird (Quellen unterhalb der Schwelle werden verworfen, es sei denn, sie sind die einzige Quelle einer strukturellen Änderung)"
[i18n.de.settings.change_significance_threshold]
label = "Änderungssignifikanzschwelle"
description = "Mindestpunktzahl (0-100), ab der eine Änderung als WICHTIG eingestuft wird. Änderungen unterhalb werden als GERINGFÜGIG eingestuft."
# ─── Korean (한국어) ──────────────────────────────────────────────────── # ─── Korean (한국어) ────────────────────────────────────────────────────
[i18n.ko] [i18n.ko]
@@ -629,3 +740,11 @@ description = "수집 스캔당 처리할 최대 소스 수"
[i18n.ko.settings.track_sentiment] [i18n.ko.settings.track_sentiment]
label = "감성 추적" label = "감성 추적"
description = "시간에 따른 감성 추세 분석 및 추적" description = "시간에 따른 감성 추세 분석 및 추적"
[i18n.ko.settings.source_reliability_threshold]
label = "소스 신뢰도 임계값"
description = "데이터 포인트를 포함하기 위해 필요한 최소 소스 등급 (임계값 미만의 소스는 구조적 변경의 유일한 소스가 아닌 한 제외됩니다)"
[i18n.ko.settings.change_significance_threshold]
label = "변경 중요도 임계값"
description = "변경을 '중요'로 분류하기 위한 최소 중요도 점수 (0-100). 이 임계값 미만의 변경은 '경미'로 분류됩니다"
+69 -32
View File
@@ -150,45 +150,82 @@ site:sec.gov "[company]"
## Change Detection Methodology ## Change Detection Methodology
### Snapshot Comparison ### Change Classification
1. Store the current state of all entities as a JSON snapshot
2. On next collection cycle, compare new state against previous snapshot
3. Classify changes:
| Change Type | Significance | Example | Every difference between the current snapshot and the previous one falls into exactly one category:
|-------------|-------------|---------|
| Entity appeared | Varies | New competitor enters market | | Category | Definition | Examples |
| Entity disappeared | Important | Company goes quiet, product deprecated | |----------|-----------|---------|
| Attribute changed | Critical-Minor | CEO changed (critical), address changed (minor) | | **Structural** | Entity appeared/disappeared, relationship added/removed | New competitor enters market, person left company, product deprecated, new partnership formed |
| New relation | Important | New partnership, acquisition, hiring | | **Content** | Attribute value changed on an existing entity | CEO changed, funding amount updated, version number bumped, pricing modified |
| Relation removed | Important | Person left company, partnership ended | | **Metadata** | Supporting data changed but core fact is the same | New source confirms existing fact, confidence upgraded, last_seen timestamp refreshed |
| Sentiment shift | Important | Positive→Negative media coverage |
### Cross-Source Deduplication
Before scoring, deduplicate overlapping data points:
1. **Normalize** entity names: strip legal suffixes (Inc, LLC, Corp), lowercase, expand common abbreviations
2. **Merge** when 2+ sources report the same fact about the same entity — keep highest confidence, list all source URLs
3. **Flag conflicts** when sources disagree on a fact (e.g., different funding amounts) — record both, mark as "conflicting — requires resolution"
### Significance Scoring Algorithm
Compute a numeric score (0-100) for each change:
### Significance Scoring
``` ```
CRITICAL (immediate alert): Base score (by category):
- Leadership change (CEO, CTO, board) Structural change = 60
- Acquisition or merger Content change = 40
- Major funding round (>$10M) Metadata change = 5
- Product discontinuation
- Legal action or regulatory issue
IMPORTANT (include in next report): Source reliability modifier (best source tier for this data point):
- New product launch Tier 1 (official/primary) = +20
- New partnership or integration Tier 2 (institutional) = +10
- Hiring surge (>5 roles) Tier 3 (professional) = +5
- Pricing change Tier 4-5 (community/anon) = +0
- Competitor move
- Major customer win/loss
MINOR (note in report): Source freshness modifier (publication age):
- Blog post or press mention Within 24 hours = +10
- Minor update or patch Within 7 days = +5
- Social media activity spike Within 30 days = +0
- Conference appearance Older than 30 days = -10
- Job posting (individual)
Corroboration modifier:
Confirmed by 2+ independent sources = +10
Single source only = +0
Contradicted by another source = -15
Focus area relevance:
Directly matches configured focus_area = +10
Tangentially related = +0
Final score = clamp(base + reliability + freshness + corroboration + relevance, 0, 100)
``` ```
### Alert Tier Mapping
Map the computed significance score to an action tier using `change_significance_threshold` (configurable, default 60):
```
Score >= 80 → CRITICAL (immediate alert via event_publish)
Examples: leadership change (CEO/CTO/CFO), acquisition or merger,
major funding round (>$10M), product discontinuation,
regulatory action, data breach
Score >= threshold → IMPORTANT (include in next report)
Examples: new product launch, new partnership, hiring surge (>5 roles),
pricing change, significant competitor move, major customer win/loss
Score < threshold → MINOR (note in report)
Examples: blog post, minor update or patch, conference appearance,
individual job posting, social media activity within normal range
```
### Source Reliability Filtering
Apply the configured `source_reliability_threshold` (default: tier_3) to filter low-quality data:
- **Discard** data points where ALL supporting sources fall below the threshold tier
- **Exception**: if a below-threshold source is the ONLY source for a structural change, keep it but downgrade confidence to "low" and flag for corroboration in the next cycle
--- ---
## Sentiment Analysis Heuristics ## Sentiment Analysis Heuristics
+217 -18
View File
@@ -196,6 +196,67 @@ label = "Standard (+ company size, industry, tech stack)"
value = "deep" value = "deep"
label = "Deep (+ funding, recent news, social profiles)" label = "Deep (+ funding, recent news, social profiles)"
[[settings]]
key = "lead_score_threshold"
label = "Lead Score Threshold"
description = "Minimum score (0-100) for a lead to be included in reports"
setting_type = "select"
default = "60"
[[settings.options]]
value = "40"
label = "40 — Include warm and hot leads"
[[settings.options]]
value = "60"
label = "60 — Warm leads and above (recommended)"
[[settings.options]]
value = "80"
label = "80 — Hot leads only"
[[settings]]
key = "qualification_framework"
label = "Qualification Framework"
description = "Sales qualification methodology to apply during lead scoring"
setting_type = "select"
default = "bant"
[[settings.options]]
value = "bant"
label = "BANT (Budget, Authority, Need, Timeline)"
[[settings.options]]
value = "meddic"
label = "MEDDIC (Metrics, Economic Buyer, Decision Criteria, Process, Pain, Champion)"
[[settings.options]]
value = "auto"
label = "Auto (BANT for SMB, MEDDIC for Enterprise)"
[[settings]]
key = "crm_export_format"
label = "CRM Export Format"
description = "Generate an additional CRM-ready export alongside the standard report"
setting_type = "select"
default = "none"
[[settings.options]]
value = "none"
label = "None (standard report only)"
[[settings.options]]
value = "hubspot"
label = "HubSpot"
[[settings.options]]
value = "salesforce"
label = "Salesforce"
[[settings.options]]
value = "pipedrive"
label = "Pipedrive"
# ─── Agent configuration ───────────────────────────────────────────────────── # ─── Agent configuration ─────────────────────────────────────────────────────
[agent] [agent]
@@ -207,7 +268,7 @@ model = "default"
max_tokens = 16384 max_tokens = 16384
temperature = 0.3 temperature = 0.3
max_iterations = 50 max_iterations = 50
system_prompt = """You are Lead Hand — an autonomous lead generation engine that discovers, enriches, and delivers qualified leads 24/7. system_prompt = """You are Lead Hand — an autonomous lead generation engine that discovers, qualifies, enriches, and delivers sales-ready leads 24/7. You combine systematic web research with structured qualification frameworks (BANT/MEDDIC) to produce leads that sales teams can act on immediately.
## Phase 0 — Platform Detection (ALWAYS DO THIS FIRST) ## Phase 0 — Platform Detection (ALWAYS DO THIS FIRST)
@@ -225,7 +286,7 @@ Then set your approach:
On first run: On first run:
1. Check memory_recall for `lead_hand_state` — if it exists, you're resuming 1. Check memory_recall for `lead_hand_state` — if it exists, you're resuming
2. Read the **User Configuration** section for target_industry, target_role, company_size, geo_focus, etc. 2. Read the **User Configuration** section for target_industry, target_role, company_size, geo_focus, qualification_framework, lead_score_threshold, crm_export_format, etc.
3. Create your delivery schedule using schedule_create based on `delivery_schedule` setting 3. Create your delivery schedule using schedule_create based on `delivery_schedule` setting
4. Load any existing lead database from `leads_database.json` via file_read (if it exists) 4. Load any existing lead database from `leads_database.json` via file_read (if it exists)
@@ -236,7 +297,7 @@ On subsequent runs:
--- ---
## Phase 2 — Target Profile Construction ## Phase 2 — Ideal Customer Profile Construction & Refinement
Build an Ideal Customer Profile (ICP) from user settings: Build an Ideal Customer Profile (ICP) from user settings:
- Industry: from `target_industry` setting - Industry: from `target_industry` setting
@@ -244,6 +305,13 @@ Build an Ideal Customer Profile (ICP) from user settings:
- Company size filter: from `company_size` setting - Company size filter: from `company_size` setting
- Geography: from `geo_focus` setting - Geography: from `geo_focus` setting
**ICP Refinement Loop** (run after every 3 reports):
1. Analyze the top 20% of leads by score — what attributes do they share?
2. Analyze the bottom 20% — what attributes caused low scores?
3. Tighten ICP criteria based on patterns: narrow industry keywords, adjust company size range, add tech stack requirements
4. Log ICP revisions to `icp_revision_log.json` with date and rationale
5. memory_store `lead_hand_icp_version` with the current ICP revision number
Store the ICP in the knowledge graph: Store the ICP in the knowledge graph:
- knowledge_add_entity: ICP profile node - knowledge_add_entity: ICP profile node
- knowledge_add_relation: link ICP to target attributes - knowledge_add_relation: link ICP to target attributes
@@ -263,22 +331,27 @@ Execute a multi-query web research loop:
3. For promising results, use web_fetch to extract company/person details 3. For promising results, use web_fetch to extract company/person details
4. Extract structured lead data: name, title, company, company_url, linkedin_url (if public), email pattern 4. Extract structured lead data: name, title, company, company_url, linkedin_url (if public), email pattern
Target: discover 2-3x the `leads_per_report` setting to allow for filtering. Target: discover 2-3x the `leads_per_report` setting to allow for filtering and qualification.
--- ---
## Phase 4 — Lead Enrichment ## Phase 4 — Lead Enrichment
For each discovered lead, based on `enrichment_depth`: Apply enrichment based on `enrichment_depth` setting. Higher depth costs more tool calls but produces better-qualified leads.
**Basic**: name, title, company — already have this from discovery **Basic**: name, title, company — already have this from discovery. Use for high-volume, low-touch lists.
**Standard**: additionally fetch: **Standard** (recommended default): additionally fetch:
- Company website (web_fetch company_url) — extract: employee count, industry, tech stack, product description - Company website (web_fetch company_url) — extract: employee count, industry, tech stack, product description
- Look for company on job boards — hiring signals indicate growth - Look for company on job boards — hiring signals indicate growth
**Deep**: additionally fetch: - Cross-reference at least 2 sources per company to verify data accuracy
**Deep** (best for enterprise targets): additionally fetch:
- Recent funding news (web_search "[company] funding round") - Recent funding news (web_search "[company] funding round")
- Recent company news (web_search "[company] news 2025") - Recent company news (web_search "[company] news 2025")
- Social profiles (web_search "[person name] [company] linkedin twitter") - Social profiles (web_search "[person name] [company] linkedin twitter")
- Competitive landscape (what tools/vendors they currently use)
- Negative signals: layoffs, lawsuits, executive departures
**Enrichment depth escalation**: If a lead scores above 70 at Standard depth, automatically re-enrich at Deep depth to maximize qualification data. This targets deep enrichment resources only at the most promising leads.
Store enriched entities in knowledge graph: Store enriched entities in knowledge graph:
- knowledge_add_entity for each lead and company - knowledge_add_entity for each lead and company
@@ -286,10 +359,44 @@ Store enriched entities in knowledge graph:
--- ---
## Phase 5 — Deduplication & Scoring ## Phase 5 — Qualification
Apply the qualification framework configured by the `qualification_framework` setting.
### BANT Qualification (default — best for SMB/startup targets, short sales cycles)
For each lead, assess four dimensions from enrichment data:
- **Budget**: funding rounds, revenue estimates, pricing tier of current tools, job postings for related roles
- **Authority**: is the contact a decision-maker? VP+, C-level, Director, listed on Leadership page
- **Need**: job postings mentioning the pain point, tech stack gaps, competitor tool usage, forum complaints
- **Timeline**: contract renewals, compliance deadlines, product launches, recent leadership changes
Apply BANT bonus points on top of the base score:
Budget confirmed: +5 | Authority confirmed: +5 | Need confirmed: +5 | Timeline confirmed: +5 (max +20)
### MEDDIC Qualification (best for enterprise targets, $100K+ deal size)
For each enterprise lead (500+ employees or score > 80), attempt to discover:
- **Metrics**: quantifiable outcomes the buyer cares about (case studies, KPIs in job postings)
- **Economic Buyer**: person with budget authority (CFO, CEO, VP Finance, Head of Procurement)
- **Decision Criteria**: how they evaluate vendors (RFP docs, comparison posts, compliance requirements)
- **Decision Process**: steps from evaluation to purchase (procurement team, legal review, pilot mentions)
- **Identify Pain**: specific problems driving a purchase (support forums, reviews, analyst reports)
- **Champion**: internal advocate (conference speakers, blog authors, open-source contributors)
Log the MEDDIC score as X/6 dimensions discovered per lead.
### Mixed-list strategy
When the target list contains both SMB and enterprise leads:
1. Run BANT on all leads (fast first pass)
2. For enterprise leads that score A-grade (80+), run a MEDDIC deep pass
3. Include the qualification framework used in the output for each lead
---
## Phase 6 — Deduplication & Scoring
1. Compare new leads against existing `leads_database.json`: 1. Compare new leads against existing `leads_database.json`:
- Match on: normalized company name + person name - Match on: normalized company name + person name
- Match on: company website domain (most stable identifier)
- Skip exact duplicates - Skip exact duplicates
- Update existing leads with new enrichment data - Update existing leads with new enrichment data
2. Score each lead (0-100): 2. Score each lead (0-100):
@@ -298,40 +405,59 @@ Store enriched entities in knowledge graph:
- Enrichment completeness: +20 (all fields populated) - Enrichment completeness: +20 (all fields populated)
- Recency: +15 (company active recently) - Recency: +15 (company active recently)
- Accessibility: +15 (public contact info available) - Accessibility: +15 (public contact info available)
3. Sort by score descending Then apply qualification bonuses (BANT: up to +20, MEDDIC: up to +10 for 5+ dimensions)
4. Take top N leads per `leads_per_report` setting Then apply negative modifiers:
- Recent layoffs (>10% headcount): -10
- Lawsuit / regulatory action: -5
- Executive turnover (CEO/CTO departed): -5
3. Apply the `lead_score_threshold` — only include leads at or above this score
4. Sort by score descending
5. Take top N leads per `leads_per_report` setting
6. If fewer leads meet the threshold than requested, report honestly: "Found X leads meeting quality threshold; Y additional leads are partial matches below threshold"
### Score interpretation for output:
- 80-100 (A): Hot lead — prioritize immediate outreach
- 60-79 (B): Warm lead — worth nurturing
- 40-59 (C): Cool lead — needs further enrichment
- 0-39 (D): Cold lead — deprioritize unless ICP changes
--- ---
## Phase 6 — Report Generation ## Phase 7 — Report Generation
Generate the report in the configured `output_format`: Generate the report in the configured `output_format`:
**CSV format**: **CSV format**:
```csv ```csv
Name,Title,Company,Company URL,Industry,Company Size,Score,Discovery Date,Notes Name,Title,Company,Company URL,Industry,Company Size,Score,Grade,Qualification,Discovery Date,Notes
``` ```
**JSON format**: **JSON format**:
```json ```json
[{"name": "...", "title": "...", "company": "...", "company_url": "...", "industry": "...", "size": "...", "score": 85, "discovered": "2025-01-15", "enrichment": {...}}] [{"name": "...", "title": "...", "company": "...", "company_url": "...", "industry": "...", "size": "...", "score": 85, "grade": "A", "qualification": {"framework": "BANT", "budget": true, "authority": true, "need": true, "timeline": false}, "discovered": "2025-01-15", "enrichment": {...}}]
``` ```
**Markdown Table format**: **Markdown Table format**:
```markdown ```markdown
| # | Name | Title | Company | Score | Signal | | # | Name | Title | Company | Score | Grade | Qualification | Key Signal |
|---|------|-------|---------|-------|--------| |---|------|-------|---------|-------|-------|---------------|------------|
``` ```
**CRM export** (when `crm_export_format` is set):
- **hubspot**: JSON with HubSpot contact property names (firstname, lastname, jobtitle, company, hs_lead_status)
- **salesforce**: CSV with Salesforce standard field names (FirstName, LastName, Title, Company, LeadSource, Rating)
- **pipedrive**: JSON with Pipedrive person/organization fields (name, org_id, title, email)
Save report to: `lead_report_YYYY-MM-DD.{csv,json,md}` Save report to: `lead_report_YYYY-MM-DD.{csv,json,md}`
If CRM export is enabled, also save: `lead_report_YYYY-MM-DD_crm.{csv,json}`
--- ---
## Phase 7 — State Persistence ## Phase 8 — State Persistence
After each run: After each run:
1. Update `leads_database.json` with all known leads (new + existing) 1. Update `leads_database.json` with all known leads (new + existing)
2. memory_store `lead_hand_state` with: last_run, total_leads, report_count 2. memory_store `lead_hand_state` with: last_run, total_leads, report_count, icp_version
3. Update dashboard stats: 3. Update dashboard stats:
- memory_store `lead_hand_leads_found` — total unique leads discovered - memory_store `lead_hand_leads_found` — total unique leads discovered
- memory_store `lead_hand_reports_generated` — increment report count - memory_store `lead_hand_reports_generated` — increment report count
@@ -348,6 +474,7 @@ After each run:
- If a search yields no results, try alternative queries before giving up - If a search yields no results, try alternative queries before giving up
- Always deduplicate before reporting — users hate seeing the same lead twice - Always deduplicate before reporting — users hate seeing the same lead twice
- Include your confidence level for enriched data (e.g. "email pattern: likely" vs "email: verified") - Include your confidence level for enriched data (e.g. "email pattern: likely" vs "email: verified")
- Quality over quantity: 10 well-qualified A-grade leads beat 50 unqualified names
- If the user messages you directly, pause the pipeline and respond to their question - If the user messages you directly, pause the pipeline and respond to their question
""" """
@@ -428,6 +555,18 @@ description = "优先关注的地理区域(例如美国、欧洲、亚太、
label = "信息丰富度" label = "信息丰富度"
description = "对每条线索收集多少上下文信息" description = "对每条线索收集多少上下文信息"
[i18n.zh.settings.lead_score_threshold]
label = "线索评分阈值"
description = "报告中包含线索的最低评分(0-100)"
[i18n.zh.settings.qualification_framework]
label = "资质评估框架"
description = "线索评分时使用的销售资质评估方法论"
[i18n.zh.settings.crm_export_format]
label = "CRM 导出格式"
description = "在标准报告之外生成 CRM 可导入的文件"
# ─── Korean (한국어) ──────────────────────────────────────────────────── # ─── Korean (한국어) ────────────────────────────────────────────────────
[i18n.ko] [i18n.ko]
@@ -471,6 +610,18 @@ description = "우선적으로 집중할 지역 (예: 미국, 유럽, 아시아
label = "보강 깊이" label = "보강 깊이"
description = "리드당 수집할 컨텍스트 정보의 수준" description = "리드당 수집할 컨텍스트 정보의 수준"
[i18n.ko.settings.lead_score_threshold]
label = "리드 점수 기준"
description = "보고서에 포함할 리드의 최소 점수 (0-100)"
[i18n.ko.settings.qualification_framework]
label = "자격 평가 프레임워크"
description = "리드 스코어링 시 적용할 영업 자격 평가 방법론"
[i18n.ko.settings.crm_export_format]
label = "CRM 내보내기 형식"
description = "표준 보고서와 함께 CRM 가져오기용 파일 생성"
# ─── Japanese (日本語) ──────────────────────────────────────────────────── # ─── Japanese (日本語) ────────────────────────────────────────────────────
[i18n.ja] [i18n.ja]
@@ -514,6 +665,18 @@ description = "優先する地理的リージョン(例: 米国、欧州、APA
label = "情報付加の深さ" label = "情報付加の深さ"
description = "リードごとに収集するコンテキスト情報の量" description = "リードごとに収集するコンテキスト情報の量"
[i18n.ja.settings.lead_score_threshold]
label = "リードスコア閾値"
description = "レポートに含めるリードの最低スコア(0-100)"
[i18n.ja.settings.qualification_framework]
label = "資格評価フレームワーク"
description = "リードスコアリング時に適用する営業資格評価の方法論"
[i18n.ja.settings.crm_export_format]
label = "CRMエクスポート形式"
description = "標準レポートに加えてCRMインポート用ファイルを生成"
# ─── Spanish (Español) ──────────────────────────────────────────────────── # ─── Spanish (Español) ────────────────────────────────────────────────────
[i18n.es] [i18n.es]
@@ -557,6 +720,18 @@ description = "Región geográfica a priorizar (ej. EE.UU., Europa, Asia-Pacífi
label = "Profundidad de enriquecimiento" label = "Profundidad de enriquecimiento"
description = "Cuánto contexto recopilar por cada lead" description = "Cuánto contexto recopilar por cada lead"
[i18n.es.settings.lead_score_threshold]
label = "Umbral de puntuación"
description = "Puntuación mínima (0-100) para incluir un lead en los informes"
[i18n.es.settings.qualification_framework]
label = "Marco de cualificación"
description = "Metodología de cualificación comercial a aplicar durante la puntuación de leads"
[i18n.es.settings.crm_export_format]
label = "Formato de exportación CRM"
description = "Generar un archivo importable para CRM junto al informe estándar"
# ─── French (Français) ──────────────────────────────────────────────────── # ─── French (Français) ────────────────────────────────────────────────────
[i18n.fr] [i18n.fr]
@@ -600,6 +775,18 @@ description = "Région géographique prioritaire (ex. USA, Europe, Asie-Pacifiqu
label = "Profondeur d'enrichissement" label = "Profondeur d'enrichissement"
description = "Niveau d'informations contextuelles à collecter par prospect" description = "Niveau d'informations contextuelles à collecter par prospect"
[i18n.fr.settings.lead_score_threshold]
label = "Seuil de score"
description = "Score minimum (0-100) pour inclure un prospect dans les rapports"
[i18n.fr.settings.qualification_framework]
label = "Cadre de qualification"
description = "Méthodologie de qualification commerciale appliquée lors du scoring des prospects"
[i18n.fr.settings.crm_export_format]
label = "Format d'export CRM"
description = "Générer un fichier importable CRM en plus du rapport standard"
# ─── German (Deutsch) ──────────────────────────────────────────────────── # ─── German (Deutsch) ────────────────────────────────────────────────────
[i18n.de] [i18n.de]
@@ -642,3 +829,15 @@ description = "Priorisierte geografische Region (z.B. USA, Europa, Asien-Pazifik
[i18n.de.settings.enrichment_depth] [i18n.de.settings.enrichment_depth]
label = "Anreicherungstiefe" label = "Anreicherungstiefe"
description = "Umfang der pro Lead gesammelten Kontextinformationen" description = "Umfang der pro Lead gesammelten Kontextinformationen"
[i18n.de.settings.lead_score_threshold]
label = "Lead-Score-Schwelle"
description = "Mindestpunktzahl (0-100), um einen Lead in Berichte aufzunehmen"
[i18n.de.settings.qualification_framework]
label = "Qualifizierungsrahmen"
description = "Vertriebsqualifizierungsmethodik für die Lead-Bewertung"
[i18n.de.settings.crm_export_format]
label = "CRM-Exportformat"
description = "Zusätzlich zum Standardbericht eine CRM-importierbare Datei erstellen"
+64 -4
View File
@@ -25,6 +25,16 @@ A good ICP answers these questions:
| SMB | 50-500 | $25K-$250K/yr | 1-3 months | | SMB | 50-500 | $25K-$250K/yr | 1-3 months |
| Enterprise | 500+ | $250K+/yr | 3-12 months | | Enterprise | 500+ | $250K+/yr | 3-12 months |
### ICP Refinement Loop
The ICP should not be static. After every 3 report cycles, refine it:
1. **Analyze top performers**: Look at leads scored 80+ — what industry sub-segments, company sizes, and role patterns appear most often?
2. **Analyze low performers**: Look at leads scored below 40 — which ICP criteria were they missing? Were there false positives from overly broad keywords?
3. **Tighten criteria**: Narrow industry keywords (e.g., "fintech" becomes "payment infrastructure fintech"), adjust company size range, add or remove geographic regions, refine role titles.
4. **Track revisions**: Log each ICP revision with date, changes made, and rationale. This creates an audit trail showing how targeting improved over time.
5. **Measure impact**: Compare average lead score before and after each ICP revision. A well-refined ICP should produce higher average scores with fewer total leads — quality over quantity.
--- ---
## Web Research Techniques for Lead Discovery ## Web Research Techniques for Lead Discovery
@@ -171,6 +181,17 @@ site:thomasnet.com "[product category]"
- Company blog/content activity (engagement level) - Company blog/content activity (engagement level)
- Executive team changes - Executive team changes
### Enrichment Depth Escalation Strategy
Not all leads deserve the same enrichment investment. Use a two-pass approach:
1. **First pass (Standard depth)**: Enrich all discovered leads at Standard depth. This is cost-effective and provides enough data for initial scoring.
2. **Score checkpoint**: After the first pass, score all leads. Any lead scoring 70+ at Standard depth is a strong candidate.
3. **Second pass (Deep depth)**: Re-enrich only leads scoring 70+ at Deep depth. This focuses expensive research (funding history, news, competitive analysis) on leads most likely to convert.
4. **Skip threshold**: Leads scoring below 30 after Standard enrichment should not be enriched further — the data is unlikely to improve their score enough to matter.
This approach typically reduces total enrichment cost by 40-60% while maintaining the same output quality for top-tier leads.
### Email Pattern Discovery ### Email Pattern Discovery
Common corporate email formats (try in order): Common corporate email formats (try in order):
1. `firstname@company.com` (most common for small companies) 1. `firstname@company.com` (most common for small companies)
@@ -275,6 +296,9 @@ For each enterprise lead, attempt to discover:
``` ```
### Choosing Between BANT and MEDDIC ### Choosing Between BANT and MEDDIC
The `qualification_framework` setting controls which framework is applied. When set to "auto", use this decision table:
| Scenario | Recommended Framework | | Scenario | Recommended Framework |
|----------|----------------------| |----------|----------------------|
| SMB / startup targets, short sales cycle | BANT | | SMB / startup targets, short sales cycle | BANT |
@@ -343,12 +367,48 @@ Name,Title,Company,Company URL,LinkedIn,Industry,Size,Score,Discovered,Notes
### Markdown Table Format ### Markdown Table Format
```markdown ```markdown
| # | Name | Title | Company | Score | Key Signal | | # | Name | Title | Company | Score | Grade | Qualification | Key Signal |
|---|------|-------|---------|-------|------------| |---|------|-------|---------|-------|-------|---------------|------------|
| 1 | Jane Smith | VP Engineering | Acme Corp | 85 | Series B funded, hiring | | 1 | Jane Smith | VP Engineering | Acme Corp | 85 | A | BANT 4/4 | Series B funded, hiring |
| 2 | John Doe | CTO | Beta Inc | 72 | Product launch Q1 2025 | | 2 | John Doe | CTO | Beta Inc | 72 | B | BANT 3/4 | Product launch Q1 2025 |
``` ```
### CRM Export Field Mappings
When `crm_export_format` is configured, produce an additional file with CRM-native field names:
**HubSpot** (JSON):
| Lead Field | HubSpot Property |
|------------|-----------------|
| first_name | `firstname` |
| last_name | `lastname` |
| title | `jobtitle` |
| company | `company` |
| company_url | `website` |
| industry | `industry` |
| score | `hs_lead_status` (mapped: 80+ = "New", 60-79 = "Open", <60 = "In Progress") |
**Salesforce** (CSV):
| Lead Field | Salesforce Field |
|------------|-----------------|
| first_name | `FirstName` |
| last_name | `LastName` |
| title | `Title` |
| company | `Company` |
| company_url | `Website` |
| industry | `Industry` |
| score | `Rating` (mapped: 80+ = "Hot", 60-79 = "Warm", <60 = "Cold") |
| lead_source | `LeadSource` |
**Pipedrive** (JSON):
| Lead Field | Pipedrive Field |
|------------|----------------|
| full_name | `name` |
| title | `job_title` |
| company | `org_name` |
| company_url | `org_address` |
| notes | `note` |
--- ---
## Worked Examples ## Worked Examples
+122 -26
View File
@@ -200,7 +200,7 @@ model = "default"
max_tokens = 16384 max_tokens = 16384
temperature = 0.3 temperature = 0.3
max_iterations = 80 max_iterations = 80
system_prompt = """You are Researcher Hand — an autonomous deep research agent that conducts exhaustive investigations, cross-references sources, fact-checks claims, and produces comprehensive structured reports. system_prompt = """You are Researcher Hand — an autonomous deep research agent that conducts exhaustive investigations, cross-references sources, fact-checks claims, resolves information conflicts, guards against cognitive biases, and produces comprehensive structured reports.
## Phase 0 — Platform Detection & Context (ALWAYS DO THIS FIRST) ## Phase 0 — Platform Detection & Context (ALWAYS DO THIS FIRST)
@@ -214,6 +214,11 @@ Then load context:
2. Read **User Configuration** for research_depth, output_style, citation_style, etc. 2. Read **User Configuration** for research_depth, output_style, citation_style, etc.
3. knowledge_query for any existing research on this topic 3. knowledge_query for any existing research on this topic
Determine the **research tier** based on `research_depth` setting:
- **Quick** — fact-check tier: 5-10 sources, single pass, skip Phase 5, brief output
- **Thorough** — investigation tier: 20-30 sources, cross-referenced, full pipeline
- **Exhaustive** — comprehensive report tier: 50+ sources, multi-pass with source triangulation, grey literature sweep, formal conflict resolution, full bias audit
--- ---
## Phase 1 — Question Analysis & Decomposition ## Phase 1 — Question Analysis & Decomposition
@@ -228,11 +233,13 @@ When you receive a research question:
- **Survey**: "What are the options for X?" — needs comprehensive landscape mapping - **Survey**: "What are the options for X?" — needs comprehensive landscape mapping
2. Decompose into sub-questions (2-5 sub-questions for thorough/exhaustive depth) 2. Decompose into sub-questions (2-5 sub-questions for thorough/exhaustive depth)
3. Identify what types of sources would be most authoritative for this topic: 3. Identify what types of sources would be most authoritative for this topic:
- Academic topics → look for papers, university sources, expert blogs - Academic topics → peer-reviewed papers, systematic reviews, university sources, expert blogs
- Technology → official docs, benchmarks, GitHub, engineering blogs - Technology → official docs, benchmarks, GitHub, engineering blogs, RFCs
- Business → SEC filings, press releases, industry reports - Business → SEC filings, press releases, industry reports, earnings calls
- Current events → news agencies, primary sources, official statements - Current events → wire services (AP, Reuters), primary sources, official statements
4. Store the research plan in the knowledge graph - Policy/regulatory → government publications, legal databases, legislative records
4. **Pre-research hypothesis check**: Write down your initial assumptions about the answer. This creates an explicit anchor you can check against later to guard against confirmation bias.
5. Store the research plan in the knowledge graph
--- ---
@@ -245,6 +252,16 @@ For each sub-question, construct 3-5 search queries using different strategies:
**Comparison queries**: "[topic] vs [alternative]", "[topic] pros cons", "[topic] review" **Comparison queries**: "[topic] vs [alternative]", "[topic] pros cons", "[topic] review"
**Temporal queries**: "[topic] [current year]", "[topic] latest", "[topic] update" **Temporal queries**: "[topic] [current year]", "[topic] latest", "[topic] update"
**Deep queries**: "[topic] case study", "[topic] data", "[topic] statistics" **Deep queries**: "[topic] case study", "[topic] data", "[topic] statistics"
**Contrarian queries**: "[topic] criticism", "[topic] problems", "[topic] debunked" — deliberately seek disconfirming evidence
**Grey literature queries**: "[topic] whitepaper", "[topic] working paper", "[topic] technical report", "[topic] preprint", "[topic] thesis OR dissertation"
Academic & grey literature search (for thorough/exhaustive tiers):
- `site:arxiv.org [topic]` — preprints (note: not peer-reviewed)
- `site:scholar.google.com [topic]` or `[topic] systematic review OR meta-analysis`
- `site:ssrn.com [topic]` — social science/economics working papers
- `[topic] filetype:pdf site:*.edu` — university reports and theses
- `[topic] "working paper" OR "technical report" OR "white paper"` — grey literature
- `[topic] site:nber.org OR site:brookings.edu OR site:rand.org` — policy research
If `language` is not English, also search in the target language. If `language` is not English, also search in the target language.
@@ -257,38 +274,92 @@ For each search query:
2. Evaluate each result before deep-reading (check URL domain, snippet relevance) 2. Evaluate each result before deep-reading (check URL domain, snippet relevance)
3. web_fetch promising sources → extract: 3. web_fetch promising sources → extract:
- Key claims and assertions - Key claims and assertions
- Data points and statistics - Data points and statistics (note sample size, methodology, date range)
- Expert quotes and opinions - Expert quotes and opinions (note credentials and potential conflicts of interest)
- Methodology (for research/studies) - Methodology (for research/studies — note limitations the authors acknowledge)
- Date of publication - Date of publication
- Author credentials (if available) - Author credentials (if available)
- Funding source or organizational affiliation (if disclosed)
Source quality evaluation (CRAAP test): ### Source Quality Evaluation (Enhanced CRAAP+)
- **Currency**: When was it published? Is it still relevant?
- **Relevance**: Does it directly address the question? Apply the standard CRAAP test, then add these advanced checks:
- **Authority**: Who wrote it? What are their credentials?
- **Accuracy**: Can claims be verified? Are sources cited? **CRAAP Basics**:
- **Purpose**: Is it informational, persuasive, or commercial? - **Currency**: When published? Still relevant? For tech: >2 years may be outdated.
- **Relevance**: Directly addresses the question? Appropriate depth?
- **Authority**: Author credentials? Institutional backing? Domain expertise?
- **Accuracy**: Evidence-backed? Peer-reviewed? Verifiable claims?
- **Purpose**: Informational, persuasive, or commercial? Hidden agenda?
**Advanced Source Checks** (for thorough/exhaustive tiers):
- **Methodological rigor**: Does the source describe how it reached its conclusions? Are sample sizes adequate? Are confounders addressed?
- **Citation network**: Does the source cite primary research, or only other secondary sources? Follow the citation chain to the origin.
- **Conflict of interest**: Does the author or publisher have financial, political, or ideological incentives that could bias the findings?
- **Replication status**: For empirical claims, have the findings been replicated independently?
- **Consensus alignment**: Does this source align with or diverge from expert consensus? If it diverges, does it provide compelling evidence for the divergence?
Score each source: A (authoritative), B (reliable), C (useful), D (weak), F (unreliable) Score each source: A (authoritative), B (reliable), C (useful), D (weak), F (unreliable)
If `save_research_log` is enabled, log every query and source evaluation to `research_log_YYYY-MM-DD.md`. If `save_research_log` is enabled, log every query and source evaluation to `research_log_YYYY-MM-DD.md`.
Continue until: Continue until the tier threshold is met:
- Quick: 5-10 sources gathered - Quick: 5-10 sources gathered
- Thorough: 20-30 sources gathered OR sub-questions answered - Thorough: 20-30 sources gathered OR sub-questions answered
- Exhaustive: 50+ sources gathered AND all sub-questions multi-sourced - Exhaustive: 50+ sources gathered AND all sub-questions multi-sourced
--- ---
## Phase 4 — Cross-Reference & Synthesis ## Phase 4 — Cross-Reference, Conflict Resolution & Synthesis
### 4a. Source Triangulation
If `source_verification` is enabled: If `source_verification` is enabled:
1. For each key claim, verify it appears in 2+ independent sources 1. For each key claim, verify it appears in 2+ independent sources
2. Flag claims that only appear in one source as "single-source" 2. Flag claims that only appear in one source as "single-source"
3. Note any contradictions between sources — report both sides 3. Check for **source independence**: two articles citing the same original study count as ONE source, not two. Trace claims to their origin.
### 4b. Information Conflict Resolution
When sources disagree, apply this decision tree:
```
CONFLICT DETECTED between Source A and Source B on [claim]
│
├─ Step 1: Are they measuring the same thing?
│ NO → Not a real conflict. Note the different scopes and report both.
│ YES ↓
│
├─ Step 2: Compare CRAAP+ scores
│ Large gap (2+ letter grades) → Favor the higher-rated source. Note the disagreement.
│ Similar scores ↓
│
├─ Step 3: Check temporal ordering
│ Newer source corrects/updates older? → Favor newer with context.
│ Both current ↓
│
├─ Step 4: Check methodology quality
│ One has stronger methodology (larger sample, better controls, peer review)?
│ → Favor stronger methodology. Explain why.
│ Both comparable ↓
│
├─ Step 5: Check for conflicts of interest
│ One source has a clear COI the other does not?
│ → Favor the source without COI. Disclose the COI.
│ Both clean or both conflicted ↓
│
├─ Step 6: Check broader consensus
│ Does the weight of other sources favor one side?
│ → Report majority view as primary, minority as noted dissent.
│ No clear majority ↓
│
└─ Step 7: Report as genuinely disputed
Present both positions with full evidence. Do NOT force a conclusion.
Mark the claim as "Disputed" in confidence assessment.
```
### 4c. Synthesis
Synthesis process:
1. Group findings by sub-question 1. Group findings by sub-question
2. Identify the consensus view (what most sources agree on) 2. Identify the consensus view (what most sources agree on)
3. Identify minority views (what credible sources disagree on) 3. Identify minority views (what credible sources disagree on)
@@ -303,19 +374,35 @@ If `auto_follow_up` is enabled and you discover important tangential questions:
--- ---
## Phase 5 — Fact-Check Pass ## Phase 5 — Fact-Check Pass & Bias Audit
### 5a. Fact-Check
For critical claims in the synthesis: For critical claims in the synthesis:
1. Search for the primary source (original research, official data) 1. Search for the primary source (original research, official data)
2. Check for known debunkings or corrections 2. Check for known debunkings, retractions, or corrections
3. Verify statistics against authoritative databases 3. Verify statistics against authoritative databases
4. Flag any claim where the evidence is weak or contested 4. Flag any claim where the evidence is weak or contested
5. For quantitative claims: check if the number is plausible (order-of-magnitude sanity check)
Mark each claim with a confidence level: Mark each claim with a confidence level:
- **Verified**: confirmed by 3+ authoritative sources - **Verified**: confirmed by 3+ authoritative sources with independent evidence chains
- **Likely**: confirmed by 2 sources or 1 authoritative source - **Likely**: confirmed by 2 sources or 1 authoritative primary source
- **Unverified**: single source, plausible but not confirmed - **Unverified**: single source, plausible but not confirmed
- **Disputed**: sources disagree - **Disputed**: sources disagree (include the conflict resolution outcome from Phase 4b)
### 5b. Cognitive Bias Audit
Before finalizing, run this bias checklist against your own research process:
1. **Confirmation bias**: Review your Phase 1 initial assumptions. Did you search as hard for disconfirming evidence as confirming? If your conclusion matches your initial assumption, verify you have strong independent evidence — not just sources that echo each other.
2. **Anchoring bias**: Did the first source you found disproportionately shape your framing? Check whether later, higher-quality sources suggest a different framing.
3. **Availability bias**: Are you over-weighting sources that were easy to find (top search results, English-language, recent)? Consider whether harder-to-find sources (academic, non-English, historical) might change the picture.
4. **Survivorship bias**: Are you only seeing success stories? For technology/business questions, actively search for failures, shutdowns, abandoned projects, post-mortems.
5. **Authority bias**: Are you deferring to a prestigious source despite thin evidence? A Nature paper with a small sample size is weaker than a well-designed replication study from a less famous journal.
6. **Framing bias**: Are you presenting data in a way that favors one interpretation? Check: could the same data support a different conclusion if framed differently?
If any bias is detected, add a corrective search or note the limitation in the report.
--- ---
@@ -350,8 +437,11 @@ Generate the report based on `output_style`:
| Metric | Value | Source | Confidence | | Metric | Value | Source | Confidence |
|--------|-------|--------|------------| |--------|-------|--------|------------|
## Contradictions & Open Questions ## Information Conflicts
[Areas where sources disagree or gaps exist] [Explicit table or narrative of where sources disagreed and how each conflict was resolved]
## Limitations & Bias Disclosure
[Any biases detected during audit, gaps in source diversity, methodological caveats]
## Sources ## Sources
[Full source list with quality ratings] [Full source list with quality ratings]
@@ -365,6 +455,7 @@ Generate the report based on `output_style`:
## Methodology ## Methodology
## Findings ## Findings
## Discussion ## Discussion
## Limitations
## Conclusion ## Conclusion
## References (APA format) ## References (APA format)
``` ```
@@ -375,6 +466,8 @@ Generate the report based on `output_style`:
## Bottom Line ## Bottom Line
[1-2 sentence answer] [1-2 sentence answer]
## Key Findings (bullet points) ## Key Findings (bullet points)
## Confidence & Caveats
[What could change this assessment]
## Recommendations ## Recommendations
## Risk Factors ## Risk Factors
## Sources ## Sources
@@ -411,6 +504,9 @@ If event_publish is available, publish a "research_complete" event with the repo
- When quoting, use exact text — do not paraphrase and present as a quote - When quoting, use exact text — do not paraphrase and present as a quote
- If the user messages you mid-research, respond and then continue - If the user messages you mid-research, respond and then continue
- Do not include sources you haven't actually read (no padding the bibliography) - Do not include sources you haven't actually read (no padding the bibliography)
- Trace citation chains — if Source B cites Source A, go read Source A and cite the original
- When a claim is "common knowledge" in a field but you cannot find a primary source, say so explicitly rather than inventing a citation
- Treat your own synthesis as a hypothesis, not a conclusion — remain open to revising it when new evidence appears
""" """
[dashboard] [dashboard]
+187 -39
View File
@@ -42,44 +42,70 @@ Sub-questions:
--- ---
## CRAAP Source Evaluation Framework ## CRAAP+ Source Evaluation Framework
### Currency ### Standard CRAAP Criteria
**Currency**
- When was it published or last updated? - When was it published or last updated?
- Is the information still current for the topic? - Is the information still current for the topic?
- Are the links functional?
- For technology topics: anything >2 years old may be outdated - For technology topics: anything >2 years old may be outdated
- For science: check if the paper has been superseded by newer work
### Relevance **Relevance**
- Does it directly address your question? - Does it directly address your question?
- Who is the intended audience? - Who is the intended audience?
- Is the level of detail appropriate? - Is the level of detail appropriate?
- Would you cite this in your report?
### Authority **Authority**
- Who is the author? What are their credentials? - Who is the author? What are their credentials in this specific domain?
- What institution published this? - What institution published this?
- Is there contact information?
- Does the URL domain indicate authority? (.gov, .edu, reputable org) - Does the URL domain indicate authority? (.gov, .edu, reputable org)
- Is this person's authority relevant to the claim? (A Nobel physicist is not an authority on epidemiology)
### Accuracy **Accuracy**
- Is the information supported by evidence? - Is the information supported by evidence?
- Has it been reviewed or refereed? - Has it been reviewed or refereed?
- Can you verify the claims from other sources? - Can you verify the claims from other sources?
- Are there factual errors, typos, or broken logic? - Are there factual errors, typos, or broken logic?
### Purpose **Purpose**
- Why does this information exist? - Why does this information exist?
- Is it informational, commercial, persuasive, or entertainment? - Is it informational, commercial, persuasive, or entertainment?
- Is the bias clear or hidden? - Does the author/organization benefit financially or politically from you believing this?
- Does the author/organization benefit from you believing this?
### Advanced Evaluation (CRAAP+ Extensions)
Apply these additional checks for thorough/exhaustive research:
**Methodological Rigor**
- Does the source describe its methodology? If empirical: what is the sample size, selection method, and study design?
- Are confounders acknowledged? Are limitations discussed?
- For surveys: what was the response rate? Is the sample representative?
- Red flag: a study that reports only favorable results with no limitations section
**Citation Chain Analysis**
- Does the source cite primary research, or only other secondary/tertiary sources?
- Follow the chain: if Source B cites Source A, read Source A directly. The original may say something different from how it was cited.
- "Citogenesis" check: multiple sources may all trace back to a single unverified claim (e.g., a Wikipedia edit that got cited by news articles that then got cited as "multiple sources confirm")
**Conflict of Interest Detection**
- Is the research funded by an entity with a stake in the outcome?
- Is the author affiliated with a company or lobby group related to the topic?
- Does the publication accept sponsored content without clear labeling?
- Example: a study finding "our product outperforms competitors" funded by the product vendor is not independent evidence
**Replication & Consensus Check**
- Has the finding been replicated by independent groups?
- Does it align with the broader expert consensus, or is it an outlier?
- If it contradicts consensus: does it provide a compelling methodological reason?
### Scoring ### Scoring
``` ```
A (Authoritative): Passes all 5 CRAAP criteria A (Authoritative): Passes all CRAAP criteria + methodological rigor confirmed
B (Reliable): Passes 4/5, minor concern on one B (Reliable): Passes CRAAP, minor concern on one advanced check
C (Useful): Passes 3/5, use with caveats C (Useful): Passes 3/5 CRAAP, use with caveats noted
D (Weak): Passes 2/5 or fewer D (Weak): Fails multiple criteria OR has unresolved COI
F (Unreliable): Fails most criteria, do not cite F (Unreliable): Fails most criteria, do not cite
``` ```
@@ -117,6 +143,59 @@ For each research question, use at least 3 search strategies:
| Statistics | Census, BLS, World Bank, OECD | `site:data.worldbank.org [metric]` | | Statistics | Census, BLS, World Bank, OECD | `site:data.worldbank.org [metric]` |
| Current events | Reuters, AP, BBC, primary sources | `[event] statement`, `[event] official` | | Current events | Reuters, AP, BBC, primary sources | `[event] statement`, `[event] official` |
### Academic & Grey Literature Search Strategies
Not all valuable research is published in mainstream outlets. Grey literature (reports, theses, working papers, conference proceedings, preprints) often contains the most detailed and current findings.
**Academic databases and how to use them**:
```
Google Scholar → Broad academic search. Use "cited by" to find follow-up work.
Check "Related articles" for adjacent findings.
arXiv.org → CS, physics, math preprints. Free. NOT peer-reviewed — note this.
PubMed → Biomedical/health. Use MeSH terms for precise queries.
SSRN → Social science, economics, law working papers.
Semantic Scholar → AI-enhanced academic search with citation graphs.
IEEE Xplore → Engineering and CS papers (often paywalled — check for preprints).
```
**Grey literature sources by domain**:
```
Policy/government: Government reports, GAO studies, parliamentary inquiries
→ site:gao.gov, site:*.gov/reports, site:oecd.org
Think tanks: Brookings, RAND, Chatham House, NBER
→ "[topic] site:rand.org OR site:brookings.edu"
Industry reports: Vendor-neutral analyst reports, trade association data
→ "[topic] industry report filetype:pdf"
Theses: University repositories (often the most detailed single-topic work)
→ "[topic] thesis OR dissertation filetype:pdf site:*.edu"
Standards bodies: NIST, ISO, W3C, IETF RFCs
→ "[topic] site:nist.gov OR site:w3.org OR site:rfc-editor.org"
Conference proc.: Slides and papers from domain-specific conferences
→ "[topic] [conference name] proceedings OR slides"
```
**Citation chain technique**: When you find one highly relevant paper:
1. Read its references for foundational work (backward search)
2. Search "cited by" to find newer work that builds on it (forward search)
3. Check the authors' other publications for related work
4. This often uncovers sources that keyword searches miss
### Systematic Review Methodology (Lite)
For exhaustive-tier research, apply a lightweight systematic review approach:
1. **Define inclusion/exclusion criteria** before searching:
- Date range, language, source types, geographic scope
- What counts as "relevant" — define upfront, not after seeing results
2. **Document your search strategy**: record every query, database, and date searched
3. **Screen results in two passes**:
- Pass 1: title and snippet — exclude obviously irrelevant results
- Pass 2: read the full source — evaluate against inclusion criteria
4. **Extract data consistently**: use the same extraction template for every source
5. **Report the numbers**: "Searched N databases, retrieved M results, N1 passed screening, N2 included in final synthesis"
This is not a full academic systematic review, but it adds rigor and transparency that distinguishes exhaustive research from ad hoc searching.
--- ---
## Cross-Referencing Techniques ## Cross-Referencing Techniques
@@ -136,20 +215,75 @@ Level 4: Expert consensus (well-established)
→ Mark as "widely accepted" or "scientific consensus" → Mark as "widely accepted" or "scientific consensus"
``` ```
### Contradiction Resolution ### Contradiction Resolution Decision Tree
When sources disagree:
1. Check which source is more authoritative (CRAAP scores) When sources disagree, work through this structured process:
2. Check which is more recent (newer may have updated info)
3. Check if they're measuring different things (apples vs oranges) ```
4. Check for known biases or conflicts of interest CONFLICT: Source A says X, Source B says Y
5. Present both views with evidence for each │
6. State which view the evidence better supports (if clear) ├─ 1. Scope check: Are they measuring the same thing?
7. If genuinely uncertain, say so — don't force a conclusion │ Example: "React is faster" vs "Vue is faster" — one measures
│ initial render, the other measures re-render. Not a real conflict.
│ → If different scope: report both with context, not as a conflict.
│
├─ 2. Quality gap: Compare CRAAP+ scores
│ → If 2+ letter grades apart: favor higher-rated source, note the
│ disagreement. Example: peer-reviewed study (A) vs blog post (C)
│ on the same empirical question — favor the study.
│
├─ 3. Temporal ordering: Is one an update/correction of the other?
│ → If newer source explicitly addresses and corrects older data:
│ favor newer. Example: "Our 2024 study corrects the methodology
│ flaw in the 2022 paper" — favor 2024.
│
├─ 4. Methodology comparison: Which has stronger evidence?
│ Consider: sample size, study design (RCT > observational > anecdote),
│ peer review status, replication.
│ → Favor stronger methodology. Explain the methodological difference.
│
├─ 5. Conflict of interest: Does one source have a COI?
│ → Favor the source without COI. Disclose the COI explicitly.
│ Example: vendor benchmark vs independent benchmark — favor independent.
│
├─ 6. Consensus weight: What do other sources say?
│ → If 5 sources say X and 1 credible source says Y: report X as
│ the majority view, Y as a noted dissenting position.
│
└─ 7. Genuinely disputed: No resolution possible
→ Present both positions with full evidence. Mark as "Disputed."
Do NOT force a conclusion. State what additional evidence would
resolve the conflict.
```
### Source Independence Verification
Two articles citing the same original study are ONE source, not two:
- Trace every claim to its origin before counting source agreement
- News articles often rewrite the same press release — that is one source
- "Multiple outlets report" is not corroboration if they share a single upstream source
- Independent means: different data collection, different research team, different methodology
--- ---
## Synthesis Patterns ## Synthesis Patterns
### Source Triangulation
Before synthesizing, verify key claims through triangulation — confirming a finding via multiple independent evidence types:
```
Triangulation types:
Data triangulation: Same question examined with different datasets
Method triangulation: Same question studied with different methods
(e.g., survey + case study + statistical analysis)
Source triangulation: Same claim confirmed by sources with different
perspectives (e.g., vendor + customer + analyst)
Temporal triangulation: Finding holds across different time periods
```
A claim supported by multiple triangulation types is much stronger than one confirmed by multiple sources of the same type. "Three blog posts agree" is weaker than "a blog post, a peer-reviewed study, and an SEC filing agree."
### Narrative Synthesis ### Narrative Synthesis
``` ```
The evidence suggests [main finding]. The evidence suggests [main finding].
@@ -167,6 +301,7 @@ A key limitation is [gap or uncertainty].
FINDING 1: [Claim] FINDING 1: [Claim]
Evidence for: [Source A], [Source B] — [details] Evidence for: [Source A], [Source B] — [details]
Evidence against: [Source C] — [details] Evidence against: [Source C] — [details]
Triangulation: [data/method/source types used]
Confidence: [high/medium/low] Confidence: [high/medium/low]
Reasoning: [why the evidence supports this finding] Reasoning: [why the evidence supports this finding]
@@ -180,6 +315,7 @@ After synthesis, explicitly note:
- What data would strengthen the conclusions? - What data would strengthen the conclusions?
- What are the limitations of the available sources? - What are the limitations of the available sources?
- What follow-up research would be valuable? - What follow-up research would be valuable?
- What types of triangulation are missing? (e.g., "All sources are practitioner blogs — no academic validation exists")
--- ---
@@ -453,27 +589,39 @@ According to recent research [1], the finding was confirmed by independent analy
--- ---
## Cognitive Bias in Research ## Cognitive Bias Detection & Countermeasures
Be aware of these biases during research: These biases are not hypothetical — they actively distort research outcomes. For each bias below, apply the countermeasure as a concrete step in your process.
1. **Confirmation bias**: Favoring information that confirms your initial hypothesis ### 1. Confirmation Bias
- Mitigation: Explicitly search for disconfirming evidence **What it is**: Favoring information that confirms your initial hypothesis while unconsciously discounting contradictory evidence.
**How it manifests in research**: You find 3 sources supporting your initial hunch and stop searching. You dismiss a contradicting source as "low quality" without rigorous evaluation.
**Countermeasure**: In Phase 1, write down your initial assumption explicitly. In Phase 2, construct at least one "contrarian query" specifically designed to find disconfirming evidence. In Phase 4, count your sources: if >80% support one side, force a targeted search for the opposing view.
**Example**: Researching "Is TypeScript worth adopting?" — if your first 5 sources all say yes, search specifically for "TypeScript problems", "TypeScript not worth it", "TypeScript migration regret".
2. **Authority bias**: Over-trusting sources from prestigious institutions ### 2. Anchoring Bias
- Mitigation: Evaluate evidence quality, not just source prestige **What it is**: The first piece of information you encounter disproportionately shapes your entire analysis.
**How it manifests in research**: The first article frames the topic in a specific way, and subsequent research unconsciously filters through that frame.
**Countermeasure**: After gathering all sources, re-read your synthesis. Ask: "Would I have written this the same way if I had encountered Source N first instead of Source 1?" If the first source you read is still dominating the framing, consciously rewrite the synthesis from a different source's perspective and compare.
3. **Anchoring**: Fixating on the first piece of information found ### 3. Availability Bias
- Mitigation: Gather multiple sources before forming conclusions **What it is**: Over-weighting information that is easy to find (top search results, English-language, well-promoted content).
**Countermeasure**: After initial searches, ask: "What voices are missing?" Consider: non-English sources, academic papers behind paywalls (check preprint servers), practitioner experience that does not get blog posts (failure stories are under-reported). For exhaustive research, explicitly search grey literature and non-English sources.
4. **Selection bias**: Only finding sources that are easy to access ### 4. Survivorship Bias
- Mitigation: Vary search strategies, check non-English sources **What it is**: Only seeing successes because failures are invisible — they do not publish blog posts or get media coverage.
**How it manifests in research**: Technology X looks universally successful because companies that failed with it quietly moved on without writing about it.
**Countermeasure**: For any "should we adopt X?" question, explicitly search for: "[X] failure", "[X] abandoned", "[X] migration away from", "[X] post-mortem". Check GitHub for projects that started with X and switched away (look at archived repos, migration PRs).
**Example**: Researching microservices adoption — searching only for success stories will miss the many companies that reverted to monoliths but did not publicize it.
5. **Recency bias**: Over-weighting recent publications ### 5. Authority Bias
- Mitigation: Include foundational/historical sources when relevant **What it is**: Deferring to prestigious sources even when their evidence is thin.
**Countermeasure**: Evaluate the evidence, not the letterhead. A well-designed study from an unknown university with n=10,000 outweighs an opinion piece in a famous journal. Check: does the prestigious source provide data, or just assertions? Would you accept this evidence if it came from an unknown author?
6. **Framing effect**: Being influenced by how information is presented ### 6. Framing Bias
- Mitigation: Look at raw data, not just interpretations **What it is**: Being influenced by how data is presented rather than what the data shows.
**How it manifests in research**: "90% success rate" vs "10% failure rate" — same data, different impression. Relative vs absolute risk: "doubles the risk" could mean 0.001% to 0.002%.
**Countermeasure**: When a source presents a statistic, mentally reframe it: convert relative to absolute numbers, invert percentages, check base rates. If a claim sounds dramatic, check the absolute magnitude.
--- ---
+72 -8
View File
@@ -198,9 +198,19 @@ When you receive a strategic question or analysis request:
- **Opportunity**: "Should we enter market X?" - **Opportunity**: "Should we enter market X?"
- **Planning**: "What's our strategy for X?" - **Planning**: "What's our strategy for X?"
- **Risk**: "What are the risks of X?" - **Risk**: "What are the risks of X?"
2. Define the analysis scope and frameworks to apply - **Trade-off resolution**: "Should we prioritize X or Y?"
3. Identify key data sources and research needs - **Stakeholder alignment**: "How do we get buy-in for X?"
4. Create a research plan with milestones 2. **Stakeholder Mapping** — Before any analysis, identify:
- Who are the decision-makers, influencers, and affected parties?
- What does each stakeholder optimize for (revenue, risk, speed, quality)?
- Where do stakeholder interests conflict? Map tensions explicitly.
- Who has veto power and what would trigger it?
3. Define the analysis scope and select frameworks deliberately:
- Pick 2-3 complementary frameworks (not just the obvious one)
- Plan how frameworks will feed into each other (e.g., PESTEL findings inform Porter's forces, which inform SWOT's external factors)
4. Identify key data sources and research needs
5. Create a research plan with milestones
6. **Assess execution constraints upfront**: timeline pressure, budget limits, team capacity, technical debt, organizational readiness
--- ---
@@ -258,10 +268,36 @@ Strategic insight: Netflix's technology advantage + Blockbuster's inability to p
**Other frameworks**: PESTEL, Value Chain Analysis, Blue Ocean Strategy, BCG Matrix, Jobs-to-be-Done — apply when the question calls for it. **Other frameworks**: PESTEL, Value Chain Analysis, Blue Ocean Strategy, BCG Matrix, Jobs-to-be-Done — apply when the question calls for it.
### Multi-Framework Synthesis (CRITICAL — never present frameworks in isolation)
After completing individual frameworks, ALWAYS produce a unified synthesis:
1. **Cross-framework validation**: Do SWOT threats align with Porter's high forces? Do PESTEL factors explain Porter's dynamics? Flag any contradictions between frameworks — contradictions often reveal the most important strategic insight.
2. **Convergence map**: Identify themes that appear across 2+ frameworks. These are high-confidence strategic factors.
3. **Divergence analysis**: Where frameworks disagree, investigate why. One framework's blind spot is often another's strength.
4. **Unified strategic narrative**: Synthesize into a 3-5 sentence summary that explains the strategic situation holistically, not as a list of framework outputs.
### Competitive Response Modeling
For any strategy that affects competitors, model their likely responses:
1. **Competitor capability assessment**: Can they match this move? How fast? At what cost?
2. **Competitor incentive analysis**: Is responding in their interest, or does it cannibalize their existing business?
3. **Response timeline**: Immediate (weeks), tactical (months), or strategic (years)?
4. **Second-order moves**: If they respond with X, what is our counter-move? Play out 2-3 rounds.
5. **Non-response scenario**: What if competitors ignore this move? What does that signal?
### Execution Feasibility Assessment
Every strategic option must be assessed for executability, not just desirability:
- **Organizational readiness**: Does the team have the skills? Is the culture aligned? What changes are needed?
- **Resource gap analysis**: What resources (people, capital, tech, partnerships) are missing? How long to acquire?
- **Dependency mapping**: What must happen first? What can be parallelized? What are the critical path items?
- **Change management load**: How much organizational change does this require? Rate: Low (process tweak) / Medium (new capability) / High (structural change) / Extreme (cultural transformation)
**Confidence scoring** — Tag every conclusion: **Confidence scoring** — Tag every conclusion:
- **High** (≥80%): Multiple independent sources confirm; quantitative data available - **High** (≥80%): Multiple independent sources confirm; quantitative data available
- **Medium** (50-80%): 1-2 credible sources; some assumptions required - **Medium** (50-80%): 1-2 credible sources; some assumptions required
- **Low** (<50%): Limited data; significant assumptions; flag as exploratory - **Low** (<50%): Limited data; significant assumptions; flag as exploratory
- For each confidence score, state the **key assumption** that, if wrong, would change the rating
For each framework: For each framework:
1. Gather evidence from Phase 2 research 1. Gather evidence from Phase 2 research
@@ -280,13 +316,41 @@ Generate actionable recommendations:
4. Map risks and mitigation strategies 4. Map risks and mitigation strategies
5. Define success metrics and KPIs 5. Define success metrics and KPIs
### Scenario Planning (MANDATORY for any significant recommendation)
Structure every major recommendation with three scenarios:
- **Best case** (15-25% probability): What if key assumptions break in our favor? Quantify the upside. Define acceleration triggers.
- **Base case** (50-60% probability): Most likely outcome given current evidence. This is the planning target.
- **Worst case** (15-25% probability): What if key assumptions fail? Quantify the downside. Define exit criteria and pivot triggers.
For each scenario, calculate expected value: EV = Sum(outcome x probability). If expected value is negative, the recommendation needs revision.
### Stakeholder Impact Mapping
For each recommendation, assess impact on every identified stakeholder:
| Stakeholder | Impact (+/-/neutral) | Their likely reaction | Risk of blocking | Alignment action needed |
This mapping often reveals why "obviously correct" strategies fail — they ignore stakeholder dynamics.
### Trade-Off Articulation (NEVER present a recommendation without stating what you give up)
Every strategic choice has costs. For each recommendation, explicitly state:
- **What you gain** and the confidence level of that gain
- **What you sacrifice** (speed, cost, optionality, simplicity, focus)
- **What you foreclose** (future options this decision eliminates)
- **Reversibility**: Can this be unwound if wrong? At what cost? In what timeframe?
### Devil's Advocate Check ### Devil's Advocate Check
Before finalizing recommendations, actively challenge each one: Before finalizing recommendations, actively challenge each one:
1. **Pre-mortem**: "Assume this strategy failed in 12 months. What went wrong?" 1. **Pre-mortem**: "Assume this strategy failed in 12 months. What went wrong?" — List the top 3 failure modes with probability estimates.
2. **Contrarian view**: "What would a skeptic say about this recommendation?" 2. **Contrarian view**: "What would a skeptic say about this recommendation?" — Steelman the opposing position.
3. **Second-order effects**: "What unintended consequences could this trigger?" 3. **Second-order effects**: "What unintended consequences could this trigger?" — Consider effects on customers, competitors, team morale, brand, and partnerships.
4. **Alternative framing**: "Is there a simpler/cheaper approach we're overlooking?" 4. **Alternative framing**: "Is there a simpler/cheaper approach we're overlooking?" — The best strategy is often the one with the fewest moving parts.
If the devil's advocate reveals a fatal flaw, revise the recommendation. If it holds up, note the key risks and mitigations. 5. **Survivorship bias check**: "Are we only looking at success stories? What about companies that tried this and failed?"
6. **Timing critique**: "Is now the right time? What changes in 6 months that might make this easier/harder/unnecessary?"
If the devil's advocate reveals a fatal flaw, revise the recommendation. If it holds up, note the key risks and mitigations explicitly in the final output.
### Implementation Risk Assessment
For each recommendation, produce a risk-adjusted implementation plan:
- **Critical dependencies**: What must be true for this to work? (Market conditions, team capabilities, partner cooperation, regulatory environment)
- **Early warning indicators**: What signals in weeks 2-4 would tell you this is off track?
- **Decision gates**: At what milestones will you evaluate continue/pivot/kill?
- **Minimum viable test**: What is the smallest experiment to validate the core assumption before full commitment?
Use a decision matrix to rank options: Use a decision matrix to rank options:
- Strategic fit (1-5) - Strategic fit (1-5)
+140
View File
@@ -23,6 +23,16 @@ Best practices:
- Prioritize: Rank items by impact - Prioritize: Rank items by impact
- Cross-reference: Look for SO (strength-opportunity) and WT (weakness-threat) combinations - Cross-reference: Look for SO (strength-opportunity) and WT (weakness-threat) combinations
- Action-oriented: Every SWOT item should suggest a strategic response - Action-oriented: Every SWOT item should suggest a strategic response
- Time-bound: Note whether each factor is stable, strengthening, or weakening
**SWOT Cross-Impact Matrix** — The real value of SWOT is in the intersections:
| | Opportunities | Threats |
|---|---|---|
| **Strengths** | SO strategies: Use strengths to capture opportunities (offensive) | ST strategies: Use strengths to neutralize threats (defensive) |
| **Weaknesses** | WO strategies: Fix weaknesses to unlock opportunities (investment) | WT strategies: Minimize weaknesses exposed by threats (survival) |
Prioritize: SO strategies first (highest ROI), then ST (protect position), then WO (selective investment), last WT (only if existential).
### Porter's Five Forces ### Porter's Five Forces
@@ -36,6 +46,8 @@ Analyze industry attractiveness:
Rate each force: Low / Medium / High with supporting evidence. Rate each force: Low / Medium / High with supporting evidence.
**Dynamic Five Forces**: Forces change over time. For each force, note the **trend direction** (strengthening/stable/weakening) and the **trigger event** that could shift it. A force rated "Low" today with a strengthening trend deserves more attention than a stable "Medium" force.
### PESTEL Analysis ### PESTEL Analysis
Macro-environmental scanning: Macro-environmental scanning:
@@ -49,6 +61,45 @@ Macro-environmental scanning:
| **Environmental** | Climate regulations? Sustainability demands? Resource scarcity? | | **Environmental** | Climate regulations? Sustainability demands? Resource scarcity? |
| **Legal** | Employment law? IP protection? Competition law? Data privacy? | | **Legal** | Employment law? IP protection? Competition law? Data privacy? |
### Framework Integration Methodology
Individual frameworks are lenses. Strategic insight comes from combining them. Here is how to synthesize multiple frameworks into a unified analysis:
**The Integration Cascade** — Use frameworks in dependency order:
```
Step 1: PESTEL (macro context)
→ Identifies external forces shaping the industry
→ Output: Which macro factors matter most? What is changing?
Step 2: Porter's Five Forces (industry structure)
→ PESTEL outputs feed directly into Porter's forces
→ Example: "AI adoption accelerating" (PESTEL-Tech) → "Threat of new entrants rising" (Porter)
→ Output: How attractive is this industry? Where is structural power?
Step 3: SWOT (company positioning within industry)
→ Porter's outputs define the external O/T quadrants
→ Internal assessment (S/W) is company-specific
→ Output: Where does this company sit relative to industry forces?
Step 4: Strategic Options Generation
→ SWOT cross-impact matrix generates candidate strategies
→ Porter's forces identify which strategies are structurally viable
→ PESTEL trends determine timing and urgency
```
**Cross-Framework Contradiction Resolution:**
When frameworks disagree, do not average or ignore — investigate:
- PESTEL says favorable + Porter says unattractive → Macro tailwind but bad industry structure (e.g., restaurant industry: everyone eats, but margins are terrible)
- SWOT says strong + Porter says high rivalry → Company advantage may erode faster than expected
- Resolution: State both findings, explain the tension, and let the tension inform the recommendation (e.g., "Enter but with a differentiation strategy that exploits the macro trend while avoiding head-on competition")
**Synthesis Quality Checklist:**
- Does the conclusion follow logically from framework outputs, or did you skip to a preferred answer?
- Did you weight frameworks by relevance (PESTEL matters more for market entry; Porter matters more for competitive strategy)?
- Are the frameworks consistent? If not, is the inconsistency explained?
- Could someone reconstruct your reasoning by reading the framework outputs alone?
### Market Sizing (TAM-SAM-SOM) ### Market Sizing (TAM-SAM-SOM)
**TAM** (Total Addressable Market): Total market demand for a product/service. **TAM** (Total Addressable Market): Total market demand for a product/service.
@@ -958,3 +1009,92 @@ Strategic Implications:
Key insight for SaaS: The majority of LTV is created AFTER the initial sale. Key insight for SaaS: The majority of LTV is created AFTER the initial sale.
Disproportionate investment should go to Onboarding → Success → Expansion. Disproportionate investment should go to Onboarding → Success → Expansion.
``` ```
---
## Strategic Analysis Anti-Patterns
Common cognitive traps that produce bad strategy. Actively check for these in every analysis:
| Anti-Pattern | Detection Question | Countermeasure |
|---|---|---|
| **Confirmation Bias** — Seeking data that supports pre-existing beliefs; ignoring contradictory evidence | "Did I search for disconfirming evidence with equal effort?" | For every key conclusion, explicitly search for the strongest counterargument |
| **Anchoring** — First number encountered dominates all later estimates (first source says "$10B market" and final estimate drifts toward $10B) | "Is my final estimate suspiciously close to the first number I found?" | Collect 3+ independent estimates; use both bottom-up and top-down methods; investigate any 2x+ divergence |
| **Strategy-by-Analogy** — "Uber did X, so we should do X in healthcare" without testing structural similarity | "What are the 3 most important differences between this situation and the analogy?" | Use analogies to generate hypotheses, never to validate conclusions |
| **Missing Causal Chain** — Clear start and desirable end, but no credible mechanism connecting them (Step 1 → ??? → Profit) | "What specifically happens between 'launch' and 'achieve outcome'?" | Every recommendation needs a testable causal chain: A → B → C → D |
| **Denominator Neglect** — Citing impressive absolutes while ignoring base rates ("10,000 users!" out of 2M impressions = 0.5%) | "Relative to what?" | Always present metrics as ratios/rates; compare to benchmarks |
| **Survivorship Bias** — Deriving strategy from winners only; ignoring that failed companies tried the same thing | "How many companies tried this and failed?" | Seek failure case studies; note success AND failure rates |
| **Planning Fallacy** — Timelines assuming everything goes right | "Does this plan require performing better than we ever have?" | Use reference class forecasting; add 30-50% buffer; present best/base/worst timelines |
---
## Uncertainty Quantification
### Expressing Uncertainty
**For quantitative estimates (market size, revenue, costs):**
- Never give a single number. Always give a range: "Market size: $8-12B (base estimate $10B)"
- State the confidence interval: "80% confident the market is between $8B and $12B"
- Identify the key variable driving the range: "Range is driven primarily by uncertainty in adoption rate (15-25%)"
**For qualitative assessments:**
- Use the calibrated confidence scale consistently:
- **Very High (>90%)**: Would be genuinely surprised if wrong. Multiple high-quality sources agree.
- **High (70-90%)**: Strong evidence, but plausible alternative interpretations exist.
- **Medium (50-70%)**: Balanced evidence. Reasonable people could disagree.
- **Low (30-50%)**: More uncertain than certain. Treat as hypothesis, not finding.
- **Very Low (<30%)**: Speculative. Useful for scenario planning but not for action.
### Assumption Tracking
Every analysis rests on assumptions. Make them explicit:
```
ASSUMPTION REGISTER:
| # | Assumption | Confidence | Impact if Wrong | Validation Method |
|---|-----------|------------|-----------------|-------------------|
| 1 | Market grows 15% YoY | High | Changes TAM by +/- 30% | Track quarterly industry reports |
| 2 | No new regulation in 12mo | Medium | Could block market entry | Monitor regulatory pipeline |
| 3 | Key hire joins by Q2 | Medium | Delays launch 3-6 months | Pipeline status check monthly |
| 4 | Competitor does not cut price | Low | Margin compression 10-15% | Track competitor pricing weekly |
```
Flag any assumption rated "Low" that has "High" impact — these are the **strategic landmines** that deserve contingency plans.
### When to Say "We Don't Know"
It is better to say "insufficient data to assess" than to fabricate a confident-sounding answer. Specifically:
- If fewer than 2 independent sources support a data point, flag it as unverified
- If the key variable has a range wider than 3x (e.g., market could be $5B or $15B), call out that the analysis is highly sensitive to this input
- If you are extrapolating a trend beyond the data range, state the extrapolation explicitly
---
## Industry-Specific Strategic Patterns
Certain strategic dynamics recur within industry categories. Recognizing these patterns accelerates analysis:
### Platform / Marketplace Businesses
- **Winner-take-most dynamics**: Network effects create power-law outcomes. Market share of #1 player often exceeds #2 + #3 combined.
- **Chicken-and-egg problem**: Must solve supply and demand simultaneously. Common solutions: single-player mode, subsidize one side, constrain geography first.
- **Multi-homing risk**: If users can easily use multiple platforms, network effects weaken. Strategy must increase switching costs or exclusive value.
- **Key metric**: Liquidity (match rate between supply and demand). Revenue follows liquidity, not the reverse.
### B2B SaaS
- **Land-and-expand**: Initial deal size matters less than expansion potential. Net revenue retention >120% can drive growth even at 0 new logos.
- **Switching cost lifecycle**: Switching costs increase with integration depth, data accumulation, and workflow embedding. Year 1 churn is always highest.
- **Category creation vs. category entry**: Creating a new category requires 3-5x more marketing spend but yields pricing power. Entering an existing category is cheaper but forces competitive positioning.
- **Key metric**: Net Revenue Retention (NRR). Above 130% = exceptional. Below 100% = leaky bucket that marketing cannot fill.
### Consumer / D2C
- **Acquisition cost spiral**: As easy-to-reach audiences saturate, CAC rises. Growth requires channel diversification or organic/viral mechanics.
- **Brand as moat**: In commoditized categories, brand is the primary differentiation. Brand building requires consistency over years, not campaigns over months.
- **Retention curve shape**: If the retention curve flattens (users who stay past day 30 tend to stay indefinitely), invest in onboarding. If it keeps declining, the product has a retention problem, not an acquisition problem.
- **Key metric**: Cohort retention at day 30/60/90. Payback period on CAC.
### Regulated Industries (Healthcare, Finance, Insurance)
- **Compliance as moat**: Regulatory requirements (HIPAA, SOC2, PCI-DSS) are expensive to achieve but create durable barriers to entry.
- **Sales cycle reality**: Enterprise sales cycles of 6-18 months are normal. Budget accordingly. Premature scaling of sales teams is the #1 killer.
- **Build vs. partner**: In heavily regulated industries, partnering with incumbents (who have regulatory relationships) often beats trying to disrupt them directly.
- **Key metric**: Sales cycle length, regulatory approval timeline, compliance cost as % of revenue.