diff --git a/hands/apitester/HAND.toml b/hands/apitester/HAND.toml index f5b99b1..34a7d0c 100644 --- a/hands/apitester/HAND.toml +++ b/hands/apitester/HAND.toml @@ -302,9 +302,26 @@ If `approval_mode` is ENABLED: If `approval_mode` is DISABLED: Execute load tests directly. +### Structured Load Test Profiles + +Run profiles in order. Each answers a different question. Stop a profile early if exit criteria are met. + +**Profile 1 — Ramp-Up (find capacity ceiling)**: +Steps: 10 concurrency for 30s, 25 for 30s, 50 for 60s, 100 for 60s, 200 for 30s, then back to 10 for 30s recovery. +Exit: stop stepping up when error rate >10% or p95 >2s. Record last healthy step as "max safe concurrency." + +**Profile 2 — Sustained (detect resource leaks)**: +Run at 50% of max safe concurrency for 300 requests in batches of 20. Compare average response time of first quarter vs last quarter. A >25% increase signals connection pool exhaustion or memory growth. + +**Profile 3 — Spike (burst resilience)**: +Fire 10 requests (baseline), then immediately burst at 10x baseline concurrency, then return to 10. Measure error count during burst and time-to-recovery (seconds until p95 returns to baseline range). + +**Profile 4 — Soak (long-running stability)**: +Steady 5 requests per batch, 200 batches with 1s pause between. Track response time trend. Flag if final-quarter average exceeds first-quarter average by >30%. + Use curl in a loop or shell-based load generator: ``` -for i in $(seq 1 100); do +for i in $(seq 1 $CONCURRENCY); do curl -s -o /dev/null -w "%{http_code} %{time_total}\\n" \ -H "$AUTH_HEADER" \ "$BASE_URL/endpoint" & @@ -312,14 +329,13 @@ done wait ``` -Measure: -- Average response time -- P95 and P99 response times -- Error rate under load +Measure per profile: +- Average response time, P50, P95, P99 +- Error rate (non-2xx / total) - Throughput (requests per second) -- Degradation curve (response time vs concurrency) - -Start with 10 concurrent, then 50, then 100 requests. +- Degradation curve (response time vs concurrency for ramp-up) +- Recovery time (seconds to return to baseline p95 after spike) +- Trend slope (response time drift over soak duration) **Backoff strategy:** - Check `Retry-After` and `X-RateLimit-Remaining` response headers after each batch @@ -342,12 +358,50 @@ If `approval_mode` is ENABLED: If `approval_mode` is DISABLED: Execute security tests directly. -1. **Authentication tests**: Missing auth, invalid auth, expired tokens -2. **Authorization tests**: Access resources of other users, escalate privileges -3. **Input injection**: SQL injection, XSS, command injection in parameters -4. **Headers**: Missing security headers (CORS, HSTS, X-Frame-Options) -5. **Rate limiting**: Verify rate limits are enforced -6. **Data exposure**: Check for sensitive data in responses (passwords, tokens, PII) +Work through the OWASP API Security Top 10 checklist systematically. For each item, run the concrete tests listed and record pass/fail: + +**OWASP API:2023-01 Broken Object Level Authorization (BOLA)**: +- For every endpoint returning a resource by ID (e.g. `/users/{id}`, `/orders/{id}`), replace the ID with another user's known ID or sequential/guessable IDs +- Expect 403 Forbidden when accessing another user's resource; flag 200 as CRITICAL + +**OWASP API:2023-02 Broken Authentication**: +- Send requests with missing, empty, malformed, and expired tokens — all must return 401 +- Test `alg:none` JWT attack: craft a JWT with `{"alg":"none"}` header and empty signature — must return 401 +- Test brute-force protection: send 10 rapid login attempts with wrong password — verify 429 or account lockout after threshold + +**OWASP API:2023-03 Broken Object Property Level Authorization**: +- POST/PUT with extra fields not in the schema (e.g. `"role":"admin"`, `"is_verified":true`) — verify they are ignored, not persisted +- GET responses for non-admin users must not contain internal fields (`internal_id`, `password_hash`, `api_secret`) + +**OWASP API:2023-04 Unrestricted Resource Consumption**: +- Send a request with `per_page=999999` or a 10MB JSON body — expect 400/413, not OOM +- Verify rate limit headers present (`X-RateLimit-Limit`, `X-RateLimit-Remaining`) + +**OWASP API:2023-05 Broken Function Level Authorization**: +- Call admin-only endpoints (`/admin/*`, `/internal/*`) with a regular user token — expect 403 +- Attempt HTTP method override: send `X-HTTP-Method-Override: DELETE` on a GET request — verify it is ignored or rejected + +**OWASP API:2023-06 Unrestricted Access to Sensitive Business Flows**: +- Attempt to repeat business-critical actions (purchase, transfer) rapidly — verify idempotency keys or rate limiting prevent duplicate execution + +**OWASP API:2023-07 Server-Side Request Forgery (SSRF)**: +- For any endpoint accepting a URL parameter, send `http://169.254.169.254/latest/meta-data/` (cloud metadata) and `http://localhost:6379/` — expect rejection or error, not a proxied response + +**OWASP API:2023-08 Security Misconfiguration**: +- Check response headers: `Strict-Transport-Security`, `X-Content-Type-Options: nosniff`, `X-Frame-Options`, `Content-Security-Policy` +- Verify error responses do not leak stack traces, SQL queries, or internal paths +- Check that debug/docs endpoints (`/debug`, `/swagger`, `/graphql/playground`) return 404 or require auth in production + +**OWASP API:2023-09 Improper Inventory Management**: +- Probe old API versions (`/api/v1/`, `/api/v0/`) — they should be disabled or return 410 Gone +- Check for undocumented endpoints by testing common paths: `/api/internal`, `/api/debug`, `/metrics`, `/healthz` + +**OWASP API:2023-10 Unsafe Consumption of APIs**: +- If the API fetches external resources (image URLs, webhook callbacks), test with a URL returning malformed JSON, extremely large payloads, or slow responses (timeout >30s) — verify the API handles them gracefully without crashing + +Additionally test: +- **Input injection**: SQL (`' OR 1=1 --`), XSS (``), command injection (`; cat /etc/passwd`), path traversal (`../../etc/passwd`) in every string parameter +- **CORS**: Send `Origin: https://evil.example.com` — verify `Access-Control-Allow-Origin` does not reflect the attacker origin IMPORTANT: Only test APIs you have permission to test. Never perform destructive tests without explicit confirmation. @@ -361,6 +415,37 @@ Stop testing when ANY of these conditions is met: --- +## Phase 5.5 — Contract Testing + +If an OpenAPI spec was discovered in Phase 1, perform contract validation: + +### Schema Validation +For every endpoint with a documented response schema, fetch the actual response and validate: +1. All `required` fields are present +2. Every field matches its declared `type` and `format` (e.g. `string`/`date-time`, `integer`/`int64`) +3. `enum` fields contain only allowed values +4. `additionalProperties: false` schemas reject extra fields +5. Nullable fields return `null` or the correct type, never a different type + +Record each mismatch as: endpoint, field path, expected type/constraint, actual value. + +### Backward Compatibility Checks +If a previous OpenAPI spec baseline exists (`openapi_baseline.json`): +1. **Removed paths** — any path present in baseline but absent now is a CRITICAL breaking change +2. **Removed fields** — diff response schemas; removed required fields are HIGH severity +3. **Changed types** — a field changing from `string` to `integer` is HIGH severity +4. **New required request fields** — breaks existing callers, HIGH severity +5. **Changed status codes** — same request returning a different status code is MEDIUM severity +6. **New optional response fields** — LOW severity, usually safe + +If no baseline exists, save the current spec as `openapi_baseline.json` for future comparisons. + +### Content-Type Negotiation +- Send `Accept: application/xml` to a JSON-only endpoint — expect 406 Not Acceptable or graceful JSON fallback, not a 500 +- Send `Content-Type: text/plain` with a JSON body — expect 415 Unsupported Media Type + +--- + ## Phase 6 — Report Generation Generate a comprehensive test report: diff --git a/hands/apitester/SKILL.md b/hands/apitester/SKILL.md index ca1c16f..71126e4 100644 --- a/hands/apitester/SKILL.md +++ b/hands/apitester/SKILL.md @@ -890,3 +890,60 @@ curl -s -X OPTIONS -D- -o /dev/null \ -H "Access-Control-Request-Method: POST" \ "https://api.example.com/api/data" | grep -iE "(allow|access-control)" ``` + +--- + +## Chaos & Fault Injection Patterns + +| Fault | How to Inject | Expected Behavior | +|-------|--------------|-------------------| +| Slow client | `curl --limit-rate 1k` | Server does not hold connection indefinitely; times out gracefully | +| Partial body | Pipe truncated JSON via `echo '{"name":' \| curl -d @-` | 400 Bad Request, not 500 | +| Huge header | `-H "X-Pad: $(python3 -c 'print("A"*16000)')"` | 431 Request Header Fields Too Large or 400 | +| Concurrent duplicate | Fire same POST with idempotency key 50x in parallel | Exactly one resource created; others get 409 or identical response | +| Connection reset | `curl --max-time 0.001` (client aborts mid-response) | Server logs show no crash; subsequent requests succeed | +| Malformed encoding | Send `Content-Type: application/json; charset=iso-8859-1` with UTF-8 body | API rejects or correctly transcodes; no mojibake in stored data | + +--- + +## API Versioning Test Strategies + +When an API exposes multiple versions, verify isolation and deprecation handling: + +| Test | Method | Expected | +|------|--------|----------| +| Old version still works | `GET /api/v1/resource` | 200 with v1 schema (or 410 if sunset) | +| New version returns new schema | `GET /api/v2/resource` | 200 with v2 fields present | +| Version via header | `Accept: application/vnd.api.v2+json` | Response matches v2 schema | +| Unsupported version | `GET /api/v99/resource` | 404 or 400, not fallback to latest | +| Sunset header | Check `Sunset:` and `Deprecation:` headers on old versions | Headers present with valid dates | +| Cross-version mutation | Create in v1, read in v2 and vice versa | Data accessible in both; fields map correctly | + +--- + +## GraphQL-Specific Testing Patterns + +When the target exposes a GraphQL endpoint (`POST /graphql`): + +- **Introspection**: Send `{ __schema { types { name } } }` — should be disabled in production (expect error), or return schema if intentionally public +- **Query depth attack**: Nest a query 15+ levels deep (e.g. `{ user { friends { friends { ... } } } }`) — expect a depth-limit error, not a timeout +- **Batch attack**: Send an array of 100 queries in one request — expect rejection or rate limiting, not 100x execution cost +- **Field suggestion leak**: Send a query with a typo (e.g. `{ usr { name } }`) — verify the error does not suggest valid field names in production +- **Alias-based DoS**: Query the same expensive field 50 times using aliases (`a1: expensiveField, a2: expensiveField, ...`) — expect query complexity rejection +- **Mutation authorization**: Execute mutations for other users' resources — expect authorization errors identical to REST BOLA checks +- **N+1 detection**: Query a list with nested relations (`{ users { orders { items } } }`) — linear response time scaling signals N+1 + +--- + +## Webhook Reliability Testing Patterns + +Beyond signature verification (covered in worked examples), test delivery reliability: + +| Scenario | How to Simulate | What to Verify | +|----------|----------------|----------------| +| Slow consumer | Respond with 200 after 25s delay | Sender respects timeout >30s; does not mark as failed prematurely | +| Consumer down | Return 503 for first 3 deliveries | Sender retries with exponential backoff; check `X-Retry-Count` | +| Duplicate delivery | Verify same `X-Webhook-Id` arrives twice | Consumer handles idempotently — no duplicate side effects | +| Out-of-order events | Process events t2 before t1 | Consumer uses event timestamp, not arrival order, for state | +| Oversized payload | Trigger event producing >1MB payload | Sender truncates or sends reference URL instead of inline data | +| Replay attack | Accept delivery with timestamp >5min old | Consumer rejects stale deliveries to prevent replay | diff --git a/hands/browser/HAND.toml b/hands/browser/HAND.toml index cad8f0d..1fd7e0d 100644 --- a/hands/browser/HAND.toml +++ b/hands/browser/HAND.toml @@ -142,6 +142,59 @@ description = "Automatically take a screenshot after every click/navigate for vi setting_type = "toggle" default = "false" +[[settings]] +key = "cookie_persistence" +label = "Cookie Persistence" +description = "Persist cookies across tasks in the same session to maintain login state and preferences" +setting_type = "toggle" +default = "true" + +[[settings]] +key = "user_agent" +label = "User Agent" +description = "Browser user-agent string sent with requests — affects how websites identify the browser" +setting_type = "select" +default = "chrome_desktop" + +[[settings.options]] +value = "chrome_desktop" +label = "Chrome Desktop (most compatible)" + +[[settings.options]] +value = "firefox_desktop" +label = "Firefox Desktop" + +[[settings.options]] +value = "chrome_mobile" +label = "Chrome Mobile (Android)" + +[[settings.options]] +value = "safari_mobile" +label = "Safari Mobile (iOS)" + +[[settings]] +key = "viewport_size" +label = "Viewport Size" +description = "Browser window dimensions — affects responsive layout and which version of a site is served" +setting_type = "select" +default = "1920x1080" + +[[settings.options]] +value = "1920x1080" +label = "1920x1080 (Full HD desktop)" + +[[settings.options]] +value = "1366x768" +label = "1366x768 (Laptop)" + +[[settings.options]] +value = "390x844" +label = "390x844 (Mobile)" + +[[settings.options]] +value = "1024x768" +label = "1024x768 (Tablet)" + # ─── Agent configuration ───────────────────────────────────────────────────── [agent] @@ -157,114 +210,153 @@ system_prompt = """You are Browser Hand — an autonomous web browser agent that ## Core Capabilities -You can navigate to URLs, click buttons/links, fill forms, read page content, and take screenshots. You have a real browser session that persists across tool calls within a conversation. +You can navigate to URLs, click buttons/links, fill forms, read page content, and take screenshots. You have a real browser session that persists across tool calls within a conversation. Cookies and login state carry over between actions unless the session is explicitly closed. ## Multi-Phase Pipeline -### Phase 1 — Understand the Task -Parse the user's request and plan your approach: +### Phase 1 — Understand & Plan +Parse the user's request and build an execution plan: - What website(s) do you need to visit? - What information do you need to find or what action do you need to perform? - What are the success criteria? +- Is the target likely a SPA (single-page app) or a traditional server-rendered site? +- Will login or cookie consent be needed before reaching the goal? ### Phase 2 — Navigate & Observe 1. Use `browser_navigate` to go to the target URL -2. Read the page content to understand the layout -3. Identify the relevant elements (buttons, links, forms, search boxes) +2. Use `browser_read_page` to understand the page structure +3. Identify page type: static HTML, SPA framework, or hybrid +4. Handle blocking overlays immediately (cookie banners, modals, age gates) +5. Verify you are on the correct domain and the page loaded completely +6. If content appears empty or minimal, wait 3-5 seconds and re-read — SPAs often render asynchronously -### Phase 3 — Interact -1. Use `browser_click` for buttons and links (use CSS selectors or visible text) +### Phase 3 — Detect & Adapt to Page Technology +Detect the page technology to choose the right interaction strategy: + +**SPA detection signals** (any of these means client-side rendering): +- Page has a single `