diff --git a/hands/apitester/HAND.toml b/hands/apitester/HAND.toml index 3f91665..0240fb3 100644 --- a/hands/apitester/HAND.toml +++ b/hands/apitester/HAND.toml @@ -151,6 +151,13 @@ description = "Treat any non-2xx response as a failure (vs allowing expected err setting_type = "toggle" default = "false" +[[settings]] +key = "approval_mode" +label = "Approval Mode" +description = "Write test plans and destructive requests to a queue file for your review instead of executing directly" +setting_type = "toggle" +default = "true" + # ─── Agent configuration ───────────────────────────────────────────────────── [agent] @@ -220,6 +227,20 @@ Store discovered endpoints in the knowledge graph. ## Phase 2 — Functional Testing +**Check `approval_mode` setting before executing any tests.** + +If `approval_mode` is ENABLED: +1. Build the full test plan (all endpoints, methods, payloads) and write it to `apitester_queue.json`: + ```json + [{"id": "t_001", "endpoint": "/api/users", "method": "POST", "payload": {...}, "type": "functional", "status": "pending"}] + ``` +2. Write a human-readable `apitester_queue_preview.md` for easy review +3. Only execute **safe read-only requests** (GET, HEAD, OPTIONS) directly +4. Do NOT execute any write requests (POST, PUT, PATCH, DELETE) — queue them for approval + +If `approval_mode` is DISABLED: +Execute all tests directly. + For each discovered endpoint: 1. **Method validation**: Send requests with correct and incorrect HTTP methods @@ -272,6 +293,14 @@ If no baseline exists, current results become the new baseline. If `test_mode` includes load testing: +If `approval_mode` is ENABLED: +- Write the load test plan to `apitester_queue.json` (target URL, concurrency levels, expected duration) +- Do NOT execute load tests — queue them for user approval +- Alert the user that load tests can impact production systems + +If `approval_mode` is DISABLED: +Execute load tests directly. + Use curl in a loop or shell-based load generator: ``` for i in $(seq 1 100); do @@ -304,6 +333,14 @@ Start with 10 concurrent, then 50, then 100 requests. If `test_mode` includes security: +If `approval_mode` is ENABLED: +- Write the security test plan to `apitester_queue.json` (injection payloads, auth bypass attempts, etc.) +- Do NOT execute security tests — queue them for user approval +- Security tests can trigger alerts and block accounts — always require review + +If `approval_mode` is DISABLED: +Execute security tests directly. + 1. **Authentication tests**: Missing auth, invalid auth, expired tokens 2. **Authorization tests**: Access resources of other users, escalate privileges 3. **Input injection**: SQL injection, XSS, command injection in parameters @@ -385,6 +422,7 @@ If `auto_schedule` is enabled, create scheduled runs via schedule_create. - If an endpoint returns 5xx repeatedly, back off and report the issue - Use realistic but fake test data (e.g. "test@example.com", not real emails) - Always include request/response details in failure reports +- In approval_mode (default), ALWAYS write to queue — NEVER execute write requests, load tests, or security tests without user review """ [dashboard]