Files
Evan d215388039 feat(devops): add auto-evolution loop (PR review + BMAD pipeline) (#94)
* feat(devops): add auto-evolution loop (PR review + BMAD bug/feature pipeline)

Extends the DevOps Hand to periodically scan configured GitHub repos and:
- review open PRs via the existing code-reviewer sub-agent, posting a
  single COMMENT review back to GitHub (never auto-APPROVE)
- triage open issues via labels first, single-prompt LLM fallback
- dispatch actionable issues (bug-fix / feature) to a new implementer
  sub-agent which runs the BMAD pipeline (Brainstorm -> Architect ->
  PRD -> Implement) scaled by bmad_strictness and produces a DRAFT PR

Safety floor (always on):
- draft PRs only, never auto-ready, never merge
- never push to main/master/protected branches
- escalates to devops_queue.json when touching workspace Cargo.toml,
  migrations, secrets, or >30 changed files
- 70% per-turn token budget cap so subsequent ticks have headroom

New settings: auto_evolve, evolution_repos, evolution_check_interval,
bmad_strictness. New sub-agent: agents.implementer. New SKILL.md
sections: Issue Triage Playbook, PR Review Automation, Bug Fix
Playbook, BMAD Feature Pipeline, Draft PR Creation. Three new
dashboard metrics: prs_reviewed, issues_processed, draft_prs_opened.

* fix(devops): address PR review — close blocking + medium + style issues

Blocking (5):
- add max_changed_files setting (was referenced in implementer prompt
  but never defined)
- drop metering_query reference (tool isn't in tools = [...] list);
  agent self-paces against budget instead
- fix \n\n literal in jq --arg for issue cross-link comment; compose
  body in shell with printf so newlines survive
- resolve BASE_BRANCH via /repos/owner/repo .default_branch instead
  of relying on an undefined variable
- complete reviewer-verdict → GitHub review-event mapping (4 cases,
  not just request_changes); block routes through REQUEST_CHANGES
  with a blocking-prefix in the body, approve downgrades to COMMENT

Medium (5):
- correct Phase 6 → Phase 7 in the auto-evolution settings comment
- remove schedule_create busy-loop confusion; Phase 7 fires per-turn
  while the Hand is already frequency = "continuous", with cadence
  enforced via devops_evolution_cursor memory key
- generalize the forbid-main-worktree wording — discover and honor
  whatever pre-commit / pre-push / commit-msg hooks the upstream
  repo configures (was librefang-specific)
- clarify the AI-attribution rule: ban LLM-vendor attribution
  (Claude, GPT, 🤖, etc.) but allow process attribution
  (DevOps Hand → implementer) for traceability
- add USER_TYPE = "Bot" short-circuit that was extracted but never
  applied (bots get a token-cheap skip, not a deep review)

Style (2):
- document the four event_publish event names (devops_evolution_*)
  in a new SKILL.md table alongside the memory-keys table
- justify implementer's max_history_messages = 100 with a comment
  (BMAD 4 phases × cargo build/test chains needs headroom)

* docs(devops): tighten evolution snippets (D1-D4 second-review nits)

D1 -- show SUMMARY_BODY (and VERDICT) assignment in PR review snippet:
add explicit jq -r .summary / .verdict extraction from reviewer_output.json
so the agent reading SKILL.md doesn't have to infer where these come from.

D2 -- reword strict-mode wait semantics in both HAND.toml and SKILL.md:
'Stop. Wait...' was misleading because the agent loop has no in-turn
pause primitive. Now spells out: end the current turn after queueing,
let the continuous tick re-read the queue, resume on approved / skip
on pending / abandon on rejected. Explicitly forbids busy-wait and
sleep loops.

D3 -- restructure bot / huge-diff short-circuit so agent-tool calls are
expressed as numbered agent steps, not as '# memory_store ...' comments
inside a bash block. The bash block now only extracts cheap signals;
the decision and the tool calls are clearly agent-level.

D4 -- remove the misleading 'exit 0' from the short-circuit bash and
add a one-liner noting that exit 0 inside shell_exec only ends one
shell session, not the Phase 7 pass; the agent must choose to move on.
2026-05-14 15:54:42 +09:00

68 lines
3.3 KiB
Markdown

# DevOps Hand
Autonomous DevOps engineer -- CI/CD management, infrastructure monitoring, deployment automation, and incident response.
## Configuration
| Field | Value |
|-------|-------|
| Category | `development` |
| Agent | `devops-hand` |
| Routing | `ci/cd`, `pipeline`, `github actions`, `infrastructure monitoring`, `deployment automation`, `incident response`, `auto evolve`, `review github prs`, `triage issues`, `implement issue`, `fix bug from issue` |
## Integrations
None required.
## Settings
- **Infrastructure Type** -- `cloud`, `kubernetes`, `docker`, `bare_metal`, `serverless`
- **CI/CD Platform** -- `github_actions`, `gitlab_ci`, `jenkins`, `circleci`, `other`
- **Monitoring Focus** -- `uptime`, `performance`, `security`, `cost`, `balanced`
- **Auto Monitor** -- Automatically monitor infrastructure (default: off)
- **Health Check Interval** -- `1min`, `5min`, `15min`, `1hour`
- **Service URLs** -- Comma-separated URLs to monitor
- **Alert on Failure** -- Publish events on health check failures (default: on)
- **Rollback Strategy** -- `manual`, `auto_previous`, `blue_green`
- **Auto Evolution** -- Periodically scan GitHub repos and run PR review / issue triage / BMAD implementation (default: off)
- **Evolution Target Repos** -- Comma-separated `owner/repo` pairs to watch
- **Evolution Check Interval** -- `5min`, `15min`, `1hour`, `6hour`, `1day`
- **BMAD Strictness** -- `light`, `standard`, `strict` -- depth of the Brainstorm-Architect-PRD-Implement pipeline before producing a draft PR
## Usage
```bash
librefang hand run devops
```
## Auto-Evolution Mode
When `auto_evolve = true` and `evolution_repos` is set, the Hand's Phase 7 loop fires on `evolution_check_interval` and, for each watched repo:
1. **Reviews open PRs** -- pulls each PR's diff, asks the `code-reviewer` sub-agent for an assessment, posts a single `COMMENT` review back on GitHub. Already-reviewed `head_sha` values are skipped.
2. **Triages open issues** -- labels first, single-prompt LLM fallback if labels are absent. Result is one of `bug-fix | feature | needs-info | skip`.
3. **Implements actionable issues** -- dispatches `bug-fix` and `feature` issues to the `implementer` sub-agent which runs the BMAD pipeline scaled by `bmad_strictness` and produces a **draft PR**.
### Safety floor (always on)
- Draft PRs only. The Hand never marks PRs ready-for-review and never merges.
- Never pushes to `main` / `master` / protected branches.
- Never `--force` / `--no-verify` / `--amend` against a remote branch.
- Stops and queues to `devops_queue.json` if the change touches `Cargo.toml` workspace members, migration files, or anything under a `secrets` / credential glob.
- Hard cap of 30 changed files per PR; larger changes get split.
- Per-tick token budget capped at 70% so subsequent ticks have headroom.
### Required GitHub token scopes
For public-repo evolution, a fine-grained token with:
- **Pull requests**: read & write (review posting, draft PR creation)
- **Issues**: read & write (triage comments, issue cross-links)
- **Contents**: read & write (branch push)
- **Metadata**: read
For private repos, add the `repo` scope and ensure the repo is listed in `evolution_repos`.
### What it does NOT do
It will never merge a PR, mark a draft as ready, or auto-approve. Human review is always required. See `SKILL.md` -> `What this Hand does NOT do` for the full list.