Files
librefang-registry/hands/devteam/HAND.toml
T
Evan 102b506b0b fix(agents,hands): per-agent/per-hand mcp_servers / skills allowlists (#87) (#92)
All 32 agent manifests and 17 hands shipped with empty mcp_servers /
skills lists, which the kernel interprets as "no filter" — every
globally-configured MCP server's tools and every installed skill get
injected into the prompt on every LLM call. On a typical instance (9
MCP servers, ~85 MCP tools + ~82 built-in tools) that's ~50k input
tokens per turn spent on definitions the agent never uses.

Changes
-------

32 agents/*/agent.toml:
  - mcp_servers: 1-4 per agent. memory wherever state persists across
    turns; fetch / exa-search / brave-search only where the prompt
    actually calls for web; git / github / filesystem on engineering
    agents; gmail / google-calendar / linear / jira on productivity
    agents whose prompts mention them.
  - skills: per-role allowlist driven by what the system_prompt names
    (e.g. coder → rust/python/typescript/git/shell-scripting; devops-
    lead → docker/kubernetes/terraform/ansible/ci-cd/helm/prometheus/
    sysadmin). Generalists (assistant) keep skills = [] (see "Open
    items" below).
  - skills_disabled = true on the four short-conversational agents
    (hello-world, recipe-assistant, health-tracker, home-automation).
    Their system prompts never instruct the LLM to consult any skill,
    so loading all 60 was pure waste. They also drop the explicit
    max_history_messages override and inherit the kernel default (60).
  - max_history_messages tiered by workload shape:
      60  short conversational (hello-world, recipe, health-tracker,
          home-automation) — inherits the rising kernel default
          (`DEFAULT_MAX_HISTORY_MESSAGES = 60`); no override needed.
      60  single-turn task agents (writer, translator, doc-writer,
          email-assistant, customer-support, sales-assistant, recruit-
          er, social-media, personal-finance, tutor, travel-planner,
          meeting-assistant, ops, devops-lead, planner) — explicit
          override at the same value to lock the cap if the kernel
          default moves again.
      80  multi-step / tool-heavy (coder, debugger, architect, code-
          reviewer, test-engineer, security-auditor, analyst, data-
          scientist, academic-researcher, researcher, legal-assistant)
      120 coordinators (assistant, orchestrator) — long multi-agent
          sessions where prompt-cache continuity is critical
    All values sit at or above the kernel default. Pinning lower
    would thrash the prompt cache (the failure mode #91 fixed for
    the creator hand by *raising* the cap, not lowering it).

17 hands/*/HAND.toml:
  - hand-level mcp_servers / skills now declared on every hand, so
    every [agents.*] inside inherits a sensible allowlist.
  - skills_disabled = true placed on each [agents.*] inside clip and
    creator (pure media pipelines that don't benefit from any skill).
    HandDefinitionRaw in librefang-hands does NOT have a top-level
    skills_disabled field — declaring it at the hand top level would
    be silently dropped by serde, so the setting must live on the
    AgentManifest of each sub-agent role.
  - devteam: expand existing mcp_servers = ["github"] to include
    memory / git / filesystem; populate skills with the expected
    dev-team expertise (replacing the placeholder skills = []).
  - wiki: replace placeholder mcp_servers = [] with [memory, fetch,
    filesystem]. Hand-level skills stays [].
  - lead: hand-level skills was originally [email-writer, writing-
    coach, interview-prep]; interview-prep is for job-interview
    preparation, not lead generation. Replaced with data-analyst
    (used by the qualification-scoring step in the prompt).

schema.toml: register mcp_servers / skills / max_history_messages on
the agent field schema so machine consumers (RegistrySchema in
librefang-types) see the new top-level fields. The
max_history_messages description now points at
librefang_runtime::agent_loop::DEFAULT_MAX_HISTORY_MESSAGES (60
today) by name, so the schema doesn't go stale when the constant
moves again.

agents/README.md: example block + "Adding a New Agent" checklist
mention the allowlists; max_history_messages example is shown
commented out with a prompt-cache caveat.

Open items
----------

`assistant` (the default user-facing agent) keeps `skills = []`
deliberately. It is the generalist entry point — capping its skill
surface at a small allowlist would defeat its "delegate to any
specialist" job. The trade-off is that this single agent still pays
the full skill-definition load on every turn; operators who want a
strict allowlist for `assistant` can override it after install.

Why not adopt PR #89's approach
-------------------------------

#89 covers similar ground but with three issues this PR avoids:

1. mcp_servers = ["_none"] sentinel. #89's body explicitly notes
   it's pending upstream librefang#4808 (mcp_disabled). Shipping a
   magic-string today means coming back later to clean it up. This
   PR uses real allowlists.
2. max_history_messages = 8 / 12 / 15 / 20. Far below today's
   kernel default (60) and #91's direction for long-workflow hands
   (80–120). Every turn that hits the cap invalidates the cached
   prompt prefix; the cost of cache misses exceeds the saving from
   shorter history. This PR uses 60–120.
3. Doubling max_llm_tokens_per_hour (coder 200k→500k, assistant
   300k→500k) widens the per-agent budget — the opposite direction
   from #87's "reduce per-call cost" goal. Left to the operator's
   instance-specific tuning.

Refs librefang/librefang-registry#87, librefang/librefang-registry#89
2026-05-12 09:30:21 +09:00

423 lines
14 KiB
TOML
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
id = "devteam"
version = "1.0.0"
name = "Dev Team"
description = "Autonomous software development team — PM triages issues, Engineer implements, QA validates"
category = "development"
icon = "lucide:construction"
# ─── Hand-level resource composition ─────────────────────────────────────────
# Tools available to ALL agents (kernel overrides per-agent capabilities.tools)
tools = [
"shell_exec",
"file_read",
"file_write",
"file_list",
"web_fetch",
"web_search",
"memory_store",
"memory_recall",
"memory_list",
"schedule_create",
"schedule_list",
"schedule_delete",
"knowledge_add_entity",
"knowledge_add_relation",
"knowledge_query",
"event_publish",
"agent_list",
"workflow_run",
]
# Per-hand resource allowlists (refs librefang/librefang-registry#87).
# Inherited by every [agents.*] in this hand unless overridden.
mcp_servers = ["memory", "github", "git", "filesystem"]
skills = [
"github",
"git-expert",
"rust-expert",
"python-expert",
"typescript-expert",
"code-reviewer",
"api-tester",
"ci-cd",
]
# Plugins: useful for dev workflow
allowed_plugins = ["todo-tracker", "auto-summarizer", "episodic-memory"]
# ─── Requirements ────────────────────────────────────────────────────────────
[[requires]]
key = "git"
label = "git"
requirement_type = "binary"
check_value = "git"
[requires.install]
macos = "brew install git"
linux_apt = "sudo apt install git"
[[requires]]
key = "gh"
label = "GitHub CLI (preferred for GitHub operations)"
requirement_type = "binary"
check_value = "gh"
optional = true
[requires.install]
macos = "brew install gh"
linux_apt = "sudo apt install gh"
manual_url = "https://cli.github.com/"
[[requires]]
key = "GITHUB_TOKEN"
label = "GitHub Token"
requirement_type = "api_key"
check_value = "GITHUB_TOKEN"
[requires.install]
signup_url = "https://github.com/settings/tokens"
env_example = "GITHUB_TOKEN=ghp_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"
steps = [
"Go to GitHub Settings → Developer settings → Fine-grained tokens",
"Grant: Issues (read/write), Pull requests (read/write), Contents (read/write)",
"Set as GITHUB_TOKEN environment variable",
]
# ─── Routing ─────────────────────────────────────────────────────────────────
[routing]
aliases = [
"dev team",
"development team",
"software team",
"build a project",
"issue triage",
]
weak_aliases = ["project management", "sprint", "kanban", "team implement"]
# ─── Settings ────────────────────────────────────────────────────────────────
[[settings]]
key = "team_size"
label = "Team Size"
description = "Lite = PM + Engineer (PM handles QA). Standard = PM + Engineer + QA."
setting_type = "select"
default = "standard"
[[settings.options]]
value = "lite"
label = "Lite (PM + Engineer)"
[[settings.options]]
value = "standard"
label = "Standard (PM + Engineer + QA)"
[[settings]]
key = "repo_url"
label = "Repository"
description = "GitHub repository (owner/repo)"
setting_type = "text"
default = ""
[[settings]]
key = "scan_interval"
label = "Issue Scan Interval"
setting_type = "select"
default = "15min"
[[settings.options]]
value = "5min"
label = "Every 5 minutes"
[[settings.options]]
value = "15min"
label = "Every 15 minutes"
[[settings.options]]
value = "1hour"
label = "Every hour"
[[settings.options]]
value = "manual"
label = "Manual only"
[[settings]]
key = "approval_mode"
label = "Approval Mode"
description = "ON = PM waits for human approval before merging. OFF = auto-merge after QA passes."
setting_type = "toggle"
default = "true"
# ─── Agents ──────────────────────────────────────────────────────────────────
# Each agent uses `base` to inherit from a registry agent template.
# Only hand-specific overrides are defined here.
[agents.pm]
coordinator = true
base = "planner"
name = "pm"
description = "Product Manager — triages issues, assigns tasks, tracks progress"
invoke_hint = "Issue triage, task assignment, progress tracking, status reports, rollback coordination"
[agents.pm.model]
max_tokens = 8192
temperature = 0.3
system_prompt = """You are the PM of an autonomous dev team. You coordinate, you do NOT write code.
## Team
Read **User Configuration** for team_size:
- **Lite**: You + Engineer. You handle QA (review PR diffs, verify acceptance criteria).
- **Standard**: You + Engineer + QA. Delegate implementation to Engineer, verification to QA.
## Repo
All agents share one checkout at `../shared/repo/`. Engineer clones it. You read from it or use GitHub MCP/gh CLI.
Read the **Repository** value from **User Configuration** (injected below). If empty, tell the user to configure it and stop.
## Workflow
1. **Startup**: memory_recall devteam_board, create scan schedule (5min=300s, 15min=900s, 1hour=3600s) if not exists
2. **Scan**: Use GitHub MCP tools to list open issues (filter out PRs). Read comments. Skip assigned/wontfix/duplicate.
3. **Triage** (Planner methodology — SCOPE, DECOMPOSE, SEQUENCE, ESTIMATE, RISK, MILESTONE):
- SCOPE: what's in/out for this issue
- Classify (bug/feature/refactor), ESTIMATE size (S/M/L/XL)
- For XL: DECOMPOSE into sub-issues, SEQUENCE by dependencies, identify RISK, set MILESTONEs. Use workflow_run product-spec for vague features.
- Label, comment (English)
4. **Delegate**: Send Engineer issue #, acceptance criteria, branch name, tech stack. Use workflow_run bug-triage for complex bugs.
5. **Review**: Check CI first. Send to QA (standard) or review yourself (lite). Max 3 rounds before escalating to user.
6. **Merge**: If approval_mode ON, wait for user. Check mergeable. Comment on issue. Merge. Close.
7. **Knowledge**: After each resolve, knowledge_add_entity for the fix. Flag modules with 3+ bugs as hotspots.
8. **Standup**: Daily summary via event_publish.
9. **Rollback**: On regression report, revert PR, re-open issue, workflow_run incident-postmortem.
## Rules
- All GitHub comments/reviews in English.
- Max 2 concurrent tasks for Engineer.
- Never close without verification.
- Prioritize: production bugs > regressions > features > tech debt.
"""
[agents.pm.capabilities]
agent_message = ["*"]
memory_read = ["*"]
memory_write = ["self.*", "shared.*"]
shell = ["gh *", "git *"]
[agents.pm.resources]
max_llm_tokens_per_hour = 200000
[agents.engineer]
base = "coder"
name = "engineer"
description = "Full-stack Engineer — designs, implements, tests, handles CI/CD"
invoke_hint = "Code implementation, bug fixing, architecture, CI/CD, tests"
[agents.engineer.model]
max_tokens = 16384
temperature = 0.2
system_prompt = """You are the Engineer of an autonomous dev team. Senior full-stack developer.
## Repo
All agents share `../shared/repo/`. You own it — clone on first task.
Read the **Repository** value from **User Configuration** (injected below your prompt) and substitute it in commands. Example for owner/repo:
```bash
REPO_DIR="../shared/repo"
[ ! -d "$REPO_DIR" ] && mkdir -p ../shared && git clone "https://x-access-token:$GITHUB_TOKEN@github.com/OWNER/REPO.git" "$REPO_DIR" && cd "$REPO_DIR" && git config user.name "DevTeam Hand" && git config user.email "devteam@librefang.ai"
cd "$REPO_DIR" && git checkout -- . 2>/dev/null; git clean -fd 2>/dev/null; git checkout main && git pull
```
## Methodology (inherited from Coder agent)
READ → PLAN → IMPLEMENT → TEST → VERIFY
## Workflow
1. Receive task from PM with issue #, acceptance criteria, branch name
2. Create branch: `git checkout -B {branch} main`
3. Read code, understand context
4. For L/XL: send design to PM for alignment first
5. Implement + write tests
6. Run build/lint/test for the detected stack
7. Commit specific files, push, create PR via GitHub MCP or gh CLI
8. Report to PM: branch, PR #, files changed, build status
## Fix Requests
Same branch. Fix. Push. Reply on PR in English. If rebase needed: `git fetch origin && git rebase origin/main && git push --force-with-lease`
## Workflows
- workflow_run test-generation: when adding test coverage
- workflow_run code-review: for external PR reviews
- workflow_run refactor-plan: for refactoring tasks
- workflow_run api-design: for new API endpoints
## Knowledge
After each task: knowledge_add_entity for what was done and why.
Before each task: knowledge_query for the affected module.
"""
[agents.engineer.capabilities]
network = ["*"]
memory_read = ["*"]
memory_write = ["self.*", "shared.*"]
shell = [
"cargo *",
"rustc *",
"npm *",
"node *",
"python *",
"pip *",
"go *",
"swift *",
"mvn *",
"git *",
"gh *",
"make *",
"docker *",
]
[agents.engineer.resources]
max_llm_tokens_per_hour = 300000
[agents.qa]
base = "code-reviewer"
name = "qa"
description = "QA Engineer — reviews code, runs tests, catches bugs"
invoke_hint = "Code review, testing, quality verification, security audit"
# QA can't modify code — only read and verify
tool_blocklist = ["file_write"]
[agents.qa.model]
max_tokens = 8192
temperature = 0.2
system_prompt = """You are the QA Engineer of an autonomous dev team. Be skeptical — assume bugs until proven otherwise.
## Repo
Shared checkout at `../shared/repo/`. Force-sync to remote:
`cd ../shared/repo && git fetch origin && git checkout -B {branch} origin/{branch}`
## Review Criteria (inherited from Code Reviewer agent)
1. CORRECTNESS: Logic errors, edge cases, error handling
2. SECURITY: Injection, auth, data exposure, input validation
3. PERFORMANCE: Complexity, allocations, I/O patterns
4. MAINTAINABILITY: Naming, structure, separation of concerns
## Severity
[MUST FIX] / [SHOULD FIX] / [NIT] / [PRAISE]
## Workflow
1. Receive branch + PR # + acceptance criteria from PM
2. Checkout branch, read changes: `git diff main...HEAD`
3. Review code against criteria above
4. Run FULL test suite (not just changed files)
5. Identify test gaps — report to PM, don't commit code yourself
6. Submit PR review via GitHub MCP or gh CLI (APPROVE or REQUEST_CHANGES with line comments)
7. Report to PM: PASS or FAIL with details
## Workflows
- workflow_run code-review: for comprehensive parallel review (correctness + security + style)
- workflow_run test-generation: when reporting test gaps, generate specific suggestions
## Principles
- Don't trust the implementation. That's why you exist.
- Test what's NOT tested, not just what is.
- Focus on behavior, not style.
- Run the FULL test suite. Check for regressions.
"""
[agents.qa.capabilities]
memory_read = ["*"]
memory_write = ["self.*", "shared.*"]
# QA has NO file_write — can't modify code, only verify
shell = [
"cargo test *",
"cargo clippy *",
"cargo audit *",
"npm test *",
"npm run lint *",
"pytest *",
"ruff *",
"go test *",
"golangci-lint *",
"swift test *",
"git *",
"gh *",
]
[agents.qa.resources]
max_llm_tokens_per_hour = 150000
# ─── Dashboard ───────────────────────────────────────────────────────────────
[dashboard]
[[dashboard.metrics]]
label = "Issues Triaged"
memory_key = "devteam_issues_triaged"
format = "number"
[[dashboard.metrics]]
label = "Tasks Completed"
memory_key = "devteam_tasks_completed"
format = "number"
[[dashboard.metrics]]
label = "Active Tasks"
memory_key = "devteam_active_tasks"
format = "number"
[[dashboard.metrics]]
label = "QA Pass Rate"
memory_key = "devteam_qa_pass_rate"
format = "percentage"
# ─── Metadata ────────────────────────────────────────────────────────────────
[metadata]
frequency = "continuous"
token_consumption = "medium"
default_active = false
# ─── i18n ────────────────────────────────────────────────────────────────────
[i18n.zh]
name = "开发团队"
description = "自主软件开发团队 — PM 分拣 Issue,工程师实现,QA 验证"
category = "开发"
[i18n.zh.agents.pm]
name = "产品经理"
description = "产品经理 — 分拣 Issue、分配任务、跟踪进度"
[i18n.zh.agents.engineer]
name = "全栈工程师"
description = "全栈工程师 — 设计、实现、测试"
[i18n.zh.agents.qa]
name = "测试工程师"
description = "QA 工程师 — 验证实现、捕获 Bug"
[i18n.ja]
name = "開発チーム"
description = "自律型開発チーム — PMがIssueをトリアージ、エンジニアが実装、QAが検証"
category = "開発"
[i18n.ko]
name = "개발팀"
description = "자율 개발팀 — PM이 이슈 분류, 엔지니어가 구현, QA가 검증"
category = "개발"
[i18n.zh-TW]
name = "開發團隊 Hand"
description = "自主軟體開發團隊 Hand —— PM 分流 Issue、工程師實作、QA 測試、Reviewer 審查,最後編排者統整結果。"
[i18n.de]
name = "Dev-Team-Hand"
description = "Autonomes Softwareentwicklungs-Team: PM triagiert Issues, Engineer implementiert, QA testet, Reviewer prüft, der Orchestrator führt alles zusammen."
[i18n.es]
name = "Hand Dev Team"
description = "Equipo de desarrollo autónomo: el PM clasifica incidencias, el ingeniero implementa, QA prueba, el revisor audita y el orquestador lo integra todo."
[i18n.fr]
name = "Hand Équipe Dev"
description = "Équipe de développement autonome : le PM trie les issues, l’ingénieur implémente, le QA teste, le relecteur audite et l’orchestrateur coordonne."