5 Commits
Author SHA1 Message Date
Evan ff69767793 fix(hands/devteam): raise default token limits (#102)
* fix(hands/devteam): raise role token limits

* fix(hands/devteam): bump patch version
2026-07-16 14:23:35 +09:00
Evan 102b506b0b fix(agents,hands): per-agent/per-hand mcp_servers / skills allowlists (#87) (#92)
All 32 agent manifests and 17 hands shipped with empty mcp_servers /
skills lists, which the kernel interprets as "no filter" — every
globally-configured MCP server's tools and every installed skill get
injected into the prompt on every LLM call. On a typical instance (9
MCP servers, ~85 MCP tools + ~82 built-in tools) that's ~50k input
tokens per turn spent on definitions the agent never uses.

Changes
-------

32 agents/*/agent.toml:
  - mcp_servers: 1-4 per agent. memory wherever state persists across
    turns; fetch / exa-search / brave-search only where the prompt
    actually calls for web; git / github / filesystem on engineering
    agents; gmail / google-calendar / linear / jira on productivity
    agents whose prompts mention them.
  - skills: per-role allowlist driven by what the system_prompt names
    (e.g. coder → rust/python/typescript/git/shell-scripting; devops-
    lead → docker/kubernetes/terraform/ansible/ci-cd/helm/prometheus/
    sysadmin). Generalists (assistant) keep skills = [] (see "Open
    items" below).
  - skills_disabled = true on the four short-conversational agents
    (hello-world, recipe-assistant, health-tracker, home-automation).
    Their system prompts never instruct the LLM to consult any skill,
    so loading all 60 was pure waste. They also drop the explicit
    max_history_messages override and inherit the kernel default (60).
  - max_history_messages tiered by workload shape:
      60  short conversational (hello-world, recipe, health-tracker,
          home-automation) — inherits the rising kernel default
          (`DEFAULT_MAX_HISTORY_MESSAGES = 60`); no override needed.
      60  single-turn task agents (writer, translator, doc-writer,
          email-assistant, customer-support, sales-assistant, recruit-
          er, social-media, personal-finance, tutor, travel-planner,
          meeting-assistant, ops, devops-lead, planner) — explicit
          override at the same value to lock the cap if the kernel
          default moves again.
      80  multi-step / tool-heavy (coder, debugger, architect, code-
          reviewer, test-engineer, security-auditor, analyst, data-
          scientist, academic-researcher, researcher, legal-assistant)
      120 coordinators (assistant, orchestrator) — long multi-agent
          sessions where prompt-cache continuity is critical
    All values sit at or above the kernel default. Pinning lower
    would thrash the prompt cache (the failure mode #91 fixed for
    the creator hand by *raising* the cap, not lowering it).

17 hands/*/HAND.toml:
  - hand-level mcp_servers / skills now declared on every hand, so
    every [agents.*] inside inherits a sensible allowlist.
  - skills_disabled = true placed on each [agents.*] inside clip and
    creator (pure media pipelines that don't benefit from any skill).
    HandDefinitionRaw in librefang-hands does NOT have a top-level
    skills_disabled field — declaring it at the hand top level would
    be silently dropped by serde, so the setting must live on the
    AgentManifest of each sub-agent role.
  - devteam: expand existing mcp_servers = ["github"] to include
    memory / git / filesystem; populate skills with the expected
    dev-team expertise (replacing the placeholder skills = []).
  - wiki: replace placeholder mcp_servers = [] with [memory, fetch,
    filesystem]. Hand-level skills stays [].
  - lead: hand-level skills was originally [email-writer, writing-
    coach, interview-prep]; interview-prep is for job-interview
    preparation, not lead generation. Replaced with data-analyst
    (used by the qualification-scoring step in the prompt).

schema.toml: register mcp_servers / skills / max_history_messages on
the agent field schema so machine consumers (RegistrySchema in
librefang-types) see the new top-level fields. The
max_history_messages description now points at
librefang_runtime::agent_loop::DEFAULT_MAX_HISTORY_MESSAGES (60
today) by name, so the schema doesn't go stale when the constant
moves again.

agents/README.md: example block + "Adding a New Agent" checklist
mention the allowlists; max_history_messages example is shown
commented out with a prompt-cache caveat.

Open items
----------

`assistant` (the default user-facing agent) keeps `skills = []`
deliberately. It is the generalist entry point — capping its skill
surface at a small allowlist would defeat its "delegate to any
specialist" job. The trade-off is that this single agent still pays
the full skill-definition load on every turn; operators who want a
strict allowlist for `assistant` can override it after install.

Why not adopt PR #89's approach
-------------------------------

#89 covers similar ground but with three issues this PR avoids:

1. mcp_servers = ["_none"] sentinel. #89's body explicitly notes
   it's pending upstream librefang#4808 (mcp_disabled). Shipping a
   magic-string today means coming back later to clean it up. This
   PR uses real allowlists.
2. max_history_messages = 8 / 12 / 15 / 20. Far below today's
   kernel default (60) and #91's direction for long-workflow hands
   (80–120). Every turn that hits the cap invalidates the cached
   prompt prefix; the cost of cache misses exceeds the saving from
   shorter history. This PR uses 60–120.
3. Doubling max_llm_tokens_per_hour (coder 200k→500k, assistant
   300k→500k) widens the per-agent budget — the opposite direction
   from #87's "reduce per-call cost" goal. Left to the operator's
   instance-specific tuning.

Refs librefang/librefang-registry#87, librefang/librefang-registry#89
2026-05-12 09:30:21 +09:00
Evan 7881d327a5 refactor: migrate icon fields from emoji to lucide:<name> tokens (#63)
* refactor: migrate icon fields from emoji to lucide:<name> tokens

Every TOML manifest's `icon = "<emoji>"` line is replaced with
`icon = "lucide:<kebab-name>"` — a reference to a lucide-react icon,
which the librefang.ai site and dashboard render as crisp SVG. Reasons
for the switch:

- Emoji render very differently across OS/browser/font stacks; the
  registry catalog looked inconsistent from one row to the next.
- Five manifests (clip / creator / linkedin / reddit / twitter) had
  their icons stored as literal Python-style escape strings
  ("\\U0001F3AC") because the TOML parser upstream never decoded
  them. Switching away from emoji drops that class of bug entirely.
- As a drive-by, also decode the \\uXXXX accent escapes in the
  [i18n.fr] block of hands/creator/HAND.toml so "Créateur" shows
  up correctly.

87 files touched. example manifests left untouched (still "TODO").

* fix: backfill i18n name + drop the single-member email category

- Every existing [i18n.<lang>] block now has a `name` field. 60 files
  previously translated description but kept the English name
  implicitly — which rendered as "some English some Chinese" in the
  registry UI. Fill in the missing name from the English brand (or a
  known localized equivalent: DingTalk→钉钉, Feishu→飞书, Email→
  电子邮件 / メール / E-Mail / Correo / Courriel, and a handful of
  hands that have Chinese product names like 视频剪辑 Hand).
- channels/email.toml was the only item under category="email";
  reclassify it as "messaging" so the sub-category filter chip list
  on the category page isn't littered with singletons.

* feat(i18n): localize 76 agents/integrations/plugins into 7 languages

Adds full [i18n.zh], [i18n.zh-TW], [i18n.ja], [i18n.ko], [i18n.de],
[i18n.es], [i18n.fr] blocks with name + description to every manifest
that previously shipped English-only.

Coverage:
- 32 agents (academic-researcher, analyst, architect, assistant,
  code-reviewer, coder, customer-support, data-scientist, debugger,
  devops-lead, doc-writer, email-assistant, health-tracker,
  hello-world, home-automation, legal-assistant, meeting-assistant,
  ops, orchestrator, personal-finance, planner, recipe-assistant,
  recruiter, researcher, sales-assistant, security-auditor,
  social-media, test-engineer, translator, travel-planner, tutor,
  writer)
- 33 integrations (AWS, Azure, Bitbucket, Brave Search, Discord,
  Dropbox, Elasticsearch, Exa Search, Fetch, Filesystem, GCP, Git,
  GitHub, GitLab, Gmail, Google Calendar, Google Drive, Google Maps,
  Jira, Linear, Memory, MongoDB, Notion, PostgreSQL, Puppeteer, Redis,
  Sentry, Sequential Thinking, Slack, SQLite, Teams, Time, Todoist) —
  brand names kept as-is across all locales, only descriptions
  translated.
- 11 plugins (auto-summarizer, context-decay, conversation-logger,
  episodic-memory, guardrails, keyword-memory, mempalace-indexer,
  sentiment-tracker, todo-tracker, topic-memory, user-profile)

The descriptions are one-line summaries — hand-translated rather than
machine-generated, so technical terms (MCP, PR, CI/CD, etc.) stay
consistent across locales.

* feat(i18n): close remaining per-lang gaps for channels, workflows, devteam

Third pass on i18n coverage. Every non-example manifest now carries a
full set of [i18n.zh], [i18n.zh-TW], [i18n.ja], [i18n.ko], [i18n.de],
[i18n.es], [i18n.fr] blocks.

- 44 channel adapters: added French descriptions (zh/zh-TW/ja/ko/de/es
  were already present). Brand names kept as-is in all locales so users
  recognize Discord / Slack / LINE / etc. consistently.
- 22 workflows: filled zh-TW / ja / ko / de / es / fr blocks. Each
  translation mirrors the existing zh one in structure and tone so the
  catalog reads consistently across locales.
- hands/devteam/HAND.toml: added the four langs that were missing
  (zh-TW, de, es, fr).

Only the 6 templates under examples/ are left without i18n blocks on
purpose — they still contain "TODO:" placeholders.
2026-04-17 22:04:26 +09:00
Evan Hu fc650e4ed3 fix(hands): resolve routing alias conflict between devteam and coder 2026-04-10 12:50:39 +09:00
Evan 2b8259a0c0 feat(hands): add devteam hand (#41)
* feat(hands): add devteam hand -- autonomous software development team

Multi-agent hand with 7 roles (PM, Architect, Frontend, Backend, DevOps, QA, Designer)
and 3 team size tiers (simple/standard/full) for different project scales.

PM coordinator auto-scans GitHub issues, triages, assigns tasks to specialists,
and tracks progress on an in-memory project board.

* refactor(hands): slim devteam to 3 agents (PM + Engineer + QA)

7 agents with serial agent_send = massive token waste and info loss at every
handoff. Merge architect/frontend/backend/devops into one Engineer with full
context. Keep QA separate for independent verification. Drop designer.

Tiers: lite (PM + Engineer) and standard (PM + Engineer + QA).

* fix(hands/devteam): fix workspace isolation and git workflow gaps

- PM uses GitHub API for code browsing, no repo clone needed
- Engineer explicitly clones repo, branches, commits, pushes, creates PR
- QA explicitly clones repo, checks out branch under review
- PM tracks last_scan timestamp to filter already-triaged issues
- approval_mode now means PR stays open for review, not skip commit

* fix(hands/devteam): use shared repo checkout instead of per-agent clones

All 3 agents share one checkout at ../shared/repo/. Engineer clones it
on the first task; PM and QA read from the same path. Eliminates
duplicate clones and cross-workspace visibility issues.

* fix(hands/devteam): read issue comments before triaging

Comments contain clarifications, reproduction steps, duplicate markers,
and resolution status. Also skip already-assigned and wontfix issues.

* fix(hands/devteam): fix interactive git add, add merge/close APIs, add fix iteration flow

- Replace git add -p (interactive) with git add <specific files>
- PM prompt now has explicit merge PR and close issue API calls
- Engineer has explicit fix-request handling (same branch, push, no new PR)

* feat(hands/devteam): add full GitHub interaction -- PR review, issue comments, labels

PM:
- Labels issues during triage, comments triage status
- Scans open PRs for external review requests
- Comments on issues linking merged PRs

Engineer:
- Replies to review comments on PR after fixing
- Reviews external PRs with APPROVE/REQUEST_CHANGES + line comments

QA:
- Leaves PR review (APPROVE or REQUEST_CHANGES with line comments)
- All findings visible on GitHub, not just via agent_send

SKILL.md:
- Added PR diff, reviews, review comments, reply, merge API references

* fix(hands/devteam): enforce English comments, line-level reviews, comment-before-close

- All GitHub comments/reviews must be in English (added global rule)
- PR reviews must use comments[] with path+line, not body-only
- Comment on issue with resolution details BEFORE closing/merging
- Improved comment templates with structured info

* fix(hands/devteam): 8 logic fixes from end-to-end workflow review

1. Filter PRs from Issues API (pull_request key)
2. PM sends PR number to QA for review
3. Deduplicate PR scanning via devteam_reviewed_prs
4. QA reports test gaps instead of pushing code to shared branch
5. branch_strategy wired into Engineer (gitflow branches from develop)
6. approval_mode: ON = wait for human, OFF = auto-merge after QA
7. scan_interval mapped to schedule_create every_secs
8. git checkout -B instead of -b to handle existing branches

* fix(hands/devteam): second-pass review — 6 more logic fixes

1. Engineer extracts PR number from create-PR API response
2. PM falls back to GitHub Contents API when shared repo not yet cloned
3. QA gets external PR review flow (was only on Engineer)
4. PM checks CI status + mergeable before merging
5. PM handles merge conflict (409) by sending back to Engineer to rebase
6. i18n approval_mode description synced with actual semantics

* fix(hands/devteam): third-pass — runtime scenarios

1. Deduplicate cron schedule on daemon restart (check schedule_list first)
2. Max 3 review rounds before escalating to user (prevent infinite loop)
3. Clean working directory before switching tasks (git checkout -- . && git clean)
4. Add user direct commands (work on #42, status, review PR #50)
5. Pass tech_stack to Engineer in task delegation
6. Fix duplicate step numbering in Review Cycle

* fix(hands/devteam): fourth-pass — state consistency and edge cases

1. QA force-syncs to remote branch (git checkout -B origin/branch) for force-push safety
2. Board sync step: reconcile with GitHub each scan cycle (catch external closes/merges)
3. Prune devteam_reviewed_prs of closed PRs, cap done list at 30
4. PM checks CI before sending to QA (don't waste QA on red builds)
5. Stop/cancel command: remove from board, comment on issue
6. Explicit rebase commands for Engineer (fetch + rebase + force-with-lease)

* fix(hands/devteam): fifth-pass — crash prevention

1. Guard empty repo_url: stop and tell user to configure it
2. Add python3 to requires (all JSON parsing depends on it)
3. Engineer git config user.name/email on first clone (prevents commit rejection)
4. Explicit build/lint/test commands per tech stack (Rust/TS/Python/Go/Java/Swift)
5. event_publish on task completion so user gets notified
6. Global rule: check API HTTP status before parsing JSON

* feat(hands/devteam): add gh CLI / MCP / curl API three-layer fallback

- Add GitHub MCP integration (mcp_servers = ["github"])
- Add gh CLI as optional requirement (preferred over curl)
- All 3 agents: gh > MCP > curl priority for GitHub operations
- Add issue_tracker setting (github/linear/jira)
- Add agent_list to shared tools
- SKILL.md: add full gh CLI reference section
- i18n: add issue_tracker translation

* feat(hands/devteam): full MCP/integration/notification layer

MCP allowlist: github, linear, jira, sentry, slack, discord
- Sentry: Engineer reads crash reports/stack traces when fixing bugs
- Slack/Discord: PM posts status updates (triaged, completed, QA results)
- Linear/Jira: alternative issue trackers

New settings: notify_channel (none/slack/discord), issue_tracker (github/linear/jira)
New optional requires: npx (MCP runtime), SENTRY_AUTH_TOKEN

PM prompt: notification section, channel-aware status posting
Engineer prompt: Sentry context lookup for bug fixes
i18n: added translations for new settings

* feat(hands/devteam): workflows, onboarding, knowledge, standup, rollback

Workflows (8 integrated):
- PM: bug-triage, product-spec, weekly-report, incident-postmortem
- Engineer: code-review, test-generation, refactor-plan, api-design
- QA: code-review, test-generation

New capabilities:
- Repo onboarding: first activation analyzes repo structure/stack/CI
- Knowledge accumulation: store lessons per issue, detect module hotspots
- Daily standup: cron schedule, board summary via notify_channel
- Rollback: gh pr revert + postmortem workflow + re-open issue

Also:
- Added workflow_run to tools, skills = [] (all allowed)
- Rewrote README with full architecture, lifecycle, workflow table
- PM prompt now has 15 sections covering full lifecycle

* feat(hands/devteam): per-agent capabilities, resources, profiles, fallbacks

Each agent now has full AgentManifest config (not just system_prompt):

PM:
- profile: automation
- capabilities: web, memory, schedule, knowledge, event, workflow, agent_send
- shell: gh, curl, cat, python3
- resources: 200k tokens/hr

Engineer:
- profile: coding
- capabilities: file r/w, shell, web, memory, knowledge, workflow
- shell: cargo, npm, python, go, swift, mvn, git, gh, docker, make
- resources: 300k tokens/hr, 10 concurrent tools
- network: * (needs to push to GitHub)

QA:
- profile: coding (read-heavy, no file_write)
- capabilities: file read, shell (test/lint commands only), web, workflow
- shell: cargo test/clippy/audit, npm test, pytest, go test, gh
- resources: 150k tokens/hr

All agents have fallback_models configured.

* feat(hands/devteam): rewrite with proper resource composition

First hand to use the new composition features:

Agents:
- PM: base=planner, capabilities restricted to gh/git shell only
- Engineer: base=coder, full shell access, network=*
- QA: base=code-reviewer, tool_blocklist=[file_write], test/lint shells only

Composition:
- base: inherit from agents/planner, agents/coder, agents/code-reviewer
- mcp_servers: github (agents interact via MCP, not curl in prompts)
- workflows: bug-triage, code-review, test-generation via workflow_run tool
- plugins: todo-tracker, auto-summarizer, episodic-memory
- per-agent skills: SKILL-pm.md, SKILL-engineer.md, SKILL-qa.md
- per-agent capabilities: QA can't write files, PM can't run builds

Prompts are clean and focused (role + methodology + principles),
not stuffed with curl commands. GitHub interaction goes through
MCP tools or gh CLI.

* fix(devteam): complete planner methodology in PM prompt

Added SCOPE/SEQUENCE/RISK/MILESTONE keywords from the planner
base template's methodology into the PM's triage workflow.

* docs: update hands README, fix repo_url reference in prompts

- hands/README.md: document full composition model (base, MCP, workflows,
  plugins, per-agent skills, per-agent capabilities)
- Updated hand count to 15 (added devteam)
- Engineer prompt: clarify repo_url comes from User Configuration, not
  a template variable
- PM prompt: same clarification

* fix(devteam): override name/description from base templates

Without explicit name, agents inherit base names (planner/coder/code-reviewer)
instead of hand-specific names (pm/engineer/qa). This affects display and
the prefixed name used in agent registry (devteam:pm vs devteam:planner).

* style: format HAND.toml with taplo
2026-04-10 10:24:26 +08:00