Commit Graph
15 Commits
Author SHA1 Message Date
Evan 102b506b0b fix(agents,hands): per-agent/per-hand mcp_servers / skills allowlists (#87) (#92)
All 32 agent manifests and 17 hands shipped with empty mcp_servers /
skills lists, which the kernel interprets as "no filter" — every
globally-configured MCP server's tools and every installed skill get
injected into the prompt on every LLM call. On a typical instance (9
MCP servers, ~85 MCP tools + ~82 built-in tools) that's ~50k input
tokens per turn spent on definitions the agent never uses.

Changes
-------

32 agents/*/agent.toml:
  - mcp_servers: 1-4 per agent. memory wherever state persists across
    turns; fetch / exa-search / brave-search only where the prompt
    actually calls for web; git / github / filesystem on engineering
    agents; gmail / google-calendar / linear / jira on productivity
    agents whose prompts mention them.
  - skills: per-role allowlist driven by what the system_prompt names
    (e.g. coder → rust/python/typescript/git/shell-scripting; devops-
    lead → docker/kubernetes/terraform/ansible/ci-cd/helm/prometheus/
    sysadmin). Generalists (assistant) keep skills = [] (see "Open
    items" below).
  - skills_disabled = true on the four short-conversational agents
    (hello-world, recipe-assistant, health-tracker, home-automation).
    Their system prompts never instruct the LLM to consult any skill,
    so loading all 60 was pure waste. They also drop the explicit
    max_history_messages override and inherit the kernel default (60).
  - max_history_messages tiered by workload shape:
      60  short conversational (hello-world, recipe, health-tracker,
          home-automation) — inherits the rising kernel default
          (`DEFAULT_MAX_HISTORY_MESSAGES = 60`); no override needed.
      60  single-turn task agents (writer, translator, doc-writer,
          email-assistant, customer-support, sales-assistant, recruit-
          er, social-media, personal-finance, tutor, travel-planner,
          meeting-assistant, ops, devops-lead, planner) — explicit
          override at the same value to lock the cap if the kernel
          default moves again.
      80  multi-step / tool-heavy (coder, debugger, architect, code-
          reviewer, test-engineer, security-auditor, analyst, data-
          scientist, academic-researcher, researcher, legal-assistant)
      120 coordinators (assistant, orchestrator) — long multi-agent
          sessions where prompt-cache continuity is critical
    All values sit at or above the kernel default. Pinning lower
    would thrash the prompt cache (the failure mode #91 fixed for
    the creator hand by *raising* the cap, not lowering it).

17 hands/*/HAND.toml:
  - hand-level mcp_servers / skills now declared on every hand, so
    every [agents.*] inside inherits a sensible allowlist.
  - skills_disabled = true placed on each [agents.*] inside clip and
    creator (pure media pipelines that don't benefit from any skill).
    HandDefinitionRaw in librefang-hands does NOT have a top-level
    skills_disabled field — declaring it at the hand top level would
    be silently dropped by serde, so the setting must live on the
    AgentManifest of each sub-agent role.
  - devteam: expand existing mcp_servers = ["github"] to include
    memory / git / filesystem; populate skills with the expected
    dev-team expertise (replacing the placeholder skills = []).
  - wiki: replace placeholder mcp_servers = [] with [memory, fetch,
    filesystem]. Hand-level skills stays [].
  - lead: hand-level skills was originally [email-writer, writing-
    coach, interview-prep]; interview-prep is for job-interview
    preparation, not lead generation. Replaced with data-analyst
    (used by the qualification-scoring step in the prompt).

schema.toml: register mcp_servers / skills / max_history_messages on
the agent field schema so machine consumers (RegistrySchema in
librefang-types) see the new top-level fields. The
max_history_messages description now points at
librefang_runtime::agent_loop::DEFAULT_MAX_HISTORY_MESSAGES (60
today) by name, so the schema doesn't go stale when the constant
moves again.

agents/README.md: example block + "Adding a New Agent" checklist
mention the allowlists; max_history_messages example is shown
commented out with a prompt-cache caveat.

Open items
----------

`assistant` (the default user-facing agent) keeps `skills = []`
deliberately. It is the generalist entry point — capping its skill
surface at a small allowlist would defeat its "delegate to any
specialist" job. The trade-off is that this single agent still pays
the full skill-definition load on every turn; operators who want a
strict allowlist for `assistant` can override it after install.

Why not adopt PR #89's approach
-------------------------------

#89 covers similar ground but with three issues this PR avoids:

1. mcp_servers = ["_none"] sentinel. #89's body explicitly notes
   it's pending upstream librefang#4808 (mcp_disabled). Shipping a
   magic-string today means coming back later to clean it up. This
   PR uses real allowlists.
2. max_history_messages = 8 / 12 / 15 / 20. Far below today's
   kernel default (60) and #91's direction for long-workflow hands
   (80–120). Every turn that hits the cap invalidates the cached
   prompt prefix; the cost of cache misses exceeds the saving from
   shorter history. This PR uses 60–120.
3. Doubling max_llm_tokens_per_hour (coder 200k→500k, assistant
   300k→500k) widens the per-agent budget — the opposite direction
   from #87's "reduce per-call cost" goal. Left to the operator's
   instance-specific tuning.

Refs librefang/librefang-registry#87, librefang/librefang-registry#89
2026-05-12 09:30:21 +09:00
Evan 651ff1b34d fix(creator): raise max_history_messages + repair refresh-cache CI (#91)
* fix(creator): raise max_history_messages to 80 for polling workflows

Creator Hand's async video_generate path polls video_status every 15-20s
until completion (1-3 min typical), consuming ~5-15 turns per video
request. Combined workflows (video + TTS + music) plus normal back-and-
forth cross the kernel default of 40 messages quickly, which surfaced
in user logs as:

  WARN run_agent_loop: Trimming old messages at safe turn boundary
    agent=creator:creator-hand total_messages=41 trimming=2
  INFO run_agent_loop: prompt cache metrics for turn
    hit_ratio=0.0 creation=0 read=0

Every turn was hitting the trim cap and invalidating the prompt-cache
prefix. 80 covers ~30 polling iterations plus a comfortable pre-context
window without runaway memory growth. Other hands keep the default 40.

* ci(refresh-cache): open PR instead of pushing directly to main

Branch protection on `main` started rejecting the workflow's auto-commit
with GH006 "Changes must be made through a pull request" — see run
25632824585 on 2026-05-10 against commit 6785807 (the first push that
hit the tightened protection). Direct push is precisely what the file's
own security comment (#1) warns against ("Compromised maintainer pushes
a malicious plugins-index.json directly to main. Mitigation: GitHub
branch protection on main requires PR review"), so the fix preserves
that gate rather than working around it.

The workflow now creates a short-lived `automation/refresh-indexes-<sha>`
branch, commits the regen there, pushes, and opens a PR back to main
via `gh pr create`. Maintainers see a one-click squash-merge.

Permissions: add `pull-requests: write` to the existing `contents: write`
so `gh pr create` can be authorised through the default GITHUB_TOKEN.

The post-merge run on the index PR is a no-op (no diff under
`hands/**`, `plugins/**`, etc. between consecutive states), so no
`[skip ci]` marker is needed and no loop is possible.

Without this fix, every content PR landing on main leaves
plugins-index.json + registry-index.json stale, blocking new agents and
hands from reaching daemons until a maintainer manually regenerates.

* fix(hands): raise max_history_messages on long-workflow coordinators

Three hand coordinators have workflows that routinely exceed the kernel
default history cap on a single user turn:

- researcher (max_iterations=80) — deep web_search → web_fetch →
  summarize loops with multi-source synthesis. 80 iterations × ~4
  messages each → 200+ messages per user turn. Set to 120.
- devops    (max_iterations=60) — incident response and CI/CD fan out
  into long shell_exec chains (logs, retries, post-mortems). Set to 80.
- predictor (max_iterations=60) — long reasoning chains accumulating
  signals across many web/knowledge queries, with scheduled re-checks
  referring back. Set to 80.

Creator's existing override is rephrased "raise above the kernel
default" so the comment stays correct regardless of the order this PR
and the upstream kernel-default bump (librefang side) land in.

Other hands (lead/linkedin/reddit/clip/analytics/apitester/browser/
collector/strategist) stay on the kernel default; the upstream bump
covers them.
2026-05-12 08:55:22 +09:00
Evan 7881d327a5 refactor: migrate icon fields from emoji to lucide:<name> tokens (#63)
* refactor: migrate icon fields from emoji to lucide:<name> tokens

Every TOML manifest's `icon = "<emoji>"` line is replaced with
`icon = "lucide:<kebab-name>"` — a reference to a lucide-react icon,
which the librefang.ai site and dashboard render as crisp SVG. Reasons
for the switch:

- Emoji render very differently across OS/browser/font stacks; the
  registry catalog looked inconsistent from one row to the next.
- Five manifests (clip / creator / linkedin / reddit / twitter) had
  their icons stored as literal Python-style escape strings
  ("\\U0001F3AC") because the TOML parser upstream never decoded
  them. Switching away from emoji drops that class of bug entirely.
- As a drive-by, also decode the \\uXXXX accent escapes in the
  [i18n.fr] block of hands/creator/HAND.toml so "Créateur" shows
  up correctly.

87 files touched. example manifests left untouched (still "TODO").

* fix: backfill i18n name + drop the single-member email category

- Every existing [i18n.<lang>] block now has a `name` field. 60 files
  previously translated description but kept the English name
  implicitly — which rendered as "some English some Chinese" in the
  registry UI. Fill in the missing name from the English brand (or a
  known localized equivalent: DingTalk→钉钉, Feishu→飞书, Email→
  电子邮件 / メール / E-Mail / Correo / Courriel, and a handful of
  hands that have Chinese product names like 视频剪辑 Hand).
- channels/email.toml was the only item under category="email";
  reclassify it as "messaging" so the sub-category filter chip list
  on the category page isn't littered with singletons.

* feat(i18n): localize 76 agents/integrations/plugins into 7 languages

Adds full [i18n.zh], [i18n.zh-TW], [i18n.ja], [i18n.ko], [i18n.de],
[i18n.es], [i18n.fr] blocks with name + description to every manifest
that previously shipped English-only.

Coverage:
- 32 agents (academic-researcher, analyst, architect, assistant,
  code-reviewer, coder, customer-support, data-scientist, debugger,
  devops-lead, doc-writer, email-assistant, health-tracker,
  hello-world, home-automation, legal-assistant, meeting-assistant,
  ops, orchestrator, personal-finance, planner, recipe-assistant,
  recruiter, researcher, sales-assistant, security-auditor,
  social-media, test-engineer, translator, travel-planner, tutor,
  writer)
- 33 integrations (AWS, Azure, Bitbucket, Brave Search, Discord,
  Dropbox, Elasticsearch, Exa Search, Fetch, Filesystem, GCP, Git,
  GitHub, GitLab, Gmail, Google Calendar, Google Drive, Google Maps,
  Jira, Linear, Memory, MongoDB, Notion, PostgreSQL, Puppeteer, Redis,
  Sentry, Sequential Thinking, Slack, SQLite, Teams, Time, Todoist) —
  brand names kept as-is across all locales, only descriptions
  translated.
- 11 plugins (auto-summarizer, context-decay, conversation-logger,
  episodic-memory, guardrails, keyword-memory, mempalace-indexer,
  sentiment-tracker, todo-tracker, topic-memory, user-profile)

The descriptions are one-line summaries — hand-translated rather than
machine-generated, so technical terms (MCP, PR, CI/CD, etc.) stay
consistent across locales.

* feat(i18n): close remaining per-lang gaps for channels, workflows, devteam

Third pass on i18n coverage. Every non-example manifest now carries a
full set of [i18n.zh], [i18n.zh-TW], [i18n.ja], [i18n.ko], [i18n.de],
[i18n.es], [i18n.fr] blocks.

- 44 channel adapters: added French descriptions (zh/zh-TW/ja/ko/de/es
  were already present). Brand names kept as-is in all locales so users
  recognize Discord / Slack / LINE / etc. consistently.
- 22 workflows: filled zh-TW / ja / ko / de / es / fr blocks. Each
  translation mirrors the existing zh one in structure and tone so the
  catalog reads consistently across locales.
- hands/devteam/HAND.toml: added the four langs that were missing
  (zh-TW, de, es, fr).

Only the 6 templates under examples/ are left without i18n blocks on
purpose — they still contain "TODO:" placeholders.
2026-04-17 22:04:26 +09:00
Evan 1e8f75b120 feat: add Chinese (zh) i18n for all hand agents (#23)
Add [i18n.zh.agents.*] sections to all 15 HAND.toml files,
providing Chinese translations for agent names and descriptions.

Total: 51 agent translations across 15 hands.
2026-03-25 12:44:20 +09:00
Evan 9b24879c4f feat: add i18n descriptions to all hands and channels (#22)
Add [i18n.zh], [i18n.zh-TW], [i18n.ja], [i18n.ko], [i18n.de], [i18n.es]
sections with translated descriptions to:
- 15 Hand TOML files (hands/*/HAND.toml)
- 44 Channel TOML files (channels/*.toml)

This enables the website to display localized Hand and Channel descriptions
based on the user's selected language.
2026-03-25 12:25:57 +09:00
Evan 945bbbd763 chore(hands): bump all HAND.toml versions to 1.1.0 (#16)
* chore(hands): bump all HAND.toml versions to 1.1.0

Triggers version-aware sync in librefang runtime (librefang/librefang#1530).
Previously sync_subdirs() skipped existing hands regardless of version.
With the runtime fix, bumping from 1.0.0 → 1.1.0 ensures users get
updated hand definitions on next registry sync.

* chore: fix taplo formatting for 4 agent.toml files

* fix(hands): fix invalid install fields in analytics and browser

- analytics: `linux` → `linux_apt`/`linux_dnf`/`linux_pacman` (parser
  only recognizes platform-specific variants, not generic `linux`)
- analytics: remove `pip = "python3 --version"` (version check, not
  an install command)
- browser: remove `pip = "python3 --version"` (same issue)

* fix: enrich sub-agent prompts and add missing requires across all hands

- analytics: fix linux → linux_apt/dnf/pacman, remove invalid pip check,
  enrich analyst and modeler sub-agent prompts
- apitester: add [[requires]] for curl
- browser: remove invalid pip check, enrich researcher and extractor prompts
- clip: enrich editor and transcriber sub-agent prompts
- collector: enrich scout, scholar, and localizer sub-agent prompts
- devops: add [[requires]] for curl, git, docker (optional), GITHUB_TOKEN
  (optional), enrich sub-agent prompts
- lead: enrich outreach, recruiter, and messenger sub-agent prompts
- linkedin: enrich content and researcher sub-agent prompts
- predictor: enrich orchestrator, planner, and modeler sub-agent prompts
- reddit: enrich monitor and composer sub-agent prompts
- strategist: enrich architect, counsel, and analyst sub-agent prompts
- trader: enrich accountant and researcher sub-agent prompts
- twitter: enrich curator and composer sub-agent prompts
2026-03-23 11:21:29 +09:00
Evan Hu 506a201329 feat: muti agent hand 2026-03-23 02:41:02 +09:00
Evan Hu 33d279889c feat(hands): complete i18n fixes, SKILL.md enhancements, and README overhaul
- Fix French accent characters (é/è/ê/ç/â/ô) across all 14 HAND.toml files
- Fix German special characters (ä/ö/ü/ß) across all 14 HAND.toml files
- Add category translations to all 6 i18n language blocks in all 14 hands
- Enhance SKILL.md content for 9 hands with practical examples and workflows
- Trim bloated SKILL.md files (apitester 1400→892, devops 1301→870)
- Rewrite root README.md with accurate stats, complete hand/integration tables
- Update hands/README.md with full 14-hand listing and i18n documentation
2026-03-23 00:18:18 +09:00
Evan Hu 315f955ce2 fix(i18n): move [i18n.zh] sections to end of HAND.toml files
The i18n table was placed before category/icon/tools, causing TOML
parser to swallow subsequent keys into the i18n table.
2026-03-22 23:04:17 +09:00
Evan Hu 123350e2b2 feat(i18n): add Chinese translations for all 14 hands
Add [i18n.zh] sections with localized name and description to every
HAND.toml in the registry.
2026-03-22 22:55:34 +09:00
Evan Hu 8f0cdfe5ab fix: format TOML files with taplo 2026-03-22 16:56:37 +09:00
Evan Hu 5d947e8da0 feat(hands): add version fields, approval_mode, and stronger routing aliases
- Add version = "1.0.0" to all 14 hands
- Add approval_mode toggle to clip and devops hands (queue actions
  for user review before executing)
- Strengthen routing aliases for analytics, apitester, browser,
  collector, devops, lead, predictor, researcher, strategist
2026-03-22 16:42:27 +09:00
EvanandClaude Opus 4.6 8f2244eb6f chore: remove router agent (#5)
* chore: remove router agent

builtin:router has been replaced by LLM intent routing in the kernel.
Assistant is now the sole entry point — see librefang/librefang#1336.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: format all TOML files with taplo

Fix CI taplo format check by running `taplo fmt` on all 132 TOML files.

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-21 03:26:11 +09:00
Evan Hu 4b38372463 docs: add README for every subdirectory
33 agent READMEs, 14 hand READMEs, 2 plugin READMEs
(echo-memory + hooks), and 2 skill READMEs
(custom-skill-prompt + custom-skill-python).

Each README documents the component's purpose, configuration,
and usage based on its TOML definition.
2026-03-21 02:27:34 +09:00
Evan Hu 17d32ed4a7 feat: sync content definitions from core repo
Copy all TOML content definitions from librefang core repo:
- 33 agent definitions (agents/*/agent.toml)
- 14 hand definitions with docs (hands/*/HAND.toml + SKILL.md)
- 25 integration templates (integrations/*.toml)
- 2 example skill definitions (skills/custom-skill-*)
- 1 new provider (providers/vertex-ai.toml)

Part of the framework-vs-content registry split (RFC v0.7).
2026-03-21 02:06:07 +09:00