Commit Graph
16 Commits
Author SHA1 Message Date
Evan 102b506b0b fix(agents,hands): per-agent/per-hand mcp_servers / skills allowlists (#87) (#92)
All 32 agent manifests and 17 hands shipped with empty mcp_servers /
skills lists, which the kernel interprets as "no filter" — every
globally-configured MCP server's tools and every installed skill get
injected into the prompt on every LLM call. On a typical instance (9
MCP servers, ~85 MCP tools + ~82 built-in tools) that's ~50k input
tokens per turn spent on definitions the agent never uses.

Changes
-------

32 agents/*/agent.toml:
  - mcp_servers: 1-4 per agent. memory wherever state persists across
    turns; fetch / exa-search / brave-search only where the prompt
    actually calls for web; git / github / filesystem on engineering
    agents; gmail / google-calendar / linear / jira on productivity
    agents whose prompts mention them.
  - skills: per-role allowlist driven by what the system_prompt names
    (e.g. coder → rust/python/typescript/git/shell-scripting; devops-
    lead → docker/kubernetes/terraform/ansible/ci-cd/helm/prometheus/
    sysadmin). Generalists (assistant) keep skills = [] (see "Open
    items" below).
  - skills_disabled = true on the four short-conversational agents
    (hello-world, recipe-assistant, health-tracker, home-automation).
    Their system prompts never instruct the LLM to consult any skill,
    so loading all 60 was pure waste. They also drop the explicit
    max_history_messages override and inherit the kernel default (60).
  - max_history_messages tiered by workload shape:
      60  short conversational (hello-world, recipe, health-tracker,
          home-automation) — inherits the rising kernel default
          (`DEFAULT_MAX_HISTORY_MESSAGES = 60`); no override needed.
      60  single-turn task agents (writer, translator, doc-writer,
          email-assistant, customer-support, sales-assistant, recruit-
          er, social-media, personal-finance, tutor, travel-planner,
          meeting-assistant, ops, devops-lead, planner) — explicit
          override at the same value to lock the cap if the kernel
          default moves again.
      80  multi-step / tool-heavy (coder, debugger, architect, code-
          reviewer, test-engineer, security-auditor, analyst, data-
          scientist, academic-researcher, researcher, legal-assistant)
      120 coordinators (assistant, orchestrator) — long multi-agent
          sessions where prompt-cache continuity is critical
    All values sit at or above the kernel default. Pinning lower
    would thrash the prompt cache (the failure mode #91 fixed for
    the creator hand by *raising* the cap, not lowering it).

17 hands/*/HAND.toml:
  - hand-level mcp_servers / skills now declared on every hand, so
    every [agents.*] inside inherits a sensible allowlist.
  - skills_disabled = true placed on each [agents.*] inside clip and
    creator (pure media pipelines that don't benefit from any skill).
    HandDefinitionRaw in librefang-hands does NOT have a top-level
    skills_disabled field — declaring it at the hand top level would
    be silently dropped by serde, so the setting must live on the
    AgentManifest of each sub-agent role.
  - devteam: expand existing mcp_servers = ["github"] to include
    memory / git / filesystem; populate skills with the expected
    dev-team expertise (replacing the placeholder skills = []).
  - wiki: replace placeholder mcp_servers = [] with [memory, fetch,
    filesystem]. Hand-level skills stays [].
  - lead: hand-level skills was originally [email-writer, writing-
    coach, interview-prep]; interview-prep is for job-interview
    preparation, not lead generation. Replaced with data-analyst
    (used by the qualification-scoring step in the prompt).

schema.toml: register mcp_servers / skills / max_history_messages on
the agent field schema so machine consumers (RegistrySchema in
librefang-types) see the new top-level fields. The
max_history_messages description now points at
librefang_runtime::agent_loop::DEFAULT_MAX_HISTORY_MESSAGES (60
today) by name, so the schema doesn't go stale when the constant
moves again.

agents/README.md: example block + "Adding a New Agent" checklist
mention the allowlists; max_history_messages example is shown
commented out with a prompt-cache caveat.

Open items
----------

`assistant` (the default user-facing agent) keeps `skills = []`
deliberately. It is the generalist entry point — capping its skill
surface at a small allowlist would defeat its "delegate to any
specialist" job. The trade-off is that this single agent still pays
the full skill-definition load on every turn; operators who want a
strict allowlist for `assistant` can override it after install.

Why not adopt PR #89's approach
-------------------------------

#89 covers similar ground but with three issues this PR avoids:

1. mcp_servers = ["_none"] sentinel. #89's body explicitly notes
   it's pending upstream librefang#4808 (mcp_disabled). Shipping a
   magic-string today means coming back later to clean it up. This
   PR uses real allowlists.
2. max_history_messages = 8 / 12 / 15 / 20. Far below today's
   kernel default (60) and #91's direction for long-workflow hands
   (80–120). Every turn that hits the cap invalidates the cached
   prompt prefix; the cost of cache misses exceeds the saving from
   shorter history. This PR uses 60–120.
3. Doubling max_llm_tokens_per_hour (coder 200k→500k, assistant
   300k→500k) widens the per-agent budget — the opposite direction
   from #87's "reduce per-call cost" goal. Left to the operator's
   instance-specific tuning.

Refs librefang/librefang-registry#87, librefang/librefang-registry#89
2026-05-12 09:30:21 +09:00
Evan Hu 98ae51444f fix: format agents/ops/agent.toml to satisfy taplo check 2026-04-27 08:54:49 +09:00
Evan d4f15fd662 feat: add web_search capability to all agents and hands (#74)
Adds the web_search tool to every agent.toml and hand HAND.toml that did
not already declare it. Without this capability the runtime gates the
tool with 'Capability denied: tool not in allowed list', leaving agents
unable to perform web searches even when a search provider is
configured.

For tools arrays that already contained web_fetch, web_search is
inserted directly after it (its natural companion). For arrays without
web_fetch, web_search is appended to the end.

24 files updated total: 21 agents and 3 hands.
2026-04-25 23:07:24 +09:00
Evan d43077afa9 fix(providers): remove ~anthropic, skip ~ prefixes in sync script (#69)
* fix(providers): remove ~anthropic, skip ~ prefixes in sync script

OpenRouter uses ~ prefixes for internal auto-routing aliases (e.g. ~anthropic).
These are not real providers — they already route through openrouter.toml.
The generated ~anthropic.toml was confusing (looked like a stale backup)
and redundant with the existing openrouter provider.

- Delete providers/~anthropic.toml
- Skip provider IDs starting with ~ in sync-pricing.py --create-missing

* fix(providers): remove morph, aider, kwaipilot

- morph: specialized code-editing/patching tool, not a general LLM provider
- aider: CLI meta-tool wrapper (base_url empty), redundant with claude-code/codex-cli/gemini-cli/qwen-code
- kwaipilot: Kwai internal coding assistant routed via OpenRouter, niche

* fix(sync): add morph/aider/kwaipilot to SKIP_PROVIDERS to prevent re-creation

* feat(sync): merge OpenRouter-only providers into openrouter.toml

Instead of generating standalone .toml files that just wrap the OpenRouter
endpoint, merge their models directly into openrouter.toml with the
standard 'openrouter/{provider}/{model}' ID convention.

- Add _build_model_fields() and _model_lines() helpers to deduplicate
  model rendering between standalone and merged paths
- Add merge_into_openrouter() that appends new models idempotently
- generate_provider_toml() now only runs for providers in PROVIDER_API
- --create-missing routes OpenRouter-only providers to merge_into_openrouter

* fix(providers): remove 14 OpenRouter-only standalone files

These providers have no direct public API and all route through
openrouter.ai/api/v1. Per the new sync-pricing.py policy, their models
will be merged into openrouter.toml on the next CI run instead of
living in separate files that just wrap the OpenRouter endpoint.

Removed: allenai, deepcogito, essentialai, inclusionai, inflection,
liquid, meituan, nex-agi, nousresearch, prime-intellect, relace,
switchpoint, tngtech, writer

* fix(providers): remove 7 niche providers with no driver support

No dedicated LLM driver code exists for these providers — they rely
purely on OpenAI-compatible passthrough with no special handling.
Removing them reduces registry noise; users can still reach them via
openrouter.toml if needed.

Removed: microsoft, ibm-granite, xiaomi, upstage, inception, aion-labs, arcee-ai

* fix(providers): remove ai21, chutes, venice

All three use ApiFormat::OpenAI with no special handling — pure passthrough.
No registry entry needed; users can reach them via openrouter.toml or by
adding a custom provider.

* docs(providers): rewrite README with full provider catalog and inclusion criteria

- List all 46 providers grouped by category with descriptions
- Document why each provider exists (direct API, unique endpoint, dedicated driver, local, CLI)
- Add inclusion criteria section explaining when to create standalone files vs merging into openrouter.toml
- Document sync script routing logic
- Update model counts: 49→46 providers, 339→232 models

* docs: add comprehensive READMEs for all registry sections + deepinfra provider

- agents/README.md: 32 agents across 7 categories with capability field reference
- channels/README.md: 44 channels across 5 categories with protocol reference table
- hands/README.md: 18 hands across 5 categories with HAND.toml format guide
- mcp/README.md: 33 MCP servers across 5 categories with transport/auth format
- plugins/README.md: 12 plugins with hook protocol documentation
- skills/README.md: 60 skills across 9 categories with SKILL.md format guide
- providers/deepinfra.toml: add DeepInfra serverless inference (5 models)
2026-04-24 00:02:33 +09:00
Evan 28ea31b073 chore(agents): drop GEMINI_API_KEY fallback from generic templates (#68)
These 11 templates are the general-purpose ones that don't depend on any
particular model at the primary level (`[model] provider = "default"`).
They also ship a secondary `[[fallback_models]]` block pointing at
`gemini-2.0-flash` with `api_key_env = "GEMINI_API_KEY"`.

That default hurts everyone who doesn't happen to have `$GEMINI_API_KEY`
set — every agent boot logs `WARN Fallback driver 'gemini' failed to
init: Missing API key`, once per turn per agent. The templates that
actually intend to use Gemini as their primary model (analyst, coder,
researcher, code-reviewer, debugger, legal-assistant, data-scientist,
academic-researcher, test-engineer) are left untouched — those
declare Gemini in `[model]`, which is an intentional design choice, not
a hidden fallback.

Users who want a Gemini fallback chain for generic agents can add
`[[fallback_models]]` themselves in `~/.librefang/workspaces/agents/...`
once they've set `$GEMINI_API_KEY`.

Removed from:
  assistant, customer-support, devops-lead, doc-writer, email-assistant,
  meeting-assistant, planner, recruiter, sales-assistant, social-media,
  writer
2026-04-21 20:10:09 +09:00
Evan 7881d327a5 refactor: migrate icon fields from emoji to lucide:<name> tokens (#63)
* refactor: migrate icon fields from emoji to lucide:<name> tokens

Every TOML manifest's `icon = "<emoji>"` line is replaced with
`icon = "lucide:<kebab-name>"` — a reference to a lucide-react icon,
which the librefang.ai site and dashboard render as crisp SVG. Reasons
for the switch:

- Emoji render very differently across OS/browser/font stacks; the
  registry catalog looked inconsistent from one row to the next.
- Five manifests (clip / creator / linkedin / reddit / twitter) had
  their icons stored as literal Python-style escape strings
  ("\\U0001F3AC") because the TOML parser upstream never decoded
  them. Switching away from emoji drops that class of bug entirely.
- As a drive-by, also decode the \\uXXXX accent escapes in the
  [i18n.fr] block of hands/creator/HAND.toml so "Créateur" shows
  up correctly.

87 files touched. example manifests left untouched (still "TODO").

* fix: backfill i18n name + drop the single-member email category

- Every existing [i18n.<lang>] block now has a `name` field. 60 files
  previously translated description but kept the English name
  implicitly — which rendered as "some English some Chinese" in the
  registry UI. Fill in the missing name from the English brand (or a
  known localized equivalent: DingTalk→钉钉, Feishu→飞书, Email→
  电子邮件 / メール / E-Mail / Correo / Courriel, and a handful of
  hands that have Chinese product names like 视频剪辑 Hand).
- channels/email.toml was the only item under category="email";
  reclassify it as "messaging" so the sub-category filter chip list
  on the category page isn't littered with singletons.

* feat(i18n): localize 76 agents/integrations/plugins into 7 languages

Adds full [i18n.zh], [i18n.zh-TW], [i18n.ja], [i18n.ko], [i18n.de],
[i18n.es], [i18n.fr] blocks with name + description to every manifest
that previously shipped English-only.

Coverage:
- 32 agents (academic-researcher, analyst, architect, assistant,
  code-reviewer, coder, customer-support, data-scientist, debugger,
  devops-lead, doc-writer, email-assistant, health-tracker,
  hello-world, home-automation, legal-assistant, meeting-assistant,
  ops, orchestrator, personal-finance, planner, recipe-assistant,
  recruiter, researcher, sales-assistant, security-auditor,
  social-media, test-engineer, translator, travel-planner, tutor,
  writer)
- 33 integrations (AWS, Azure, Bitbucket, Brave Search, Discord,
  Dropbox, Elasticsearch, Exa Search, Fetch, Filesystem, GCP, Git,
  GitHub, GitLab, Gmail, Google Calendar, Google Drive, Google Maps,
  Jira, Linear, Memory, MongoDB, Notion, PostgreSQL, Puppeteer, Redis,
  Sentry, Sequential Thinking, Slack, SQLite, Teams, Time, Todoist) —
  brand names kept as-is across all locales, only descriptions
  translated.
- 11 plugins (auto-summarizer, context-decay, conversation-logger,
  episodic-memory, guardrails, keyword-memory, mempalace-indexer,
  sentiment-tracker, todo-tracker, topic-memory, user-profile)

The descriptions are one-line summaries — hand-translated rather than
machine-generated, so technical terms (MCP, PR, CI/CD, etc.) stay
consistent across locales.

* feat(i18n): close remaining per-lang gaps for channels, workflows, devteam

Third pass on i18n coverage. Every non-example manifest now carries a
full set of [i18n.zh], [i18n.zh-TW], [i18n.ja], [i18n.ko], [i18n.de],
[i18n.es], [i18n.fr] blocks.

- 44 channel adapters: added French descriptions (zh/zh-TW/ja/ko/de/es
  were already present). Brand names kept as-is in all locales so users
  recognize Discord / Slack / LINE / etc. consistently.
- 22 workflows: filled zh-TW / ja / ko / de / es / fr blocks. Each
  translation mirrors the existing zh one in structure and tone so the
  catalog reads consistently across locales.
- hands/devteam/HAND.toml: added the four langs that were missing
  (zh-TW, de, es, fr).

Only the 6 templates under examples/ are left without i18n blocks on
purpose — they still contain "TODO:" placeholders.
2026-04-17 22:04:26 +09:00
Evan 4f8dd2404d feat(providers): expand ollama catalog with thinking-capable models (#58)
* feat(providers): expand ollama model catalog with thinking-capable models

Add commonly used local models with accurate capability flags:
- gemma4, gemma3: supports_thinking, supports_vision
- deepseek-r1: supports_thinking (fix missing flag)
- deepseek-v3: supports_tools
- qwen3, qwq: supports_thinking
- llama4: supports_vision
- llama3.3: supports_tools
- phi4: supports_tools

Previously only 6 models were listed and none had supports_thinking
(except deepseek-r1), causing the dashboard to hide thinking toggles
for models that actually support it.

* chore(providers): major cleanup — remove defunct providers and old models

Delete 21 defunct/obscure providers:
aion-labs, arcee-ai, deepcogito, eleutherai, essentialai, ibm-granite,
inception, inflection, kwaipilot, lemonade, liquid, morph, nex-agi,
nousresearch, prime-intellect, reka, relace, switchpoint, tngtech,
upstage, writer

Clean up 10 major providers — keep only latest generation models:
- anthropic: remove claude-3.5-sonnet (superseded by 4.x)
- openai: remove gpt-4o/4-turbo/3.5/o1/o3-mini (superseded by gpt-5/4.1/o3/o4-mini)
- gemini: remove 1.5-*/2.0-flash (superseded by 2.5/3.x)
- deepseek: remove coder/chat-v3-0324 (superseded by r1/v3)
- qwen: remove turbo/2.5-coder (superseded by qwen3)
- groq: remove old llama/mixtral/gemma (keep latest only)
- mistral: remove medium/nemo/pixtral-large (keep large/small/codestral)
- xai: remove grok-2 (superseded by grok-3/4)
- meta-llama: remove 3.x/guard (keep llama-4 + 3.3)
- ollama: rewrite with current models (gemma4, qwen3, qwq, llama4, etc)

Total: 90 → 48 models across major providers. All thinking-capable
models now have supports_thinking = true.

* chore: add pre-commit hook for automatic TOML formatting

- .githooks/pre-commit: runs taplo fmt on staged .toml files
- Makefile: add setup target + auto-configure hooks on first make
- .gitignore: add .make-setup-done and .sync_marker
2026-04-15 22:18:01 +09:00
Evan 733c35df5b fix: add missing tools to hello-world and test-engineer templates
- hello-world: add file_write (LLM needs it for writing files)
- test-engineer: add web_fetch, web_search (needed for researching test patterns)
2026-03-23 14:42:25 +00:00
Evan 945bbbd763 chore(hands): bump all HAND.toml versions to 1.1.0 (#16)
* chore(hands): bump all HAND.toml versions to 1.1.0

Triggers version-aware sync in librefang runtime (librefang/librefang#1530).
Previously sync_subdirs() skipped existing hands regardless of version.
With the runtime fix, bumping from 1.0.0 → 1.1.0 ensures users get
updated hand definitions on next registry sync.

* chore: fix taplo formatting for 4 agent.toml files

* fix(hands): fix invalid install fields in analytics and browser

- analytics: `linux` → `linux_apt`/`linux_dnf`/`linux_pacman` (parser
  only recognizes platform-specific variants, not generic `linux`)
- analytics: remove `pip = "python3 --version"` (version check, not
  an install command)
- browser: remove `pip = "python3 --version"` (same issue)

* fix: enrich sub-agent prompts and add missing requires across all hands

- analytics: fix linux → linux_apt/dnf/pacman, remove invalid pip check,
  enrich analyst and modeler sub-agent prompts
- apitester: add [[requires]] for curl
- browser: remove invalid pip check, enrich researcher and extractor prompts
- clip: enrich editor and transcriber sub-agent prompts
- collector: enrich scout, scholar, and localizer sub-agent prompts
- devops: add [[requires]] for curl, git, docker (optional), GITHUB_TOKEN
  (optional), enrich sub-agent prompts
- lead: enrich outreach, recruiter, and messenger sub-agent prompts
- linkedin: enrich content and researcher sub-agent prompts
- predictor: enrich orchestrator, planner, and modeler sub-agent prompts
- reddit: enrich monitor and composer sub-agent prompts
- strategist: enrich architect, counsel, and analyst sub-agent prompts
- trader: enrich accountant and researcher sub-agent prompts
- twitter: enrich curator and composer sub-agent prompts
2026-03-23 11:21:29 +09:00
Evan d778da72a2 fix(validate): check for [agents] instead of [agent] in HAND.toml (#15)
* fix(validate): check for [agents] instead of [agent] in HAND.toml

All 14 hands use [agents.main] (plural) for multi-agent config,
but the validator was checking for [agent] (singular), causing
all hands to fail validation.

* fix(routing): resolve 19 routing alias collisions

Agent is a sub-unit of hand, so hands take priority for routing.
Remove conflicting aliases from agent side when hand already owns them.

- analyst: remove data analysis, analyze data, dashboard (owned by hand/analytics)
- data-scientist: remove statistical analysis, forecast, prediction (owned by hand/analytics, hand/predictor)
- sales-assistant: remove prospecting, sales, pipeline (owned by hand/lead, hand/devops)
- devops-lead: remove incident response, kubernetes, terraform (owned by hand/devops)
- researcher: remove deep research, research, literature review (owned by hand/researcher)
- academic-researcher: remove literature review, systematic review (owned by hand/researcher)
- social-media: remove duplicate content calendar from weak_aliases
- hand/collector: remove competitive analysis (owned by hand/strategist)
2026-03-23 09:41:29 +09:00
EvanandClaude Opus 4.6 8f2244eb6f chore: remove router agent (#5)
* chore: remove router agent

builtin:router has been replaced by LLM intent routing in the kernel.
Assistant is now the sole entry point — see librefang/librefang#1336.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: format all TOML files with taplo

Fix CI taplo format check by running `taplo fmt` on all 132 TOML files.

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-21 03:26:11 +09:00
Evan Hu 4b38372463 docs: add README for every subdirectory
33 agent READMEs, 14 hand READMEs, 2 plugin READMEs
(echo-memory + hooks), and 2 skill READMEs
(custom-skill-prompt + custom-skill-python).

Each README documents the component's purpose, configuration,
and usage based on its TOML definition.
2026-03-21 02:27:34 +09:00
Evan Hu 206169c1d7 docs: add README for every content directory
Each directory (agents, hands, integrations, plugins, providers,
scripts, skills) now has a README documenting its TOML format,
current contents, and contribution steps.
2026-03-21 02:24:22 +09:00
Evan Hu d1bc8ead69 chore: cleanup repo and enhance validation
- Add .gitignore (.DS_Store, .vscode, __pycache__)
- Remove stale .gitkeep files (directories have content now)
- Expand schema.toml to document all 6 content types (agent, hand, integration, skill, plugin)
- Add plugin validation and contribution guide
- Add id/name vs directory name consistency checks
- Add cross-type routing alias collision detection (14 warnings found)
2026-03-21 02:21:24 +09:00
Evan Hu 17d32ed4a7 feat: sync content definitions from core repo
Copy all TOML content definitions from librefang core repo:
- 33 agent definitions (agents/*/agent.toml)
- 14 hand definitions with docs (hands/*/HAND.toml + SKILL.md)
- 25 integration templates (integrations/*.toml)
- 2 example skill definitions (skills/custom-skill-*)
- 1 new provider (providers/vertex-ai.toml)

Part of the framework-vs-content registry split (RFC v0.7).
2026-03-21 02:06:07 +09:00
Evan Hu ded26ce300 feat: add dir 2026-03-21 01:57:04 +09:00