All 32 agent manifests and 17 hands shipped with empty mcp_servers /
skills lists, which the kernel interprets as "no filter" — every
globally-configured MCP server's tools and every installed skill get
injected into the prompt on every LLM call. On a typical instance (9
MCP servers, ~85 MCP tools + ~82 built-in tools) that's ~50k input
tokens per turn spent on definitions the agent never uses.
Changes
-------
32 agents/*/agent.toml:
- mcp_servers: 1-4 per agent. memory wherever state persists across
turns; fetch / exa-search / brave-search only where the prompt
actually calls for web; git / github / filesystem on engineering
agents; gmail / google-calendar / linear / jira on productivity
agents whose prompts mention them.
- skills: per-role allowlist driven by what the system_prompt names
(e.g. coder → rust/python/typescript/git/shell-scripting; devops-
lead → docker/kubernetes/terraform/ansible/ci-cd/helm/prometheus/
sysadmin). Generalists (assistant) keep skills = [] (see "Open
items" below).
- skills_disabled = true on the four short-conversational agents
(hello-world, recipe-assistant, health-tracker, home-automation).
Their system prompts never instruct the LLM to consult any skill,
so loading all 60 was pure waste. They also drop the explicit
max_history_messages override and inherit the kernel default (60).
- max_history_messages tiered by workload shape:
60 short conversational (hello-world, recipe, health-tracker,
home-automation) — inherits the rising kernel default
(`DEFAULT_MAX_HISTORY_MESSAGES = 60`); no override needed.
60 single-turn task agents (writer, translator, doc-writer,
email-assistant, customer-support, sales-assistant, recruit-
er, social-media, personal-finance, tutor, travel-planner,
meeting-assistant, ops, devops-lead, planner) — explicit
override at the same value to lock the cap if the kernel
default moves again.
80 multi-step / tool-heavy (coder, debugger, architect, code-
reviewer, test-engineer, security-auditor, analyst, data-
scientist, academic-researcher, researcher, legal-assistant)
120 coordinators (assistant, orchestrator) — long multi-agent
sessions where prompt-cache continuity is critical
All values sit at or above the kernel default. Pinning lower
would thrash the prompt cache (the failure mode #91 fixed for
the creator hand by *raising* the cap, not lowering it).
17 hands/*/HAND.toml:
- hand-level mcp_servers / skills now declared on every hand, so
every [agents.*] inside inherits a sensible allowlist.
- skills_disabled = true placed on each [agents.*] inside clip and
creator (pure media pipelines that don't benefit from any skill).
HandDefinitionRaw in librefang-hands does NOT have a top-level
skills_disabled field — declaring it at the hand top level would
be silently dropped by serde, so the setting must live on the
AgentManifest of each sub-agent role.
- devteam: expand existing mcp_servers = ["github"] to include
memory / git / filesystem; populate skills with the expected
dev-team expertise (replacing the placeholder skills = []).
- wiki: replace placeholder mcp_servers = [] with [memory, fetch,
filesystem]. Hand-level skills stays [].
- lead: hand-level skills was originally [email-writer, writing-
coach, interview-prep]; interview-prep is for job-interview
preparation, not lead generation. Replaced with data-analyst
(used by the qualification-scoring step in the prompt).
schema.toml: register mcp_servers / skills / max_history_messages on
the agent field schema so machine consumers (RegistrySchema in
librefang-types) see the new top-level fields. The
max_history_messages description now points at
librefang_runtime::agent_loop::DEFAULT_MAX_HISTORY_MESSAGES (60
today) by name, so the schema doesn't go stale when the constant
moves again.
agents/README.md: example block + "Adding a New Agent" checklist
mention the allowlists; max_history_messages example is shown
commented out with a prompt-cache caveat.
Open items
----------
`assistant` (the default user-facing agent) keeps `skills = []`
deliberately. It is the generalist entry point — capping its skill
surface at a small allowlist would defeat its "delegate to any
specialist" job. The trade-off is that this single agent still pays
the full skill-definition load on every turn; operators who want a
strict allowlist for `assistant` can override it after install.
Why not adopt PR #89's approach
-------------------------------
#89 covers similar ground but with three issues this PR avoids:
1. mcp_servers = ["_none"] sentinel. #89's body explicitly notes
it's pending upstream librefang#4808 (mcp_disabled). Shipping a
magic-string today means coming back later to clean it up. This
PR uses real allowlists.
2. max_history_messages = 8 / 12 / 15 / 20. Far below today's
kernel default (60) and #91's direction for long-workflow hands
(80–120). Every turn that hits the cap invalidates the cached
prompt prefix; the cost of cache misses exceeds the saving from
shorter history. This PR uses 60–120.
3. Doubling max_llm_tokens_per_hour (coder 200k→500k, assistant
300k→500k) widens the per-agent budget — the opposite direction
from #87's "reduce per-call cost" goal. Left to the operator's
instance-specific tuning.
Refs librefang/librefang-registry#87, librefang/librefang-registry#89
* fix(creator): raise max_history_messages to 80 for polling workflows
Creator Hand's async video_generate path polls video_status every 15-20s
until completion (1-3 min typical), consuming ~5-15 turns per video
request. Combined workflows (video + TTS + music) plus normal back-and-
forth cross the kernel default of 40 messages quickly, which surfaced
in user logs as:
WARN run_agent_loop: Trimming old messages at safe turn boundary
agent=creator:creator-hand total_messages=41 trimming=2
INFO run_agent_loop: prompt cache metrics for turn
hit_ratio=0.0 creation=0 read=0
Every turn was hitting the trim cap and invalidating the prompt-cache
prefix. 80 covers ~30 polling iterations plus a comfortable pre-context
window without runaway memory growth. Other hands keep the default 40.
* ci(refresh-cache): open PR instead of pushing directly to main
Branch protection on `main` started rejecting the workflow's auto-commit
with GH006 "Changes must be made through a pull request" — see run
25632824585 on 2026-05-10 against commit 6785807 (the first push that
hit the tightened protection). Direct push is precisely what the file's
own security comment (#1) warns against ("Compromised maintainer pushes
a malicious plugins-index.json directly to main. Mitigation: GitHub
branch protection on main requires PR review"), so the fix preserves
that gate rather than working around it.
The workflow now creates a short-lived `automation/refresh-indexes-<sha>`
branch, commits the regen there, pushes, and opens a PR back to main
via `gh pr create`. Maintainers see a one-click squash-merge.
Permissions: add `pull-requests: write` to the existing `contents: write`
so `gh pr create` can be authorised through the default GITHUB_TOKEN.
The post-merge run on the index PR is a no-op (no diff under
`hands/**`, `plugins/**`, etc. between consecutive states), so no
`[skip ci]` marker is needed and no loop is possible.
Without this fix, every content PR landing on main leaves
plugins-index.json + registry-index.json stale, blocking new agents and
hands from reaching daemons until a maintainer manually regenerates.
* fix(hands): raise max_history_messages on long-workflow coordinators
Three hand coordinators have workflows that routinely exceed the kernel
default history cap on a single user turn:
- researcher (max_iterations=80) — deep web_search → web_fetch →
summarize loops with multi-source synthesis. 80 iterations × ~4
messages each → 200+ messages per user turn. Set to 120.
- devops (max_iterations=60) — incident response and CI/CD fan out
into long shell_exec chains (logs, retries, post-mortems). Set to 80.
- predictor (max_iterations=60) — long reasoning chains accumulating
signals across many web/knowledge queries, with scheduled re-checks
referring back. Set to 80.
Creator's existing override is rephrased "raise above the kernel
default" so the comment stays correct regardless of the order this PR
and the upstream kernel-default bump (librefang side) land in.
Other hands (lead/linkedin/reddit/clip/analytics/apitester/browser/
collector/strategist) stay on the kernel default; the upstream bump
covers them.
Refs librefang/librefang#4842 — long-term replacement for the substring
match that the OpenAI driver currently uses to decide how to handle
`reasoning_content` on historical assistant turns.
Three provider-specific behaviours that the driver must distinguish at
wire time, now expressed as catalog metadata:
* `strip` — DeepSeek R1 / deepseek-reasoner. The API rejects requests
that carry reasoning_content on previous assistant messages.
* `echo` — DeepSeek V4 Flash. Thinking mode is on by default and the
API rejects multi-turn requests when assistant turns containing
tool_calls don't echo back the original reasoning text. This is the
bug surfaced in librefang/librefang#4842.
* `empty_string` — Moonshot / Kimi K2 family. The field must be present
(empty string) on tool_calls turns, with thinking disabled wire-side
for multi-turn compatibility.
* `none` (default) — most providers; field is omitted entirely.
V4 Pro is intentionally NOT marked `echo` — librefang#4842 reports it
working out-of-the-box; flip when there's an empirical reproducer.
Marks affected models:
providers/deepseek.toml
deepseek-v4-flash → echo
deepseek-reasoner → strip
providers/moonshot.toml
kimi-k2.6, kimi-k2.5, kimi-k2 → empty_string
providers/kimi-coding.toml
kimi-for-coding → empty_string
providers/byteplus-coding.toml
kimi-k2.5 → empty_string
providers/novita.toml
moonshotai/kimi-k2-thinking → empty_string
Tooling:
* schema.toml registers the field with the four enum options and a
`none` default so existing TOML files keep parsing unchanged.
* scripts/validate.py rejects unknown enum values; verified with a
hand-crafted negative case (`reasoning_echo_policy = "bogus"` →
validation fails with the expected message).
* `python3 scripts/validate.py` passes (267 models).
The librefang side that consumes this field will land in a follow-up
PR — until then, registry consumers ignore the field via
`#[serde(default)]` and the existing substring fallback continues to
work, so this commit is safe to ship independently.
GitHub flagged the previous owner as "unknown" — CODEOWNERS only
honors entries that resolve to a user with repo write access.
@houko is the actual repo owner.
Round-3 PR re-review follow-ups:
LOW — CODEOWNERS missed /scripts/build-plugins-index.mjs and
/wrangler.toml. Both can change the bytes that get signed without
touching the sign step. Replaced the per-file enumeration with /scripts/
catch-all and added wrangler.toml.
HIGH — workflow comment said the post-sign verify step "catches a buggy
or tampered sign-script run". The "tampered" claim was wrong: any
attacker who can edit sign-plugins-index.mjs in a PR can edit the
verify step (and the embedded pubkey) in the same diff. Re-stated as
"catches accidental regressions only — CODEOWNERS is what stops
adversarial edits".
Branch protection on main has been enabled separately via gh API
(force-push + deletion blocked, PR review required, CODEOWNERS
enforced; bypass for repo owner and github-actions[bot] so the
auto-publish workflow keeps working).
The first iteration moved signing from the Cloudflare worker into this
repo's CI to fix the worker-as-sign-oracle defect. The re-review pointed
out that this just relocated the trust problem: anyone who lands a
commit on main gets the resulting bytes signed automatically. Mitigations:
* CODEOWNERS — sign-script, build scripts, the workflow itself, and
the auto-generated index/sig files are owned by the registry owner.
Combined with branch protection requiring CODEOWNERS review, a PR
touching the signing infrastructure or the artefacts it produces
cannot land without explicit owner sign-off. Plugin contributions
under plugins/<name>/ are covered by the standard PR-review rules
but don't trip CODEOWNERS unless they touch the signing path.
* SHA-pinned actions — actions/checkout@v4 and actions/setup-node@v4
replaced with full-SHA refs (v4.2.2 / v4.1.0). Blocks
action-supply-chain swaps (a malicious mutable tag move on a
popular action would otherwise execute in the same job that has
REGISTRY_PRIVATE_KEY in scope).
* Post-sign self-verify — a new step verifies plugins-index.json.sig
against the committed pubkey (not a secret) BEFORE the .sig hits
main. Catches a buggy sign-script run, an env-leak that produces
zero-bytes output, or a half-applied edit.
* REGISTRY_PRIVATE_KEY env scope — explicitly noted in workflow
comment that the secret is on the sign step ONLY (where it was
already), not the job. Prevents future contributors lifting it
to job-level out of convenience.
* In-workflow security model docstring — enumerates the residual
threat model (compromised maintainer, malicious PR, sign-step
bug, leaked refresh token) and what each defense addresses.
Trust root is now: GitHub branch protection on main + CODEOWNERS on
sign infrastructure + SHA-pinned actions + post-sign verification.
A maintainer with push rights can still ship malicious bytes through
a code-review bypass — that residual is the same as for any signed
package registry and falls outside what CI alone can mitigate.
Branch protection on `main` MUST be configured by an org admin to
match the assumptions in CODEOWNERS:
- require pull request reviews (at least 1)
- require review from CODEOWNERS
- dismiss stale approvals on new commits
- restrict who can push directly to main (org admins only)
Pair with the worker-side simplification on the librefang PR — the
worker is now a pure transport (no key material, no signing) and this
repo's CI takes over signature production.
scripts/sign-plugins-index.mjs reads REGISTRY_PRIVATE_KEY from a GitHub
Actions secret, signs plugins-index.json with Ed25519, and writes
plugins-index.json.sig alongside it. Aborts loudly when the secret is
missing so a misconfigured CI can't silently ship an unsigned payload.
The workflow now runs build → sign → commit (.json + .sig) → push →
poke worker /refresh. The worker fetches the committed .json + .sig
verbatim and stores both — the daemon then verifies against the
embedded pubkey it ships with.
Closes PR review CRITICAL #1: the worker is no longer a sign-anything
oracle reachable via REGISTRY_REFRESH_TOKEN. Trust root is now this
repo's branch protection + Actions secret scope, not a token any CI
job that can talk to stats.librefang.ai can use to mint signatures.
Note: the keypair was rotated as part of this change (PR not yet
merged so no daemon TOFU pins exist). New pubkey:
ClGa0Ucap8NdrKAy1rw9Tt6A9I8eg4zJ53+xIuKMuq0=
The plugins-index.json.sig committed here is signed with the matching
new private key, in lockstep with the daemon EMBEDDED_REGISTRY_PUBKEY
constant and all three worker [vars] entries.
Pair with the existing plugins-index.json path (which feeds the
daemon's signed install lane). registry-index.json mirrors the
dict-shaped payload the registry-worker's cron currently builds via
40+ GitHub Contents API calls — but built locally from the checked-out
tree by scripts/build-registry-index.mjs, so the worker only fetches
ONE file per category type (2 total: plugins + registry) on refresh.
Workflow now picks up content changes across all 8 category dirs
(was: plugins/ only) so dashboard updates land within seconds of a
push instead of waiting for the 02:00 UTC cron tick.
The dashboard's /api/registry endpoint reads kv_store('registry_data');
the worker's forced-refresh now writes that key with these bytes and
purges the Cache-API entry, so the next dashboard hit sees fresh
data instead of the 1h-cached previous payload.
Generated counts on first build: 11p 17h 32a 60s 44c 57pr 22w 33mcp.
Walking ~40+ plugin TOMLs via the GitHub Contents API from the worker
exceeded the Workers Free 50-subrequest-per-invocation limit, leaving
the daemon's signed plugins index either empty or partial after every
forced refresh.
Move the walk into the repo: scripts/build-plugins-index.mjs reads each
plugins/<name>/plugin.toml directly from the checked-out tree and emits
a sorted flat array (name, version?, description?, needs?) at
plugins-index.json. The CI workflow regenerates and commits this file
on every push under plugins/, then pokes the worker's
/api/registry/refresh — which now fetches the single committed
plugins-index.json (1 subrequest), validates the JSON shape, and
re-signs it with Ed25519. Refresh cost is now constant in registry
size, not linear.
The dashboard's dict-shaped /api/registry payload is unchanged — that
still rebuilds via the daily 02:00 UTC cron.
Without this, dashboard / daemon see content changes only after the
next 02:00 UTC cron tick (up to ~24h delay). The action POSTs to
stats.librefang.ai/api/registry/refresh with a bearer token shared
with the worker secret of the same name; until both secrets exist on
their respective sides, the endpoint returns 503 and the action fails
loud — no silent half-deploy.
Filters on directories the worker actually reads (plugins/, agents/,
skills/, hands/, channels/, providers/, workflows/, mcp/) so README
edits don't burn worker invocations.
Known limitation: the worker rebuild walks ~40+ GitHub Contents API
subrequests, which on Workers Free truncates partway through and
leaves some categories empty. This action fires the right path; the
underlying budget fix (Workers Paid, or pre-building the index in
this repo) is tracked separately.
The daemon's `manifest_missing_integrity_hooks` check (#3804) hard-fails
any registry install whose plugin.toml declares hooks but lacks an
[integrity] entry for each one. All 11 plugins under plugins/ had the
[hooks] table but were missing [integrity], so `librefang plugin install
<name>` would error after download with "missing [integrity] hashes for
hook script(s): ...".
Compute and pin SHA-256 over each `hooks/*.py` under every plugin dir.
The mempalace-indexer entries were already present and are unchanged
(reordered alphabetically by the regenerator).
Verified each plugin.toml still parses (Python tomllib).
The `librefang` dashboard's federated catalog UI surfaces every
optional SKILL.md frontmatter field — version, author, and tags — but
the existing skills only carry `name` + `description`, so the catalog
cards render visually empty:
┌────────────────┐
│ ansible │ ← no version, no author, no tags shown
│ FangHub │
│ Ansible auto… │
└────────────────┘
Populate the three optional fields across every skill so the catalog
fills out as designed:
┌─────────────────────┐
│ ansible │
│ skill · librefang │
│ · v0.1.0 │
│ Ansible auto… │
│ [devops][automation]│
│ [infra] │
└─────────────────────┘
Choices
- author = `librefang`. Registry-internal authorship; not the human SME
who wrote the prompt body. Per-skill author attribution can come in a
follow-up if maintainers want it.
- version = `0.1.0` baseline. Future content updates bump per-skill.
- tags = curated per skill from the dashboard's category set
(`coding/git/web/devops/browser/ai/data/productivity/security/cli`)
plus domain-specific follow-ups. First tag is the primary category.
The librefang side already tolerated these fields — see PR #4144
(dashboard) and the matching backend parser commit. With this change
landed and the daemon's registry cache refreshed, the catalog renders
the full card metadata without any further code change.
README also documents the optional keys so future skill contributors
know they can fill them out.
The previous "BytePlus ModelArk Coding Plan" overflows in dashboard
provider cards and CLI lists. Drop the redundant "ModelArk" — it's
already implied by the parent product line, and the file's header
comment + docs explain the relationship.
`zai` (api.z.ai, the international front-end) and `zhipu` (open.bigmodel.cn,
the China front-end) target the same Zhipu account system but are surfaced
as distinct providers in the dashboard. They previously both declared
`api_key_env = "ZHIPU_API_KEY"`, so configuring a single Zhipu credential
silently activated both — same root cause as the cross-product collisions
fixed in librefang-registry#82 and librefang-registry#83.
Rename `zai` to use its own `ZAI_API_KEY`. `zhipu` keeps `ZHIPU_API_KEY`
as the more-established name.
This is a breaking change: existing users authenticating `zai` via
`ZHIPU_API_KEY` need to also export `ZAI_API_KEY` (same key value works).
The companion daemon-side change will be filed against librefang/librefang.
Refs librefang/librefang#3282
Both `microsoft` (GitHub Models / Azure AI Inference at
models.inference.ai.azure.com) and `github-copilot` (the IDE subscription
product) previously declared `api_key_env = "GITHUB_TOKEN"`, so a single
PAT silently activated both providers — users with only the IDE
subscription saw GitHub Models entries appear in their model picker
without intent, and vice versa.
Rename `microsoft` to `GITHUB_MODELS_TOKEN`. `github-copilot` keeps
`GITHUB_TOKEN` as the established convention for the IDE side.
Existing users who set `GITHUB_TOKEN` to use the GitHub Models endpoint
will see the `microsoft` provider become unavailable until they set the
new env var. The librefang daemon will be updated separately to
recognize `GITHUB_TOKEN` as a deprecated fallback for `microsoft` during
a migration window, mirroring the legacy-fallback infrastructure added
in librefang/librefang#3279.
Refs librefang/librefang#3278
Each `_coding` provider entry previously declared the same `api_key_env`
as its standard counterpart, causing one credential to silently activate
two providers. Users who only signed up for the per-token API endpoint
saw the Coding Plan endpoint's models mixed into their model picker
without intent.
Rename so each pair uses an independent env var:
byteplus-coding BYTEPLUS_API_KEY -> BYTEPLUS_CODING_API_KEY
volcengine-coding VOLCENGINE_API_KEY -> VOLCENGINE_CODING_API_KEY
zai-coding ZHIPU_API_KEY -> ZAI_CODING_API_KEY
zhipu-coding ZHIPU_API_KEY -> ZHIPU_CODING_API_KEY
This is a breaking change: existing users must set the new env vars.
The librefang daemon will be updated separately to recognize the old
env vars as a deprecated fallback during a migration window. See
librefang/librefang#3278 for the design discussion and rollout plan.
Out of scope (different problem class, tracked in the same issue):
- zai/zhipu cross-region key sharing (both still share ZHIPU_API_KEY)
- github-copilot/microsoft cross-product GITHUB_TOKEN reuse
Refs librefang/librefang#3278
* docs(byteplus): warn that `byteplus` is per-token, not Coding Plan
Self-followup on PR #78. The BytePlus official docs explicitly warn:
Do not use the standard model endpoint
(https://ark.ap-southeast.bytepluses.com/api/v3) [for Coding Plan
workloads], as requests there bypass Coding Plan quota and incur
separate charges.
— https://docs.byteplus.com/en/docs/ModelArk/1928261
Without this warning visible to users, anyone with a Coding Plan
subscription who picks `provider = "byteplus"` will silently bill
against their per-token USD balance instead of consuming the
subscription quota they paid for.
Adds a prominent "BILLING — READ BEFORE USE" block to byteplus.toml
making the per-token vs subscription split unambiguous, and a
companion note in byteplus-coding.toml pointing back so the choice is
discoverable from either side. Also calls out which capabilities are
unique to the standard endpoint (image / video) so the choice isn't
"just use Coding Plan for everything".
No model definitions or pricing changed.
* chore(byteplus): drop superseded model entries (21 → 11) (#81)
The original byteplus.toml from PR #78 enumerated every BytePlus
ModelArk endpoint that returned HTTP 200, regardless of whether a
sane user would still pick it. Trims to the current per-family
flagship plus useful fast/preview tiers.
Removed (10):
Text:
seed-2-0-lite-260228 — superseded by seed-2-0-mini in the
fast/cheap niche
seed-1-8-251228 — superseded by seed-2-0 family
seed-translation-250915 — too narrow; chat models cover this
deepseek-v3-1-250821 — superseded by deepseek-v3-2
Image:
seedream-3-0-t2i-250415 — superseded by seedream-4-5 / 5-0-lite
seedream-4-0-250828 — same
Video:
seedance-1-0-lite-i2v-250428 — superseded by 1-5 / dreamina-2-0
seedance-1-0-lite-t2v-250428 — same
seedance-1-0-pro-250528 — same
seedance-1-0-pro-fast-251015 — same
Kept (11):
Text (6): seed-2-0-pro/mini/code-preview, glm-4-7,
deepseek-v3-2, gpt-oss-120b
Image (2): seedream-4-5, seedream-5-0 (lite)
Video (3): seedance-1-5-pro, dreamina-seedance-2-0,
dreamina-seedance-2-0-fast
Also corrects two per-piece / per-K price comments that the trim
left orphaned over the wrong [[models]] block (4-5 = $0.0400, not
$0.0300; 1-5-pro = $0.0024/$0.0012 with/without audio, not
$0.0018/K) and updates the header capability list to match the
new model set.
Video and music modality models are billed per generation rather than
per token. Without this field the librefang runtime metering layer
records every call as $0 and emits a warning. Sourced from
https://platform.minimax.io/docs/guides/pricing-paygo.
- Hailuo 2.3 Fast: $0.33 per call (1080P/6s upper bound)
- Hailuo 2.3: $0.56 per call (1080P/6s upper bound)
- Hailuo 02: $0.56 per call (1080P/6s upper bound)
- Music 2.6: $0.15 per up-to-5-minute track
- Lyrics gen: $0.01 per song
Also adds per_call_cost to schema.toml so the field is documented
alongside the other cost fields.
Adds two providers for BytePlus ModelArk, the international edition of
Volcano Engine, distinct from the existing cn-only `volcengine` provider:
- `byteplus`: standard `/api/v3` endpoint, 21 models total
- Text (10): Seed 2.0 Pro/Mini/Lite/Code, Seed 1.8, Seed Translation,
GLM-4.7, DeepSeek V3.2/V3.1, GPT-OSS-120B
- Image (4): Seedream 3.0/4.0/4.5/5.0-lite (per-piece pricing in comments)
- Video (7): Seedance 1.0/1.5 family + Dreamina Seedance 2.0/fast
- `byteplus_coding`: `/api/coding` Anthropic-compatible endpoint with
9 friendly aliases (ark-code-latest auto-router, bytedance-seed-code,
dola-seed-2.0-{pro,code,lite}, kimi-k2.5, glm-4.7, glm-5.1, gpt-oss-120b).
Both use BYTEPLUS_API_KEY env var. All listed model IDs were verified
to return HTTP 200 against the live ap-southeast endpoint. Pricing is the
standard real-time tier from the BytePlus console (a discounted batch tier
exists at roughly half the rate, not modeled here).
Excluded for follow-up (schema doesn't currently support these modalities):
- Skylark embedding-vision (no `embedding` modality in schema)
- Hyper3D-Gen2, Hitem3D-2.0 (no `3d` modality in schema)
Extend modality enum to support video and music, then register the
non-text MiniMax models that were already declared in
media_capabilities but had no concrete entries:
- image-01 ($0.0035/image)
- speech-2.8/2.6 hd & turbo ($60-$100 per 1M chars)
- Hailuo 2.3 Fast / 2.3 / 02 video models ($0.10-$0.56 per video)
- music-2.6, lyrics_generation
Per-call pricing is documented in inline comments since the schema's
token-based cost fields don't naturally fit per-call billing.
schema.toml and scripts/validate.py both updated; the change is
additive (existing modality values remain valid).
* chore: prune deprecated models across providers
Remove old-generation models that are strictly superseded by current versions
on the same provider/family. Affected providers: anthropic, bedrock, vertex-ai,
xai, moonshot, zhipu, baichuan, stepfun, volcengine, minimax, cohere, together,
fireworks, deepinfra, openrouter. Also clean up orphan aliases (grok3, grok-mini,
minimax-m2.1) and remap moonshot alias to kimi-k2.5.
Net: -42 model entries across 18 files. Provider model counts and README rows
updated accordingly.
* chore: remove redundant and orphan aliases from aliases.toml
Provider TOML files auto-register their model.aliases at load time, so
re-declaring them globally is duplication. Also drop entries pointing to
models that no longer exist after the prune.
- 45 redundant entries duplicating provider-defined aliases
- 11 orphan targets (gpt-4o, gpt-4o-mini, grok-2-mini, grok-3,
mixtral-8x7b-32768, copilot/gpt-4, open-mistral-nemo,
pixtral-large-latest, jamba-1.5-large, palmyra-x5, venice-uncensored)
Net: 100 lines down to 21. The file is now what the header comment
always claimed it was: 'additional global aliases not tied to a specific
model entry.'
* chore: second pass — prune more deprecated models
Apply the same 'strictly superseded by same-provider/family successor'
rule to providers missed in the first pass:
- openai: gpt-4.1 / -mini / -nano, o3, o4-mini (5)
- meta-llama: llama-3.3-70b-instruct (1)
- zhipu: glm-4v-plus (1)
- together: Llama-3.3-70B-Instruct-Turbo (1)
- xiaomi: mimo-v2-flash / -omni / -pro (3)
- aion-labs: aion-1.0 / -mini (2)
- qianfan: ernie-speed-128k, ernie-4.0-turbo-8k (2)
- cerebras: cerebras/llama3.1-8b (1)
- qwen-code: qwen-code/qwq-32b (1)
- nvidia-nim: 12 models (llama-3.1/3.2 series, mixtral-8x22b,
mistral-small-3.1, phi-4-mini, qwq-32b, r1-distill-32b,
qwen2.5-coder, nemotron-mini-4b, nemotron-70b-instruct)
- openrouter: meta-llama/llama-3.3-70b (paid), rekaai/reka-edge (2)
- alibaba-coding-plan: qwen3.5-plus, qwen3-max-2026-01-23,
MiniMax-M2.5, kimi-k2.5 (4)
Net: -35 model entries. providers/README.md model counts updated.
220 models remain.
Adds the web_search tool to every agent.toml and hand HAND.toml that did
not already declare it. Without this capability the runtime gates the
tool with 'Capability denied: tool not in allowed list', leaving agents
unable to perform web searches even when a search provider is
configured.
For tools arrays that already contained web_fetch, web_search is
inserted directly after it (its natural companion). For arrays without
web_fetch, web_search is appended to the end.
24 files updated total: 21 agents and 3 hands.
On dual-stack hosts (notably macOS), `localhost` resolves to both ::1
and 127.0.0.1 with IPv6 tried first. Local LLM servers (Ollama, vLLM,
LM Studio) installed via the standard scripts bind IPv4 only, so the
IPv6 connection attempt fails immediately and Happy Eyeballs fallback
to IPv4 isn't reliably triggered for connection-refused errors,
producing spurious "Configured local provider offline" warnings in
the daemon even when the server is up and reachable via curl.
Companion to librefang/librefang#3112 which fixes the hardcoded URL
constants in the main repo. After both land, existing installs pick
up the fix on their next registry sync.
OpenAI-compatible LLM gateway. Pairs with the librefang-llm-drivers
registration (librefang PR #3076) so Novita models surface in the
dashboard model picker without each user having to add them via
/api/models/custom.
Pricing and limits sourced from GET /openai/v1/models on 2026-04-25
(input_token_price_per_m / output_token_price_per_m, divided by 10000
to get USD per million tokens). Curated to a small popular subset:
- deepseek/deepseek-v3.2 (frontier)
- moonshotai/kimi-k2-thinking (frontier, supports_thinking)
- minimax/minimax-m2 (smart)
- meta-llama/llama-3.3-70b-instruct (smart)
- qwen/qwen3-coder-30b-a3b-instruct (balanced)
- zai-org/glm-4.7-flash (balanced)
Introduces image-generation models as a first-class [[models]] entry via
a new `modality` field on the model schema ("text" default, "image",
"audio"). When modality != "text", context_window / max_output_tokens
are optional since no conventional context gate exists — OpenAI's
gpt-image-2 docs omit them.
Adds `image_input_cost_per_m` / `image_output_cost_per_m` alongside
existing text token cost fields to cover the 4-price structure OpenAI
uses for image generation (text $5/$10, image $8/$30 per 1M tokens).
Validator updated to:
- accept any modality in {text, image, audio}
- require context_window/max_output_tokens only for modality=text
- range-check the two new cost fields
gpt-image-2 entry added to providers/openai.toml with pricing sourced
from https://developers.openai.com/api/docs/pricing. Snapshot
gpt-image-2-2026-04-21 listed as alias.
* fix(providers): remove ~anthropic, skip ~ prefixes in sync script
OpenRouter uses ~ prefixes for internal auto-routing aliases (e.g. ~anthropic).
These are not real providers — they already route through openrouter.toml.
The generated ~anthropic.toml was confusing (looked like a stale backup)
and redundant with the existing openrouter provider.
- Delete providers/~anthropic.toml
- Skip provider IDs starting with ~ in sync-pricing.py --create-missing
* fix(providers): remove morph, aider, kwaipilot
- morph: specialized code-editing/patching tool, not a general LLM provider
- aider: CLI meta-tool wrapper (base_url empty), redundant with claude-code/codex-cli/gemini-cli/qwen-code
- kwaipilot: Kwai internal coding assistant routed via OpenRouter, niche
* fix(sync): add morph/aider/kwaipilot to SKIP_PROVIDERS to prevent re-creation
* feat(sync): merge OpenRouter-only providers into openrouter.toml
Instead of generating standalone .toml files that just wrap the OpenRouter
endpoint, merge their models directly into openrouter.toml with the
standard 'openrouter/{provider}/{model}' ID convention.
- Add _build_model_fields() and _model_lines() helpers to deduplicate
model rendering between standalone and merged paths
- Add merge_into_openrouter() that appends new models idempotently
- generate_provider_toml() now only runs for providers in PROVIDER_API
- --create-missing routes OpenRouter-only providers to merge_into_openrouter
* fix(providers): remove 14 OpenRouter-only standalone files
These providers have no direct public API and all route through
openrouter.ai/api/v1. Per the new sync-pricing.py policy, their models
will be merged into openrouter.toml on the next CI run instead of
living in separate files that just wrap the OpenRouter endpoint.
Removed: allenai, deepcogito, essentialai, inclusionai, inflection,
liquid, meituan, nex-agi, nousresearch, prime-intellect, relace,
switchpoint, tngtech, writer
* fix(providers): remove 7 niche providers with no driver support
No dedicated LLM driver code exists for these providers — they rely
purely on OpenAI-compatible passthrough with no special handling.
Removing them reduces registry noise; users can still reach them via
openrouter.toml if needed.
Removed: microsoft, ibm-granite, xiaomi, upstage, inception, aion-labs, arcee-ai
* fix(providers): remove ai21, chutes, venice
All three use ApiFormat::OpenAI with no special handling — pure passthrough.
No registry entry needed; users can reach them via openrouter.toml or by
adding a custom provider.
* docs(providers): rewrite README with full provider catalog and inclusion criteria
- List all 46 providers grouped by category with descriptions
- Document why each provider exists (direct API, unique endpoint, dedicated driver, local, CLI)
- Add inclusion criteria section explaining when to create standalone files vs merging into openrouter.toml
- Document sync script routing logic
- Update model counts: 49→46 providers, 339→232 models
* docs: add comprehensive READMEs for all registry sections + deepinfra provider
- agents/README.md: 32 agents across 7 categories with capability field reference
- channels/README.md: 44 channels across 5 categories with protocol reference table
- hands/README.md: 18 hands across 5 categories with HAND.toml format guide
- mcp/README.md: 33 MCP servers across 5 categories with transport/auth format
- plugins/README.md: 12 plugins with hook protocol documentation
- skills/README.md: 60 skills across 9 categories with SKILL.md format guide
- providers/deepinfra.toml: add DeepInfra serverless inference (5 models)
These 11 templates are the general-purpose ones that don't depend on any
particular model at the primary level (`[model] provider = "default"`).
They also ship a secondary `[[fallback_models]]` block pointing at
`gemini-2.0-flash` with `api_key_env = "GEMINI_API_KEY"`.
That default hurts everyone who doesn't happen to have `$GEMINI_API_KEY`
set — every agent boot logs `WARN Fallback driver 'gemini' failed to
init: Missing API key`, once per turn per agent. The templates that
actually intend to use Gemini as their primary model (analyst, coder,
researcher, code-reviewer, debugger, legal-assistant, data-scientist,
academic-researcher, test-engineer) are left untouched — those
declare Gemini in `[model]`, which is an intentional design choice, not
a hidden fallback.
Users who want a Gemini fallback chain for generic agents can add
`[[fallback_models]]` themselves in `~/.librefang/workspaces/agents/...`
once they've set `$GEMINI_API_KEY`.
Removed from:
assistant, customer-support, devops-lead, doc-writer, email-assistant,
meeting-assistant, planner, recruiter, sales-assistant, social-media,
writer
Add `claude-opus-4-7`, Anthropic's current flagship model (per
https://platform.claude.com/docs/en/docs/about-claude/models/overview).
- Context window: 1,000,000 tokens (Opus 4.7 ships with a new tokenizer)
- Max output: 128,000 tokens
- Pricing: $5 / input MTok, $25 / output MTok (unchanged from 4.6)
- Tier: frontier
- Supports tools, vision, streaming, and adaptive thinking
Move the `opus` / `claude-opus` aliases from 4.6 to 4.7 so a user asking
for "opus" gets the current flagship. Opus 4.6 is now in Anthropic's
"Legacy" section of the models overview; keeping it in the registry
entry (for existing callers that pin the exact ID) but without the
generic aliases.
Also **correct Opus 4.6's `context_window`**: it was listed as 200,000
tokens but the official model page has shown 1M tokens since release.
That was a pre-existing bug this PR fixes in passing since it directly
affects anyone who'd have routed queries to 4.6 expecting 1M.
Add a header comment pinning the source URL and explaining the
"latest-first" ordering convention so future additions don't silently
rebind aliases to a previous-generation snapshot.
Add the two GPT-5.4 variants exposed by OpenAI's API (source:
https://developers.openai.com/api/docs/models/gpt-5.4 and
https://developers.openai.com/api/docs/models/gpt-5.4-mini):
- **gpt-5.4** — frontier tier, 1,050,000 context window, 128k max output,
$2.50 / input MTok, $15.00 / output MTok.
- **gpt-5.4-mini** — balanced tier (matching the naming convention used
by `gpt-5-mini`, `gpt-4.1-mini`, etc.), 400,000 context window,
128k max output, $0.75 / input MTok, $4.50 / output MTok.
Ordered right after the GPT-5.2 family and before the Codex variants
section to keep the frontier-GPT chain in release order.
Addresses librefang/librefang-registry#65 — the original request filed
these under the `codex-cli` provider, but Codex CLI's upstream
`models.json` doesn't list `gpt-5.4-mini` (only the full `gpt-5.4`
slug is list-visible there, and it's already tracked in
`providers/codex-cli.toml`). The correct home for OpenAI-API-direct
access is this file; users who want `gpt-5.4-mini` should configure
`provider = "openai"` rather than `provider = "codex-cli"`.
* refactor: migrate icon fields from emoji to lucide:<name> tokens
Every TOML manifest's `icon = "<emoji>"` line is replaced with
`icon = "lucide:<kebab-name>"` — a reference to a lucide-react icon,
which the librefang.ai site and dashboard render as crisp SVG. Reasons
for the switch:
- Emoji render very differently across OS/browser/font stacks; the
registry catalog looked inconsistent from one row to the next.
- Five manifests (clip / creator / linkedin / reddit / twitter) had
their icons stored as literal Python-style escape strings
("\\U0001F3AC") because the TOML parser upstream never decoded
them. Switching away from emoji drops that class of bug entirely.
- As a drive-by, also decode the \\uXXXX accent escapes in the
[i18n.fr] block of hands/creator/HAND.toml so "Créateur" shows
up correctly.
87 files touched. example manifests left untouched (still "TODO").
* fix: backfill i18n name + drop the single-member email category
- Every existing [i18n.<lang>] block now has a `name` field. 60 files
previously translated description but kept the English name
implicitly — which rendered as "some English some Chinese" in the
registry UI. Fill in the missing name from the English brand (or a
known localized equivalent: DingTalk→钉钉, Feishu→飞书, Email→
电子邮件 / メール / E-Mail / Correo / Courriel, and a handful of
hands that have Chinese product names like 视频剪辑 Hand).
- channels/email.toml was the only item under category="email";
reclassify it as "messaging" so the sub-category filter chip list
on the category page isn't littered with singletons.
* feat(i18n): localize 76 agents/integrations/plugins into 7 languages
Adds full [i18n.zh], [i18n.zh-TW], [i18n.ja], [i18n.ko], [i18n.de],
[i18n.es], [i18n.fr] blocks with name + description to every manifest
that previously shipped English-only.
Coverage:
- 32 agents (academic-researcher, analyst, architect, assistant,
code-reviewer, coder, customer-support, data-scientist, debugger,
devops-lead, doc-writer, email-assistant, health-tracker,
hello-world, home-automation, legal-assistant, meeting-assistant,
ops, orchestrator, personal-finance, planner, recipe-assistant,
recruiter, researcher, sales-assistant, security-auditor,
social-media, test-engineer, translator, travel-planner, tutor,
writer)
- 33 integrations (AWS, Azure, Bitbucket, Brave Search, Discord,
Dropbox, Elasticsearch, Exa Search, Fetch, Filesystem, GCP, Git,
GitHub, GitLab, Gmail, Google Calendar, Google Drive, Google Maps,
Jira, Linear, Memory, MongoDB, Notion, PostgreSQL, Puppeteer, Redis,
Sentry, Sequential Thinking, Slack, SQLite, Teams, Time, Todoist) —
brand names kept as-is across all locales, only descriptions
translated.
- 11 plugins (auto-summarizer, context-decay, conversation-logger,
episodic-memory, guardrails, keyword-memory, mempalace-indexer,
sentiment-tracker, todo-tracker, topic-memory, user-profile)
The descriptions are one-line summaries — hand-translated rather than
machine-generated, so technical terms (MCP, PR, CI/CD, etc.) stay
consistent across locales.
* feat(i18n): close remaining per-lang gaps for channels, workflows, devteam
Third pass on i18n coverage. Every non-example manifest now carries a
full set of [i18n.zh], [i18n.zh-TW], [i18n.ja], [i18n.ko], [i18n.de],
[i18n.es], [i18n.fr] blocks.
- 44 channel adapters: added French descriptions (zh/zh-TW/ja/ko/de/es
were already present). Brand names kept as-is in all locales so users
recognize Discord / Slack / LINE / etc. consistently.
- 22 workflows: filled zh-TW / ja / ko / de / es / fr blocks. Each
translation mirrors the existing zh one in structure and tone so the
catalog reads consistently across locales.
- hands/devteam/HAND.toml: added the four langs that were missing
(zh-TW, de, es, fr).
Only the 6 templates under examples/ are left without i18n blocks on
purpose — they still contain "TODO:" placeholders.
The librefang MCP security check blocks shell interpreters (sh, bash)
as commands. Use `command = "npx"` with `$HOME` in args — the runtime
now expands env vars in args natively.