`zai` (api.z.ai, the international front-end) and `zhipu` (open.bigmodel.cn,
the China front-end) target the same Zhipu account system but are surfaced
as distinct providers in the dashboard. They previously both declared
`api_key_env = "ZHIPU_API_KEY"`, so configuring a single Zhipu credential
silently activated both — same root cause as the cross-product collisions
fixed in librefang-registry#82 and librefang-registry#83.
Rename `zai` to use its own `ZAI_API_KEY`. `zhipu` keeps `ZHIPU_API_KEY`
as the more-established name.
This is a breaking change: existing users authenticating `zai` via
`ZHIPU_API_KEY` need to also export `ZAI_API_KEY` (same key value works).
The companion daemon-side change will be filed against librefang/librefang.
Refs librefang/librefang#3282
Both `microsoft` (GitHub Models / Azure AI Inference at
models.inference.ai.azure.com) and `github-copilot` (the IDE subscription
product) previously declared `api_key_env = "GITHUB_TOKEN"`, so a single
PAT silently activated both providers — users with only the IDE
subscription saw GitHub Models entries appear in their model picker
without intent, and vice versa.
Rename `microsoft` to `GITHUB_MODELS_TOKEN`. `github-copilot` keeps
`GITHUB_TOKEN` as the established convention for the IDE side.
Existing users who set `GITHUB_TOKEN` to use the GitHub Models endpoint
will see the `microsoft` provider become unavailable until they set the
new env var. The librefang daemon will be updated separately to
recognize `GITHUB_TOKEN` as a deprecated fallback for `microsoft` during
a migration window, mirroring the legacy-fallback infrastructure added
in librefang/librefang#3279.
Refs librefang/librefang#3278
Each `_coding` provider entry previously declared the same `api_key_env`
as its standard counterpart, causing one credential to silently activate
two providers. Users who only signed up for the per-token API endpoint
saw the Coding Plan endpoint's models mixed into their model picker
without intent.
Rename so each pair uses an independent env var:
byteplus-coding BYTEPLUS_API_KEY -> BYTEPLUS_CODING_API_KEY
volcengine-coding VOLCENGINE_API_KEY -> VOLCENGINE_CODING_API_KEY
zai-coding ZHIPU_API_KEY -> ZAI_CODING_API_KEY
zhipu-coding ZHIPU_API_KEY -> ZHIPU_CODING_API_KEY
This is a breaking change: existing users must set the new env vars.
The librefang daemon will be updated separately to recognize the old
env vars as a deprecated fallback during a migration window. See
librefang/librefang#3278 for the design discussion and rollout plan.
Out of scope (different problem class, tracked in the same issue):
- zai/zhipu cross-region key sharing (both still share ZHIPU_API_KEY)
- github-copilot/microsoft cross-product GITHUB_TOKEN reuse
Refs librefang/librefang#3278
* docs(byteplus): warn that `byteplus` is per-token, not Coding Plan
Self-followup on PR #78. The BytePlus official docs explicitly warn:
Do not use the standard model endpoint
(https://ark.ap-southeast.bytepluses.com/api/v3) [for Coding Plan
workloads], as requests there bypass Coding Plan quota and incur
separate charges.
— https://docs.byteplus.com/en/docs/ModelArk/1928261
Without this warning visible to users, anyone with a Coding Plan
subscription who picks `provider = "byteplus"` will silently bill
against their per-token USD balance instead of consuming the
subscription quota they paid for.
Adds a prominent "BILLING — READ BEFORE USE" block to byteplus.toml
making the per-token vs subscription split unambiguous, and a
companion note in byteplus-coding.toml pointing back so the choice is
discoverable from either side. Also calls out which capabilities are
unique to the standard endpoint (image / video) so the choice isn't
"just use Coding Plan for everything".
No model definitions or pricing changed.
* chore(byteplus): drop superseded model entries (21 → 11) (#81)
The original byteplus.toml from PR #78 enumerated every BytePlus
ModelArk endpoint that returned HTTP 200, regardless of whether a
sane user would still pick it. Trims to the current per-family
flagship plus useful fast/preview tiers.
Removed (10):
Text:
seed-2-0-lite-260228 — superseded by seed-2-0-mini in the
fast/cheap niche
seed-1-8-251228 — superseded by seed-2-0 family
seed-translation-250915 — too narrow; chat models cover this
deepseek-v3-1-250821 — superseded by deepseek-v3-2
Image:
seedream-3-0-t2i-250415 — superseded by seedream-4-5 / 5-0-lite
seedream-4-0-250828 — same
Video:
seedance-1-0-lite-i2v-250428 — superseded by 1-5 / dreamina-2-0
seedance-1-0-lite-t2v-250428 — same
seedance-1-0-pro-250528 — same
seedance-1-0-pro-fast-251015 — same
Kept (11):
Text (6): seed-2-0-pro/mini/code-preview, glm-4-7,
deepseek-v3-2, gpt-oss-120b
Image (2): seedream-4-5, seedream-5-0 (lite)
Video (3): seedance-1-5-pro, dreamina-seedance-2-0,
dreamina-seedance-2-0-fast
Also corrects two per-piece / per-K price comments that the trim
left orphaned over the wrong [[models]] block (4-5 = $0.0400, not
$0.0300; 1-5-pro = $0.0024/$0.0012 with/without audio, not
$0.0018/K) and updates the header capability list to match the
new model set.
Video and music modality models are billed per generation rather than
per token. Without this field the librefang runtime metering layer
records every call as $0 and emits a warning. Sourced from
https://platform.minimax.io/docs/guides/pricing-paygo.
- Hailuo 2.3 Fast: $0.33 per call (1080P/6s upper bound)
- Hailuo 2.3: $0.56 per call (1080P/6s upper bound)
- Hailuo 02: $0.56 per call (1080P/6s upper bound)
- Music 2.6: $0.15 per up-to-5-minute track
- Lyrics gen: $0.01 per song
Also adds per_call_cost to schema.toml so the field is documented
alongside the other cost fields.
Adds two providers for BytePlus ModelArk, the international edition of
Volcano Engine, distinct from the existing cn-only `volcengine` provider:
- `byteplus`: standard `/api/v3` endpoint, 21 models total
- Text (10): Seed 2.0 Pro/Mini/Lite/Code, Seed 1.8, Seed Translation,
GLM-4.7, DeepSeek V3.2/V3.1, GPT-OSS-120B
- Image (4): Seedream 3.0/4.0/4.5/5.0-lite (per-piece pricing in comments)
- Video (7): Seedance 1.0/1.5 family + Dreamina Seedance 2.0/fast
- `byteplus_coding`: `/api/coding` Anthropic-compatible endpoint with
9 friendly aliases (ark-code-latest auto-router, bytedance-seed-code,
dola-seed-2.0-{pro,code,lite}, kimi-k2.5, glm-4.7, glm-5.1, gpt-oss-120b).
Both use BYTEPLUS_API_KEY env var. All listed model IDs were verified
to return HTTP 200 against the live ap-southeast endpoint. Pricing is the
standard real-time tier from the BytePlus console (a discounted batch tier
exists at roughly half the rate, not modeled here).
Excluded for follow-up (schema doesn't currently support these modalities):
- Skylark embedding-vision (no `embedding` modality in schema)
- Hyper3D-Gen2, Hitem3D-2.0 (no `3d` modality in schema)
Extend modality enum to support video and music, then register the
non-text MiniMax models that were already declared in
media_capabilities but had no concrete entries:
- image-01 ($0.0035/image)
- speech-2.8/2.6 hd & turbo ($60-$100 per 1M chars)
- Hailuo 2.3 Fast / 2.3 / 02 video models ($0.10-$0.56 per video)
- music-2.6, lyrics_generation
Per-call pricing is documented in inline comments since the schema's
token-based cost fields don't naturally fit per-call billing.
schema.toml and scripts/validate.py both updated; the change is
additive (existing modality values remain valid).
* chore: prune deprecated models across providers
Remove old-generation models that are strictly superseded by current versions
on the same provider/family. Affected providers: anthropic, bedrock, vertex-ai,
xai, moonshot, zhipu, baichuan, stepfun, volcengine, minimax, cohere, together,
fireworks, deepinfra, openrouter. Also clean up orphan aliases (grok3, grok-mini,
minimax-m2.1) and remap moonshot alias to kimi-k2.5.
Net: -42 model entries across 18 files. Provider model counts and README rows
updated accordingly.
* chore: remove redundant and orphan aliases from aliases.toml
Provider TOML files auto-register their model.aliases at load time, so
re-declaring them globally is duplication. Also drop entries pointing to
models that no longer exist after the prune.
- 45 redundant entries duplicating provider-defined aliases
- 11 orphan targets (gpt-4o, gpt-4o-mini, grok-2-mini, grok-3,
mixtral-8x7b-32768, copilot/gpt-4, open-mistral-nemo,
pixtral-large-latest, jamba-1.5-large, palmyra-x5, venice-uncensored)
Net: 100 lines down to 21. The file is now what the header comment
always claimed it was: 'additional global aliases not tied to a specific
model entry.'
* chore: second pass — prune more deprecated models
Apply the same 'strictly superseded by same-provider/family successor'
rule to providers missed in the first pass:
- openai: gpt-4.1 / -mini / -nano, o3, o4-mini (5)
- meta-llama: llama-3.3-70b-instruct (1)
- zhipu: glm-4v-plus (1)
- together: Llama-3.3-70B-Instruct-Turbo (1)
- xiaomi: mimo-v2-flash / -omni / -pro (3)
- aion-labs: aion-1.0 / -mini (2)
- qianfan: ernie-speed-128k, ernie-4.0-turbo-8k (2)
- cerebras: cerebras/llama3.1-8b (1)
- qwen-code: qwen-code/qwq-32b (1)
- nvidia-nim: 12 models (llama-3.1/3.2 series, mixtral-8x22b,
mistral-small-3.1, phi-4-mini, qwq-32b, r1-distill-32b,
qwen2.5-coder, nemotron-mini-4b, nemotron-70b-instruct)
- openrouter: meta-llama/llama-3.3-70b (paid), rekaai/reka-edge (2)
- alibaba-coding-plan: qwen3.5-plus, qwen3-max-2026-01-23,
MiniMax-M2.5, kimi-k2.5 (4)
Net: -35 model entries. providers/README.md model counts updated.
220 models remain.
Adds the web_search tool to every agent.toml and hand HAND.toml that did
not already declare it. Without this capability the runtime gates the
tool with 'Capability denied: tool not in allowed list', leaving agents
unable to perform web searches even when a search provider is
configured.
For tools arrays that already contained web_fetch, web_search is
inserted directly after it (its natural companion). For arrays without
web_fetch, web_search is appended to the end.
24 files updated total: 21 agents and 3 hands.
On dual-stack hosts (notably macOS), `localhost` resolves to both ::1
and 127.0.0.1 with IPv6 tried first. Local LLM servers (Ollama, vLLM,
LM Studio) installed via the standard scripts bind IPv4 only, so the
IPv6 connection attempt fails immediately and Happy Eyeballs fallback
to IPv4 isn't reliably triggered for connection-refused errors,
producing spurious "Configured local provider offline" warnings in
the daemon even when the server is up and reachable via curl.
Companion to librefang/librefang#3112 which fixes the hardcoded URL
constants in the main repo. After both land, existing installs pick
up the fix on their next registry sync.
OpenAI-compatible LLM gateway. Pairs with the librefang-llm-drivers
registration (librefang PR #3076) so Novita models surface in the
dashboard model picker without each user having to add them via
/api/models/custom.
Pricing and limits sourced from GET /openai/v1/models on 2026-04-25
(input_token_price_per_m / output_token_price_per_m, divided by 10000
to get USD per million tokens). Curated to a small popular subset:
- deepseek/deepseek-v3.2 (frontier)
- moonshotai/kimi-k2-thinking (frontier, supports_thinking)
- minimax/minimax-m2 (smart)
- meta-llama/llama-3.3-70b-instruct (smart)
- qwen/qwen3-coder-30b-a3b-instruct (balanced)
- zai-org/glm-4.7-flash (balanced)
Introduces image-generation models as a first-class [[models]] entry via
a new `modality` field on the model schema ("text" default, "image",
"audio"). When modality != "text", context_window / max_output_tokens
are optional since no conventional context gate exists — OpenAI's
gpt-image-2 docs omit them.
Adds `image_input_cost_per_m` / `image_output_cost_per_m` alongside
existing text token cost fields to cover the 4-price structure OpenAI
uses for image generation (text $5/$10, image $8/$30 per 1M tokens).
Validator updated to:
- accept any modality in {text, image, audio}
- require context_window/max_output_tokens only for modality=text
- range-check the two new cost fields
gpt-image-2 entry added to providers/openai.toml with pricing sourced
from https://developers.openai.com/api/docs/pricing. Snapshot
gpt-image-2-2026-04-21 listed as alias.
* fix(providers): remove ~anthropic, skip ~ prefixes in sync script
OpenRouter uses ~ prefixes for internal auto-routing aliases (e.g. ~anthropic).
These are not real providers — they already route through openrouter.toml.
The generated ~anthropic.toml was confusing (looked like a stale backup)
and redundant with the existing openrouter provider.
- Delete providers/~anthropic.toml
- Skip provider IDs starting with ~ in sync-pricing.py --create-missing
* fix(providers): remove morph, aider, kwaipilot
- morph: specialized code-editing/patching tool, not a general LLM provider
- aider: CLI meta-tool wrapper (base_url empty), redundant with claude-code/codex-cli/gemini-cli/qwen-code
- kwaipilot: Kwai internal coding assistant routed via OpenRouter, niche
* fix(sync): add morph/aider/kwaipilot to SKIP_PROVIDERS to prevent re-creation
* feat(sync): merge OpenRouter-only providers into openrouter.toml
Instead of generating standalone .toml files that just wrap the OpenRouter
endpoint, merge their models directly into openrouter.toml with the
standard 'openrouter/{provider}/{model}' ID convention.
- Add _build_model_fields() and _model_lines() helpers to deduplicate
model rendering between standalone and merged paths
- Add merge_into_openrouter() that appends new models idempotently
- generate_provider_toml() now only runs for providers in PROVIDER_API
- --create-missing routes OpenRouter-only providers to merge_into_openrouter
* fix(providers): remove 14 OpenRouter-only standalone files
These providers have no direct public API and all route through
openrouter.ai/api/v1. Per the new sync-pricing.py policy, their models
will be merged into openrouter.toml on the next CI run instead of
living in separate files that just wrap the OpenRouter endpoint.
Removed: allenai, deepcogito, essentialai, inclusionai, inflection,
liquid, meituan, nex-agi, nousresearch, prime-intellect, relace,
switchpoint, tngtech, writer
* fix(providers): remove 7 niche providers with no driver support
No dedicated LLM driver code exists for these providers — they rely
purely on OpenAI-compatible passthrough with no special handling.
Removing them reduces registry noise; users can still reach them via
openrouter.toml if needed.
Removed: microsoft, ibm-granite, xiaomi, upstage, inception, aion-labs, arcee-ai
* fix(providers): remove ai21, chutes, venice
All three use ApiFormat::OpenAI with no special handling — pure passthrough.
No registry entry needed; users can reach them via openrouter.toml or by
adding a custom provider.
* docs(providers): rewrite README with full provider catalog and inclusion criteria
- List all 46 providers grouped by category with descriptions
- Document why each provider exists (direct API, unique endpoint, dedicated driver, local, CLI)
- Add inclusion criteria section explaining when to create standalone files vs merging into openrouter.toml
- Document sync script routing logic
- Update model counts: 49→46 providers, 339→232 models
* docs: add comprehensive READMEs for all registry sections + deepinfra provider
- agents/README.md: 32 agents across 7 categories with capability field reference
- channels/README.md: 44 channels across 5 categories with protocol reference table
- hands/README.md: 18 hands across 5 categories with HAND.toml format guide
- mcp/README.md: 33 MCP servers across 5 categories with transport/auth format
- plugins/README.md: 12 plugins with hook protocol documentation
- skills/README.md: 60 skills across 9 categories with SKILL.md format guide
- providers/deepinfra.toml: add DeepInfra serverless inference (5 models)
These 11 templates are the general-purpose ones that don't depend on any
particular model at the primary level (`[model] provider = "default"`).
They also ship a secondary `[[fallback_models]]` block pointing at
`gemini-2.0-flash` with `api_key_env = "GEMINI_API_KEY"`.
That default hurts everyone who doesn't happen to have `$GEMINI_API_KEY`
set — every agent boot logs `WARN Fallback driver 'gemini' failed to
init: Missing API key`, once per turn per agent. The templates that
actually intend to use Gemini as their primary model (analyst, coder,
researcher, code-reviewer, debugger, legal-assistant, data-scientist,
academic-researcher, test-engineer) are left untouched — those
declare Gemini in `[model]`, which is an intentional design choice, not
a hidden fallback.
Users who want a Gemini fallback chain for generic agents can add
`[[fallback_models]]` themselves in `~/.librefang/workspaces/agents/...`
once they've set `$GEMINI_API_KEY`.
Removed from:
assistant, customer-support, devops-lead, doc-writer, email-assistant,
meeting-assistant, planner, recruiter, sales-assistant, social-media,
writer
Add `claude-opus-4-7`, Anthropic's current flagship model (per
https://platform.claude.com/docs/en/docs/about-claude/models/overview).
- Context window: 1,000,000 tokens (Opus 4.7 ships with a new tokenizer)
- Max output: 128,000 tokens
- Pricing: $5 / input MTok, $25 / output MTok (unchanged from 4.6)
- Tier: frontier
- Supports tools, vision, streaming, and adaptive thinking
Move the `opus` / `claude-opus` aliases from 4.6 to 4.7 so a user asking
for "opus" gets the current flagship. Opus 4.6 is now in Anthropic's
"Legacy" section of the models overview; keeping it in the registry
entry (for existing callers that pin the exact ID) but without the
generic aliases.
Also **correct Opus 4.6's `context_window`**: it was listed as 200,000
tokens but the official model page has shown 1M tokens since release.
That was a pre-existing bug this PR fixes in passing since it directly
affects anyone who'd have routed queries to 4.6 expecting 1M.
Add a header comment pinning the source URL and explaining the
"latest-first" ordering convention so future additions don't silently
rebind aliases to a previous-generation snapshot.
Add the two GPT-5.4 variants exposed by OpenAI's API (source:
https://developers.openai.com/api/docs/models/gpt-5.4 and
https://developers.openai.com/api/docs/models/gpt-5.4-mini):
- **gpt-5.4** — frontier tier, 1,050,000 context window, 128k max output,
$2.50 / input MTok, $15.00 / output MTok.
- **gpt-5.4-mini** — balanced tier (matching the naming convention used
by `gpt-5-mini`, `gpt-4.1-mini`, etc.), 400,000 context window,
128k max output, $0.75 / input MTok, $4.50 / output MTok.
Ordered right after the GPT-5.2 family and before the Codex variants
section to keep the frontier-GPT chain in release order.
Addresses librefang/librefang-registry#65 — the original request filed
these under the `codex-cli` provider, but Codex CLI's upstream
`models.json` doesn't list `gpt-5.4-mini` (only the full `gpt-5.4`
slug is list-visible there, and it's already tracked in
`providers/codex-cli.toml`). The correct home for OpenAI-API-direct
access is this file; users who want `gpt-5.4-mini` should configure
`provider = "openai"` rather than `provider = "codex-cli"`.
* refactor: migrate icon fields from emoji to lucide:<name> tokens
Every TOML manifest's `icon = "<emoji>"` line is replaced with
`icon = "lucide:<kebab-name>"` — a reference to a lucide-react icon,
which the librefang.ai site and dashboard render as crisp SVG. Reasons
for the switch:
- Emoji render very differently across OS/browser/font stacks; the
registry catalog looked inconsistent from one row to the next.
- Five manifests (clip / creator / linkedin / reddit / twitter) had
their icons stored as literal Python-style escape strings
("\\U0001F3AC") because the TOML parser upstream never decoded
them. Switching away from emoji drops that class of bug entirely.
- As a drive-by, also decode the \\uXXXX accent escapes in the
[i18n.fr] block of hands/creator/HAND.toml so "Créateur" shows
up correctly.
87 files touched. example manifests left untouched (still "TODO").
* fix: backfill i18n name + drop the single-member email category
- Every existing [i18n.<lang>] block now has a `name` field. 60 files
previously translated description but kept the English name
implicitly — which rendered as "some English some Chinese" in the
registry UI. Fill in the missing name from the English brand (or a
known localized equivalent: DingTalk→钉钉, Feishu→飞书, Email→
电子邮件 / メール / E-Mail / Correo / Courriel, and a handful of
hands that have Chinese product names like 视频剪辑 Hand).
- channels/email.toml was the only item under category="email";
reclassify it as "messaging" so the sub-category filter chip list
on the category page isn't littered with singletons.
* feat(i18n): localize 76 agents/integrations/plugins into 7 languages
Adds full [i18n.zh], [i18n.zh-TW], [i18n.ja], [i18n.ko], [i18n.de],
[i18n.es], [i18n.fr] blocks with name + description to every manifest
that previously shipped English-only.
Coverage:
- 32 agents (academic-researcher, analyst, architect, assistant,
code-reviewer, coder, customer-support, data-scientist, debugger,
devops-lead, doc-writer, email-assistant, health-tracker,
hello-world, home-automation, legal-assistant, meeting-assistant,
ops, orchestrator, personal-finance, planner, recipe-assistant,
recruiter, researcher, sales-assistant, security-auditor,
social-media, test-engineer, translator, travel-planner, tutor,
writer)
- 33 integrations (AWS, Azure, Bitbucket, Brave Search, Discord,
Dropbox, Elasticsearch, Exa Search, Fetch, Filesystem, GCP, Git,
GitHub, GitLab, Gmail, Google Calendar, Google Drive, Google Maps,
Jira, Linear, Memory, MongoDB, Notion, PostgreSQL, Puppeteer, Redis,
Sentry, Sequential Thinking, Slack, SQLite, Teams, Time, Todoist) —
brand names kept as-is across all locales, only descriptions
translated.
- 11 plugins (auto-summarizer, context-decay, conversation-logger,
episodic-memory, guardrails, keyword-memory, mempalace-indexer,
sentiment-tracker, todo-tracker, topic-memory, user-profile)
The descriptions are one-line summaries — hand-translated rather than
machine-generated, so technical terms (MCP, PR, CI/CD, etc.) stay
consistent across locales.
* feat(i18n): close remaining per-lang gaps for channels, workflows, devteam
Third pass on i18n coverage. Every non-example manifest now carries a
full set of [i18n.zh], [i18n.zh-TW], [i18n.ja], [i18n.ko], [i18n.de],
[i18n.es], [i18n.fr] blocks.
- 44 channel adapters: added French descriptions (zh/zh-TW/ja/ko/de/es
were already present). Brand names kept as-is in all locales so users
recognize Discord / Slack / LINE / etc. consistently.
- 22 workflows: filled zh-TW / ja / ko / de / es / fr blocks. Each
translation mirrors the existing zh one in structure and tone so the
catalog reads consistently across locales.
- hands/devteam/HAND.toml: added the four langs that were missing
(zh-TW, de, es, fr).
Only the 6 templates under examples/ are left without i18n blocks on
purpose — they still contain "TODO:" placeholders.
The librefang MCP security check blocks shell interpreters (sh, bash)
as commands. Use `command = "npx"` with `$HOME` in args — the runtime
now expands env vars in args natively.
The provider creation form only collected api_key_env (the env var
name) but not the actual key value. New providers were always created
as "unconfigured" because no key was stored.
Add an optional secret api_key field so the dashboard can pass the
key value during creation. The backend strips it from the TOML and
saves it to secrets.env instead.
- Move example skills to examples/skills/
- Move scaffolding templates to examples/{agents,hands,channels,plugins,providers,integrations,skills}/
- Remove legacy skill.toml from example skills; SKILL.md is the canonical format
- Delete top-level templates/ directory
The prior entries (o4-mini, o3, gpt-4.1) no longer appear in
upstream openai/codex's codex-rs/models-manager/models.json and the
Codex CLI actively migrates users away from the old gpt-5 /
gpt-5-codex slugs via tui/src/model_migration.rs.
Replace with the currently-visible ("visibility": "list") slugs:
- gpt-5.4
- gpt-5.3-codex
- gpt-5.2
- gpt-5.2-codex
- gpt-5.1-codex-max
- gpt-5.1-codex-mini
All six share a 272k context window per upstream models.json.
Marking supports_tools / supports_vision / supports_streaming = true
to match the gpt-5-family capability envelope.
Closeslibrefang/librefang#2347
Squashed replay of the original 9-commit branch onto current main. The
original branch was 30+ commits behind, forked from before the skills
refactor (PR #42) and workflow template expansion (PR #36), so a
standard rebase hit heavy add/add conflicts on workflows/*.toml that
are unrelated to the wiki hand.
This replay keeps only the final hands/wiki/ tree state, which is the
actual intent of the PR (the author iterated several times on the same
files; squashing matches that).
Implements the "LLM Wiki" pattern (Andrej Karpathy) for building a
personal, Obsidian-compatible knowledge base. Instead of on-the-fly
RAG, the wiki hand incrementally maintains a Markdown vault:
hands/wiki/
├── HAND.toml # hand manifest + [agents.*] sections
├── README.md # user-facing docs
├── SKILL-main.md # Librarian (coordinator) routing + FS ops
├── SKILL-ingestor.md # Source extraction + [[wikilink]] writing
├── SKILL-analyst.md # Synthesis with provenance citations
└── SKILL-linter.md # Broken link / orphan / contradiction audit
Closeslibrefang/librefang-registry#44 (via replay, not merge).
Standardize on Claude Code's SKILL.md format as every skill's source of
truth. skill.toml becomes an optional metadata layer for runtime, input
schema, and versioning — never for the prompt body.
- validate.py: require SKILL.md in every skill dir; when skill.toml also
exists, cross-check name/description consistency to prevent drift
- Add SKILL.md to the two custom-skill examples
- Move the meeting-agenda prompt body out of skill.toml into SKILL.md
- Rewrite skills/README.md to document the md-first, toml-as-metadata convention
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Claude Code-style skills use SKILL.md with YAML frontmatter instead of
skill.toml. Validator now accepts either form, unblocking the 60 bundled
skills restored in #42.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(hands): add devteam hand -- autonomous software development team
Multi-agent hand with 7 roles (PM, Architect, Frontend, Backend, DevOps, QA, Designer)
and 3 team size tiers (simple/standard/full) for different project scales.
PM coordinator auto-scans GitHub issues, triages, assigns tasks to specialists,
and tracks progress on an in-memory project board.
* refactor(hands): slim devteam to 3 agents (PM + Engineer + QA)
7 agents with serial agent_send = massive token waste and info loss at every
handoff. Merge architect/frontend/backend/devops into one Engineer with full
context. Keep QA separate for independent verification. Drop designer.
Tiers: lite (PM + Engineer) and standard (PM + Engineer + QA).
* fix(hands/devteam): fix workspace isolation and git workflow gaps
- PM uses GitHub API for code browsing, no repo clone needed
- Engineer explicitly clones repo, branches, commits, pushes, creates PR
- QA explicitly clones repo, checks out branch under review
- PM tracks last_scan timestamp to filter already-triaged issues
- approval_mode now means PR stays open for review, not skip commit
* fix(hands/devteam): use shared repo checkout instead of per-agent clones
All 3 agents share one checkout at ../shared/repo/. Engineer clones it
on the first task; PM and QA read from the same path. Eliminates
duplicate clones and cross-workspace visibility issues.
* fix(hands/devteam): read issue comments before triaging
Comments contain clarifications, reproduction steps, duplicate markers,
and resolution status. Also skip already-assigned and wontfix issues.
* fix(hands/devteam): fix interactive git add, add merge/close APIs, add fix iteration flow
- Replace git add -p (interactive) with git add <specific files>
- PM prompt now has explicit merge PR and close issue API calls
- Engineer has explicit fix-request handling (same branch, push, no new PR)
* feat(hands/devteam): add full GitHub interaction -- PR review, issue comments, labels
PM:
- Labels issues during triage, comments triage status
- Scans open PRs for external review requests
- Comments on issues linking merged PRs
Engineer:
- Replies to review comments on PR after fixing
- Reviews external PRs with APPROVE/REQUEST_CHANGES + line comments
QA:
- Leaves PR review (APPROVE or REQUEST_CHANGES with line comments)
- All findings visible on GitHub, not just via agent_send
SKILL.md:
- Added PR diff, reviews, review comments, reply, merge API references
* fix(hands/devteam): enforce English comments, line-level reviews, comment-before-close
- All GitHub comments/reviews must be in English (added global rule)
- PR reviews must use comments[] with path+line, not body-only
- Comment on issue with resolution details BEFORE closing/merging
- Improved comment templates with structured info
* fix(hands/devteam): 8 logic fixes from end-to-end workflow review
1. Filter PRs from Issues API (pull_request key)
2. PM sends PR number to QA for review
3. Deduplicate PR scanning via devteam_reviewed_prs
4. QA reports test gaps instead of pushing code to shared branch
5. branch_strategy wired into Engineer (gitflow branches from develop)
6. approval_mode: ON = wait for human, OFF = auto-merge after QA
7. scan_interval mapped to schedule_create every_secs
8. git checkout -B instead of -b to handle existing branches
* fix(hands/devteam): second-pass review — 6 more logic fixes
1. Engineer extracts PR number from create-PR API response
2. PM falls back to GitHub Contents API when shared repo not yet cloned
3. QA gets external PR review flow (was only on Engineer)
4. PM checks CI status + mergeable before merging
5. PM handles merge conflict (409) by sending back to Engineer to rebase
6. i18n approval_mode description synced with actual semantics
* fix(hands/devteam): third-pass — runtime scenarios
1. Deduplicate cron schedule on daemon restart (check schedule_list first)
2. Max 3 review rounds before escalating to user (prevent infinite loop)
3. Clean working directory before switching tasks (git checkout -- . && git clean)
4. Add user direct commands (work on #42, status, review PR #50)
5. Pass tech_stack to Engineer in task delegation
6. Fix duplicate step numbering in Review Cycle
* fix(hands/devteam): fourth-pass — state consistency and edge cases
1. QA force-syncs to remote branch (git checkout -B origin/branch) for force-push safety
2. Board sync step: reconcile with GitHub each scan cycle (catch external closes/merges)
3. Prune devteam_reviewed_prs of closed PRs, cap done list at 30
4. PM checks CI before sending to QA (don't waste QA on red builds)
5. Stop/cancel command: remove from board, comment on issue
6. Explicit rebase commands for Engineer (fetch + rebase + force-with-lease)
* fix(hands/devteam): fifth-pass — crash prevention
1. Guard empty repo_url: stop and tell user to configure it
2. Add python3 to requires (all JSON parsing depends on it)
3. Engineer git config user.name/email on first clone (prevents commit rejection)
4. Explicit build/lint/test commands per tech stack (Rust/TS/Python/Go/Java/Swift)
5. event_publish on task completion so user gets notified
6. Global rule: check API HTTP status before parsing JSON
* feat(hands/devteam): add gh CLI / MCP / curl API three-layer fallback
- Add GitHub MCP integration (mcp_servers = ["github"])
- Add gh CLI as optional requirement (preferred over curl)
- All 3 agents: gh > MCP > curl priority for GitHub operations
- Add issue_tracker setting (github/linear/jira)
- Add agent_list to shared tools
- SKILL.md: add full gh CLI reference section
- i18n: add issue_tracker translation
* feat(hands/devteam): full MCP/integration/notification layer
MCP allowlist: github, linear, jira, sentry, slack, discord
- Sentry: Engineer reads crash reports/stack traces when fixing bugs
- Slack/Discord: PM posts status updates (triaged, completed, QA results)
- Linear/Jira: alternative issue trackers
New settings: notify_channel (none/slack/discord), issue_tracker (github/linear/jira)
New optional requires: npx (MCP runtime), SENTRY_AUTH_TOKEN
PM prompt: notification section, channel-aware status posting
Engineer prompt: Sentry context lookup for bug fixes
i18n: added translations for new settings
* feat(hands/devteam): workflows, onboarding, knowledge, standup, rollback
Workflows (8 integrated):
- PM: bug-triage, product-spec, weekly-report, incident-postmortem
- Engineer: code-review, test-generation, refactor-plan, api-design
- QA: code-review, test-generation
New capabilities:
- Repo onboarding: first activation analyzes repo structure/stack/CI
- Knowledge accumulation: store lessons per issue, detect module hotspots
- Daily standup: cron schedule, board summary via notify_channel
- Rollback: gh pr revert + postmortem workflow + re-open issue
Also:
- Added workflow_run to tools, skills = [] (all allowed)
- Rewrote README with full architecture, lifecycle, workflow table
- PM prompt now has 15 sections covering full lifecycle
* feat(hands/devteam): per-agent capabilities, resources, profiles, fallbacks
Each agent now has full AgentManifest config (not just system_prompt):
PM:
- profile: automation
- capabilities: web, memory, schedule, knowledge, event, workflow, agent_send
- shell: gh, curl, cat, python3
- resources: 200k tokens/hr
Engineer:
- profile: coding
- capabilities: file r/w, shell, web, memory, knowledge, workflow
- shell: cargo, npm, python, go, swift, mvn, git, gh, docker, make
- resources: 300k tokens/hr, 10 concurrent tools
- network: * (needs to push to GitHub)
QA:
- profile: coding (read-heavy, no file_write)
- capabilities: file read, shell (test/lint commands only), web, workflow
- shell: cargo test/clippy/audit, npm test, pytest, go test, gh
- resources: 150k tokens/hr
All agents have fallback_models configured.
* feat(hands/devteam): rewrite with proper resource composition
First hand to use the new composition features:
Agents:
- PM: base=planner, capabilities restricted to gh/git shell only
- Engineer: base=coder, full shell access, network=*
- QA: base=code-reviewer, tool_blocklist=[file_write], test/lint shells only
Composition:
- base: inherit from agents/planner, agents/coder, agents/code-reviewer
- mcp_servers: github (agents interact via MCP, not curl in prompts)
- workflows: bug-triage, code-review, test-generation via workflow_run tool
- plugins: todo-tracker, auto-summarizer, episodic-memory
- per-agent skills: SKILL-pm.md, SKILL-engineer.md, SKILL-qa.md
- per-agent capabilities: QA can't write files, PM can't run builds
Prompts are clean and focused (role + methodology + principles),
not stuffed with curl commands. GitHub interaction goes through
MCP tools or gh CLI.
* fix(devteam): complete planner methodology in PM prompt
Added SCOPE/SEQUENCE/RISK/MILESTONE keywords from the planner
base template's methodology into the PM's triage workflow.
* docs: update hands README, fix repo_url reference in prompts
- hands/README.md: document full composition model (base, MCP, workflows,
plugins, per-agent skills, per-agent capabilities)
- Updated hand count to 15 (added devteam)
- Engineer prompt: clarify repo_url comes from User Configuration, not
a template variable
- PM prompt: same clarification
* fix(devteam): override name/description from base templates
Without explicit name, agents inherit base names (planner/coder/code-reviewer)
instead of hand-specific names (pm/engineer/qa). This affects display and
the prefixed name used in agent registry (devteam:pm vs devteam:planner).
* style: format HAND.toml with taplo