`zai` (api.z.ai, the international front-end) and `zhipu` (open.bigmodel.cn,
the China front-end) target the same Zhipu account system but are surfaced
as distinct providers in the dashboard. They previously both declared
`api_key_env = "ZHIPU_API_KEY"`, so configuring a single Zhipu credential
silently activated both — same root cause as the cross-product collisions
fixed in librefang-registry#82 and librefang-registry#83.
Rename `zai` to use its own `ZAI_API_KEY`. `zhipu` keeps `ZHIPU_API_KEY`
as the more-established name.
This is a breaking change: existing users authenticating `zai` via
`ZHIPU_API_KEY` need to also export `ZAI_API_KEY` (same key value works).
The companion daemon-side change will be filed against librefang/librefang.
Refs librefang/librefang#3282
Both `microsoft` (GitHub Models / Azure AI Inference at
models.inference.ai.azure.com) and `github-copilot` (the IDE subscription
product) previously declared `api_key_env = "GITHUB_TOKEN"`, so a single
PAT silently activated both providers — users with only the IDE
subscription saw GitHub Models entries appear in their model picker
without intent, and vice versa.
Rename `microsoft` to `GITHUB_MODELS_TOKEN`. `github-copilot` keeps
`GITHUB_TOKEN` as the established convention for the IDE side.
Existing users who set `GITHUB_TOKEN` to use the GitHub Models endpoint
will see the `microsoft` provider become unavailable until they set the
new env var. The librefang daemon will be updated separately to
recognize `GITHUB_TOKEN` as a deprecated fallback for `microsoft` during
a migration window, mirroring the legacy-fallback infrastructure added
in librefang/librefang#3279.
Refs librefang/librefang#3278
Each `_coding` provider entry previously declared the same `api_key_env`
as its standard counterpart, causing one credential to silently activate
two providers. Users who only signed up for the per-token API endpoint
saw the Coding Plan endpoint's models mixed into their model picker
without intent.
Rename so each pair uses an independent env var:
byteplus-coding BYTEPLUS_API_KEY -> BYTEPLUS_CODING_API_KEY
volcengine-coding VOLCENGINE_API_KEY -> VOLCENGINE_CODING_API_KEY
zai-coding ZHIPU_API_KEY -> ZAI_CODING_API_KEY
zhipu-coding ZHIPU_API_KEY -> ZHIPU_CODING_API_KEY
This is a breaking change: existing users must set the new env vars.
The librefang daemon will be updated separately to recognize the old
env vars as a deprecated fallback during a migration window. See
librefang/librefang#3278 for the design discussion and rollout plan.
Out of scope (different problem class, tracked in the same issue):
- zai/zhipu cross-region key sharing (both still share ZHIPU_API_KEY)
- github-copilot/microsoft cross-product GITHUB_TOKEN reuse
Refs librefang/librefang#3278
* docs(byteplus): warn that `byteplus` is per-token, not Coding Plan
Self-followup on PR #78. The BytePlus official docs explicitly warn:
Do not use the standard model endpoint
(https://ark.ap-southeast.bytepluses.com/api/v3) [for Coding Plan
workloads], as requests there bypass Coding Plan quota and incur
separate charges.
— https://docs.byteplus.com/en/docs/ModelArk/1928261
Without this warning visible to users, anyone with a Coding Plan
subscription who picks `provider = "byteplus"` will silently bill
against their per-token USD balance instead of consuming the
subscription quota they paid for.
Adds a prominent "BILLING — READ BEFORE USE" block to byteplus.toml
making the per-token vs subscription split unambiguous, and a
companion note in byteplus-coding.toml pointing back so the choice is
discoverable from either side. Also calls out which capabilities are
unique to the standard endpoint (image / video) so the choice isn't
"just use Coding Plan for everything".
No model definitions or pricing changed.
* chore(byteplus): drop superseded model entries (21 → 11) (#81)
The original byteplus.toml from PR #78 enumerated every BytePlus
ModelArk endpoint that returned HTTP 200, regardless of whether a
sane user would still pick it. Trims to the current per-family
flagship plus useful fast/preview tiers.
Removed (10):
Text:
seed-2-0-lite-260228 — superseded by seed-2-0-mini in the
fast/cheap niche
seed-1-8-251228 — superseded by seed-2-0 family
seed-translation-250915 — too narrow; chat models cover this
deepseek-v3-1-250821 — superseded by deepseek-v3-2
Image:
seedream-3-0-t2i-250415 — superseded by seedream-4-5 / 5-0-lite
seedream-4-0-250828 — same
Video:
seedance-1-0-lite-i2v-250428 — superseded by 1-5 / dreamina-2-0
seedance-1-0-lite-t2v-250428 — same
seedance-1-0-pro-250528 — same
seedance-1-0-pro-fast-251015 — same
Kept (11):
Text (6): seed-2-0-pro/mini/code-preview, glm-4-7,
deepseek-v3-2, gpt-oss-120b
Image (2): seedream-4-5, seedream-5-0 (lite)
Video (3): seedance-1-5-pro, dreamina-seedance-2-0,
dreamina-seedance-2-0-fast
Also corrects two per-piece / per-K price comments that the trim
left orphaned over the wrong [[models]] block (4-5 = $0.0400, not
$0.0300; 1-5-pro = $0.0024/$0.0012 with/without audio, not
$0.0018/K) and updates the header capability list to match the
new model set.
Video and music modality models are billed per generation rather than
per token. Without this field the librefang runtime metering layer
records every call as $0 and emits a warning. Sourced from
https://platform.minimax.io/docs/guides/pricing-paygo.
- Hailuo 2.3 Fast: $0.33 per call (1080P/6s upper bound)
- Hailuo 2.3: $0.56 per call (1080P/6s upper bound)
- Hailuo 02: $0.56 per call (1080P/6s upper bound)
- Music 2.6: $0.15 per up-to-5-minute track
- Lyrics gen: $0.01 per song
Also adds per_call_cost to schema.toml so the field is documented
alongside the other cost fields.
Adds two providers for BytePlus ModelArk, the international edition of
Volcano Engine, distinct from the existing cn-only `volcengine` provider:
- `byteplus`: standard `/api/v3` endpoint, 21 models total
- Text (10): Seed 2.0 Pro/Mini/Lite/Code, Seed 1.8, Seed Translation,
GLM-4.7, DeepSeek V3.2/V3.1, GPT-OSS-120B
- Image (4): Seedream 3.0/4.0/4.5/5.0-lite (per-piece pricing in comments)
- Video (7): Seedance 1.0/1.5 family + Dreamina Seedance 2.0/fast
- `byteplus_coding`: `/api/coding` Anthropic-compatible endpoint with
9 friendly aliases (ark-code-latest auto-router, bytedance-seed-code,
dola-seed-2.0-{pro,code,lite}, kimi-k2.5, glm-4.7, glm-5.1, gpt-oss-120b).
Both use BYTEPLUS_API_KEY env var. All listed model IDs were verified
to return HTTP 200 against the live ap-southeast endpoint. Pricing is the
standard real-time tier from the BytePlus console (a discounted batch tier
exists at roughly half the rate, not modeled here).
Excluded for follow-up (schema doesn't currently support these modalities):
- Skylark embedding-vision (no `embedding` modality in schema)
- Hyper3D-Gen2, Hitem3D-2.0 (no `3d` modality in schema)
Extend modality enum to support video and music, then register the
non-text MiniMax models that were already declared in
media_capabilities but had no concrete entries:
- image-01 ($0.0035/image)
- speech-2.8/2.6 hd & turbo ($60-$100 per 1M chars)
- Hailuo 2.3 Fast / 2.3 / 02 video models ($0.10-$0.56 per video)
- music-2.6, lyrics_generation
Per-call pricing is documented in inline comments since the schema's
token-based cost fields don't naturally fit per-call billing.
schema.toml and scripts/validate.py both updated; the change is
additive (existing modality values remain valid).
* chore: prune deprecated models across providers
Remove old-generation models that are strictly superseded by current versions
on the same provider/family. Affected providers: anthropic, bedrock, vertex-ai,
xai, moonshot, zhipu, baichuan, stepfun, volcengine, minimax, cohere, together,
fireworks, deepinfra, openrouter. Also clean up orphan aliases (grok3, grok-mini,
minimax-m2.1) and remap moonshot alias to kimi-k2.5.
Net: -42 model entries across 18 files. Provider model counts and README rows
updated accordingly.
* chore: remove redundant and orphan aliases from aliases.toml
Provider TOML files auto-register their model.aliases at load time, so
re-declaring them globally is duplication. Also drop entries pointing to
models that no longer exist after the prune.
- 45 redundant entries duplicating provider-defined aliases
- 11 orphan targets (gpt-4o, gpt-4o-mini, grok-2-mini, grok-3,
mixtral-8x7b-32768, copilot/gpt-4, open-mistral-nemo,
pixtral-large-latest, jamba-1.5-large, palmyra-x5, venice-uncensored)
Net: 100 lines down to 21. The file is now what the header comment
always claimed it was: 'additional global aliases not tied to a specific
model entry.'
* chore: second pass — prune more deprecated models
Apply the same 'strictly superseded by same-provider/family successor'
rule to providers missed in the first pass:
- openai: gpt-4.1 / -mini / -nano, o3, o4-mini (5)
- meta-llama: llama-3.3-70b-instruct (1)
- zhipu: glm-4v-plus (1)
- together: Llama-3.3-70B-Instruct-Turbo (1)
- xiaomi: mimo-v2-flash / -omni / -pro (3)
- aion-labs: aion-1.0 / -mini (2)
- qianfan: ernie-speed-128k, ernie-4.0-turbo-8k (2)
- cerebras: cerebras/llama3.1-8b (1)
- qwen-code: qwen-code/qwq-32b (1)
- nvidia-nim: 12 models (llama-3.1/3.2 series, mixtral-8x22b,
mistral-small-3.1, phi-4-mini, qwq-32b, r1-distill-32b,
qwen2.5-coder, nemotron-mini-4b, nemotron-70b-instruct)
- openrouter: meta-llama/llama-3.3-70b (paid), rekaai/reka-edge (2)
- alibaba-coding-plan: qwen3.5-plus, qwen3-max-2026-01-23,
MiniMax-M2.5, kimi-k2.5 (4)
Net: -35 model entries. providers/README.md model counts updated.
220 models remain.
On dual-stack hosts (notably macOS), `localhost` resolves to both ::1
and 127.0.0.1 with IPv6 tried first. Local LLM servers (Ollama, vLLM,
LM Studio) installed via the standard scripts bind IPv4 only, so the
IPv6 connection attempt fails immediately and Happy Eyeballs fallback
to IPv4 isn't reliably triggered for connection-refused errors,
producing spurious "Configured local provider offline" warnings in
the daemon even when the server is up and reachable via curl.
Companion to librefang/librefang#3112 which fixes the hardcoded URL
constants in the main repo. After both land, existing installs pick
up the fix on their next registry sync.
OpenAI-compatible LLM gateway. Pairs with the librefang-llm-drivers
registration (librefang PR #3076) so Novita models surface in the
dashboard model picker without each user having to add them via
/api/models/custom.
Pricing and limits sourced from GET /openai/v1/models on 2026-04-25
(input_token_price_per_m / output_token_price_per_m, divided by 10000
to get USD per million tokens). Curated to a small popular subset:
- deepseek/deepseek-v3.2 (frontier)
- moonshotai/kimi-k2-thinking (frontier, supports_thinking)
- minimax/minimax-m2 (smart)
- meta-llama/llama-3.3-70b-instruct (smart)
- qwen/qwen3-coder-30b-a3b-instruct (balanced)
- zai-org/glm-4.7-flash (balanced)
Introduces image-generation models as a first-class [[models]] entry via
a new `modality` field on the model schema ("text" default, "image",
"audio"). When modality != "text", context_window / max_output_tokens
are optional since no conventional context gate exists — OpenAI's
gpt-image-2 docs omit them.
Adds `image_input_cost_per_m` / `image_output_cost_per_m` alongside
existing text token cost fields to cover the 4-price structure OpenAI
uses for image generation (text $5/$10, image $8/$30 per 1M tokens).
Validator updated to:
- accept any modality in {text, image, audio}
- require context_window/max_output_tokens only for modality=text
- range-check the two new cost fields
gpt-image-2 entry added to providers/openai.toml with pricing sourced
from https://developers.openai.com/api/docs/pricing. Snapshot
gpt-image-2-2026-04-21 listed as alias.
* fix(providers): remove ~anthropic, skip ~ prefixes in sync script
OpenRouter uses ~ prefixes for internal auto-routing aliases (e.g. ~anthropic).
These are not real providers — they already route through openrouter.toml.
The generated ~anthropic.toml was confusing (looked like a stale backup)
and redundant with the existing openrouter provider.
- Delete providers/~anthropic.toml
- Skip provider IDs starting with ~ in sync-pricing.py --create-missing
* fix(providers): remove morph, aider, kwaipilot
- morph: specialized code-editing/patching tool, not a general LLM provider
- aider: CLI meta-tool wrapper (base_url empty), redundant with claude-code/codex-cli/gemini-cli/qwen-code
- kwaipilot: Kwai internal coding assistant routed via OpenRouter, niche
* fix(sync): add morph/aider/kwaipilot to SKIP_PROVIDERS to prevent re-creation
* feat(sync): merge OpenRouter-only providers into openrouter.toml
Instead of generating standalone .toml files that just wrap the OpenRouter
endpoint, merge their models directly into openrouter.toml with the
standard 'openrouter/{provider}/{model}' ID convention.
- Add _build_model_fields() and _model_lines() helpers to deduplicate
model rendering between standalone and merged paths
- Add merge_into_openrouter() that appends new models idempotently
- generate_provider_toml() now only runs for providers in PROVIDER_API
- --create-missing routes OpenRouter-only providers to merge_into_openrouter
* fix(providers): remove 14 OpenRouter-only standalone files
These providers have no direct public API and all route through
openrouter.ai/api/v1. Per the new sync-pricing.py policy, their models
will be merged into openrouter.toml on the next CI run instead of
living in separate files that just wrap the OpenRouter endpoint.
Removed: allenai, deepcogito, essentialai, inclusionai, inflection,
liquid, meituan, nex-agi, nousresearch, prime-intellect, relace,
switchpoint, tngtech, writer
* fix(providers): remove 7 niche providers with no driver support
No dedicated LLM driver code exists for these providers — they rely
purely on OpenAI-compatible passthrough with no special handling.
Removing them reduces registry noise; users can still reach them via
openrouter.toml if needed.
Removed: microsoft, ibm-granite, xiaomi, upstage, inception, aion-labs, arcee-ai
* fix(providers): remove ai21, chutes, venice
All three use ApiFormat::OpenAI with no special handling — pure passthrough.
No registry entry needed; users can reach them via openrouter.toml or by
adding a custom provider.
* docs(providers): rewrite README with full provider catalog and inclusion criteria
- List all 46 providers grouped by category with descriptions
- Document why each provider exists (direct API, unique endpoint, dedicated driver, local, CLI)
- Add inclusion criteria section explaining when to create standalone files vs merging into openrouter.toml
- Document sync script routing logic
- Update model counts: 49→46 providers, 339→232 models
* docs: add comprehensive READMEs for all registry sections + deepinfra provider
- agents/README.md: 32 agents across 7 categories with capability field reference
- channels/README.md: 44 channels across 5 categories with protocol reference table
- hands/README.md: 18 hands across 5 categories with HAND.toml format guide
- mcp/README.md: 33 MCP servers across 5 categories with transport/auth format
- plugins/README.md: 12 plugins with hook protocol documentation
- skills/README.md: 60 skills across 9 categories with SKILL.md format guide
- providers/deepinfra.toml: add DeepInfra serverless inference (5 models)
Add `claude-opus-4-7`, Anthropic's current flagship model (per
https://platform.claude.com/docs/en/docs/about-claude/models/overview).
- Context window: 1,000,000 tokens (Opus 4.7 ships with a new tokenizer)
- Max output: 128,000 tokens
- Pricing: $5 / input MTok, $25 / output MTok (unchanged from 4.6)
- Tier: frontier
- Supports tools, vision, streaming, and adaptive thinking
Move the `opus` / `claude-opus` aliases from 4.6 to 4.7 so a user asking
for "opus" gets the current flagship. Opus 4.6 is now in Anthropic's
"Legacy" section of the models overview; keeping it in the registry
entry (for existing callers that pin the exact ID) but without the
generic aliases.
Also **correct Opus 4.6's `context_window`**: it was listed as 200,000
tokens but the official model page has shown 1M tokens since release.
That was a pre-existing bug this PR fixes in passing since it directly
affects anyone who'd have routed queries to 4.6 expecting 1M.
Add a header comment pinning the source URL and explaining the
"latest-first" ordering convention so future additions don't silently
rebind aliases to a previous-generation snapshot.
Add the two GPT-5.4 variants exposed by OpenAI's API (source:
https://developers.openai.com/api/docs/models/gpt-5.4 and
https://developers.openai.com/api/docs/models/gpt-5.4-mini):
- **gpt-5.4** — frontier tier, 1,050,000 context window, 128k max output,
$2.50 / input MTok, $15.00 / output MTok.
- **gpt-5.4-mini** — balanced tier (matching the naming convention used
by `gpt-5-mini`, `gpt-4.1-mini`, etc.), 400,000 context window,
128k max output, $0.75 / input MTok, $4.50 / output MTok.
Ordered right after the GPT-5.2 family and before the Codex variants
section to keep the frontier-GPT chain in release order.
Addresses librefang/librefang-registry#65 — the original request filed
these under the `codex-cli` provider, but Codex CLI's upstream
`models.json` doesn't list `gpt-5.4-mini` (only the full `gpt-5.4`
slug is list-visible there, and it's already tracked in
`providers/codex-cli.toml`). The correct home for OpenAI-API-direct
access is this file; users who want `gpt-5.4-mini` should configure
`provider = "openai"` rather than `provider = "codex-cli"`.
The prior entries (o4-mini, o3, gpt-4.1) no longer appear in
upstream openai/codex's codex-rs/models-manager/models.json and the
Codex CLI actively migrates users away from the old gpt-5 /
gpt-5-codex slugs via tui/src/model_migration.rs.
Replace with the currently-visible ("visibility": "list") slugs:
- gpt-5.4
- gpt-5.3-codex
- gpt-5.2
- gpt-5.2-codex
- gpt-5.1-codex-max
- gpt-5.1-codex-mini
All six share a 272k context window per upstream models.json.
Marking supports_tools / supports_vision / supports_streaming = true
to match the gpt-5-family capability envelope.
Closeslibrefang/librefang#2347
Providers with known public APIs use their official endpoints:
- meta-llama → api.llama.com/v1
- microsoft → models.inference.ai.azure.com (GitHub Models)
- ibm-granite → us-south.ml.cloud.ibm.com/ml/v1 (watsonx)
- tencent → api.hunyuan.cloud.tencent.com/v1
- morph → api.morphllm.com/v1
16 remaining providers without known public APIs route through
OpenRouter (base_url = openrouter.ai/api/v1, OPENROUTER_API_KEY).
sync-pricing.py updated with PROVIDER_API mapping.
Providers without their own public API now use OpenRouter as their
base_url with OPENROUTER_API_KEY, making them testable and usable
when the user has an OpenRouter key configured.
- 20 OpenRouter-only providers: set base_url to openrouter.ai/api/v1
- morph: set correct official API (api.morphllm.com/v1)
- sync-pricing.py: default to OpenRouter routing for new providers
- Merge unique models from duplicate providers into their hand-written
counterparts and remove the duplicates:
- alibaba (tongyi-deepresearch) → qwen
- amazon (nova-2-lite, nova-micro, nova-premier) → bedrock
- bytedance (ui-tars) → volcengine
- nvidia (nemotron-3-nano, nemotron-3-super, etc.) → nvidia-nim
- rekaai (reka-flash-3) → reka
- Set correct official API base_url for providers with public APIs:
arcee-ai, inception, morph, reka, upstage
- Set key_required=false for 20 providers only accessible through
hosting platforms (no public API)
- Update sync-pricing.py with SKIP_DUPLICATES, PROVIDER_API mapping,
and default key_required=false for future auto-generated providers
- Replace tier "free" with "fast" (valid tiers: frontier/smart/balanced/fast/local)
- Remove version suffix from teams-mcp integration id field
- Update sync-pricing.py to not generate invalid tier values
* fix: pin npm package versions in MCP integration templates
Prevent supply chain attacks by pinning exact versions instead of
using unpinned `npx -y @package` which pulls latest on every run.
23 of 25 integrations pinned. sqlite-mcp and aws skipped (packages
not found on npm registry).
* fix: use stable azure/mcp version instead of beta
* feat: add pricing sync script and update model prices from OpenRouter API
- scripts/sync-pricing.py fetches real-time pricing from OpenRouter
- Updated 64 price fields across 13 provider files
- Run periodically or in CI to keep prices current
* feat: add Qwen International and US provider catalogs
* refactor: merge qwen-us into qwen-intl with regions support
* refactor: merge regional providers into single files with regions
Merge qwen-intl.toml into qwen.toml with [provider.regions] for
intl (Singapore) and us (Virginia) endpoints.
Merge minimax-cn.toml into minimax.toml with [provider.regions.china]
including separate api_key_env for China endpoint.
Uses new RegionConfig table format instead of simple string URLs.