* docs(byteplus): warn that `byteplus` is per-token, not Coding Plan
Self-followup on PR #78. The BytePlus official docs explicitly warn:
Do not use the standard model endpoint
(https://ark.ap-southeast.bytepluses.com/api/v3) [for Coding Plan
workloads], as requests there bypass Coding Plan quota and incur
separate charges.
— https://docs.byteplus.com/en/docs/ModelArk/1928261
Without this warning visible to users, anyone with a Coding Plan
subscription who picks `provider = "byteplus"` will silently bill
against their per-token USD balance instead of consuming the
subscription quota they paid for.
Adds a prominent "BILLING — READ BEFORE USE" block to byteplus.toml
making the per-token vs subscription split unambiguous, and a
companion note in byteplus-coding.toml pointing back so the choice is
discoverable from either side. Also calls out which capabilities are
unique to the standard endpoint (image / video) so the choice isn't
"just use Coding Plan for everything".
No model definitions or pricing changed.
* chore(byteplus): drop superseded model entries (21 → 11) (#81)
The original byteplus.toml from PR #78 enumerated every BytePlus
ModelArk endpoint that returned HTTP 200, regardless of whether a
sane user would still pick it. Trims to the current per-family
flagship plus useful fast/preview tiers.
Removed (10):
Text:
seed-2-0-lite-260228 — superseded by seed-2-0-mini in the
fast/cheap niche
seed-1-8-251228 — superseded by seed-2-0 family
seed-translation-250915 — too narrow; chat models cover this
deepseek-v3-1-250821 — superseded by deepseek-v3-2
Image:
seedream-3-0-t2i-250415 — superseded by seedream-4-5 / 5-0-lite
seedream-4-0-250828 — same
Video:
seedance-1-0-lite-i2v-250428 — superseded by 1-5 / dreamina-2-0
seedance-1-0-lite-t2v-250428 — same
seedance-1-0-pro-250528 — same
seedance-1-0-pro-fast-251015 — same
Kept (11):
Text (6): seed-2-0-pro/mini/code-preview, glm-4-7,
deepseek-v3-2, gpt-oss-120b
Image (2): seedream-4-5, seedream-5-0 (lite)
Video (3): seedance-1-5-pro, dreamina-seedance-2-0,
dreamina-seedance-2-0-fast
Also corrects two per-piece / per-K price comments that the trim
left orphaned over the wrong [[models]] block (4-5 = $0.0400, not
$0.0300; 1-5-pro = $0.0024/$0.0012 with/without audio, not
$0.0018/K) and updates the header capability list to match the
new model set.
Video and music modality models are billed per generation rather than
per token. Without this field the librefang runtime metering layer
records every call as $0 and emits a warning. Sourced from
https://platform.minimax.io/docs/guides/pricing-paygo.
- Hailuo 2.3 Fast: $0.33 per call (1080P/6s upper bound)
- Hailuo 2.3: $0.56 per call (1080P/6s upper bound)
- Hailuo 02: $0.56 per call (1080P/6s upper bound)
- Music 2.6: $0.15 per up-to-5-minute track
- Lyrics gen: $0.01 per song
Also adds per_call_cost to schema.toml so the field is documented
alongside the other cost fields.
Adds two providers for BytePlus ModelArk, the international edition of
Volcano Engine, distinct from the existing cn-only `volcengine` provider:
- `byteplus`: standard `/api/v3` endpoint, 21 models total
- Text (10): Seed 2.0 Pro/Mini/Lite/Code, Seed 1.8, Seed Translation,
GLM-4.7, DeepSeek V3.2/V3.1, GPT-OSS-120B
- Image (4): Seedream 3.0/4.0/4.5/5.0-lite (per-piece pricing in comments)
- Video (7): Seedance 1.0/1.5 family + Dreamina Seedance 2.0/fast
- `byteplus_coding`: `/api/coding` Anthropic-compatible endpoint with
9 friendly aliases (ark-code-latest auto-router, bytedance-seed-code,
dola-seed-2.0-{pro,code,lite}, kimi-k2.5, glm-4.7, glm-5.1, gpt-oss-120b).
Both use BYTEPLUS_API_KEY env var. All listed model IDs were verified
to return HTTP 200 against the live ap-southeast endpoint. Pricing is the
standard real-time tier from the BytePlus console (a discounted batch tier
exists at roughly half the rate, not modeled here).
Excluded for follow-up (schema doesn't currently support these modalities):
- Skylark embedding-vision (no `embedding` modality in schema)
- Hyper3D-Gen2, Hitem3D-2.0 (no `3d` modality in schema)
Extend modality enum to support video and music, then register the
non-text MiniMax models that were already declared in
media_capabilities but had no concrete entries:
- image-01 ($0.0035/image)
- speech-2.8/2.6 hd & turbo ($60-$100 per 1M chars)
- Hailuo 2.3 Fast / 2.3 / 02 video models ($0.10-$0.56 per video)
- music-2.6, lyrics_generation
Per-call pricing is documented in inline comments since the schema's
token-based cost fields don't naturally fit per-call billing.
schema.toml and scripts/validate.py both updated; the change is
additive (existing modality values remain valid).
* chore: prune deprecated models across providers
Remove old-generation models that are strictly superseded by current versions
on the same provider/family. Affected providers: anthropic, bedrock, vertex-ai,
xai, moonshot, zhipu, baichuan, stepfun, volcengine, minimax, cohere, together,
fireworks, deepinfra, openrouter. Also clean up orphan aliases (grok3, grok-mini,
minimax-m2.1) and remap moonshot alias to kimi-k2.5.
Net: -42 model entries across 18 files. Provider model counts and README rows
updated accordingly.
* chore: remove redundant and orphan aliases from aliases.toml
Provider TOML files auto-register their model.aliases at load time, so
re-declaring them globally is duplication. Also drop entries pointing to
models that no longer exist after the prune.
- 45 redundant entries duplicating provider-defined aliases
- 11 orphan targets (gpt-4o, gpt-4o-mini, grok-2-mini, grok-3,
mixtral-8x7b-32768, copilot/gpt-4, open-mistral-nemo,
pixtral-large-latest, jamba-1.5-large, palmyra-x5, venice-uncensored)
Net: 100 lines down to 21. The file is now what the header comment
always claimed it was: 'additional global aliases not tied to a specific
model entry.'
* chore: second pass — prune more deprecated models
Apply the same 'strictly superseded by same-provider/family successor'
rule to providers missed in the first pass:
- openai: gpt-4.1 / -mini / -nano, o3, o4-mini (5)
- meta-llama: llama-3.3-70b-instruct (1)
- zhipu: glm-4v-plus (1)
- together: Llama-3.3-70B-Instruct-Turbo (1)
- xiaomi: mimo-v2-flash / -omni / -pro (3)
- aion-labs: aion-1.0 / -mini (2)
- qianfan: ernie-speed-128k, ernie-4.0-turbo-8k (2)
- cerebras: cerebras/llama3.1-8b (1)
- qwen-code: qwen-code/qwq-32b (1)
- nvidia-nim: 12 models (llama-3.1/3.2 series, mixtral-8x22b,
mistral-small-3.1, phi-4-mini, qwq-32b, r1-distill-32b,
qwen2.5-coder, nemotron-mini-4b, nemotron-70b-instruct)
- openrouter: meta-llama/llama-3.3-70b (paid), rekaai/reka-edge (2)
- alibaba-coding-plan: qwen3.5-plus, qwen3-max-2026-01-23,
MiniMax-M2.5, kimi-k2.5 (4)
Net: -35 model entries. providers/README.md model counts updated.
220 models remain.
On dual-stack hosts (notably macOS), `localhost` resolves to both ::1
and 127.0.0.1 with IPv6 tried first. Local LLM servers (Ollama, vLLM,
LM Studio) installed via the standard scripts bind IPv4 only, so the
IPv6 connection attempt fails immediately and Happy Eyeballs fallback
to IPv4 isn't reliably triggered for connection-refused errors,
producing spurious "Configured local provider offline" warnings in
the daemon even when the server is up and reachable via curl.
Companion to librefang/librefang#3112 which fixes the hardcoded URL
constants in the main repo. After both land, existing installs pick
up the fix on their next registry sync.
OpenAI-compatible LLM gateway. Pairs with the librefang-llm-drivers
registration (librefang PR #3076) so Novita models surface in the
dashboard model picker without each user having to add them via
/api/models/custom.
Pricing and limits sourced from GET /openai/v1/models on 2026-04-25
(input_token_price_per_m / output_token_price_per_m, divided by 10000
to get USD per million tokens). Curated to a small popular subset:
- deepseek/deepseek-v3.2 (frontier)
- moonshotai/kimi-k2-thinking (frontier, supports_thinking)
- minimax/minimax-m2 (smart)
- meta-llama/llama-3.3-70b-instruct (smart)
- qwen/qwen3-coder-30b-a3b-instruct (balanced)
- zai-org/glm-4.7-flash (balanced)
Introduces image-generation models as a first-class [[models]] entry via
a new `modality` field on the model schema ("text" default, "image",
"audio"). When modality != "text", context_window / max_output_tokens
are optional since no conventional context gate exists — OpenAI's
gpt-image-2 docs omit them.
Adds `image_input_cost_per_m` / `image_output_cost_per_m` alongside
existing text token cost fields to cover the 4-price structure OpenAI
uses for image generation (text $5/$10, image $8/$30 per 1M tokens).
Validator updated to:
- accept any modality in {text, image, audio}
- require context_window/max_output_tokens only for modality=text
- range-check the two new cost fields
gpt-image-2 entry added to providers/openai.toml with pricing sourced
from https://developers.openai.com/api/docs/pricing. Snapshot
gpt-image-2-2026-04-21 listed as alias.
* fix(providers): remove ~anthropic, skip ~ prefixes in sync script
OpenRouter uses ~ prefixes for internal auto-routing aliases (e.g. ~anthropic).
These are not real providers — they already route through openrouter.toml.
The generated ~anthropic.toml was confusing (looked like a stale backup)
and redundant with the existing openrouter provider.
- Delete providers/~anthropic.toml
- Skip provider IDs starting with ~ in sync-pricing.py --create-missing
* fix(providers): remove morph, aider, kwaipilot
- morph: specialized code-editing/patching tool, not a general LLM provider
- aider: CLI meta-tool wrapper (base_url empty), redundant with claude-code/codex-cli/gemini-cli/qwen-code
- kwaipilot: Kwai internal coding assistant routed via OpenRouter, niche
* fix(sync): add morph/aider/kwaipilot to SKIP_PROVIDERS to prevent re-creation
* feat(sync): merge OpenRouter-only providers into openrouter.toml
Instead of generating standalone .toml files that just wrap the OpenRouter
endpoint, merge their models directly into openrouter.toml with the
standard 'openrouter/{provider}/{model}' ID convention.
- Add _build_model_fields() and _model_lines() helpers to deduplicate
model rendering between standalone and merged paths
- Add merge_into_openrouter() that appends new models idempotently
- generate_provider_toml() now only runs for providers in PROVIDER_API
- --create-missing routes OpenRouter-only providers to merge_into_openrouter
* fix(providers): remove 14 OpenRouter-only standalone files
These providers have no direct public API and all route through
openrouter.ai/api/v1. Per the new sync-pricing.py policy, their models
will be merged into openrouter.toml on the next CI run instead of
living in separate files that just wrap the OpenRouter endpoint.
Removed: allenai, deepcogito, essentialai, inclusionai, inflection,
liquid, meituan, nex-agi, nousresearch, prime-intellect, relace,
switchpoint, tngtech, writer
* fix(providers): remove 7 niche providers with no driver support
No dedicated LLM driver code exists for these providers — they rely
purely on OpenAI-compatible passthrough with no special handling.
Removing them reduces registry noise; users can still reach them via
openrouter.toml if needed.
Removed: microsoft, ibm-granite, xiaomi, upstage, inception, aion-labs, arcee-ai
* fix(providers): remove ai21, chutes, venice
All three use ApiFormat::OpenAI with no special handling — pure passthrough.
No registry entry needed; users can reach them via openrouter.toml or by
adding a custom provider.
* docs(providers): rewrite README with full provider catalog and inclusion criteria
- List all 46 providers grouped by category with descriptions
- Document why each provider exists (direct API, unique endpoint, dedicated driver, local, CLI)
- Add inclusion criteria section explaining when to create standalone files vs merging into openrouter.toml
- Document sync script routing logic
- Update model counts: 49→46 providers, 339→232 models
* docs: add comprehensive READMEs for all registry sections + deepinfra provider
- agents/README.md: 32 agents across 7 categories with capability field reference
- channels/README.md: 44 channels across 5 categories with protocol reference table
- hands/README.md: 18 hands across 5 categories with HAND.toml format guide
- mcp/README.md: 33 MCP servers across 5 categories with transport/auth format
- plugins/README.md: 12 plugins with hook protocol documentation
- skills/README.md: 60 skills across 9 categories with SKILL.md format guide
- providers/deepinfra.toml: add DeepInfra serverless inference (5 models)
Add `claude-opus-4-7`, Anthropic's current flagship model (per
https://platform.claude.com/docs/en/docs/about-claude/models/overview).
- Context window: 1,000,000 tokens (Opus 4.7 ships with a new tokenizer)
- Max output: 128,000 tokens
- Pricing: $5 / input MTok, $25 / output MTok (unchanged from 4.6)
- Tier: frontier
- Supports tools, vision, streaming, and adaptive thinking
Move the `opus` / `claude-opus` aliases from 4.6 to 4.7 so a user asking
for "opus" gets the current flagship. Opus 4.6 is now in Anthropic's
"Legacy" section of the models overview; keeping it in the registry
entry (for existing callers that pin the exact ID) but without the
generic aliases.
Also **correct Opus 4.6's `context_window`**: it was listed as 200,000
tokens but the official model page has shown 1M tokens since release.
That was a pre-existing bug this PR fixes in passing since it directly
affects anyone who'd have routed queries to 4.6 expecting 1M.
Add a header comment pinning the source URL and explaining the
"latest-first" ordering convention so future additions don't silently
rebind aliases to a previous-generation snapshot.
Add the two GPT-5.4 variants exposed by OpenAI's API (source:
https://developers.openai.com/api/docs/models/gpt-5.4 and
https://developers.openai.com/api/docs/models/gpt-5.4-mini):
- **gpt-5.4** — frontier tier, 1,050,000 context window, 128k max output,
$2.50 / input MTok, $15.00 / output MTok.
- **gpt-5.4-mini** — balanced tier (matching the naming convention used
by `gpt-5-mini`, `gpt-4.1-mini`, etc.), 400,000 context window,
128k max output, $0.75 / input MTok, $4.50 / output MTok.
Ordered right after the GPT-5.2 family and before the Codex variants
section to keep the frontier-GPT chain in release order.
Addresses librefang/librefang-registry#65 — the original request filed
these under the `codex-cli` provider, but Codex CLI's upstream
`models.json` doesn't list `gpt-5.4-mini` (only the full `gpt-5.4`
slug is list-visible there, and it's already tracked in
`providers/codex-cli.toml`). The correct home for OpenAI-API-direct
access is this file; users who want `gpt-5.4-mini` should configure
`provider = "openai"` rather than `provider = "codex-cli"`.
The prior entries (o4-mini, o3, gpt-4.1) no longer appear in
upstream openai/codex's codex-rs/models-manager/models.json and the
Codex CLI actively migrates users away from the old gpt-5 /
gpt-5-codex slugs via tui/src/model_migration.rs.
Replace with the currently-visible ("visibility": "list") slugs:
- gpt-5.4
- gpt-5.3-codex
- gpt-5.2
- gpt-5.2-codex
- gpt-5.1-codex-max
- gpt-5.1-codex-mini
All six share a 272k context window per upstream models.json.
Marking supports_tools / supports_vision / supports_streaming = true
to match the gpt-5-family capability envelope.
Closeslibrefang/librefang#2347
Providers with known public APIs use their official endpoints:
- meta-llama → api.llama.com/v1
- microsoft → models.inference.ai.azure.com (GitHub Models)
- ibm-granite → us-south.ml.cloud.ibm.com/ml/v1 (watsonx)
- tencent → api.hunyuan.cloud.tencent.com/v1
- morph → api.morphllm.com/v1
16 remaining providers without known public APIs route through
OpenRouter (base_url = openrouter.ai/api/v1, OPENROUTER_API_KEY).
sync-pricing.py updated with PROVIDER_API mapping.
Providers without their own public API now use OpenRouter as their
base_url with OPENROUTER_API_KEY, making them testable and usable
when the user has an OpenRouter key configured.
- 20 OpenRouter-only providers: set base_url to openrouter.ai/api/v1
- morph: set correct official API (api.morphllm.com/v1)
- sync-pricing.py: default to OpenRouter routing for new providers
- Merge unique models from duplicate providers into their hand-written
counterparts and remove the duplicates:
- alibaba (tongyi-deepresearch) → qwen
- amazon (nova-2-lite, nova-micro, nova-premier) → bedrock
- bytedance (ui-tars) → volcengine
- nvidia (nemotron-3-nano, nemotron-3-super, etc.) → nvidia-nim
- rekaai (reka-flash-3) → reka
- Set correct official API base_url for providers with public APIs:
arcee-ai, inception, morph, reka, upstage
- Set key_required=false for 20 providers only accessible through
hosting platforms (no public API)
- Update sync-pricing.py with SKIP_DUPLICATES, PROVIDER_API mapping,
and default key_required=false for future auto-generated providers
- Replace tier "free" with "fast" (valid tiers: frontier/smart/balanced/fast/local)
- Remove version suffix from teams-mcp integration id field
- Update sync-pricing.py to not generate invalid tier values
* fix: pin npm package versions in MCP integration templates
Prevent supply chain attacks by pinning exact versions instead of
using unpinned `npx -y @package` which pulls latest on every run.
23 of 25 integrations pinned. sqlite-mcp and aws skipped (packages
not found on npm registry).
* fix: use stable azure/mcp version instead of beta
* feat: add pricing sync script and update model prices from OpenRouter API
- scripts/sync-pricing.py fetches real-time pricing from OpenRouter
- Updated 64 price fields across 13 provider files
- Run periodically or in CI to keep prices current
* feat: add Qwen International and US provider catalogs
* refactor: merge qwen-us into qwen-intl with regions support
* refactor: merge regional providers into single files with regions
Merge qwen-intl.toml into qwen.toml with [provider.regions] for
intl (Singapore) and us (Virginia) endpoints.
Merge minimax-cn.toml into minimax.toml with [provider.regions.china]
including separate api_key_env for China endpoint.
Uses new RegionConfig table format instead of simple string URLs.
* feat: add 4 context engine plugins
- topic-memory: keyword clustering for topic-aware memory recall
- episodic-memory: conversation segmentation and cross-session recall
- user-profile: persistent user profiling from conversation patterns
- context-decay: time-based memory decay with reinforcement dynamics
All plugins use the ingest/after_turn hook protocol with stdin/stdout JSON.
* chore: add plugin scaffolding, update docs and templates
- Add plugin.toml template with {{NAME}} placeholder
- Add new-plugin Makefile target with hooks/ scaffolding
- Update plugins/README.md with all 10 plugins
- Update README.md stats (10 plugins, 220+ models)
- Add Plugin checkbox and checklist to PR template
- Add Plugin to issue template content type dropdown
- Fix CONTRIBUTING.md: last_verified is recommended, not required
* fix: correct model pricing and remove deprecated entries
- openrouter/gemma-2-9b-it: fix pricing from 0.0 to 0.03/0.09 per M tokens
(free variant correctly stays at 0.0)
- github-copilot: remove deprecated copilot/gpt-4 model entry
(GPT-4 retired in favor of GPT-4o for Copilot)
* docs: annotate kimi-coding as membership-gated
Kimi Code CLI uses quota-based membership model (not per-token billing).
Free tier has limited weekly requests; underlying model is K2.5.
Pricing kept at 0.0 consistent with other subscription providers
(chatgpt, github-copilot) but with explanatory comments.
* style: fix trailing newline in github-copilot.toml
* fix: correct Moonshot/Kimi model pricing from official sources
All 5 models had incorrect pricing:
- moonshot-v1-8k: 0.10/0.10 → 0.20/2.00
- moonshot-v1-32k: 0.30/0.30 → 1.00/3.00
- moonshot-v1-128k: 0.80/0.80 → 2.00/5.00
- kimi-k2: 2.00/8.00 → 0.60/2.50
- kimi-k2.5: 2.00/8.00 → 0.45/2.20
Sources: platform.moonshot.ai/docs/pricing/chat, costgoat.com, getmaxim.ai
* feat: add MiniMax M2.7 and M2.7-highspeed models
Released 2026-03-18, MiniMax's latest flagship text model.
10B activated params, 200K context, 128K output, tool use, streaming.
Pricing: $0.30/$1.20 per M tokens (input/output).
Added to both international (minimax.io) and China (minimaxi.com) providers.
* chore: remove router agent
builtin:router has been replaced by LLM intent routing in the kernel.
Assistant is now the sole entry point — see librefang/librefang#1336.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* style: format all TOML files with taplo
Fix CI taplo format check by running `taplo fmt` on all 132 TOML files.
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>