Commit Graph
59 Commits
Author SHA1 Message Date
Evan 9a7a2751e4 feat(providers): split microsoft api_key_env from github-copilot (#83)
Both `microsoft` (GitHub Models / Azure AI Inference at
models.inference.ai.azure.com) and `github-copilot` (the IDE subscription
product) previously declared `api_key_env = "GITHUB_TOKEN"`, so a single
PAT silently activated both providers — users with only the IDE
subscription saw GitHub Models entries appear in their model picker
without intent, and vice versa.

Rename `microsoft` to `GITHUB_MODELS_TOKEN`. `github-copilot` keeps
`GITHUB_TOKEN` as the established convention for the IDE side.

Existing users who set `GITHUB_TOKEN` to use the GitHub Models endpoint
will see the `microsoft` provider become unavailable until they set the
new env var. The librefang daemon will be updated separately to
recognize `GITHUB_TOKEN` as a deprecated fallback for `microsoft` during
a migration window, mirroring the legacy-fallback infrastructure added
in librefang/librefang#3279.

Refs librefang/librefang#3278
2026-04-27 16:34:37 +09:00
Evan 95d29be8a6 feat(providers): split _coding api_key_env from main provider (#82)
Each `_coding` provider entry previously declared the same `api_key_env`
as its standard counterpart, causing one credential to silently activate
two providers. Users who only signed up for the per-token API endpoint
saw the Coding Plan endpoint's models mixed into their model picker
without intent.

Rename so each pair uses an independent env var:

  byteplus-coding    BYTEPLUS_API_KEY   -> BYTEPLUS_CODING_API_KEY
  volcengine-coding  VOLCENGINE_API_KEY -> VOLCENGINE_CODING_API_KEY
  zai-coding         ZHIPU_API_KEY      -> ZAI_CODING_API_KEY
  zhipu-coding       ZHIPU_API_KEY      -> ZHIPU_CODING_API_KEY

This is a breaking change: existing users must set the new env vars.
The librefang daemon will be updated separately to recognize the old
env vars as a deprecated fallback during a migration window. See
librefang/librefang#3278 for the design discussion and rollout plan.

Out of scope (different problem class, tracked in the same issue):

  - zai/zhipu cross-region key sharing (both still share ZHIPU_API_KEY)
  - github-copilot/microsoft cross-product GITHUB_TOKEN reuse

Refs librefang/librefang#3278
2026-04-27 15:12:14 +09:00
Evan 23621ef25e docs(byteplus): warn that byteplus is per-token, not Coding Plan (#80)
* docs(byteplus): warn that `byteplus` is per-token, not Coding Plan

Self-followup on PR #78. The BytePlus official docs explicitly warn:

  Do not use the standard model endpoint
  (https://ark.ap-southeast.bytepluses.com/api/v3) [for Coding Plan
  workloads], as requests there bypass Coding Plan quota and incur
  separate charges.
  — https://docs.byteplus.com/en/docs/ModelArk/1928261

Without this warning visible to users, anyone with a Coding Plan
subscription who picks `provider = "byteplus"` will silently bill
against their per-token USD balance instead of consuming the
subscription quota they paid for.

Adds a prominent "BILLING — READ BEFORE USE" block to byteplus.toml
making the per-token vs subscription split unambiguous, and a
companion note in byteplus-coding.toml pointing back so the choice is
discoverable from either side. Also calls out which capabilities are
unique to the standard endpoint (image / video) so the choice isn't
"just use Coding Plan for everything".

No model definitions or pricing changed.

* chore(byteplus): drop superseded model entries (21 → 11) (#81)

The original byteplus.toml from PR #78 enumerated every BytePlus
ModelArk endpoint that returned HTTP 200, regardless of whether a
sane user would still pick it. Trims to the current per-family
flagship plus useful fast/preview tiers.

Removed (10):
  Text:
    seed-2-0-lite-260228       — superseded by seed-2-0-mini in the
                                 fast/cheap niche
    seed-1-8-251228            — superseded by seed-2-0 family
    seed-translation-250915    — too narrow; chat models cover this
    deepseek-v3-1-250821       — superseded by deepseek-v3-2
  Image:
    seedream-3-0-t2i-250415    — superseded by seedream-4-5 / 5-0-lite
    seedream-4-0-250828        — same
  Video:
    seedance-1-0-lite-i2v-250428    — superseded by 1-5 / dreamina-2-0
    seedance-1-0-lite-t2v-250428    — same
    seedance-1-0-pro-250528         — same
    seedance-1-0-pro-fast-251015    — same

Kept (11):
  Text (6):  seed-2-0-pro/mini/code-preview, glm-4-7,
             deepseek-v3-2, gpt-oss-120b
  Image (2): seedream-4-5, seedream-5-0 (lite)
  Video (3): seedance-1-5-pro, dreamina-seedance-2-0,
             dreamina-seedance-2-0-fast

Also corrects two per-piece / per-K price comments that the trim
left orphaned over the wrong [[models]] block (4-5 = $0.0400, not
$0.0300; 1-5-pro = $0.0024/$0.0012 with/without audio, not
$0.0018/K) and updates the header capability list to match the
new model set.
2026-04-27 12:03:40 +09:00
Evan 2c3487a2f8 feat(minimax): add per_call_cost for video and music models (#79)
Video and music modality models are billed per generation rather than
per token. Without this field the librefang runtime metering layer
records every call as $0 and emits a warning. Sourced from
https://platform.minimax.io/docs/guides/pricing-paygo.

- Hailuo 2.3 Fast:  $0.33 per call (1080P/6s upper bound)
- Hailuo 2.3:       $0.56 per call (1080P/6s upper bound)
- Hailuo 02:        $0.56 per call (1080P/6s upper bound)
- Music 2.6:        $0.15 per up-to-5-minute track
- Lyrics gen:       $0.01 per song

Also adds per_call_cost to schema.toml so the field is documented
alongside the other cost fields.
2026-04-27 10:55:08 +09:00
Evan 62bdd04901 feat(providers): add BytePlus ModelArk (international) + coding endpoint (#78)
Adds two providers for BytePlus ModelArk, the international edition of
Volcano Engine, distinct from the existing cn-only `volcengine` provider:

- `byteplus`: standard `/api/v3` endpoint, 21 models total
  - Text (10): Seed 2.0 Pro/Mini/Lite/Code, Seed 1.8, Seed Translation,
    GLM-4.7, DeepSeek V3.2/V3.1, GPT-OSS-120B
  - Image (4): Seedream 3.0/4.0/4.5/5.0-lite (per-piece pricing in comments)
  - Video (7): Seedance 1.0/1.5 family + Dreamina Seedance 2.0/fast
- `byteplus_coding`: `/api/coding` Anthropic-compatible endpoint with
  9 friendly aliases (ark-code-latest auto-router, bytedance-seed-code,
  dola-seed-2.0-{pro,code,lite}, kimi-k2.5, glm-4.7, glm-5.1, gpt-oss-120b).

Both use BYTEPLUS_API_KEY env var. All listed model IDs were verified
to return HTTP 200 against the live ap-southeast endpoint. Pricing is the
standard real-time tier from the BytePlus console (a discounted batch tier
exists at roughly half the rate, not modeled here).

Excluded for follow-up (schema doesn't currently support these modalities):
- Skylark embedding-vision (no `embedding` modality in schema)
- Hyper3D-Gen2, Hitem3D-2.0 (no `3d` modality in schema)
2026-04-27 10:52:57 +09:00
Evan 82d5a6ecd5 feat(minimax): add image/audio/video/music model entries (#77)
Extend modality enum to support video and music, then register the
non-text MiniMax models that were already declared in
media_capabilities but had no concrete entries:

- image-01 ($0.0035/image)
- speech-2.8/2.6 hd & turbo ($60-$100 per 1M chars)
- Hailuo 2.3 Fast / 2.3 / 02 video models ($0.10-$0.56 per video)
- music-2.6, lyrics_generation

Per-call pricing is documented in inline comments since the schema's
token-based cost fields don't naturally fit per-call billing.

schema.toml and scripts/validate.py both updated; the change is
additive (existing modality values remain valid).
2026-04-27 09:44:14 +09:00
Evan d3b9814fb1 feat: add 2026 Q2 flagship models (#76)
Add latest flagships released in April 2026 that registry missed:
- deepseek: V4-Pro, V4-Flash (2026-04-24, 1M context)
- qwen: qwen3.6-max-preview (2026-04-20, 256K context)
- moonshot: kimi-k2.6 (2026-04-20, 256K context)
- zhipu: glm-5.1 (2026-04-08), glm-4.7-flash (free tier)

Update default aliases: deepseek -> v4-pro, kimi -> k2.6, glm -> 5.1.
2026-04-27 09:43:52 +09:00
Evan 541052dc79 chore: prune deprecated models across providers (#75)
* chore: prune deprecated models across providers

Remove old-generation models that are strictly superseded by current versions
on the same provider/family. Affected providers: anthropic, bedrock, vertex-ai,
xai, moonshot, zhipu, baichuan, stepfun, volcengine, minimax, cohere, together,
fireworks, deepinfra, openrouter. Also clean up orphan aliases (grok3, grok-mini,
minimax-m2.1) and remap moonshot alias to kimi-k2.5.

Net: -42 model entries across 18 files. Provider model counts and README rows
updated accordingly.

* chore: remove redundant and orphan aliases from aliases.toml

Provider TOML files auto-register their model.aliases at load time, so
re-declaring them globally is duplication. Also drop entries pointing to
models that no longer exist after the prune.

- 45 redundant entries duplicating provider-defined aliases
- 11 orphan targets (gpt-4o, gpt-4o-mini, grok-2-mini, grok-3,
  mixtral-8x7b-32768, copilot/gpt-4, open-mistral-nemo,
  pixtral-large-latest, jamba-1.5-large, palmyra-x5, venice-uncensored)

Net: 100 lines down to 21. The file is now what the header comment
always claimed it was: 'additional global aliases not tied to a specific
model entry.'

* chore: second pass — prune more deprecated models

Apply the same 'strictly superseded by same-provider/family successor'
rule to providers missed in the first pass:

- openai: gpt-4.1 / -mini / -nano, o3, o4-mini (5)
- meta-llama: llama-3.3-70b-instruct (1)
- zhipu: glm-4v-plus (1)
- together: Llama-3.3-70B-Instruct-Turbo (1)
- xiaomi: mimo-v2-flash / -omni / -pro (3)
- aion-labs: aion-1.0 / -mini (2)
- qianfan: ernie-speed-128k, ernie-4.0-turbo-8k (2)
- cerebras: cerebras/llama3.1-8b (1)
- qwen-code: qwen-code/qwq-32b (1)
- nvidia-nim: 12 models (llama-3.1/3.2 series, mixtral-8x22b,
  mistral-small-3.1, phi-4-mini, qwq-32b, r1-distill-32b,
  qwen2.5-coder, nemotron-mini-4b, nemotron-70b-instruct)
- openrouter: meta-llama/llama-3.3-70b (paid), rekaai/reka-edge (2)
- alibaba-coding-plan: qwen3.5-plus, qwen3-max-2026-01-23,
  MiniMax-M2.5, kimi-k2.5 (4)

Net: -35 model entries. providers/README.md model counts updated.
220 models remain.
2026-04-27 09:36:43 +09:00
Evan 7398983350 fix: use 127.0.0.1 instead of localhost for local provider base URLs (#73)
On dual-stack hosts (notably macOS), `localhost` resolves to both ::1
and 127.0.0.1 with IPv6 tried first. Local LLM servers (Ollama, vLLM,
LM Studio) installed via the standard scripts bind IPv4 only, so the
IPv6 connection attempt fails immediately and Happy Eyeballs fallback
to IPv4 isn't reliably triggered for connection-refused errors,
producing spurious "Configured local provider offline" warnings in
the daemon even when the server is up and reachable via curl.

Companion to librefang/librefang#3112 which fixes the hardcoded URL
constants in the main repo. After both land, existing installs pick
up the fix on their next registry sync.
2026-04-25 18:58:59 +09:00
github-actions[bot] 2097fc05bc chore: sync model pricing from OpenRouter API 2026-04-25 07:26:35 +00:00
Evan 12d19943c5 feat(novita): add Novita AI provider with 6 popular models (#72)
OpenAI-compatible LLM gateway. Pairs with the librefang-llm-drivers
registration (librefang PR #3076) so Novita models surface in the
dashboard model picker without each user having to add them via
/api/models/custom.

Pricing and limits sourced from GET /openai/v1/models on 2026-04-25
(input_token_price_per_m / output_token_price_per_m, divided by 10000
to get USD per million tokens). Curated to a small popular subset:

- deepseek/deepseek-v3.2 (frontier)
- moonshotai/kimi-k2-thinking (frontier, supports_thinking)
- minimax/minimax-m2 (smart)
- meta-llama/llama-3.3-70b-instruct (smart)
- qwen/qwen3-coder-30b-a3b-instruct (balanced)
- zai-org/glm-4.7-flash (balanced)
2026-04-25 14:16:04 +09:00
Evan 5909b024c1 feat(openai): add GPT Image 2 (image-generation modality) (#71)
Introduces image-generation models as a first-class [[models]] entry via
a new `modality` field on the model schema ("text" default, "image",
"audio"). When modality != "text", context_window / max_output_tokens
are optional since no conventional context gate exists — OpenAI's
gpt-image-2 docs omit them.

Adds `image_input_cost_per_m` / `image_output_cost_per_m` alongside
existing text token cost fields to cover the 4-price structure OpenAI
uses for image generation (text $5/$10, image $8/$30 per 1M tokens).

Validator updated to:
- accept any modality in {text, image, audio}
- require context_window/max_output_tokens only for modality=text
- range-check the two new cost fields

gpt-image-2 entry added to providers/openai.toml with pricing sourced
from https://developers.openai.com/api/docs/pricing. Snapshot
gpt-image-2-2026-04-21 listed as alias.
2026-04-25 13:30:07 +09:00
Evan 65cb852632 feat(openai): add GPT-5.5 and GPT-5.5 Pro (#70)
Source: https://openai.com/index/introducing-gpt-5-5/ (announcement 2026-04-23)

Adds 4 new model entries across 3 provider files:

- providers/openai.toml:
    gpt-5.5      — 1M context,  $5/$30 per 1M tokens (input/output)
    gpt-5.5-pro  — 1M context, $30/$180 per 1M tokens
- providers/codex-cli.toml:
    codex-cli/gpt-5.5  — 400K context (Codex subscription limit),
                         $0/$0 (covered by subscription)
- providers/chatgpt.toml:
    gpt-5.5-codex  — 400K context, session-auth (subscription)

Notes:
- API availability announced as 'very soon'; pricing confirmed in the
  announcement. tools/vision/streaming flags match GPT-5.4 family.
- max_output_tokens retained at 128000 (openai/codex-cli) and 65536
  (chatgpt) — matches sibling 5.4 entries; announcement does not
  specify a new output cap.
- Header comment in openai.toml updated (15 → 17 models).
2026-04-25 00:08:46 +09:00
github-actions[bot] b1ebaf6d34 chore: sync model pricing from OpenRouter API 2026-04-24 08:11:38 +00:00
Evan d43077afa9 fix(providers): remove ~anthropic, skip ~ prefixes in sync script (#69)
* fix(providers): remove ~anthropic, skip ~ prefixes in sync script

OpenRouter uses ~ prefixes for internal auto-routing aliases (e.g. ~anthropic).
These are not real providers — they already route through openrouter.toml.
The generated ~anthropic.toml was confusing (looked like a stale backup)
and redundant with the existing openrouter provider.

- Delete providers/~anthropic.toml
- Skip provider IDs starting with ~ in sync-pricing.py --create-missing

* fix(providers): remove morph, aider, kwaipilot

- morph: specialized code-editing/patching tool, not a general LLM provider
- aider: CLI meta-tool wrapper (base_url empty), redundant with claude-code/codex-cli/gemini-cli/qwen-code
- kwaipilot: Kwai internal coding assistant routed via OpenRouter, niche

* fix(sync): add morph/aider/kwaipilot to SKIP_PROVIDERS to prevent re-creation

* feat(sync): merge OpenRouter-only providers into openrouter.toml

Instead of generating standalone .toml files that just wrap the OpenRouter
endpoint, merge their models directly into openrouter.toml with the
standard 'openrouter/{provider}/{model}' ID convention.

- Add _build_model_fields() and _model_lines() helpers to deduplicate
  model rendering between standalone and merged paths
- Add merge_into_openrouter() that appends new models idempotently
- generate_provider_toml() now only runs for providers in PROVIDER_API
- --create-missing routes OpenRouter-only providers to merge_into_openrouter

* fix(providers): remove 14 OpenRouter-only standalone files

These providers have no direct public API and all route through
openrouter.ai/api/v1. Per the new sync-pricing.py policy, their models
will be merged into openrouter.toml on the next CI run instead of
living in separate files that just wrap the OpenRouter endpoint.

Removed: allenai, deepcogito, essentialai, inclusionai, inflection,
liquid, meituan, nex-agi, nousresearch, prime-intellect, relace,
switchpoint, tngtech, writer

* fix(providers): remove 7 niche providers with no driver support

No dedicated LLM driver code exists for these providers — they rely
purely on OpenAI-compatible passthrough with no special handling.
Removing them reduces registry noise; users can still reach them via
openrouter.toml if needed.

Removed: microsoft, ibm-granite, xiaomi, upstage, inception, aion-labs, arcee-ai

* fix(providers): remove ai21, chutes, venice

All three use ApiFormat::OpenAI with no special handling — pure passthrough.
No registry entry needed; users can reach them via openrouter.toml or by
adding a custom provider.

* docs(providers): rewrite README with full provider catalog and inclusion criteria

- List all 46 providers grouped by category with descriptions
- Document why each provider exists (direct API, unique endpoint, dedicated driver, local, CLI)
- Add inclusion criteria section explaining when to create standalone files vs merging into openrouter.toml
- Document sync script routing logic
- Update model counts: 49→46 providers, 339→232 models

* docs: add comprehensive READMEs for all registry sections + deepinfra provider

- agents/README.md: 32 agents across 7 categories with capability field reference
- channels/README.md: 44 channels across 5 categories with protocol reference table
- hands/README.md: 18 hands across 5 categories with HAND.toml format guide
- mcp/README.md: 33 MCP servers across 5 categories with transport/auth format
- plugins/README.md: 12 plugins with hook protocol documentation
- skills/README.md: 60 skills across 9 categories with SKILL.md format guide
- providers/deepinfra.toml: add DeepInfra serverless inference (5 models)
2026-04-24 00:02:33 +09:00
github-actions[bot] dbfb32d9d4 chore: sync model pricing from OpenRouter API 2026-04-22 08:03:48 +00:00
github-actions[bot] 1ecca29fee chore: sync model pricing from OpenRouter API 2026-04-21 08:00:48 +00:00
Evan 80c6ee79cd feat(anthropic): add Claude Opus 4.7 and fix Opus 4.6 context window (#66)
Add `claude-opus-4-7`, Anthropic's current flagship model (per
https://platform.claude.com/docs/en/docs/about-claude/models/overview).

- Context window: 1,000,000 tokens (Opus 4.7 ships with a new tokenizer)
- Max output: 128,000 tokens
- Pricing: $5 / input MTok, $25 / output MTok (unchanged from 4.6)
- Tier: frontier
- Supports tools, vision, streaming, and adaptive thinking

Move the `opus` / `claude-opus` aliases from 4.6 to 4.7 so a user asking
for "opus" gets the current flagship. Opus 4.6 is now in Anthropic's
"Legacy" section of the models overview; keeping it in the registry
entry (for existing callers that pin the exact ID) but without the
generic aliases.

Also **correct Opus 4.6's `context_window`**: it was listed as 200,000
tokens but the official model page has shown 1M tokens since release.
That was a pre-existing bug this PR fixes in passing since it directly
affects anyone who'd have routed queries to 4.6 expecting 1M.

Add a header comment pinning the source URL and explaining the
"latest-first" ordering convention so future additions don't silently
rebind aliases to a previous-generation snapshot.
2026-04-20 14:15:45 +09:00
Evan 396c88dc49 feat(openai): add GPT-5.4 and GPT-5.4-mini (#67)
Add the two GPT-5.4 variants exposed by OpenAI's API (source:
https://developers.openai.com/api/docs/models/gpt-5.4 and
https://developers.openai.com/api/docs/models/gpt-5.4-mini):

- **gpt-5.4** — frontier tier, 1,050,000 context window, 128k max output,
  $2.50 / input MTok, $15.00 / output MTok.
- **gpt-5.4-mini** — balanced tier (matching the naming convention used
  by `gpt-5-mini`, `gpt-4.1-mini`, etc.), 400,000 context window,
  128k max output, $0.75 / input MTok, $4.50 / output MTok.

Ordered right after the GPT-5.2 family and before the Codex variants
section to keep the frontier-GPT chain in release order.

Addresses librefang/librefang-registry#65 — the original request filed
these under the `codex-cli` provider, but Codex CLI's upstream
`models.json` doesn't list `gpt-5.4-mini` (only the full `gpt-5.4`
slug is list-visible there, and it's already tracked in
`providers/codex-cli.toml`). The correct home for OpenAI-API-direct
access is this file; users who want `gpt-5.4-mini` should configure
`provider = "openai"` rather than `provider = "codex-cli"`.
2026-04-20 14:15:20 +09:00
github-actions[bot] dc23636656 chore: sync model pricing from OpenRouter API 2026-04-19 07:25:57 +00:00
github-actions[bot] e859083cac chore: sync model pricing from OpenRouter API 2026-04-18 07:15:25 +00:00
github-actions[bot] 9f740e4843 chore: sync model pricing from OpenRouter API 2026-04-17 07:56:38 +00:00
github-actions[bot] 6d44be4277 chore: sync model pricing from OpenRouter API 2026-04-16 07:55:20 +00:00
Evan 4f8dd2404d feat(providers): expand ollama catalog with thinking-capable models (#58)
* feat(providers): expand ollama model catalog with thinking-capable models

Add commonly used local models with accurate capability flags:
- gemma4, gemma3: supports_thinking, supports_vision
- deepseek-r1: supports_thinking (fix missing flag)
- deepseek-v3: supports_tools
- qwen3, qwq: supports_thinking
- llama4: supports_vision
- llama3.3: supports_tools
- phi4: supports_tools

Previously only 6 models were listed and none had supports_thinking
(except deepseek-r1), causing the dashboard to hide thinking toggles
for models that actually support it.

* chore(providers): major cleanup — remove defunct providers and old models

Delete 21 defunct/obscure providers:
aion-labs, arcee-ai, deepcogito, eleutherai, essentialai, ibm-granite,
inception, inflection, kwaipilot, lemonade, liquid, morph, nex-agi,
nousresearch, prime-intellect, reka, relace, switchpoint, tngtech,
upstage, writer

Clean up 10 major providers — keep only latest generation models:
- anthropic: remove claude-3.5-sonnet (superseded by 4.x)
- openai: remove gpt-4o/4-turbo/3.5/o1/o3-mini (superseded by gpt-5/4.1/o3/o4-mini)
- gemini: remove 1.5-*/2.0-flash (superseded by 2.5/3.x)
- deepseek: remove coder/chat-v3-0324 (superseded by r1/v3)
- qwen: remove turbo/2.5-coder (superseded by qwen3)
- groq: remove old llama/mixtral/gemma (keep latest only)
- mistral: remove medium/nemo/pixtral-large (keep large/small/codestral)
- xai: remove grok-2 (superseded by grok-3/4)
- meta-llama: remove 3.x/guard (keep llama-4 + 3.3)
- ollama: rewrite with current models (gemma4, qwen3, qwq, llama4, etc)

Total: 90 → 48 models across major providers. All thinking-capable
models now have supports_thinking = true.

* chore: add pre-commit hook for automatic TOML formatting

- .githooks/pre-commit: runs taplo fmt on staged .toml files
- Makefile: add setup target + auto-configure hooks on first make
- .gitignore: add .make-setup-done and .sync_marker
2026-04-15 22:18:01 +09:00
github-actions[bot] 4979c355b2 chore: sync model pricing from OpenRouter API 2026-04-15 07:55:35 +00:00
Evan 85db9c8d78 feat(providers): add supports_thinking to thinking-capable models (#53)
* feat(providers): add supports_thinking field to thinking-capable models

Mark models that support extended thinking / reasoning with
supports_thinking = true so the dashboard can conditionally show
thinking toggles.

Providers updated: anthropic (8), codex-cli (7), gemini (6),
openai (3), qwen (2), deepseek (1). Schema updated accordingly.

* feat(providers): add supports_thinking to remaining thinking-capable models

Cover 19 additional providers: alibaba-coding-plan, allenai, arcee-ai,
fireworks, gemini-cli, groq, liquid, nvidia-nim, ollama, openai (codex),
openrouter, perplexity, qwen-code, replicate, sambanova, tngtech,
venice, vertex-ai, xai.

Total: 62 models across 24 providers now have supports_thinking = true.

* feat(providers): add supports_thinking to chutes, huggingface, together

Missed in prior commits: DeepSeek-R1 on chutes/huggingface/together,
Qwen3-235B on chutes. Total now 66 models across 27 providers.

* feat(providers): add supports_thinking to bedrock, claude-code, aider, moonshot, stepfun

- bedrock: all 5 Claude models
- claude-code: all 3 models (opus/sonnet/haiku wrappers)
- aider: aider/sonnet (Claude-backed)
- moonshot: kimi-k2.5 (reasoning mode)
- stepfun: step-1o-turbo-vision (reasoning model)

Total: 77 models across 32 providers.

* feat(providers): add supports_thinking to alibaba kimi-k2.5, openrouter claude-sonnet-4
2026-04-15 00:46:51 +09:00
Evan 2f4f94b45b fix(codex-cli): refresh model list to match upstream (#2347) (#50)
The prior entries (o4-mini, o3, gpt-4.1) no longer appear in
upstream openai/codex's codex-rs/models-manager/models.json and the
Codex CLI actively migrates users away from the old gpt-5 /
gpt-5-codex slugs via tui/src/model_migration.rs.

Replace with the currently-visible ("visibility": "list") slugs:

- gpt-5.4
- gpt-5.3-codex
- gpt-5.2
- gpt-5.2-codex
- gpt-5.1-codex-max
- gpt-5.1-codex-mini

All six share a 272k context window per upstream models.json.
Marking supports_tools / supports_vision / supports_streaming = true
to match the gpt-5-family capability envelope.

Closes librefang/librefang#2347
2026-04-14 19:40:56 +09:00
Joshua Chong 49019d4c42 add qwen3.6-plus from coding plan 2026-04-13 20:07:54 +08:00
github-actions[bot] 07772c22ea chore: sync model pricing from OpenRouter API 2026-04-11 07:06:23 +00:00
Evan 2468e87968 fix(openrouter): use :free variant for qwen3.6-plus default (#46) 2026-04-10 22:12:30 +08:00
Evan 919ac1acdf fix(openrouter): replace delisted models and set qwen3.6-plus as default (#45) 2026-04-10 21:26:06 +08:00
github-actions[bot] 3b5bde3486 chore: sync model pricing from OpenRouter API 2026-04-03 07:16:43 +00:00
Evan 5967c09c1b fix: correct morph API base_url to api.morphllm.com (#40) 2026-04-02 23:59:19 +08:00
Evan c192f493b4 fix: route providers through correct APIs (#39)
Providers with known public APIs use their official endpoints:
- meta-llama → api.llama.com/v1
- microsoft → models.inference.ai.azure.com (GitHub Models)
- ibm-granite → us-south.ml.cloud.ibm.com/ml/v1 (watsonx)
- tencent → api.hunyuan.cloud.tencent.com/v1
- morph → api.morphllm.com/v1

16 remaining providers without known public APIs route through
OpenRouter (base_url = openrouter.ai/api/v1, OPENROUTER_API_KEY).

sync-pricing.py updated with PROVIDER_API mapping.
2026-04-02 23:55:42 +08:00
Evan aa5822992a fix: route OpenRouter-only providers through OpenRouter API (#38)
Providers without their own public API now use OpenRouter as their
base_url with OPENROUTER_API_KEY, making them testable and usable
when the user has an OpenRouter key configured.

- 20 OpenRouter-only providers: set base_url to openrouter.ai/api/v1
- morph: set correct official API (api.morphllm.com/v1)
- sync-pricing.py: default to OpenRouter routing for new providers
2026-04-02 23:41:25 +08:00
Evan 2fa48bfdcb fix: clean up OpenRouter-generated provider configs (#37)
- Merge unique models from duplicate providers into their hand-written
  counterparts and remove the duplicates:
  - alibaba (tongyi-deepresearch) → qwen
  - amazon (nova-2-lite, nova-micro, nova-premier) → bedrock
  - bytedance (ui-tars) → volcengine
  - nvidia (nemotron-3-nano, nemotron-3-super, etc.) → nvidia-nim
  - rekaai (reka-flash-3) → reka
- Set correct official API base_url for providers with public APIs:
  arcee-ai, inception, morph, reka, upstage
- Set key_required=false for 20 providers only accessible through
  hosting platforms (no public API)
- Update sync-pricing.py with SKIP_DUPLICATES, PROVIDER_API mapping,
  and default key_required=false for future auto-generated providers
2026-04-02 23:30:56 +08:00
github-actions[bot] 8104d86caa chore: sync model pricing from OpenRouter API 2026-04-02 07:20:15 +00:00
Joshua Chong 60313a7009 feat: add alibaba coding plan as provider (#32)
* add alibaba coding plan as provider

* add alibaba coding plan as provider
2026-03-31 22:55:53 +08:00
github-actions[bot] 479c73832b chore: sync model pricing from OpenRouter API 2026-03-31 07:23:23 +00:00
github-actions[bot] 8e8aaf202d chore: sync model pricing from OpenRouter API 2026-03-29 07:11:03 +00:00
github-actions[bot] 609306b919 chore: sync model pricing from OpenRouter API 2026-03-28 07:04:27 +00:00
github-actions[bot] f7251982d9 chore: sync model pricing from OpenRouter API 2026-03-27 07:15:47 +00:00
github-actions[bot] 3f56adc784 chore: sync model pricing from OpenRouter API 2026-03-26 07:16:28 +00:00
github-actions[bot] 310c03b1e6 chore: sync model pricing from OpenRouter API 2026-03-25 14:50:12 +00:00
Evan fc37ce253b fix: correct invalid tier "free" and teams-mcp id with version (#29)
- Replace tier "free" with "fast" (valid tiers: frontier/smart/balanced/fast/local)
- Remove version suffix from teams-mcp integration id field
- Update sync-pricing.py to not generate invalid tier values
2026-03-25 23:49:39 +09:00
Evan 553ecc6947 feat: add pricing sync script and update model prices from OpenRouter (#27)
* fix: pin npm package versions in MCP integration templates

Prevent supply chain attacks by pinning exact versions instead of
using unpinned `npx -y @package` which pulls latest on every run.

23 of 25 integrations pinned. sqlite-mcp and aws skipped (packages
not found on npm registry).

* fix: use stable azure/mcp version instead of beta

* feat: add pricing sync script and update model prices from OpenRouter API

- scripts/sync-pricing.py fetches real-time pricing from OpenRouter
- Updated 64 price fields across 13 provider files
- Run periodically or in CI to keep prices current
2026-03-25 23:43:13 +09:00
Evan a00e813bdf feat: add 22 mainstream models to nvidia-nim provider catalog (#26)
Add model entries for the most popular models available on NVIDIA NIM:
- NVIDIA: Nemotron Ultra 253B, Super 49B, 70B, Mini 4B
- Meta: Llama 4 Maverick/Scout, Llama 3.3 70B, 3.1 405B/8B, 3.2 Vision
- Mistral: Large 3 675B, Small 3.1 24B, Mixtral 8x22B
- DeepSeek: V3.2, R1 Distill 32B
- Qwen: 3.5 397B, 2.5 Coder 32B, QWQ 32B
- Google: Gemma 3 27B
- Microsoft: Phi-4 Multimodal, Phi-4 Mini

Refs librefang/librefang#1621
2026-03-25 23:36:00 +09:00
Evan 9bcbe04d08 feat: add media_capabilities to provider definitions (#18) 2026-03-23 14:13:13 +09:00
Evan 31f17ba369 feat: add Qwen International provider with regional endpoints (#8)
* feat: add Qwen International and US provider catalogs

* refactor: merge qwen-us into qwen-intl with regions support

* refactor: merge regional providers into single files with regions

Merge qwen-intl.toml into qwen.toml with [provider.regions] for
intl (Singapore) and us (Virginia) endpoints.

Merge minimax-cn.toml into minimax.toml with [provider.regions.china]
including separate api_key_env for China endpoint.

Uses new RegionConfig table format instead of simple string URLs.
2026-03-21 17:23:14 +09:00
Evan 078218c70f feat: add Gemini CLI, Codex CLI, and Aider provider catalogs (#7)
- gemini-cli: Google Gemini CLI (gemini-2.5-pro, gemini-2.5-flash)
- codex-cli: OpenAI Codex CLI (o4-mini, o3, gpt-4.1)
- aider: Aider AI coding assistant (sonnet, gpt-4o)
2026-03-21 07:14:18 +09:00