Refs librefang/librefang#4842 — long-term replacement for the substring
match that the OpenAI driver currently uses to decide how to handle
`reasoning_content` on historical assistant turns.
Three provider-specific behaviours that the driver must distinguish at
wire time, now expressed as catalog metadata:
* `strip` — DeepSeek R1 / deepseek-reasoner. The API rejects requests
that carry reasoning_content on previous assistant messages.
* `echo` — DeepSeek V4 Flash. Thinking mode is on by default and the
API rejects multi-turn requests when assistant turns containing
tool_calls don't echo back the original reasoning text. This is the
bug surfaced in librefang/librefang#4842.
* `empty_string` — Moonshot / Kimi K2 family. The field must be present
(empty string) on tool_calls turns, with thinking disabled wire-side
for multi-turn compatibility.
* `none` (default) — most providers; field is omitted entirely.
V4 Pro is intentionally NOT marked `echo` — librefang#4842 reports it
working out-of-the-box; flip when there's an empirical reproducer.
Marks affected models:
providers/deepseek.toml
deepseek-v4-flash → echo
deepseek-reasoner → strip
providers/moonshot.toml
kimi-k2.6, kimi-k2.5, kimi-k2 → empty_string
providers/kimi-coding.toml
kimi-for-coding → empty_string
providers/byteplus-coding.toml
kimi-k2.5 → empty_string
providers/novita.toml
moonshotai/kimi-k2-thinking → empty_string
Tooling:
* schema.toml registers the field with the four enum options and a
`none` default so existing TOML files keep parsing unchanged.
* scripts/validate.py rejects unknown enum values; verified with a
hand-crafted negative case (`reasoning_echo_policy = "bogus"` →
validation fails with the expected message).
* `python3 scripts/validate.py` passes (267 models).
The librefang side that consumes this field will land in a follow-up
PR — until then, registry consumers ignore the field via
`#[serde(default)]` and the existing substring fallback continues to
work, so this commit is safe to ship independently.
The previous "BytePlus ModelArk Coding Plan" overflows in dashboard
provider cards and CLI lists. Drop the redundant "ModelArk" — it's
already implied by the parent product line, and the file's header
comment + docs explain the relationship.
Each `_coding` provider entry previously declared the same `api_key_env`
as its standard counterpart, causing one credential to silently activate
two providers. Users who only signed up for the per-token API endpoint
saw the Coding Plan endpoint's models mixed into their model picker
without intent.
Rename so each pair uses an independent env var:
byteplus-coding BYTEPLUS_API_KEY -> BYTEPLUS_CODING_API_KEY
volcengine-coding VOLCENGINE_API_KEY -> VOLCENGINE_CODING_API_KEY
zai-coding ZHIPU_API_KEY -> ZAI_CODING_API_KEY
zhipu-coding ZHIPU_API_KEY -> ZHIPU_CODING_API_KEY
This is a breaking change: existing users must set the new env vars.
The librefang daemon will be updated separately to recognize the old
env vars as a deprecated fallback during a migration window. See
librefang/librefang#3278 for the design discussion and rollout plan.
Out of scope (different problem class, tracked in the same issue):
- zai/zhipu cross-region key sharing (both still share ZHIPU_API_KEY)
- github-copilot/microsoft cross-product GITHUB_TOKEN reuse
Refs librefang/librefang#3278
* docs(byteplus): warn that `byteplus` is per-token, not Coding Plan
Self-followup on PR #78. The BytePlus official docs explicitly warn:
Do not use the standard model endpoint
(https://ark.ap-southeast.bytepluses.com/api/v3) [for Coding Plan
workloads], as requests there bypass Coding Plan quota and incur
separate charges.
— https://docs.byteplus.com/en/docs/ModelArk/1928261
Without this warning visible to users, anyone with a Coding Plan
subscription who picks `provider = "byteplus"` will silently bill
against their per-token USD balance instead of consuming the
subscription quota they paid for.
Adds a prominent "BILLING — READ BEFORE USE" block to byteplus.toml
making the per-token vs subscription split unambiguous, and a
companion note in byteplus-coding.toml pointing back so the choice is
discoverable from either side. Also calls out which capabilities are
unique to the standard endpoint (image / video) so the choice isn't
"just use Coding Plan for everything".
No model definitions or pricing changed.
* chore(byteplus): drop superseded model entries (21 → 11) (#81)
The original byteplus.toml from PR #78 enumerated every BytePlus
ModelArk endpoint that returned HTTP 200, regardless of whether a
sane user would still pick it. Trims to the current per-family
flagship plus useful fast/preview tiers.
Removed (10):
Text:
seed-2-0-lite-260228 — superseded by seed-2-0-mini in the
fast/cheap niche
seed-1-8-251228 — superseded by seed-2-0 family
seed-translation-250915 — too narrow; chat models cover this
deepseek-v3-1-250821 — superseded by deepseek-v3-2
Image:
seedream-3-0-t2i-250415 — superseded by seedream-4-5 / 5-0-lite
seedream-4-0-250828 — same
Video:
seedance-1-0-lite-i2v-250428 — superseded by 1-5 / dreamina-2-0
seedance-1-0-lite-t2v-250428 — same
seedance-1-0-pro-250528 — same
seedance-1-0-pro-fast-251015 — same
Kept (11):
Text (6): seed-2-0-pro/mini/code-preview, glm-4-7,
deepseek-v3-2, gpt-oss-120b
Image (2): seedream-4-5, seedream-5-0 (lite)
Video (3): seedance-1-5-pro, dreamina-seedance-2-0,
dreamina-seedance-2-0-fast
Also corrects two per-piece / per-K price comments that the trim
left orphaned over the wrong [[models]] block (4-5 = $0.0400, not
$0.0300; 1-5-pro = $0.0024/$0.0012 with/without audio, not
$0.0018/K) and updates the header capability list to match the
new model set.
Adds two providers for BytePlus ModelArk, the international edition of
Volcano Engine, distinct from the existing cn-only `volcengine` provider:
- `byteplus`: standard `/api/v3` endpoint, 21 models total
- Text (10): Seed 2.0 Pro/Mini/Lite/Code, Seed 1.8, Seed Translation,
GLM-4.7, DeepSeek V3.2/V3.1, GPT-OSS-120B
- Image (4): Seedream 3.0/4.0/4.5/5.0-lite (per-piece pricing in comments)
- Video (7): Seedance 1.0/1.5 family + Dreamina Seedance 2.0/fast
- `byteplus_coding`: `/api/coding` Anthropic-compatible endpoint with
9 friendly aliases (ark-code-latest auto-router, bytedance-seed-code,
dola-seed-2.0-{pro,code,lite}, kimi-k2.5, glm-4.7, glm-5.1, gpt-oss-120b).
Both use BYTEPLUS_API_KEY env var. All listed model IDs were verified
to return HTTP 200 against the live ap-southeast endpoint. Pricing is the
standard real-time tier from the BytePlus console (a discounted batch tier
exists at roughly half the rate, not modeled here).
Excluded for follow-up (schema doesn't currently support these modalities):
- Skylark embedding-vision (no `embedding` modality in schema)
- Hyper3D-Gen2, Hitem3D-2.0 (no `3d` modality in schema)