Commit Graph
7 Commits
Author SHA1 Message Date
Evan 6785807633 feat(providers): add reasoning_echo_policy field for OpenAI-compat reasoning_content handling (#90)
Refs librefang/librefang#4842 — long-term replacement for the substring
match that the OpenAI driver currently uses to decide how to handle
`reasoning_content` on historical assistant turns.

Three provider-specific behaviours that the driver must distinguish at
wire time, now expressed as catalog metadata:

* `strip`  — DeepSeek R1 / deepseek-reasoner. The API rejects requests
  that carry reasoning_content on previous assistant messages.
* `echo`   — DeepSeek V4 Flash. Thinking mode is on by default and the
  API rejects multi-turn requests when assistant turns containing
  tool_calls don't echo back the original reasoning text. This is the
  bug surfaced in librefang/librefang#4842.
* `empty_string` — Moonshot / Kimi K2 family. The field must be present
  (empty string) on tool_calls turns, with thinking disabled wire-side
  for multi-turn compatibility.
* `none` (default) — most providers; field is omitted entirely.

V4 Pro is intentionally NOT marked `echo` — librefang#4842 reports it
working out-of-the-box; flip when there's an empirical reproducer.

Marks affected models:

  providers/deepseek.toml
    deepseek-v4-flash → echo
    deepseek-reasoner → strip
  providers/moonshot.toml
    kimi-k2.6, kimi-k2.5, kimi-k2 → empty_string
  providers/kimi-coding.toml
    kimi-for-coding → empty_string
  providers/byteplus-coding.toml
    kimi-k2.5 → empty_string
  providers/novita.toml
    moonshotai/kimi-k2-thinking → empty_string

Tooling:

* schema.toml registers the field with the four enum options and a
  `none` default so existing TOML files keep parsing unchanged.
* scripts/validate.py rejects unknown enum values; verified with a
  hand-crafted negative case (`reasoning_echo_policy = "bogus"` →
  validation fails with the expected message).
* `python3 scripts/validate.py` passes (267 models).

The librefang side that consumes this field will land in a follow-up
PR — until then, registry consumers ignore the field via
`#[serde(default)]` and the existing substring fallback continues to
work, so this commit is safe to ship independently.
2026-05-11 00:41:24 +09:00
github-actions[bot] a8854345fb chore: sync model pricing from OpenRouter API 2026-05-01 09:34:41 +00:00
github-actions[bot] b0e0d2e8f4 chore: sync model pricing from OpenRouter API 2026-04-27 08:33:11 +00:00
Evan fb4bebdc1e docs(byteplus_coding): shorten display name to "BytePlus Coding Plan" (#85)
The previous "BytePlus ModelArk Coding Plan" overflows in dashboard
provider cards and CLI lists. Drop the redundant "ModelArk" — it's
already implied by the parent product line, and the file's header
comment + docs explain the relationship.
2026-04-27 16:34:56 +09:00
Evan 95d29be8a6 feat(providers): split _coding api_key_env from main provider (#82)
Each `_coding` provider entry previously declared the same `api_key_env`
as its standard counterpart, causing one credential to silently activate
two providers. Users who only signed up for the per-token API endpoint
saw the Coding Plan endpoint's models mixed into their model picker
without intent.

Rename so each pair uses an independent env var:

  byteplus-coding    BYTEPLUS_API_KEY   -> BYTEPLUS_CODING_API_KEY
  volcengine-coding  VOLCENGINE_API_KEY -> VOLCENGINE_CODING_API_KEY
  zai-coding         ZHIPU_API_KEY      -> ZAI_CODING_API_KEY
  zhipu-coding       ZHIPU_API_KEY      -> ZHIPU_CODING_API_KEY

This is a breaking change: existing users must set the new env vars.
The librefang daemon will be updated separately to recognize the old
env vars as a deprecated fallback during a migration window. See
librefang/librefang#3278 for the design discussion and rollout plan.

Out of scope (different problem class, tracked in the same issue):

  - zai/zhipu cross-region key sharing (both still share ZHIPU_API_KEY)
  - github-copilot/microsoft cross-product GITHUB_TOKEN reuse

Refs librefang/librefang#3278
2026-04-27 15:12:14 +09:00
Evan 23621ef25e docs(byteplus): warn that byteplus is per-token, not Coding Plan (#80)
* docs(byteplus): warn that `byteplus` is per-token, not Coding Plan

Self-followup on PR #78. The BytePlus official docs explicitly warn:

  Do not use the standard model endpoint
  (https://ark.ap-southeast.bytepluses.com/api/v3) [for Coding Plan
  workloads], as requests there bypass Coding Plan quota and incur
  separate charges.
  — https://docs.byteplus.com/en/docs/ModelArk/1928261

Without this warning visible to users, anyone with a Coding Plan
subscription who picks `provider = "byteplus"` will silently bill
against their per-token USD balance instead of consuming the
subscription quota they paid for.

Adds a prominent "BILLING — READ BEFORE USE" block to byteplus.toml
making the per-token vs subscription split unambiguous, and a
companion note in byteplus-coding.toml pointing back so the choice is
discoverable from either side. Also calls out which capabilities are
unique to the standard endpoint (image / video) so the choice isn't
"just use Coding Plan for everything".

No model definitions or pricing changed.

* chore(byteplus): drop superseded model entries (21 → 11) (#81)

The original byteplus.toml from PR #78 enumerated every BytePlus
ModelArk endpoint that returned HTTP 200, regardless of whether a
sane user would still pick it. Trims to the current per-family
flagship plus useful fast/preview tiers.

Removed (10):
  Text:
    seed-2-0-lite-260228       — superseded by seed-2-0-mini in the
                                 fast/cheap niche
    seed-1-8-251228            — superseded by seed-2-0 family
    seed-translation-250915    — too narrow; chat models cover this
    deepseek-v3-1-250821       — superseded by deepseek-v3-2
  Image:
    seedream-3-0-t2i-250415    — superseded by seedream-4-5 / 5-0-lite
    seedream-4-0-250828        — same
  Video:
    seedance-1-0-lite-i2v-250428    — superseded by 1-5 / dreamina-2-0
    seedance-1-0-lite-t2v-250428    — same
    seedance-1-0-pro-250528         — same
    seedance-1-0-pro-fast-251015    — same

Kept (11):
  Text (6):  seed-2-0-pro/mini/code-preview, glm-4-7,
             deepseek-v3-2, gpt-oss-120b
  Image (2): seedream-4-5, seedream-5-0 (lite)
  Video (3): seedance-1-5-pro, dreamina-seedance-2-0,
             dreamina-seedance-2-0-fast

Also corrects two per-piece / per-K price comments that the trim
left orphaned over the wrong [[models]] block (4-5 = $0.0400, not
$0.0300; 1-5-pro = $0.0024/$0.0012 with/without audio, not
$0.0018/K) and updates the header capability list to match the
new model set.
2026-04-27 12:03:40 +09:00
Evan 62bdd04901 feat(providers): add BytePlus ModelArk (international) + coding endpoint (#78)
Adds two providers for BytePlus ModelArk, the international edition of
Volcano Engine, distinct from the existing cn-only `volcengine` provider:

- `byteplus`: standard `/api/v3` endpoint, 21 models total
  - Text (10): Seed 2.0 Pro/Mini/Lite/Code, Seed 1.8, Seed Translation,
    GLM-4.7, DeepSeek V3.2/V3.1, GPT-OSS-120B
  - Image (4): Seedream 3.0/4.0/4.5/5.0-lite (per-piece pricing in comments)
  - Video (7): Seedance 1.0/1.5 family + Dreamina Seedance 2.0/fast
- `byteplus_coding`: `/api/coding` Anthropic-compatible endpoint with
  9 friendly aliases (ark-code-latest auto-router, bytedance-seed-code,
  dola-seed-2.0-{pro,code,lite}, kimi-k2.5, glm-4.7, glm-5.1, gpt-oss-120b).

Both use BYTEPLUS_API_KEY env var. All listed model IDs were verified
to return HTTP 200 against the live ap-southeast endpoint. Pricing is the
standard real-time tier from the BytePlus console (a discounted batch tier
exists at roughly half the rate, not modeled here).

Excluded for follow-up (schema doesn't currently support these modalities):
- Skylark embedding-vision (no `embedding` modality in schema)
- Hyper3D-Gen2, Hitem3D-2.0 (no `3d` modality in schema)
2026-04-27 10:52:57 +09:00