Compare commits

..

326 Commits

Author SHA1 Message Date
Aiden Cline 96f71c533e fix(vercel): sync audio model types 2026-06-22 16:15:50 -05:00
Aiden Cline 4d50c8b588 Merge pull request #2728 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-06-22 11:04:20 -05:00
github-actions[bot] 3792481dd6 chore(sync): update OpenRouter model catalog 2026-06-22 15:07:20 +00:00
Aiden Cline e3300474ee Merge pull request #2699 from quantverse/dev
Add GLM-5.2 for novita-ai provider
2026-06-22 09:29:19 -05:00
Aiden Cline 1e9875220e Merge pull request #2717 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-06-22 09:24:35 -05:00
Aiden Cline b1e39e81d3 Merge pull request #2723 from anomalyco/automation/sync-models-vercel
chore(sync): update Vercel AI Gateway model catalog
2026-06-22 09:24:20 -05:00
github-actions[bot] 0ce055cc57 chore(sync): update OpenRouter model catalog 2026-06-22 13:11:11 +00:00
github-actions[bot] d7ca0b2627 chore(sync): update Vercel AI Gateway model catalog 2026-06-22 13:11:11 +00:00
Karel Vavra d6e9aee388 Add GLM-5.2 for novita-ai provider 2026-06-22 10:24:49 +02:00
Aiden Cline 6421137686 Merge pull request #2716 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-06-21 22:38:48 -05:00
github-actions[bot] d55e91a7ed chore(sync): update OpenRouter model catalog 2026-06-22 03:27:02 +00:00
Aiden Cline 2b9886f76b Merge pull request #2695 from leszek3737/zenmuz-glm52
Zenmux add GLM 5.2 and GLM 5.2 (Free) models
2026-06-21 21:59:17 -05:00
Aiden Cline f027b11048 Merge pull request #2708 from tonimelisma/add-zai-glm-5.2-local
feat(zai): add GLM-5.2 to Z.AI API provider
2026-06-21 21:55:21 -05:00
Aiden Cline 641d790970 Merge pull request #2710 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-06-21 21:50:59 -05:00
Aiden Cline fcebda34d4 Merge pull request #2712 from anomalyco/automation/sync-models-vercel
chore(sync): update Vercel AI Gateway model catalog
2026-06-21 21:50:10 -05:00
github-actions[bot] 3262e4ca30 chore(sync): update Vercel AI Gateway model catalog 2026-06-22 01:30:25 +00:00
github-actions[bot] 656c744502 chore(sync): update OpenRouter model catalog 2026-06-22 01:30:23 +00:00
Toni Melisma ccc1375e0a feat(zai): add GLM-5.2 to Z.AI API provider
Add metered GLM-5.2 for the standard Z.AI API endpoint, matching
zhipuai pricing and reasoning_options and using base_model inheritance
like other zai models.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-06-20 23:43:37 -07:00
Leszek 8065aa3c9f Add reasoning effort options to Zenmux GLM 5.2 models 2026-06-21 01:51:54 +02:00
Aiden Cline 363e0e6f3d Merge pull request #2705 from shzdehmd/dev
chore(firepass): remove the Fireworks AI Firepass provider
2026-06-20 18:00:42 -05:00
Ahmad Shahzad e3758e83d8 chore: remove Firepass provider 2026-06-21 03:52:09 +05:00
Aiden Cline 88046a33d3 Merge pull request #2700 from Tavernari/feat/add-claudius-model
chore(sync): add claudius model with audio/video input + add audio/video to claudinio
2026-06-20 16:12:51 -05:00
Aiden Cline 950b283605 Merge pull request #2701 from anomalyco/automation/sync-models-vercel
chore(sync): update Vercel AI Gateway model catalog
2026-06-20 16:12:27 -05:00
Aiden Cline 685635d0ca Merge pull request #2702 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-06-20 16:12:14 -05:00
Aiden Cline edf9c72753 Merge pull request #2703 from anomalyco/automation/sync-models-venice
chore(sync): update Venice model catalog
2026-06-20 16:12:03 -05:00
github-actions[bot] 3ad03b27c9 chore(sync): update Vercel AI Gateway model catalog 2026-06-20 20:44:46 +00:00
github-actions[bot] 2cb040b524 chore(sync): update OpenRouter model catalog 2026-06-20 20:44:45 +00:00
github-actions[bot] e8cab955f6 chore(sync): update Venice model catalog 2026-06-20 20:44:43 +00:00
Victor Carvalho Tavernari f759c801f6 feat: add audio+video input modalities to claudinio and claudius 2026-06-20 21:30:14 +01:00
Victor Carvalho Tavernari 539f58605e fix: inline claudius model instead of extends to fix CI validation 2026-06-20 21:28:24 +01:00
Victor Carvalho Tavernari ed3264b049 feat: add claudius model extending claudinio with / pricing 2026-06-20 20:04:53 +01:00
Aiden Cline e3df94e9a1 Merge pull request #2623 from smakosh/feat/llmgateway-add-gemma4-kimi-highspeed-qwen35-glm52
feat: add LLM Gateway gemma-4, kimi-k2.7-code-highspeed, qwen3.5-9b, glm-5.2
2026-06-20 14:53:38 -04:00
Aiden Cline 67b48ce993 Merge pull request #2679 from mitjap/remove-cortecs-devstral-small-2512
remove deprecated model cortecs/devstral-small-2512
2026-06-20 14:47:11 -04:00
Aiden Cline 9a0e70541c Merge pull request #2694 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-06-20 14:45:28 -04:00
Aiden Cline cce974a4ad Merge pull request #2698 from howmanysmall/feat/crof-deepseek-v4-pro-lightning-name
feat(crof): name DeepSeek V4 Pro Lightning
2026-06-20 14:43:36 -04:00
github-actions[bot] 1e9d80827c chore(sync): update OpenRouter model catalog 2026-06-20 17:47:50 +00:00
howmanysmall 12c5588b63 feat(models): add name to crof DeepSeek V4 Pro Lightning model
This helps distinguish it from the cheaper DeepSeek V4 Pro on crof
2026-06-19 22:35:32 -06:00
Leszek f2e9ca7166 Update Zenmux GLM 5.2 base model provider alias 2026-06-20 02:01:57 +02:00
Leszek 7e36f36eae Zenmux add GLM 5.2 and GLM 5.2 (Free) models
Adds configurations for the GLM 5.2 model under the Zenmux provider, including a distinct free-tier variant with zero cost.
2026-06-20 01:59:52 +02:00
Aiden Cline 28d4dbbd4c Merge pull request #2693 from anomalyco/automation/sync-models-venice
chore(sync): update Venice model catalog
2026-06-19 19:08:39 -04:00
Aiden Cline fbac01e55b Merge pull request #2584 from anomalyco/automation/sync-models-google
chore(sync): update Google model catalog
2026-06-19 19:08:26 -04:00
Aiden Cline 5fb63a5fab Merge pull request #2646 from BlockListed/cortecs-add-glm-5-2-kimi-k2-7
Cortecs add glm 5.2 and kimi k2.7
2026-06-19 18:36:50 -04:00
github-actions[bot] 32aaa20233 chore(sync): update Venice model catalog 2026-06-19 22:35:02 +00:00
github-actions[bot] 828e41d9fc chore(sync): update Google model catalog 2026-06-19 22:34:58 +00:00
Aiden Cline bca5c31ce3 Merge pull request #2644 from anomalyco/automation/sync-models-vercel
chore(sync): update Vercel AI Gateway model catalog
2026-06-19 18:33:13 -04:00
Aiden Cline c02d341c32 Merge pull request #2669 from anomalyco/automation/sync-models-venice
chore(sync): update Venice model catalog
2026-06-19 18:32:41 -04:00
Aiden Cline 5df780da5d Merge pull request #2668 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-06-19 18:32:35 -04:00
Aiden Cline 6462c50ac4 Merge pull request #2680 from jcraftsman/umans-curated-reasoning-options
umans-ai: curate reasoning_options to match each model's real reasoning support
2026-06-19 18:20:13 -04:00
Aiden Cline ff9ba6944e Merge pull request #2684 from Omee11/feat/token-plan-glm5.2-kimi-k2.7-code
feat(alibaba-token-plan): add glm-5.2 and kimi-k2.7-code
2026-06-19 18:19:40 -04:00
Aiden Cline e3d50a0856 Merge pull request #2686 from patrik-kuehl/mark-glm-5.2-as-open-weighted
chore(models): mark GLM 5.2 as open-weighted
2026-06-19 18:16:55 -04:00
Aiden Cline 8c56ecaed4 Merge pull request #2683 from skyitachi/add-siliconflow-glm-5.2
feat(siliconflow): add zai-org/GLM-5.2
2026-06-19 18:16:14 -04:00
github-actions[bot] f47485ca2a chore(sync): update Vercel AI Gateway model catalog 2026-06-19 21:41:53 +00:00
github-actions[bot] 5a665a775e chore(sync): update OpenRouter model catalog 2026-06-19 21:41:52 +00:00
github-actions[bot] 35e34181ea chore(sync): update Venice model catalog 2026-06-19 21:41:50 +00:00
BlockListed 20cce673df add kimi k2.7 to cortecs 2026-06-19 10:29:55 +02:00
BlockListed 29ae2fac59 add glm 5.2 to cortecs 2026-06-19 10:29:51 +02:00
Patrik Kühl 0705837fc7 chore(models): mark GLM 5.2 as open-weighted 2026-06-19 09:27:22 +02:00
Oliver Mee 5cd309aa57 feat(alibaba-token-plan): add glm-5.2 and kimi-k2.7-code
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-19 12:51:34 +08:00
skyitachi 3cc1c1861a feat: add zai-org/GLM-5.2 to siliconflow and siliconflow-cn 2026-06-19 11:14:44 +08:00
wassel alazhar c03cb0723d umans-ai: curate reasoning_options to match each model's real support
Align the umans-ai and umans-ai-coding-plan reasoning_options with the
levels each model actually exposes via the Umans gateway:

- GLM 5.1: toggle only (reasoning is on/off; effort is not meaningful)
- GLM 5.2: toggle + effort high/max (only high/max are real levels)
- Umans Coder / Kimi K2.7: [] (always-on; no toggle, no effort tiers)

Flash and the Qwen alias are unchanged (off + low/medium/high).
2026-06-18 23:28:33 +02:00
Mitja Puzigaća fe73d94598 remove deprecated model cortecs/devstral-small-2512 2026-06-18 19:58:45 +02:00
Aiden Cline a571899ac0 Merge pull request #2670 from xhml-tangf/dev
feat: add GLM-5.2 to Zhipu AI provider models
2026-06-18 14:09:15 +02:00
Edward d00ab87a87 feat: add GLM-5.2 to Zhipu AI provider models 2026-06-18 19:50:33 +08:00
Aiden Cline 760fbc07d6 Merge pull request #2665 from jcraftsman/feat/umans-coding-plan-coder-glm52-reasoning
umans-ai + umans-ai-coding-plan: repoint coder to K2.7, add GLM 5.2, drop K2.6, normalise reasoning
2026-06-18 12:36:17 +02:00
Aiden Cline 8bc9cfafe3 Merge pull request #2289 from anomalyco/split/opencode-anthropic-reasoning-options
[opencode/anthropic] Add reasoning options
2026-06-18 12:21:42 +02:00
Aiden Cline 67c8d0971e Merge pull request #2293 from anomalyco/split/opencode-minimax-reasoning-options
[opencode/minimax] Add reasoning options
2026-06-18 12:21:05 +02:00
Aiden Cline 6da1466c7c Merge pull request #2467 from anomalyco/split/nano-gpt-moonshotai-reasoning-options
[nano-gpt/moonshotai] Add reasoning options
2026-06-18 12:20:05 +02:00
Aiden Cline bac480d051 Merge pull request #2483 from anomalyco/split/nano-gpt-x-ai-reasoning-options
[nano-gpt/x-ai] Add reasoning options
2026-06-18 12:19:46 +02:00
Aiden Cline 4f254bda4e Merge pull request #2511 from anomalyco/split/siliconflow-tencent-reasoning-options
[siliconflow/tencent] Add reasoning options
2026-06-18 12:19:21 +02:00
Aiden Cline 633540f206 Merge pull request #2512 from anomalyco/split/siliconflow-thudm-reasoning-options
[siliconflow/THUDM] Add reasoning options
2026-06-18 12:18:56 +02:00
wassel alazhar 964bf76999 umans-ai + coding-plan: repoint coder to K2.7, add GLM 5.2, drop K2.6, normalise reasoning
Brings both umans providers in line with what umans.ai serves today, with identical
model structure across them. Per-token [cost] lives on the pay-by-token provider
(umans-ai) only; the coding plan is a flat subscription, so its models stay at [cost] = 0.

Both providers (umans-ai and umans-ai-coding-plan):
- umans-coder: base_model -> moonshotai/kimi-k2.7-code (inherits the kimi-k2 family).
  Always reasons, so it exposes effort levels only (no on/off toggle).
- add umans-glm-5.2 (reasoning toggle + effort, 405504 context).
- drop umans-kimi-k2.6 (no longer published in the catalogue).
- reasoning_options: effort (low/medium/high) everywhere; the on/off toggle is kept only
  on models that can disable reasoning (flash, glm-5.1, glm-5.2, qwen3.6-35b-a3b).
  kimi-k2.7 and coder always reason, so no toggle.

Pricing (umans-ai / pay-by-token only, $/M in / out / cache-read):
    umans-coder, umans-kimi-k2.7        0.95 / 4.00 / 0.19
    umans-glm-5.2                       1.40 / 4.40 / 0.26
    umans-glm-5.1                       1.40 / 4.40 / 0.29
    umans-flash                         0.15 / 1.00 / 0.05
umans-ai-coding-plan keeps [cost] = 0 (subscription, no per-token charge).
2026-06-18 12:15:09 +02:00
Aiden Cline a9b9f2998e Merge pull request #2652 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-06-18 12:14:17 +02:00
Aiden Cline 689126a1a2 Merge pull request #2526 from anomalyco/split/vercel-minimax-reasoning-options
[vercel/minimax] Add reasoning options
2026-06-18 12:13:46 +02:00
Aiden Cline 5668077eae Merge pull request #2567 from anomalyco/consolidate/github-copilot-google-router-reasoning-options
[github-copilot/google] Add reasoning options and remove Raptor Mini
2026-06-18 12:11:58 +02:00
Aiden Cline 09a783c54f chore(github-copilot): remove Raptor Mini 2026-06-18 12:11:08 +02:00
Aiden Cline 8d635f97a8 fix(github-copilot): complete Google reasoning options 2026-06-18 12:06:32 +02:00
Aiden Cline 02fc312c05 Merge pull request #2656 from anomalyco/automation/sync-models-venice
chore(sync): update Venice model catalog
2026-06-18 12:00:47 +02:00
Aiden Cline 2780a25242 Merge pull request #2658 from InfHorus/dev
Add Latest LucidQuery models
2026-06-18 12:00:37 +02:00
Aiden Cline 0f26261671 Merge pull request #2655 from KTibow/automation/sync-models-crof
chore(sync): update CrofAI model catalog
2026-06-18 12:00:16 +02:00
Aiden Cline 684b1f37aa Merge pull request #2654 from RISHIKREDDYL/fix/azure-cognitive-services-env-var
fix(azure-cognitive-services): correct env var in kimi model API URLs
2026-06-18 11:59:28 +02:00
Aiden Cline c0341f0ce8 refactor(azure-cognitive-services): use Kimi base models 2026-06-18 11:57:15 +02:00
Aiden Cline 389fefa6df Merge pull request #2648 from houtanb/dev
Use base model metadata for GLM-5 and GLM-5.1, and fix release dates
2026-06-18 11:54:42 +02:00
Aiden Cline d78e43f537 Merge pull request #2633 from anomalyco/automation/sync-models-baseten
chore(sync): update Baseten model catalog
2026-06-18 11:51:49 +02:00
github-actions[bot] dd725fcb4b chore(sync): update Baseten model catalog 2026-06-18 09:47:33 +00:00
github-actions[bot] dc85ae0999 chore(sync): update OpenRouter model catalog 2026-06-18 09:47:31 +00:00
github-actions[bot] 4db533d460 chore(sync): update Venice model catalog 2026-06-18 09:47:30 +00:00
Aiden Cline abcb2424ca Merge pull request #2659 from thehaseebahmed/azure-gpt-image-models
feat(azure): add gpt-image-1, 1.5, and 2 models with pricing
2026-06-18 11:46:32 +02:00
Aiden Cline 0248ace087 Merge pull request #2627 from JoshuaDietz/dev
feat(ollama-cloud): add glm-5.2
2026-06-18 11:45:33 +02:00
Aiden Cline fc05522afe feat(openai): add GPT Image 2 2026-06-18 11:43:28 +02:00
Aiden Cline 5ae1dc5ff8 fix(azure): use base models for GPT Image 2026-06-18 11:40:46 +02:00
Aiden Cline 13e826f763 Merge pull request #2666 from heimoshuiyu/add-alibaba-token-plan-cn-glm-5.2
feat(alibaba-token-plan-cn): add GLM-5.2 model
2026-06-18 11:37:09 +02:00
heimoshuiyu e543afc6cf feat(alibaba-token-plan-cn): add GLM-5.2 model 2026-06-18 17:27:19 +08:00
Haseeb Ahmed c0b530099b feat(azure): add gpt-image-1, 1.5, and 2 models with pricing 2026-06-18 01:24:33 +02:00
InfHorus 7da1e391f4 Add 'agi' to the family list 2026-06-18 00:51:29 +02:00
InfHorus 441920b865 Add support for LucidQuery AGI-01 family 2026-06-18 00:39:31 +02:00
InfHorus b81c4c47fd Update LucidQuery API 2026-06-18 00:28:41 +02:00
KTibow a671ff0c50 chore(sync): update CrofAI model catalog 2026-06-17 14:18:03 -07:00
RISHIKREDDYL 87114fccb6 Fix incorrect env var in Azure Cognitive Services kimi models
The kimi-k2.5.toml and kimi-k2.6.toml files in azure-cognitive-services used
AZURE_RESOURCE_NAME in their API URLs, but the provider declares
AZURE_COGNITIVE_SERVICES_RESOURCE_NAME as the expected environment variable.

Changes:
- kimi-k2.5.toml: converted from symlink (pointing to azure/models/) to
  standalone real file with the corrected env var
- kimi-k2.6.toml: replaced AZURE_RESOURCE_NAME with
  AZURE_COGNITIVE_SERVICES_RESOURCE_NAME in the API URL

This matches the pattern used by other models with provider overrides in
azure-cognitive-services (e.g. claude-haiku-4-5, claude-opus-4-1, etc.).
2026-06-17 23:30:08 +05:30
Houtan Bastani 2bed70cca9 Use base model metadata for GLM-5 and GLM-5.1, and fix release dates
glm-5 release date: https://docs.z.ai/release-notes/new-released?utm_source=chatgpt.com#2026-02-12
glm-5.1 release date: https://docs.z.ai/release-notes/new-released?utm_source=chatgpt.com#2026-04-07
2026-06-17 13:57:11 +02:00
Frank 3f537855c3 update go models 2026-06-17 13:22:27 +02:00
Aiden Cline 8f5ae25daf Merge pull request #2645 from monotykamary/neuralwatt-glm-5-2-reasoning-efforts
feat(neuralwatt): expose full GLM 5.2 reasoning effort scale
2026-06-17 13:05:36 +02:00
Tom X Nguyen c5b3973a25 feat(neuralwatt): expose full GLM 5.2 reasoning effort scale
GLM-5.2 accepts the OpenAI-standard reasoning_effort field and supports
a wider depth range than the three levels previously advertised. Per
the Neuralwatt chat-completions docs [1], the gateway accepts and
normalizes the full scale:

  minimal -> skips the reasoning phase entirely (eq enable_thinking: false)
  low     -> mapped to high
  medium  -> mapped to high
  high    -> enhanced reasoning (balanced)
  xhigh   -> mapped to max (deepest; best for math/planning/agentic tasks)

The provider's thinkingLevelMap (pi-neuralwatt-provider/patch.json) already
exposes all five pi tiers, so mirror that here by adding minimal and xhigh
to the effort values for glm-5.2.

[1] https://portal.neuralwatt.com/docs/api/chat-completions
2026-06-17 17:57:22 +07:00
Joshua Dietz 5bf8d5a2c4 fix reasoning options
I'm unsure about the possible values, but the zai-coding-plan version uses the same high/max options that I've added now. This seems to be confirmed by https://huggingface.co/zai-org/GLM-5.2/discussions/1
2026-06-17 12:50:28 +02:00
Aiden Cline 553cde57ce Merge pull request #2563 from anomalyco/consolidate/aihubmix-small-labs-reasoning-options
[aihubmix/multiple labs] Add reasoning options
2026-06-17 12:26:46 +02:00
Aiden Cline 5094f20a1a Merge pull request #2529 from anomalyco/split/vercel-nvidia-reasoning-options
[vercel/nvidia] Add reasoning options
2026-06-17 12:26:29 +02:00
Aiden Cline 2d77e101e0 Merge pull request #2643 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-06-17 12:24:43 +02:00
Aiden Cline be1481eb27 Delete providers/openrouter/models/anthropic/claude-fable-5.toml 2026-06-17 12:24:28 +02:00
Aiden Cline 39f2a40a75 Merge pull request #2642 from anomalyco/fix/openrouter-blacklist-fable-5
fix(openrouter): blacklist Fable 5 models
2026-06-17 12:24:12 +02:00
Aiden Cline 5e9701a219 test: remove sync test suites 2026-06-17 12:18:19 +02:00
github-actions[bot] 3d0fbe7f20 chore(sync): update OpenRouter model catalog 2026-06-17 09:55:47 +00:00
Aiden Cline a8dd73ac1e fix(openrouter): blacklist Fable 5 models 2026-06-17 11:44:29 +02:00
Aiden Cline 96518b7942 Merge pull request #2615 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-06-17 05:42:14 -04:00
Aiden Cline ab2e1ed68d Delete providers/openrouter/models/anthropic/claude-fable-5.toml 2026-06-17 11:42:04 +02:00
Aiden Cline f8abb306ee Merge pull request #2640 from anomalyco/audit/vercel-mistral-small-capability
[vercel/mistral] Correct Mistral Small reasoning
2026-06-17 05:40:33 -04:00
Aiden Cline 6f3fb7e69a Merge pull request #2639 from anomalyco/audit/vercel-alibaba-openai-followup
[vercel/alibaba openai] Add tested reasoning controls
2026-06-17 05:40:16 -04:00
Aiden Cline 1309db94e6 [vercel/mistral] Correct Mistral Small reasoning 2026-06-17 11:30:40 +02:00
Joshua Dietz dc07908c6e implement PR feedback 2026-06-17 11:13:10 +02:00
github-actions[bot] 571188e70a chore(sync): update OpenRouter model catalog 2026-06-17 08:13:08 +00:00
Aiden Cline dee4628f3e Merge pull request #2636 from monotykamary/add-neuralwatt-glm-5-2
feat(neuralwatt): add GLM 5.2 and retire MiniMax M2.5, Devstral, GPT OSS 20B
2026-06-17 03:06:33 -04:00
Tom X Nguyen 34bbbc4b5f feat(neuralwatt): add GLM 5.2 and retire MiniMax M2.5, Devstral, GPT OSS 20B
Sync neuralwatt provider with the current Neuralwatt API data (from
../pi-neuralwatt-provider: models.json -> patch.json -> custom-models.json).

Added:
- glm-5.2: GLM 5.2 (family glm, 1_048_560 context/output, 1.45/4.5 cost,
  reasoning via effort [low,medium,high] — provider sets
  supportsReasoningEffort with no reasoning_content interleaving)

Removed (no longer in the provider API):
- MiniMaxAI/MiniMax-M2.5.toml
- mistralai/Devstral-Small-2-24B-Instruct-2512.toml
- openai/gpt-oss-20b.toml

README: added GLM 5.2 to the reasoning list; dropped the MiniMax M2.5,
GPT OSS 20B lines and the now-empty Devstral section.

opus/flex/long and canary variants excluded by request.
2026-06-17 11:32:01 +07:00
Aiden Cline 0eba09c28e Merge pull request #2634 from shzdehmd/dev
feat(fireworks-ai): add GLM-5.2 and fix Kimi/DeepSeek/GPT pricing
2026-06-17 00:17:07 -04:00
Ahmad Shahzad b4bf6468b4 feat(fireworks-ai): add GLM-5.2 and fix Kimi/DeepSeek/GPT pricing
- Add GLM-5.2 (accounts/fireworks/models/glm-5p2) with 1M context and

  Fireworks serverless pricing ($1.40 / $0.26 / $4.40).

- Normalize Kimi K2.7 Code and Kimi K2.7 Code Fast TOML files to be

  self-contained and follow the same metadata pattern as Kimi K2.6.

- Fix Kimi K2.7 Code Fast input price ($2.00 -> $1.90).

- Fix DeepSeek V4 Flash cache read price ($0.03 -> $0.028).

- Fix GPT OSS 120B cache read price ($0.01 -> $0.015).

- Set last_updated to 2026-06-16 for all touched provider files.
2026-06-17 08:13:52 +05:00
Aiden Cline eb89d9b2ad Merge pull request #2624 from anomalyco/automation/sync-models-venice
chore(sync): update Venice model catalog
2026-06-16 19:41:08 -04:00
Aiden Cline 89857613f2 Merge pull request #2629 from pat-baseten/add-glm-5.2-baseten
Add GLM-5.2 to Baseten provider
2026-06-16 19:09:00 -04:00
Aiden Cline 30cff10862 Merge pull request #2630 from pat-baseten/fix-kimi-k2.7-code-pricing-baseten
Fix Kimi K2.7 Code pricing for Baseten provider
2026-06-16 19:08:42 -04:00
github-actions[bot] eec148d635 chore(sync): update Venice model catalog 2026-06-16 22:54:48 +00:00
Pat c2d5869400 Fix Kimi K2.7 Code pricing for Baseten provider
Correct input and cache read costs to match published Baseten pricing
($0.95 input / $0.16 cached input / $4.00 output per 1M tokens).
2026-06-16 14:34:59 -07:00
Pat caa20e6a67 Add GLM-5.2 pricing from Baseten Model APIs
Set input, cache read, and output costs to match the published
Baseten pricing page ($1.50 / $0.30 / $4.50 per 1M tokens).
2026-06-16 14:34:57 -07:00
Pat c85b741815 Add GLM-5.2 to Baseten provider
Configure Baseten serving metadata for zai-org/GLM-5.2 using the
zhipuai/glm-5.2 base model. Limits and reasoning options are sourced
from the Baseten Model APIs catalog; cost is omitted until pricing is
published in the /v1/models endpoint.
2026-06-16 14:34:57 -07:00
Joshua Dietz 4fce8b4df6 feat(ollama-cloud): add glm-5.2 2026-06-16 22:37:10 +02:00
Aiden Cline 2655f319f0 Merge pull request #2622 from cline/saoudrizwan/add-openrouter-glm-5.2
feat: add z-ai/glm-5.2 model on OpenRouter
2026-06-16 14:48:14 -04:00
smakosh 805aababcb feat: add LLM Gateway gemma-4, kimi-k2.7-code-highspeed, qwen3.5-9b, glm-5.2
Add provider entries for newly available LLM Gateway text models:
- gemma-4-31b-it, gemma-4-26b-a4b-it (Google, reasoning)
- kimi-k2.7-code-highspeed (Moonshot, highspeed tier of kimi-k2.7-code)
- qwen3.5-9b (Alibaba)
- glm-5.2 (Z.AI)

Adds base model metadata for kimi-k2.7-code-highspeed and qwen3.5-9b.
Pricing for gemma/kimi/qwen taken from the api.llmgateway.io catalog;
glm-5.2 pricing from the Z.AI docs (input $1.4, cache_read $0.26, output $4.4).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-16 20:10:59 +02:00
Saoud Rizwan d8e8b4ec82 feat: add z-ai/glm-5.2 model on OpenRouter 2026-06-16 11:00:27 -07:00
Aiden Cline cbe5e319dd Merge pull request #2619 from anomalyco/automation/sync-models-vercel
chore(sync): update Vercel AI Gateway model catalog
2026-06-16 13:06:44 -04:00
Aiden Cline b6c8f645fb Merge pull request #2620 from anomalyco/automation/sync-models-cloudflare-workers-ai
chore(sync): update Cloudflare Workers AI model catalog
2026-06-16 13:06:19 -04:00
github-actions[bot] fb62ddc484 chore(sync): update Cloudflare Workers AI model catalog 2026-06-16 16:56:06 +00:00
github-actions[bot] 0f41070eb8 chore(sync): update Vercel AI Gateway model catalog 2026-06-16 16:55:59 +00:00
Aiden Cline 484985736c Merge pull request #2565 from anomalyco/consolidate/cortecs-small-labs-reasoning-options
[cortecs/multiple labs] Add reasoning options
2026-06-16 12:31:28 -04:00
Aiden Cline 1f43cb15ef Merge pull request #2608 from anomalyco/audit/vercel-other-labs-reasoning-options
[vercel/multiple labs] Add verified reasoning options
2026-06-16 12:31:07 -04:00
Aiden Cline e6b8ec45e1 Merge pull request #2617 from anomalyco/automation/sync-models-baseten
chore(sync): update Baseten model catalog
2026-06-16 12:30:47 -04:00
Aiden Cline 2d12b0d3ca Merge pull request #2609 from anomalyco/audit/vercel-xai-reasoning-options
[vercel/xai] Add verified reasoning options
2026-06-16 12:25:04 -04:00
github-actions[bot] b44440b6af chore(sync): update Baseten model catalog 2026-06-16 15:06:04 +00:00
Aiden Cline a53102dc3c [vercel/alibaba] Add tested Qwen 3.7 Plus budget 2026-06-16 16:28:45 +02:00
Aiden Cline 98ccd21e83 [vercel/openai] Correct tested effort ranges 2026-06-16 16:17:43 +02:00
Aiden Cline 50c93de146 [vercel/openai] Add tested chat model efforts 2026-06-16 16:15:11 +02:00
Aiden Cline f0c8295802 [vercel/xai] Remove ineffective Grok none effort 2026-06-16 16:12:52 +02:00
Aiden Cline 2f9470a3b1 Merge pull request #2611 from anomalyco/audit/vercel-alibaba-openai-followup
[vercel/alibaba openai] Complete reasoning audit
2026-06-16 10:08:33 -04:00
Aiden Cline efe09a008e Merge pull request #2610 from anomalyco/audit/vercel-zai-reasoning-options
[vercel/zai] Add reasoning toggles
2026-06-16 10:08:18 -04:00
Aiden Cline a2a0de474b Merge pull request #2612 from anomalyco/audit/vercel-capability-reconciliation
[vercel] Reconcile reasoning capabilities
2026-06-16 07:41:52 -04:00
Aiden Cline cfe25d7eb2 Merge pull request #2595 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-06-16 07:00:33 -04:00
Aiden Cline 625519a252 Delete providers/openrouter/models/anthropic/claude-fable-5.toml 2026-06-16 12:59:59 +02:00
Aiden Cline a4a0e09c88 Merge pull request #2596 from anomalyco/automation/sync-models-vercel
chore(sync): update Vercel AI Gateway model catalog
2026-06-16 06:58:22 -04:00
github-actions[bot] a91f972ac7 chore(sync): update OpenRouter model catalog 2026-06-16 10:57:27 +00:00
github-actions[bot] ff652bba52 chore(sync): update Vercel AI Gateway model catalog 2026-06-16 10:57:27 +00:00
Aiden Cline abb6c053f4 Merge pull request #2613 from anomalyco/feat/moonshot-kimi-k2.7-code-highspeed
feat(moonshotai): add Kimi K2.7 Code HighSpeed
2026-06-16 06:40:56 -04:00
Aiden Cline 837d9f414e feat(moonshotai): add Kimi K2.7 Code HighSpeed 2026-06-16 12:39:59 +02:00
Aiden Cline 4358b05cac Merge pull request #2593 from houtanb/dev
Reuse base model metadata for Gemini and Mistral provider entries
2026-06-16 06:25:17 -04:00
Aiden Cline d0089030e3 Merge pull request #2585 from cline/saoudrizwan/remove-openrouter-fable-5
chore: remove Claude Fable 5 from OpenRouter
2026-06-16 06:24:53 -04:00
Aiden Cline e6ee64384c Merge pull request #2594 from hqrrr/moonshotai-cn-kimi-k2.7-code
[moonshotai-cn] Add kimi-k2.7-code.toml symlink
2026-06-16 06:24:33 -04:00
Aiden Cline a1d7729b1c Merge pull request #2597 from SvanBoxel/patch-1
Update context lengths for Poolside Laguna models
2026-06-16 06:24:14 -04:00
Aiden Cline d20ce82097 Merge pull request #2598 from anomalyco/automation/sync-models-venice
chore(sync): update Venice model catalog
2026-06-16 06:23:44 -04:00
Aiden Cline a6cd9749ef Merge pull request #2600 from oskarkocol/feat/update-togetherai-prices
chore: update TogetherAI prices 20260615
2026-06-16 06:23:35 -04:00
Aiden Cline 909041d729 Merge pull request #2590 from nikosch86/add/cortecs-minimax-m3
add llama-4-maverick and minimax-m3 to cortecs
2026-06-16 06:23:22 -04:00
Aiden Cline 118828ca54 [vercel] Reconcile reasoning capabilities 2026-06-16 12:20:05 +02:00
Aiden Cline f76e595b63 Merge remote-tracking branch 'origin/dev' into audit/vercel-other-labs-reasoning-options
# Conflicts:
#	providers/vercel/models/bytedance/seed-1.6.toml
#	providers/vercel/models/bytedance/seed-1.8.toml
#	providers/vercel/models/mistral/mistral-medium-3.5.toml
#	providers/vercel/models/perplexity/sonar-reasoning-pro.toml
#	providers/vercel/models/stepfun/step-3.5-flash.toml
2026-06-16 12:18:26 +02:00
Aiden Cline 1ea9929364 Merge remote-tracking branch 'origin/dev' into audit/vercel-xai-reasoning-options
# Conflicts:
#	providers/vercel/models/xai/grok-4.1-fast-reasoning.toml
#	providers/vercel/models/xai/grok-4.20-multi-agent-beta.toml
#	providers/vercel/models/xai/grok-4.20-multi-agent.toml
#	providers/vercel/models/xai/grok-4.20-reasoning-beta.toml
#	providers/vercel/models/xai/grok-4.20-reasoning.toml
#	providers/vercel/models/xai/grok-4.3.toml
2026-06-16 12:18:00 +02:00
Aiden Cline b0cfceeff3 [vercel/alibaba openai] Complete reasoning audit 2026-06-16 12:16:34 +02:00
Aiden Cline 05db497663 [vercel/multiple labs] Add verified reasoning options 2026-06-16 12:16:28 +02:00
Aiden Cline ae393aaca8 [vercel/zai] Add reasoning toggles 2026-06-16 12:16:20 +02:00
Aiden Cline 82ddea90f3 [vercel/xai] Add verified reasoning options 2026-06-16 12:16:14 +02:00
Aiden Cline f3070c436e Merge pull request #2601 from oskarkocol/chore/update-stepfun-20260615
chore: update stepai prices 20260615
2026-06-16 06:14:52 -04:00
Aiden Cline 3f0df86ec4 Merge pull request #2599 from maxlang/update-ambient-glm51-kimi-k27
chore(ambient): add Kimi K2.7 Code, refresh GLM 5.1
2026-06-16 06:13:13 -04:00
Aiden Cline 43e1010e1e [vercel/minimax] Add M3 reasoning toggle 2026-06-16 12:12:32 +02:00
Aiden Cline 0886fc4e11 Merge pull request #2602 from oskarkocol/chore/update-siliconflow-20260615
chore: update siliconflow prices 20260615
2026-06-16 05:49:17 -04:00
Aiden Cline 7a7276123a Merge pull request #2603 from oskarkocol/chore/update-novitaai-20260615
chore: update novita pricing 20260615
2026-06-16 05:46:26 -04:00
Aiden Cline 374135b350 Merge pull request #2606 from JDinABox/dev
Add Neuralwatt Kimi K2.7 Code model configuration
2026-06-16 05:46:00 -04:00
Aiden Cline 87ba6613d2 Merge pull request #2605 from oskarkocol/chore/update-fireworks-20260615
chore: update fireworks pricing 20260615
2026-06-16 05:45:45 -04:00
Aiden Cline 57d1b2489a Merge pull request #2607 from BlockListed/cortecs-add-glm-5v
add glm-5*-turbo to cortecs
2026-06-16 05:45:23 -04:00
Aiden Cline ee243e06b4 Merge pull request #2591 from vglafirov/remove-gitlab-fable-5
Remove GitLab Duo Chat Fable 5 model
2026-06-16 11:29:35 +02:00
github-actions[bot] d7f8f4f40a chore(sync): update Venice model catalog 2026-06-16 08:20:39 +00:00
BlockListed e35a772633 add glm-5*-turbo to cortecs 2026-06-16 10:04:51 +02:00
JD Crawford 20056e2c02 feat(neuralwatt): add Kimi K2.7 Code model support 2026-06-16 03:38:30 -04:00
oskar 45c6ab5999 update the last_updated date 2026-06-16 13:54:51 +07:00
oskar 99c9baa635 update fireworks pricing 2026-06-16 13:51:35 +07:00
oskar 0e8c0b79f3 update novita pricing 2026-06-16 13:36:24 +07:00
oskar d99ba71ad0 update siliconflow models 2026-06-16 13:03:40 +07:00
oskar 2dce213dbd chore: update stepai prices 2026-06-16 12:17:43 +07:00
oskar 79389f68b7 chore: update last_updated 2026-06-16 11:39:38 +07:00
oskar cd95e58488 update togetherai prices 2026-06-16 11:35:44 +07:00
Max Lang 623ab9c61c chore(ambient): add Kimi K2.7 Code, refresh GLM 5.1
Update the Ambient catalog for two models from the live
api.ambient.xyz/v1/models endpoint:

- add moonshotai/kimi-k2.7-code (base_model: moonshotai/kimi-k2.7-code)
- refresh zai-org/GLM-5.1-FP8 display name

Both inherit canonical metadata via base_model and override only the
fields Ambient's API reports (pricing, capabilities).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-15 14:25:02 -07:00
Claude ffec078bbb Set output token limit to 32768 for Laguna M.1 and XS.2 2026-06-15 20:30:15 +00:00
Sebass van Boxel 03d0d79709 Update context limit foe XS.2 in kilo 2026-06-15 21:47:34 +02:00
Sebass van Boxel 987800ea87 Update and context limit for laguna m1 in kilo 2026-06-15 21:47:06 +02:00
Sebass van Boxel 555ca498e7 Update last_updated date and context limit for laguna m.1 2026-06-15 21:40:34 +02:00
Sebass van Boxel 37eacd2574 Update last_updated date and context limit for laguna.xs2 2026-06-15 21:39:08 +02:00
hqr 800e7404ef [moonshotai-cn] Add kimi-k2.7-code.toml symlink
Link providers/moonshotai-cn/models/kimi-k2.7-code.toml to providers/moonshotai/models/kimi-k2.7-code.toml.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-06-15 18:21:31 +02:00
Houtan Bastani f5c3437d74 Reuse base model metadata for Gemini and Mistral provider entries
Replace duplicated provider-agnostic metadata with base_model references for `Gemini 2.5 Flash`, `Gemini 2.5 Pro`, and `mistral-large-2411`.

Follow on to 5a8f9d4, 61a153e and PR #2251
2026-06-15 17:36:18 +02:00
Vladimir Glafirov b4c236a4b3 Remove GitLab Duo Chat Fable 5 model 2026-06-15 15:07:34 +02:00
Niko 6552f7489f add llama-4-maverick and minimax-m3 to cortecs 2026-06-15 15:13:29 +04:00
Saoud Rizwan 28dcf895db chore: remove Claude Fable 5 from OpenRouter 2026-06-14 21:38:52 -07:00
Aiden Cline 351541ea2c Merge pull request #2566 from anomalyco/consolidate/frogbot-small-labs-reasoning-options
[frogbot/multiple labs] Add reasoning options
2026-06-14 21:24:59 -05:00
Aiden Cline a1f3591660 [frogbot] Remove options from non-reasoning models 2026-06-14 22:23:55 -04:00
Aiden Cline a3c3e97556 Merge pull request #2568 from anomalyco/consolidate/github-models-small-labs-reasoning-options
[github-models/multiple labs] Add reasoning options
2026-06-14 21:12:34 -05:00
Aiden Cline ff4f61c81f Merge pull request #2570 from anomalyco/consolidate/kilo-small-labs-1-reasoning-options
[kilo/multiple labs 1] Add reasoning options
2026-06-14 21:12:22 -05:00
Aiden Cline 84c75799b8 Merge pull request #2571 from anomalyco/consolidate/kilo-small-labs-2-reasoning-options
[kilo/multiple labs 2] Add reasoning options
2026-06-14 21:12:09 -05:00
Aiden Cline 78cb09383d Merge pull request #2572 from anomalyco/consolidate/kilo-small-labs-3-reasoning-options
[kilo/multiple labs 3] Add reasoning options
2026-06-14 21:11:57 -05:00
Aiden Cline c7601d0f5a Merge pull request #2573 from anomalyco/consolidate/kilo-small-labs-4-reasoning-options
[kilo/multiple labs 4] Add reasoning options
2026-06-14 21:03:41 -05:00
Aiden Cline befaefc783 Merge pull request #2569 from anomalyco/consolidate/jiekou-small-labs-reasoning-options
[jiekou/multiple labs] Add reasoning options
2026-06-14 21:03:11 -05:00
Aiden Cline ca44c698aa Merge pull request #2574 from anomalyco/consolidate/llmgateway-search-xai-reasoning-options
[llmgateway/search and xAI] Add reasoning options
2026-06-14 21:00:49 -05:00
Aiden Cline 69613e89c6 Merge pull request #2577 from anomalyco/consolidate/nano-gpt-small-labs-2-reasoning-options
[nano-gpt/multiple labs 2] Add reasoning options
2026-06-14 21:00:36 -05:00
Aiden Cline 304e56b702 Merge pull request #2575 from anomalyco/consolidate/merge-gateway-small-labs-reasoning-options
[merge-gateway/multiple labs] Add reasoning options
2026-06-14 21:00:20 -05:00
Aiden Cline 1e1224b3a6 Merge pull request #2576 from anomalyco/consolidate/nano-gpt-small-labs-1-reasoning-options
[nano-gpt/multiple labs 1] Add reasoning options
2026-06-14 20:56:09 -05:00
Aiden Cline 39747a0c4c Merge pull request #2580 from anomalyco/consolidate/opencode-small-labs-reasoning-options
[opencode/multiple labs] Add reasoning options
2026-06-14 20:51:33 -05:00
Aiden Cline 3c03d0af77 Merge pull request #2578 from anomalyco/consolidate/nano-gpt-small-labs-3-reasoning-options
[nano-gpt/multiple labs 3] Add reasoning options
2026-06-14 20:50:28 -05:00
Aiden Cline 760f814f20 Merge pull request #2579 from anomalyco/consolidate/nearai-google-qwen-zai-reasoning-options
[nearai/google, Qwen, and Z.AI] Add reasoning options
2026-06-14 20:50:08 -05:00
Aiden Cline bc4b4af78e Merge pull request #2581 from anomalyco/consolidate/poe-small-labs-reasoning-options
[poe/multiple labs] Add reasoning options
2026-06-14 20:50:00 -05:00
Aiden Cline 87e5357f02 [nearai/google] Remove unsupported reasoning controls 2026-06-14 21:27:02 -04:00
Aiden Cline f3a85a45db Merge pull request #2582 from anomalyco/consolidate/siliconflow-small-labs-reasoning-options
[siliconflow/multiple labs] Add reasoning options
2026-06-14 20:23:25 -05:00
Aiden Cline de08ce69dc Merge pull request #2583 from anomalyco/consolidate/vercel-small-labs-reasoning-options
[vercel/multiple labs] Add reasoning options
2026-06-14 20:16:02 -05:00
Aiden Cline e9bea3caa7 Merge pull request #2562 from anomalyco/consolidate/302ai-small-labs-reasoning-options
[302ai/multiple labs] Add reasoning options
2026-06-14 20:15:42 -05:00
Aiden Cline 484ee191e7 [vercel/multiple labs] Add reasoning options 2026-06-14 21:04:26 -04:00
Aiden Cline d25df3464c [siliconflow/multiple labs] Add reasoning options 2026-06-14 21:04:22 -04:00
Aiden Cline 0f1ef5df74 [poe/multiple labs] Add reasoning options 2026-06-14 21:04:17 -04:00
Aiden Cline 0a757f8f3c [opencode/multiple labs] Add reasoning options 2026-06-14 21:04:15 -04:00
Aiden Cline a683e15e05 [nearai/google, Qwen, and Z.AI] Add reasoning options 2026-06-14 21:04:11 -04:00
Aiden Cline c96e3a9a1c [nano-gpt/multiple labs 3] Add reasoning options 2026-06-14 21:04:09 -04:00
Aiden Cline 07c3d34cab [nano-gpt/multiple labs 2] Add reasoning options 2026-06-14 21:04:05 -04:00
Aiden Cline 2f1141725e [nano-gpt/multiple labs 1] Add reasoning options 2026-06-14 21:04:01 -04:00
Aiden Cline aaa7f0225c [merge-gateway/multiple labs] Add reasoning options 2026-06-14 21:03:57 -04:00
Aiden Cline 7964fde548 [llmgateway/search and xAI] Add reasoning options 2026-06-14 21:03:54 -04:00
Aiden Cline a2981ede7a [kilo/multiple labs 4] Add reasoning options 2026-06-14 21:03:52 -04:00
Aiden Cline 4e6b11d780 [kilo/multiple labs 3] Add reasoning options 2026-06-14 21:03:50 -04:00
Aiden Cline 1537342ee8 [kilo/multiple labs 2] Add reasoning options 2026-06-14 21:03:46 -04:00
Aiden Cline f8ac69d04d [kilo/multiple labs 1] Add reasoning options 2026-06-14 21:03:43 -04:00
Aiden Cline 9ed0691a8f [jiekou/multiple labs] Add reasoning options 2026-06-14 21:03:39 -04:00
Aiden Cline a181661717 [github-models/multiple labs] Add reasoning options 2026-06-14 21:03:35 -04:00
Aiden Cline 25e84df306 [github-copilot/google and router] Add reasoning options 2026-06-14 21:03:33 -04:00
Aiden Cline 483483a548 [frogbot/multiple labs] Add reasoning options 2026-06-14 21:03:31 -04:00
Aiden Cline 6b4fc2da6c [cortecs/multiple labs] Add reasoning options 2026-06-14 21:03:28 -04:00
Aiden Cline 9178b8d96a [aihubmix/multiple labs] Add reasoning options 2026-06-14 21:03:21 -04:00
Aiden Cline 97e4f410f8 [302ai/multiple labs] Add reasoning options 2026-06-14 21:03:19 -04:00
Aiden Cline 61c9292dd2 Merge pull request #2541 from anomalyco/split/zenmux-minimax-reasoning-options
[zenmux/minimax] Add reasoning options
2026-06-14 19:57:18 -05:00
Aiden Cline 515cbe55e4 Merge pull request #2537 from anomalyco/split/zenmux-baidu-reasoning-options
[zenmux/baidu] Add reasoning options
2026-06-14 19:57:07 -05:00
Aiden Cline 4443540d24 Merge pull request #2538 from anomalyco/split/zenmux-deepseek-reasoning-options
[zenmux/deepseek] Add reasoning options
2026-06-14 19:56:58 -05:00
Aiden Cline 5142d98eac Merge pull request #2536 from anomalyco/split/zenmux-anthropic-reasoning-options
[zenmux/anthropic] Add reasoning options
2026-06-14 19:56:44 -05:00
Aiden Cline 75fc4a0341 Merge pull request #2525 from anomalyco/split/vercel-meituan-reasoning-options
[vercel/meituan] Add reasoning options
2026-06-14 19:56:30 -05:00
Aiden Cline 55234f593d Merge pull request #2535 from anomalyco/split/vercel-zai-reasoning-options
[vercel/zai] Add reasoning options
2026-06-14 19:56:20 -05:00
Aiden Cline d38f09549c Merge pull request #2542 from anomalyco/split/zenmux-moonshotai-reasoning-options
[zenmux/moonshotai] Add reasoning options
2026-06-14 19:56:07 -05:00
Aiden Cline 128a8da199 Merge pull request #2543 from anomalyco/split/zenmux-openai-reasoning-options
[zenmux/openai] Add reasoning options
2026-06-14 19:55:57 -05:00
Aiden Cline 2240450c73 Merge pull request #2556 from anomalyco/automation/sync-models-baseten
chore(sync): update Baseten model catalog
2026-06-14 19:55:06 -05:00
Aiden Cline 1e63debae9 Merge pull request #2560 from zainhas/dev
[Together AI] add kimi k2.7
2026-06-14 19:54:21 -05:00
Aiden Cline 72d8a5773a Merge pull request #2534 from anomalyco/split/vercel-xai-reasoning-options
[vercel/xai] Add reasoning options
2026-06-14 19:54:02 -05:00
Zain Hasan 16d6022afc fix family 2026-06-14 17:29:01 -07:00
Zain Hasan 9478cd312d Merge branch 'dev' into dev 2026-06-14 17:27:27 -07:00
Zain Hasan 5559feb253 add k2.7 to enum 2026-06-14 17:26:27 -07:00
Aiden Cline 389f551f32 Merge pull request #2558 from patrik-kuehl/add-minimax-m3-to-synthetic-provider
feat(providers): add MiniMax M3 to Synthetic provider
2026-06-14 18:53:30 -05:00
Aiden Cline f487c9692f Merge pull request #2561 from jpetrina/add-gemma4-e2b-e4b
feat(models): add Gemma 4 E2B and E4B variants
2026-06-14 18:52:08 -05:00
Aiden Cline 89282134fd Merge pull request #2551 from smakosh/feat/llmgateway-newest-text-models
feat: add LLM Gateway kimi-k2.7-code, nemotron-3-ultra-550b, grok-build-0-1
2026-06-14 18:51:42 -05:00
github-actions[bot] 247ffb8207 chore(sync): update Baseten model catalog 2026-06-14 23:42:27 +00:00
Jakov Petrina cb96a2e701 feat(models): add Gemma 4 E2B and E4B variants
Signed-off-by: Jakov Petrina <jkv.petrina@gmail.com>
2026-06-15 00:03:04 +02:00
Patrik Kühl 0e53645dce chore(models): add MiniMax M3 weights URL 2026-06-15 00:00:31 +02:00
Patrik Kühl 626700e268 chore: provide empty reasoning options 2026-06-14 23:54:59 +02:00
Zain Hasan c5bc7e9e3e [Together AI] add kimi k2.7 2026-06-14 13:40:50 -07:00
smakosh fb2b96a4b7 feat: add reasoning_options to new LLM Gateway models
Addresses review feedback: kimi-k2.7-code and grok-build-0-1 use the
effort (low/medium/high) option matching the kimi/grok gateway models;
nemotron-3-ultra-550b uses a reasoning toggle per its nvidia source.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-14 18:15:02 +02:00
Patrik Kühl 953b651adc feat(providers): add MiniMax M3 model to Synthetic provider 2026-06-14 14:22:59 +02:00
Patrik Kühl bf1e39cd19 chore(models): mark MiniMax M3 as open-weighted 2026-06-14 14:09:28 +02:00
Aiden Cline ce14787192 Merge pull request #2554 from dsingal0/fix-dsv4-context
fix(baseten): update DeepSeek V4 Pro context length to 1,048,576
2026-06-14 05:05:17 -05:00
Dhruv Singal 401da7398d fix(baseten): remove incorrect context limits from DeepSeek V4 Pro, inherit from base model 2026-06-14 05:20:52 +00:00
Aiden Cline 883951b6ae Merge pull request #2553 from JSap0914/fix/command-r7b-release-date
fix(cohere): correct Command R7B release date to 2024-12-02
2026-06-13 23:53:13 -05:00
JSap0914 0d09d0f2a2 fix(cohere): correct Command R7B release date to 2024-12-02
command-r7b-12-2024 had release_date/last_updated set to 2024-02-27,
which predates the model — its id encodes December 2024, and 02-27 was
evidently copied from the sibling command-r7b-arabic-02-2025 entry.
Cohere's official announcement is dated December 2, 2024.
2026-06-14 13:05:37 +09:00
Aiden Cline 3b642e68c1 [zenmux/minimax] Add MiniMax M3 thinking toggle 2026-06-13 19:16:05 -05:00
Aiden Cline f0cfea9185 Merge pull request #2544 from anomalyco/split/zenmux-qwen-reasoning-options
[zenmux/qwen] Add reasoning options
2026-06-13 19:14:40 -05:00
Aiden Cline dc4f59bf13 Merge pull request #2516 from anomalyco/split/vercel-amazon-reasoning-options
[vercel/amazon] Add reasoning options
2026-06-13 19:08:41 -05:00
Aiden Cline 101a1c3771 Merge pull request #2539 from anomalyco/split/zenmux-google-reasoning-options
[zenmux/google] Add reasoning options
2026-06-13 19:06:06 -05:00
Aiden Cline 9344d01b8b Merge pull request #2540 from anomalyco/split/zenmux-inclusionai-reasoning-options
[zenmux/inclusionai] Add reasoning options
2026-06-13 19:05:53 -05:00
Aiden Cline b096a9f0d2 Merge pull request #2518 from anomalyco/split/vercel-arcee-ai-reasoning-options
[vercel/arcee-ai] Add reasoning options
2026-06-13 19:05:45 -05:00
Aiden Cline 1bf16b9774 Merge pull request #2425 from anomalyco/split/kilo-stepfun-reasoning-options
[kilo/stepfun] Add reasoning options
2026-06-13 19:05:35 -05:00
Aiden Cline b303848e33 Merge pull request #2546 from anomalyco/split/zenmux-stepfun-reasoning-options
[zenmux/stepfun] Add reasoning options
2026-06-13 19:04:30 -05:00
Aiden Cline 0440528e10 Merge pull request #2549 from anomalyco/split/zenmux-x-ai-reasoning-options
[zenmux/x-ai] Add reasoning options
2026-06-13 19:04:16 -05:00
Aiden Cline 3bbab9fd50 Merge pull request #2545 from anomalyco/split/zenmux-sapiens-ai-reasoning-options
[zenmux/sapiens-ai] Add reasoning options
2026-06-13 19:01:50 -05:00
Aiden Cline 78f0824557 [zenmux/x-ai] Correct Grok reasoning controls 2026-06-13 19:01:50 -05:00
Aiden Cline 15a29aabfc Merge pull request #2523 from anomalyco/split/vercel-interfaze-reasoning-options
[vercel/interfaze] Add reasoning options
2026-06-13 19:01:41 -05:00
Aiden Cline 5dbbd02f35 Merge pull request #2531 from anomalyco/split/vercel-openai-reasoning-options-part-2
[vercel/openai part 2] Add reasoning options
2026-06-13 19:01:29 -05:00
Aiden Cline a34573e367 Merge pull request #2530 from anomalyco/split/vercel-openai-reasoning-options-part-1
[vercel/openai part 1] Add reasoning options
2026-06-13 19:01:16 -05:00
Aiden Cline 9f6f058562 Merge pull request #2547 from anomalyco/split/zenmux-tencent-reasoning-options
[zenmux/tencent] Add reasoning options
2026-06-13 19:00:54 -05:00
Aiden Cline 8ed57cde03 Merge pull request #2548 from anomalyco/split/zenmux-volcengine-reasoning-options
[zenmux/volcengine] Add reasoning options
2026-06-13 19:00:46 -05:00
Aiden Cline 0383342620 Merge pull request #2550 from anomalyco/split/zenmux-z-ai-reasoning-options
[zenmux/z-ai] Add reasoning options
2026-06-13 19:00:09 -05:00
Aiden Cline 4c645691d7 Merge pull request #2552 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-06-13 18:58:59 -05:00
Aiden Cline be9e01d10c Merge pull request #2288 from anomalyco/split/opencode-alibaba-reasoning-options
[opencode/alibaba] Add reasoning options
2026-06-13 18:58:46 -05:00
github-actions[bot] ad68e2b348 chore(sync): update OpenRouter model catalog 2026-06-13 23:40:33 +00:00
smakosh 57940ad416 feat: add LLM Gateway kimi-k2.7-code, nemotron-3-ultra-550b, grok-build-0-1
Newest text models from the LLM Gateway catalog, using the base_model
structure to inherit from the canonical model registry with gateway-specific
cost overrides.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-13 23:02:24 +01:00
Aiden Cline 7a9e515f24 [siliconflow/THUDM] Add reasoning options 2026-06-13 16:11:23 -05:00
Aiden Cline 3f0ba493d0 [siliconflow/tencent] Add reasoning options 2026-06-13 16:11:21 -05:00
Aiden Cline 383a36f646 [nano-gpt/x-ai] Add reasoning options 2026-06-13 16:09:14 -05:00
Aiden Cline e6305607a4 [nano-gpt/moonshotai] Add reasoning options 2026-06-13 16:08:44 -05:00
Aiden Cline c0d3207da7 [zenmux/z-ai] Add reasoning options 2026-06-13 16:07:58 -05:00
Aiden Cline e2fbc080ca [zenmux/x-ai] Add reasoning options 2026-06-13 16:07:56 -05:00
Aiden Cline 23563d1b2d [zenmux/volcengine] Add reasoning options 2026-06-13 16:07:55 -05:00
Aiden Cline ce4594421c [zenmux/tencent] Add reasoning options 2026-06-13 16:07:53 -05:00
Aiden Cline 50e3a7d946 [zenmux/stepfun] Add reasoning options 2026-06-13 16:07:51 -05:00
Aiden Cline bdf25065cf [zenmux/sapiens-ai] Add reasoning options 2026-06-13 16:07:49 -05:00
Aiden Cline 9f82b646a8 [zenmux/qwen] Add reasoning options 2026-06-13 16:07:47 -05:00
Aiden Cline 67ba921b70 [zenmux/openai] Add reasoning options 2026-06-13 16:07:45 -05:00
Aiden Cline 937948ac05 [zenmux/moonshotai] Add reasoning options 2026-06-13 16:07:43 -05:00
Aiden Cline b4ad1b5e7c [zenmux/minimax] Add reasoning options 2026-06-13 16:07:41 -05:00
Aiden Cline 8855f33982 [zenmux/inclusionai] Add reasoning options 2026-06-13 16:07:39 -05:00
Aiden Cline d580d186f4 [zenmux/google] Add reasoning options 2026-06-13 16:07:37 -05:00
Aiden Cline 338a3ba4cc [zenmux/deepseek] Add reasoning options 2026-06-13 16:07:35 -05:00
Aiden Cline aa29468222 [zenmux/baidu] Add reasoning options 2026-06-13 16:07:34 -05:00
Aiden Cline 45f0268363 [zenmux/anthropic] Add reasoning options 2026-06-13 16:07:32 -05:00
Aiden Cline 6eb4986851 [kilo/stepfun] Add reasoning options 2026-06-13 16:06:44 -05:00
Aiden Cline dd0988cde0 [vercel/zai] Add reasoning options 2026-06-13 16:04:50 -05:00
Aiden Cline 631d348d75 [vercel/xai] Add reasoning options 2026-06-13 16:04:48 -05:00
Aiden Cline 3eb0985188 [vercel/openai part 2] Add reasoning options 2026-06-13 16:04:43 -05:00
Aiden Cline b69a4fc71e [vercel/openai part 1] Add reasoning options 2026-06-13 16:04:41 -05:00
Aiden Cline cbb47c5fb7 [vercel/nvidia] Add reasoning options 2026-06-13 16:04:39 -05:00
Aiden Cline 57319b2086 [vercel/minimax] Add reasoning options 2026-06-13 16:04:33 -05:00
Aiden Cline 2eef2259c5 [vercel/meituan] Add reasoning options 2026-06-13 16:04:31 -05:00
Aiden Cline debfd6339c [vercel/interfaze] Add reasoning options 2026-06-13 16:04:27 -05:00
Aiden Cline 6543300a5d [vercel/arcee-ai] Add reasoning options 2026-06-13 16:04:18 -05:00
Aiden Cline e6b575adf1 [vercel/amazon] Add reasoning options 2026-06-13 16:04:14 -05:00
Aiden Cline 57a8c746e5 [opencode/minimax] Add reasoning options 2026-06-13 16:00:55 -05:00
Aiden Cline 3e3918929c [opencode/anthropic] Add reasoning options 2026-06-13 16:00:47 -05:00
Aiden Cline 4d9a365f36 [opencode/alibaba] Add reasoning options 2026-06-13 16:00:45 -05:00
Aiden Cline 4dff8372f3 Merge pull request #2287 from anomalyco/automation/sync-models-venice
chore(sync): update Venice model catalog
2026-06-13 15:59:36 -05:00
github-actions[bot] e05c2a09a7 chore(sync): update Venice model catalog 2026-06-13 20:43:48 +00:00
606 changed files with 1776 additions and 1895 deletions
+22
View File
@@ -0,0 +1,22 @@
name = "Qwen3.5 9B"
family = "qwen"
release_date = "2026-02-23"
last_updated = "2026-02-23"
attachment = false
reasoning = true
temperature = true
tool_call = true
structured_output = true
open_weights = true
[limit]
context = 262_144
output = 65_536
[modalities]
input = ["text"]
output = ["text"]
[[weights]]
label = "Hugging Face"
url = "https://huggingface.co/Qwen/Qwen3.5-9B"
+2 -2
View File
@@ -1,7 +1,7 @@
name = "Command R7B"
family = "command-r"
release_date = "2024-02-27"
last_updated = "2024-02-27"
release_date = "2024-12-02"
last_updated = "2024-12-02"
attachment = false
reasoning = false
temperature = true
+22
View File
@@ -0,0 +1,22 @@
name = "Gemma 4 E2B IT"
family = "gemma"
release_date = "2026-04-02"
last_updated = "2026-04-02"
attachment = true
reasoning = true
temperature = true
tool_call = true
structured_output = true
open_weights = true
[limit]
context = 131_072
output = 8_192
[modalities]
input = ["text", "image", "audio"]
output = ["text"]
[[weights]]
label = "Hugging Face"
url = "https://huggingface.co/google/gemma-4-E2B-it"
+22
View File
@@ -0,0 +1,22 @@
name = "Gemma 4 E4B IT"
family = "gemma"
release_date = "2026-04-02"
last_updated = "2026-04-02"
attachment = true
reasoning = true
temperature = true
tool_call = true
structured_output = true
open_weights = true
[limit]
context = 131_072
output = 8_192
[modalities]
input = ["text", "image", "audio"]
output = ["text"]
[[weights]]
label = "Hugging Face"
url = "https://huggingface.co/google/gemma-4-E4B-it"
+5 -1
View File
@@ -6,7 +6,7 @@ attachment = true
reasoning = true
temperature = true
tool_call = true
open_weights = false
open_weights = true
[limit]
context = 512_000
@@ -15,3 +15,7 @@ output = 128_000
[modalities]
input = ["text", "image", "video"]
output = ["text"]
[[weights]]
label = "Hugging Face"
url = "https://huggingface.co/MiniMaxAI/MiniMax-M3"
@@ -0,0 +1,23 @@
name = "Kimi K2.7 Code Highspeed"
family = "kimi-k2"
release_date = "2026-06-12"
last_updated = "2026-06-12"
attachment = true
reasoning = true
temperature = false
tool_call = true
structured_output = true
knowledge = "2025-01"
open_weights = true
[limit]
context = 262_144
output = 262_144
[modalities]
input = ["text", "image", "video"]
output = ["text"]
[[weights]]
label = "Hugging Face"
url = "https://huggingface.co/moonshotai/Kimi-K2.7-Code"
+17
View File
@@ -0,0 +1,17 @@
name = "GPT-Image-1.5"
family = "gpt-image"
release_date = "2025-11-25"
last_updated = "2025-11-25"
attachment = true
reasoning = false
temperature = false
tool_call = false
open_weights = false
[limit]
context = 0
output = 0
[modalities]
input = ["text", "image"]
output = ["text", "image"]
+17
View File
@@ -0,0 +1,17 @@
name = "GPT-Image-1"
family = "gpt-image"
release_date = "2025-04-24"
last_updated = "2025-04-24"
attachment = true
reasoning = false
temperature = false
tool_call = false
open_weights = false
[limit]
context = 0
output = 0
[modalities]
input = ["text", "image"]
output = ["image"]
+17
View File
@@ -0,0 +1,17 @@
name = "GPT-Image-2"
family = "gpt-image"
release_date = "2026-04-21"
last_updated = "2026-04-21"
attachment = true
reasoning = false
temperature = false
tool_call = false
open_weights = false
[limit]
context = 0
output = 0
[modalities]
input = ["text", "image"]
output = ["image"]
+2 -2
View File
@@ -1,7 +1,7 @@
name = "GLM-5.1"
family = "glm"
release_date = "2026-03-27"
last_updated = "2026-03-27"
release_date = "2026-04-07"
last_updated = "2026-04-07"
attachment = false
reasoning = true
temperature = true
+5 -1
View File
@@ -7,7 +7,7 @@ reasoning = true
temperature = true
tool_call = true
structured_output = true
open_weights = false
open_weights = true
[limit]
context = 1_000_000
@@ -16,3 +16,7 @@ output = 131_072
[modalities]
input = ["text"]
output = ["text"]
[[weights]]
label = "Hugging Face"
url = "https://huggingface.co/zai-org/GLM-5.2"
+2 -2
View File
@@ -1,7 +1,7 @@
name = "GLM-5"
family = "glm"
release_date = "2026-02-11"
last_updated = "2026-02-11"
release_date = "2026-02-12"
last_updated = "2026-02-12"
attachment = false
reasoning = true
temperature = true
+3
View File
@@ -246,6 +246,9 @@ export const ModelFamilyValues = [
// Lucid
"lucid",
// LucidQuery
"agi",
// Intellect
"intellect",
@@ -7,6 +7,7 @@ import type { ExistingModel, SyncProvider, SyncedFullModel, SyncedModel } from "
const API_ENDPOINT = "https://openrouter.ai/api/v1/models";
const MODELS_DIR = path.join(import.meta.dirname, "..", "..", "..", "..", "..", "models");
const MODEL_NAME_BLACKLIST = ["fable-5"];
const modelMetadataByID = new Map<string, Record<string, unknown>>();
const modelMetadataFilesByProvider = new Map<string, Set<string>>();
@@ -79,7 +80,10 @@ export const openrouter = {
return response.json();
},
parseModels(raw) {
return OpenRouterResponse.parse(raw).data;
return OpenRouterResponse.parse(raw).data.filter((model) => {
const name = `${model.id} ${model.name}`.toLowerCase();
return MODEL_NAME_BLACKLIST.every((value) => !name.includes(value));
});
},
translateModel(model, context) {
return {
+23 -6
View File
@@ -6,7 +6,16 @@ import { factorBaseModel, resolveCanonicalBaseModel } from "./openrouter.js";
const API_ENDPOINT = "https://ai-gateway.vercel.sh/v1/models";
const ModelType = z.enum(["language", "embedding", "image", "video", "reranking"]);
const ModelType = z.enum([
"language",
"embedding",
"image",
"video",
"reranking",
"transcription",
"speech",
"realtime",
]);
const PricingTier = z.object({
cost: z.string(),
@@ -30,8 +39,8 @@ export const VercelModel = z.object({
name: z.string(),
created: z.number(),
released: z.number().optional(),
context_window: z.number(),
max_tokens: z.number(),
context_window: z.number().optional().default(0),
max_tokens: z.number().optional().default(0),
type: ModelType,
tags: z.array(z.string()).optional().default([]),
pricing: Pricing.optional(),
@@ -107,9 +116,17 @@ export function buildVercelModel(model: VercelModel, existing: ExistingModel | u
cost,
limit: { context, input, output },
modalities: {
input: ["text", tags.has("vision") ? "image" : undefined, tags.has("file-input") ? "pdf" : undefined]
.filter((value): value is "text" | "image" | "pdf" => value !== undefined),
output: model.type === "image"
input: model.type === "transcription"
? ["audio"]
: model.type === "realtime"
? ["text", "audio"]
: ["text", tags.has("vision") ? "image" : undefined, tags.has("file-input") ? "pdf" : undefined]
.filter((value): value is "text" | "image" | "pdf" => value !== undefined),
output: model.type === "speech"
? ["audio"]
: model.type === "realtime"
? ["text", "audio"]
: model.type === "image"
? ["image"]
: model.type === "video"
? ["video"]
-150
View File
@@ -1,150 +0,0 @@
import { afterEach, expect, test } from "bun:test";
import path from "node:path";
import { mkdir, mkdtemp } from "node:fs/promises";
import os from "node:os";
import { syncProvider } from "../src/sync/index.js";
import {
BasetenResponse,
baseten,
buildBasetenModel,
fetchBasetenModels,
type BasetenModel,
} from "../src/sync/providers/baseten.js";
const catalogModel: BasetenModel = {
id: "zai-org/GLM-5.1",
name: "GLM 5.1",
context_length: 128_000,
max_completion_tokens: 32_000,
input_modalities: ["text"],
output_modalities: ["text"],
pricing: {
prompt: "0.00000012",
completion: "0.0000005",
},
supported_features: ["reasoning", "reasoning_effort", "tools", "structured_outputs"],
supported_sampling_parameters: ["temperature", "top_p"],
};
const newCatalogModel: BasetenModel = {
...catalogModel,
};
afterEach(() => {
baseten.modelsDir = "providers/baseten/models";
baseten.fetchModels = async () => {
const key = process.env.BASETEN_API_KEY;
if (key === undefined) throw new Error("Baseten sync requires BASETEN_API_KEY");
return fetchBasetenModels(key);
};
});
test("Baseten maps authoritative fields and preserves curated metadata", () => {
const synced = buildBasetenModel(catalogModel, {
name: "Old name",
release_date: "2025-08-05",
last_updated: "2025-09-01",
attachment: false,
reasoning: true,
reasoning_options: [{ type: "effort", values: ["low", "high"] }],
tool_call: true,
open_weights: true,
status: "deprecated",
interleaved: { field: "reasoning_content" },
base_model: "zhipuai/glm-5.1",
base_model_omit: ["limit.input"],
cost: { input: 0.1, output: 0.4, cache_write: 0.2 },
limit: { context: 64_000, output: 16_000 },
modalities: { input: ["text"], output: ["text"] },
});
expect(synced).toMatchObject({
base_model: "zhipuai/glm-5.1",
base_model_omit: ["limit.input"],
reasoning_options: [{ type: "effort", values: ["low", "high"] }],
status: "deprecated",
interleaved: { field: "reasoning_content" },
cost: { input: 0.12, output: 0.5, cache_write: 0.2 },
limit: { context: 128_000, output: 32_000 },
});
});
test("Baseten preserves curated reasoning when an opt-in capability is omitted", () => {
const synced = buildBasetenModel({
...catalogModel,
supported_features: ["tools", "structured_outputs"],
}, {
name: "GLM 5.1",
release_date: "2026-05-20",
last_updated: "2026-05-20",
attachment: false,
reasoning: true,
reasoning_options: [{ type: "toggle" }],
tool_call: true,
open_weights: true,
cost: { input: 1, output: 4 },
limit: { context: 100_000, output: 50_000 },
modalities: { input: ["text"], output: ["text"] },
});
expect(synced).toMatchObject({
reasoning: true,
reasoning_options: [{ type: "toggle" }],
});
});
test("Baseten sync adds exact base models, retains missing entries, and is idempotent", async () => {
const root = await mkdtemp(path.join(os.tmpdir(), "models-dev-baseten-"));
const modelsDir = path.join(root, "providers", "baseten", "models");
const metadataDir = path.join(root, "models", "zhipuai");
await mkdir(path.join(modelsDir, "stale"), { recursive: true });
await mkdir(metadataDir, { recursive: true });
await Bun.write(
path.join(metadataDir, "glm-5.1.toml"),
Bun.file(path.join(import.meta.dirname, "../../../models/zhipuai/glm-5.1.toml")),
);
await Bun.write(path.join(modelsDir, "stale", "model.toml"), [
'name = "Retained"',
'release_date = "2025-01-01"',
'last_updated = "2025-01-01"',
"attachment = false",
"reasoning = false",
"tool_call = false",
"open_weights = false",
"[cost]",
"input = 1",
"output = 1",
"[limit]",
"context = 1000",
"output = 100",
"[modalities]",
'input = ["text"]',
'output = ["text"]',
"",
].join("\n"));
baseten.modelsDir = modelsDir;
baseten.fetchModels = async () => ({ data: [newCatalogModel] });
const first = await syncProvider(baseten);
const second = await syncProvider(baseten);
expect(first.created).toBe(1);
expect(first.deleted).toBe(0);
expect(first.notices.join(" ")).toContain("stale/model.toml");
expect(second).toMatchObject({ created: 0, updated: 0, deleted: 0 });
});
test("Baseten rejects malformed catalog responses", () => {
expect(() => BasetenResponse.parse({ data: "broken" })).toThrow();
});
test("Baseten rejects non-success API responses", async () => {
const fetcher = async () => new Response("unauthorized", {
status: 401,
statusText: "Unauthorized",
});
expect(fetchBasetenModels("fixture-key", fetcher as typeof fetch))
.rejects.toThrow("Baseten models request failed: 401 Unauthorized");
});
@@ -1,35 +0,0 @@
import { expect, test } from "bun:test";
import { buildWorkersAiModel } from "../src/sync/providers/cloudflare-workers-ai.js";
import type { OpenRouterModel } from "../src/sync/providers/openrouter.js";
test("Cloudflare Workers AI sync preserves reasoning options", () => {
const model: OpenRouterModel = {
id: "@cf/nvidia/nemotron-3-120b-a12b",
name: "Nemotron 3 Super 120B",
created: 1_773_187_200,
hugging_face_id: "nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16",
knowledge_cutoff: null,
context_length: 256_000,
architecture: {
input_modalities: ["text"],
output_modalities: ["text"],
},
pricing: {
prompt: "0.0000005",
completion: "0.0000015",
},
top_provider: {
context_length: 256_000,
max_completion_tokens: 256_000,
},
supported_parameters: ["reasoning", "tools", "temperature"],
};
const synced = buildWorkersAiModel(model, {
base_model: "nvidia/nemotron-3-super-120b-a12b",
reasoning_options: [{ type: "toggle" }],
});
expect(synced.reasoning_options).toEqual([{ type: "toggle" }]);
});
-35
View File
@@ -1,35 +0,0 @@
import { expect, test } from "bun:test";
import { buildGoogleModel } from "../src/sync/providers/google.js";
test("Google sync keeps base models compact", () => {
const synced = buildGoogleModel({
name: "models/gemini-3-pro-image-preview",
displayName: "Nano Banana Pro",
inputTokenLimit: 131_072,
outputTokenLimit: 32_768,
temperature: 1,
thinking: true,
}, {
base_model: "google/gemini-3-pro-image-preview",
name: "Nano Banana Pro",
family: "gemini-pro",
release_date: "2025-11-20",
last_updated: "2025-11-20",
attachment: true,
reasoning: true,
temperature: true,
tool_call: false,
knowledge: "2025-01",
open_weights: false,
cost: { input: 2, output: 120 },
limit: { context: 65_536, output: 32_768 },
modalities: { input: ["text", "image"], output: ["text", "image"] },
});
expect(synced).toEqual({
base_model: "google/gemini-3-pro-image-preview",
cost: { input: 2, output: 120 },
limit: { context: 131_072 },
});
});
-121
View File
@@ -1,121 +0,0 @@
import { expect, test } from "bun:test";
import { preserveBaseModel } from "../src/sync/index.js";
import { resolveCloudflareBaseModel } from "../src/sync/providers/cloudflare-workers-ai.js";
import { buildOpenRouterModel, type OpenRouterModel } from "../src/sync/providers/openrouter.js";
test("OpenRouter z-ai models inherit from zhipuai metadata", () => {
const model: OpenRouterModel = {
id: "z-ai/glm-5.1",
name: "Z.AI: GLM-5.1",
created: 1_777_680_000,
hugging_face_id: "zai-org/GLM-5.1",
knowledge_cutoff: null,
context_length: 200_000,
architecture: {
input_modalities: ["text"],
output_modalities: ["text"],
},
pricing: {
prompt: "0.0000014",
completion: "0.0000044",
},
top_provider: {
context_length: 200_000,
max_completion_tokens: 131_072,
},
supported_parameters: ["tools", "tool_choice", "temperature", "structured_outputs"],
};
const synced = buildOpenRouterModel(model, undefined);
expect("base_model" in synced ? synced.base_model : undefined).toBe("zhipuai/glm-5.1");
});
test("OpenRouter-derived syncs preserve existing base model links", () => {
const model: OpenRouterModel = {
id: "@cf/nvidia/nemotron-3-120b-a12b",
name: "Nemotron 3 Super 120B",
created: 1_773_187_200,
hugging_face_id: "nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16",
knowledge_cutoff: null,
context_length: 256_000,
architecture: {
input_modalities: ["text"],
output_modalities: ["text"],
},
pricing: {
prompt: "0.0000005",
completion: "0.0000015",
},
top_provider: {
context_length: 256_000,
max_completion_tokens: 256_000,
},
supported_parameters: ["reasoning", "tools", "temperature", "structured_outputs"],
};
const synced = preserveBaseModel(buildOpenRouterModel(model, undefined), {
base_model: "nvidia/nemotron-3-super-120b-a12b",
base_model_omit: ["limit.input"],
});
expect("base_model" in synced ? synced.base_model : undefined)
.toBe("nvidia/nemotron-3-super-120b-a12b");
expect("base_model_omit" in synced ? synced.base_model_omit : undefined)
.toEqual(["limit.input"]);
});
test("newly detected base models do not replace existing links", () => {
const model: OpenRouterModel = {
id: "z-ai/glm-5.1",
name: "Z.AI: GLM-5.1",
created: 1_777_680_000,
hugging_face_id: null,
knowledge_cutoff: null,
context_length: 200_000,
architecture: { input_modalities: ["text"], output_modalities: ["text"] },
pricing: { prompt: "0.0000014", completion: "0.0000044" },
top_provider: { context_length: 200_000, max_completion_tokens: 131_072 },
supported_parameters: ["tools"],
};
const synced = buildOpenRouterModel(model, {
base_model: "zhipuai/glm-5",
});
expect("base_model" in synced ? synced.base_model : undefined).toBe("zhipuai/glm-5");
});
test("undefined translated links preserve existing base model fields", () => {
const synced = preserveBaseModel({
base_model: undefined,
} as never, {
base_model: "nvidia/nemotron-3-super-120b-a12b",
base_model_omit: ["limit.input"],
});
expect("base_model" in synced ? synced.base_model : undefined)
.toBe("nvidia/nemotron-3-super-120b-a12b");
expect("base_model_omit" in synced ? synced.base_model_omit : undefined)
.toEqual(["limit.input"]);
});
test("new Cloudflare models discover a unique metadata base model", () => {
const model: OpenRouterModel = {
id: "@cf/nvidia/nemotron-3-120b-a12b",
name: "Nemotron 3 Super 120B",
created: 1_773_187_200,
hugging_face_id: null,
knowledge_cutoff: null,
context_length: 256_000,
architecture: { input_modalities: ["text"], output_modalities: ["text"] },
pricing: { prompt: "0.0000005", completion: "0.0000015" },
top_provider: { context_length: 256_000, max_completion_tokens: 256_000 },
supported_parameters: ["reasoning"],
};
expect(resolveCloudflareBaseModel(model)).toBe("nvidia/nemotron-3-super-120b-a12b");
const synced = buildOpenRouterModel(model, undefined, resolveCloudflareBaseModel(model));
expect("base_model" in synced ? synced.base_model : undefined)
.toBe("nvidia/nemotron-3-super-120b-a12b");
});
-227
View File
@@ -1,227 +0,0 @@
import { expect, test } from "bun:test";
import path from "node:path";
import { mkdtemp, mkdir, readlink, symlink } from "node:fs/promises";
import os from "node:os";
import { AuthoredModelShape } from "../src/schema.js";
import { syncProvider, type SyncProvider, type SyncedFullModel } from "../src/sync/index.js";
const model: SyncedFullModel = {
name: "Test model",
release_date: "2026-01-01",
last_updated: "2026-01-01",
attachment: false,
reasoning: false,
tool_call: false,
open_weights: false,
cost: { input: 1, output: 2 },
limit: { context: 1_000, output: 100 },
modalities: { input: ["text"], output: ["text"] },
};
test("reasoning budgets allow only the -1 negative sentinel", () => {
const authored = { id: "model", ...model };
expect(AuthoredModelShape.safeParse({
...authored,
reasoning_options: [{ type: "budget_tokens", min: -1, max: 32_768 }],
}).success).toBe(true);
expect(AuthoredModelShape.safeParse({
...authored,
reasoning_options: [{ type: "budget_tokens", min: -2, max: 32_768 }],
}).success).toBe(false);
});
test("reasoning efforts accept the provider default value", () => {
expect(AuthoredModelShape.safeParse({
id: "model",
...model,
reasoning: true,
reasoning_options: [{ type: "effort", values: ["none", "default"] }],
}).success).toBe(true);
});
async function fixture() {
const root = await mkdtemp(path.join(os.tmpdir(), "models-dev-sync-"));
const modelsDir = path.join(root, "providers", "test", "models");
await mkdir(modelsDir, { recursive: true });
return { root, modelsDir };
}
function provider(
modelsDir: string,
ids: string[],
deleteMissing = true,
preserveSymlinks = false,
): SyncProvider<string> {
return {
id: "test",
name: "Test",
modelsDir,
deleteMissing,
preserveSymlinks,
missingNotice: (paths) => paths.map((item) => `missing: ${item}`),
async fetchModels() {
return ids;
},
parseModels(raw) {
return raw as string[];
},
translateModel(id) {
return { id, model };
},
};
}
test("sync repairs a broken symlink returned by the source", async () => {
const { modelsDir } = await fixture();
const filePath = path.join(modelsDir, "model.toml");
await symlink("missing.toml", filePath);
const result = await syncProvider(provider(modelsDir, ["model"]));
expect(result.created).toBe(1);
expect(await Bun.file(filePath).text()).toContain('name = "Test model"');
expect(readlink(filePath)).rejects.toThrow();
});
test("sync preserves valid symlink aliases when configured", async () => {
const { root, modelsDir } = await fixture();
const targetPath = path.join(root, "target.toml");
const filePath = path.join(modelsDir, "model.toml");
await Bun.write(targetPath, `name = "Alias target"\n`);
await symlink(targetPath, filePath);
const result = await syncProvider(provider(modelsDir, ["model"], true, true));
expect(result.updated).toBe(0);
expect(await readlink(filePath)).toBe(targetPath);
});
test("sync removes a broken symlink absent from the source", async () => {
const { modelsDir } = await fixture();
const filePath = path.join(modelsDir, "model.toml");
await symlink("missing.toml", filePath);
const result = await syncProvider(provider(modelsDir, []));
expect(result.deleted).toBe(1);
expect(await Bun.file(filePath).exists()).toBe(false);
});
test("non-deleting sync reports missing broken symlinks", async () => {
const { modelsDir } = await fixture();
const filePath = path.join(modelsDir, "model.toml");
await symlink("missing.toml", filePath);
const result = await syncProvider(provider(modelsDir, [], false));
expect(result.deleted).toBe(0);
expect(result.notices).toEqual(["missing: model.toml"]);
expect(await readlink(filePath)).toBe("missing.toml");
});
test("sync preserves authored reasoning options omitted by a translator", async () => {
const { modelsDir } = await fixture();
const filePath = path.join(modelsDir, "model.toml");
await Bun.write(filePath, `name = "Old name"
release_date = "2026-01-01"
last_updated = "2026-01-01"
attachment = false
reasoning = true
tool_call = false
open_weights = false
[[reasoning_options]]
type = "effort"
values = ["low", "high"]
[[reasoning_options]]
type = "budget_tokens"
min = -1
max = 32768
[cost]
input = 1
output = 2
[limit]
context = 1000
output = 100
[modalities]
input = ["text"]
output = ["text"]
`);
const sync = provider(modelsDir, ["model"]);
sync.translateModel = (id) => ({
id,
model: { ...model, reasoning: true },
});
const first = await syncProvider(sync);
const content = await Bun.file(filePath).text();
const second = await syncProvider(sync);
expect(first.updated).toBe(1);
expect(content).toContain("[[reasoning_options]]");
expect(content).toContain('values = ["low", "high"]');
expect(content).toContain("min = -1");
expect(content).toContain("max = 32_768");
expect(second.updated).toBe(0);
expect(second.unchanged).toBe(1);
});
test("sync writes metadata returned by a provider translator", async () => {
const root = await mkdtemp(path.join(os.tmpdir(), "models-dev-sync-metadata-"));
const modelsDir = path.join(root, "providers", "test", "models");
await mkdir(modelsDir, { recursive: true });
const sync = provider(modelsDir, ["model"]);
sync.translateModel = () => ({
id: "model",
model: {
base_model: "test/model",
reasoning_options: [],
cost: { input: 1, output: 2 },
},
metadata: {
id: "test/model",
model: {
name: "Model",
release_date: "2026-06-10",
last_updated: "2026-06-10",
attachment: false,
reasoning: false,
tool_call: true,
open_weights: false,
limit: { context: 1_000, output: 100 },
modalities: { input: ["text"], output: ["text"] },
},
},
});
const first = await syncProvider(sync);
const second = await syncProvider(sync);
expect(first).toMatchObject({ created: 2, updated: 0 });
expect(second).toMatchObject({ created: 0, updated: 0 });
expect(await Bun.file(path.join(root, "models", "test", "model.toml")).text()).toContain('name = "Model"');
});
test("sync removes missing metadata only from its owned namespace", async () => {
const { root, modelsDir } = await fixture();
const ownedDir = path.join(root, "models", "test");
const otherDir = path.join(root, "models", "other");
await mkdir(ownedDir, { recursive: true });
await mkdir(otherDir, { recursive: true });
await Bun.write(path.join(ownedDir, "stale.toml"), 'name = "Stale"\n');
await Bun.write(path.join(otherDir, "retained.toml"), 'name = "Retained"\n');
const sync = provider(modelsDir, []);
sync.metadataNamespace = "test";
const result = await syncProvider(sync);
expect(result.deleted).toBe(1);
expect(await Bun.file(path.join(ownedDir, "stale.toml")).exists()).toBe(false);
expect(await Bun.file(path.join(otherDir, "retained.toml")).exists()).toBe(true);
});
-167
View File
@@ -1,167 +0,0 @@
import { expect, test } from "bun:test";
import { readdirSync } from "node:fs";
import path from "node:path";
import {
buildVeniceModel,
resolveVeniceBaseModel,
venice,
VeniceResponse,
type VeniceModel,
} from "../src/sync/providers/venice.js";
const catalogModel: VeniceModel = {
id: "openai-gpt-54",
created: 1_772_668_800,
model_spec: {
name: "GPT-5.4",
availableContextTokens: 400_000,
maxCompletionTokens: 128_000,
modelSource: "OpenAI",
capabilities: {
supportsVision: true,
supportsReasoning: true,
supportsReasoningEffort: true,
reasoningEffortOptions: ["none", "low", "medium", "high"],
supportsFunctionCalling: true,
supportsResponseSchema: true,
},
pricing: {
input: { usd: 3.13 },
output: { usd: 18.75 },
cache_input: { usd: 0.313 },
extended: {
context_token_threshold: 200_000,
input: { usd: 6.26 },
output: { usd: 28.125 },
},
},
},
};
test("Venice resolves flattened IDs to canonical metadata", () => {
expect(resolveVeniceBaseModel("openai-gpt-54", "GPT-5.4")).toBe("openai/gpt-5.4");
expect(resolveVeniceBaseModel("claude-opus-4-8-fast", "Claude Opus 4.8 Fast"))
.toBe("anthropic/claude-opus-4-8");
});
test("Venice emits empty reasoning options when efforts are unavailable", () => {
const synced = buildVeniceModel({
...catalogModel,
id: "reasoning-without-efforts",
model_spec: {
...catalogModel.model_spec,
name: "Reasoning Without Efforts",
capabilities: {
...catalogModel.model_spec.capabilities,
reasoningEffortOptions: [],
},
},
}, undefined, undefined, "2026-06-10");
expect(synced).toMatchObject({ reasoning: true, reasoning_options: [] });
});
test("Venice does not infer temperature support", () => {
const synced = buildVeniceModel(catalogModel, undefined, null, "2026-06-10");
expect(synced.temperature).toBeUndefined();
});
test("Venice skips E2EE models", () => {
const translated = venice.translateModel({
...catalogModel,
id: "e2ee-test-model",
model_spec: {
...catalogModel.model_spec,
capabilities: { ...catalogModel.model_spec.capabilities, supportsE2EE: true },
},
}, { existing: () => undefined });
expect(translated).toBeUndefined();
});
test("Venice uses boundary-aware family matching", () => {
const synced = buildVeniceModel({
...catalogModel,
id: "google-gemma-4-31b-it",
model_spec: { ...catalogModel.model_spec, name: "Google Gemma 4 31B Instruct" },
}, undefined, null, "2026-06-10");
expect(synced).toMatchObject({ family: "gemma" });
});
test("Venice maps API fields without bumping inherited model timestamps", () => {
const synced = buildVeniceModel(catalogModel, {
base_model: "openai/gpt-5.4",
name: "GPT-5.4",
family: "gpt",
release_date: "2026-03-05",
last_updated: "2026-03-09",
attachment: true,
reasoning: true,
tool_call: true,
structured_output: true,
temperature: true,
open_weights: false,
interleaved: { field: "reasoning_content" },
cost: { input: 3, output: 18, input_audio: 4 },
limit: { context: 400_000, output: 128_000 },
modalities: { input: ["text", "image", "pdf"], output: ["text"] },
}, "openai/gpt-5.4", "2026-06-10");
expect(synced).toMatchObject({
base_model: "openai/gpt-5.4",
base_model_omit: ["limit.input"],
last_updated: "2026-03-09",
reasoning_options: [{ type: "effort", values: ["none", "low", "medium", "high"] }],
interleaved: { field: "reasoning_content" },
cost: {
input: 3.13,
output: 18.75,
cache_read: 0.313,
input_audio: 4,
tiers: [{ tier: { type: "context", size: 200_000 }, input: 6.26, output: 28.125 }],
},
});
expect(synced).not.toHaveProperty("family");
expect(synced).not.toHaveProperty("release_date");
expect(synced).not.toHaveProperty("open_weights");
expect(synced).not.toHaveProperty("modalities");
expect(synced).not.toHaveProperty("temperature");
});
test("Venice preserves last_updated when authoritative data is unchanged", () => {
const providerModel = {
...catalogModel,
id: "venice-only-test-model",
model_spec: { ...catalogModel.model_spec, name: "Venice Only Test Model" },
};
const full = buildVeniceModel(providerModel, undefined, undefined, "2026-06-10");
if ("base_model" in full) throw new Error("Expected a full provider model fixture");
const synced = buildVeniceModel(providerModel, full, undefined, "2026-06-11");
expect(synced).toMatchObject({ last_updated: "2026-06-10" });
});
test("Venice rejects malformed responses", () => {
expect(() => VeniceResponse.parse({ data: [{ id: "broken" }] })).toThrow();
});
test("Venice models use only canonical metadata and declare reasoning options", async () => {
const root = path.join(import.meta.dirname, "..", "..", "..");
const modelsDir = path.join(root, "providers", "venice", "models");
for (const file of readdirSync(modelsDir).filter((item) => item.endsWith(".toml"))) {
const model = Bun.TOML.parse(await Bun.file(path.join(modelsDir, file)).text()) as {
base_model?: string;
reasoning_options?: unknown[];
};
expect(model.reasoning_options, file).toBeDefined();
if (model.base_model !== undefined) {
expect(model.base_model.startsWith("venice/"), file).toBe(false);
expect(await Bun.file(path.join(root, "models", `${model.base_model}.toml`)).exists(), file).toBe(true);
}
expect(file.startsWith("e2ee-"), file).toBe(false);
}
});
-113
View File
@@ -1,113 +0,0 @@
import { expect, test } from "bun:test";
import { buildVercelModel, type VercelModel, vercel } from "../src/sync/providers/vercel.js";
const model: VercelModel = {
id: "openai/gpt-test",
name: "GPT Test",
created: 1_700_000_000,
released: 1_710_000_000,
context_window: 128_000,
max_tokens: 32_000,
type: "language",
tags: ["reasoning", "tool-use", "vision", "file-input"],
pricing: {
input: "0.000001",
output: "0.000004",
input_cache_read: "0.0000001",
},
};
test("Vercel models translate gateway metadata", () => {
const synced = buildVercelModel(model, undefined);
expect(synced).toMatchObject({
name: "GPT Test",
release_date: "2024-03-09",
last_updated: "2024-03-09",
attachment: true,
reasoning: true,
tool_call: true,
open_weights: false,
cost: { input: 1, output: 4, cache_read: 0.1 },
limit: { context: 128_000, input: 96_000, output: 32_000 },
modalities: { input: ["text", "image", "pdf"], output: ["text"] },
});
});
test("Vercel models preserve curated metadata and missing limits", () => {
const synced = buildVercelModel({
...model,
context_window: 0,
max_tokens: 0,
}, {
name: "Curated name",
release_date: "2024-01-01",
last_updated: "2025-01-01",
reasoning_options: [{ type: "effort", values: ["low", "high"] }],
cost: {
input: 2,
output: 8,
tiers: [{
tier: { type: "context", size: 200_000 },
input: 3,
output: 12,
}],
},
limit: { context: 64_000, input: 48_000, output: 16_000 },
});
expect(synced.name).toBe("Curated name");
expect(synced.last_updated).toBe("2025-01-01");
expect(synced.reasoning_options).toEqual([{ type: "effort", values: ["low", "high"] }]);
expect(synced.cost?.tiers).toHaveLength(1);
expect(synced.limit).toEqual({ context: 64_000, input: 48_000, output: 16_000 });
});
test("Vercel non-language models use API tool capabilities", () => {
const synced = buildVercelModel({
...model,
type: "image",
tags: [],
}, {
tool_call: true,
});
expect(synced.tool_call).toBe(false);
});
test("Vercel sync includes non-language model types", () => {
for (const [type, output] of [
["image", ["image"]],
["video", ["video"]],
["reranking", ["text"]],
] as const) {
const source = {
...model,
id: `test/${type}`,
type,
tags: [],
context_window: 0,
max_tokens: 0,
pricing: undefined,
};
expect(vercel.translateModel(source, { existing: () => undefined })).toBeDefined();
expect(buildVercelModel(source, undefined)).toMatchObject({
tool_call: false,
modalities: { input: ["text"], output },
});
}
});
test("Vercel models use canonical metadata when available", () => {
const synced = buildVercelModel({
...model,
id: "nvidia/nemotron-3-ultra-550b-a55b",
name: "Nemotron 3 Ultra",
}, undefined);
expect("base_model" in synced ? synced.base_model : undefined)
.toBe("nvidia/nemotron-3-ultra-550b-a55b");
expect("last_updated" in synced ? synced.last_updated : undefined).toBeUndefined();
});
-30
View File
@@ -1,30 +0,0 @@
import { expect, test } from "bun:test";
import { buildXAIModel, type XAIModel } from "../src/sync/providers/xai.js";
const model: XAIModel = {
id: "grok-test",
created: 1_700_000_000,
input_modalities: ["text"],
output_modalities: ["text"],
prompt_text_token_price: 10_000,
completion_text_token_price: 20_000,
};
test("xAI sync preserves reasoning options", () => {
const synced = buildXAIModel(model, {
name: "Grok Test",
release_date: "2024-01-01",
last_updated: "2024-01-01",
attachment: false,
reasoning: true,
reasoning_options: [{ type: "effort", values: ["none", "high"] }],
tool_call: true,
open_weights: false,
limit: { context: 2_000_000, output: 30_000 },
modalities: { input: ["text"], output: ["text"] },
});
expect(synced.reasoning_options).toEqual([{ type: "effort", values: ["none", "high"] }]);
expect(synced.limit?.context).toBe(2_000_000);
});
@@ -3,6 +3,7 @@ release_date = "2026-02-18"
last_updated = "2026-03-13"
attachment = true
reasoning = true
reasoning_options = []
temperature = true
tool_call = true
open_weights = false
@@ -4,6 +4,7 @@ release_date = "2026-02-18"
last_updated = "2026-03-13"
attachment = true
reasoning = true
reasoning_options = [{ type = "toggle" }, { type = "budget_tokens", min = 1024, max = 63999 }]
temperature = true
tool_call = true
open_weights = false
@@ -4,6 +4,7 @@ release_date = "2025-01-20"
last_updated = "2025-01-20"
attachment = false
reasoning = true
reasoning_options = []
temperature = true
tool_call = true
open_weights = false
@@ -3,6 +3,7 @@ release_date = "2025-12-01"
last_updated = "2025-12-01"
attachment = false
reasoning = true
reasoning_options = []
temperature = true
tool_call = true
open_weights = false
@@ -3,6 +3,7 @@ release_date = "2025-07-15"
last_updated = "2025-07-15"
attachment = true
reasoning = true
reasoning_options = []
temperature = true
tool_call = true
open_weights = false
@@ -3,6 +3,7 @@ release_date = "2025-09-05"
last_updated = "2025-09-05"
attachment = false
reasoning = true
reasoning_options = []
temperature = true
tool_call = true
open_weights = false
@@ -3,6 +3,7 @@ release_date = "2025-09-05"
last_updated = "2025-09-05"
attachment = false
reasoning = true
reasoning_options = []
temperature = true
tool_call = true
open_weights = false
+1
View File
@@ -4,6 +4,7 @@ release_date = "2026-05-01"
last_updated = "2026-05-01"
attachment = true
reasoning = true
reasoning_options = [{ type = "effort", values = ["none", "low", "medium", "high"] }]
temperature = true
tool_call = true
structured_output = true
+1
View File
@@ -4,6 +4,7 @@ release_date = "2026-01"
last_updated = "2026-01"
attachment = true
reasoning = true
reasoning_options = [{ type = "toggle" }]
temperature = false
tool_call = true
structured_output = true
+1
View File
@@ -4,6 +4,7 @@ release_date = "2026-04-21"
last_updated = "2026-04-21"
attachment = true
reasoning = true
reasoning_options = [{ type = "toggle" }]
temperature = false
tool_call = true
structured_output = true
@@ -4,6 +4,7 @@ release_date = "2026-04-02"
last_updated = "2026-04-02"
attachment = true
reasoning = true
reasoning_options = [{ type = "toggle" }]
temperature = true
tool_call = true
structured_output = true
@@ -4,6 +4,7 @@ release_date = "2026-05-09"
last_updated = "2026-05-09"
attachment = false
reasoning = true
reasoning_options = [{ type = "toggle" }]
temperature = true
tool_call = true
structured_output = true
@@ -4,6 +4,7 @@ release_date = "2026-05-09"
last_updated = "2026-05-09"
attachment = true
reasoning = true
reasoning_options = [{ type = "toggle" }]
temperature = true
tool_call = true
structured_output = true
@@ -0,0 +1,14 @@
base_model = "zhipuai/glm-5.2"
[[reasoning_options]]
type = "effort"
values = ["high", "max"]
[interleaved]
field = "reasoning_content"
[cost]
input = 0
output = 0
cache_read = 0
cache_write = 0
@@ -0,0 +1,13 @@
base_model = "moonshotai/kimi-k2.7-code"
base_model_omit = ["structured_output"]
family = "kimi-k2"
reasoning_options = [{ type = "toggle" }, { type = "budget_tokens" }]
[interleaved]
field = "reasoning_content"
[cost]
input = 0
output = 0
cache_read = 0
cache_write = 0
@@ -0,0 +1,14 @@
base_model = "zhipuai/glm-5.2"
[[reasoning_options]]
type = "effort"
values = ["high", "max"]
[interleaved]
field = "reasoning_content"
[cost]
input = 0
output = 0
cache_read = 0
cache_write = 0
@@ -0,0 +1,13 @@
base_model = "moonshotai/kimi-k2.7-code"
base_model_omit = ["structured_output"]
family = "kimi-k2"
reasoning_options = [{ type = "toggle" }, { type = "budget_tokens" }]
[interleaved]
field = "reasoning_content"
[cost]
input = 0
output = 0
cache_read = 0
cache_write = 0
@@ -0,0 +1,8 @@
base_model = "moonshotai/kimi-k2.7-code"
temperature = true
[cost]
input = 0.75
output = 3.5
cache_read = 0.16
cache_write = 0
@@ -1,5 +1,6 @@
reasoning_options = []
base_model = "zhipuai/glm-5.1"
name = "GLM 5.1"
reasoning_options = []
[interleaved]
field = "reasoning_content"
@@ -1 +0,0 @@
../../azure/models/kimi-k2.5.toml
@@ -0,0 +1,16 @@
base_model = "moonshotai/kimi-k2.5"
reasoning_options = [{ type = "toggle" }]
temperature = true
interleaved = true
[cost]
input = 0.60
output = 3.00
[modalities]
input = ["text", "image"]
[provider]
shape = "completions"
npm = "@ai-sdk/openai-compatible"
api = "https://${AZURE_COGNITIVE_SERVICES_RESOURCE_NAME}.services.ai.azure.com/models"
@@ -1,30 +1,16 @@
name = "Kimi K2.6"
family = "kimi-k2"
release_date = "2026-04-22"
last_updated = "2026-04-22"
base_model = "moonshotai/kimi-k2.6"
attachment = false
reasoning = true
reasoning_options = [{ type = "toggle" }]
structured_output = true
temperature = true
tool_call = true
knowledge = "2025-01"
interleaved = true
open_weights = true
[cost]
input = 0.95
output = 4.00
[limit]
context = 262_144
output = 262_144
[modalities]
input = ["text", "image"]
output = ["text"]
[provider]
shape = "completions"
npm = "@ai-sdk/openai-compatible"
api = "https://${AZURE_RESOURCE_NAME}.services.ai.azure.com/models"
api = "https://${AZURE_COGNITIVE_SERVICES_RESOURCE_NAME}.services.ai.azure.com/models"
@@ -0,0 +1,7 @@
base_model = "openai/gpt-image-1.5"
reasoning_options = []
[cost]
input = 5.00
output = 32.00
cache_read = 1.25
+7
View File
@@ -0,0 +1,7 @@
base_model = "openai/gpt-image-1"
reasoning_options = []
[cost]
input = 5.00
output = 40.00
cache_read = 1.25
+7
View File
@@ -0,0 +1,7 @@
base_model = "openai/gpt-image-2"
reasoning_options = []
[cost]
input = 5.00
output = 30.00
cache_read = 1.25
@@ -0,0 +1,14 @@
base_model = "moonshotai/kimi-k2.7-code"
temperature = true
[cost]
input = 0.95
output = 4
cache_read = 0.16
[limit]
context = 262_000
output = 262_000
[modalities]
input = ["text", "image"]
@@ -10,7 +10,7 @@ type = "toggle"
field = "reasoning_content"
[cost]
input = 0.06
input = 0.3
output = 0.75
cache_read = 0.06
@@ -17,7 +17,7 @@ type = "toggle"
field = "reasoning_content"
[cost]
input = 0.12
input = 0.6
output = 2.2
cache_read = 0.12
@@ -0,0 +1,16 @@
base_model = "zhipuai/glm-5.2"
name = "GLM 5.2"
[[reasoning_options]]
type = "toggle"
[interleaved]
field = "reasoning_content"
[cost]
input = 1.5
output = 4.5
cache_read = 0.3
[limit]
context = 131_072
+21
View File
@@ -0,0 +1,21 @@
name = "Claudius"
release_date = "2026-05-12"
last_updated = "2026-05-12"
attachment = true
reasoning = true
tool_call = true
open_weights = false
knowledge = "2026-05"
[cost]
input = 3.00
output = 8.00
cache_read = 0.900
[limit]
context = 256_000
output = 64_000
[modalities]
input = ["text", "image", "audio", "video"]
output = ["text"]
@@ -0,0 +1,11 @@
base_model = "zhipuai/glm-5.2"
name = "Glm 5.2"
[cost]
input = 1.4
output = 4.4
cache_read = 0.26
[limit]
context = 262_144
output = 262_144
@@ -1,7 +1,7 @@
name = "Command R7B"
family = "command-r"
release_date = "2024-02-27"
last_updated = "2024-02-27"
release_date = "2024-12-02"
last_updated = "2024-12-02"
attachment = false
reasoning = false
temperature = true
@@ -1,22 +0,0 @@
name = "Devstral Small 2 2512"
release_date = "2025-12-09"
last_updated = "2025-12-09"
knowledge = "2025-12"
attachment = false
reasoning = false
tool_call = true
temperature = true
open_weights = true
[cost]
input = 0
output = 0
[limit]
context = 262_000
output = 262_000
[modalities]
input = ["text", "image"]
output = ["text"]
@@ -0,0 +1,7 @@
base_model = "zhipuai/glm-5-turbo"
[cost]
input = 1.235
output = 4.118
cache_read = 0.308
cache_write = 1.544
+10
View File
@@ -0,0 +1,10 @@
base_model = "zhipuai/glm-5.2"
reasoning_options = [{ type = "effort", values = ["high", "max"]}]
[interleaved]
field = "reasoning_content"
[cost]
input = 1.44
output = 4.53
cache_read = 0.39
@@ -0,0 +1,7 @@
base_model = "zhipuai/glm-5v-turbo"
[cost]
input = 1.235
output = 4.118
cache_read = 0.308
cache_write = 1.544
+1
View File
@@ -1,5 +1,6 @@
base_model = "openai/gpt-5.4"
base_model_omit = ["limit.input"]
reasoning_options = [{ type = "effort", values = ["low", "medium", "high"] }]
[cost]
input = 3
+2 -1
View File
@@ -5,6 +5,7 @@ last_updated = "2025-08-05"
knowledge = "2024-01"
attachment = false
reasoning = false
reasoning_options = [{ type = "effort", values = ["low", "medium", "high"] }]
tool_call = true
temperature = true
open_weights = true
@@ -20,4 +21,4 @@ output = 128_000
[modalities]
input = ["text"]
output = ["text"]
output = ["text"]
@@ -4,6 +4,7 @@ last_updated = "2025-08-26"
knowledge = "2023-12"
attachment = false
reasoning = true
reasoning_options = []
tool_call = true
temperature = true
open_weights = true
+2 -1
View File
@@ -4,6 +4,7 @@ last_updated = "2025-11-26"
knowledge = "2025-11"
attachment = true
reasoning = true
reasoning_options = []
tool_call = true
temperature = true
open_weights = true
@@ -19,4 +20,4 @@ output = 128_000
[modalities]
input = ["text"]
output = ["text"]
output = ["text"]
@@ -4,6 +4,7 @@ last_updated = "2025-12-08"
knowledge = "2025-12"
attachment = true
reasoning = true
reasoning_options = []
tool_call = true
temperature = true
open_weights = true
@@ -21,4 +22,4 @@ output = 262_000
[modalities]
input = ["text"]
output = ["text"]
output = ["text"]
+1
View File
@@ -5,6 +5,7 @@ last_updated = "2026-01-27"
knowledge = "2025-01"
attachment = false
reasoning = true
reasoning_options = [{ type = "toggle" }]
temperature = true
tool_call = true
open_weights = true
+1
View File
@@ -4,6 +4,7 @@ release_date = "2026-04-17"
last_updated = "2026-04-17"
attachment = false
reasoning = true
reasoning_options = [{ type = "toggle" }]
temperature = true
tool_call = true
open_weights = true
@@ -0,0 +1,10 @@
base_model = "moonshotai/kimi-k2.7-code"
reasoning_options = []
[interleaved]
field = "reasoning_content"
[cost]
input = 1.28
output = 4.63
cache_read = 0.32
@@ -2,6 +2,7 @@ base_model = "meta/llama-3.3-70b-instruct"
name = "Llama 3.3 70B Instruct"
attachment = false
reasoning = true
reasoning_options = []
[cost]
input = 0.089
@@ -0,0 +1,7 @@
base_model = "meta/llama-4-maverick-17b-instruct"
[cost]
input = 0.124
output = 0.603
cache_read = 0.03
cache_write = 0.151
+6
View File
@@ -0,0 +1,6 @@
base_model = "minimax/MiniMax-M3"
[cost]
input = 0.355
output = 1.775
cache_read = 0.089
@@ -4,6 +4,7 @@ last_updated = "2023-12-11"
knowledge = "2023-09"
attachment = false
reasoning = true
reasoning_options = []
tool_call = false
temperature = true
open_weights = true
@@ -5,6 +5,7 @@ last_updated = "2026-03-11"
knowledge = "2025-12"
attachment = false
reasoning = true
reasoning_options = []
tool_call = true
temperature = true
open_weights = true
@@ -1,4 +1,5 @@
base_model = "deepseek/deepseek-v4-pro"
name = "DeepSeek V4 Pro Lightning"
reasoning_options = [{ type = "effort", values = ["none", "low", "medium", "high"] }]
[interleaved]
+16
View File
@@ -0,0 +1,16 @@
base_model = "zhipuai/glm-5.2"
reasoning_options = [{ type = "effort", values = ["none", "low", "medium", "high"] }]
[interleaved]
field = "reasoning_content"
[cost]
input = 0.5
output = 2.2
cache_read = 0.08
[limit]
output = 131_072
[provider]
npm = "@ai-sdk/openai-compatible"
@@ -1,6 +1,6 @@
name = "Greg 1 Super"
release_date = "2026-01-27"
last_updated = "2026-01-27"
name = "Greg 2 Super"
release_date = "2026-06-14"
last_updated = "2026-06-14"
attachment = false
reasoning = false
temperature = true
@@ -8,9 +8,9 @@ tool_call = false
open_weights = false
[cost]
input = 1.00
output = 5.00
cache_read = 0.20
input = 1.5
output = 5
cache_read = 0.25
[limit]
context = 229_376
@@ -1,6 +1,6 @@
name = "Greg 1 Normal"
release_date = "2026-01-27"
last_updated = "2026-01-27"
name = "Greg 2 Ultra"
release_date = "2026-06-14"
last_updated = "2026-06-14"
attachment = false
reasoning = false
temperature = true
@@ -8,9 +8,9 @@ tool_call = false
open_weights = false
[cost]
input = 0.10
output = 0.30
cache_read = 0.02
input = 3
output = 10
cache_read = 0.5
[limit]
context = 229_376
@@ -18,4 +18,4 @@ output = 229_376
[modalities]
input = ["text"]
output = ["text"]
output = ["text"]
+13
View File
@@ -0,0 +1,13 @@
base_model = "moonshotai/kimi-k2.7-code"
reasoning_options = [{ type = "effort", values = ["none", "low", "medium", "high"] }]
[interleaved]
field = "reasoning_content"
[cost]
input = 0.55
output = 2.25
cache_read = 0.05
[provider]
npm = "@ai-sdk/openai-compatible"
-3
View File
@@ -1,3 +0,0 @@
<svg width="24" height="24" viewBox="0 0 24 24" fill="none" xmlns="http://www.w3.org/2000/svg">
<path d="M14.45 5.74999L11.9991 11.6956L9.54562 5.74999H7.97237L10.6604 12.2495C10.7678 12.5138 10.9515 12.7402 11.1882 12.8996C11.4248 13.059 11.7036 13.1442 11.9889 13.1444C12.2742 13.1446 12.5531 13.0597 12.79 12.9006C13.0268 12.7415 13.2109 12.5154 13.3186 12.2512L16.0232 5.74999H14.45ZM15.4965 14.808L19.98 10.2195L19.3684 8.75911L14.4719 13.7807C14.272 13.9856 14.1369 14.2448 14.0835 14.5261C14.0301 14.8073 14.0608 15.098 14.1718 15.3619C14.2807 15.624 14.4648 15.848 14.7008 16.0056C14.9369 16.1632 15.2144 16.2473 15.4983 16.2474L15.5 16.25L22.5 16.2325L21.8884 14.7721L15.4983 14.808H15.4965ZM4.02 10.216L4.63162 8.75561L9.52813 13.7772C9.93763 14.1964 10.0557 14.8176 9.82825 15.3584C9.71925 15.6204 9.53511 15.8443 9.29905 16.0019C9.06299 16.1595 8.78557 16.2437 8.50175 16.2439L1.50175 16.2281L1.5 16.2299L2.11163 14.7695L8.50175 14.8062L4.02 10.216Z" fill="currentColor"/>
</svg>

Before

Width:  |  Height:  |  Size: 992 B

@@ -1,28 +0,0 @@
name = "Kimi K2.6 Turbo"
family = "kimi-thinking"
release_date = "2026-04-17"
last_updated = "2026-04-17"
attachment = false
reasoning = true
temperature = true
tool_call = true
open_weights = true
[[reasoning_options]]
type = "toggle"
[cost]
input = 0
output = 0
cache_read = 0
[limit]
context = 262_000
output = 262_000
[modalities]
input = ["text", "image"]
output = ["text"]
[interleaved]
field = "reasoning_content"
-5
View File
@@ -1,5 +0,0 @@
name = "Fireworks (Firepass)"
env = ["FIREPASS_API_KEY"]
npm = "@ai-sdk/openai-compatible"
api = "https://api.fireworks.ai/inference/v1/"
doc = "https://docs.fireworks.ai/firepass"
@@ -1,4 +1,5 @@
base_model = "deepseek/deepseek-v4-flash"
last_updated = "2026-06-16"
[[reasoning_options]]
type = "toggle"
@@ -13,4 +14,4 @@ field = "reasoning_content"
[cost]
input = 0.14
output = 0.28
cache_read = 0.03
cache_read = 0.028
@@ -0,0 +1,32 @@
name = "GLM 5.2"
family = "glm"
release_date = "2026-06-16"
last_updated = "2026-06-16"
attachment = false
reasoning = true
temperature = true
tool_call = true
open_weights = true
[[reasoning_options]]
type = "toggle"
[[reasoning_options]]
type = "effort"
values = ["high", "max"]
[cost]
input = 1.40
output = 4.40
cache_read = 0.26
[limit]
context = 1_048_576
output = 131_072
[modalities]
input = ["text"]
output = ["text"]
[interleaved]
field = "reasoning_content"
@@ -1,7 +1,7 @@
name = "GPT OSS 120B"
family = "gpt-oss"
release_date = "2025-08-05"
last_updated = "2025-08-05"
last_updated = "2026-06-16"
attachment = false
reasoning = true
temperature = true
@@ -1,5 +1,12 @@
base_model = "moonshotai/kimi-k2.7-code"
name = "Kimi K2.7 Code"
family = "kimi-k2"
release_date = "2026-06-12"
last_updated = "2026-06-16"
attachment = true
reasoning = true
temperature = true
tool_call = true
open_weights = true
[[reasoning_options]]
type = "toggle"
@@ -1,13 +1,19 @@
base_model = "moonshotai/kimi-k2.7-code"
name = "Kimi K2.7 Code Fast"
family = "kimi-k2"
release_date = "2026-06-12"
last_updated = "2026-06-16"
attachment = true
reasoning = true
temperature = true
tool_call = true
open_weights = true
[[reasoning_options]]
type = "toggle"
[cost]
cache_read = 0.38
input = 2.00
input = 1.90
output = 8.00
[limit]
+1
View File
@@ -3,6 +3,7 @@ release_date = "1970-01-01"
last_updated = "1970-01-01"
attachment = false
reasoning = true
reasoning_options = []
temperature = true
tool_call = true
structured_output = true
+1
View File
@@ -3,6 +3,7 @@ release_date = "1970-01-01"
last_updated = "1970-01-01"
attachment = false
reasoning = true
reasoning_options = []
temperature = true
tool_call = true
structured_output = true
@@ -2,6 +2,7 @@ name = "Qwen 3.6 Plus"
family = "qwen"
attachment = true
reasoning = true
reasoning_options = []
tool_call = true
temperature = true
release_date = "2026-04-02"
@@ -1,5 +1,4 @@
base_model = "anthropic/claude-fable-5"
reasoning_options = [{ type = "effort", values = ["low", "medium", "high", "xhigh", "max"] }]
[cost]
input = 10
@@ -1,5 +1,4 @@
base_model = "anthropic/claude-haiku-4-5"
reasoning_options = [{ type = "budget_tokens", min = 1_024, max = 32_000 }]
[cost]
input = 1
@@ -1,5 +1,4 @@
base_model = "anthropic/claude-opus-4-5"
reasoning_options = [{ type = "budget_tokens", min = 1_024, max = 32_000 }]
[cost]
input = 5
@@ -1,5 +1,4 @@
base_model = "anthropic/claude-opus-4-6"
reasoning_options = [{ type = "effort", values = ["low", "medium", "high", "max"] }]
[cost]
input = 5
@@ -1,5 +1,4 @@
base_model = "anthropic/claude-opus-4-7"
reasoning_options = [{ type = "effort", values = ["low", "medium", "high", "xhigh", "max"] }]
[cost]
input = 5
@@ -1,5 +1,4 @@
base_model = "anthropic/claude-opus-4-8"
reasoning_options = [{ type = "effort", values = ["low", "medium", "high", "xhigh", "max"] }]
[cost]
input = 5
@@ -1,5 +1,4 @@
base_model = "anthropic/claude-sonnet-4-5"
reasoning_options = [{ type = "budget_tokens", min = 1_024, max = 32_000 }]
[cost]
input = 3
@@ -1,5 +1,4 @@
base_model = "anthropic/claude-sonnet-4-6"
reasoning_options = [{ type = "effort", values = ["low", "medium", "high", "max"] }, { type = "budget_tokens", min = 1_024, max = 32_000 }]
[cost]
input = 3
@@ -1,5 +1,4 @@
base_model = "anthropic/claude-sonnet-4-0"
reasoning_options = []
[cost]
input = 3
@@ -1,4 +1,5 @@
base_model = "google/gemini-2.5-pro"
reasoning_options = [{ type = "budget_tokens", min = 128, max = 32_768 }]
[cost]
input = 1.25
@@ -1,4 +1,5 @@
base_model = "google/gemini-3-flash-preview"
reasoning_options = [{ type = "effort", values = ["low", "medium", "high"] }, { type = "budget_tokens", min = 256, max = 32_000 }]
[cost]
input = 0.5
@@ -1,4 +1,5 @@
base_model = "google/gemini-3.1-pro-preview"
reasoning_options = [{ type = "effort", values = ["low", "medium", "high"] }, { type = "budget_tokens", min = 256, max = 32_000 }]
[cost]
input = 2

Some files were not shown because too many files have changed in this diff Show More