Compare commits

...

1385 Commits

Author SHA1 Message Date
Jack f4f210986c fix opencode go MiMo pricing conflict resolution 2026-05-28 00:36:48 +08:00
Jack f4206f4eaa Merge branch 'dev' into update/opencode-go-mimo-v2-5-pricing 2026-05-28 00:34:02 +08:00
Jack d11b0e448c update opencode go MiMo V2.5 pricing 2026-05-27 23:54:58 +08:00
Aiden Cline ec4ec6d441 Merge pull request #1867 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-05-27 00:00:17 -05:00
Aiden Cline d96da5ece6 Merge pull request #1868 from Ardakilic/chore/sync-kilo-models-20260527-1
Sync Kilo models with upstream gateway
2026-05-26 23:59:26 -05:00
github-actions[bot] 0342c79d03 chore(sync): update OpenRouter model catalog 2026-05-27 03:26:29 +00:00
Arda Kilicdagi 37136fc2f3 chore: sync upstream kilo api gateway models 2026-05-27 02:27:53 +04:00
Aiden Cline 989939773b Merge pull request #1863 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-05-26 13:04:26 -05:00
Frank 4f1d5c511a update go models 2026-05-26 13:54:19 -04:00
github-actions[bot] 34a21e07f2 chore(sync): update OpenRouter model catalog 2026-05-26 17:26:26 +00:00
Aiden Cline 979f5da0ca Merge pull request #1864 from oskarkocol/update/openai-cache-rates
chore: update openai cache rates
2026-05-26 12:06:15 -05:00
oskar 73b357e2ab update cache rates 2026-05-26 23:28:11 +07:00
Aiden Cline 97c71f1663 Merge pull request #1862 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-05-26 09:11:54 -05:00
github-actions[bot] 59c270bc85 chore(sync): update OpenRouter model catalog 2026-05-26 13:23:20 +00:00
Aiden Cline a4c18f88ed Merge pull request #1860 from bas3line/sync-routing-run-models-2
Sync routing.run model catalog
2026-05-25 23:44:24 -05:00
bas3line 3343a585ac fix(routing-run): sync model catalog 2026-05-26 09:51:14 +05:30
Aiden Cline f23db95550 Merge pull request #1858 from smakosh/fix/llmgateway-qwen3.7-max-id
fix(llmgateway): correct Qwen3.7 Max model id to qwen3.7-max
2026-05-25 17:22:56 -05:00
Aiden Cline 556d9a9045 Merge pull request #1859 from Suat-B/update-xpersona-frieren-coder-limits
Update Xpersona Frieren Coder limits
2026-05-25 17:22:46 -05:00
Frank 49991c8f8f update zen models 2026-05-25 17:56:52 -04:00
Suat-B 20c1e810ce Update Xpersona Frieren Coder limits 2026-05-25 13:23:59 -05:00
smakosh 15d08b4540 fix(llmgateway): correct Qwen3.7 Max model id to qwen3.7-max
Rename qwen37-max.toml to qwen3.7-max.toml so the model id matches
the canonical alibaba/qwen3.7-max definition (name: Qwen3.7 Max).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 18:48:12 +01:00
Aiden Cline 14bbc303e5 Merge pull request #1852 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-05-25 09:49:26 -05:00
Aiden Cline 02da63ed48 Merge pull request #1853 from dpuyosa/chore/venice-pricing
Venice: Update pricing and limits for 4 models
2026-05-25 09:49:09 -05:00
github-actions[bot] 07daef8c54 chore(sync): update OpenRouter model catalog 2026-05-25 13:27:22 +00:00
dpuyosa b6ebe8696c [venice] Update pricing, limits, and last_updated for 4 models
- Reduce input/output/cache prices for gemini-3-5-flash, google-gemma-4-31b-it, and qwen-3-7-max
- Lower max output tokens from 65,536 to 16,384 for qwen3-5-35b-a3b
- Bump last_updated to 2026-05-25 for all 4 models
2026-05-25 10:07:32 +02:00
Aiden Cline ad654e71da Merge pull request #1850 from huxeon/dev
fix: modify the deepseek v4 flash/pro price
2026-05-25 00:18:18 -05:00
Aiden Cline 508f4d48e1 Merge pull request #1847 from ceyhanmolla/poolside/laguna-direct
Add Poolside provider with Laguna M.1 and XS.2 models
2026-05-24 23:26:34 -05:00
opencode-agent[bot] cd7c70b4fe revert: remove opencode-go deepseek-v4-pro price changes
Keep only the deepseek provider price updates as intended.
2026-05-25 04:26:14 +00:00
Aiden Cline 10e752ea84 Merge pull request #1845 from yukoba/vultr
Update Vultr models
2026-05-24 23:26:12 -05:00
huxeon 4cdb4c700b fix: modify the deepseek v4 flash/pro price in provider deepseek and opencode-go 2026-05-24 13:33:56 +08:00
Aiden Cline d497a446eb Merge pull request #1849 from technoabsurdist/add-wafer-ai-qwen3.6-35b-a3b-and-kimi-k2.6
providers/wafer.ai: add Qwen3.6-35B-A3B and Kimi-K2.6
2026-05-23 16:41:24 -05:00
Emilio Andere 745cf557b3 feat(wafer.ai): add Qwen3.6-35B-A3B and Kimi-K2.6
Both models are public serverless on pass.wafer.ai/v1/models but were
missing from the wafer.ai provider in models.dev, so OpenCode and other
tools that pull from the registry could not discover them.

- Qwen3.6 35B A3B: compact MoE, 32K context, vision-capable, $0.19/M in,
  $1.25/M out (NVFP4 on AMD MI355X — see wafer.ai/blog/qwen36-mi355x).
- Kimi K2.6: 1T sparse MoE, 262K context, vision-capable, $1.10/M in,
  $4.80/M out (NVFP4 on Blackwell — see wafer.ai/blog/kimi-k26-nvfp4).

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-05-23 13:26:18 -04:00
Yu Kobayashi db295d6842 Refactor to use the extends syntax 2026-05-24 01:20:31 +09:00
ceyhanmolla 47c0eed41e Add Poolside provider with Laguna M.1 and XS.2 models 2026-05-23 17:09:33 +02:00
Yu Kobayashi e690980857 Update Vultr models 2026-05-23 16:25:11 +09:00
Aiden Cline f5f7d1a167 Merge pull request #1840 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-05-22 16:58:23 -05:00
Aiden Cline cfbbb55c4d Merge pull request #1839 from anthraxx/alibaba-qwen3.7-plus
Alibaba: Add Qwen 3.7 Max and 3.6 Flash to all regions and plans
2026-05-22 16:58:08 -05:00
github-actions[bot] 1956762cdf chore(sync): update OpenRouter model catalog 2026-05-22 21:41:35 +00:00
Levente Polyak 28571433f7 feat(alibaba): add Qwen3.7 Max model configuration to all regions
Link: https://bailian.console.alibabacloud.com/cn-beijing?tab=model#/model-market/detail/qwen3.7-max?serviceSite=asia-pacific-china
2026-05-22 19:00:29 +02:00
Levente Polyak aef5e48bae feat(alibaba): add Qwen3.6 Flash model configuration to all regions
Link: https://bailian.console.alibabacloud.com/cn-beijing?tab=model#/model-market/detail/qwen3.6-flash?serviceSite=asia-pacific-china
2026-05-22 18:54:22 +02:00
Aiden Cline 8ba19639a7 Merge pull request #1825 from shzdehmd/dev
update(fireworks): sync models and pricing with current offerings
2026-05-22 11:45:45 -05:00
Aiden Cline 0f9b4c9edc Merge pull request #1834 from monotykamary/chore/update-neuralwatt-qwen3.6-pricing
fix(neuralwatt): update Qwen3.6 pricing to match API
2026-05-22 09:13:02 -05:00
Aiden Cline 2a9b6256dc Merge pull request #1836 from PierreLeGuen/nearai-provider
Add current NEAR AI Cloud models
2026-05-22 09:12:36 -05:00
Aiden Cline 51633fe106 Merge pull request #1838 from Quentinchampenois/fix/update-models-scaleway
fix: update Scaleway provider models list
2026-05-22 09:12:26 -05:00
Aiden Cline 169b5c4331 Merge pull request #1833 from NicoAvanzDev/add-copilot-gemini-3-5-flash
[GitHub Copilot] add Gemini 3.5 Flash
2026-05-22 09:09:25 -05:00
Aiden Cline 9d0f0d6d56 Merge pull request #1837 from fydrah/fix/google-vertex-gemini-3.5-flash
fix: missing google-vertex gemini 3.5 flash extend
2026-05-22 09:07:34 -05:00
Aiden Cline 1d2af0c97b Merge pull request #1835 from dpuyosa/feat/venice-models
Venice: Add Gemini 3.5 Flash and Qwen 3.7 Max, update Grok Build 0.1 pricing
2026-05-22 09:07:03 -05:00
Aiden Cline cc14ae4370 Merge pull request #1832 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-05-22 09:06:16 -05:00
github-actions[bot] ad9a3448d2 chore(sync): update OpenRouter model catalog 2026-05-22 14:03:32 +00:00
Quentin Champenois c6f9cf58e4 fix(scaleway): add new models gemma-4-26b-a4b-it and qwen3.6-35b-a3b 2026-05-22 15:43:11 +02:00
Quentin Champenois bba5809471 fix(scaleway): Extends existing models 2026-05-22 15:42:36 +02:00
Quentin Champenois c0852f1b4f fix(scaleway): Clear removed models from list 2026-05-22 15:11:44 +02:00
Flavien Hardy 0fc3a3f635 fix: missing google-vertex gemini 3.5 flash extend 2026-05-22 08:56:44 -04:00
Pierre LE GUEN 68936dc062 Add current NEAR AI Cloud models 2026-05-22 09:27:57 +00:00
dpuyosa ade9760060 [venice] Add Gemini 3.5 Flash and Qwen 3.7 Max, update Grok Build 0.1 pricing
- Add Gemini 3.5 Flash model (1M context, multimodal input)
- Add Qwen 3.7 Max model (1M context, text-only)
- Update Grok Build 0.1 cost tiers and pricing
2026-05-22 10:57:38 +02:00
Tom X Nguyen 2a8b90b197 fix(neuralwatt): update Qwen3.6 pricing to match API
Updates the per-token cost for Qwen3.6-35B-A3B and qwen3.6-35b-fast from /bin/bash.05//bin/bash.10 to /bin/bash.29/.15 (input/output per million tokens), matching the actual Neuralwatt API pricing as reflected in pi-neuralwatt-provider commit f634286.
2026-05-22 15:19:51 +07:00
NicoAvanzDev 8569f0dfef [GitHub Copilot] add Gemini 3.5 Flash 2026-05-22 07:57:10 +00:00
Aiden Cline fc98ceb72e fix sync workflow force lease 2026-05-21 23:52:34 -05:00
Aiden Cline be7c5afc94 Merge pull request #1831 from anomalyco/fix-vertex-sonnet-4-6-limits
Fix Vertex Sonnet 4.6 token limits
2026-05-21 23:38:18 -05:00
Aiden Cline 57e62c43b0 fix vertex sonnet 4.6 limits 2026-05-21 23:37:30 -05:00
Aiden Cline 0898c35c9f Merge pull request #1830 from zainhas/dev
[Together AI] add Qwen3.7 max
2026-05-21 21:00:51 -05:00
Zain Hasan 46b23fb313 Merge branch 'anomalyco:dev' into dev 2026-05-21 17:47:35 -07:00
Zain Hasan 05fedc76cc [Together AI] add Qwen3.7 2026-05-21 17:47:19 -07:00
Aiden Cline 2738f81d1a Merge pull request #1828 from anomalyco/refactor/sync-core-layout
refactor: move sync implementation into core src
2026-05-21 18:10:09 -05:00
Aiden Cline 89b834086a refactor: move sync implementation into core src 2026-05-21 18:06:25 -05:00
Aiden Cline 1ab2ff8163 Merge pull request #1826 from smakosh/add-llmgateway-models
feat: add new LLM Gateway text models
2026-05-21 17:58:29 -05:00
Frank 9468676683 update zen models 2026-05-21 18:42:36 -04:00
Claude 6cdd2f054b Merge upstream/dev into add-llmgateway-models; resolve gemini-3.5-flash conflict
# Conflicts:
#	providers/google/models/gemini-3.5-flash.toml
2026-05-21 21:48:48 +00:00
Aiden Cline b13abc9141 Merge pull request #1827 from anomalyco/update-xai-pricing
fix xAI long-context pricing
2026-05-21 16:45:43 -05:00
Aiden Cline e5ba264751 fix xAI long-context pricing 2026-05-21 16:41:36 -05:00
smakosh a7811fb522 refactor: use extends for gemini and qwen models
Add canonical google/gemini-3.5-flash and alibaba/qwen3.7-max defs and
have the llmgateway entries extend them, per PR review.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 23:38:57 +02:00
smakosh 605fae75d9 feat: add new LLM Gateway text models
Add Grok 4.20 (reasoning/non-reasoning), Gemini 3.5 Flash, and Qwen3.7 Max to the llmgateway provider.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-21 23:04:47 +02:00
Ahmad Shahzad bcab0885bc update(fireworks): sync models and pricing with current offerings
Removed — 11 deprecated models:
- deepseek-v3p1
- deepseek-v3p2
- glm-4p5
- glm-4p5-air
- glm-4p7
- glm-5
- kimi-k2-instruct
- kimi-k2-thinking
- minimax-m2p1
- routers/kimi-k2p5-turbo

Added — 2 new Turbo (routers) models:
- routers/glm-5p1-fast
- routers/kimi-k2p6-turbo

Modified — pricing fixes:
- deepseek-v4-pro: cache_read 0.15 → 0.145
- gpt-oss-120b: added cache_read = 0.015
- gpt-oss-20b: input 0.05 → 0.07, output 0.20 → 0.30, added cache_read = 0.035
- minimax-m2p7: cache_read 0.03 → 0.06
2026-05-22 01:55:31 +05:00
Aiden Cline 26b05268ae Merge pull request #1824 from anomalyco/fix/vercel-gemini-35-flash
Add new Vercel AI Gateway models
2026-05-21 15:30:35 -05:00
Aiden Cline 1aee13d2e5 Add new Vercel AI Gateway models 2026-05-21 13:21:25 -05:00
Aiden Cline 0a924e6bb2 Merge pull request #1822 from anomalyco/automation/sync-models-xai
chore(sync): update xAI model catalog
2026-05-21 13:05:20 -05:00
Aiden Cline 9769b2b11d Merge pull request #1823 from anomalyco/automation/sync-models-openrouter
chore(sync): update OpenRouter model catalog
2026-05-21 13:05:13 -05:00
github-actions[bot] 1b0db099cf chore(sync): update OpenRouter model catalog 2026-05-21 17:56:13 +00:00
github-actions[bot] 16ed78587c chore(sync): update xAI model catalog 2026-05-21 17:56:11 +00:00
Frank 0b88965165 update zen models 2026-05-21 13:41:55 -04:00
Aiden Cline acc704ce39 Merge pull request #1797 from arnavchachra/add-crof-provider
add crof.ai provider with 21 models
2026-05-21 11:43:02 -05:00
Aiden Cline 51ad3b264e Merge pull request #1821 from anomalyco/sync-provider-ci
chore: automate provider sync jobs
2026-05-21 11:23:49 -05:00
Aiden Cline 146b6c7084 Merge pull request #1819 from Inceptron-Software/add_inceptron_provider
Add Inceptron provider
2026-05-21 11:21:27 -05:00
Aiden Cline 0e3cbe3c64 chore: automate provider sync jobs 2026-05-21 11:21:12 -05:00
Aiden Cline 604d4d66a4 Merge pull request #1820 from Suat-B/codex/xpersona-image-input-20260521
Add image input modality to Xpersona model
2026-05-21 11:00:49 -05:00
SuatB f5090028b8 Add image input modality to Xpersona model 2026-05-21 09:46:28 -05:00
Frank 4bad8faf29 update zen models 2026-05-21 09:05:13 -04:00
Oskar Gustafsson 0df2ccf586 Add Inceptron provider 2026-05-21 09:24:43 +02:00
Aiden Cline bafdc00b45 Merge pull request #1812 from anomalyco/openrouter-extends-sync
Sync OpenRouter models with extends
2026-05-20 21:07:21 -05:00
Aiden Cline 49840c013b Merge pull request #1814 from neonn0d/feat/stepfun-ai
feat(stepfun-ai): add international StepFun platform
2026-05-20 20:58:37 -05:00
Aiden Cline eccae0b54e sync openrouter models with extends 2026-05-20 20:32:11 -05:00
Aiden Cline 4cca29405f Merge pull request #1817 from dpuyosa/dev
Venice: Remove Grok 4.1 Fast and add Grok Build 0.1
2026-05-20 20:26:45 -05:00
Aiden Cline e40d9dd338 Merge pull request #1818 from anomalyco/cloudflare-sync-env
chore(sync): isolate cloudflare credentials
2026-05-20 20:26:22 -05:00
Aiden Cline 6a74991397 chore(sync): isolate cloudflare credentials 2026-05-20 20:19:47 -05:00
dpuyosa 035999cb58 [venice] Replace Grok 4.1 Fast with Grok Build 0.1
- Remove deprecated grok-41-fast model entry
- Add grok-build-0-1 with 200K token tiered pricing
- Update context to 256K and output limit to 65,536
2026-05-21 02:36:19 +02:00
Frank cec56bf1bc update zen models 2026-05-20 19:43:25 -04:00
Aiden Cline 85f0cdcb2f Merge pull request #1816 from anomalyco/xai-sync
Add PDF input modality to Grok models
2026-05-20 18:12:47 -05:00
Aiden Cline ef80d4df4e Infer PDF modality for xAI image models 2026-05-20 18:12:11 -05:00
Aiden Cline af0ef00109 Update xAI Grok PDF modalities 2026-05-20 18:06:33 -05:00
Aiden Cline 5fdcea6b36 Merge pull request #1815 from anomalyco/cloudflare-ai-gateway
chore(sync): add cloudflare workers ai sync
2026-05-20 18:02:28 -05:00
Aiden Cline 31e56480b4 chore(sync): add cloudflare workers ai sync 2026-05-20 16:55:29 -05:00
Aiden Cline 92a621594e Merge pull request #1813 from anomalyco/sync-xai
chore(sync): add xai model sync
2026-05-20 16:02:38 -05:00
Aiden Cline d2db353ceb chore: ignore sync reports 2026-05-20 16:01:51 -05:00
Aiden Cline 900ae509d2 Merge pull request #1808 from ajussak/scaleway
Added Mistral Medium 3.5 128B from Scaleway
2026-05-20 15:58:00 -05:00
neo 9d60164243 feat(stepfun-ai): add international StepFun platform
StepFun runs two separate platforms with distinct accounts/keys:
platform.stepfun.com (China, already covered by providers/stepfun) and
platform.stepfun.ai (international). Keys are not interchangeable
across the two — .ai keys are rejected by api.stepfun.com as
invalid_api_key.

Stepfun's own opencode integration guide instructs users to point at
https://api.stepfun.ai/step_plan/v1. This adds providers/stepfun-ai
for that endpoint, symlinking the shared chat models. Follows the
moonshotai / moonshotai-cn pattern.
2026-05-20 20:43:47 +02:00
Adrien Jussak cd3e99025f Update Mistral Medium 3.5 128B model configuration to extend from mistral-medium-2604 and adjust context window size. 2026-05-20 20:43:45 +02:00
Aiden Cline 1098981eb6 chore(sync): add xai model sync 2026-05-20 13:23:02 -05:00
Aiden Cline 27a151cf53 Merge pull request #1811 from anomalyco/xai-grok-build-model
Add xAI Grok Build model
2026-05-20 13:05:06 -05:00
Aiden Cline 41ff42ab7f Merge pull request #1810 from fhennerkes/dev
poe: add Gemini-3.5-Flash model
2026-05-20 13:04:47 -05:00
Aiden Cline adf1cbdecd add xai grok build model 2026-05-20 13:04:13 -05:00
Frank 10ddc78ce0 update zen models 2026-05-20 14:02:01 -04:00
fhennerkes 7ae897e440 poe: add Gemini-3.5-Flash model
Add new Google model from Poe API (released 2026-05-19).
Uses extends format inheriting from google/gemini-3.5-flash with
Poe-specific overrides (name format, no temperature, markup pricing,
limited input modalities).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-20 10:52:15 -07:00
Aiden Cline 02e452c2e8 Merge pull request #1805 from anomalyco/sync-google
sync google models
2026-05-20 10:53:57 -05:00
Aiden Cline f05b63fff5 Merge pull request #1807 from anomalyco/automation/sync-models-aggregators
chore(sync): update aggregator model catalogs
2026-05-20 10:34:30 -05:00
Adrien Jussak e277d60236 Add Mistral Medium 3.5 128B to Scaleway 2026-05-20 16:14:37 +02:00
github-actions[bot] b01277a737 chore(sync): update aggregator model catalogs 2026-05-20 09:25:37 +00:00
Frank a5da5aa429 update zen models 2026-05-20 04:14:28 -04:00
Aiden Cline 5ceac8a58b updates 2026-05-19 23:43:01 -05:00
Aiden Cline 11e1d5623a Merge pull request #1806 from Cahl-Dee/grid-model-updates-2026-05
Grid model updates 2026-05
2026-05-19 23:42:26 -05:00
Aiden Cline f107afc57c sync 2026-05-19 22:56:50 -05:00
Carl DiClementi e789d7c1f2 Merge branch 'anomalyco:dev' into grid-model-updates-2026-05 2026-05-19 16:06:23 -05:00
Cahl-Dee 899668ad49 added new code and agent models and updated existing text models 2026-05-19 16:04:58 -05:00
Aiden Cline 8c677f0134 sync google models 2026-05-19 15:58:06 -05:00
Aiden Cline 462c7877d9 add gemini 3.5 flash 2026-05-19 15:43:38 -05:00
Aiden Cline 55871f9dca Merge pull request #1804 from Adanlink/dev
feat: add deepseek-v4-flash to the fireworks-ai provider
2026-05-19 15:29:33 -05:00
Aiden Cline 4f7194a3c8 test 2026-05-19 15:09:11 -05:00
Aiden Cline f65f0148da add sync guide 2026-05-19 15:08:44 -05:00
Adán 14952f8855 Rename deepseek-v4-flash to deepseek-v4-flash.toml 2026-05-19 18:56:43 +02:00
Adán d7c6d3ad12 Add deepseek-v4-flash model configuration 2026-05-19 18:53:48 +02:00
Aiden Cline 356bc79d08 Merge pull request #1637 from elvexai/fix/amazon-bedrock-kimi-token-limits
fix: Token limits for Amazon Bedrock Kimi K2 models
2026-05-19 09:42:30 -05:00
Aiden Cline a89b1ed726 Merge pull request #1801 from bas3line/sync-routing-run-models
Sync routing.run model catalog
2026-05-19 09:41:38 -05:00
bas3line a998576773 fix(routing-run): match live model metadata 2026-05-19 10:38:58 +05:30
bas3line fbe842bbea fix(routing-run): expose reasoning metadata 2026-05-19 08:39:51 +05:30
bas3line 6c0c3d1b10 fix(routing-run): sync model catalog 2026-05-19 07:37:53 +05:30
Aiden Cline db0a7cf611 Merge pull request #1798 from anomalyco/rework-sync-logic
sync: centralize aggregator model updates
2026-05-18 20:12:50 -05:00
Aiden Cline d775e37e3b Merge pull request #1800 from jerome-benoit/feat/sap-ai-core-gpt-5.4
feat(sap-ai-core): add GPT-5.4
2026-05-18 20:12:21 -05:00
Jérôme Benoit 36753063d9 feat(sap-ai-core): add GPT-5.4
Add gpt-5.4 with availability date from official SAP source.

Drop [[cost.tiers]] from gemini-2.5-pro pending SAP-side tiered
pricing confirmation; sap-ai-core now declares no per-model tiers
(SAP Note 3437766 is login-gated and authoritative for capacity
unit conversion rates).
2026-05-19 02:58:29 +02:00
Aiden Cline 5ee955297a sync: drop vercel catalog updates 2026-05-18 19:07:14 -05:00
Aiden Cline 1b77511903 Merge pull request #1799 from vglafirov/add-gitlab-gpt-5-5
feat(gitlab): add Agentic Chat (GPT-5.5) model
2026-05-18 15:24:18 -05:00
Aiden Cline 8896ead7bf sync: fix vercel pricing tiers 2026-05-18 14:52:40 -05:00
Vladimir Glafirov eb96594d47 feat(gitlab): add Agentic Chat (GPT-5.5) model
Adds duo-chat-gpt-5-5 to the GitLab provider. The GitLab AI Gateway
proxies this model to OpenAI's gpt-5.5-2026-04-23 backend with a
1.05M token context window (922k input + 128k output).

Source: gitlab-org/modelops/applied-ml/code-suggestions/ai-assist
models.yml (gpt_5_5 entry with proxy_provider: openai).

The gitlab-ai-provider npm package exposes this model id starting in
v6.7.0.
2026-05-18 20:44:34 +02:00
Aiden Cline 327332efe3 Merge pull request #1794 from bas3line/add-routing-run-provider
Add routing.run provider
2026-05-18 12:32:34 -05:00
Aiden Cline 5020951745 sync: refresh openrouter after dev merge 2026-05-18 12:30:29 -05:00
Aiden Cline cb6f97774e Merge remote-tracking branch 'origin/dev' into rework-sync-logic 2026-05-18 12:29:28 -05:00
Aiden Cline 7f8b493b0c Merge pull request #1795 from delafthi/delafthi/lxxqxzktnozv
fix(providers/novita-ai): use lowercase model names
2026-05-18 12:28:32 -05:00
Aiden Cline d65a862533 sync: centralize aggregator model updates 2026-05-18 12:12:15 -05:00
arnavchachra 9420048dfe fix crof model limits and reasoning flag to match Crof API 2026-05-18 21:51:19 +05:30
arnavchachra 8db6c27634 add crof provider with 21 models 2026-05-18 21:40:11 +05:30
Victor Navarro 8e710e19ea bring back old bick-pickle
Added interleaved section with reasoning_content field and removed provider section.
2026-05-18 11:44:14 +02:00
Frank 36c6896e97 update zen models 2026-05-17 22:58:06 -04:00
Aiden Cline a8be548a5d Merge pull request #1416 from Luew2/add-lilac-provider
Add Lilac provider
2026-05-17 19:23:07 -05:00
Luew2 8c2fae4ab0 Keep exact Lilac Gemma model name 2026-05-17 17:19:53 -07:00
Luew2 5e7fad350d Align Lilac Gemma display name 2026-05-17 17:19:04 -07:00
Luew2 feb85ef2c9 Align Lilac provider with registry conventions 2026-05-17 17:15:40 -07:00
Luew2 d4161ebf24 Follow models.dev conventions for Lilac provider 2026-05-17 17:11:03 -07:00
Luew2 b91ab02e2b Add Lilac MiniMax M2.7 model 2026-05-17 17:07:44 -07:00
Luew2 dd09d07f75 Update Lilac Kimi model to K2.6 2026-05-17 17:07:44 -07:00
Luew2 4515f85d47 Add Lilac cache pricing 2026-05-17 17:07:44 -07:00
Luew2 97572240e1 Add Gemma 4 31B IT model
Adds google/gemma-4-31b-it to the Lilac provider ($0.11/M input,
$0.35/M output, 262K context, native multimodal with image/video).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-17 17:07:44 -07:00
Luew2 80cc04e9e5 Fix Kimi K2.5 output limit to 262,144 tokens 2026-05-17 17:07:44 -07:00
Luew2 d532ebb89b Use purple gradient for Lilac logo (brand colors #6451dc → #b6a6f9) 2026-05-17 17:07:44 -07:00
Luew2 9bae3e887f Replace placeholder logo with Lilac icon mark (currentColor) 2026-05-17 17:07:44 -07:00
Luew2 0be69bf872 Add Lilac provider
Add Lilac as an OpenAI-compatible provider serving:
- z-ai/glm-5.1: Z.ai's flagship agentic model (754B MoE, 202.8K context)
- moonshotai/kimi-k2.5: Moonshot AI's multimodal reasoning model (1T MoE, 262K context)

API: https://api.getlilac.com/v1
Docs: https://docs.getlilac.com
2026-05-17 17:07:44 -07:00
Thierry Delafontaine 0ed38cbecf fix(providers/novita-ai): use lowercase model names
Mixed case naming causes conflicts on case-insensitive filesystems like macOS.
2026-05-17 21:17:46 +02:00
bas3line e65703382a feat: add routing.run provider 2026-05-17 20:27:56 +05:30
Aiden Cline 748754c99b Merge pull request #1792 from monotykamary/fix/neuralwatt-context-limits
fix(neuralwatt): sync context window and output limits with upstream API
2026-05-16 13:14:05 -05:00
Tom X Nguyen ec9c12d0fc fix(neuralwatt): sync context window and output limits with upstream API
Updates all 14 neuralwatt model TOML files with corrected context window
and max output token values as reported by the Neuralwatt API:

- Devstral-Small-2-24B-Instruct-2512: 262,144 -> 262,128
- GLM-5/GLM-5.1 variants: 200,000 -> 202,736
- GPT-OSS-20B: 16,384 -> 16,368
- Kimi-K2.5/K2.6 variants: 262,144 -> 262,128
- MiniMax-M2.5: 196,608 -> 196,592
- Qwen3.5-397B variants: 262,144 -> 262,128
- Qwen3.6-35B variants: 131,072 -> 131,056

Also fixes the README: moves kimi-k2.6-fast from 'Reasoning Models' to
'Fast Variants' and removes incorrect claim that fast variants support
reasoning.
2026-05-16 23:16:34 +07:00
Aiden Cline ac81822c89 Merge pull request #1777 from berget-ai/feat/berget-kimi-k2.6
feat: add Kimi K2.6 to berget.ai
2026-05-16 06:09:00 -05:00
Aiden Cline d32ed764bd Merge pull request #1791 from anomalyco/automation/sync-openrouter-models
Sync OpenRouter models
2026-05-16 06:08:09 -05:00
Christian Landgren 45fb951c42 feat: add Kimi K2.6 to berget.ai
Add Moonshot AI Kimi K2.6 model to berget.ai provider catalog.

- 262K context window
- 16K output tokens
- Text input/output
- Supports: reasoning, structured output, tool calling
- Pricing: /bin/zsh.83/M input, .85/M output (EUR-based)
- Open weights
2026-05-16 12:46:35 +02:00
github-actions[bot] 260d79b2d5 Sync OpenRouter models 2026-05-16 08:54:41 +00:00
Aiden Cline 746b9caf79 Merge pull request #1785 from jerome-benoit/feat/sap-ai-core-opus-4-7
feat(sap-ai-core): add Claude Opus 4.7 and sync model specs
2026-05-15 23:24:47 -05:00
Aiden Cline 8362b55503 Merge pull request #1786 from Ardakilic/chore/kilo-sync-20260516
providers(kilo): sync upstream
2026-05-15 23:24:34 -05:00
Aiden Cline dde3953a9f Merge pull request #1787 from Suat-B/codex/xpersona-www-api-url
Fix Xpersona API base URL
2026-05-15 23:24:06 -05:00
Aiden Cline 0a1695212c Merge pull request #1788 from Jaaneek/xai-may-15-2026-retirement
xai: drop models retired May 15, 2026 + add Grok Imagine models
2026-05-15 23:23:52 -05:00
Jaaneek 89fbb6bb69 xai: drop models retired May 15, 2026 + add Grok Imagine models 2026-05-16 01:57:15 +01:00
SuatB 005fe0fb5a Fix Xpersona provider API URL 2026-05-15 18:37:41 -05:00
Frank e283875ce7 update zen models 2026-05-15 17:24:53 -04:00
Arda Kilicdagi 3598019251 providers(kilo): sync upstream 2026-05-16 01:11:25 +04:00
Jérôme Benoit e9ad8b0a3f feat(sap-ai-core): add Claude Opus 4.7 and sync model specs 2026-05-15 22:16:07 +02:00
Aiden Cline 0ee78eeda5 sync: openrouter models 2026-05-15 10:20:36 -05:00
Aiden Cline 0a5b33e518 Merge pull request #1778 from zhenjunchen-png/add-orcarouter
feat: add OrcaRouter as a new provider
2026-05-15 10:16:29 -05:00
Aiden Cline 22416dda64 Merge pull request #1783 from anomalyco/sync-openrouter
add sync script for openrouter, sync openrouter models
2026-05-15 10:12:02 -05:00
Aiden Cline eba2702e3a fix: families 2026-05-15 10:03:39 -05:00
Aiden Cline c2c5cc8f21 add sync script for openrouter, sync openrouter models 2026-05-15 10:00:32 -05:00
zhenjun.chen 022b1b9946 feat(orcarouter): expand to 80 chat models and add brand logo
Adds 55 additional upstream-mirrored models alongside the existing 25,
covering the full OrcaRouter chat catalog as exposed by
https://www.orcarouter.ai/api/pricing (text-only chat — TTS, embeddings,
video, and image generation are filtered out).

Per-namespace upstream mappings used by [extends]:

  OrcaRouter ns  -> models.dev provider
  ---------------- + ---------------
  openai         -> openai
  anthropic      -> anthropic   (dot version -> dash, e.g. opus-4.7 -> opus-4-7)
  google         -> google
  deepseek       -> deepseek
  qwen           -> alibaba
  grok           -> xai
  kimi           -> moonshotai
  minimax        -> minimax     (minimax-m2.7 -> MiniMax-M2.7)
  z-ai           -> zai

OrcaRouter-specific aliases (dated snapshots like gpt-5-2025-08-07,
search-preview variants, qwen3-vl-* visual variants) are excluded from
v1 because their upstream canonical files do not yet exist in models.dev.

Also adds providers/orcarouter/logo.svg.
2026-05-15 14:48:31 +08:00
Aiden Cline 8269e04222 Merge pull request #1782 from Suat-B/codex/xpersona-limits-logo
Update Xpersona limits, cutoff, and logo
2026-05-15 00:29:32 -05:00
Aiden Cline a75cf2ed1c Merge pull request #1552 from aredridel/as/add-umans
feat(models): add umans.ai coding plan
2026-05-14 22:40:39 -05:00
Aria Stewart 7b00aafa79 feat(models): umans.ai definitions 2026-05-14 23:39:30 -04:00
Suat-B e54e0fc8b5 Update Xpersona limits, cutoff, and logo 2026-05-15 03:22:16 +00:00
Aiden Cline 38611e75fa Merge pull request #1781 from dpuyosa/feat/add-venice-claude-opus-4-7-fast-model
Venice: Add Claude Opus 4.7 Fast model
2026-05-14 22:08:14 -05:00
dpuyosa e62c1e973e [venice] Add Claude Opus 4.7 Fast model
- New pricing with 36/180 input/output per million tokens
- 1M context window with 128K output limit
- Text+image input, text-only output
2026-05-15 01:32:24 +02:00
Aiden Cline c2b3c601e4 Merge pull request #1724 from isaachuangGMICLOUD/feat/add-gmicloud-provider
providers(gmicloud): add GMI Cloud provider
2026-05-14 17:26:50 -05:00
Aiden Cline 9ff1d36a21 Merge pull request #1476 from Vect0rM/feat/add-atomic-chat-provider
feat: add Atomic Chat provider
2026-05-14 17:18:21 -05:00
Aiden Cline e0f4042ad1 Merge pull request #1776 from kapelame/docs/minimax-token-plan-rename
providers(minimax): rename Coding Plan → Token Plan in display labels
2026-05-14 17:16:28 -05:00
Frank 14736ba4b6 update zen models 2026-05-14 17:15:53 -04:00
Aiden Cline 9351d68731 Merge pull request #1772 from nearai/add-nearai
Add NEAR AI Cloud provider
2026-05-14 10:30:28 -05:00
Aiden Cline 7a5a1d2aff Merge pull request #1779 from NameIsHiki/siliconflow-deepseek-v4
feat(siliconflow): add DeepSeek v4 models
2026-05-14 10:29:52 -05:00
zhenjun.chen 699284ce91 chore(orcarouter): drop oversize logo, fall back to models.dev default
The previously committed logo is ~100KB; existing wrapper-provider logos
(openrouter, llmgateway, kilo, aihubmix, ambient) are all 0.3-6KB and use
`currentColor`. Falling back to the default logo per README:

  > If we don't have a provider's logo, a default logo is served instead.

A properly-sized currentColor logo will follow in a separate PR.
2026-05-14 21:37:48 +08:00
Hiki 21ce5c3ac8 Create deepseek-v4-flash.toml 2026-05-14 15:26:37 +02:00
Hiki 3485cf52d0 Create deepseek-v4-pro.toml 2026-05-14 15:22:53 +02:00
Frank 99e8f25c78 update zen models 2026-05-14 08:59:42 -04:00
zhenjun.chen 7102978cb4 feat: add OrcaRouter provider
OrcaRouter is an OpenAI-compatible meta-router aggregating 150+ LLMs
(OpenAI, Anthropic, Google, xAI, DeepSeek, Qwen, Kimi, MiniMax, ...)
behind a single API key, with a virtual orcarouter/auto smart-routing
entry that picks an upstream per request.

This initial scope covers 26 models (1 AUTO router + 25 upstream mirrors
using [extends]). Pricing computed from https://www.orcarouter.ai/api/pricing
on 2026-05-14: input = model_ratio * $2, output = model_ratio *
completion_ratio * $2 (USD per 1M tokens).

Disclosure: I'm an engineer on the OrcaRouter team.
2026-05-14 20:55:47 +08:00
kapelame 530f60c69c providers(minimax): rename Coding Plan → Token Plan in display labels
The product was renamed from "Coding Plan" to "Token Plan" when its
scope expanded beyond coding to cover all MiniMax modalities (text,
speech, video, music, image). Per
https://platform.minimax.io/docs/token-plan/intro:
"Token Plan extends upon our former Coding Plan."

Updates display name and doc URL for the two affected provider
catalog entries. Provider IDs (minimax-coding-plan,
minimax-cn-coding-plan) are intentionally unchanged for backward
compatibility — anyone with these IDs in opencode.json or
elsewhere keeps working. Old /coding-plan/* URLs still 307-redirect
to the new /token-plan/* paths upstream.

Region disambiguation stays as the URL in parens (matching the
existing minimax / minimax-cn naming convention) — no "China" word
added, since the URL already conveys the region cleanly in the
provider picker.
2026-05-14 19:18:03 +08:00
Misha Skvortsov 2415c5be21 fix(atomic-chat): update logo.svg with new design
Replaces the existing logo.svg file with an updated design for the Atomic Chat provider. This change enhances the visual branding of the application.
2026-05-14 10:58:36 +03:00
Aiden Cline 85aba468cd Merge pull request #1767 from Suat-B/codex/xpersona-provider-20260513
Add Xpersona provider
2026-05-13 23:20:20 -05:00
Aiden Cline d1ec1ba777 Merge pull request #1773 from ambient-gregory/dev
feat: add Ambient provider with GLM-5.1 and Kimi K2.6
2026-05-13 19:20:20 -05:00
Gregory 0f94bf16ec fix(ambient): shrink logo display size to match other providers 2026-05-13 19:33:26 -04:00
Aiden Cline 506e8f48a9 Merge pull request #1770 from EriDeLee/dev
chore(aihubmix): sync model catalog
2026-05-13 17:43:30 -05:00
Aiden Cline 3480bc5992 Merge pull request #1775 from michaelnchin/fix/amazon-bedrock-gpt-oss-tokens
fix: Output tokens for Bedrock GPT-OSS models
2026-05-13 17:36:42 -05:00
Michael Chin fde97814ef fix: Output tokens for Bedrock GPT-OSS models 2026-05-13 14:45:17 -07:00
Gregory ff7eddcb70 feat: add Ambient provider with GLM-5.1 and Kimi K2.6
Adds the Ambient inference provider (api.ambient.xyz) with an initial
catalog of GLM-5.1 and Kimi K2.6, plus a generator script that pulls
from /v1/models so pricing and limits stay in sync with the upstream API.

Run `bun run ambient:generate` to refresh model TOMLs.
2026-05-13 11:30:45 -04:00
Evrard-Nil Daillet 5cbab85b8d Add nearai logo.svg from cloud.near.ai favicon 2026-05-13 16:32:08 +02:00
Evrard-Nil Daillet 6f9820de9f Add NEAR AI Cloud provider
Adds nearai as an OpenAI-compatible provider at https://cloud-api.near.ai/v1
serving 33 models. First-party mirrors (anthropic/openai/google) use `extends`;
NEAR-hosted open-weight models (Qwen, GLM-5.1-FP8, gpt-oss, whisper, FLUX) have
full definitions.

Pricing and context limits sourced from cloud-api.near.ai/v1/models.
2026-05-13 16:32:08 +02:00
EriDeLee cdfb429098 chore(aihubmix): sync model catalog 2026-05-13 21:55:32 +08:00
Victor Navarro 1c2546af8a perf: virtualize models table and other improvements 2026-05-13 12:29:27 +02:00
vimtor 2288a1626b trim search index to essential fields and remove dead code 2026-05-13 12:21:52 +02:00
vimtor b3ecfc3d70 improve row scanning 2026-05-13 11:20:42 +02:00
Suat-B 30b3e677fd Add Xpersona provider 2026-05-13 00:39:09 -05:00
Suat-B 3b37eee86e Add Xpersona provider 2026-05-13 00:39:08 -05:00
Suat-B 71f069670e Add Xpersona provider 2026-05-13 00:39:07 -05:00
Aiden Cline f401672689 Merge pull request #1766 from michaelnchin/fix/amazon-bedrock-structured-output-05122026-2
fix: add structured_output=True for more supported Bedrock models
2026-05-12 23:24:39 -05:00
Michael Chin 2a0d86a034 update structured_output for more Bedrock models 2026-05-12 20:52:36 -07:00
Aiden Cline d9439cdf2f Merge pull request #1762 from zxyaction/feat/add-auriko-provider
feat: add Auriko provider with 15 models
2026-05-12 22:16:02 -05:00
Aiden Cline 3d443d568d Merge pull request #1765 from michaelnchin/fix/amazon-bedrock-structured-output-05122026
fix: update structured_output for Bedrock Claude 4.x models
2026-05-12 22:15:51 -05:00
Michael Chin a76c8fe9dd fix: update structured_output for Bedrock Claude 4.x models 2026-05-12 20:07:12 -07:00
Aiden Cline d08e8d6cc1 Merge pull request #1763 from Tavernari/feat/add-claudinio-provider
feat: add Claudinio provider
2026-05-12 19:13:54 -05:00
Victor Carvalho Tavernari 4c06e44047 fix: use currentColor in logo SVG per contributing guidelines 2026-05-12 23:59:03 +01:00
Aiden Cline 5e344ded49 Merge pull request #1755 from anomalyco/correct-context-tracking
feat: add new context pricing tiers
2026-05-12 17:40:48 -05:00
Aiden Cline 458b7f4d1a use Venice context tier thresholds 2026-05-12 17:39:40 -05:00
Aiden Cline baf4432140 Merge pull request #1759 from NameIsHiki/deepinfra-xiaomi-mimo-models
feat(deepinfra): add Xiaomi MiMo v2.5 and v2.5 Pro
2026-05-12 17:10:59 -05:00
Aiden Cline bbf72ea4e0 Merge pull request #1764 from anomalyco/add-anthropic-opus-4-7-fast-mode
Add fast mode for Anthropic Opus 4.7
2026-05-12 17:10:49 -05:00
Aiden Cline 8f9adc7567 fix generated tier change detection 2026-05-12 17:01:35 -05:00
Aiden Cline addaf1c036 add fast mode for anthropic opus 4.7 2026-05-12 17:01:05 -05:00
Victor Carvalho Tavernari 55d16a58b6 feat: add claudinio provider (OpenAI-compatible, 256K ctx, $0.50/$2.00 per MTok) 2026-05-12 22:33:45 +01:00
Hiki 72a4deab66 Update mimo-v2.5.toml 2026-05-12 23:23:53 +02:00
Hiki 8dd829a187 Update mimo-v2.5-pro.toml 2026-05-12 23:23:01 +02:00
Aiden Cline a671cc05d5 align cost tiers with model schema 2026-05-12 16:01:49 -05:00
Aiden Cline 656c6f08a7 Merge pull request #1761 from Sewer56/deprecate-wafer-models
providers/wafer.ai: Remove DeepSeek-V4-Pro and MiniMax-M2.7 models
2026-05-12 15:51:10 -05:00
Aiden Cline 8979741a32 Merge pull request #1760 from Ardakilic/fix/kilo/kimik26
Fix: Kimi k2.6 definition on Kilo Gateway
2026-05-12 15:50:53 -05:00
Frank 82851b9a3d update zen models 2026-05-12 16:44:41 -04:00
zxy_action ae511892d7 feat: add Auriko provider with 15 models
All models use [extends] to inherit from canonical definitions,
overriding only Auriko-specific pricing. Omits remove cost tiers
and features Auriko doesn't carry.

Models: claude-opus-4-{6,7}, claude-sonnet-4-6, deepseek-v4-{pro,flash},
gemini-{2.5-pro,2.5-flash,3.1-pro-preview}, grok-4.3, kimi-k2.{5,6},
minimax-m2-7{,-highspeed}, glm-5.1, qwen-3.6-plus
2026-05-12 13:34:34 -07:00
Sewer56 9312418242 Changed: Remove DSv4 Pro & MiniMax M2.7 from models.dev 2026-05-12 20:16:08 +01:00
vimtor 9828a0177d improve empty row 2026-05-12 19:49:03 +02:00
Arda Kılıçdağı 122627a852 fix: Kimi k2.6 definition on Kilo Gateway 2026-05-12 20:35:58 +03:00
vimtor 4dee8d0d34 minor improvements 2026-05-12 19:10:21 +02:00
Hiki 6883e793ce Create mimo-v2.5-pro.toml 2026-05-12 18:22:41 +02:00
Hiki e69064709b Update mimo-v2.5.toml 2026-05-12 18:19:52 +02:00
Hiki bd8e582b96 Create mimo-v2.5.toml 2026-05-12 18:03:09 +02:00
Shoubhit Dash 21945db90f Merge pull request #1662 from anomalyco/nxl/add-sarvam-provider
provider(sarvam): add chat models
2026-05-12 13:31:59 +05:30
Aiden Cline 1771e02be8 Merge pull request #1660 from Alex-wuhu/feat/novita-ai-sync-models
provider(novita-ai): sync latest models
2026-05-11 23:15:44 -05:00
Aiden Cline bb08fc26e9 Merge pull request #1706 from rohita5l/rohit/addDatabricks
Add Databricks as a provider
2026-05-11 19:28:34 -05:00
Aiden Cline 2e015de42d preserve generated tier thresholds 2026-05-11 17:07:19 -05:00
Rohit Agrawal 914a3d9d18 fix: restore bun.lock to use default registry instead of Databricks npm proxy
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-11 17:54:26 -04:00
Rohit Agrawal bab01dd9ab refactor: move databricks generate to packages/core/script following repo conventions
Addresses review feedback by removing AI SDK dependencies from package.json
and aligning with the Vercel/Helicone/Wandb pattern.

- Move generate-databricks.ts to packages/core/script/
- Add databricks:generate to root scripts
- Remove smoke test and runtime filtering (catalog should reflect what the
  upstream API exposes; AI SDK compatibility is a downstream concern)
- Add --dry-run and --new-only flags
- Merge with existing TOMLs instead of nuking them; warn about orphans
- Restore databricks-gemini-3-pro and databricks-gemini-3-1-pro
- Drop @ai-sdk/openai-compatible, ai, zod from root dependencies

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-11 17:54:04 -04:00
Rohit Agrawal bdaae956af feat: add AI SDK compatibility test to generate script, remove incompatible models
Generate script now smoke-tests each model with streamText after writing TOMLs
and removes any that return empty responses (incompatible with @ai-sdk/openai-compatible).
Removes databricks-gemini-3-pro and databricks-gemini-3-1-pro which return content
as array with thoughtSignature that the AI SDK cannot parse.

Also adds test-databricks.ts for standalone smoke testing.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-11 17:53:44 -04:00
Rohit Agrawal 6f118145c0 fix: inline gpt-oss model metadata instead of invalid extends path
openrouter models in subdirectories can't use extends (schema requires
provider/model format); resolve() now inlines the source TOML content directly.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-11 17:53:44 -04:00
Rohit Agrawal 386eaed119 add databricks 2026-05-11 17:53:44 -04:00
Aiden Cline 151e9c9071 fix tiered cost generation 2026-05-11 16:43:32 -05:00
Isaac ba99f1edce Add GMI Cloud GLM models 2026-05-11 14:40:03 -07:00
Aiden Cline b96b074a3b fix long-context cost tier omissions 2026-05-11 16:31:07 -05:00
Aiden Cline 2593e131a1 wip 2026-05-11 16:11:19 -05:00
Aiden Cline 4aebbe5ca3 Merge pull request #1754 from BruceMacD/brucemacd/fix-ollama-kimi-k2-6-model-id
fix ollama cloud kimi k2.6 model id
2026-05-11 14:57:03 -05:00
Bruce MacDonald b98495e9c8 fix ollama cloud kimi k2.6 model id 2026-05-11 12:40:17 -07:00
Frank bc95b42ccd update zen models 2026-05-11 11:26:08 -04:00
Aiden Cline 6139fb8c69 Merge pull request #1598 from mugnimaestra/feat/chutes-generate-script
feat(chutes): add API-driven model generator script
2026-05-11 09:30:38 -05:00
Aiden Cline 3070758007 Merge pull request #1678 from 5kahoisaac/chore/nvidia-models
Sync NVIDIA endpoint model catalog
2026-05-11 09:29:54 -05:00
Frank 5525e83de4 update zen models 2026-05-11 10:00:09 -04:00
Frank 359fd879b8 update zen models 2026-05-10 03:54:03 -04:00
Frank 01b5a1a656 update zen models 2026-05-10 02:52:44 -04:00
Frank b1958be099 update zen models 2026-05-10 02:42:44 -04:00
Aiden Cline 08aa068523 Temporarily remove kiro provider and models 2026-05-10 01:19:47 -05:00
Aiden Cline f31ad0b02f Merge pull request #1738 from mattiacerutti/chore/remove-gh-copilot-deprecated
chore(copilot): mark deprecated models
2026-05-09 15:42:19 -05:00
Aiden Cline 585aa7fa1b Merge pull request #1741 from EriDeLee/dev
Update aihubmix models
2026-05-09 15:42:03 -05:00
Aiden Cline c42a327b3e Merge pull request #1745 from mads-digitial-solutions/patch-1
Update Google provider docs url from pricing page to models page
2026-05-09 15:41:51 -05:00
mads-digitial-solutions 92ebbfb5c4 Update provider.toml
Update Google provider docs URL from the pricing page to the models page
2026-05-09 19:49:19 +01:00
Aiden Cline 535fe8c971 Merge pull request #1744 from OpeOginni/fix/bedrock-model-ids
chore(bedrock): Getting rid of legacy Amazon Bedrock model offerings
2026-05-09 13:48:52 -05:00
OpeOginni a3b4bfc16c fix(bedrock): remove uneeded model configurations 2026-05-09 20:35:34 +02:00
OpeOginni d0fcd6f11f fix(bedrock): align models with current docs 2026-05-09 20:29:11 +02:00
OpeOginni e55cd54218 fix(bedrock): remove legacy model entries 2026-05-09 20:21:33 +02:00
OpeOginni 0d73b82b9f fix(bedrock): restore regional model IDs 2026-05-09 20:18:59 +02:00
Aiden Cline 83c7e2b63f Merge pull request #1742 from Adam8234/add-firepass-provider
feat: add Fireworks (Firepass) provider
2026-05-09 12:45:09 -05:00
Adam 83ae4cf813 feat: add Fireworks (Firepass) provider
Adds the Fireworks AI Firepass subscription provider.
- Provider uses a dedicated FIREPASS_API_KEY
- Uses @ai-sdk/openai-compatible SDK
- Includes Kimi K2.6 Turbo (accounts/fireworks/routers/kimi-k2p6-turbo)
- Zero per-token cost since covered by subscription
2026-05-09 12:22:13 -05:00
EriDeLee 77eae6eef7 Update aihubmix models 2026-05-09 20:53:57 +08:00
Aiden Cline 8cbf6ed10e Merge pull request #1736 from vercel/update-vercel-models-1778258030
Update Vercel models
2026-05-08 21:58:40 -05:00
Mattia Cerutti 06d87e4411 chore(copilot): remove deprecated models 2026-05-09 00:19:50 +02:00
Frank 2cb3832618 update zen models 2026-05-08 17:11:05 -04:00
vimtor 950a0446d4 bring back svg loading 2026-05-08 20:18:12 +02:00
vimtor 41fbfb1a17 change sst mention 2026-05-08 20:13:09 +02:00
vimtor cf5045a90b move copy button next to name 2026-05-08 20:12:39 +02:00
vimtor ef739220de lock virtualized table column widths to prevent scroll jitter 2026-05-08 19:07:25 +02:00
github-actions[bot] df960d1a90 chore(vercel): update Vercel model definitions
Auto-generated by weekly workflow from Vercel AI Gateway API.

Co-Authored-By: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-08 16:33:52 +00:00
vimtor 7294efee15 move row-render.ts to shared.ts 2026-05-08 18:30:14 +02:00
vimtor 1e537147eb remove table row tuple optimization 2026-05-08 18:24:18 +02:00
Aiden Cline 8f2f83ef61 Merge pull request #1735 from slpdy/dev
Create DeepSeek-V4-Pro.toml
2026-05-08 10:59:33 -05:00
Aiden Cline b133426465 Merge pull request #1734 from oskarkocol/chore/update-novita-deepseek-prices
chore: update novita deepseek-v4-pro prices
2026-05-08 10:59:25 -05:00
Aiden Cline dafff5a770 Merge pull request #1730 from smakosh/add-llmgateway-models
Add new LLM Gateway text models (gpt-5.5, grok-4-3, gemini-3.1-flash-lite, qwen3.6, MiMo v2)
2026-05-08 10:58:20 -05:00
smakosh dd894f077f Merge remote-tracking branch 'upstream/dev' into add-llmgateway-models
# Conflicts:
#	providers/google/models/gemini-3.1-flash-lite.toml
2026-05-08 17:46:52 +02:00
smakosh 91590874e7 Revert "fix(models): use canonical entries for qwen3.6-max-preview and grok-4.3"
This reverts commit 70ac6fccda.
2026-05-08 17:43:07 +02:00
vimtor e9f4cecc54 extract shared row rendering module 2026-05-08 17:29:04 +02:00
vimtor fb1ac09883 virtualize model table 2026-05-08 17:07:28 +02:00
Jj a436236146 Create DeepSeek-V4-Pro.toml
Added DeepSeek-v4-Pro model to Nebius provider
2026-05-08 11:03:16 +01:00
oskar 1415b4be97 chore: update novita deepseek prices 2026-05-08 14:21:27 +07:00
Aiden Cline 1437da86e7 Merge pull request #1685 from 8023/dev
Add kimi-k2.6/deepseek-v4 model and EmpirioLabs AI integration for poe.com
2026-05-07 22:38:51 -05:00
Aiden Cline e7d57885d1 Merge pull request #1733 from chl-0537/feature/add-tencent
add model by openrouter
2026-05-07 22:21:07 -05:00
Aiden Cline 8157916515 fix: attachment 2026-05-07 22:12:11 -05:00
Aiden Cline 2e5b87a9c2 Merge pull request #1731 from mikeyp/chore/update-digitalocean-models
Add script to generate/update Digitalocean models
2026-05-07 16:42:09 -05:00
Aiden Cline 92e19432d3 add google gemini 3.1 flash lite 2026-05-07 15:44:25 -05:00
smakosh 70ac6fccda fix(models): use canonical entries for qwen3.6-max-preview and grok-4.3
Apply the canonical TOML provided by the LLM Gateway team for the
Qwen3.6 Max Preview and Grok 4.3 parent definitions, replacing the
upstream-merged variants whose dates and pricing did not match.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-07 22:34:01 +02:00
smakosh dc3283417d Merge remote-tracking branch 'upstream/dev' into add-llmgateway-models
# Conflicts:
#	providers/alibaba/models/qwen3.6-max-preview.toml
#	providers/llmgateway/models/qwen3.6-max-preview.toml
2026-05-07 22:28:23 +02:00
smakosh 34fd6673e5 chore(llmgateway): add new text models from llmgateway catalog
Add gemini-3.1-flash-lite, grok-4-3, gpt-5.5, gpt-5.5-pro, qwen3.6
and MiMo v2 models that exist in llmgateway.io but were missing
from models.dev. Adds parent definitions for grok-4-3,
gemini-3.1-flash-lite, and qwen3.6-max-preview where they did not
already exist.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-07 22:20:52 +02:00
Mike Prasuhn 5df9293314 Add script to generate/update Digitalocean models 2026-05-07 15:28:07 -04:00
Aiden Cline a81a9559d7 Merge pull request #1726 from Ardakilic/chore/sync-kilo-models
Chore: Sync Kilo models with upstream gateway
2026-05-07 13:04:51 -05:00
Aiden Cline 4197bf57a6 Merge pull request #1725 from dpuyosa/chore/venice-grok-costs
Venice: Update grok-4-20 pricing
2026-05-07 13:04:26 -05:00
Aiden Cline d0ac772507 Merge pull request #1717 from juls0730/refactor/mimo-extends/token-plan
refactor(xiaomi-token-plan): extends xiaomi base provider for xiaomi token plans
2026-05-07 13:04:05 -05:00
Aiden Cline 1afaa053e7 Merge pull request #1727 from sergeykonkin/update-nebius-models-2026-05
chore: update nebius provider models
2026-05-07 13:03:46 -05:00
Aiden Cline 9b77ce1c9e Merge pull request #1729 from Sewer56/add-minimax-m27-wafer
Added: MiniMax-M2.7 model for wafer.ai
2026-05-07 12:59:49 -05:00
Aiden Cline 68dc6d1820 Merge pull request #1728 from arafatkatze/codex/openrouter-qwen-cache-pricing
Add OpenRouter Qwen cache pricing
2026-05-07 12:59:29 -05:00
Arafatkatze a1eb5eece3 Add OpenRouter Qwen cache pricing 2026-05-07 10:29:48 -07:00
Sewer56 f7ec2c517f Added: wafer.ai/MiniMax-M2.7 model 2026-05-07 18:12:50 +01:00
Sergey Konkin f98e8ec793 chore: update nebius provider models 2026-05-07 15:37:18 +02:00
Arda Kilicdagi d210066793 chore: sync kilo gw models 2026-05-07 14:17:15 +03:00
Frank 06908cbf36 update zen models 2026-05-07 04:47:22 -04:00
dpuyosa b0614d2088 [venice] Update grok-4-20 pricing
- Lower grok-4-20 and multi-agent input/output costs to latest Venice pricing
2026-05-07 10:29:48 +02:00
mickalchen 32c1c45c52 add openrouter model 2026-05-07 11:22:11 +08:00
Frank 7c033f27e6 update zen models 2026-05-06 23:00:15 -04:00
mickalchen d648e63499 Merge remote-tracking branch 'origin/dev' into feature/add-tencent 2026-05-07 10:54:56 +08:00
Zoe 1353f965b9 refactor(xiaomi-token-plan): extends xiaomi base provider for xiaomi token plan 2026-05-06 18:53:49 -05:00
Isaac Huang 175082d43f Add GMI Cloud provider 2026-05-06 15:12:07 -07:00
Aiden Cline bba0a9c3f4 Merge pull request #1312 from NachoFLizaur/feat/kiro-provider
feat: add Kiro provider with 12 models
2026-05-06 12:16:47 -05:00
Frank 4bdb195178 update zen models 2026-05-06 12:56:58 -04:00
Aiden Cline 12706b7652 Merge pull request #1716 from juls0730/refactor/mimo-extends/zenmux
refactor(zenmux): extends mimo models from xiaomi provider
2026-05-06 10:45:01 -05:00
Aiden Cline d8c76d0c67 Merge pull request #1715 from juls0730/refactor/mimo-extends/qiniu-ai
refactor(qiniu-ai): extends mimo models from xiaomi provider
2026-05-06 10:30:46 -05:00
Aiden Cline 3963dd13d5 Merge pull request #1714 from juls0730/refactor/mimo-extends/openrouter
refactor(openrouter): extends mimo models from xiaomi provider
2026-05-06 10:30:32 -05:00
Aiden Cline 8749a56efa Merge pull request #1711 from juls0730/refactor/mimo-extends/kilo
refactor(kilo/xiaomi): extends mimo models from xiaomi provider
2026-05-06 10:29:31 -05:00
Aiden Cline 25de2ee27d Merge pull request #1718 from juls0730/refactor/mimo-extends/xiaomi
refactor(xiaomi): round xiaomi models to powers of 2 & fix small errors
2026-05-06 10:27:36 -05:00
Aiden Cline 438df7f03c Merge pull request #1723 from CloudFerro/fix/cloudferro-sherlock/minimax-m2.5
fix: cloudferro sherlock - update context values for minimax m2.5
2026-05-06 10:25:54 -05:00
Aiden Cline 7d18558aa3 Merge pull request #1722 from dpuyosa/chore/venice-model-updates
Venice: Remove deprecated models, enable reasoning on gpt-oss-120b
2026-05-06 10:25:36 -05:00
Jan Szypulski c82f736fcc fix: cloudferro sherlock - update context values for minimax m2.5 2026-05-06 10:52:39 +02:00
dpuyosa 6a436805b3 [venice] Remove deprecated models, enable reasoning on gpt-oss-120b
- Remove kimi-k2-thinking, qwen3-coder-480b-a35b-instruct, and venice-uncensored models
- Enable reasoning capability on openai-gpt-oss-120b
2026-05-06 09:44:24 +02:00
Jack ce7823f073 Merge pull request #1720 from anomalyco/fix/opencode-go-kimi-k26-pricing
fix(opencode-go): restore kimi k2.6 pricing
2026-05-06 12:33:24 +08:00
Jack 033efdb7d4 fix(opencode-go): restore kimi k2.6 pricing 2026-05-06 12:32:06 +08:00
Alex-wuhu 0f1855c0a7 provider(novita-ai): use extends for kimi k2.6 2026-05-06 10:41:46 +08:00
Zoe bd9e0c2677 refactor(xiaomi): round xiaomi models to powers of 2 & fix small errors 2026-05-05 18:40:05 -05:00
Zoe 38545d63a5 refactor(zenmux): extends mimo models from xiaomi provider 2026-05-05 18:17:21 -05:00
Zoe 014be328d6 refactor(qiniu-ai): extends mimo models from xiaomi provider 2026-05-05 18:05:39 -05:00
Zoe c139147540 refactor(openrouter): extends mimo models from xiaomi provider 2026-05-05 17:49:39 -05:00
Zoe 59cd93cafc refactor(kilo/xiaomi): extends mimo models from xiaomi provider 2026-05-05 17:13:06 -05:00
Aiden Cline e91db96d83 Merge pull request #1710 from Spherrrical/add-digitalocean-kimi-2-6
feat(digitalocean): add kimi-k2.6 model
2026-05-05 14:07:50 -05:00
Spherrrical d70dd8dcdc feat(digitalocean): add kimi-k2.6 model 2026-05-05 12:05:16 -07:00
Frank b18e681457 update deepseek flash on deepinfra 2026-05-05 14:24:03 -04:00
Aiden Cline b73a6a2130 Merge pull request #1709 from xiaomochn/fix/xiaomi-mimo-v2.5-modalities
fix(xiaomi): swap modalities for MiMo-V2.5 and MiMo-V2.5-Pro
2026-05-05 11:36:43 -05:00
Aiden Cline 153c1cc420 Merge pull request #1614 from Yashwanth-Kumar-26/patch-1
Add Qwen 3.6 27B model configuration
2026-05-05 11:22:43 -05:00
Aiden Cline ca0b30569e update google vertex to include all anthropic models 2026-05-05 11:12:23 -05:00
xiaomochn 8345bfbd06 fix(xiaomi): correct modalities for MiMo V2.5 models across providers
Issues fixed:
1. MiMo-V2.5 and MiMo-V2.5-Pro had their modalities swapped in the
   xiaomi canonical source (affects OpenRouter/ZenMux via extends)
2. Removed 'pdf' from V2.5 models — not a supported input modality
3. Fixed vercel provider: V2.5-Pro incorrectly marked as multimodal
4. Fixed opencode-go and vercel V2.5: added missing 'video', removed pdf

Correct modalities:
- MiMo-V2.5: input = ["text", "image", "audio", "video"]
- MiMo-V2.5-Pro: input = ["text"]

Affected providers: xiaomi, opencode-go, vercel, openrouter (extends),
zenmux (extends)

Fixes #1708
2026-05-05 22:46:44 +08:00
Shoubhit Dash 28c0d9ce23 fix(sarvam): correct output limits 2026-05-05 15:30:47 +05:30
Aiden Cline c16f3da694 Merge pull request #1707 from deathbeam/revert-1664-fix/glm-qwen-ollama-output-limit
Revert "fix(ollama): set glm-5.1 and qwen3.5:397b output limits to match context"
2026-05-04 23:51:53 -05:00
Tomas Slusny 6aa6e55d60 fix(ollam): use correct output limit for qwen3.5:397b
{"error":"max_tokens (262144) exceeds model's maximum output tokens (65536) for model qwen3.5:397b (ref: 8554a681-e6a8-45d7-9fdd-433785eb6c67)"}

Signed-off-by: Tomas Slusny <slusnucky@gmail.com>
2026-05-05 01:31:20 +02:00
Tomas Slusny 9efaf754a5 Revert "fix(ollama): set glm-5.1 and qwen3.5:397b output limits to match context" 2026-05-05 01:04:11 +02:00
Aiden Cline 104e4bdc1f Merge pull request #1684 from TheBaconWizard/add-clarifai-kimi-k2.6
provider(clarifai): add Kimi-K2.6 (moonshotai/chat-completion)
2026-05-04 12:00:33 -05:00
Aiden Cline db16c113f7 Merge pull request #1704 from stylings/stylings/add-openrouter-grok-4-3
feat: add OpenRouter Grok 4.3
2026-05-04 12:00:08 -05:00
Alex 64a122a13d fix: update OpenRouter Grok 4.3 file 2026-05-04 12:44:07 -04:00
Alex fcc6521d1d fix: simplify OpenRouter Grok 4.3 file 2026-05-04 12:41:56 -04:00
Alex c623b4a55f feat: add OpenRouter Grok 4.3 2026-05-04 12:36:33 -04:00
Aiden Cline 1d730fea16 Merge pull request #1696 from cgilly2fast/dev
chore(frogbot): convert firmware provider to frogbot
2026-05-04 10:26:36 -05:00
Aiden Cline 5457215e29 Merge pull request #1650 from PedroACosta/feat/add-dinference-models
feat(dinference): add GLM-5.1 and MiniMax-M2.5 models
2026-05-04 10:26:02 -05:00
Aiden Cline b3ab45990f Merge pull request #1697 from rocuevas9511/feat/add-deepinfra-gemma4
add gemma4 26b a4b and 31b to deepinfra
2026-05-04 10:25:31 -05:00
Aiden Cline d906a07e31 Merge pull request #1703 from dpuyosa/feat/venice-grok-4-3
Venice: Add Grok 4.3 model configuration
2026-05-04 10:25:16 -05:00
dpuyosa 4b1f6edd52 [venice] Add Grok 4.3 model configuration
- Add Grok 4.3 model with 1M context and 32K output
- Configure standard and >200K cost tiers
- Enable text+image input with text output modalities
2026-05-04 09:54:01 +02:00
Aiden Cline a92a2cfe3d Merge pull request #1702 from langyo/fix/glm-5v-turbo-naming
fix: use proper GLM family casing for GLM-5V-Turbo
2026-05-03 16:53:44 -05:00
Aiden Cline 70891f58e5 Merge pull request #1701 from kaeltrn/add-perplexity-agent-opus-4-7-gpt-5-5
Add Claude Opus 4.7 and GPT-5.5 models for perplexity-agent
2026-05-03 16:53:26 -05:00
Aiden Cline 1600c827fc Merge pull request #1698 from tim-mcdonald/add-kimi-k2.6-nvidia
Add Kimi K2.6 model for NVIDIA provider
2026-05-03 16:53:01 -05:00
Aiden Cline 5851cdc136 Merge pull request #1700 from JDinABox/dev
Add Synthetic Kimi-K2.6 model configuration
2026-05-03 16:52:46 -05:00
langyo 4fd0e38c58 fix: use proper GLM family casing for GLM-5V-Turbo
- Rename glm-5v-turbo to GLM-5V-Turbo in zai, zhipuai, and 302ai providers
- Add GLM-5V-Turbo back to zhipuai-coding-plan (removed in #1589)

Ref: #1589
2026-05-04 00:51:54 +08:00
Pedro dd685ea42c refactor(dinference): use extends format for GLM and MiniMax models 2026-05-03 17:54:18 +02:00
kaeltrn 7835298241 Add Claude Opus 4.7 and GPT-5.5 models for perplexity-agent 2026-05-03 20:09:14 +07:00
Isaac Ng c5fbcc2c9b 📦 CHORE: remove senera 2026-05-03 16:17:37 +08:00
Isaac Ng ab2eb51b4e chore(nvidia): align Nemotron endpoint slugs
Replace stale NVIDIA Nemotron entries with the live Build catalog slugs so the local provider catalog matches current free and partner endpoints.
2026-05-03 15:47:19 +08:00
Isaac Ng 8e19ec580c 📦 CHORE: sync latest nvidia model 2026-05-03 15:19:17 +08:00
Isaac Ng 3aecc94c46 chore(nvidia): sync endpoint model catalog
Update NVIDIA model TOMLs to match the live Build endpoint list by removing stale entries and adding missing ones.

This keeps the provider catalog aligned with the current free and partner endpoint inventory.
2026-05-03 14:47:30 +08:00
JD Crawford 2ac7912ee7 use extends format 2026-05-03 01:04:53 -04:00
JD Crawford 9ac5b5f625 Add Synthetic Kimi-K2.6 model configuration 2026-05-03 00:17:37 -04:00
Aiden Cline 8c4d9f4696 Merge pull request #1699 from cfal/qwen3.6-max-preview
providers/alibaba/models/qwen3.6-max-preview.toml: add qwen 3.6 max
2026-05-02 21:35:02 -05:00
cfal c4bb0b4b11 providers/alibaba/models/qwen3.6-max-preview.toml: add qwen 3.6 max 2026-05-03 09:25:31 +08:00
Tim McDonald db4d03c171 Add Kimi K2.6 model for NVIDIA provider 2026-05-02 16:09:37 -06:00
rocuevas9511 60d1d4df77 add gemma4 26b a4b and 31b to deepinfra 2026-05-02 14:26:19 -06:00
Colby Gilbert 31654fc2ef chore(frogbot): convert firmware provider to frogbot 2026-05-02 12:36:20 -07:00
Aiden Cline 01e56b1e1e Merge pull request #1682 from varunrandery/poolside/laguna-openrouter
Add Poolside Laguna series models (OpenRouter)
2026-05-02 14:30:20 -05:00
Aiden Cline 4a7d9275d6 Merge pull request #1695 from hgraca/cortecs
Cortecs
2026-05-02 14:04:00 -05:00
Aiden Cline 978fe11e0c Merge pull request #1694 from cgilly2fast/dev
feat(firmware): add deepseek v4, remove gemini 3 pro add gpt 5.4 min,…
2026-05-02 14:03:14 -05:00
Herberto Graca 0b4198955f refactor(cortecs): use extends for models with canonical bases
Convert 7 Cortecs models to extends format, inheriting from their
canonical provider definitions (deepseek, llama, mistral, alibaba).
Reduces duplication by ~55 lines while preserving Cortecs-specific
cost overrides.
2026-05-02 20:56:11 +02:00
Herberto Graca 66b568b681 Add deepseek-v4-pro model for cortecs 2026-05-02 20:56:11 +02:00
Herberto Graca 589fddeed7 Add codestral-2508 model for cortecs 2026-05-02 20:56:10 +02:00
Herberto Graca 6f75266d42 Add deepseek-v3.2 model for cortecs 2026-05-02 20:56:10 +02:00
Herberto Graca 0f03115ee0 Add deepseek-r1-0528 model for cortecs 2026-05-02 20:56:09 +02:00
Herberto Graca ae2564edc2 Add mixtral-8x7B-instruct-v0.1 model for cortecs 2026-05-02 20:56:09 +02:00
Herberto Graca 629856cf41 Add hermes-4-70b model for cortecs 2026-05-02 20:56:08 +02:00
Herberto Graca 567bc13ab4 Add llama-3.3-70b-instruct model for cortecs 2026-05-02 20:56:08 +02:00
Herberto Graca 5eac173bc5 Add qwen3-235b-a22b-instruct-2507 model for cortecs 2026-05-02 20:56:07 +02:00
Herberto Graca cb0d640afd Add qwen3-coder-30b-a3b-instruct model for cortecs 2026-05-02 20:56:07 +02:00
Herberto Graca b5e27e652f Add nemotron-3-super-120b-a12b model for cortecs 2026-05-02 20:56:06 +02:00
Herberto Graca 650c27411f Add qwen3.5-122b-a10b model for cortecs 2026-05-02 20:56:06 +02:00
Herberto Graca 98c55325bf Add mistral-large-2512 model for cortecs 2026-05-02 20:56:05 +02:00
Herberto Graca 2fdd33b8a7 Add qwen3.5-397b-a17b model for cortecs 2026-05-02 20:55:57 +02:00
Colby Gilbert e4dea0a3d8 feat(firmware): grok 4.3 2026-05-02 10:53:24 -07:00
Colby Gilbert b4b3622bfd feat(firmware): add deepseek v4, remove gemini 3 pro add gpt 5.4 min, add gpt 5.4 nano, add gpt 5.5, and minimax m2.7 2026-05-02 09:23:42 -07:00
Aiden Cline 3b35b5598a Merge pull request #1693 from hgraca/add-deepseek-v4-flash-cortecs
Add deepseek-v4-flash model for cortecs
2026-05-02 10:46:57 -05:00
Aiden Cline d2e16bab34 Merge pull request #1610 from monotykamary/feat/neuralwatt-provider
feat: add neuralwatt provider with 14 models
2026-05-02 10:45:50 -05:00
Tom X Nguyen 051dc6236a fix(neuralwatt): sync model capabilities with provider API 2026-05-02 20:42:00 +07:00
Herberto Graca 5b9dee1f3d Add deepseek-v4-flash model for cortecs 2026-05-02 12:50:59 +02:00
8023 426dfe7be5 fix validate error 2026-05-02 12:15:50 +08:00
8023 fddfb083d1 add EmpirioLabs AI and deepseek v4 2026-05-02 11:14:51 +08:00
8023 4eccfaba87 Add Kimi-K2.6 model configuration file 2026-05-02 11:08:32 +08:00
Jeff Lim ea46d016d3 fix(clarifai): drop redundant cost override for Kimi-K2.6
Clarifai's published pricing ($0.95 input, $4.00 output) matches the
canonical moonshotai/kimi-k2.6, so the explicit [cost] block was just
duplicating upstream. Inherit it via extends instead.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-01 18:20:39 -07:00
Jeff Lim 1bd1449ec5 provider(clarifai): add Kimi-K2.6 (moonshotai/chat-completion)
Extends moonshotai/kimi-k2.6 with Clarifai-specific cost and modalities
(text+image only on Clarifai; cache pricing not exposed).

Model URL: https://clarifai.com/moonshotai/chat-completion/models/Kimi-K2_6

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-01 18:14:19 -07:00
Varun Randery 7d334a51c9 Add Laguna models 2026-05-01 23:35:13 +01:00
Aiden Cline 10ae0b5a1d Merge pull request #1664 from fernandoenzo/fix/glm-qwen-ollama-output-limit
fix(ollama): set glm-5.1 and qwen3.5:397b output limits to match context
2026-05-01 15:50:22 -05:00
Aiden Cline d25cce170d Merge pull request #1665 from Ardakilic/chore/cleanup-kilo-provider
Chore: Sync Kilo Gateway provider models with upstream API changes and add Owl Alpha model.
2026-05-01 15:49:53 -05:00
Aiden Cline 3e097a5d89 Merge pull request #1671 from hgraca/add-qwen-2.5-72b-instruct-cortecs
Add qwen-2.5-72b-instruct model for cortecs
2026-05-01 15:49:28 -05:00
Aiden Cline 074b38eb98 Merge pull request #1663 from fernandoenzo/fix/minimax-m2.7-ollama-context-output-limit
fix(ollama): set minimax-m2.7 context and output limits to match Ollama API
2026-05-01 14:24:35 -05:00
Aiden Cline d474922588 Merge pull request #1679 from Spherrrical/add-digitalocean-deepseek-v4-pro
feat(digitalocean): add deepseek-v4-pro model
2026-05-01 13:46:50 -05:00
Spherrrical 56ddb9017c feat(digitalocean): add deepseek-v4-pro model 2026-05-01 11:20:09 -07:00
Rohan Taneja a565bbebc0 Merge pull request #1677 from vercel/update-vercel-models-1777652649 2026-05-01 10:47:39 -07:00
github-actions[bot] 5b97fca592 chore(vercel): update Vercel model definitions
Auto-generated by weekly workflow from Vercel AI Gateway API.

Co-Authored-By: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-05-01 16:24:16 +00:00
Arda Kilicdagi aec3f94081 feat: owl alpha, chore: sync kilo code upstream
feat: owl alpha, chore: sync kilo code upstream
2026-05-01 14:09:51 +03:00
Fernando Guarini e5c63d671e fix(ollama): set glm-5.1 and qwen3.5:397b output limits to match context 2026-05-01 11:29:49 +02:00
Fernando Guarini 22e87ab31b fix(ollama): set minimax-m2.7 context and output limits to match Ollama API 2026-05-01 11:29:41 +02:00
Shoubhit Dash 90515f1913 provider(sarvam): add chat models 2026-05-01 14:41:47 +05:30
Herberto Graca 97d93676bc Add qwen-2.5-72b-instruct model for cortecs 2026-05-01 09:36:16 +02:00
Alex-wuhu 6c3c4a721c provider(novita-ai): sync latest models 2026-05-01 14:19:17 +08:00
Aiden Cline 692fbd0f19 Merge pull request #1628 from fanweixiao/dev
provider(vivgrid): remove GLM-5, add GPT-5.5 and DeepSeek-v4-Pro model
2026-04-30 23:57:52 -05:00
C.C. Fan c3bd263a67 provider(vivgrid): add deepseek-v4-pro model
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-01 11:46:39 +08:00
Aiden Cline 9b9685c297 Merge pull request #1658 from v1gnesh/dev
Create grok-4.3.toml
2026-04-30 22:46:08 -05:00
C.C. Fan b3f063da79 provider(vivgrid): use extends for gpt-5.5 instead of duplicating fields
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-01 11:45:13 +08:00
C.C. a33d0aea81 Merge branch 'anomalyco:dev' into dev 2026-05-01 11:38:42 +08:00
v1gnesh ca380a136e Create grok-4.3.toml 2026-05-01 08:17:12 +05:30
Arda Kılıçdağı dfa186a768 feat: owl alpha 2026-05-01 01:47:25 +03:00
Aiden Cline d63ffa53e8 Merge pull request #1654 from zainhas/dev
[Together AI] add qwen 3.6 plus
2026-04-30 16:34:28 -05:00
Aiden Cline 284def86ef Merge pull request #1653 from smakosh/feat/llmgateway-add-gpt-5-5-and-qwen3-6
feat(llmgateway): add gpt-5.5, gpt-5.5-pro, qwen3.6-35b-a3b, qwen3.6-plus, qwen3.6-max-preview
2026-04-30 16:34:18 -05:00
Siddharth Dhulipalla c19936565d Remove Fire Pass from Fireworks Kimi K2.5 Turbo description (#1655) 2026-04-30 17:28:36 -04:00
Zain Hasan 73b872e141 output 500_000 2026-04-30 14:22:41 -07:00
Zain Hasan 7373bb5878 [Together AI] add qwen 3.6 plus 2026-04-30 14:21:40 -07:00
smakosh d11c151c25 feat(llmgateway): add gpt-5.5, gpt-5.5-pro, qwen3.6-35b-a3b, qwen3.6-plus, qwen3.6-max-preview
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 22:05:40 +02:00
Aiden Cline f3f4fea66c Merge pull request #1645 from stylings/stylings/add-mistral-medium-3-5
feat: add Mistral Medium 3.5
2026-04-30 14:37:16 -05:00
Aiden Cline 90116256a6 Merge pull request #1652 from Spherrrical/add-digitalocean-provider
feat(digitalocean): sync model catalog
2026-04-30 14:32:07 -05:00
Spherrrical 5008df8bcc feat(digitalocean): sync model catalog with /v1/models API
Add 16 new models (anthropic, openai, alibaba, deepseek, google,
meta, mistral, nvidia, baai, intfloat, fal-hosted) to match the
current DigitalOcean Gradient AI Platform catalog, and rename
openai-gpt-5-2-pro to openai-gpt-5.2-pro to match the API id.
2026-04-30 12:08:57 -07:00
Alex 696aa80dec revert(openrouter): remove Mistral Medium 3.5 stub 2026-04-30 14:59:56 -04:00
Aiden Cline 7d00d863dc Merge pull request #1647 from dpuyosa/chore/venice-kimi-pricing
Venice: Update Kimi K2.5 and K2.6 pricing and dates
2026-04-30 11:29:08 -05:00
Aiden Cline ad9eb83b8c Merge pull request #1648 from Snat3r/patch-1
Fix casing in model name MiniMax m2.7 foor cortects provider
2026-04-30 11:28:55 -05:00
Aiden Cline 467da9a82c Merge pull request #1649 from berget-ai/feat/berget-mistral-medium-3.5
feat: add Mistral Medium 3.5 128B to berget.ai
2026-04-30 11:28:43 -05:00
Pedro 22786bcf4b feat(dinference): add GLM-5.1 and MiniMax-M2.5 models 2026-04-30 13:43:44 +02:00
Christian Landgren c8d258b7cb feat: add Mistral Medium 3.5 128B to berget.ai 2026-04-30 12:31:28 +02:00
Snat3r 70309829ca Fix casing in model name and update output limit 2026-04-30 11:59:44 +02:00
dpuyosa 71e00f193b [venice] Update Kimi K2.5 and K2.6 pricing and dates
- Bump kimi-k2-5 cache_read from 0.11 to 0.22
- Bump kimi-k2-6 input from 0.7448 to 0.85 and cache_read from 0.1463 to 0.22
- Update last_updated dates to 2026-04-30
2026-04-30 11:18:54 +02:00
Alex d44e724170 feat(openrouter): add Mistral Medium 3.5 2026-04-29 18:58:26 -04:00
Alex 414695db9f feat(mistral): add Mistral Medium 3.5 2026-04-29 18:44:28 -04:00
Aiden Cline e8b5a27723 Merge pull request #1642 from Sewer56/add-wafer-deepseek-v4-pro
feat(wafer.ai): add DeepSeek V4 Pro
2026-04-29 17:14:17 -05:00
Aiden Cline 6dc9d805db Merge pull request #1643 from Sawyerb/patch-1
Delete providers/vercel/models/inception/mercury-coder-small.toml
2026-04-29 17:14:00 -05:00
Aiden Cline bdf56c9111 add kimi k2.6 to azure cognitive services 2026-04-29 17:13:19 -05:00
Sewer56 293cc3e82f feat(wafer.ai): add DeepSeek V4 Pro 2026-04-29 23:12:04 +01:00
Sawyer Birnbaum 7f5bd4f231 Delete providers/vercel/models/inception/mercury-coder-small.toml
Mercury Coder Small has been deprecated. People should use Mercury Edit 2 instead.
2026-04-29 14:50:35 -07:00
Aiden Cline c4826babc5 Merge pull request #1497 from Lydanne/fix/302ai-models
Update 302ai model metadata and add GPT-5.4 configs
2026-04-29 14:26:01 -05:00
Rohan Taneja 91e8bb985e Merge pull request #1641 from vercel/update-vercel-models-1777480545
Update Vercel models
2026-04-29 11:29:45 -07:00
Aiden Cline 56723051d4 Merge pull request #1615 from deaquino/dev
Add Qwen3.5-9B model configuration file to OVHCloud
2026-04-29 13:09:19 -05:00
Aiden Cline 0bb3e55c08 Merge pull request #1640 from dpuyosa/chore/venice-update-pricing
Venice: Update DeepSeek v4 and Qwen 3.6 model pricing and metadata
2026-04-29 13:02:40 -05:00
github-actions[bot] 080b7a328b chore(vercel): update Vercel model definitions
Auto-generated by weekly workflow from Vercel AI Gateway API.

Co-Authored-By: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-29 16:35:47 +00:00
Misha Skvortsov c59c4aae6c feat(atomic-chat): add curated initial model list
Re-introduces a small curated list of models that ship preconfigured
in Atomic Chat, so opencode users get a working `models.dev` entry
out of the box instead of an empty `models: {}`.

Models (ids match the normalized form returned by Atomic Chat's
/v1/models endpoint, i.e. dots replaced with underscores):

- gemma-4-E4B-it-IQ4_XS
- gemma-4-E4B-it-MLX-4bit
- Qwen3_5-9B-Q4_K_M
- Qwen3_5-9B-MLX-4bit
- Meta-Llama-3_1-8B-Instruct-GGUF

Qwen 3.5 9B dates are taken from the verified providers/venice entry
for the same base model; quantization does not change release dates.

Made-with: Cursor
2026-04-29 17:52:56 +03:00
Frank c5d696583e update zen models 2026-04-29 09:38:00 -04:00
Mike Sukmanowsky 2cb5a99b98 fix: add model card links for Kimi K2 and Kimi K2.5 2026-04-29 09:29:36 -04:00
dpuyosa 70490d892f [venice] Update DeepSeek v4 and Qwen 3.6 model pricing and metadata
- Reduce DeepSeek v4 Flash/Pro input and output pricing
- Add cache_read pricing for DeepSeek v4 models
- Fix Qwen 3.6 27B model name formatting
2026-04-29 09:39:11 +02:00
Aiden Cline f858a85ba6 Merge pull request #1625 from xinrui-z/fix-aihubmix-2026-04-28
fix: sync AIHubMix models (2026-04-28)
2026-04-28 23:03:28 -05:00
Aiden Cline 44e4c92882 Merge pull request #1620 from kill74/add-zai-coding-plan-glm-5v-turbo
Add GLM-5V-Turbo to Z.ai coding plan
2026-04-28 19:28:48 -05:00
Aiden Cline 43eafcb258 Merge pull request #1581 from xiaojiezj/zenmux_0425
feat: add models for zenmux provider
2026-04-28 19:27:38 -05:00
Aiden Cline b382ac7af9 Merge branch 'dev' into zenmux_0425 2026-04-28 19:08:09 -05:00
Tom X Nguyen ca21110644 fix: sync neuralwatt models with updated API pricing and capabilities
The Neuralwatt API now returns accurate pricing and capabilities,
eliminating the need for manual patches (patch.json is now empty).

Changes:
- Update pricing for all 14 models from API (significant changes for
  GLM, GPT-OSS, Qwen, and MiniMax models)
- Devstral Small 2 now supports image input (vision)
- kimi-k2.5-fast now supports image input (vision)
- kimi-k2.6-fast now supports reasoning + image input (was non-reasoning)
- Qwen3.6-35B-A3B now supports reasoning (was non-reasoning)
- GLM models context window: 202,752 → 200,000
- Rename fast variant model IDs to match API (dropped org prefix):
  zai-org/glm-5-fast → glm-5-fast
  zai-org/glm-5.1-fast → glm-5.1-fast
  moonshotai/kimi-k2.5-fast → kimi-k2.5-fast
  moonshotai/kimi-k2.6-fast → kimi-k2.6-fast
  Qwen/qwen3.5-397b-fast → qwen3.5-397b-fast
  Qwen/qwen3.6-35b-fast → qwen3.6-35b-fast
2026-04-29 07:05:23 +07:00
Mike Sukmanowsky 0c2e47e8ba Fix token limits for Amazon Bedrock Kimi K2 models
Correct context and output limits for moonshot.kimi-k2-thinking and
moonshotai.kimi-k2.5 on Amazon Bedrock:
- context: 256_000 → 262_143
- output: 256_000 → 16_000
2026-04-28 17:45:59 -04:00
Aiden Cline 6a0704574b Merge pull request #1621 from YuzhongHuangCS/dev
feat(wandb): Add GLM-5.1
2026-04-28 15:51:43 -05:00
Aiden Cline 1cb1341516 Merge pull request #1635 from stylings/feat/nemotron-3-nano-omni
feat: add Nemotron 3 Nano Omni model
2026-04-28 15:04:07 -05:00
Alex bc47e95427 fix: rename Nemotron Omni metadata 2026-04-28 15:20:24 -04:00
Aiden Cline b071e8add8 Merge pull request #1619 from fernandoenzo/fix/deepseek-v4-pro-ollama-cloud
fix(ollama-cloud): correct deepseek-v4-pro model config
2026-04-28 14:00:42 -05:00
Aiden Cline 23e527753e Merge pull request #1623 from itsnebulalol/dev
feat: add gpt-5.5 pro on openai and openrouter
2026-04-28 14:00:34 -05:00
Aiden Cline e81c045ed0 Merge pull request #1636 from dsingal0/feat/openrouter-deepseek-v4
feat(baseten): add DeepSeek V4 Pro
2026-04-28 13:49:11 -05:00
Dhruv Singal 0c602ca936 feat(baseten): update DeepSeek V4 Pro pricing 2026-04-28 11:45:02 -07:00
Dhruv Singal 9866f84989 feat(baseten): add DeepSeek V4 Pro 2026-04-28 11:38:22 -07:00
Alex 6b397ffe37 fix: align nvidia output limit 2026-04-28 14:35:31 -04:00
Alex bc21596889 fix: drop openrouter provider prefix 2026-04-28 14:22:04 -04:00
Alex 20abba3190 feat: add Nemotron 3 Nano Omni 2026-04-28 14:17:22 -04:00
Dominic Frye 5881bf98a0 fix: enable pdf input modality for gpt-5.5 pro 2026-04-28 13:33:31 -04:00
Aiden Cline 595f7d028c Merge pull request #1632 from rocuevas9511/feat/deepinfra-deepseek-v4-pro
feat: add DeepSeek-V4-Pro to deepinfra
2026-04-28 12:07:22 -05:00
rocuevas9511 c81dec9c5d feat: add DeepSeek-V4-Pro to deepinfra 2026-04-28 10:56:44 -06:00
Guiii 4d45ed25d2 Use extended GLM-5V-Turbo config
Removed various fields and added extends section.
2026-04-28 17:49:34 +01:00
Yuzhong Huang 03cf48de53 use extends instead 2026-04-28 09:19:26 -07:00
Aiden Cline 0d3a284395 Merge pull request #1622 from eduqr/feat/fireworks-ai-deepseek-v4-pro
feat(fireworks-ai): add deepseek-v4-pro
2026-04-28 10:52:54 -05:00
Aiden Cline 5b1bb0fc80 Merge pull request #1624 from shelvick/add-azure-kimi-k2-6
Add Kimi K2.6 to Azure
2026-04-28 10:38:07 -05:00
Aiden Cline 332ebb8811 Merge pull request #1627 from ceoAppsknight/kilo/add-mimo-models
Add Kilo Mimo v2.5 models
2026-04-28 10:37:40 -05:00
Aiden Cline 2111813bd4 Merge pull request #1629 from ndeybach/PR-azure-5.4-limits
fix(azure): correct GPT-5.4 series limits and cleanup
2026-04-28 10:37:01 -05:00
Nils DEYBACH f5b8521af6 fix: use extends and not symlinks 2026-04-28 17:34:44 +02:00
Nils DEYBACH 79481cff40 fix(azure): update GPT-5.4 metadata
Use `extends` for Azure GPT-5.4 variants and keep Azure-specific overrides for
PDF input and omitted fast mode.

Validated with `bun validate`.

Azure runtime manual probing confirmed GPT-5.4 uses the documented 1.05M context /
922K input / 128K output limits.
2026-04-28 14:00:20 +02:00
Nils DEYBACH 40dc356d4c fix(azure): correct GPT-5.4 and GPT-5.4 Pro limits (and convert to extend)
Correct Azure GPT-5.4 and GPT-5.4 Pro limits to `1_050_000` context,
`922_000` input, and `128_000` output based on Azure runtime results and
Microsoft Learn docs. Mini and Nano already matched and are unchanged.

The limits were tested directly (see script at : https://github.com/ndeybach/Azure_endpoint_limit_test_script )
2026-04-28 12:33:27 +02:00
C.C. Fan f929fe89e7 provider(vivgrid): remove GLM-5, add GPT-5.5 model 2026-04-28 16:44:59 +08:00
Syed Assadullah Shah cbd245d454 add Kilo Mimo v2.5 models 2026-04-28 13:04:36 +05:00
xinrui e5289e9b3a fix: sync AIHubMix models (2026-04-28) 2026-04-28 11:33:25 +08:00
Scott Helvick bb623f3ff9 Add Kimi K2.6 to Azure 2026-04-28 02:28:46 +00:00
Dominic Frye 37fffafafc feat: add gpt-5.5 pro on openai and openrouter 2026-04-27 22:24:32 -04:00
eduqr d329310745 feat(fireworks-ai): add deepseek-v4-pro 2026-04-27 21:05:55 -05:00
Yuzhong Huang 8016a6c45a Add GLM-5.1 to wandb provider 2026-04-27 17:35:38 -07:00
Guilherme Sales 1eeaa0b756 Add GLM-5V-Turbo to Z.ai coding plan 2026-04-28 00:47:15 +01:00
Frank dd3533b4e0 update zen models 2026-04-27 19:31:24 -04:00
Fernando Guarini 3a5867834f fix(ollama-cloud): correct deepseek-v4-pro model config
- Remove fields that don't belong in ollama-cloud: temperature, structured_output, knowledge, interleaved
- Set output = context (1048576) per ollama-cloud convention
- Set name to lowercase per ollama-cloud convention
- Reorder fields to match existing ollama-cloud model files
2026-04-28 00:24:54 +02:00
Aiden Cline 1e83bca7a3 Merge pull request #1617 from JoshuaDietz/dev
feat(ollama cloud): add deepseek v4 pro
2026-04-27 16:42:56 -05:00
Aiden Cline cd8853f88b Merge pull request #1616 from fhennerkes/dev
poe: add GPT-5.5 and GPT-5.5-Pro models
2026-04-27 16:08:36 -05:00
Joshua Dietz 23b290c6a8 fix(ollama cloud): fix model name
Model name was inconsistent with naming schema of flash model on ollama cloud
2026-04-27 21:47:30 +02:00
fhennerkes 4e7849cee7 poe: reduce omits in gpt-5.5 extends configs
Inherit family, knowledge, and structured_output from base models
instead of omitting them. Only omit fields that genuinely don't
apply to Poe (provider-specific pricing tiers, different context
limits, opencode-specific provider config).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-27 12:44:49 -07:00
Joshua Dietz e3e63a7247 feat(ollama cloud): add deepseek v4 pro 2026-04-27 21:42:10 +02:00
Aiden Cline fb297153e4 Merge pull request #1572 from YoshiTabletopGamer/qwen3.5-3.6-alibaba-open
[alibaba] Add remaining open Qwen 3.5 and 3.6 models, fix Qwen-3.5 397B-A17B
2026-04-27 14:36:51 -05:00
fhennerkes e8dd06e0ce poe: use extends format for gpt-5.5-pro
Address PR review comment to use extends format. Inherit from
opencode/gpt-5.5-pro since no openai/gpt-5.5-pro base exists yet.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-27 10:49:36 -07:00
fhennerkes bdee3d438b poe: use extends format for gpt-5.5
Address PR review comment to use extends format and inherit from
openai/gpt-5.5 base model.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-27 10:36:10 -07:00
fhennerkes b410bc3ef2 poe: add GPT-5.5 and GPT-5.5-Pro models 2026-04-27 10:34:52 -07:00
Aiden Cline 870e3d1d26 Merge pull request #1612 from hanouticelina/add-deepseek-v4-for-huggingface
feat(huggingface): add DeepSeek V4 Pro
2026-04-27 11:20:12 -05:00
Aiden Cline 5f21f3b603 Merge pull request #1586 from fernandoenzo/add-ollama-cloud-deepseek-v4-flash
feat(ollama-cloud): add deepseek-v4-flash model
2026-04-27 10:42:53 -05:00
Celina Hanouti 27dad05e25 extend deepseek/deepseek-v4-pro 2026-04-27 16:38:36 +01:00
Aiden Cline 2ed88cbcc0 Merge pull request #1604 from ndeybach/PR-gpt-5.5
feat(azure): add GPT-5.5 model metadata
2026-04-27 10:20:32 -05:00
Nils DEYBACH 4d199c932e fix: base azure-cognitive-services model not on azure
extend of extend does not seem to be supported
2026-04-27 17:04:20 +02:00
Jaime de Aquino 2d142f920c Add Qwen3.5-9B model configuration file 2026-04-27 16:46:04 +02:00
Celina Hanouti 47dea9e551 fix 2026-04-27 15:40:35 +01:00
Celina Hanouti 0d25c3dcac use extends 2026-04-27 15:37:31 +01:00
Aiden Cline 3d7f9256cb Merge pull request #1583 from abliteration-ai/codex/add-abliteration-provider
Add abliteration.ai provider
2026-04-27 09:29:26 -05:00
Aiden Cline 6aa1ebd4be Merge pull request #1595 from Contraboi/contra/add-openrouter-nano-banana-2
feat(openrouter): add Gemini 3.1 flash image preview (Nano Banana 2)
2026-04-27 09:28:23 -05:00
Aiden Cline f9ebebaffd Merge pull request #1601 from shikbupt/alibaba-deepseek
add alibaba-cn deepseek-v4
2026-04-27 09:28:08 -05:00
sk 7b3fe83c09 use extend format 2026-04-27 21:48:26 +08:00
Yashwanth Kumar 0a06b3efc2 Update Qwen model configuration in TOML file 2026-04-27 16:40:28 +05:30
Yashwanth Kumar 90dcbbbcc5 Add Qwen3.6 27B model configuration 2026-04-27 16:29:56 +05:30
Yashwanth Kumar b17f5fd8ae Delete providers/openrouter/models/qwen/qwen-3.6-27b.toml 2026-04-27 16:28:19 +05:30
Yashwanth Kumar 68691ac3f9 Update providers/openrouter/models/qwen/qwen-3.6-27b.toml
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2026-04-27 16:27:56 +05:30
Yashwanth Kumar 78fe905fcb Add Qwen 3.6 27B model configuration 2026-04-27 16:22:24 +05:30
Nils DEYBACH 6b488802cf fix: parsing error and add context over price
- adds the context over X price capability from parent (awaiting refactor to be correct on exact limit threashold)
- fix parsing since anything must be before extends.
2026-04-27 12:27:47 +02:00
Jack 4d0505b70e Merge pull request #1611 from anomalyco/fix/opencode-go-deepseek-v4-flash-cache-read-20260427
fix(opencode-go): update deepseek v4 flash cache pricing in Go
2026-04-27 17:34:52 +08:00
Jack b729923bd9 fix(opencode-go): correct deepseek v4 flash cache pricing 2026-04-27 17:31:41 +08:00
Celina Hanouti 34c7aa7dfe update context limit 2026-04-27 09:33:57 +01:00
Celina Hanouti 09b4d3548b add support for DeepSeek V4 Pro for Hugging Face provider 2026-04-27 09:31:55 +01:00
xiaojie.zj f1cad8fdc0 feat: add zenmux models 2026-04-27 16:27:46 +08:00
Tom X Nguyen b2f7f57f26 feat: add neuralwatt provider with 14 models
Add Neuralwatt as an OpenAI-compatible inference provider with
energy-aware GPU optimization. Includes 14 models across 6
sub-providers (Mistral, ZAI, OpenAI, Moonshot, MiniMax, Qwen).

Models include reasoning variants (Kimi K2.5/K2.6, GLM 5.1 FP8,
MiniMax M2.5, Qwen3.5 397B, GPT OSS 20B) and fast non-reasoning
variants (Kimi K2.5/K2.6 Fast, GLM 5/5.1 Fast, Qwen3.5/3.6 Fast),
plus Devstral Small 2 and Qwen3.6 35B A3B.

Logo derived from official Neuralwatt favicon (currentColor variant).
Pricing sourced from Neuralwatt's published rates.
2026-04-27 15:01:56 +07:00
Aiden Cline 925d4eba1f Merge pull request #1536 from philipmat/add-openrouter-pareto-code-router
Adds support for openrouter/pareto-code
2026-04-26 23:52:04 -05:00
Aiden Cline bc1e4b870b Merge pull request #1608 from Alex-wuhu/dev
add deepseek-v4, qwen3.6 on novita
2026-04-26 23:16:56 -05:00
Alex-wuhu ef913f9645 refactor: use extends format for novita deepseek v4 models
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-04-27 12:00:04 +08:00
Alex-wuhu cec747ade3 fix: add cache_read pricing for deepseek v4 models 2026-04-27 11:09:15 +08:00
Aiden Cline 2ced32d52f Merge pull request #1596 from NathanDrake2406/add-cf-ai-gateway-gpt-5.5
feat(cloudflare-ai-gateway): add openai/gpt-5.5
2026-04-26 21:57:13 -05:00
Alex-wuhu 92cc5a79df feat: add deepseek-v4, qwen3.6 on novita 2026-04-27 10:49:13 +08:00
Aiden Cline 4cf8661f92 Merge pull request #1602 from shikbupt/alibaba-qwen3.6-max
add alibaba-cn qwen3.6 max
2026-04-26 17:48:14 -05:00
Aiden Cline ebe431c0dc Merge pull request #1606 from LightAndy1/dev
Add gemini-3.1-flash-preview for google-vertex
2026-04-26 17:30:05 -05:00
LightAndy 8b2c5f30a0 ✏️ Fix typo 2026-04-26 21:48:54 +03:00
Nils DEYBACH 0fc92e1c47 fix: align limit on base 5.5 model
now that limits were fixed in base, we align azure on it
2026-04-26 20:48:00 +02:00
Nils DEYBACH d73ca9024a Merge remote-tracking branch 'upstream/dev' into PR-gpt-5.5 2026-04-26 20:45:20 +02:00
LightAndy bff48f9fde Merge branch 'anomalyco:dev' into dev 2026-04-26 21:43:13 +03:00
Aiden Cline c5c803f415 fix: ensure openai gpt-5.5 limits are exact 2026-04-26 13:38:56 -05:00
Nils DEYBACH cd69509b48 fix: simplify by extending the azure model from opeani 2026-04-26 20:04:27 +02:00
Aiden Cline 83d15fd756 Merge pull request #1574 from juls0730/dev
feat: add mimo v2.5/pro to xiaomi and openrouter providers
2026-04-26 13:55:57 -04:00
LightAndy b3c09451d9 Add gemini-3.1-flash-preview model configuration 2026-04-26 20:27:17 +03:00
Frank d98bdb5eff Merge pull request #1560 from TigerBeanst/patch-1
fix: opencode go mimo-v2.5 context limit to 1,000,000
2026-04-26 13:16:33 -04:00
Frank ea205913ce update zen models 2026-04-26 12:49:37 -04:00
Frank 3532801639 update zen models 2026-04-26 11:48:32 -04:00
Nils DEYBACH 1bb50141c2 feat(azure): add GPT-5.5 model metadata
## Summary

Adds GPT-5.5 metadata for:

- Azure
- Azure Cognitive Services

The Azure Cognitive Services entry mirrors the existing local convention of full TOML model definitions.

## Sources

- Microsoft Learn lists `gpt-5.5` for Azure OpenAI / Microsoft Foundry with version `2026-04-24`, `1,050,000` context, `922,000` input, `128,000` output, structured outputs, tools, image input, and December 2025 training data.
- Microsoft’s Azure GPT-5.5 announcement lists pricing at `$5.00` input, `$0.50` cached input, and `$30.00` output per 1M tokens.
- Azure Responses API docs list PDF input support for vision-capable models and include `gpt-5.5` version `2026-04-24`.

## Notes

This intentionally does not add `gpt-5.5-pro`, since Azure Learn currently lists `gpt-5.5` but not `gpt-5.5-pro` in the Azure model catalog.

This also intentionally omits OpenAI-specific `context_over_200k` and `experimental.modes.fast` metadata because the Azure sources confirm the standard pricing and limits, but not those OpenAI-specific fields.
2026-04-26 17:45:25 +02:00
sk fa8bdd63fb add alibaba-cn qwen3.6 max 2026-04-26 19:39:41 +08:00
sk 56f64577ce add alibaba-cn deepseek-v4 2026-04-26 19:23:04 +08:00
Nathan Nguyen f726af5767 refactor(cloudflare-ai-gateway): use [extends] for openai/gpt-5.5
The model entry duplicated every field from providers/openai/models/gpt-5.5.toml,
so any future change to the upstream OpenAI definition would silently drift here.

Switch to the `[extends] from = "openai/gpt-5.5"` form already used by sibling
providers (openrouter, requesty), omitting `experimental.modes.fast` since the
gateway does not surface the OpenAI priority service tier. Validation output is
byte-identical to the prior expanded form.
2026-04-26 13:33:39 +10:00
Zoe 4838e3cb9b feat: add mimo v2.5/pro to xiaomi and openrouter providers 2026-04-25 21:32:24 -05:00
Muhammad Mugni Hadi dbe92646c3 chore(chutes): add header comments to generated TOML files
Each generated TOML now includes a comment noting which fields are
auto-managed vs manually overridable on re-run.
2026-04-26 06:52:21 +07:00
Muhammad Mugni Hadi 4717c67054 feat(chutes): add API-driven model generator script
Add generate-chutes.ts that fetches models from https://llm.chutes.ai/v1/models
and generates/updates TOML files, following the same pattern as generate-vercel.ts.

Supports --dry-run, --new-only, and --keep-orphans flags. Auto-deletes TOML files
for models no longer in the API (with empty directory cleanup).

Preserves manually-set fields (family, knowledge, interleaved, status) when merging
with API data. Also syncs current models from the API.
2026-04-26 06:51:16 +07:00
Nathan Nguyen d4c77c14fd feat(cloudflare-ai-gateway): add openai/gpt-5.5
Mirrors the existing direct openai/gpt-5.5 entry under the
cloudflare-ai-gateway provider so opencode and other consumers can
route GPT-5.5 traffic through Cloudflare AI Gateway without hitting
ProviderModelNotFoundError.

Pricing, limits, modalities, and dates copied from
providers/openai/models/gpt-5.5.toml; provider stanza follows the
sibling gpt-5.4 entry (npm = "ai-gateway-provider").
2026-04-26 05:22:29 +10:00
Selmir Nedzibi 96b3d65307 feat(openrouter): add Gemini 3.1 flash image preview (Nano Banana 2) 2026-04-25 21:14:44 +02:00
Aiden Cline b491c29cf9 Merge pull request #1573 from zainhas/dev
[Together AI] add deepseek-v4
2026-04-25 13:47:14 -04:00
Aiden Cline d937abd849 Merge pull request #1539 from manascb1344/fix-xiaomi-provider-ids
feat: add MiMo-V2.5 and MiMo-V2.5-Pro to xiaomi-token-plan providers
2026-04-25 13:33:42 -04:00
Aiden Cline df52175b0c Merge pull request #1580 from LeGazeon/add-nvidia-deepseek-v4-pro/flash
Add NVIDIA DeepSeek-V4 models
2026-04-25 13:32:41 -04:00
Aiden Cline bee8339c07 Merge pull request #1589 from MiyakoMeow/feat/restrict-zai-zhipuai-coding-plan-models
rm: unavailable models in zai/zhipuai coding plan
2026-04-25 13:30:39 -04:00
Aiden Cline 181bf96fa3 Merge pull request #1585 from saju01/add-copilot-gpt-5.5
feat(github-copilot): add gpt-5.5
2026-04-25 13:30:16 -04:00
Aiden Cline 9d49d2fd52 Merge pull request #1587 from smakosh/claude/rebase-add-llmgateway-models-yqKLn
feat(llmgateway): add deepseek-v4-pro, deepseek-v4-flash, kimi-k2.6
2026-04-25 13:29:39 -04:00
Aiden Cline 648776aa85 Merge pull request #1590 from dpuyosa/feat/venice-models
Venice: Add GPT-5.5 and Qwen3.6 model configs
2026-04-25 13:29:03 -04:00
Aiden Cline 421cb099b0 Merge pull request #1591 from dpuyosa/fix/venice-deepseek-family
Venice: Fix DeepSeek V4 Flash family classification
2026-04-25 13:28:54 -04:00
Aiden Cline f458b19994 Merge pull request #1592 from MiyakoMeow/feat/deepseek-1m-context
fix(deepseek): all has 1M context / 384k output / adjusted price
2026-04-25 13:28:45 -04:00
MiyakoMeow d347093b03 feat(deepseek): 1M context / 384k output 2026-04-25 18:58:12 +08:00
MiyakoMeow 3328712262 feat: restrict zai/zhipuai coding plan models to glm-5.1, glm-5-turbo, glm-4.7, glm-4.5-air only
Based on official documentation:
- ZAI DevPack Coding Plan: https://docs.z.ai/devpack/overview
- Zhipu AI BigModel Coding Plan: https://docs.bigmodel.cn/cn/coding-plan/overview

Both providers only officially support the following GLM models for coding plans:
- glm-5.1
- glm-5-turbo
- glm-4.7
- glm-4.5-air

Removed unsupported models from zai-coding-plan:
- glm-4.5, glm-4.5-flash, glm-4.5v
- glm-4.6, glm-4.6v
- glm-4.7-flash, glm-4.7-flashx
- glm-5, glm-5v-turbo

Removed unsupported models from zhipuai-coding-plan:
- glm-4.5, glm-4.5-flash, glm-4.5v
- glm-4.6, glm-4.6v, glm-4.6v-flash
- glm-4.7-flash, glm-4.7-flashx
- glm-5, glm-5v-turbo
2026-04-25 18:49:10 +08:00
dpuyosa 60edc1b52d [venice] Add GPT-5.5 and Qwen3.6 model configs
- Add OpenAI GPT-5.5 with 1M context window and tiered pricing
- Add OpenAI GPT-5.5 Pro with premium pricing and 128K output limit
- Add Qwen3.6 27B with text, image, and video input modalities
2026-04-25 12:41:31 +02:00
dpuyosa eee44cd080 [venice] Fix DeepSeek V4 Flash family classification
- Correct family from "deepseek" to "deepseek-flash" for accurate model categorization
2026-04-25 12:36:06 +02:00
smakosh 048a3235e8 feat(llmgateway): add deepseek-v4-pro, deepseek-v4-flash, kimi-k2.6 2026-04-25 12:19:40 +02:00
Fernando Guarini 1db03ec1e6 feat(ollama-cloud): add deepseek-v4-flash model 2026-04-25 11:24:12 +02:00
Saju Sarangdharan 6d283349ad feat(github-copilot): add gpt-5.5
GitHub Copilot now serves gpt-5.5 (verified via GET https://api.githubcopilot.com/models with a Copilot Enterprise token). Adding the catalog row so downstream consumers (e.g. pi-ai) can route requests.
2026-04-25 10:31:21 +02:00
Abliteration.ai 7e07302ecd add abliteration.ai provider 2026-04-24 22:46:54 -07:00
LeGazeon 8bc407a617 chore: remove deepseek-v4-pro config (duplicated by #1578)
The Pro model configuration was already added via #1578 which
has been merged. Removing the duplicate from this branch to
keep only the Flash variant.
2026-04-25 13:17:12 +08:00
LeGazeon 5305d2bae2 refactor: extend flash config from deepseek base
Remove duplicated fields by inheriting common settings
from providers/deepseek base config via [extends].

This addresses the review comment in #1580
2026-04-25 13:11:26 +08:00
Aiden Cline fee96c27b9 Merge pull request #1578 from panwar-stack/dev
feat(nvidia): add DeepSeek V4 model
2026-04-25 00:41:03 -04:00
Aiden Cline 66520adbc6 Merge pull request #1577 from ezShroom/dev
add openrouter gpt-5.5
2026-04-25 00:40:34 -04:00
Zain Hasan 40714995cc Add interleaved section to DeepSeek-V4-Pro.toml 2026-04-24 18:59:45 -07:00
LeGazeon 66c4896003 Add NVIDIA DeepSeek-V4 models
Add model entries for DeepSeek V4 Pro and DeepSeek V4 Flash to the NVIDIA NIM provider.

## Changes
- Added `providers/nvidia/deepseek-v4-pro.toml`
- Added `providers/nvidia/deepseek-v4-flash.toml`

## Data Sources
- NVIDIA NIM Model Cards:
  - DeepSeek V4 Pro: https://build.nvidia.com/deepseek-ai/deepseek-v4-pro/modelcard
  - DeepSeek V4 Flash: https://build.nvidia.com/deepseek-ai/deepseek-v4-flash/modelcard
2026-04-25 09:57:43 +08:00
panwar-stack 31091f3d4e Rename deepseek-v4.toml to deepseek-v4-pro.toml 2026-04-24 17:09:12 -07:00
panwar-stack 5fe512c1b1 Follow extends pattern
Follow extends pattern
2026-04-24 17:08:49 -07:00
panwar-stack 63efa131e7 feat(nvidia): add DeepSeek V4 model
add DeepSeek V4 model
2026-04-24 17:05:12 -07:00
Shroom 29cd503070 Add gpt-5.5.toml configuration file 2026-04-25 00:08:04 +01:00
Rohan Taneja a9b704c656 Merge pull request #1575 from vercel/update-vercel-models-1777063875 2026-04-24 15:29:39 -07:00
Aiden Cline 0a88e412e5 Merge pull request #1576 from dsingal0/feat/openrouter-deepseek-v4
Add OpenRouter DeepSeek V4 models
2026-04-24 17:42:48 -04:00
Dhruv Singal b46e29ccd5 fix(openrouter): use DeepSeek reasoning content field 2026-04-24 14:06:05 -07:00
Jerilyn Zheng a181b770d6 Update kimi-k2.6.toml 2026-04-24 13:55:27 -07:00
Jerilyn Zheng fa71201f20 Update deepseek-v4-pro.toml 2026-04-24 13:54:56 -07:00
Jerilyn Zheng 3ac17aefb6 Enable open_weights in deepseek-v4-flash configuration 2026-04-24 13:54:34 -07:00
Jerilyn Zheng 1f2ceb91a5 Update qwen-3.6-max-preview.toml 2026-04-24 13:53:57 -07:00
github-actions[bot] d82681900c chore(vercel): update Vercel model definitions
Auto-generated by weekly workflow from Vercel AI Gateway API.

Co-Authored-By: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-24 20:51:22 +00:00
Zain Hasan 7e296f4cbd ds recommend 384,000 2026-04-24 12:39:09 -07:00
Zain Hasan f202ff51e4 Reduce output limit from 512000 to 300000 2026-04-24 12:13:53 -07:00
Zain Hasan 4085a0b536 Merge branch 'anomalyco:dev' into dev 2026-04-24 12:12:28 -07:00
Zain Hasan 935fdeca65 [Together AI] Add deepseekv4 pro 2026-04-24 12:08:33 -07:00
Frank cef8828fbe update zen models 2026-04-24 14:51:59 -04:00
Frank 9277a23a29 update zen models 2026-04-24 14:50:50 -04:00
Frank d7bfb16b0f update zen models 2026-04-24 14:42:25 -04:00
Frank 37ae6fe7c8 update zen models 2026-04-24 12:11:33 -04:00
Aiden Cline 87ce527476 Merge pull request #1571 from cgilly2fast/dev
fix(firmware): glm 5.1 name
2026-04-24 11:49:51 -04:00
Aiden Cline 34bed30fa1 Merge pull request #1567 from dsingal0/feat/openrouter-deepseek-v4
feat(openrouter): add DeepSeek V4 Pro and V4 Flash
2026-04-24 11:49:35 -04:00
Dhruv Singal 8a0c3cb75f refactor(openrouter): extend official DeepSeek V4 Pro/Flash
Use [extends] from deepseek/ with OpenRouter-specific overrides
(attachment, interleaved reasoning_details, limits).

Made-with: Cursor
2026-04-24 08:43:38 -07:00
Colby Gilbert f65460ac16 fix(firmware): glm 5.1 name 2026-04-24 08:40:19 -07:00
YoshiTabletopGamer 8ea92aed9a [alibaba] Add remaining open Qwen 3.5 models, fix Qwen-3.5 397B-A17B, add open Qwen 3.6 models
- Added Qwen 3.5 122B-A10B
- Added Qwen 3.5 27B
- Added Qwen 3.5 35B-A3B
- Fixed Qwen 3.5 397B-A17B (see below)
- Added Qwen 3.6 27b
- Added Qwen 3.6-35B-A3B

I was not able to find a reliable source for the knowledge cutoff of any of these models.
2025-04 was already set as the cutoff for Qwen 3, and Qwen 3.5 is newer.
All data is from the ModelStudio webpage.
It seems to not include audio, but the ModelStudio page clearly has an audio symbol and the model is capable of this.
And I found no data for a price for reasoning tokens in particular, unlike what was in the file for Qwen 3.5 397B-A17B.
The models are all capable of structured output.
2026-04-24 12:38:52 -03:00
Dhruv Singal f340d82fc3 feat(openrouter): add DeepSeek V4 Pro and V4 Flash
Add model configs aligned with OpenRouter pricing and limits
(1M context, 384K max output, cache read rates from provider page).

Made-with: Cursor
2026-04-24 08:25:17 -07:00
Frank c7431ae24c update zen models 2026-04-24 10:53:10 -04:00
Frank 3d1888b7b5 update zen models 2026-04-24 10:24:34 -04:00
Misha Skvortsov 16a8fa5c20 improve(atomic-chat): drop hardcoded model list per maintainer feedback
Made-with: Cursor
2026-04-24 17:20:26 +03:00
Aiden Cline dcd37ccdbb add deepseek v4 flash 2026-04-24 08:34:03 -04:00
Aiden Cline d18c3f910c Merge pull request #1562 from dpuyosa/update/venice-kimi-pricing
Venice: Update kimi-k2-6 pricing
2026-04-24 08:08:23 -04:00
Aiden Cline 2cec5a492c Merge pull request #1563 from dpuyosa/feat/venice-deepseek-v4
Venice: Add DeepSeek V4 Flash and Pro models
2026-04-24 08:08:13 -04:00
dpuyosa 61dd0ec489 [venice] Add DeepSeek V4 Flash and Pro models
- Add DeepSeek V4 Flash with 1M context, reasoning, and tool support
- Add DeepSeek V4 Pro with 1M context, reasoning, and tool support
- Set pricing and interleaved reasoning_content field for both
2026-04-24 12:21:07 +02:00
dpuyosa c7758204b5 [venice] Update kimi-k2-6 pricing
- Update input, output, and cache_read costs to current rates
- Update last_updated timestamp to 2026-04-24
2026-04-24 12:17:51 +02:00
manascb1344 ed91520aa2 feat: add MiMo-V2.5 and MiMo-V2.5-Pro to xiaomi-token-plan providers 2026-04-24 15:38:55 +05:30
Frank 3e82669a82 Merge pull request #1561 from wenbindu/dev
add deepseek new moels
2026-04-24 03:04:52 -04:00
Frank 1cc0c9c074 sync 2026-04-24 03:03:06 -04:00
TigerBeanst d73d7f6453 fix: opencode go mimo-v2.5 context limit to 1,000,000
https://platform.xiaomimimo.com/docs/pricing
2026-04-24 12:48:34 +08:00
Aiden Cline afb59f86ee Merge pull request #1557 from seffhunnn/dev
feat: add AU Sonnet and Opus models for Amazon Bedrock
2026-04-24 00:30:51 -04:00
wenbindu 05242f68d4 add deepseek new moel 2026-04-24 12:06:19 +08:00
Mohd Saif c1b029dcc1 feat: add AU Opus model for Amazon Bedrock 2026-04-24 03:26:28 +05:30
Mohd Saif bc2dd5137a feat: add AU Sonnet model for Amazon Bedrock 2026-04-24 03:25:38 +05:30
Aiden Cline 99ec4900c7 Merge pull request #1555 from brentdurksen/add-azure-claude-sonnet-4-6
feat(azure): add Claude Sonnet 4.6 model
2026-04-23 17:33:44 -04:00
Brent Durksen a8c124ac9e refactor: use extends to inherit from anthropic/claude-sonnet-4-6 2026-04-23 15:16:38 -06:00
Aiden Cline 0d20a363a9 Merge pull request #1556 from fhennerkes/dev
poe: add GPT-Image-2 model
2026-04-23 17:12:20 -04:00
fhennerkes 3ac613678b poe: add GPT-Image-2 model 2026-04-23 12:38:22 -07:00
Brent Durksen e2ead1b4e6 feat(azure): add Claude Sonnet 4.6 model 2026-04-23 13:37:50 -06:00
Aiden Cline be53c33588 Merge pull request #1550 from BlockListed/cortecs-kimi-k2.6
Add kimi k2.6 to cortecs
2026-04-23 15:26:58 -04:00
Aiden Cline 55cf5fa310 Merge pull request #1554 from mattyatea/add-gpt-5-5
[codex] Add GPT-5.5
2026-04-23 15:17:20 -04:00
mattyatea 3e0fe362f2 add gpt-5.5 model 2026-04-24 04:13:51 +09:00
BlockListed 89d06ae31f add kimi k2.6 to cortecs 2026-04-23 19:49:18 +02:00
Aiden Cline c994b116ae Merge pull request #1542 from u007/patch-1
Add Chutes: Kimi K2.6 TEE
2026-04-23 12:47:49 -04:00
Aiden Cline 833e8f7a66 Merge pull request #1548 from fernandoenzo/fix/gemma4-ollama-output-limit
fix(ollama): set gemma4:31b output limit to match context
2026-04-23 12:46:00 -04:00
Aiden Cline 32bd1427fb Merge pull request #1545 from Alex-wuhu/dev
Add deepseek, gemma, ling, llama, kimi on NovitaAI
2026-04-23 12:36:06 -04:00
Frank ae7672b87e update zen models 2026-04-23 11:11:30 -04:00
Fernando Guarini 9b27cc5a76 fix(ollama): set gemma4:31b output limit to match context
Ollama does not impose official output limits. The existing convention for Gemma models on Ollama (gemma3:4b, gemma3:12b, gemma3:27b) is to set output equal to context. gemma4:31b was the only exception with output=8192 vs context=262144.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-04-23 11:42:30 +02:00
Alex-wuhu 7014e3e318 feat: add missing Novita AI model configurations
Add 6 models served by Novita:
- deepseek/deepseek-r1-distill-qwen-14b
- deepseek/deepseek-r1-distill-qwen-32b
- google/gemma-3-12b-it
- inclusionai/ling-2.6-1t
- meta-llama/llama-3.2-3b-instruct
- moonshotai/kimi-k2.6

Capabilities, pricing, context, and modalities sourced from Novita's
/v1/models API; family slugs and release dates aligned with existing
same-model entries in the repo.
2026-04-23 14:58:43 +08:00
Frank e0e153e8d6 update zen models 2026-04-23 02:51:28 -04:00
mickalchen a0fbef93e4 revert 2026-04-23 14:41:05 +08:00
mickalchen e4a89229be add model by openrouter 2026-04-23 14:23:06 +08:00
mickalchen 9ad15e97cd add model by openrouter 2026-04-23 14:18:51 +08:00
James e3b4df4e87 Update Kimi-K2.6-TEE.toml
fix reasoning
2026-04-23 13:53:37 +08:00
Aiden Cline e3e2066c83 Merge pull request #1544 from GodTamIt/deepinfra/kimi-k2.6
deepinfra: Add Kimi-K2.6 support
2026-04-23 00:40:50 -04:00
Aiden Cline 208dcd12a2 Merge pull request #1533 from qychen2001/dev
Add Kimi-K2.6 and Qwen3.6-35B-A3B, update Kimi-K2.5 config for siliconflow and siliconflow-cn
2026-04-23 00:40:40 -04:00
Aiden Cline d166444aa1 Merge pull request #1541 from zainhas/dev
[Together AI] add Kimi k2.6 support
2026-04-23 00:40:15 -04:00
Aiden Cline cd35f96e73 Merge pull request #1528 from seffhunnn/dev
Fix incorrect model ID for Gemma 4 26B (Google provider)
2026-04-23 00:39:12 -04:00
Christopher Tam 2b0c21e0b4 deepinfra: Add Kimi-K2.6 support 2026-04-22 23:39:42 -04:00
James 787b0bf9e9 Update Kimi K2.5 TEE to Kimi K2.6 TEE 2026-04-23 10:58:39 +08:00
Zain Hasan 77fbf02ff6 remove interleaved 2026-04-22 16:19:10 -07:00
Zain Hasan 89973fd122 [Together AI] add Kimi k2.6 support 2026-04-22 16:16:08 -07:00
Frank e458a9f5b8 Merge pull request #1540 from dsingal0/feat/baseten-kimi-k2.6
feat(baseten): add Kimi K2.6
2026-04-22 16:59:51 -04:00
Dhruv Singal 61d95afae2 feat(baseten): add Kimi K2.6 2026-04-22 13:53:46 -07:00
Jack 4a2df5e008 Merge pull request #1537 from anomalyco/feat/opencode-go-mimo-v2.5
Feat/opencode go mimo-v2.5-pro & mimo-v2.5
2026-04-23 00:51:49 +08:00
Aiden Cline 3db907f3ce Merge pull request #758 from regolo-ai/dev
Add Regolo-ai Provider
2026-04-22 12:33:02 -04:00
Jack 70d8f9cc6e update mimo v2 output limits to 128k 2026-04-22 23:32:16 +08:00
Jack b9b354ada0 update mimo v2.5 output limits to 128k 2026-04-22 23:17:40 +08:00
Jack 95632ad376 remove mimo-v2.5-omni (renamed to mimo-v2.5) 2026-04-22 23:15:00 +08:00
Philip M 7153506989 Adds support for openrouter/pareto-code
The Pareto Router is a way to have OpenRouter always pick a strong coding model for your needs without committing to a specific one. You express a single min_coding_score preference between 0 and 1, and the router routes your request to a coding model that meets that bar.

The Pareto Router is tuned for coding use cases. Under the hood it keeps a curated shortlist of strong coding models currently available on OpenRouter. The exact shortlist and selection logic evolve over time as new models land and benchmarks shift.
2026-04-22 10:14:46 -05:00
Jack c73bab2e7e providers(opencode-go): rename mimo-v2.5-omni to mimo-v2.5 2026-04-22 23:14:27 +08:00
Jack a783dc808d providers(opencode-go): add mimo v2.5 models and separate v2 families 2026-04-22 23:07:40 +08:00
Daniele Scasciafratte 58e72802bc feat(models): update 2026-04-22 16:05:48 +02:00
Mohd Saif 19233d93c4 fix: remove unnecessary id field 2026-04-22 15:19:47 +05:30
QiyuanChen 7eea45e078 feat(siliconflow-cn): add Kimi-K2.6, Qwen3.6-35B-A3B and update Kimi-K2.5 config 2026-04-22 13:38:34 +08:00
QiyuanChen 001ec226f6 feat(siliconflow): add Kimi-K2.6 and update Kimi-K2.5 config 2026-04-22 13:37:59 +08:00
Jack 32461d5b44 Merge pull request #1532 from chl-0537/feature/add-tencent
Remove tencent token plan
2026-04-22 13:12:06 +08:00
mickalchen 9f078294c0 Remove tencent token plan 2026-04-22 13:07:03 +08:00
Aiden Cline c885ed49cd Merge pull request #1531 from zhiyuan1024/zhiyuan/alibaba-cn_kimi-k2.6
feat(alibaba-cn): add Kimi K2.6 model configuration
2026-04-21 23:41:29 -04:00
Aiden Cline 2fc434062f Merge pull request #1529 from compumike/compumike/fix-openrouter-openai-gpt-5.4-pricing
Fix pricing for openrouter/openai gpt-5.4-[mini,nano] off by 10^6
2026-04-21 23:40:51 -04:00
Aiden Cline dbcb7e6d69 Merge pull request #1530 from cgilly2fast/dev
feat(firmware): kimi k2.6 model
2026-04-21 23:40:22 -04:00
Zhiyuan Hou b58392fc62 feat(alibaba-cn): add Kimi K2.6 model configuration
Signed-off-by: Zhiyuan Hou <zhiyuan2048@outlook.com>
2026-04-22 10:37:37 +08:00
Frank a4818c90ca update zen models 2026-04-21 20:18:35 -04:00
Colby Gilbert 49fbdba49f feat(firmware): kimi k2.6 model 2026-04-21 17:14:46 -07:00
Mike Robbins f2dd4da7f9 Fix pricing for openrouter/openai gpt-5.4-[mini,nano] off by 10^6 2026-04-21 18:41:02 -04:00
Frank a3ed215038 update zen models 2026-04-21 17:43:38 -04:00
Mohd Saif d45df0530b fix: correct Gemma 4 26B model ID for Google provider
Updated model ID from gemma-4-26b-it to gemma-4-26b-a4b-it to match actual Gemini API. Also added missing id field and renamed the file accordingly.
2026-04-22 01:51:23 +05:30
Jack 990531258b Merge pull request #1527 from anomalyco/feat/opencode-go-kimi-k2.6-3x-name
providers(opencode-go): rename kimi k2.6
2026-04-21 22:59:26 +08:00
Jack cd2c9e3b62 providers(opencode-go): rename kimi k2.6 2026-04-21 22:54:53 +08:00
Aiden Cline a114991278 Merge pull request #1505 from rocuevas9511/feat/deepinfra-qwen-3.5-35b
feat: add Qwen 3.5 35B A3B to deepinfra
2026-04-21 10:02:27 -04:00
Aiden Cline 5ff2035cee Merge pull request #1520 from Marenz/add-deepinfra-qwen3.6-35b-a3b
Add Qwen3.6-35B-A3B to Deep Infra
2026-04-21 10:01:56 -04:00
Aiden Cline a5993cc140 Merge pull request #1515 from llc1123/chore/zenmux-update
providers(zenmux): add support for kimi k2.6
2026-04-21 10:00:06 -04:00
Aiden Cline aebe4b6cd0 Merge pull request #1517 from otterDeveloper/kimi2.6-pull
add Firework's kimi k2.6
2026-04-21 09:59:54 -04:00
Aiden Cline 214adb1154 Merge pull request #1522 from ceoAppsknight/kilo/kimi-k2.6
Added kilo/kimi-k2.6
2026-04-21 09:59:33 -04:00
Aiden Cline b301c1f8b6 Merge pull request #1523 from sk0x0y/feature/nanogpt-kimi-k2.6-qwen-3.6
feat(nano-gpt): add Kimi K2.6 and Qwen 3.6 models
2026-04-21 09:58:45 -04:00
Aiden Cline b81c8b385b Merge pull request #1507 from rocuevas9511/feat/deepinfra-qwen-3.5-397b
feat: add Qwen 3.5 397B A17B to deepinfra
2026-04-21 09:58:28 -04:00
rocuevas9511 1caa438b3e fix: remove id field (per Marenz feedback) 2026-04-21 07:35:52 -06:00
rocuevas9511 673bc92f4d fix: remove id field (per Marenz feedback) 2026-04-21 07:35:38 -06:00
Jack 0159eaa158 Merge pull request #1526 from anomalyco/feat/moonshotai-cn-kimi-k2.6
providers(moonshotai-cn): add kimi k2.6
2026-04-21 20:39:05 +08:00
Jack efdc7b9a54 providers(moonshotai-cn): add kimi k2.6 2026-04-21 20:26:19 +08:00
Jack b08721206d Merge pull request #1525 from anomalyco/feat/moonshotai-kimi-k2.6
providers(moonshotai): add kimi k2.6
2026-04-21 19:35:13 +08:00
Jack fe0d4cd9fa providers(moonshotai): add kimi k2.6 2026-04-21 19:33:06 +08:00
sk0x0y 738ad6ed70 feat(nano-gpt): add Kimi K2.6 and Qwen 3.6 model family 2026-04-21 19:22:38 +09:00
Syed Assadullah Shah 19b66fbaf2 Added kilo/kimi-k2.6 2026-04-21 15:19:17 +05:00
Mathias L. Baumann c416836497 Add Qwen3.6-35B-A3B to Deep Infra
35B-total / 3B-active MoE (256 experts, 8 routed + 1 shared).
262K native context, vision + video input, thinking mode, tool calls.
Apache 2.0, $0.20 in / $1.00 out per 1M tokens.
2026-04-21 11:58:04 +02:00
rocuevas9511 059cc4b91f fix: update model id to match DeepInfra API 2026-04-21 00:33:20 -06:00
rocuevas9511 116bd328a6 fix: update model id to match DeepInfra API 2026-04-21 00:31:22 -06:00
Frank aa30ce3ef2 update zen models 2026-04-21 02:12:27 -04:00
Frank 9533a47906 update zen models 2026-04-21 01:20:47 -04:00
Miguel Medina ad18f780f0 add firework's kimi 2.6 2026-04-20 23:12:11 -06:00
粒粒橙 a48519e557 providers(zenmux): add support for kimi k2.6 2026-04-21 10:13:22 +08:00
Aiden Cline 23f5e74392 Merge pull request #1514 from mfbalestra/add/kimi-k2.6-ollama-cloud
providers/ollama-cloud: add kimi-k2.6:cloud
2026-04-20 21:47:51 -04:00
Aiden Cline 7a50ea28a1 Merge pull request #1510 from dpuyosa/feat/venice-add-kimi-k2-6
Venice: Add Kimi K2.6 model configuration
2026-04-20 21:45:49 -04:00
mfbalestra 5944f94197 providers(ollama-cloud): add kimi-k2.6:cloud 2026-04-20 22:45:01 -03:00
Aiden Cline 98732b7d76 Merge pull request #1512 from SomeoneWithOptions/dev
add kimi-K2.6 for OpenRouter provider
2026-04-20 21:44:39 -04:00
SomeoneWithOptions c6412d7e59 add kimi-K2.6 for OpenRouter provider 2026-04-20 18:59:49 -05:00
dpuyosa b5a7a6e974 [venice] Add Kimi K2.6 model configuration
- Add new model definition for Venice provider
- Include cost, limits, and modality specs
- Enable reasoning, tool calling, and image input
2026-04-21 00:40:19 +02:00
Aiden Cline ba7c3d7b0b Merge pull request #1509 from kostiak/patch-1
Add support for Kimi-K2.6 in Kimi For Coding provider
2026-04-20 18:02:04 -04:00
Aiden Cline 53a9a2a36d Merge pull request #1508 from hanouticelina/add-kimi-k2.6-modeling
feat(huggingface): add Kimi K2.6
2026-04-20 18:01:32 -04:00
kostiak 3b6bca9b90 Add support for Kimi-K2.6 for Kimi For Coding provider 2026-04-21 00:24:51 +03:00
Celina Hanouti f5a048060f add support for Kimi-K2.6 for Hugging Face provider 2026-04-20 21:41:14 +01:00
rocuevas9511 9e6178a5e3 fix: update cost for Qwen 3.5 397B A17B 2026-04-20 13:38:45 -06:00
rocuevas9511 e4150361b3 fix: update cost for Qwen 3.5 35B A3B 2026-04-20 13:38:06 -06:00
rocuevas9511 8de0fc059d add Qwen 3.5 397B A17B to deepinfra 2026-04-20 13:34:56 -06:00
rocuevas9511 054733884f add Qwen 3.5 35B A3B to deepinfra 2026-04-20 13:32:54 -06:00
Aiden Cline 3d09981eda Merge pull request #1499 from rovo89/patch-1
[google] Fix cache_read cost in gemini-2.5-flash model
2026-04-20 14:43:53 -04:00
Aiden Cline 6877af7770 Merge pull request #1501 from mchenco/kimi-k2.6
Add Kimi K2.6 to Workers AI and AI Gateway
2026-04-20 14:41:58 -04:00
Nacho F. Lizaur 802985f76c feat: update kiro provider to use kiro-acp-ai-provider, add opus 4.7 2026-04-20 20:13:43 +02:00
mchen 794993fd48 Add Kimi K2.6 to Workers AI and AI Gateway 2026-04-20 13:54:31 -04:00
Jack 00b53a422a separate opencode-go kimi k2 families 2026-04-21 01:00:50 +08:00
Jack 2ccecb6011 Merge pull request #1500 from chl-0537/feature/add-tencent
feat: rename model
2026-04-20 22:49:09 +08:00
mickalchen 9b7e3abf00 rename model 2026-04-20 22:02:37 +08:00
Robert Vollmer 27d6a3d503 [google] Fix cache_read cost in gemini-2.5-flash model
https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-flash

There's no cache_read_audio, is there?
2026-04-20 15:02:54 +02:00
Lyda 6beb1f8be2 feat(302ai): standardize Claude model metadata and capabilities
- Add family field for all Claude models (claude-haiku, claude-opus, claude-sonnet)
- Standardize knowledge cutoff dates to full date format (YYYY-MM-DD)
- Enable reasoning capability for Claude Opus 4.x and Sonnet 4.x series models
- Add PDF input modality support for claude-opus-4-1-20250805
- Update claude-opus-4-7 context limit to 1,000,000 tokens
2026-04-20 16:10:52 +08:00
Lyda e0c1124fe2 feat(302ai): update GPT model capabilities and specifications
- Add structured_output capability for GPT-4.1, GPT-4o, and GPT-5 series models
- Enable reasoning capability and disable temperature for GPT-5 series models
- Update context limits: GPT-4.1 series to 1,047,576 tokens, GPT-5.4 series to 1,050,000 tokens
- Add input token limits for GPT-5 series models (272,000 or 922,000 tokens)
- Update knowledge cutoffs across GPT-5 series (2024-05-30 to 2025-08-31)
- Add PDF input modality support for GPT-4
2026-04-20 15:58:32 +08:00
Lyda bf4ceb7baa feat(302ai): update GLM model capabilities and knowledge cutoffs
- Enable reasoning capability for GLM-4.5-air, GLM-4.5, GLM-4.5V, and GLM-4.6V models
- Update knowledge cutoff to 2025-04 for GLM-4.5, GLM-4.5V, GLM-4.6, GLM-4.6V, and GLM-4.7
- Add video input modality support for GLM-4.5V and GLM-4.6V
- Add structured_output capability for GLM-5-turbo and GLM-5.1
- Add interleaved reasoning_content field for GLM-4.7, GLM-5, GLM-5-turbo, GLM-5.1, and GLM-5V-turbo
2026-04-20 15:48:29 +08:00
Aiden Cline 1a41934e55 Merge pull request #1448 from Lydanne/dev
feat(302ai): supplement commonly missing models
2026-04-19 22:19:18 -05:00
Aiden Cline aeb4caec9f Merge pull request #1488 from Sewer56/add-wafer-provider
Add wafer.ai provider
2026-04-19 22:19:06 -05:00
Aiden Cline ccb8dcc65f Merge pull request #1493 from rovo89/patch-1
Add context_over_200k for gemini-2.5-pro and adjust cache_read costs
2026-04-19 22:17:45 -05:00
Lyda 435ec1df7b feat(302ai): add claude-opus-4-7 model 2026-04-20 11:14:27 +08:00
Robert Vollmer 7c9d609143 Add context_over_200k for gemini-2.5-pro and adjust cache_read costs
https://ai.google.dev/gemini-api/docs/pricing#gemini-2.5-pro
https://cloud.google.com/vertex-ai/generative-ai/pricing#gemini-models-2.5 (rounds 0.125 to 0.13)
2026-04-20 00:09:11 +02:00
Aiden Cline add7947164 Merge pull request #1492 from dpuyosa/feat/venice-add-gemma4-uncensored
Venice: Add Gemma 4 and Venice Uncensored 1.2 models
2026-04-19 16:52:28 -05:00
Aiden Cline dd0c1af12a Merge pull request #1491 from dpuyosa/chore/venice-pricing-update
Venice: Update pricing for Grok 4.20 and Qwen3.5 9B
2026-04-19 16:52:15 -05:00
Aiden Cline 0c93cc03be Merge pull request #1485 from anomalyco/more-extends-cases
migrate more providers to extends format
2026-04-19 16:51:58 -05:00
Aiden Cline 62e25b73f8 Merge branch 'dev' into more-extends-cases 2026-04-19 16:47:05 -05:00
Aiden Cline 17093e0031 Merge pull request #1489 from berget-ai/update/berget-prices-gemma4
chore: update berget.ai models - prices and Gemma 4
2026-04-19 16:42:06 -05:00
Aiden Cline ae542978f0 Merge pull request #1490 from BlockListed/cortecs-add-claude-opus-4-7
add claude opus 4.7
2026-04-19 16:41:36 -05:00
dpuyosa 7f16117bba [venice] Add Gemma 4 and Venice Uncensored 1.2 models
- Add Gemma 4 Uncensored with 256K context, image support
- Add Venice Uncensored 1.2 with 128K context, image support
- Both models support tool calls and structured output
2026-04-19 23:41:17 +02:00
dpuyosa 2b96a2d3d6 [venice-models] Update pricing for Grok 4.20 and Qwen3.5 9B
- Update cache_read pricing for Grok 4.20 context_over_200k (0.23 → 0.45)
- Update input cost for Qwen3.5 9B (0.05 → 0.1)
2026-04-19 23:38:20 +02:00
BlockListed 9799a841c6 add claude opus 4.7
yes this model id is correct, cortecs is weird.
2026-04-19 23:13:58 +02:00
Christian Landgren 71c59b4235 chore: update berget.ai models - prices and Gemma 4
- Add Google Gemma 4 31B Instruct model
- Update prices for existing models (EUR to USD conversion)
- Remove non-coding models (bge-reranker, multilingual-e5 embeddings, kb-whisper)
- Remove deprecated Llama-3.1-8B-Instruct

Updated models:
- GLM-4.7: 0.77/2.75 USD/M (was 0.7/2.3)
- Llama-3.3-70B: 0.99/0.99 USD/M (was 0.9/0.9)
- Mistral-Small-3.2: 0.33/0.33 USD/M (was 0.3/0.3)
- GPT-OSS-120B: 0.44/0.99 USD/M (was 0.3/0.9)

New models:
- Gemma-4-31B-it: 0.275/0.55 USD/M

Removed models (not relevant for coding):
- BAAI/bge-reranker-v2-m3 (reranker)
- intfloat/multilingual-e5-large/* (embeddings)
- KBLab/kb-whisper-large (speech-to-text)
- meta-llama/Llama-3.1-8B-Instruct (deprecated)
2026-04-19 12:19:32 +02:00
Sewer56 0568b412aa Add wafer.ai provider with GLM-5.1 and Qwen3.5-397B-A17B models 2026-04-19 03:02:04 +01:00
Aiden Cline 812612465a Merge pull request #1484 from smakosh/feat/llmgateway-new-models
feat: update LLM Gateway to 182 models
2026-04-18 18:37:26 -05:00
Aiden Cline ac8a79dd74 Merge pull request #1487 from WJQSERVER/add/nvidia(nim)-z-ai-glm-5.1
Add Z.AI GLM-5.1 to NVIDIA(NIM)
2026-04-18 18:36:55 -05:00
smakosh 9460d981d4 fix: replace broken extends with concrete glm-4.6v-flash def
zhipuai/glm-4.6v-flash.toml is a symlink to zai/models/glm-4.6v-flash.toml
which does not exist, causing validate to fail with 'Unable to resolve
extends.from'. Inline the concrete definition instead.
2026-04-18 13:36:52 +02:00
smakosh 91db9d81ec feat: update LLM Gateway to 182 models
- Uses extends to reference canonical providers where
  possible (116 models), keeping 66 full definitions
- Only includes active text/chat models
- Adds new models: Claude Opus 4.7, Grok 4 Fast,
  Kimi K2, Mimo V2, GLM 5.1, Qwen 3 Coder, and more

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-18 13:20:41 +02:00
Jack d8d5f08df6 Merge pull request #1486 from chl-0537/feature/add-tencent
feat: add new provider and model
2026-04-18 17:30:53 +08:00
WJQSERVER 11d5443ab1 follow the nim modelcard change context length to 131072
https://build.nvidia.com/z-ai/glm-5.1/modelcard
Other Properties Related to Input: Supports multi-turn conversations, tool calling, system prompts, and extended agentic sessions. Input context length: 131,072 tokens.
2026-04-18 16:48:29 +08:00
wjqserver 90558e9eed add glm-5.1 2026-04-18 16:41:02 +08:00
Aiden Cline cb7d258e33 migrate more providers to extends format 2026-04-17 23:02:34 -05:00
Frank 2af43dc4f8 update zen models 2026-04-17 19:08:05 -04:00
Aiden Cline 93ddb6b131 Merge pull request #1482 from sopial42/ovhcloud/update-models-clean
chore(ovhcloud): remove 3 models no longer available in AI Endpoints
2026-04-17 16:50:53 -05:00
Aiden Cline 04bf671f18 Merge pull request #1481 from Spherrrical/add-digitalocean-provider
feat(provider): add DigitalOcean provider
2026-04-17 16:49:44 -05:00
Aiden Cline 1f3ba4ba21 Merge pull request #1483 from anomalyco/add-extends-support
feat: add extends support
2026-04-17 16:49:23 -05:00
Aiden Cline 305bdb6cdc Merge branch 'dev' into add-extends-support 2026-04-17 16:22:41 -05:00
Aiden Cline 2e5b4b44ae update some modes 2026-04-17 16:22:14 -05:00
aadhondt bb1c08dd41 chore(ovhcloud): remove 3 models no longer available in AI Endpoints
- deepseek-r1-distill-llama-70b
- mixtral-8x7b-instruct-v0.1
- qwen2.5-coder-32b-instruct

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-17 22:38:35 +02:00
Spherrrical 86d0f14397 feat(digitalocean): add DigitalOcean Gradient AI Platform provider
Adds the DigitalOcean provider with 46 models (Anthropic, OpenAI,
Arcee, fal, and DO-hosted open-source/embedding models) served via
the OpenAI-compatible endpoint at https://inference.do-ai.run/v1.
2026-04-17 13:07:47 -07:00
Aiden Cline f93c1a8998 Merge pull request #1478 from Kaspazza/dev
Add Github Copilot Claude Opus 4.7
2026-04-17 15:01:04 -05:00
Aiden Cline 27a6758eb7 new gen script 2026-04-17 14:57:57 -05:00
Aiden Cline 9dbafb81fa add script 2026-04-17 14:57:41 -05:00
Aiden Cline 419f8a3a20 add migration checker script 2026-04-17 14:57:32 -05:00
Aiden Cline 785a091073 Merge pull request #1480 from nicocasaisd/openai/remove-deprecated-codex-mini-latest
chore(openai): remove deprecated model codex-mini-latest
2026-04-17 14:52:34 -05:00
nicocasaisd ad2409dce1 chore(openai): remove deprecated model codex-mini-latest 2026-04-17 15:35:43 -03:00
kaspazza 63450ab67d Add Github Copilot Claude Opus 4.7 2026-04-17 20:19:36 +02:00
Aiden Cline 4002cb6739 model 2026-04-17 13:11:51 -05:00
Aiden Cline 72fa27a91e Merge pull request #1474 from vglafirov/add-gitlab-duo-chat-opus-4-7
feat(gitlab): add duo-chat-opus-4-7 model definition
2026-04-17 12:37:36 -05:00
Aiden Cline ddb3a0ff05 update agents.md 2026-04-17 12:14:04 -05:00
Aiden Cline 96c12042bc remeda 2026-04-17 12:13:54 -05:00
Misha Skvortsov 7336b3619c atomic-chat: add provider with initial blessed models
Adds Atomic Chat as a local OpenAI-compatible provider at
http://127.0.0.1:1337/v1. Includes logo and three curated models:

- unsloth/Qwen3.5-9B-IQ4_XS  (id: Qwen3_5-9B-IQ4_XS)
- unsloth/gemma-4-E4B-it-IQ4_XS  (id: gemma-4-E4B-it-IQ4_XS)
- unsloth/MiniMax-M2.5-UD-TQ1_0  (id: MiniMax-M2_5-UD-TQ1_0)

Model ids match the normalized form returned by Atomic Chat's
/v1/models endpoint (dots replaced with underscores).

Made-with: Cursor
2026-04-17 13:03:53 +03:00
mickalchen 73b81ac027 add tencent provider 2026-04-17 17:27:35 +08:00
Vladimir Glafirov 9c5839a414 feat(gitlab): add duo-chat-opus-4-7 model definition 2026-04-17 08:54:07 +02:00
Aiden Cline 721464bc3c Merge pull request #1469 from GrahamCampbell/ops-4-7-fixes
Corrected and normalized claude opus 4.7 knowledge cut-off dates
2026-04-16 22:30:13 -05:00
Aiden Cline b6b45a9d25 Merge pull request #1463 from GrahamCampbell/claude-4-6
Correct Anthropic Claude 4.6 model knowledge cut-off dates
2026-04-16 21:36:19 -05:00
Aiden Cline 7b8f98bb23 Merge pull request #1471 from cfbender/fix/openrouter-opus-4-7
feat: add openrouter opus 4.7
2026-04-16 21:35:52 -05:00
Aiden Cline 92ac48b07f Merge pull request #1473 from fhennerkes/dev
Poe: add Claude-Opus-4.7
2026-04-16 20:57:21 -05:00
fhennerkes 36a455ce4a Merge branch 'anomalyco:dev' into dev 2026-04-16 18:11:32 -07:00
fhennerkes 0bf5c60319 poe: add Claude-Opus-4.7 model
Add new Anthropic model from Poe API (released 2026-04-15):
- Reasoning support
- 1M context window with 128K output
- Cost: $4.3/M input, $21/M output, $0.43/M cache read, $5.4/M cache write
- Modalities: text, image, pdf
2026-04-16 18:08:02 -07:00
Kit Langton b123711494 Merge pull request #1472 from elithrar/patch-4
cloudflare: add opus 4.7
2026-04-16 19:38:46 -04:00
Matt Silverlock 832064c1d1 cloudflare: add opus 4.7 2026-04-16 18:51:54 -04:00
Cody Bender c16e1c817b fix: add openrouter opus 4.7 2026-04-16 18:38:36 -04:00
Aiden Cline 8aaf31711b Merge pull request #1468 from heimoshuiyu/fix/opus-4-7-temperature
fix: set temperature=false for Claude Opus 4.7
2026-04-16 14:34:52 -05:00
Graham Campbell ca6acf0b3e Corrected and normalized claude opus 4.7 knowledge cut-off dates 2026-04-16 20:33:49 +01:00
heimoshuiyu 660a672647 fix: set temperature=false for firmware and venice Opus 4.7 2026-04-17 03:23:11 +08:00
heimoshuiyu 147cb3138a fix: set temperature=false for Claude Opus 4.7 across all providers 2026-04-17 03:22:33 +08:00
Aiden Cline 5c1fb729fd Merge pull request #1462 from dpuyosa/dev
Venice: Add Claude Opus 4.7 and remove deprecated models
2026-04-16 14:10:03 -05:00
Aiden Cline 0d34900078 Merge pull request #1464 from cgilly2fast/dev
feat(firmware): opus 4.7 remove old claude models
2026-04-16 14:09:39 -05:00
Aiden Cline 4bc71918f5 Merge pull request #1467 from vercel/update-vercel-models-20260416-1812
Update Vercel models
2026-04-16 13:34:40 -05:00
Jerilyn Zheng 6f62522505 Set temperature to false in claude-opus-4.7 configuration
Changed temperature setting from true to false.
2026-04-16 11:23:29 -07:00
Aiden Cline 7680a1d169 Merge pull request #1466 from vercel/fix-reranking-type-upstream
fix(vercel): accept reranking model type from API
2026-04-16 13:22:29 -05:00
github-actions[bot] bdc15a57ce chore(vercel): update Vercel model definitions
Auto-generated by weekly workflow from Vercel AI Gateway API.

Co-Authored-By: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-16 18:12:52 +00:00
R-Taneja c7d324fed8 fix(vercel): accept reranking model type from API
The Vercel AI Gateway API now returns models with type "reranking",
which caused the generate-vercel script to fail schema validation.

Add "reranking" to the ModelType enum and skip these models in the
main loop, matching the existing pattern for image/video types that
OpenCode does not consume.
2026-04-16 11:05:26 -07:00
Colby Gilbert 863bada2cf feat(firmware): opus 4.7 remove old claude models 2026-04-16 10:37:52 -07:00
dpuyosa 28c4af0631 [venice] Add Claude Opus 4.7 and update Qwen models
- Add Claude Opus 4.7 with 1M context, 128K output, multimodal support
- Update Qwen 3.5 35B and 397B to open_weights=true and refresh last_updated
- Remove deprecated models: Grok Code Fast 1, Mercury Edit 2, MiniMax M2.1
2026-04-16 18:30:45 +02:00
Aiden Cline 38f6b7dfc5 Merge pull request #1456 from Snat3r/dev
Add GML5.1.toml model to cortecs provider
2026-04-16 11:16:51 -05:00
Aiden Cline 1ba010f66e Merge pull request #1461 from llc1123/chore/zenmux-update
chore(zenmux): add claude-opus-4.7
2026-04-16 11:16:41 -05:00
粒粒橙 132a5ade8d chore(zenmux): add claude-opus-4.7 2026-04-17 00:02:58 +08:00
Graham Campbell 2ef2b10847 Correct anthropic 4.6 knowledge cut-off dates 2026-04-16 16:52:15 +01:00
Frank 87e1dcb70f update zen modles 2026-04-16 11:31:20 -04:00
Aiden Cline 43d2e058d9 Merge pull request #1449 from shikbupt/alibaba-glm5.1
add alibaba-cn glm5.1
2026-04-16 10:25:58 -05:00
Aiden Cline 0a6c2ca49d Merge pull request #1459 from itsnebulalol/dev
feat: add anthropic claude opus 4.7 models
2026-04-16 10:25:04 -05:00
Dominic Frye 2f954fc838 feat: add anthropic claude opus 4.7 models 2026-04-16 11:19:19 -04:00
Frank 91b7851971 update zen models 2026-04-16 04:51:30 -04:00
Jack 80e23ab90e Merge pull request #1266 from lioZ129/feature/add-hpc-ai-provider
feat: add HPC-AI model provider support
2026-04-16 15:09:44 +08:00
Snat3r 6e574da2e4 Update input modalities in minimax-M2.7.toml 2026-04-16 08:21:20 +02:00
Snat3r 5b68565783 Add MiniMax-M2.7 model configuration file cortecs 2026-04-16 08:19:51 +02:00
lioZ129 65dafb56a3 add [cost] and glm5.1 support 2026-04-16 14:13:27 +08:00
Snat3r 32c0c88600 Update context and output limits in glm-5.1.toml 2026-04-16 08:08:09 +02:00
Snat3r d911b6f610 Add GLM-5.1 model configuration file 2026-04-16 08:06:37 +02:00
Aiden Cline 5cf28a566f Merge pull request #1424 from WJQSERVER/feat/nvidia-minimax-m2.7
Add MiniMax M2.7 to NVIDIA(NIM)
2026-04-15 20:19:01 -05:00
Aiden Cline 52cdb783a1 Merge pull request #1452 from wwth8819/dev
For aihubmix add GPT-5.4 \ GPT-5.4-mini, remove Incorrect value from old models, update cost
2026-04-15 20:18:24 -05:00
wwth8819 4a14ce5ae6 Remove temperature setting from gpt-5.2-codex.toml
Removed the temperature setting from the configuration.
2026-04-16 03:12:14 +08:00
wwth8819 f17c352027 Update cost values in coding-glm-4.7.toml 2026-04-16 03:11:16 +08:00
wwth8819 46ec19ab02 add gpt-5.4-mini gpt-5.4 2026-04-16 03:09:26 +08:00
Frank f12aae094e update zen models 2026-04-15 10:55:02 -04:00
sk 377d0f1c8d add alibaba-cn glm5.1 2026-04-15 22:02:52 +08:00
WJQSERVER 884b799012 Merge branch 'dev' into feat/nvidia-minimax-m2.7 2026-04-15 21:53:54 +08:00
Lyda ac7e35af4e feat(302ai): supplement commonly missing models 2026-04-15 17:32:18 +08:00
Frank 6f04d267cf update zen models 2026-04-15 02:17:31 -04:00
Aiden Cline a0b89e739b Merge pull request #1425 from ceyhanmolla/add-minimax-m2.7-nvidia
feat(nvidia): add MiniMax-M2.7
2026-04-14 22:59:38 -05:00
Frank 0ce000a521 update go models 2026-04-14 23:09:06 -04:00
Frank 4b7dda6cc6 update go models 2026-04-14 22:50:12 -04:00
Aiden Cline 7220310828 Merge pull request #1447 from Sawyerb/dev
Removed deprecated models and added ME2 to all relevant providers.
2026-04-14 21:49:22 -05:00
Aiden Cline 15746b9845 Merge pull request #1444 from wwth8819/dev
Add glm-5.1,  coding-glm-5.1  TO  AiHubMix
2026-04-14 21:49:06 -05:00
Aiden Cline e9be42b4bc Merge pull request #1443 from teodortomas/add-minimax-m2.7
add minimax-m2p7 to fireworks-ai provider
2026-04-14 17:11:09 -05:00
Aiden Cline c0d21d802f Merge pull request #1429 from Lee-Si-Yoon/fix/cache-read-friendli
fix: cache read costs for friendliAI models
2026-04-14 17:10:56 -05:00
Aiden Cline 4937952a52 Merge pull request #1437 from Ardakilic/dev
feat: kilo gateway: elephant alpha
2026-04-14 17:10:37 -05:00
Aiden Cline 1bb5deadea Merge pull request #1439 from fhennerkes/dev
poe: update models with pricing, deprecations, and display name fixes
2026-04-14 17:10:26 -05:00
Aiden Cline c2ad18c87d Merge pull request #1445 from cantalupo555/chore/openrouter-remove-deprecated-free-models
chore(openrouter): remove deprecated free-tier models no longer available via API
2026-04-14 17:10:01 -05:00
Sawyer e4b0a53e26 Removed deprecated models and added ME2 to all relevant providers. 2026-04-14 12:33:35 -07:00
cantalupo555 61ea2e2093 chore(openrouter): remove deprecated free-tier models no longer available via API 2026-04-14 08:11:45 -03:00
wwth8819 162fd8b72b Add configuration for Coding-GLM-5.1 model 2026-04-14 17:35:23 +08:00
wwth8819 b8fae048db Add GLM-5.1 model configuration file 2026-04-14 17:32:59 +08:00
Teodor Tomáš ca6cf794c1 add minimax-m2p7 to fireworks-ai provider 2026-04-14 11:06:47 +02:00
fhennerkes 152b401976 poe: update models with pricing, deprecations, and display name fixes
Mark 11 models no longer on the Poe API as deprecated:
- anthropic: claude-sonnet-3.5, claude-sonnet-3.5-june
- cerebras: qwen3-235b-2507-cs, qwen3-32b-cs, llama-3.3-70b-cs
- google: gemini-3-pro, gemini-deep-research
- novita: glm-4.7
- openai: chatgpt-4o-latest, gpt-4-classic-0314, gpt-4-classic

Update existing models with latest API data:
- Cerebras (gpt-oss-120b-cs, llama-3.1-8b-cs): Add pricing and correct context (128K)
- kimi-k2.5: Add pricing, fix display name, temperature/reasoning from API
- gpt-4o: Fix output formatting (8_192)
- gpt-5.1-codex-max: Fix display name

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-13 16:55:18 -07:00
Nacho F. Lizaur 8f340e1eb2 feat: update kiro provider to use kiro-ai-provider npm package 2026-04-13 23:12:11 +02:00
Aiden Cline d99a1581df Merge pull request #1435 from cantalupo555/feat/add-openrouter-elephant-alpha
feat(openrouter): add Elephant Alpha model
2026-04-13 15:01:55 -05:00
Nacho F. Lizaur 62da5b0cbd feat: enable reasoning on Kiro Claude models 2026-04-13 19:49:25 +02:00
Nacho F. Lizaur 4e49abc10c feat: add Kiro provider 2026-04-13 19:49:25 +02:00
Arda Kilicdagi 21cf1fb1e9 feat: kilo gateway: elephant alpha 2026-04-13 20:39:54 +03:00
cantalupo555 970a00c3ad feat(openrouter): add Elephant Alpha model 2026-04-13 13:43:56 -03:00
cantalupo555 d611fc2ef7 feat(core): add elephant model family 2026-04-13 13:35:10 -03:00
Aiden Cline 7d7711878d Merge pull request #1434 from cantalupo555/chore/openrouter-step-3.5-flash-free
chore(openrouter): remove Step 3.5 Flash free-tier model
2026-04-13 10:45:25 -05:00
cantalupo555 fc5d5ec67f chore(openrouter): remove deprecated step-3.5-flash free-tier variant
- Remove the free-tier variant of Step 3.5 Flash as it is no longer needed
2026-04-13 11:51:10 -03:00
Aiden Cline bd5050a15f Merge pull request #1423 from spiffytech/dev
Add Synthetic support for GLM-5.1
2026-04-13 09:18:49 -05:00
Aiden Cline ac67c648a2 Merge pull request #1422 from hanouticelina/add-minimax-2.7-huggingface
feat(huggingface): add MiniMax-M2.7
2026-04-13 09:18:35 -05:00
Aiden Cline f58038fc21 Merge pull request #1421 from zainhas/dev
[Together AI]add minimax m2.7
2026-04-13 09:18:14 -05:00
Aiden Cline 3ea7d56b96 Merge pull request #1428 from dpuyosa/feat/venice-glm-5-models
Venice: Add Z-AI GLM-5 Turbo and GLM-5V Turbo models
2026-04-13 09:17:45 -05:00
Aiden Cline e23657368e Merge pull request #1426 from dpuyosa/venice/update-model-configs-0412
Venice: Update model configs and pricing
2026-04-13 09:17:33 -05:00
Aiden Cline d59ad3ffa3 Merge pull request #1427 from dpuyosa/refactor/venice-model-naming
Venice: Update model naming convention
2026-04-13 09:17:08 -05:00
siyoon dcbb415960 fix: add cache_read parameter to cost section 2026-04-13 16:05:36 +09:00
dpuyosa e302956dc4 [venice] Update model naming convention
- Rename files to use dashes instead of dots
- Rename opus/sonnet 45 to 4-5 format
- Remove beta suffix from Grok 4.20 models
- Update family from grok-beta to grok
- Remove knowledge field from Claude models
- Update output limits
2026-04-13 01:10:06 +02:00
dpuyosa 5d100a8243 [venice] Add Z-AI GLM-5 Turbo and GLM-5V Turbo models
- Add Z-AI GLM-5 Turbo (text-only, reasoning, tool_call)

- Add Z-AI GLM-5V Turbo (vision, reasoning, tool_call)
2026-04-13 01:05:47 +02:00
dpuyosa 9ff52ee539 [venice] Update model configs and pricing
- Update last_updated dates for 5 models

- Set open_weights to false for 4 models

- Add context_over_200k cache pricing for qwen-3-6-plus
2026-04-13 00:59:21 +02:00
ceyhanmolla 8114522671 feat(nvidia): add MiniMax-M2.7 2026-04-12 16:56:04 +02:00
wjqserver 5e534d84bc feat: add NVIDIA MiniMax M2.7 model 2026-04-12 22:51:46 +08:00
spiffytech 1f79cd8fd0 Added Synthetic support for GLM-5.1 2026-04-12 08:46:10 -04:00
Celina Hanouti 462b822e60 add minimax M2.7 for hugging face provider 2026-04-12 11:06:30 +02:00
Zain Hasan 8021e1afd0 Merge branch 'anomalyco:dev' into dev 2026-04-11 22:23:20 -07:00
Zain Hasan 304d85fa43 [Together AI]add minimax m2.7 2026-04-11 22:14:05 -07:00
Aiden Cline f07262370f Merge pull request #1418 from Ardakilic/dev
Kilo Gateway: Sync model list with upstream (2026-04-11)
2026-04-11 16:48:46 -05:00
Aiden Cline 7cfaa0393e Merge pull request #1419 from nicopujia/add-openrouter-deepseek-r1
Add OpenRouter support for DeepSeek R1
2026-04-11 16:46:58 -05:00
Arda Kilicdagi c93de9f54f chore: sync all kilo gw models 2026-04-11 03:31:00 +03:00
Aiden Cline 2b17d9efc2 Merge pull request #1414 from Ardakilic/dev
Feat: Kilo Gateway: Add MiniMax M2.7
2026-04-10 13:31:50 -05:00
Arda Kilicdagi 8d8521e0e1 feat: kilo gateway: MiniMax M2.7 2026-04-10 20:42:45 +03:00
Aiden Cline c9f22d26b9 Merge pull request #1413 from Ardakilic/dev
Kilo Gateway: Add GLM 5.1, Remove MiniMax M2.5 Free
2026-04-10 10:33:18 -05:00
Aiden Cline 61e8ea2a2e Merge pull request #1410 from gjtiquia/dev
Poe: add GLM-5 model
2026-04-10 10:26:57 -05:00
Aiden Cline 7d2cd9818f Merge pull request #1411 from Alex-wuhu/dev
novita-ai: add 8 new models and remove 2 deprecated models
2026-04-10 10:26:45 -05:00
Arda Kilicdagi 385b4fbcf6 feat: kilo gateway: glm-5.1 2026-04-10 15:11:37 +03:00
Alex-wuhu 2f0890c9a9 Add new model configurations for Gemma, MiniMax, Qwen, and GLM 2026-04-10 16:28:34 +08:00
GJ Tiquia 9d377b768f poe: GLM-5 temperature set to true 2026-04-10 15:57:25 +08:00
GJ Tiquia aa4f3f550a poe: add GLM-5 model 2026-04-10 14:13:03 +08:00
Aiden Cline f82d6fc61a Merge pull request #1404 from nicopujia/add-openrouter-qwen3.5-flash-02-23
Add OpenRouter support for Qwen3.5 Flash 2026-02-23
2026-04-09 22:37:05 -05:00
Aiden Cline 4222b040b7 Merge pull request #1407 from line72/deepinfra-glm-5.1
[DeepInfra] Add GLM 5.1
2026-04-09 22:36:53 -05:00
Aiden Cline d2e4174103 Merge pull request #1397 from qychen2001/dev
Add GLM-5.1 and GLM-5V-Turbo model configurations for siliconflow
2026-04-09 22:35:52 -05:00
Aiden Cline f30b5fc754 Merge pull request #1408 from nanai10a/dev
Add MiniMax M2.5 (free) configuration file
2026-04-09 22:35:37 -05:00
Aiden Cline 10239c95e2 Merge pull request #1401 from dpuyosa/venice-open-weights-fix
Venice: Remove private field fallback for open weights
2026-04-09 20:05:03 -05:00
Aiden Cline 59831a9a0e Merge pull request #1399 from dpuyosa/venice-model-updates-2026-04-09
Venice: Update model configs with refreshed pricing and limits
2026-04-09 20:04:55 -05:00
Aiden Cline c7552d0e00 Merge pull request #1396 from shelvick/add-azure-grok-4-20
Add Grok 4.20 reasoning and non-reasoning to Azure
2026-04-09 20:04:04 -05:00
Aiden Cline 76d53c9e96 Merge pull request #1382 from cgilly2fast/dev
feat(firmware): add zai 5.1 and qwen 3.6 plus
2026-04-09 20:03:50 -05:00
Aiden Cline 61573f666a Merge pull request #1395 from zainhas/dev
[Together AI] add GLM-5.1 + Gemma 4 31B it
2026-04-09 20:03:13 -05:00
Aiden Cline 1d30640400 Merge branch 'dev' into dev 2026-04-09 20:02:59 -05:00
Aiden Cline 1f1eafe173 Merge pull request #1406 from riccardogiorato/dev
Update GLM to version 5.1 for together provider
2026-04-09 20:02:08 -05:00
Aiden Cline a75c0f0fe9 Merge pull request #1400 from dpuyosa/venice-add-new-models
Venice: Add 4 new AI models
2026-04-09 17:34:55 -05:00
Marcus Dillavou 57c6d817d5 DeepInfra: Add GLM 5.1 2026-04-09 15:50:04 -05:00
Riccardo Giorato 3d1d77da76 Merge pull request #2 from riccardogiorato/orchestrator/add-glm-5-1-together-r8k9f
add GLM-5.1 for together provider
2026-04-09 22:42:54 +02:00
orchestrator-build[bot] e7cfed72ff fix GLM-5.1 open_weights to true 2026-04-09 20:40:20 +00:00
orchestrator-build[bot] b748364a20 fix GLM-5.1 pricing for together provider 2026-04-09 20:39:49 +00:00
orchestrator-build[bot] 30439f4409 replace GLM-5 with GLM-5.1 for together provider 2026-04-09 20:37:34 +00:00
orchestrator-build[bot] 7baad3cc22 add GLM-5.1 for together provider 2026-04-09 20:35:51 +00:00
Nicolás Pujia 564992885b Add OpenRouter support for DeepSeek R1 2026-04-09 11:54:16 -03:00
Nicolás Pujia 57a53889db Add OpenRouter support for Qwen3.5 Flash 2026-02-23 2026-04-09 11:53:53 -03:00
Nanai Jua 0f019a5e9a Add MiniMax M2.5 (free) configuration file
https://openrouter.ai/provider/open-inference
2026-04-09 18:11:40 +09:00
dpuyosa d6fec11252 [venice] Remove private field fallback for open weights
- Rely solely on modelSource for open weights detection
- Remove privacy field fallback per new ZDR policies
2026-04-09 11:03:03 +02:00
dpuyosa 7604313114 [venice] Add 4 new AI models
- Add Mercury 2 (reasoning model)
- Add Mistral Small 4 (multimodal)
- Add Nemotron Cascade 2 30B A3B
- Add Qwen 3.5 397B (multimodal)
2026-04-09 10:31:24 +02:00
dpuyosa c3a2b1a1db [venice] Update model configs with refreshed pricing and limits
- Update last_updated dates to 2026-04-09 across 6 models
- Adjust Grok pricing to reflect current rates
- Add context_over_200k pricing for Qwen 3.6 Plus
- Correct Gemma 4 output limits from 12288 to 8192
- Rename Qwen 3.6 Plus to "Uncensored" variant
2026-04-09 10:17:05 +02:00
QiyuanChen 68294f3bd6 Add GLM-5V-Turbo model configuration for siliconflow provider 2026-04-09 11:18:25 +08:00
QiyuanChen ad10ce6232 Add GLM-5.1 model configuration files for siliconflow provider 2026-04-09 11:08:53 +08:00
Scott Helvick 5a09420d65 Add Grok 4.20 reasoning and non-reasoning to Azure 2026-04-09 02:30:55 +00:00
Zain Hasan f1b3177ff7 add gemma 4 31b instruct 2026-04-08 18:49:38 -07:00
Zain Hasan a7d0152fd1 [Together AI] add GLM-5.1 2026-04-08 17:55:11 -07:00
Aiden Cline 46c6aef51b Merge pull request #1393 from spiffytech/dev
Add Synthetic support for GLM-5 and Nemotron 3 Super
2026-04-08 16:01:12 -05:00
Aiden Cline 7c34bf01b5 Merge pull request #1394 from spiffytech/ollama-changes
Add Ollama Cloud support for Gemma 4. Updated properties on Gemini 3 Flash
2026-04-08 16:00:12 -05:00
spiffytech 5ce20c0978 Added Ollama Cloud support for Gemma 4. Updated properties on Gemini 3 Flash. 2026-04-08 15:03:29 -04:00
spiffytech 6055551b33 Added Synthetic support for GLM-5 and Nemotron 3 Super 2026-04-08 14:57:51 -04:00
Aiden Cline a96094c059 Merge pull request #1390 from GoGoris/add-qwen3-coder-next-cortecs
Add qwen3-coder-next model for cortecs
2026-04-08 11:26:42 -05:00
Aiden Cline 39e86eb055 Merge pull request #1384 from dpuyosa/feat/add-venice-claude-opus-4-6-fast-glm-5-1
Venice: Add Claude Opus 4.6 Fast and GLM 5.1 models
2026-04-08 11:25:56 -05:00
Aiden Cline 119f421437 Merge pull request #1387 from cantalupo555/feat/add-gemma-4-free-openrouter
feat(openrouter): add Gemma 4 free models
2026-04-08 11:25:30 -05:00
Aiden Cline 6ab4c1be04 Merge pull request #1386 from cantalupo555/chore/remove-qwen3.6-plus-free-openrouter
chore(openrouter): remove discontinued Qwen3.6 Plus free
2026-04-08 11:25:15 -05:00
Aiden Cline f1eaa4bd9d Merge pull request #1392 from Solidsilver/feat/fireworks-add-glm-5-1-qwen-3-6-plus
feat(fireworks): add GLM 5.1 and Qwen 3.6 Plus models
2026-04-08 11:24:52 -05:00
Aiden Cline 2f4693b9cf Merge pull request #1391 from CassiusXiang/fix/openrouter-qwen3.6-plus
fix(openrouter): replace discontinued qwen3.6-plus free with paid model
2026-04-08 11:24:41 -05:00
Luke M 5c867d5fba feat(fireworks): add GLM 5.1 and Qwen 3.6 Plus models 2026-04-08 09:06:50 -07:00
XiangChang 682a486f09 fix(openrouter): replace discontinued qwen3.6-plus free with paid model 2026-04-08 22:31:52 +08:00
Steven Goris ea10c5674b Add qwen3-coder-next model for cortecs 2026-04-08 15:14:31 +02:00
cantalupo555 5829a9174c feat(openrouter): add Gemma 4 26B A4B free
- Add google/gemma-4-26b-a4b-it:free (MoE, 256K context, multimodal, reasoning)
2026-04-08 08:32:07 -03:00
cantalupo555 4717e5a902 feat(openrouter): add Gemma 4 31B free
- Add google/gemma-4-31b-it:free (256K context, multimodal, reasoning)
2026-04-08 08:31:55 -03:00
cantalupo555 20f225d6a6 chore(openrouter): remove discontinued Qwen3.6 Plus free
- Model no longer available on OpenRouter API
2026-04-08 08:16:47 -03:00
dpuyosa 04ada2c06e [venice] Add Claude Opus 4.6 Fast and GLM 5.1 models
- Add claude-opus-4-6-fast model with 1M context
- Add zai-org-glm-5-1 model with reasoning and tool_call
2026-04-08 11:40:46 +02:00
Frank 23fd440f57 update zen models 2026-04-08 02:20:44 -04:00
Colby Gilbert a4c09d58c3 feat(firmware): add zai 5.1 and qwen 3.6 plus 2026-04-07 22:18:54 -07:00
Aiden Cline 82924aa6f2 Merge pull request #1376 from mugnimaestra/feat/add-glm-5.1-tee-chutes
feat(chutes): add zai-org/GLM-5.1-TEE model
2026-04-07 23:55:23 -05:00
Aiden Cline c8f0b6d573 Merge pull request #1377 from mchenco/mchen/update-cf-workers-ai-models
update cloudflare-workers-ai: add gemma-4, remove non-LLMs, fix metadata
2026-04-07 23:55:07 -05:00
Aiden Cline 09c3cb3e0a Merge pull request #1379 from zhongruan0522/dev
add GLM-5.1 to zhipuai and zai providers
2026-04-07 23:54:32 -05:00
Aiden Cline 555662ec80 Merge pull request #1381 from llc1123/chore/zenmux-update
chore(zenmux): add GLM-5.1 model configuration
2026-04-07 23:54:20 -05:00
Aiden Cline fc63cc19c3 feat: add experimental modes to models to express things like "fast" that induce additional price changes 2026-04-07 23:47:19 -05:00
Aiden Cline 26d2f3e8e9 Merge pull request #1378 from friendliai/minpeter/add-friendli-glm-5.1
feat(friendli): add GLM-5.1 and remove deprecated models
2026-04-07 22:30:44 -05:00
粒粒橙 92374dd74d chore(zenmux): add GLM-5.1 model configuration 2026-04-08 11:30:14 +08:00
minpeter ee4de44fbe fix(friendli): align model display names with cross-provider majority convention 2026-04-08 11:39:05 +09:00
zhongruan0522 5513af5b6c add GLM-5.1 to zhipuai and zai providers 2026-04-08 02:34:12 +00:00
minpeter 755509be95 feat(friendli): add GLM-5.1 and remove deprecated models 2026-04-08 11:31:50 +09:00
mchen 392b0988f2 update cloudflare-workers-ai: add gemma-4, remove non-LLMs, fix metadata
- Add gemma-4-27b-a4b-it (multimodal, reasoning, tool calling)

- Remove non-LLM models: embeddings, TTS, translation, sentiment analysis

- Remove redundant models: Llama 2/3.x variants, Qwen, Mistral, DeepSeek, Gemma 3

- Update metadata: tool_call, reasoning, open_weights, attachment for remaining models

- Final models: gemma-4, llama-4-scout, kimi-k2.5, nemotron-3, gpt-oss-20b/120b, glm-4.7-flash
2026-04-07 21:00:44 -04:00
Muhammad Mugni Hadi 7f1e94f571 feat(chutes): add zai-org/GLM-5.1-TEE model 2026-04-08 06:40:30 +07:00
Aiden Cline 61c596874c feat: add new provider.body and provider.headers support 2026-04-07 17:36:40 -05:00
Frank ca0e64451d update zen models 2026-04-07 18:01:31 -04:00
Aiden Cline d2870fcfeb Merge pull request #1374 from fhennerkes/dev
Poe: adding Gemma-4-31B (free model)
2026-04-07 16:48:52 -05:00
fhennerkes 9c8e1e0fa0 Merge branch 'anomalyco:dev' into dev 2026-04-07 14:02:06 -07:00
Frank 30d42207cb update zen models 2026-04-07 16:48:49 -04:00
fhennerkes c450a6ffaf poe: add Gemma-4-31B model
Add new Google model from Poe API (released 2026-04-02):
- Free during preview
- 262K context, 8K output
- Modalities: text, image

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-07 13:04:03 -07:00
Frank e889bd7c3b update zen models 2026-04-07 13:43:40 -04:00
Aiden Cline dcd3803dea Merge pull request #1359 from rdbisme/dev
Add missing Qwen3 Coder Next model to Amazon Bedrock
2026-04-07 12:41:41 -05:00
Aiden Cline 05b22d8f8e Merge pull request #1372 from JoshuaDietz/dev
feat(ollama-cloud): add GLM-5.1
2026-04-07 12:30:52 -05:00
Aiden Cline 5c1a00bbe8 Merge pull request #1373 from cantalupo555/feat/openrouter-glm-5.1
feat(openrouter): add z-ai/glm-5.1 model
2026-04-07 12:30:12 -05:00
cantalupo555 27cbd6557d feat(openrouter): add z-ai/glm-5.1 model
- Add GLM-5.1 with 202K context, reasoning, tool call, and structured output

- Pricing: $1.40/M input, $4.40/M output, $0.26/M cache read
2026-04-07 14:02:55 -03:00
JoshuaDietz 8114fa2f13 Merge branch 'anomalyco:dev' into dev 2026-04-07 19:01:23 +02:00
Joshua Dietz 5dc325bf09 feat(ollama-cloud): add GLM-5.1
unsure about temperature=true which is not set for GLM-5 but is set for GLM-5.1 huggingface.
2026-04-07 19:00:54 +02:00
Aiden Cline d27ce785fe Merge pull request #1370 from gary149/feat/huggingface-glm-5.1
feat(huggingface): add GLM-5.1
2026-04-07 11:55:12 -05:00
Aiden Cline f47f9e8414 Merge pull request #1225 from mixlayer/add_mixlayer
New provider: Mixlayer
2026-04-07 11:46:19 -05:00
Victor Muštar df41a38c6f feat(huggingface): add GLM-5.1 2026-04-07 18:33:53 +02:00
Aiden Cline 5b37d05f82 Merge pull request #1367 from cantalupo555/feat/stepfun-step-3.5-flash-2603
feat(stepfun): add step-3.5-flash-2603 model
2026-04-07 11:21:02 -05:00
Aiden Cline 99e046916c Merge pull request #1360 from jonathancaevans/update-kimi-k2p5-turbo-name
Update Kimi K2.5 Turbo display name for Firepass clarity
2026-04-07 11:09:02 -05:00
Aiden Cline 98baf7eaca Merge pull request #1364 from seffhunnn/fix-openrouter-glm-5-turbo-web
fix: correct glm-5-turbo pricing and context for openrouter
2026-04-07 11:08:29 -05:00
Aiden Cline 462d7fa620 Merge pull request #1365 from dpuyosa/feat/venice-add-qwen-3-6-plus
Venice: Add Qwen 3.6 Plus model
2026-04-07 11:08:15 -05:00
cantalupo555 7eea1dec18 feat(stepfun): add step-3.5-flash-2603 model
- Add Step 3.5 Flash 2603 model optimized for agent workflows

- Released April 2, 2026, same pricing as step-3.5-flash
2026-04-07 12:01:26 -03:00
dpuyosa 10652fb8dc feat(venice): add Qwen 3.6 Plus model
- Add Qwen 3.6 Plus with 1M context window

- Support text, image, and video input modalities

- Enable reasoning, tool calling, and structured output
2026-04-07 13:41:42 +02:00
Mohd Saif 816f9bb585 fix: correct glm-5-turbo pricing and context 2026-04-07 14:45:31 +05:30
Jonathan Evans 03376e9986 Update Kimi K2.5 Turbo display name
Add (firepass) suffix to clarify this is the Firepass router endpoint.

Follow-up to #1256
2026-04-06 13:10:01 -07:00
Ruben Di Battista 91d73da942 Enable reasoning capability for Qwen3 Coder Next model 2026-04-06 21:51:16 +02:00
Ruben Di Battista db7e4ff9ba Add missing Qwen3 Coder Next model to Amazon Bedrock 2026-04-06 21:07:36 +02:00
Aiden Cline d6145d1479 Merge pull request #1354 from llc1123/chore/zenmux-updates
zenmux: remove deprecated models and add Agnes 1.5 entries
2026-04-06 08:27:09 -07:00
粒粒橙 ec314aa0f1 fix(zenmux): add image support for agnes-1.5-lite 2026-04-06 14:28:22 +08:00
粒粒橙 188c36696e zenmux: remove deprecated models and add models from sapiens-ai 2026-04-06 14:08:08 +08:00
Aiden Cline 2b5f3f961d Merge pull request #1340 from battall/patch-1
fix: google/gemma-4 -it suffixes
2026-04-05 21:10:43 -07:00
Aiden Cline bf8ce0b8f8 Merge pull request #1342 from seffhunnn/fix-deepseek-context-window
fix: correct deepseek-chat context window to 131072
2026-04-05 21:06:53 -07:00
Aiden Cline 5e3eb74da2 Merge pull request #1341 from u1630022/feat-openrouter-trinity-large-thinking
add trinity large thinking to openrouter provider
2026-04-05 21:06:19 -07:00
Aiden Cline 542620b288 Merge pull request #1343 from spyridonas/patch-1
Fix capitalization in model name
2026-04-05 21:05:15 -07:00
Aiden Cline f6576ffb0f Merge pull request #1344 from dpuyosa/venice/gemma4-trinity
Venice: Add Gemma 4 and Arcee Trinity models, enable GLM 4.6 reasoning
2026-04-05 21:05:00 -07:00
Aiden Cline 88649810a6 Merge pull request #1346 from GHagui/add-gemma-4-openrouter
feat(openrouter): add Gemma 4 26B A4B and Gemma 4 31B models
2026-04-05 21:04:42 -07:00
Aiden Cline 1367d90f32 Merge pull request #1353 from cyberofficial/vultr
[Vultr] Update Inference Models
2026-04-05 21:03:58 -07:00
Cyber Official 322b1154be Update Inference Models
Vultr Updated Inference API endpoint with different versions of models, this commit adds in new models and corrects some information
2026-04-05 17:12:13 -04:00
Gabriel Hagui 7336867f65 Apply suggestions from code review
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2026-04-04 21:17:55 -03:00
GHagui caaf4d8b0a Merge branch 'add-gemma-4-openrouter' of https://github.com/GHagui/models.dev into add-gemma-4-openrouter 2026-04-05 00:09:19 +00:00
GHagui 08c2b928b6 fix(openrouter) Normalize formatting in Gemma 4 26B A4B and Gemma 4 31B TOML files 2026-04-05 00:03:44 +00:00
Gabriel Hagui 3a4b5e50cb Add files via upload
Fixing CRLF to LF
2026-04-04 20:51:25 -03:00
Gabriel Hagui 3244ef835e feat(openrouter) Add Gemma 4 26B A4B and Gemma 4 31B 2026-04-04 20:43:15 -03:00
GHagui 346b12f42f feat(openrouter) Add Gemma 4 26B A4B and Gemma 4 31B 2026-04-04 23:38:56 +00:00
dpuyosa 116d91a0cc [venice] Add Gemma 4 and Arcee Trinity models, enable GLM 4.6 reasoning
- Add Google Gemma 4 26B A4B and 31B instruct models with multimodal support
- Add Arcee Trinity Large Thinking reasoning model
- Enable reasoning capability for GLM 4.6
- Reduce Qwen3 5-9B output limit to 32K
2026-04-05 00:08:47 +02:00
Spyros Sakellaropoulos 778036c53c Fix capitalization in model name 2026-04-05 00:10:17 +03:00
Mohd Saif Ansari e86cc85dd4 fix: update deepseek-chat context window 2026-04-05 01:59:40 +05:30
Eavan Pattie a6f030fe1c add trinity large thinking to openrouter provider
* fixes trinity-large-thinking erroneously marked as not open_weight in
  vercel provider
2026-04-04 22:44:25 +03:00
Aiden Cline 1eb0b8c8e1 Merge pull request #1338 from anthraxx/alibaba-qwen3.6-plus
feat(alibaba): add Qwen3.6 Plus model configuration to all regions
2026-04-04 11:55:23 -07:00
Aiden Cline 1bc0b9d81f Merge pull request #1335 from branchgrove/dev
Add google-vertex DeepSeek V3.2 model
2026-04-04 11:51:37 -07:00
Battal Doğukan Hazar eca166ed4d fix: -it suffix for google/gemma-4 2026-04-04 21:48:31 +03:00
Aiden Cline e64404b173 Merge pull request #1336 from WJQSERVER/dev
Add Google Gemma 4 31B IT to NVIDIA(NIM) provider
2026-04-04 11:48:19 -07:00
Battal Doğukan Hazar e9c8425b32 fix: google/gemma-4 -it suffix 2026-04-04 21:47:27 +03:00
Aiden Cline 8618de3429 Merge pull request #1333 from riccardogiorato/dev
Remove deprecated models from TogetherAI provider
2026-04-04 11:28:46 -07:00
Aiden Cline 6e64316225 Merge pull request #1337 from fanweixiao/vivgrd/gpt-5.4
provider(vivgrid): add GPT-5.3 Codex, GPT-5.4 Mini, and GPT-5.4 Nano models
2026-04-04 11:28:33 -07:00
Aiden Cline a36d032e93 Merge pull request #1339 from cantalupo555/remove/qwen3.6-plus-preview-free
chore(openrouter): remove discontinued Qwen3.6 Plus Preview free
2026-04-04 11:28:18 -07:00
Aiden Cline e74dce023f Merge pull request #1315 from seffhunnn/fix-pdf-modalities
fix: add missing pdf modality to supported OpenAI models
2026-04-04 11:28:07 -07:00
cantalupo555 47f9b2b910 chore(openrouter): remove discontinued Qwen3.6 Plus Preview free
- Remove qwen3.6-plus-preview:free model after Qwen3.6 Plus free replaced it
2026-04-04 15:08:11 -03:00
Levente Polyak 6ba0af61d6 feat(alibaba): add Qwen3.6 Plus model configuration to all regions
- Add missing regions
- Add coding-plan variants
- Use pricing from model info page

Link: https://bailian.console.alibabacloud.com/cn-beijing?tab=model#/model-market/detail/qwen3.6-plus
2026-04-04 19:52:15 +02:00
C.C. Fan 24d4a9b8dd provider(vivgrid): add GPT-5.3 Codex, GPT-5.4 Mini, and GPT-5.4 Nano models
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-04 23:07:49 +08:00
wjqserver d5b420768a feat: add Google Gemma 4 31B IT to NVIDIA provider 2026-04-04 22:26:39 +08:00
Riccardo Giorato 66ada85298 more deprecations 2026-04-04 14:18:08 +02:00
Elias Lundgren 2ffb1d9181 Add google-vertex DeepSeek V3.2 model 2026-04-04 13:57:42 +02:00
Riccardo Giorato 04768ff141 Merge pull request #1 from riccardogiorato/orchestrator/remove-deprecated-models-g5nx
Remove deprecated models from togetherai provider
2026-04-03 21:51:02 +02:00
orchestrator-dev[bot] 9560d51908 Remove deprecated models from togetherai provider 2026-04-03 19:49:40 +00:00
Aiden Cline 6a41e31306 Merge pull request #1326 from michaelnchin/fix/amazon-bedrock-structured-output
fix: set correct structured output values for Amazon Bedrock models
2026-04-03 14:03:05 -05:00
Aiden Cline 133c529ebf Merge pull request #1327 from llc1123/chore/zenmux-new-models
zenmux: add KAT-Coder-Pro-V2, Qwen3.6-Plus, and GLM 5V Turbo
2026-04-03 14:02:40 -05:00
Aiden Cline 0a6b828e42 Merge pull request #1331 from zhongruan0522/feat/xiaomi-token-plan
feat: add Xiaomi Token Plan providers (cn/sgp/ams)
2026-04-03 14:00:46 -05:00
Aiden Cline 406f2f66c6 Merge pull request #1323 from Pxys-io/fix-xiaomi-mimo-cache-pricing
fix(openrouter): add missing cache_read pricing for xiaomi/mimo-v2-pro and xiaomi/mimo-v2-omni
2026-04-03 13:59:51 -05:00
Aiden Cline 948ce76d8d Merge pull request #1325 from DEAN-Cherry/dev
revert: alibaba-cn MiniMax-M2.5 to MiniMax-M2.7
2026-04-03 13:47:53 -05:00
Aiden Cline e3dd89ba2e Merge pull request #1332 from vercel/update-vercel-models-20260403-1639
Update Vercel models
2026-04-03 13:47:33 -05:00
Aiden Cline 7e5ae3bb06 Merge pull request #1329 from Track07-cda/alibaba-cn-qwen3.6
feat(alibaba-cn): add Qwen3.6 Plus model configuration
2026-04-03 13:47:19 -05:00
Aiden Cline 5392c185d2 Merge pull request #1328 from sadoclaw/add-gemma-4-models
Add Gemma 4 26B and 31B models
2026-04-03 13:47:00 -05:00
github-actions[bot] 72613f5dbf chore(vercel): update Vercel model definitions
Auto-generated by weekly workflow from Vercel AI Gateway API.

Co-Authored-By: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-03 16:39:32 +00:00
zhongruan0522 01098f81a9 feat: add Xiaomi Token Plan providers (cn/sgp/ams) 2026-04-03 11:22:04 +00:00
Track07-cda 65e798c8f2 feat(alibaba-cn): add Qwen3.6 Plus model configuration 2026-04-03 15:58:53 +08:00
sadoclaw 28a75176c4 Add Gemma 4 26B and 31B models 2026-04-03 08:34:41 +03:00
粒粒橙 61e63db8e3 zenmux: add new models 2026-04-03 13:09:53 +08:00
Michael Chin 95cd6196fa set correct structured_output values for Amazon Bedrock models 2026-04-02 21:15:18 -07:00
Bryan 19c1f2ebe2 revert: alibaba-cn MiniMax-M2.5 to MiniMax-M2.7
Revert PR #1002 - MiniMax-M2.5 is no longer available, now using MiniMax-M2.7
2026-04-03 12:09:05 +08:00
pxys-io 4d25bb5622 fix(openrouter): add missing cache_read pricing for xiaomi/mimo-v2-pro and xiaomi/mimo-v2-omni
OpenRouter charges bash.20/M cache_read tokens for mimo-v2-pro and
bash.08/M for mimo-v2-omni, but these were missing from the cost section.

Source: https://openrouter.ai/api/v1/models

Co-authored-by: Qwen-Coder <qwen-coder@alibabacloud.com>
2026-04-03 04:08:55 +02:00
Aiden Cline 8845bf3f3b Merge pull request #1321 from BlockListed/cortecs-add-glm-5
Add glm-5 to cortecs
2026-04-02 19:28:50 -05:00
Aiden Cline 9b5bcde109 Merge pull request #1322 from fhennerkes/dev
poe: add GPT-5.3-Codex-Spark and Kimi-K2.5-FW models
2026-04-02 19:28:41 -05:00
Frank fa75002f19 update zen models 2026-04-02 19:01:00 -04:00
fhennerkes 69b6f3e94a poe: add GPT-5.3-Codex-Spark and Kimi-K2.5-FW models
Add 2 new free models
2026-04-02 15:11:34 -07:00
BlockListed 26f6b602fc add glm-5 to cortecs 2026-04-02 21:45:20 +02:00
Aiden Cline 287c69acaf Merge pull request #1314 from dpark01/add-kimi-k2-thinking-vertex
feat(google-vertex): add moonshotai/kimi-k2-thinking-maas model
2026-04-02 11:30:08 -05:00
Aiden Cline 2b46c3aef2 Merge pull request #1320 from cantalupo555/feature/qwen3.6-plus-free
feat: add Qwen3.6 Plus free on OpenRouter
2026-04-02 11:29:46 -05:00
cantalupo555 a9b7faa409 feat(openrouter): add qwen3.6-plus free model configuration
- Add Qwen3.6 Plus (free) provider configuration
- $0 pricing with 1M context window
- Multimodal input support (text, image, video)
- Full capabilities: reasoning, tool calls, structured output
- Attachment enabled for vision modality
2026-04-02 13:23:09 -03:00
Jack ad3305bc08 Merge branch 'dev' of github.com:anomalyco/models.dev into dev 2026-04-03 00:12:23 +08:00
Jack 0198228bbb Add MiMo V2 Pro/Omni models and family entries
Register two new MiMo V2 models and update model family values. Added "mimo-pro" and "mimo-omni" to ModelFamilyValues in packages/core/src/family.ts, and added corresponding TOML model descriptors under providers/opencode-go/models: mimo-v2-pro.toml and mimo-v2-omni.toml. The Pro model includes very large context (1,048,576) and tiered costs for >200k context, while the Omni model exposes multimodal input (text, image, audio, pdf) and a large 262,144 context. Both files set metadata (release_date, last_updated, knowledge cutoff, open_weights) and define interleaved reasoning field, costs, limits, and modalities.
2026-04-03 00:11:58 +08:00
Aiden Cline 165bc7df94 Merge pull request #1316 from NIKU-SINGH/remove-claude-3-7-sonnet-latest
Remove invalid claude-3-7-sonnet-latest model entry
2026-04-02 10:51:44 -05:00
NIKU-SINGH 36956dc4f4 Remove invalid claude-3-7-sonnet-latest model entry
Anthropic's API does not accept claude-3-7-sonnet-latest as a model ID
(returns 404). The versioned alias claude-3-7-sonnet-20250219 should be
used instead.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-04-02 16:27:15 +05:30
Mohd Saif Ansari 5ec371653b fix: add missing pdf modality to supported OpenAI models 2026-04-02 12:40:50 +05:30
Frank bb62fc43c1 update zen models 2026-04-01 22:59:53 -04:00
Frank c1465447ab update zen models 2026-04-01 17:53:17 -04:00
Daniel Park 471e9ec492 feat(google-vertex): add moonshotai/kimi-k2-thinking-maas model 2026-04-01 16:24:14 -04:00
Aiden Cline 57c3d38c21 Merge pull request #1313 from zhongruan0522/dev
Add GLM-5V-Turbo
2026-04-01 13:08:38 -05:00
Aiden Cline d19d2508b0 Merge pull request #1273 from Cahl-Dee/add-the-grid-ai
feat: add thegrid.ai provider and associated models
2026-04-01 12:25:35 -05:00
zhongruan0522 d0d2fffcfb Add GLM-5V-Turbo 2026-04-01 16:30:38 +00:00
Aiden Cline 07f48b1f2a Merge pull request #1311 from Moniet/feat/add-gpt-image-models
feat(openai): add openai gpt-image models
2026-04-01 10:52:12 -05:00
Moniet a33e15d095 feat(openai): add openai gpt-image models 2026-04-01 17:03:36 +05:30
Aiden Cline b64c2100f4 Merge pull request #1306 from dinhkim/feat/add-openrouter-glm-5-turbo
feat: add GLM-5-Turbo model in OpenRouter AI provider
2026-03-31 23:11:14 -05:00
Kim Truong 7ba7237633 feat: add GLM-5-Turbo model in OpenRouter AI provider 2026-03-31 23:18:02 +07:00
Aiden Cline 6e1ca23e6c Merge pull request #1305 from marcelarie/dev
Update synthetic.new model: Qwen3.5-397B
2026-03-31 10:43:23 -05:00
Aiden Cline 8ffb4ed5a7 Merge pull request #1301 from xinrui-z/fix/aihubmix-zod-validation-provider
fix(aihubmix): zod-validation-error
2026-03-31 10:43:12 -05:00
Xinrui 6b12398083 Replace @ai-sdk/openai-compatible with the official aihubmix/ai-sdk-provider package and remove the hardcoded api URL, as the dedicated package handles schema validation and endpoint configuration internally. 2026-03-31 23:29:10 +08:00
Aiden Cline 751745f200 Merge pull request #1304 from dpuyosa/add-gpt-54-mini-venice
Venice: Add GPT-5.4 Mini and remove discontinued models
2026-03-31 10:19:03 -05:00
marcelarie 593308b596 Merge branch 'dev' of github.com:marcelarie/models.dev into dev 2026-03-31 13:05:29 +02:00
marcelarie 54386f35b7 Added: Missing synthetic.new Qwen3.5-397B model 2026-03-31 13:02:19 +02:00
dpuyosa 73a83971e8 [venice] Add GPT-5.4 Mini and remove discontinued models
- Add GPT-5.4 Mini with reasoning and tool_call
- Update Aion 2.0 with reasoning capability
- Remove discontinued mistral-31-24b and qwen3-4b
2026-03-31 09:49:09 +02:00
Xinrui b5d8da29bb fix(aihubmix): switch to dedicated provider package to resolve Zod validation error
Replace @ai-sdk/openai-compatible with the official aihubmix/ai-sdk-provider
package and remove the hardcoded api URL, as the dedicated package handles
schema validation and endpoint configuration internally.
2026-03-31 12:37:56 +08:00
Aiden Cline 798538f9ae Merge pull request #1254 from llc1123/dev
feat(zenmux): route models through protocol-specific SDKs
2026-03-30 18:50:02 -05:00
Aiden Cline d4ce566f27 Merge pull request #1296 from sylviezhang37/update-vercel-models-20260330-1655
Update Vercel models
2026-03-30 18:49:47 -05:00
Aiden Cline 2042e3dd71 Merge pull request #1298 from cantalupo555/feature/qwen3.6-plus-preview-free
feat: add Qwen3.6 Plus Preview free on OpenRouter
2026-03-30 18:49:34 -05:00
Aiden Cline 08b577b6c9 Merge pull request #1299 from cyberofficial/vultr
Remove discontinued Vultr models
2026-03-30 18:49:24 -05:00
Cyber Official 01c3d44f99 Remove discontinued Vultr models
Vultr no longer supports these models on Serverless Inference:
- DeepSeek R1 Distill Llama 70B
- DeepSeek R1 Distill Qwen 32B
- GPT OSS 120B
- Llama 3.1 Nemotron Ultra 253B v1
- NVIDIA Nemotron 3 Super 120B A12B NVFP4

Remaining active models:
- MiniMax-M2.5: $0.30/M in, $1.20/M out
- DeepSeek-V3.2: $0.55/M in, $1.65/M out
- Kimi-K2.5: $0.55/M in, $2.75/M out
- GLM-5-FP8: $0.85/M in, $3.10/M out
2026-03-30 17:39:00 -04:00
cantalupo555 1695bfb958 feat: add Qwen3.6 Plus Preview free on OpenRouter
- Add free variant of Qwen3.6 Plus Preview to OpenRouter provider
- 1M context, 65K max output, /bin/bash.00 pricing
- Text-only modality (OpenRouter API reports text->text)
- Source: OpenRouter /api/v1/models API

---
Co-Authored-By: opencode https://opencode.ai
2026-03-30 18:08:54 -03:00
Frank bad8bedd25 update zen models 2026-03-30 16:50:10 -04:00
github-actions[bot] ad26e881ce chore(vercel): update Vercel model definitions
Auto-generated by weekly workflow from Vercel AI Gateway API.

Co-Authored-By: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-03-30 16:55:28 +00:00
Aiden Cline 348932f00c Merge pull request #1290 from zhongruan0522/dev
feat: add gpt-5.3-chat-latest model to openai
2026-03-29 22:25:52 -05:00
Aiden Cline d564c80b78 Merge pull request #1289 from pedrxd/mistral-add-mistral-small-4
Adding mistral small 4
2026-03-29 22:23:57 -05:00
Aiden Cline cba38e4707 Merge pull request #1287 from aeonzh/patch-1
Use correct name for Nemotron 3 Super (free) on OpenRouter
2026-03-29 12:35:18 -05:00
Aiden Cline 7ce9e95d62 Merge pull request #1294 from sk0x0y/add-glm-5.1-nanogpt
feat(nano-gpt): add glm-5.1 and glm-5.1:thinking models
2026-03-29 12:34:31 -05:00
sk0x0y a05b2a6be7 add glm-5.1 model to nano-gpt provider 2026-03-29 23:12:38 +09:00
zhongruan0522 17f41b72ad feat: add gpt-5.3-chat-latest model to openai 2026-03-29 11:18:44 +00:00
Pedro Ruiz 16cc382572 feat(mistral): Adding mistral small 4 2026-03-29 09:46:58 +02:00
Aiden Cline 3d456e3798 Merge pull request #1286 from khda-tech/dev
Add gemma family for google provider
2026-03-28 20:07:16 -05:00
Aiden Cline 0223ab3107 Merge pull request #1288 from cgilly2fast/dev
fix(firmware): incorrect model name for glm-5
2026-03-28 20:06:57 -05:00
Colby Gilbert 28ad50f6e9 fix(firmware): incorrect model name for glm-5 2026-03-27 22:26:19 -07:00
Zheng He Hu bf8fb378a4 Rename nemotron-3-super-120b-a12b-free.toml to nemotron-3-super-120b-a12b:free.toml 2026-03-28 02:54:35 +01:00
Aiden Cline b74242fdbf Merge pull request #1278 from fhennerkes/dev
poe: add Grok-4.20-Multi-Agent and DeepSeek-V3.2 models
2026-03-27 15:48:08 -05:00
Aiden Cline 8131cc947c Update providers/poe/models/novita/deepseek-v3.2.toml
Co-authored-by: Oleg Voronkovich <oleg-voronkovich@yandex.ru>
2026-03-27 15:21:16 -05:00
Khrulev Danil 95a73581cc Add gemma family for google provider 2026-03-27 21:39:52 +03:00
Zack Angelo 39ee133c98 mixlayer: adhere to logo standards, remove fill and size attributes 2026-03-27 09:23:09 -07:00
Aiden Cline c5e4e2c740 Merge pull request #1276 from voronkovich/feat-update-groq
feat(groq): Update Groq models
2026-03-27 10:51:40 -05:00
Aiden Cline 357c3021fb Merge pull request #1281 from zhongruan0522/dev
Added support for Zhipu AI's official CodingPlan GLM-5.1 model.
2026-03-27 09:49:49 -05:00
Aiden Cline 6afb0fea06 Merge pull request #1280 from dpuyosa/dev
Venice: Add Aion 2.0, update DeepSeek V3.2, remove Gemini 3 Pro Preview
2026-03-27 09:46:58 -05:00
dd781c4c15 feat: add glm-5.1 model to zai-coding-plan and zhipuai-coding-plan 2026-03-27 11:51:14 +00:00
dpuyosa 04e3b7c008 [venice] Add Aion 2.0, update DeepSeek V3.2, remove Gemini 3 Pro Preview
- feat(venice): add Aion 2.0 model
- fix(venice): enable tool_call and structured_output on DeepSeek V3.2
- fix(venice): remove deprecated Gemini 3 Pro Preview
2026-03-27 09:19:28 +01:00
Oleg Voronkovich ca40cb538d Updates 2026-03-27 00:37:40 +03:00
Aiden Cline f03de60559 Merge pull request #1259 from smakosh/llmgateway-models
feat: update LLM Gateway to 204 models
2026-03-26 15:25:22 -05:00
fhennerkes 0226a37a51 poe: add Grok-4.20-Multi-Agent and DeepSeek-V3.2 models 2026-03-26 12:18:13 -07:00
smakosh c2ba40d7d5 fix: logo format, remove README, minimax weights
- Normalize logo to 24x24, viewBox 0 0 40 40, currentColor
- Remove README.md (other providers don't have one)
- Mark all MiniMax models as open_weights = true

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-26 20:02:58 +01:00
smakosh d77bee4f29 fix: correct reasoning, vision, tools flags
The export script only checked the first active
provider for capabilities. Now checks all providers
and uses model ID patterns for reasoning detection.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-26 19:55:28 +01:00
Oleg Voronkovich 5ec0ac2807 feat(groq): Update Groq models 2026-03-26 18:59:09 +03:00
Aiden Cline 62015086c6 Merge pull request #1274 from petit-blaireau/copilot/add-openrouter-mistral-small-4
Add Mistral Small 4 for OpenRouter
2026-03-25 19:47:51 -05:00
copilot-swe-agent[bot] 2d86bafb20 feat(openrouter): add Mistral Small 4 (mistral-small-2603)
Co-authored-by: petit-blaireau <1893252+petit-blaireau@users.noreply.github.com>
Agent-Logs-Url: https://github.com/petit-blaireau/models.dev/sessions/871bbf50-3ef0-4926-afd1-e95c79f6bd57
2026-03-26 00:03:38 +00:00
Cahl-Dee bf4e5aab17 remove family property and add open_weight 2026-03-25 16:57:02 -05:00
Cahl-Dee 1590791225 adding thegrid.ai provider and associated models 2026-03-25 16:09:54 -05:00
Aiden Cline 047f3356d6 Merge pull request #1265 from MiyakoMeow/add-glm-4.7-flashx
Add glm-4.7-flashx to ZAI/ZhipuAI
2026-03-25 16:09:24 -05:00
Aiden Cline 1394d2ca4d Merge pull request #1270 from NachoFLizaur/feat/bedrock-add-nemotron-super-3-120b
feat(amazon-bedrock): add NVIDIA Nemotron 3 Super 120B
2026-03-25 15:05:31 -05:00
Aiden Cline 5b73677b33 Merge pull request #1272 from NachoFLizaur/fix/bedrock-minimax-m2.5-glm-5-limits
fix(amazon-bedrock): correct MiniMax M2.5 and GLM-5 context/output limits
2026-03-25 15:05:18 -05:00
Nacho F. Lizaur 780db038c4 fix(amazon-bedrock): correct MiniMax M2.5 and GLM-5 context/output limits 2026-03-25 16:49:31 +01:00
Nacho F. Lizaur 29c1249d64 feat(amazon-bedrock): add NVIDIA Nemotron 3 Super 120B 2026-03-25 15:44:56 +01:00
lioZ129 1b0633b172 Update providers/hpc-ai/models/moonshotai/kimi-k2.5.toml
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2026-03-25 17:15:16 +08:00
lioZ129 aed0ee3bb7 Update providers/hpc-ai/models/moonshotai/kimi-k2.5.toml
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2026-03-25 17:15:03 +08:00
Contributor 9fa7856fc6 feat: add HPC-AI model provider support 2026-03-25 16:55:45 +08:00
MiyakoMeow 8b80a2b34b Add glm-4.7-flashx to ZAI/ZhipuAI 2026-03-25 06:13:03 +08:00
Aiden Cline 897aa53905 Merge pull request #1257 from fhennerkes/dev
poe: add GPT-5.4-Nano and GPT-5.4-Mini models
2026-03-24 15:19:13 -05:00
Aiden Cline 5c36e54a43 Merge pull request #1260 from Happily-Coding/dev
Add MiniMax 2.5 to siliconflow
2026-03-24 15:18:39 -05:00
Aiden Cline a38f9373ab Merge pull request #1261 from cyberofficial/vultr
[Vultr] Delete Qwen2.5-Coder-32B-Instruct.toml
2026-03-24 10:14:38 -05:00
Aiden Cline fb72c181f8 Merge pull request #1263 from fanweixiao/vivgrd/gpt-5.4
provider(vivgrid): add gpt-5.4 and upgrade gemini-3 to gemini-3.1
2026-03-24 10:14:20 -05:00
C.C. Fan 601300c7e7 provider(vivgrid): add gpt-5.4 and upgrade gemini-3 to gemini-3.1 2026-03-24 21:04:59 +08:00
Cyber Official 8d9e966867 Delete Qwen2.5-Coder-32B-Instruct.toml
Model no longer offered
2026-03-24 02:48:38 -04:00
UrielS cf8f12b375 Add MiniMax 2.5 to siliconflow
Add MiniMax 2.5 to siliconflow
2026-03-24 01:14:36 -03:00
UrielS 116e35a32a Add MiniMax 2.5 to siliconflow 2026-03-24 01:13:37 -03:00
Aiden Cline 87a02b897a Merge pull request #1258 from vglafirov/feat/gitlab-gpt-5-4-models
feat(gitlab): add GPT-5.4, GPT-5.4 Mini, GPT-5.4 Nano, and GPT-5.3 Codex models
2026-03-23 22:00:07 -05:00
Vladimir Glafirov 849a529a42 fix(gitlab): use unscoped gitlab-ai-provider npm package name 2026-03-23 23:52:13 +01:00
smakosh c2f0bd08e6 feat: update LLM Gateway models to 204
Regenerated model exports from latest LLM Gateway
source. Adds 66 new models including Claude 4.6,
GPT-5.x, Gemini 3.1, Grok 4, and more. Removes
deprecated model aliases.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-23 22:28:46 +01:00
Vladimir Glafirov 5257d0919a feat(gitlab): add GPT-5.4, GPT-5.4 Mini, GPT-5.4 Nano, and GPT-5.3 Codex models 2026-03-23 21:41:11 +01:00
fhennerkes 379ab2757f poe: add GPT-5.4-Nano and GPT-5.4-Mini models
Add 2 new OpenAI models from Poe API:

GPT-5.4-Nano (released 2026-03-11):
- Reasoning support, 400K context, 128K output
- Cost: $0.18/M input, $1.1/M output
- Modalities: text, image

GPT-5.4-Mini (released 2026-03-12):
- Reasoning support, 400K context, 128K output
- Cost: $0.68/M input, $4/M output
- Modalities: text, image

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-03-23 13:13:10 -07:00
Aiden Cline 9838c55e29 Merge pull request #1256 from jonathancaevans/add-kimi-k2p5-turbo-router
Add Fireworks Kimi K2.5 Turbo router
2026-03-23 15:08:32 -05:00
Aiden Cline 4235bc6432 Merge pull request #1249 from tobwen/cleanup/openrouter-deprecated-models
chore(openrouter): remove deprecated and unavailable models
2026-03-23 15:07:27 -05:00
Jonathan Evans 29c3e9cbf4 Add Kimi K2.5 Turbo router for Fireworks
- Model ID: accounts/fireworks/routers/kimi-k2p5-turbo

- Pricing set to 0 (handled at subscription layer)

- Supports text and image input, text output

- Includes reasoning capabilities
2026-03-23 16:01:08 -04:00
Aiden Cline b0c1f37aad Merge pull request #1255 from sylviezhang37/update-vercel-models-20260323-1954
Update Vercel models
2026-03-23 15:00:14 -05:00
github-actions[bot] bd07f155f5 chore(vercel): update Vercel model definitions
Auto-generated by weekly workflow from Vercel AI Gateway API.

Co-Authored-By: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-03-23 19:54:44 +00:00
粒粒橙 0f95849d38 fix(zenmux): align model metadata with runtime support 2026-03-23 23:06:36 +08:00
粒粒橙 8dc90a4097 feat(zenmux): route models through protocol-specific SDKs 2026-03-23 21:48:00 +08:00
Aiden Cline 0a19559e4e Merge pull request #1252 from tarun1793/add-glm-5-fastrouter
Add GLM-5 model to fastrouter provider
2026-03-23 08:13:45 -05:00
Aiden Cline f27f83bdf9 Merge pull request #1253 from jacksonwilliamsva/add-bedrock-minimax-m2.5-glm-5
feat(amazon-bedrock): add MiniMax M2.5 and GLM-5 models
2026-03-23 08:13:35 -05:00
Aiden Cline 7b050ec719 Merge pull request #1244 from BlockListed/cortecs-add-claude-models
Add more claude models to cortecs
2026-03-23 08:13:27 -05:00
Aiden Cline 443cfd7703 Merge pull request #1247 from wojons/dev
Add Nemotron 3 Super model to OpenRouter and NVIDIA providers
2026-03-23 08:11:12 -05:00
Aiden Cline 6d648bf471 Merge pull request #1250 from dsingal0/dev
correct model name for nemotron super 3 on baseten
2026-03-23 08:10:45 -05:00
Jackson Williams e1a83a6812 feat(amazon-bedrock): add MiniMax M2.5 and GLM-5 models
Add two newly available Amazon Bedrock models:

- minimax.minimax-m2.5: 1M context, /bin/bash.30/.20 per 1M tokens
- zai.glm-5: 200K context, .00/.20 per 1M tokens

Both models were added to Amazon Bedrock on March 18, 2026.
Specs sourced from AWS Bedrock pricing page and vendor documentation.
2026-03-23 11:41:21 +11:00
Tarun b8606ed0e4 use latest price from fastrouter 2026-03-22 23:47:13 +00:00
Tarun 69c4600842 Override fastrouter glm-5 with zai glm-5 values 2026-03-22 23:43:09 +00:00
Tarun 5eb4a369f3 Add GLM-5 model to fastrouter provider 2026-03-22 23:31:53 +00:00
Dhruv Singal 17df1b49ce Update Baseten Nemotron model name 2026-03-22 13:04:40 -07:00
Dhruv Singal f18bc9b95b Update Baseten Nemotron display name 2026-03-22 13:00:49 -07:00
Dhruv Singal 485c37e862 Rename Baseten Nemotron model to match API ID 2026-03-22 12:57:10 -07:00
tobwen 0f092e3f62 chore(openrouter): remove expired/revealed/ended endpoints 2026-03-22 12:28:08 +00:00
tobwen a7bcd7e632 chore(openrouter): remove models without endpoints 2026-03-22 12:27:59 +00:00
Alexis Okuwa 6d081af472 Add Nemotron 3 Super model to OpenRouter and NVIDIA providers 2026-03-22 05:45:00 -05:00
Aiden Cline 8ee9ea1d96 Merge pull request #1243 from v1gnesh/dev
Update Grok 4.2 model names
2026-03-21 11:56:37 -05:00
Aiden Cline c71a365320 Merge pull request #1245 from Daltonganger/add-nanogpt-minimax-m2-7
Add NanoGPT MiniMax M2.7 model metadata
2026-03-21 11:55:15 -05:00
Ruben Beuker bb0e828b77 add NanoGPT MiniMax M2.7 model metadata 2026-03-21 14:58:35 +01:00
BlockListed 58cb222125 add more claude models to cortecs 2026-03-21 09:18:00 +01:00
v1gnesh 33f67289ee Rename model and remove beta status 2026-03-21 11:29:13 +05:30
v1gnesh ad34d7948c Update model name and status in TOML file 2026-03-21 11:28:22 +05:30
v1gnesh 7132293513 Add grok-4.20-0309-non-reasoning.toml file 2026-03-21 11:27:53 +05:30
Aiden Cline 495bc263e7 Merge pull request #1241 from BlockListed/add-minimax-2.5-cortecs
Add minimax M2.5 to cortecs
2026-03-20 15:55:28 -05:00
Aiden Cline 3811a45efe Merge pull request #1235 from anomalyco/github-sync
sync github copilot limits
2026-03-20 15:55:09 -05:00
BlockListed c4d04d2ed9 add minimax m2.5 to cortecs 2026-03-20 21:51:30 +01:00
Aiden Cline 20fcbbc336 Merge pull request #1240 from sylviezhang37/update-vercel-models-20260320-1642
Update Vercel models
2026-03-20 13:08:30 -05:00
github-actions[bot] 482b8ed69d chore(vercel): update Vercel model definitions
Auto-generated by weekly workflow from Vercel AI Gateway API.

Co-Authored-By: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-03-20 16:42:15 +00:00
Aiden Cline b77e95b84e Merge pull request #1236 from LYY/update/zenmux-sync
Sync ZenMux models with latest website data
2026-03-20 10:23:40 -05:00
Aiden Cline 463006f80c Merge pull request #1237 from vincentbernat/fix/scaleway-qwen3.5
fix(scaleway): set the correct family for Qwen 3.5 for Scaleway
2026-03-20 10:23:17 -05:00
Vincent Bernat 45b5abb875 fix(scaleway): set the correct family for Qwen 3.5 for Scaleway 2026-03-20 06:32:51 +01:00
LYY 4255e420ec Sync zenmux models with website
Add 17 new models found on zenmux.ai:
- google/gemini-3-pro-image-preview
- google/gemini-3.1-flash-lite-preview
- minimax/minimax-m2.7
- minimax/minimax-m2.7-highspeed
- openai/gpt-5.3-chat
- openai/gpt-5.3-codex
- openai/gpt-5.4
- openai/gpt-5.4-mini
- openai/gpt-5.4-nano
- openai/gpt-5.4-pro
- qwen/qwen3.5-flash
- qwen/qwen3.5-plus
- volcengine/doubao-seed-2.0-code
- x-ai/grok-4.2-fast
- x-ai/grok-4.2-fast-non-reasoning
- xiaomi/mimo-v2-omni
- xiaomi/mimo-v2-pro
- z-ai/glm-5-turbo

Mark anthropic/claude-3.5-sonnet as deprecated (not found on website)
2026-03-20 12:46:16 +08:00
Aiden Cline cda892e0d7 sync github copilot limits 2026-03-19 21:54:42 -05:00
Aiden Cline 098ff4f5bf Merge pull request #1227 from Verizane/dev
add OpenRouter models for gpt-5.4 mini and gpt-5.4 nano
2026-03-19 21:23:15 -05:00
Aiden Cline 0f70b8959f Merge pull request #1234 from mchenco/dev
Add Workers AI models: kimi-k2.5, nemotron-3-120b-a12b, glm-4.7-flash
2026-03-19 15:11:19 -05:00
mchen b8e6d58e5b add workers-ai models: kimi-k2.5, nemotron-3-120b-a12b, glm-4.7-flash 2026-03-19 14:59:17 -04:00
Roman Koslowski a855001a7e apply changes from review 2026-03-19 17:20:26 +01:00
Aiden Cline ac760b2268 Merge pull request #1230 from SamizuHM/feature/zhipuai-coding-plan-add-glm-5-turbo
zhipuai-coding-plan: Add glm-5-turbo.toml and replace symlink
2026-03-19 10:42:43 -05:00
Aiden Cline d4a5ea7ae7 Merge pull request #1226 from spiffytech/dev
Add Ollama Cloud support for Minimax M2.7
2026-03-19 10:41:47 -05:00
Aiden Cline 434ed89ba2 Merge pull request #1228 from dpuyosa/minimax_m2_7
Venice: Add MiniMax M2.7 and update DeepSeek V3.2 pricing
2026-03-19 10:41:16 -05:00
Aiden Cline 6d7719a62a Merge pull request #1229 from 0b1000/dev
Xiaomi: Add MiMo-V2-Pro and MiMo-V2-Omni
2026-03-19 10:41:06 -05:00
Aiden Cline 93637039ef Merge pull request #1231 from ariane-emory/feat/feat/add-xiaomi-mimo-v2-pro-and-omni
feat: add the Xiaomi MiMo V2 Pro and Xiaomi MiMo V2 Omni models to the OpenRouter provide
2026-03-19 10:40:44 -05:00
Ariane Emory 9c95f796c0 Merge remote-tracking branch 'upstream/dev' into feat/feat/add-xiaomi-mimo-v2-pro 2026-03-19 11:22:43 -04:00
Ariane Emory e8650b6073 feat: add xiaomi mimo-v2-pro and mimo-v2-omni models to openrouter 2026-03-19 11:18:46 -04:00
SamizuHM 23c2be6ff7 feat(zhipuai-coding-plan): add glm-5-turbo.toml and replace glm-5-turbo with symlink 2026-03-19 18:09:17 +08:00
Frank 913a63dbe6 update zen models 2026-03-19 00:33:45 -04:00
0b1000 503087e99b Merge branch 'anomalyco:dev' into dev 2026-03-19 12:28:38 +08:00
0b1000 48150f09d3 Xiaomi: Add MiMo-V2-Pro and MiMo-V2-Omni 2026-03-19 12:27:00 +08:00
Aiden Cline 5fef681657 Disable tool_call in grok model configuration 2026-03-18 23:09:30 -05:00
Frank 123054ae0c update zen models 2026-03-18 20:45:44 -04:00
Frank 03060d154b update zen models 2026-03-18 20:37:47 -04:00
dpuyosa 5c9b8108e0 Update minimax-m27.toml 2026-03-19 01:02:24 +01:00
dpuyosa c8084681f9 [venice] Add MiniMax M2.7 and update DeepSeek V3.2 pricing
- Add MiniMax M2.7 model with reasoning and tool_call support
 - Update DeepSeek V3.2 pricing (input: $0.33, output: $0.48, cache: $0.16)
2026-03-19 00:58:50 +01:00
Roman Koslowski 352ab4ae1b add gpt-5.4 mini and gpt-5.4 nano 2026-03-18 22:16:55 +01:00
spiffytech cf0b416b15 Added Ollama Cloud support for Minimax M2.7 2026-03-18 16:15:07 -04:00
Aiden Cline 38339a2a90 Merge pull request #1224 from APonce911/minimax-m2.7-openrouter
add MiniMax M2.7 to OpenRouter
2026-03-18 14:10:13 -05:00
Aiden Cline ff9040bf52 Update minimax-m2.7.toml 2026-03-18 14:09:26 -05:00
Aiden Cline 3039804af4 Delete providers/opencode/models/minimax-m2.7.toml 2026-03-18 14:08:55 -05:00
Frank 7a4ad7bec8 update go models 2026-03-18 14:40:25 -04:00
Zack Angelo 7a2ec5ab95 New provider: Mixlayer 2026-03-18 11:19:35 -07:00
airton 721cc122bc add MiniMax M2.7 to OpenRouter and OpenCode 2026-03-18 18:57:09 +01:00
Aiden Cline 0527f019af Merge pull request #1221 from sergical/fix/bedrock-claude-4-6-context-window-and-pricing
fix(amazon-bedrock): set Claude Sonnet 4.6 and Opus 4.6 context window to 1M
2026-03-18 12:17:57 -05:00
Aiden Cline c89371de50 Merge pull request #1223 from sylviezhang37/update-vercel-models-20260318-1659
Update Vercel models
2026-03-18 12:17:22 -05:00
Sylvie Zhang 6d6d4220d8 Enable open_weights in minimax-m2.7.toml 2026-03-18 10:12:37 -07:00
Sylvie Zhang 8b984eeec1 Enable open_weights in minimax-m2.7-highspeed model 2026-03-18 10:12:21 -07:00
github-actions[bot] 586027c8f1 chore(vercel): update Vercel model definitions
Auto-generated by weekly workflow from Vercel AI Gateway API.

Co-Authored-By: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-03-18 16:59:24 +00:00
Sergiy Dybskiy 343b5f87ef fix(amazon-bedrock): set Claude Sonnet 4.6 and Opus 4.6 context window to 1M
Both models support a 1M token context window natively on Bedrock via the
Converse API with no beta headers required. Verified empirically via the
AWS CLI (bedrock-runtime converse): 950K tokens succeeds, >1M returns
'prompt is too long: N tokens > 1000000 maximum'.

The AWS Bedrock pricing page confirms long context pricing for these two
models is identical to standard pricing (no surcharge), so the
[cost.context_over_200k] section is removed as it was incorrect.
2026-03-18 12:20:38 -04:00
Aiden Cline 955b773ee5 Merge pull request #1218 from pomidornijfrukt/azure/5.4-mini-nano
Add GPT-5.4 Mini and Nano models for Azure providers
2026-03-18 10:31:13 -05:00
Aiden Cline 98559071f0 Merge pull request #1217 from cgilly2fast/dev
chore(firmware): update base url and docs url
2026-03-18 10:30:44 -05:00
eCube-cachy 0660308816 add: GPT-5.4 Mini and Nano model configurations for Azure providers 2026-03-18 15:17:52 +02:00
Jack 380f9dd8eb Merge pull request #1216 from no1wudi/dev
Add MiniMax M2.7 and M2.7-highspeed models to 4 official providers
2026-03-18 16:29:59 +08:00
Jack 1cfdab1b18 update MiniMax-M2.7 cache_read to 0.06 2026-03-18 16:27:53 +08:00
Colby Gilbert 75a981f957 chore(firmware): update base url and docs url 2026-03-18 00:41:05 -07:00
Huang Qi 7fadbcadc8 Add MiniMax M2.7 and M2.7-highspeed models to 4 official providers 2026-03-18 15:21:06 +08:00
Frank 38f9092292 update zen models 2026-03-18 02:30:18 -04:00
Aiden Cline 92149b9eaa rm nonexistant github model 2026-03-17 21:41:51 -05:00
Aiden Cline b614f0e69c Merge pull request #1214 from luisrudge/dev
Add GPT-5.4 mini and nano to GitHub Copilot provider
2026-03-17 20:13:46 -05:00
Luís Rudge 67d6dac5c5 Add GPT-5.4 mini and nano to GitHub Copilot provider 2026-03-17 18:44:38 -06:00
Aiden Cline 7d3cc61a48 Merge pull request #1207 from PedroACosta/feat/add-dinference-provider
feat(providers): add dinference provider
2026-03-17 14:51:31 -05:00
Aiden Cline f02ea6c4d2 Merge pull request #1115 from skywalker512/feat/add-tencent-coding-plan
feat: add Tencent Coding Plan provider
2026-03-17 14:51:19 -05:00
Aiden Cline 0cb50eeece Merge pull request #1208 from scwgoire/march-update
Scaleway 26-03 model updates
2026-03-17 14:48:12 -05:00
Aiden Cline 878311d2e0 Merge pull request #1210 from dm-cohere/dm/fix-update-cohere-model-capabilities
fix(models): update cohere model capabilities
2026-03-17 14:32:27 -05:00
Aiden Cline a0e89f65d6 Merge pull request #1206 from 0b1000/dev
Rename minimax-m2.5.toml to MiniMax-M2.5.toml
2026-03-17 14:32:19 -05:00
Aiden Cline 74099b7c9c Merge pull request #1213 from smrdotgg/add-openai-gpt-5-4-mini-and-nano
Add OpenAI GPT-5.4 mini and nano
2026-03-17 14:31:24 -05:00
Aiden Cline ec522435c3 Merge pull request #1211 from sylviezhang37/update-vercel-models-20260317-1807
Update Vercel models
2026-03-17 14:30:24 -05:00
smr d839cd37d4 Add OpenAI GPT-5.4 mini and nano
Capture the newly released mini and nano model metadata so models.dev reflects OpenAI's latest GPT-5.4 lineup with current pricing, limits, and knowledge cutoff.
2026-03-17 22:09:13 +03:00
github-actions[bot] ecb6ef7f93 chore(vercel): update Vercel model definitions
Auto-generated by weekly workflow from Vercel AI Gateway API.

Co-Authored-By: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-03-17 18:07:06 +00:00
Deirdre Meehan 8b4d341054 fix: cohere models on non-cohere providers 2026-03-17 16:51:24 +00:00
Deirdre Meehan 62f4a28308 fix: cohere provider models 2026-03-17 16:44:55 +00:00
Pedro 2fb8ef0dc8 feat(providers): add dinference provider 2026-03-17 14:13:36 +01:00
Gregoire de Turckheim 96968e2bf8 feat: Scaleway 26-03 model updates 2026-03-17 12:15:52 +01:00
0b1000 ee9d7879ce Rename minimax-m2.5.toml to MiniMax-M2.5.toml 2026-03-17 14:50:35 +08:00
Frank 71283512a6 update zen models 2026-03-17 02:21:13 -04:00
Frank cd4afd7e7c update zen models 2026-03-17 02:19:17 -04:00
Aiden Cline 1239d0190b Merge pull request #1204 from cyberofficial/vultr
VULTR: Updated Vultr model pricing to reflect current serverless inference rates
2026-03-16 16:10:39 -05:00
Aiden Cline 491bf6ccba Merge pull request #1202 from RaviTharuma/fix/chutes-pricing-update-2026-03
fix(chutes): update pricing and limits from live API
2026-03-16 16:10:25 -05:00
Cyber Official c993d0c121 Updated Vultr model pricing to reflect current serverless inference rates
Updated Vultr model pricing to reflect current serverless inference rates

This commit updates the cost configuration for all Vultr models to align with their latest pricing tiers:

**Cost Reductions:**
- DeepSeek-R1-Distill-Qwen-32B: Input $0.55→$0.30, Output $2.75→$0.30 (73% reduction)
- NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4: Input $0.55→$0.20, Output $2.75→$0.80 (64% input, 71% output reduction)
- Qwen2.5-Coder-32B-Instruct: Input $0.55→$0.20, Output $2.75→$0.60 (64% input, 78% output reduction)
- gpt-oss-120b: Input $0.55→$0.15, Output $2.75→$0.60 (73% input, 78% output reduction)
- MiniMax-M2.5: Input $0.55→$0.30, Output $2.75→$1.20 (45% input, 56% output reduction)

**Cost Adjustments:**
- DeepSeek-R1-Distill-Llama-70B: Input $0.55→$2.00, Output $2.75→$2.00 (significant increase)
- DeepSeek-V3.2: Output $2.75→$1.65 (40% reduction)
- Llama-3.1-Nemotron-Ultra-253B-v1: Output $2.75→$1.80 (35% reduction)
- GLM-5-FP8: Input $0.55→$0.85, Output $2.75→$3.10 (55% input, 13% output increase)
2026-03-16 13:58:28 -04:00
Aiden Cline e55c39a83d Merge pull request #1141 from sk0x0y/feature/nanogpt-confirmed-suffix2-fixes
fix(nano-gpt): rename confirmed 2-suffix model ids
2026-03-16 10:57:44 -05:00
Aiden Cline 6dea000e25 Merge pull request #1148 from sk0x0y/feature/nanogpt-bundled-confirmed-suffix2-fixes
fix(nano-gpt): rename bundled confirmed 2-suffix model ids
2026-03-16 10:57:30 -05:00
Aiden Cline 3f7a757b3f Merge pull request #1142 from sk0x0y/feature/nanogpt-more-confirmed-suffix2-fixes
fix(nano-gpt): rename more confirmed 2-suffix model ids
2026-03-16 10:56:36 -05:00
Aiden Cline c693fd71e2 Merge pull request #1194 from cyberofficial/vultr
Update Vultr model list with 10 new models and updated pricing
2026-03-16 10:55:19 -05:00
Aiden Cline 54e04e288a Merge pull request #1198 from amritbanerjee/add-glm-5-turbo
Add GLM-5-Turbo model support
2026-03-16 10:47:08 -05:00
Aiden Cline 462a179eee Merge pull request #1203 from jerome-benoit/fix/sap-ai-core-model-specs
fix(sap-ai-core): align model specs with official sources
2026-03-16 10:46:40 -05:00
Aiden Cline 95db59034d Merge pull request #1201 from dpuyosa/venice-new-models
Venice: Add new provider models
2026-03-16 10:45:59 -05:00
Aiden Cline 74dcc74e32 Merge pull request #1200 from dpuyosa/venice/pricing-update
Venice: Update model pricing for 7 models
2026-03-16 10:45:47 -05:00
Jérôme Benoit 57975f5f25 fix(sap-ai-core): align model specs with official sources 2026-03-16 13:59:06 +01:00
Ravi Tharuma ad7b063747 fix(chutes): update pricing and limits from live API
Synced 6 Chutes model definitions against the live API at
https://llm.chutes.ai/v1/models (queried 2026-03-16).

Models updated:
- deepseek-ai/DeepSeek-V3.2-TEE: cost 0.25/0.38→0.28/0.42, cache 0.125→0.14, context 163840→131072
- zai-org/GLM-5-TEE: cost 0.75/2.5→0.95/3.15, added cache_read 0.475
- zai-org/GLM-4.6-TEE: cost 0.35/1.5→0.4/1.7, added cache_read 0.2
- zai-org/GLM-4.6V: added cache_read 0.15
- MiniMaxAI/MiniMax-M2.5-TEE: cost 0.15/0.6→0.3/1.1, added cache_read 0.15
- Qwen/Qwen3.5-397B-A17B-TEE: cost 0.3/1.2→0.39/2.34, cache 0.15→0.195
2026-03-16 11:42:29 +01:00
dpuyosa f76e9f0551 [venice] Add new provider models
- Add mistral-small-3.2-24b-instruct, qwen3-5-9b, venice-uncensored-role-play, zai-org-glm-4.6
2026-03-16 09:37:41 +01:00
dpuyosa d70a49b36f [venice] Update model pricing for 7 models
- Remove context_over_200k pricing from Claude models
- Update Grok cache_read pricing from 0.5 to 0.25
- Update Kimi, MiniMax input/output pricing
2026-03-16 09:05:37 +01:00
amrit 3487135f9f Add GLM-5-Turbo model support 2026-03-16 12:14:50 +11:00
Aiden Cline 458a66c766 Merge pull request #1197 from kesku/update-perplexity-agent-models
Update Perplexity Agent API models
2026-03-15 10:59:23 -05:00
Frank d3a84dc7ec update zen models 2026-03-15 10:59:52 -04:00
Kesku ae61b25583 update perplexity-agent: add gpt-5.4 & nemotron, remove gemini-3-pro 2026-03-15 03:46:50 +00:00
Aiden Cline 74be576eda Merge pull request #1178 from Sewer56/change-synthetic-endpoint
Add OpenAI and Anthropic compatible endpoints
2026-03-14 20:55:30 -05:00
Aiden Cline 164df2cda0 Merge pull request #1191 from Alcatraz-Zhang/update/kilo-models
Sync Kilo model definitions with latest gateway catalog
2026-03-14 20:54:45 -05:00
Cyber Official 2cd7908369 Update Vultr model list with 10 new models and updated pricing
- Updated pricing to $0.55/M input tokens, $2.75/M output tokens
- Updated context limits to safe floor values from official testing
- Added accurate output token limits from official model documentation
- Added 5 new models: MiniMax M2.5, DeepSeek V3.2, GLM-5 FP8, Llama 3.1 Nemotron Ultra 253B, NVIDIA Nemotron 3 Super 120B A12B NVFP4
- Updated existing models: DeepSeek R1 Distill variants, GPT OSS 120B, Kimi K2.5, Qwen2.5 Coder 32B

Model specifications:
- MiniMax M2.5: 196K context, 4,096 output
- Qwen2.5-Coder-32B: 15K context, 256 output (notable low default)
- DeepSeek R1 Distill Llama 70B: 130K context, 4,096 output
- DeepSeek R1 Distill Qwen 32B: 130K context, 4,096 output
- DeepSeek V3.2: 163K context, 4,096 output
- Kimi K2.5: 261K context, 32,768 output (high output limit)
- GPT OSS 120B: 130K context, 8,192 output
- GLM-5 FP8: 202K context, 131,072 output (exceptionally high)
- Llama 3.1 Nemotron Ultra 253B: 32K context, 4,096 output
- NVIDIA Nemotron 3 Super 120B A12B NVFP4: 260K context, 8,192 output

All models set to text-only (no vision support) as confirmed.
2026-03-14 19:47:01 -04:00
Alcatraz-Zhang cc667340f5 Sync Kilo model definitions with latest gateway catalog
Refresh the Kilo provider catalog so models.dev matches the current gateway inventory, pricing, and availability.
2026-03-15 04:35:38 +08:00
Sewer56 f2cfc1435d Changed: Synthetic to use newer openai endpoint 2026-03-14 17:09:44 +00:00
Aiden Cline 35bb8cca47 Merge pull request #1172 from bigfluffycookie/add-deepinfra-llama-models
Add deepinfra llama models
2026-03-14 10:55:13 -05:00
Aiden Cline 3468a410e1 Merge pull request #1177 from ar27111994/dev
Add Grok 4.1 Fast configurations for reasoning and non-reasoning
2026-03-14 10:54:57 -05:00
Aiden Cline b1b5e3c5cd Merge pull request #1174 from dacbd/patch-1
fix(wandb): fix k2.5 settings
2026-03-14 10:54:35 -05:00
Aiden Cline 97f03ec672 Merge pull request #1175 from dacbd/patch-2
chore(docs): add note for manual testing with opencode
2026-03-14 10:54:22 -05:00
BigFluffyCookie 9b516924aa Add limit output for llama models 2026-03-14 11:49:57 +01:00
Ahmed Rehan 929a39600b feat(models): add Grok 4.1 Fast (Reasoning and Non-Reasoning) configurations 2026-03-14 14:27:24 +05:00
Daniel Barnes a87d8bb8cc chore(docs): add note for manual testing with opencode 2026-03-14 13:42:57 +09:00
Daniel Barnes 574139eb49 fix(wandb): fix k2.5 settings 2026-03-14 13:07:16 +09:00
Aiden Cline 1e3bc38b31 Merge pull request #1137 from mcowger/mcowger/correct-gemini-flash-lite-pricing
Fix incorrect pricing for gemini-3.1-flash-lite-preview
2026-03-13 18:41:41 -05:00
Aiden Cline 8916fe9874 Merge pull request #1171 from stephenkuhn214/dev
fix(amazon-bedrock): Remove deprecated and add missing models
2026-03-13 18:26:25 -05:00
BigFluffyCookie 5d956b41a6 Rename llama models to remove "Meta" prefix 2026-03-13 23:15:27 +01:00
BigFluffyCookie 42a7a14f69 Add Meta Llama models to DeepInfra provider 2026-03-13 22:53:39 +01:00
Stephen Kuhn f24ee000d7 fix(amazon-bedrock): update and add models
- Remove 19 deprecated/EOL models
- Add 7 new models: DeepSeek V3.2, Llama 3.1 405B, Magistral Small 1.2, Ministral 3 3B, Mistral Large 3, Pixtral Large, NVIDIA Nemotron Nano 3 30B
- Fix Devstral 2 123B: correct name, family, and open_weights
- Set accurate Bedrock launch dates for all new models
2026-03-13 16:02:04 -04:00
Aiden Cline 7196b1fb2c Merge pull request #1170 from anomalyco/revert-1166-fix/update-gpt53-codex-spark-preview
Revert "fix(openai): rename gpt-5.3-codex-spark to gpt-5.3-codex-spark-preview"
2026-03-13 14:31:33 -05:00
Aiden Cline f6c0d5a29d Revert "fix(openai): rename gpt-5.3-codex-spark to gpt-5.3-codex-spark-preview" 2026-03-13 14:30:58 -05:00
Aiden Cline ee63449aa5 sonnet 4.6 and opus 4.6 1M context 2026-03-13 14:27:55 -05:00
Aiden Cline 92aa44ec00 Merge pull request #1166 from rluisr/fix/update-gpt53-codex-spark-preview
fix(openai): rename gpt-5.3-codex-spark to gpt-5.3-codex-spark-preview
2026-03-13 14:18:41 -05:00
Aiden Cline 477284535c Rename model from 'GPT-5.3 Codex Spark Preview' to 'GPT-5.3 Codex Spark' 2026-03-13 14:17:44 -05:00
Aiden Cline 304233bdda Merge pull request #1169 from mdrxy/mdrxy/anthropic-token-limits
Update Claude 4.6 context/pricing
2026-03-13 14:13:40 -05:00
Aiden Cline 25d782ee2c Reduce context limit from 1,000,000 to 200,000 2026-03-13 14:13:30 -05:00
Aiden Cline 0f63393d51 Update context limit in claude-opus-4-6.toml 2026-03-13 14:12:56 -05:00
rluisr e780eefce2 fix(openai): rename gpt-5.3-codex-spark to gpt-5.3-codex-spark-preview
The OpenAI API expects model ID 'gpt-5.3-codex-spark-preview', not
'gpt-5.3-codex-spark'. Rename model files in both openai and opencode
providers so the generated model ID matches the actual API.
2026-03-14 03:59:03 +09:00
Aiden Cline a79585fa83 Merge pull request #1163 from micuintus/feature/Kimi2.5-fast
feat(nebius): add Kimi-K2.5-fast model
2026-03-13 13:14:38 -05:00
Aiden Cline 00801f74f2 Merge pull request #1164 from butyess/dev
Openrouter models: gemini 3.1 flash lite preview, grok 4.20 beta models.
2026-03-13 13:14:22 -05:00
Aiden Cline 185f6731ee Merge pull request #1162 from dpuyosa/feature/venice-grok-4-20-beta
Venice: Add Grok 4.20 Beta models
2026-03-13 12:53:28 -05:00
Aiden Cline d291b0575c Merge pull request #1167 from sylviezhang37/update-vercel-models-20260313-1639
Update Vercel models
2026-03-13 12:53:11 -05:00
Mason Daugherty 382d9f3e7d Update Claude 4.6 context/pricing 2026-03-13 13:53:04 -04:00
Aiden Cline e64f5fe963 Merge pull request #1168 from mdrxy/mdrxy/update-baseten
Update Baseten models
2026-03-13 12:51:56 -05:00
Mason Daugherty ea57ddfe7e Update Baseten models 2026-03-13 13:48:41 -04:00
github-actions[bot] 29463d7fa8 chore(vercel): update Vercel model definitions
Auto-generated by weekly workflow from Vercel AI Gateway API.

Co-Authored-By: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-03-13 16:39:30 +00:00
Jack bcc8db49ee Merge pull request #1165 from anomalyco/chore/openrouter-alpha-reasoning-details-20260313
feat(openrouter): add interleaved reasoning details for alpha models
2026-03-13 22:25:25 +08:00
Jack c8521d70f3 feat(openrouter): add interleaved reasoning details for alpha models 2026-03-13 22:20:54 +08:00
Federico Masi 490cd249e4 Openrouter models: gemini 3.1 flash lite preview, grok 4.20 beta models. 2026-03-13 15:11:12 +01:00
Michael Voigt dbc636f5f3 feat(nebius): add Kimi-K2.5-fast model 2026-03-13 12:35:13 +01:00
Michael Voigt 9a32f671a1 fix(nebius): lowercase model ID for Nemotron-3-Super-120B-A12B
The filename must match the API casing (lowercase) to avoid 'model does not exist' errors.
2026-03-13 12:35:08 +01:00
dpuyosa 856d925eda [venice] Add Grok 4.20 Beta models
- Add Grok 4.20 Beta model configuration (2M context, 128K output)
- Add Grok 4.20 Multi-Agent Beta model configuration
2026-03-13 10:48:38 +01:00
Aiden Cline 066a425917 Merge pull request #1158 from micuintus/feature/Nebius_Nemotron-3-Super-120b-a12b
feat(nebius): Add support for Nemotron-3-Super-120B-A12B
2026-03-12 22:20:20 -05:00
Aiden Cline 6df7f20cdc Merge pull request #1156 from dsingal0/dev
added nemotron super on baseten
2026-03-12 22:20:06 -05:00
Aiden Cline 78bb47b90e Merge pull request #1151 from dacbd/dacbd
fix(wandb): update models
2026-03-12 22:19:43 -05:00
Aiden Cline c121d86419 Merge pull request #1160 from kreatoo/dev
feat: add zai-org/glm-4.7 and zai-org/glm-4.7-flash to NanoGPT
2026-03-12 22:11:18 -05:00
Aiden Cline ab148eeb14 Merge pull request #1161 from Grin1024/dev
Add Claude Opus 4.6 and Sonnet 4.6 models to RequestY provider
2026-03-12 22:11:07 -05:00
lihui 49d196d326 Add Claude Opus 4.6 and Sonnet 4.6 models to RequestY provider 2026-03-13 09:00:54 +08:00
Kreato 8899b390ef feat: add zai-org/glm-4.7 and zai-org/glm-4.7-flash to NanoGPT 2026-03-13 00:27:09 +03:00
Michael Voigt 5217f62ddf fix(nebius): Follow context updates for Kimi 2.5 and GLM-5 2026-03-12 20:22:48 +01:00
Michael Voigt 55eaff9af1 feat(nebius): Add support for Nemotron-3-Super-120B-A12B 2026-03-12 20:22:21 +01:00
Dhruv Singal 7557c06ac0 update output length 2026-03-12 09:41:25 -07:00
Dhruv Singal e85d820121 fix input output 2026-03-12 08:29:01 -07:00
Dhruv Singal 499d3a39ef remove cache pricing 2026-03-12 08:21:22 -07:00
Dhruv Singal b9b38d6e33 added nemotron super on baseten 2026-03-12 08:18:46 -07:00
Aiden Cline ca24ac14fa Merge pull request #1153 from dpuyosa/dev
Venice: Update model output token limits
2026-03-12 10:08:46 -05:00
Aiden Cline 822546fc67 Merge pull request #1155 from spiffytech/dev
Add Ollama Cloud support for Nemotron 3 Super
2026-03-12 10:08:31 -05:00
Aiden Cline 4555195b71 Merge pull request #1152 from v1gnesh/dev
Update grok-4.20 model defs
2026-03-12 10:08:15 -05:00
spiffytech 5eae8effc6 Added Ollama Cloud support for Nemotron 3 Super 2026-03-12 09:28:47 -04:00
dpuyosa c1801aef87 [venice] Normalize model output token limits
- Update output limits to standard values across all models
2026-03-12 10:08:39 +01:00
v1gnesh 5e6464b272 Update grok-4.20-beta-reasoning 2026-03-12 10:27:40 +05:30
v1gnesh e1a4f23332 Update grok-4.20-beta-non-reasoning 2026-03-12 10:26:03 +05:30
v1gnesh 753e1f9f0c grok-multi-agent-beta update 2026-03-12 10:23:57 +05:30
Daniel Barnes 123ecd2ba5 docs url 2026-03-12 13:27:56 +09:00
Daniel Barnes f15cda9fcb remove old 2026-03-12 13:26:08 +09:00
Daniel Barnes 0205debbd3 fix values 2026-03-12 13:22:29 +09:00
Daniel Barnes 0059766509 number formating 2026-03-12 13:17:22 +09:00
Daniel Barnes be81b02916 additional model files 2026-03-12 13:02:17 +09:00
Daniel Barnes 2dab141166 initial script & model updates 2026-03-12 13:01:35 +09:00
Aiden Cline 45aa49af25 tweak: azure kimi k2.5 2026-03-11 22:35:20 -05:00
Aiden Cline 781fad3ad4 Merge pull request #1150 from cau1k/5.4-family
feat(azure): add 5.4/pro families
2026-03-11 22:14:08 -05:00
cau1k 99d2ffcfdd feat(azure): add 5.4/pro families 2026-03-11 20:59:11 -04:00
Aiden Cline 381d7cc19d Merge pull request #1149 from ariane-emory/fear/add-march-or-stealth-models
Add OpenRouter stealth models: Hunter Alpha and Healer Alpha
2026-03-11 18:07:50 -05:00
Ariane Emory 7482e22458 Fix family field to use 'alpha' for stealth models 2026-03-11 18:49:32 -04:00
Ariane Emory f5e6a402e6 Add OpenRouter stealth models: Hunter Alpha and Healer Alpha 2026-03-11 18:41:58 -04:00
Aiden Cline 9265852852 tweak: adjust some gh limits to align better w/ api 2026-03-11 15:23:44 -05:00
Aiden Cline dc98a32996 Merge pull request #1018 from Sewer56/add-synthetic-missing-models
Update synthetic.new models: promote MiniMax-M2.5, add GLM-4.7-Flash
2026-03-11 14:55:50 -05:00
Aiden Cline 56c39ae0f6 Merge pull request #1140 from sk0x0y/feature/nanogpt-thudm-id-fixes
fix(nano-gpt): rename THUDM 2 ids to canonical THUDM ids
2026-03-11 14:55:07 -05:00
Aiden Cline b1f43a7595 Merge pull request #1147 from msadiks/fix/alibaba-coding-minimax
fix: alibaba-coding-plan MiniMax-M2.5 context window
2026-03-11 14:54:37 -05:00
Matt Cowger fed8bcae19 Merge branch 'dev' into mcowger/correct-gemini-flash-lite-pricing 2026-03-11 12:23:42 -07:00
sk0x0y fb95150d02 fix(nano-gpt): rename VongolaChouko model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:20:39 +09:00
sk0x0y a7c9a240b4 fix(nano-gpt): rename Steelskull model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:20:39 +09:00
sk0x0y 6432a4a3e6 fix(nano-gpt): rename Sao10K model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:20:38 +09:00
sk0x0y f2e4a249fe fix(nano-gpt): rename NeverSleep model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:20:38 +09:00
sk0x0y 8667a6eed8 fix(nano-gpt): rename MarinaraSpaghetti model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:19:56 +09:00
sk0x0y 429554397a fix(nano-gpt): rename LatitudeGames model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:19:56 +09:00
sk0x0y a64e6ad0ac fix(nano-gpt): rename LLM360 model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:19:56 +09:00
sk0x0y d68d79888c fix(nano-gpt): rename Infermatic model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:19:56 +09:00
sk0x0y 6c52905c6a fix(nano-gpt): rename Gryphe model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:19:55 +09:00
sk0x0y 62410b8f26 fix(nano-gpt): rename GalrionSoftworks model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:19:55 +09:00
sk0x0y 50ce68ccab fix(nano-gpt): rename Envoid model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:19:20 +09:00
sk0x0y d1c6a6b873 fix(nano-gpt): rename EVA-UNIT-01 model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:19:20 +09:00
Frank 7193b068a5 update zen models 2026-03-11 13:52:50 -04:00
Sadik 79a8a06bd7 fix MiniMax-M2.5 context window 2026-03-11 20:50:33 +03:00
Aiden Cline b60c03e11c Merge pull request #1139 from zainhas/dev
[Together AI] add prompt caching pricing for MiniMax m2.5
2026-03-11 12:31:56 -05:00
Aiden Cline 15cf98d57b Merge pull request #1146 from gotjoshua/patch-1
Rename step-3-5-flash.toml to step-3.5-flash.toml
2026-03-11 12:31:39 -05:00
Aiden Cline b2ee6c407b Merge pull request #1144 from micuintus/feature/update-nebius-changes
Feat: update Nebius changes
2026-03-11 12:31:29 -05:00
gotjoshua 96a14a06e7 Rename step-3-5-flash.toml to step-3.5-flash.toml
on nvidia it is 3.5 not 3-5
2026-03-11 11:41:36 +00:00
Michael Voigt adc358606d fix(nebius): update model context limits per API 2026-03-11 11:33:14 +01:00
Michael Voigt 63d52adf6f feat(nebius): add GLM-5 model 2026-03-11 11:33:14 +01:00
sk0x0y 9a31387766 fix(nano-gpt): rename Salesforce model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 18:17:41 +09:00
sk0x0y 735157b837 fix(nano-gpt): rename ReadyArt model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 18:17:41 +09:00
sk0x0y d75b46fb37 fix(nano-gpt): rename Doctor-Shotgun model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 18:17:41 +09:00
sk0x0y cc555f8482 fix(nano-gpt): rename CrucibleLab model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 18:17:13 +09:00
sk0x0y 7fbbcf2b49 fix(nano-gpt): rename MiniMaxAI model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 16:04:28 +09:00
sk0x0y 14c8ec8ca5 fix(nano-gpt): rename Tongyi-Zhiwen model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 16:04:28 +09:00
sk0x0y 72568bbdb3 fix(nano-gpt): rename Alibaba-NLP model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 16:03:57 +09:00
sk0x0y c2225b715f fix(nano-gpt): rename THUDM GLM-Z1 rumination id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 15:49:20 +09:00
sk0x0y 7ce25e3742 fix(nano-gpt): rename THUDM GLM-Z1 model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 15:49:20 +09:00
sk0x0y 427868604b fix(nano-gpt): rename THUDM GLM-4 model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 15:49:20 +09:00
Zain Hasan 247cd801a8 add prompt caching pricing for MiniMax m2.5 2026-03-10 22:54:42 -07:00
Aiden Cline 1aa2ee22b1 Merge pull request #1134 from sk0x0y/feature/nanogpt-catalog-fixes
fix(nano-gpt): correct TEE path ids and add missing canonical entries
2026-03-10 22:02:52 -05:00
Aiden Cline 0f57233eff Merge pull request #1105 from sylviezhang37/add-vercel-input-context-and-new-models
feat(vercel): add input context calculation + new models
2026-03-10 22:01:52 -05:00
Aiden Cline 73a78eebfc Merge pull request #1138 from mugnimaestra/feat/add-glm-5-turbo-chutes
feat: add GLM-5-Turbo to Chutes provider listings
2026-03-10 22:01:08 -05:00
Sylvie Zhang 3a6789b819 Merge branch 'dev' into add-vercel-input-context-and-new-models 2026-03-10 17:44:14 -07:00
Sylvie Zhang f7c505e140 remove context from gemini models 2026-03-10 17:43:08 -07:00
Sylvie Zhang 6bb36806d6 only calc input context for openai models 2026-03-10 17:40:46 -07:00
Sylvie Zhang 20a404eb88 revert non openai changes 2026-03-10 17:38:46 -07:00
Muhammad Mugni Hadi 65ecb5cd4a feat: add GLM-5-Turbo to Chutes provider listings 2026-03-11 05:26:11 +07:00
Matt Cowger 56062a9129 Fix incorrect pricing 2026-03-10 14:57:44 -07:00
Aiden Cline d3d9c580d4 Merge pull request #1135 from gitpush-gitpaid/fix/gpt-5-4-pdf-input-modalities
Added PDF to input modalities for GPT-5.4
2026-03-10 13:53:42 -05:00
gitpush-gitpaid ef98d8a9cb Updated GPT-5.4 PDF input modalities 2026-03-10 13:59:29 -04:00
sk0x0y 9d17752b88 fix(nano-gpt): add missing GLM 5 thinking model
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 00:56:22 +09:00
sk0x0y b5a838fe8b fix(nano-gpt): add missing TEE qwen3.5 model
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 00:56:22 +09:00
sk0x0y aa1ac39ee6 fix(nano-gpt): rename TEE gemma and minimax ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 00:56:22 +09:00
sk0x0y 4bc17ccf96 fix(nano-gpt): rename TEE oss and llama ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 00:56:22 +09:00
sk0x0y 08c1899bfe fix(nano-gpt): rename TEE deepseek model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 00:56:02 +09:00
sk0x0y ad50e4a5ed fix(nano-gpt): rename TEE qwen model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 00:56:02 +09:00
sk0x0y 730915a123 fix(nano-gpt): rename TEE kimi model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 00:56:02 +09:00
sk0x0y 6f12d18cb8 fix(nano-gpt): rename TEE glm model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 00:56:02 +09:00
Aiden Cline bd8774db99 Merge pull request #1132 from sk0x0y/feature/nanogpt-model-sync
feat(nano-gpt): add text and image models
2026-03-10 10:31:50 -05:00
Aiden Cline 88fbea52a4 Merge pull request #1133 from anomalyco/fix-model
fix: bedrock devstral
2026-03-10 10:31:08 -05:00
sk0x0y 898b3c18b7 feat(nano-gpt): add image models
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-10 22:13:48 +09:00
sk0x0y 6316e543ef feat(nano-gpt): add text models
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-10 22:13:48 +09:00
Sewer56 7a02946620 Update synthetic models: promote MiniMax-M2.5, add GLM-4.7-Flash, remove deprecated Qwen3.5 2026-03-08 22:56:31 +00:00
skywalker512 236af40da3 feat: add Tencent Coding Plan provider
Add support for Tencent Coding Plan with 8 models:
- Auto (tc-code-latest)
- Hunyuan 2.0 Instruct
- Hunyuan 2.0 Think
- Hunyuan-T1
- Hunyuan-TurboS
- MiniMax-M2.5
- Kimi-K2.5
- GLM-5

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-08 15:34:36 +08:00
Sylvie Zhang 26465319d6 Merge branch 'dev' into add-vercel-input-context-and-new-models 2026-03-07 11:39:53 -08:00
Sylvie Zhang 7a11ef241d update more models 2026-03-06 08:24:49 -08:00
Sylvie Zhang 145862315d add input calculation + new models 2026-03-06 08:07:11 -08:00
Sewer56 0428299773 Added: Qwen3.5-397B natively supports image, MM2.5 No Image as it was a mistake. 2026-02-25 08:10:42 +00:00
Sewer56 eee3303df0 Add missing synthetic.new models
Add configuration for hf:Qwen/Qwen3.5-397B-A17B and hf:MiniMaxAI/MiniMax-M2.5
to the synthetic provider, based on API specs from synthetic.new.

Note: API reports image support but these models may not natively support
images (likely rerouted/proxied through vision-capable infrastructure).
2026-02-24 09:19:57 +00:00
massaindustries 6ecd9ec509 add qwen-next-coder-2 2026-02-05 09:11:08 +00:00
massaindustries bb42d0b855 add qwen-next-coder 2026-02-05 09:09:30 +00:00
massaindustries a46966efef add-regolo-02 2026-01-29 15:14:02 +00:00
massaindustries 19ace4436f add-regolo-01 2026-01-29 12:36:44 +00:00
Luca Steeb 92269282eb fix: use correct family for gemma and gpt-oss models
- Gemma models now use "gemma" family instead of "gemini"
- GPT OSS models now use "gpt-oss" family instead of "gpt"

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-23 01:36:29 +00:00
Luca Steeb 6b9b340fbc fix: map llmgateway family to auto
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-23 01:28:38 +00:00
Luca Steeb e8a6793654 fix: use valid models.dev family enum values
Maps internal family names to valid models.dev families:
- moonshot → kimi
- bytedance → seed
- zai → glm
- nvidia → nemotron

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-23 01:27:15 +00:00
Luca Steeb b22ff136a8 fix: add required output limit to all models
models.dev schema requires limit.output field.
Defaults to 16384 when not specified.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-23 01:25:53 +00:00
Luca Steeb e28d4e387f chore: trigger CI 2026-01-23 01:21:06 +00:00
Luca Steeb 09b5dd4d84 refactor: remove scripts/ dir, link to repo script
Removes empty generate.ts file and scripts/ directory.
README now links to llmgateway repo for regeneration.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-23 01:18:53 +00:00
Luca Steeb 24e575a86e refactor: flatten model structure to models/ directory
Removes provider subdirectories, exports all models directly
to models/ folder for simpler structure.

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-23 01:15:27 +00:00
Luca Steeb 3549abec36 feat: add LLM Gateway provider with 153 models
Add LLM Gateway (llmgateway.io) as a new provider with all supported models
organized by upstream provider subdirectory.

LLM Gateway is an OpenAI-compatible API gateway that provides unified
access to 40+ LLM providers through a single API endpoint.

Directory structure:
  providers/llmgateway/
  ├── provider.toml
  ├── README.md
  ├── scripts/
  │   └── generate.ts
  └── models/
      ├── anthropic/ (16 models)
      ├── openai/ (28 models)
      ├── google/ (19 models)
      ├── zai/ (17 models - GLM, CogView)
      ├── alibaba/ (27 models - Qwen)
      ├── meta/ (12 models - Llama)
      ├── xai/ (9 models - Grok)
      ├── deepseek/ (5 models)
      ├── bytedance/ (6 models - Seed)
      ├── moonshot/ (4 models - Kimi)
      ├── mistral/ (3 models)
      ├── perplexity/ (3 models - Sonar)
      ├── minimax/ (1 model)
      ├── nvidia/ (1 model)
      └── llmgateway/ (2 models - auto, custom)

Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
2026-01-23 01:10:53 +00:00
2654 changed files with 33928 additions and 10086 deletions
+1
View File
@@ -35,3 +35,4 @@ jobs:
- run: bun sst deploy --stage=dev
env:
CLOUDFLARE_API_TOKEN: ${{ secrets.CLOUDFLARE_API_TOKEN }}
CLOUDFLARE_DEFAULT_ACCOUNT_ID: ${{ secrets.CLOUDFLARE_DEFAULT_ACCOUNT_ID }}
+108
View File
@@ -0,0 +1,108 @@
name: Sync Model Catalogs
on:
schedule:
- cron: "17 * * * *"
workflow_dispatch:
permissions:
contents: write
issues: write
pull-requests: write
concurrency: ${{ github.workflow }}-${{ github.ref }}
jobs:
providers:
runs-on: ubuntu-latest
outputs:
matrix: ${{ steps.providers.outputs.matrix }}
steps:
- name: Checkout code
uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5
with:
ref: dev
- name: Setup Bun
uses: oven-sh/setup-bun@f4d14e03ff726c06358e5557344e1da148b56cf7
with:
bun-version: latest
- name: Install dependencies
run: bun install
- name: List sync providers
id: providers
run: echo "matrix=$(bun models:sync --list-providers)" >> "$GITHUB_OUTPUT"
sync:
needs: providers
runs-on: ubuntu-latest
strategy:
fail-fast: false
matrix: ${{ fromJSON(needs.providers.outputs.matrix) }}
steps:
- name: Checkout code
uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5
with:
ref: dev
- name: Setup Bun
uses: oven-sh/setup-bun@f4d14e03ff726c06358e5557344e1da148b56cf7
with:
bun-version: latest
- name: Install dependencies
run: bun install
- name: Sync model catalogs
run: bun models:sync ${{ matrix.provider }}
env:
OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}
GOOGLE_API_KEY: ${{ secrets.GOOGLE_API_KEY }}
GEMINI_API_KEY: ${{ secrets.GEMINI_API_KEY }}
GOOGLE_GENERATIVE_AI_API_KEY: ${{ secrets.GOOGLE_GENERATIVE_AI_API_KEY }}
XAI_API_KEY: ${{ secrets.XAI_API_KEY }}
CLOUDFLARE_WORKERS_AI_SYNC_ACCOUNT_ID: ${{ secrets.CLOUDFLARE_WORKERS_AI_SYNC_ACCOUNT_ID }}
CLOUDFLARE_WORKERS_AI_SYNC_API_TOKEN: ${{ secrets.CLOUDFLARE_WORKERS_AI_SYNC_API_TOKEN }}
- name: Validate models
run: bun validate
- name: Create pull request
env:
GH_TOKEN: ${{ github.token }}
BRANCH: automation/sync-models-${{ matrix.provider }}
LABELS: automation,model-sync,provider:${{ matrix.provider }}
TITLE: "chore(sync): update ${{ matrix.name }} model catalog"
run: |
if [ -z "$(git status --porcelain -- providers)" ]; then
echo "No model catalog changes found."
exit 0
fi
git config user.name "github-actions[bot]"
git config user.email "41898282+github-actions[bot]@users.noreply.github.com"
git fetch --no-tags --depth=1 origin "+refs/heads/$BRANCH:refs/remotes/origin/$BRANCH" || true
git checkout -B "$BRANCH"
git add providers
git commit -m "$TITLE"
git push --force-with-lease origin "$BRANCH"
label_args=()
IFS=',' read -ra labels <<< "$LABELS"
for label in "${labels[@]}"; do
gh label create "$label" --color "0E8A16" --description "Automated model catalog sync" >/dev/null 2>&1 || true
label_args+=(--label "$label")
done
pr_number="$(gh pr list --head "$BRANCH" --base dev --json number --jq '.[0].number')"
if [ -n "$pr_number" ]; then
gh pr edit "$pr_number" --title "$TITLE" --body-file .sync/model-sync-report.md
for label in "${labels[@]}"; do
gh pr edit "$pr_number" --add-label "$label"
done
else
gh pr create --base dev --head "$BRANCH" --title "$TITLE" --body-file .sync/model-sync-report.md "${label_args[@]}"
fi
+1
View File
@@ -3,6 +3,7 @@
.idea
dist
.DS_Store
.sync/
node_modules
data/tokenspeed-monitor.sqlite
data/tokenspeed-monitor.sqlite-shm
+19 -1
View File
@@ -31,9 +31,27 @@
## Model Configuration
- Model `id` is **auto-injected** from filename (minus `.toml`) — never put `id` in TOML files
- Same model is duplicated across provider directories with no cross-referencing
- Models may reuse another model's definition via `extends` (see below); otherwise the full definition must be present in the file
- Schema uses `.strict()` — extra fields cause validation errors
### `[extends]` (inheritance between models)
- Syntax — a table at the top of the TOML:
```toml
[extends]
from = "<provider-id>/<model-id>" # required
omit = ["experimental.modes.fast"] # optional, dot-path strings
```
Example: `from = "anthropic/claude-opus-4-6"`
- Resolved at parse time in `generate()`; the final JSON output contains **no** `extends` field — it exists only to cut duplication in the TOMLs
- Merge semantics:
- Plain objects (`[cost]`, `[limit]`, `[modalities]`, `[provider]`, `[experimental]`, …) are **deep-merged**
- Arrays (e.g. `modalities.input`) and primitives are **replaced** wholesale by the child
- Any field the child omits is inherited verbatim from the base
- `omit` runs **after** the merge and deletes each dot-path from the result (used when the child needs to *remove* something the base defines, e.g. a provider-specific experimental mode). Every listed path must exist in the merged model, else an error is thrown. Ancestor tables that become empty as a result are also pruned, so `omit = ["experimental.modes.fast"]` yields no `experimental` key in the final JSON when `fast` was the only mode.
- Chains are allowed (A extends B extends C); cycles throw
- The base model must exist; `[extends.from]` pointing at a missing provider/model is an error
- The `extends` table is stripped before schema validation, so the merged result must still satisfy the strict `Model` schema
### Bedrock Naming Patterns
- Dated models: `-v1:0` suffix (`anthropic.claude-3-5-sonnet-20241022-v1:0.toml`)
- Latest/undated models: bare `-v1` (`anthropic.claude-opus-4-6-v1.toml`)
+45 -1
View File
@@ -120,6 +120,31 @@ output = ["text"] # Supported output modalities
field = "reasoning_content" # Name of the interleaved field "reasoning_content" or "reasoning_details"
```
#### 3a. Reuse an Existing Model with `extends`
For wrapper providers that mirror a model from another provider, prefer reusing the canonical model definition instead of duplicating the whole file.
Use `extends` only for non-first-party wrappers and mirrors. Do not use it inside the actual lab provider directories that act as the canonical source for a model family, for example `providers/anthropic/`, `providers/openai/`, `providers/google/`, `providers/xai/`, `providers/minimax/`, or `providers/moonshot/`.
```toml
[extends]
from = "anthropic/claude-opus-4-6"
omit = ["experimental.modes.fast"]
[provider]
npm = "@ai-sdk/anthropic"
```
Rules:
- `from` must point to another model using `<provider>/<model-id>`.
- `omit` is optional and removes fields after the inherited model and local overrides are merged.
- You can override any top-level model field locally.
- If you override a nested table like `[cost]`, `[limit]`, or `[modalities]`, include the full values needed for that table.
- `id` still comes from the filename; do not add it to the TOML.
Use `extends` when the wrapper model is materially the same as the source model and only differs by a small set of overrides or omitted fields.
#### 4. Submit a Pull Request
1. Fork this repo
@@ -136,9 +161,17 @@ There's a GitHub Action that will automatically validate your submission against
- Values are within acceptable ranges
- TOML syntax is valid
When converting existing wrapper models to `extends`, compare generated output before and after the change:
```bash
bun run compare:migrations
```
This prints a diff for each changed model TOML so you can confirm the generated JSON only changed where you intended.
### Schema Reference
Models must conform to the following schema, as defined in `app/schemas.ts`.
Models must conform to the following schema, as defined in `packages/core/src/schema.ts`.
**Provider Schema:**
@@ -199,6 +232,17 @@ $ bun run dev
And it'll open the frontend at http://localhost:3000
### Manual testing with opencode
You can manually check provider changes with opencode by:
```bash
$ bun install
$ cd packages/web
$ bun run build
$ OPENCODE_MODELS_PATH="dist/_api.json" opencode
```
### Questions?
Open an issue if you need help or have questions about contributing.
+10 -12
View File
@@ -5,7 +5,7 @@
"": {
"name": "models.dev",
"dependencies": {
"@cloudflare/workers-types": "^4.20250801.0",
"@cloudflare/workers-types": "^4.20260424.1",
"sst": "3.17.23",
},
},
@@ -13,6 +13,7 @@
"name": "models.dev",
"version": "0.0.0",
"dependencies": {
"remeda": "^2.33.7",
"zod": "catalog:",
},
"devDependencies": {
@@ -31,6 +32,7 @@
"packages/web": {
"name": "@models.dev/web",
"dependencies": {
"@tanstack/virtual-core": "^3.14.0",
"hono": "^4.8.0",
"models.dev": "workspace:*",
},
@@ -48,7 +50,7 @@
"zod": "3.24.2",
},
"packages": {
"@cloudflare/workers-types": ["@cloudflare/workers-types@4.20250801.0", "", {}, "sha512-BQmMdoOGClY23TesgkR1PeGrPvPsSFD/zW7pDzWZHkOEsqkPk2A91h52bP8GbtKYTl1vdaYjQgJlGsP6Ih4G0w=="],
"@cloudflare/workers-types": ["@cloudflare/workers-types@4.20260424.1", "", {}, "sha512-0DLJ9yEk1KKzPbqop80Gw/P1wkKKzawmipULiJWdBXIBCoMvE0OVWms3IrL/Q/G7tfmPop9yF4XlZ69k9JLYng=="],
"@modelcontextprotocol/sdk": ["@modelcontextprotocol/sdk@1.6.1", "", { "dependencies": { "content-type": "^1.0.5", "cors": "^2.8.5", "eventsource": "^3.0.2", "express": "^5.0.1", "express-rate-limit": "^7.5.0", "pkce-challenge": "^4.1.0", "raw-body": "^3.0.0", "zod": "^3.23.8", "zod-to-json-schema": "^3.24.1" } }, "sha512-oxzMzYCkZHMntzuyerehK3fV6A2Kwh5BD6CGEJSVDU2QNEhfLOptf2X7esQgaHZXHZY0oHmMsOtIDLP71UJXgA=="],
@@ -56,9 +58,11 @@
"@models.dev/web": ["@models.dev/web@workspace:packages/web"],
"@tanstack/virtual-core": ["@tanstack/virtual-core@3.14.0", "", {}, "sha512-JLANqGy/D6k4Ujmh8Tr25lGimuOXNiaVyXaCAZS0W+1390sADdGnyUdSWNIfd49gebtIxGMij4IktRVzrdr12Q=="],
"@tsconfig/bun": ["@tsconfig/bun@1.0.8", "", {}, "sha512-JlJaRaS4hBTypxtFe8WhnwV8blf0R+3yehLk8XuyxUYNx6VXsKCjACSCvOYEFUiqlhlBWxtYCn/zRlOb8BzBQg=="],
"@types/bun": ["@types/bun@1.2.16", "", { "dependencies": { "bun-types": "1.2.16" } }, "sha512-1aCZJ/6nSiViw339RsaNhkNoEloLaPzZhxMOYEa7OzRzO41IGg5n/7I43/ZIAW/c+Q6cT12Vf7fOZOoVIzb5BQ=="],
"@types/bun": ["@types/bun@1.3.0", "", { "dependencies": { "bun-types": "1.3.0" } }, "sha512-+lAGCYjXjip2qY375xX/scJeVRmZ5cY0wyHYyCYxNcdEXrQ4AOe3gACgd4iQ8ksOslJtW4VNxBJ8llUwc3a6AA=="],
"@types/node": ["@types/node@22.13.9", "", { "dependencies": { "undici-types": "~6.20.0" } }, "sha512-acBjXdRJ3A6Pb3tqnw9HZmyR3Fiol3aGxRCK1x3d+6CDAMjl7I649wpSd+yNURCjbOUGu9tqtLKnTGxmK6CyGw=="],
@@ -78,7 +82,7 @@
"buffer": ["buffer@4.9.2", "", { "dependencies": { "base64-js": "^1.0.2", "ieee754": "^1.1.4", "isarray": "^1.0.0" } }, "sha512-xq+q3SRMOxGivLhBNaUdC64hDTQwejJ+H0T/NB1XMtTVEwNTrfFF3gAxiyW0Bu/xWEGhjVKgUcMhCrUy2+uCWg=="],
"bun-types": ["bun-types@1.2.16", "", { "dependencies": { "@types/node": "*" } }, "sha512-ciXLrHV4PXax9vHvUrkvun9VPVGOVwbbbBF/Ev1cXz12lyEZMoJpIJABOfPcN9gDJRaiKF9MVbSygLg4NXu3/A=="],
"bun-types": ["bun-types@1.3.0", "", { "dependencies": { "@types/node": "*" }, "peerDependencies": { "@types/react": "^19" } }, "sha512-u8X0thhx+yJ0KmkxuEo9HAtdfgCBaM/aI9K90VQcQioAmkVp3SG3FkwWGibUFz3WdXAdcsqOcbU40lK7tbHdkQ=="],
"bytes": ["bytes@3.1.2", "", {}, "sha512-/Nf7TyzTx6S3yRJObOAV7956r8cr2+Oj8AC5dt8wSP3BQAoeX58NoHyCU8P8zGkNXStjTSi6fzO6F0pBdcYbEg=="],
@@ -240,6 +244,8 @@
"raw-body": ["raw-body@3.0.0", "", { "dependencies": { "bytes": "3.1.2", "http-errors": "2.0.0", "iconv-lite": "0.6.3", "unpipe": "1.0.0" } }, "sha512-RmkhL8CAyCRPXCE28MMH0z2PNWQBNk2Q09ZdxM9IOOXwxwZbN+qbWaatPkdkWIKL2ZVDImrN/pK5HTRz2PcS4g=="],
"remeda": ["remeda@2.33.7", "", {}, "sha512-cXlyjevWx5AcslOUEETG4o8XYi9UkoCXcJmj7XhPFVbla+ITuOBxv6ijBrmbeg+ZhzmDThkNdO+iXKUfrJep1w=="],
"router": ["router@2.2.0", "", { "dependencies": { "debug": "^4.4.0", "depd": "^2.0.0", "is-promise": "^4.0.0", "parseurl": "^1.3.3", "path-to-regexp": "^8.0.0" } }, "sha512-nLTrUKm2UyiL7rlhapu/Zl45FwNgkZGaCpZbIHajDYgwlJCOzLSk+cIPAnsEqV955GjILJnKbdQC1nVPz+gAYQ=="],
"safe-buffer": ["safe-buffer@5.2.1", "", {}, "sha512-rp3So07KcdmmKbGvgaNxQSJr7bGVSVk5S9Eq1F+ppbRo70+YeaDxkw5Dd8NPN+GD6bjnYm2VuPuCXmpuYvmCXQ=="],
@@ -322,8 +328,6 @@
"http-errors/statuses": ["statuses@2.0.1", "", {}, "sha512-RwNA9Z/7PrK06rYLIzFMlaF+l73iwpzsqRIFgbMLbTcLD6cOao82TaWefPXQvB2fOC4AjuYSEndS7N/mTCbkdQ=="],
"models.dev/@types/bun": ["@types/bun@1.3.0", "", { "dependencies": { "bun-types": "1.3.0" } }, "sha512-+lAGCYjXjip2qY375xX/scJeVRmZ5cY0wyHYyCYxNcdEXrQ4AOe3gACgd4iQ8ksOslJtW4VNxBJ8llUwc3a6AA=="],
"opencontrol/@tsconfig/bun": ["@tsconfig/bun@1.0.7", "", {}, "sha512-udGrGJBNQdXGVulehc1aWT73wkR9wdaGBtB6yL70RJsqwW/yJhIg6ZbRlPOfIUiFNrnBuYLBi9CSmMKfDC7dvA=="],
"opencontrol/hono": ["hono@4.7.4", "", {}, "sha512-Pst8FuGqz3L7tFF+u9Pu70eI0xa5S3LPUmrNd5Jm8nTHze9FxLTK9Kaj5g/k4UcwuJSXTP65SyHOPLrffpcAJg=="],
@@ -331,11 +335,5 @@
"openid-client/jose": ["jose@4.15.9", "", {}, "sha512-1vUQX+IdDMVPj4k8kOxgUqlcK518yluMuGZwqlr44FS1ppZB/5GWh4rZG89erpOBOJjU/OBsnCVFfapsRz6nEA=="],
"bun-types/@types/node/undici-types": ["undici-types@7.8.0", "", {}, "sha512-9UJ2xGDvQ43tYyVMpuHlsgApydB8ZKfVYTsLDhXkFL/6gfkp+U8xTGdh8pMJv1SpZna0zxG1DwsKZsreLbXBxw=="],
"models.dev/@types/bun/bun-types": ["bun-types@1.3.0", "", { "dependencies": { "@types/node": "*" }, "peerDependencies": { "@types/react": "^19" } }, "sha512-u8X0thhx+yJ0KmkxuEo9HAtdfgCBaM/aI9K90VQcQioAmkVp3SG3FkwWGibUFz3WdXAdcsqOcbU40lK7tbHdkQ=="],
"models.dev/@types/bun/bun-types/@types/node": ["@types/node@24.0.3", "", { "dependencies": { "undici-types": "~7.8.0" } }, "sha512-R4I/kzCYAdRLzfiCabn9hxWfbuHS573x+r0dJMkkzThEa7pbrcDWK+9zu3e7aBOouf+rQAciqPFMnxwr0aWgKg=="],
"models.dev/@types/bun/bun-types/@types/node/undici-types": ["undici-types@7.8.0", "", {}, "sha512-9UJ2xGDvQ43tYyVMpuHlsgApydB8ZKfVYTsLDhXkFL/6gfkp+U8xTGdh8pMJv1SpZna0zxG1DwsKZsreLbXBxw=="],
}
}
+1
View File
File diff suppressed because one or more lines are too long
+10 -2
View File
@@ -16,12 +16,20 @@
},
"scripts": {
"validate": "bun ./packages/core/script/validate.ts",
"compare:migrations": "bun ./packages/core/script/compare-model-migrations.ts",
"cloudflare:sync": "bun ./packages/core/script/sync-models.ts cloudflare-workers-ai",
"chutes:generate": "bun ./packages/core/script/generate-chutes.ts",
"databricks:generate": "bun ./packages/core/script/generate-databricks.ts",
"helicone:generate": "bun ./packages/core/script/generate-helicone.ts",
"venice:generate": "bun ./packages/core/script/generate-venice.ts",
"vercel:generate": "bun ./packages/core/script/generate-vercel.ts"
"vercel:generate": "bun ./packages/core/script/generate-vercel.ts",
"wandb:generate": "bun ./packages/core/script/generate-wandb.ts",
"digitalocean:generate": "bun ./packages/core/script/generate-digitalocean.ts",
"ambient:generate": "bun ./packages/core/script/generate-ambient.ts",
"models:sync": "bun ./packages/core/script/sync-models.ts"
},
"dependencies": {
"@cloudflare/workers-types": "^4.20250801.0",
"@cloudflare/workers-types": "^4.20260424.1",
"sst": "3.17.23"
}
}
+1
View File
@@ -4,6 +4,7 @@
"$schema": "https://json.schemastore.org/package.json",
"type": "module",
"dependencies": {
"remeda": "^2.33.7",
"zod": "catalog:"
},
"main": "./src/index.ts",
@@ -0,0 +1,89 @@
#!/usr/bin/env bun
import path from "node:path";
import { cp, mkdir, rm, writeFile } from "node:fs/promises";
import { tmpdir } from "node:os";
import { generate } from "../src/generate.js";
const root = path.join(import.meta.dirname, "..", "..", "..");
const providersPath = path.join(root, "providers");
const diffOutput = await Bun.$`git diff --name-only HEAD -- providers`.cwd(root).text();
const changedProviderPaths = diffOutput
.split("\n")
.filter(Boolean)
.filter((filePath) => /^providers\/[^/]+\/models\/.+\.toml$/.test(filePath));
if (changedProviderPaths.length === 0) {
process.exit(0);
}
const baselineRoot = path.join(tmpdir(), `models-dev-compare-${Date.now()}`);
await mkdir(baselineRoot, { recursive: true });
try {
const baselineProvidersPath = path.join(baselineRoot, "providers");
await cp(providersPath, baselineProvidersPath, { recursive: true });
for (const filePath of changedProviderPaths) {
const tempFilePath = path.join(baselineRoot, filePath);
const show = Bun.spawn(["git", "show", `HEAD:${filePath}`], {
cwd: root,
stdout: "pipe",
stderr: "pipe",
});
const exitCode = await show.exited;
if (exitCode !== 0) {
await rm(tempFilePath, { force: true });
continue;
}
const contents = await new Response(show.stdout).text();
await mkdir(path.dirname(tempFilePath), { recursive: true });
await writeFile(tempFilePath, contents);
}
const before = await generate(baselineProvidersPath);
const after = await generate(providersPath);
for (const filePath of changedProviderPaths) {
const match = /^providers\/([^/]+)\/models\/(.+)\.toml$/.exec(filePath);
if (!match) continue;
const [, providerID, modelID] = match;
const beforeModel = before[providerID]?.models[modelID];
const afterModel = after[providerID]?.models[modelID];
const beforeJson = JSON.stringify(beforeModel, null, 2);
const afterJson = JSON.stringify(afterModel, null, 2);
if (beforeJson === afterJson) {
continue;
}
const beforeFilePath = path.join(baselineRoot, "before.json");
const afterFilePath = path.join(baselineRoot, "after.json");
await writeFile(beforeFilePath, `${beforeJson}\n`);
await writeFile(afterFilePath, `${afterJson}\n`);
const diff = Bun.spawn(
[
"diff",
"-u",
"-L",
`${filePath} (before)`,
"-L",
`${filePath} (after)`,
beforeFilePath,
afterFilePath,
],
{
stdout: "pipe",
stderr: "pipe",
},
);
const output = await new Response(diff.stdout).text();
process.stdout.write(output);
}
} finally {
await rm(baselineRoot, { recursive: true, force: true });
}
+172
View File
@@ -0,0 +1,172 @@
#!/usr/bin/env bun
/**
* Generates Ambient model TOML files from https://api.ambient.xyz/v1/models.
*
* Emits `[extends]`-format TOMLs that inherit upstream metadata
* (family, release_date, knowledge, capabilities) from the canonical
* provider model, and override only the fields Ambient's API reports:
* cost, limit, modalities.
*
* Flags:
* --dry-run Preview generated TOMLs without writing files.
*/
import { z } from "zod";
import path from "node:path";
import { mkdir } from "node:fs/promises";
const API_ENDPOINT = "https://api.ambient.xyz/v1/models";
// Allowlist for the initial rollout.
const ALLOWLIST = new Set<string>([
"zai-org/GLM-5.1-FP8",
"moonshotai/kimi-k2.6",
]);
// Maps Ambient model IDs to canonical <provider>/<model> in this repo.
// The generated TOML uses this path as `[extends].from` so capabilities
// and metadata propagate from the upstream provider automatically.
const EXTENDS_MAP: Record<string, string> = {
"zai-org/GLM-5.1-FP8": "zai/glm-5.1",
"moonshotai/kimi-k2.6": "moonshotai/kimi-k2.6",
};
const Pricing = z
.object({
prompt: z.string(),
completion: z.string(),
input_cache_read: z.string().optional(),
input_cache_write: z.string().optional(),
})
.passthrough();
const AmbientModel = z
.object({
id: z.string(),
name: z.string(),
context_length: z.number(),
max_output_length: z.number(),
input_modalities: z.array(z.string()),
output_modalities: z.array(z.string()),
pricing: Pricing,
})
.passthrough();
const AmbientResponse = z
.object({
object: z.literal("list"),
data: z.array(AmbientModel),
})
.passthrough();
const ALLOWED_MODALITIES = new Set(["text", "audio", "image", "video", "pdf"]);
function modalities(values: string[]): string[] {
return values
.map((v) => v.toLowerCase())
.filter((v) => ALLOWED_MODALITIES.has(v));
}
function perMTok(price: string): number {
const n = parseFloat(price);
if (!Number.isFinite(n)) {
throw new Error(`Invalid price: ${price}`);
}
// Round to 6 decimals to absorb float noise from per-token strings.
return Math.round(n * 1_000_000 * 1_000_000) / 1_000_000;
}
function formatToml(
model: z.infer<typeof AmbientModel>,
extendsFrom: string,
): string {
const lines: string[] = [];
lines.push("[extends]");
lines.push(`from = "${extendsFrom}"`);
lines.push("");
lines.push("[cost]");
lines.push(`input = ${perMTok(model.pricing.prompt)}`);
lines.push(`output = ${perMTok(model.pricing.completion)}`);
if (model.pricing.input_cache_read !== undefined) {
lines.push(`cache_read = ${perMTok(model.pricing.input_cache_read)}`);
}
if (model.pricing.input_cache_write !== undefined) {
lines.push(`cache_write = ${perMTok(model.pricing.input_cache_write)}`);
}
lines.push("");
lines.push("[limit]");
lines.push(`context = ${model.context_length}`);
lines.push(`output = ${model.max_output_length}`);
lines.push("");
const input = modalities(model.input_modalities);
const output = modalities(model.output_modalities);
lines.push("[modalities]");
lines.push(`input = [${input.map((m) => `"${m}"`).join(", ")}]`);
lines.push(`output = [${output.map((m) => `"${m}"`).join(", ")}]`);
return lines.join("\n") + "\n";
}
async function main() {
const dryRun = process.argv.includes("--dry-run");
const outDir = path.join(
import.meta.dirname,
"..",
"..",
"..",
"providers",
"ambient",
"models",
);
const res = await fetch(API_ENDPOINT);
if (!res.ok) {
console.error(`Fetch failed: ${res.status} ${res.statusText}`);
process.exit(1);
}
const parsed = AmbientResponse.safeParse(await res.json());
if (!parsed.success) {
console.error("Invalid Ambient response:", parsed.error.issues);
process.exit(1);
}
const selected = parsed.data.data.filter((m) => ALLOWLIST.has(m.id));
const missing = [...ALLOWLIST].filter(
(id) => !selected.some((m) => m.id === id),
);
if (missing.length > 0) {
console.error(`Allowlisted models missing from API: ${missing.join(", ")}`);
process.exit(1);
}
let count = 0;
for (const model of selected) {
const extendsFrom = EXTENDS_MAP[model.id];
if (!extendsFrom) {
console.error(`No EXTENDS_MAP entry for ${model.id}; skipping`);
continue;
}
const filePath = path.join(outDir, `${model.id}.toml`);
const toml = formatToml(model, extendsFrom);
if (dryRun) {
console.log(`--- ${path.relative(process.cwd(), filePath)} ---`);
console.log(toml);
} else {
await mkdir(path.dirname(filePath), { recursive: true });
await Bun.write(filePath, toml);
}
count++;
}
console.log(
`${dryRun ? "Previewed" : "Wrote"} ${count} model file(s) under providers/ambient/models/`,
);
}
await main();
+589
View File
@@ -0,0 +1,589 @@
#!/usr/bin/env bun
/**
* Generates Chutes model TOML files from the Chutes LLM API.
*
* Flags:
* --dry-run: Preview changes without writing files
* --new-only: Only create new models, skip updating existing ones
* --keep-orphans: Don't delete TOML files for models no longer in the API
*/
import { z } from "zod";
import path from "node:path";
import { mkdir } from "node:fs/promises";
import { ModelFamilyValues } from "../src/family.js";
const API_ENDPOINT = "https://llm.chutes.ai/v1/models";
enum SkipZeroFields {
LimitContext = "limit.context",
LimitOutput = "limit.output",
}
const Pricing = z.object({
prompt: z.number().optional(),
completion: z.number().optional(),
input_cache_read: z.number().optional(),
}).passthrough();
const ChutesModel = z.object({
id: z.string(),
created: z.number(),
pricing: Pricing.optional(),
context_length: z.number().optional(),
max_output_length: z.number().optional(),
max_model_len: z.number().optional(),
input_modalities: z.array(z.string()).optional(),
output_modalities: z.array(z.string()).optional(),
supported_features: z.array(z.string()).optional(),
supported_sampling_parameters: z.array(z.string()).optional(),
quantization: z.string().optional(),
}).passthrough();
const ChutesResponse = z.object({
data: z.array(ChutesModel),
}).passthrough();
interface ExistingModel {
name?: string;
family?: string;
attachment?: boolean;
reasoning?: boolean;
tool_call?: boolean;
structured_output?: boolean;
temperature?: boolean;
knowledge?: string;
release_date?: string;
last_updated?: string;
open_weights?: boolean;
interleaved?: boolean | { field: string };
status?: string;
cost?: {
input?: number;
output?: number;
cache_read?: number;
};
limit?: {
context?: number;
output?: number;
};
modalities?: {
input?: string[];
output?: string[];
};
}
interface MergedModel {
name: string;
family?: string;
attachment: boolean;
reasoning: boolean;
tool_call: boolean;
structured_output?: boolean;
temperature: boolean;
knowledge?: string;
release_date: string;
last_updated: string;
open_weights: boolean;
interleaved?: boolean | { field: string };
status?: string;
cost?: {
input: number;
output: number;
cache_read?: number;
};
limit: {
context: number;
output: number;
};
modalities: {
input: string[];
output: string[];
};
}
interface Changes {
field: string;
oldValue: string;
newValue: string;
}
// ── Utility functions ────────────────────────────────────────────────
function timestampToDate(timestamp: number): string {
const date = new Date(timestamp * 1000);
return date.toISOString().slice(0, 10);
}
function getTodayDate(): string {
return new Date().toISOString().slice(0, 10);
}
function formatNumber(n: number): string {
if (n >= 1000) {
return n.toString().replace(/\B(?=(\d{3})+(?!\d))/g, "_");
}
return n.toString();
}
/**
* Humanize a model ID into a readable name.
* Strips the org prefix and replaces hyphens with spaces.
* e.g. "Qwen/Qwen3-32B-TEE" → "Qwen3 32B TEE"
*/
function humanizeModelName(modelId: string): string {
const parts = modelId.split("/");
const modelPart = parts[parts.length - 1];
return modelPart.replace(/-/g, " ");
}
// ── Family inference ───────────
function isSubstring(target: string, family: string): boolean {
return target.toLowerCase().includes(family.toLowerCase());
}
function matchesFamily(target: string, family: string): boolean {
const targetLower = target.toLowerCase();
const familyLower = family.toLowerCase();
let familyIdx = 0;
for (let i = 0; i < targetLower.length && familyIdx < familyLower.length; i++) {
if (targetLower[i] === familyLower[familyIdx]) {
familyIdx++;
}
}
return familyIdx === familyLower.length;
}
function inferFamily(modelId: string, modelName: string): string | undefined {
const sortedFamilies = [...ModelFamilyValues].sort((a, b) => b.length - a.length);
// First pass: try exact substring matches
for (const family of sortedFamilies) {
if (isSubstring(modelId, family)) {
return family;
}
}
for (const family of sortedFamilies) {
if (isSubstring(modelName, family)) {
return family;
}
}
// Second pass: fall back to subsequence matching
for (const family of sortedFamilies) {
if (matchesFamily(modelId, family)) {
return family;
}
}
for (const family of sortedFamilies) {
if (matchesFamily(modelName, family)) {
return family;
}
}
return undefined;
}
// ── Load existing TOML ───────────────────────────────────────────────
async function loadExistingModel(filePath: string): Promise<ExistingModel | null> {
try {
const file = Bun.file(filePath);
if (!(await file.exists())) {
return null;
}
const toml = await import(filePath, { with: { type: "toml" } }).then(
(mod) => mod.default,
);
return toml as ExistingModel;
} catch (e) {
console.warn(`Warning: Failed to parse existing file ${filePath}:`, e);
return null;
}
}
// ── Merge API data with existing TOML ────────────────────────────────
function mergeModel(
apiModel: z.infer<typeof ChutesModel>,
existing: ExistingModel | null,
): MergedModel {
const features = new Set(apiModel.supported_features ?? []);
const samplingParams = new Set(apiModel.supported_sampling_parameters ?? []);
const inputMods = apiModel.input_modalities ?? ["text"];
const outputMods = apiModel.output_modalities ?? ["text"];
// Capabilities from API features
const hasAttachment = inputMods.some((m) =>
m === "image" || m === "video" || m === "pdf",
);
const hasReasoning = features.has("reasoning");
const hasToolCall = features.has("tools");
const hasStructuredOutput = features.has("structured_outputs");
const hasTemperature = samplingParams.size > 0
? samplingParams.has("temperature")
: true; // default true if no sampling params info
// Preserve existing values when available (manually specified)
const modelName = existing?.name ?? humanizeModelName(apiModel.id);
const family = existing?.family ?? inferFamily(apiModel.id, modelName);
const knowledge = existing?.knowledge;
const interleaved = existing?.interleaved;
const status = existing?.status;
// Release date: existing > API created timestamp > today
const releaseDate = existing?.release_date
?? timestampToDate(apiModel.created)
?? getTodayDate();
// Context limit: prefer context_length, fallback to max_model_len
const apiContext = apiModel.context_length ?? apiModel.max_model_len ?? 0;
const contextLimit = apiContext > 0
? apiContext
: (existing?.limit?.context ?? 0);
// Output limit: prefer max_output_length, fallback to existing
const apiOutput = apiModel.max_output_length ?? 0;
const outputLimit = apiOutput > 0
? apiOutput
: (existing?.limit?.output ?? 0);
const merged: MergedModel = {
name: modelName,
family,
attachment: hasAttachment,
reasoning: hasReasoning,
tool_call: hasToolCall,
temperature: hasTemperature,
release_date: releaseDate,
last_updated: getTodayDate(),
open_weights: true, // Chutes hosts open-weight models
...(hasStructuredOutput && { structured_output: hasStructuredOutput }),
...(knowledge && { knowledge }),
...(interleaved !== undefined && { interleaved }),
...(status && { status }),
limit: {
context: contextLimit,
output: outputLimit,
},
modalities: {
input: inputMods,
output: outputMods,
},
};
// Cost: API values are already in USD per 1M tokens — use directly
if (apiModel.pricing) {
const inputPrice = apiModel.pricing.prompt;
const outputPrice = apiModel.pricing.completion;
const cacheReadPrice = apiModel.pricing.input_cache_read;
if (inputPrice !== undefined && outputPrice !== undefined) {
merged.cost = {
input: inputPrice,
output: outputPrice,
...(cacheReadPrice !== undefined && { cache_read: cacheReadPrice }),
};
}
}
return merged;
}
// ── TOML formatting ──────────────────────────────────────────────────
function formatToml(model: MergedModel): string {
const lines: string[] = [];
lines.push(`# Auto-generated by generate-chutes.ts — do not edit pricing, limits, or capabilities.`);
lines.push(`# Manual overrides preserved on re-run: name, family, knowledge, interleaved, status`);
lines.push(`name = "${model.name.replace(/"/g, '\\"')}"`);
if (model.family) {
lines.push(`family = "${model.family}"`);
}
lines.push(`release_date = "${model.release_date}"`);
lines.push(`last_updated = "${model.last_updated}"`);
lines.push(`attachment = ${model.attachment}`);
lines.push(`reasoning = ${model.reasoning}`);
lines.push(`temperature = ${model.temperature}`);
lines.push(`tool_call = ${model.tool_call}`);
if (model.structured_output !== undefined) {
lines.push(`structured_output = ${model.structured_output}`);
}
lines.push(`open_weights = ${model.open_weights}`);
if (model.knowledge) {
lines.push(`knowledge = "${model.knowledge}"`);
}
if (model.status) {
lines.push(`status = "${model.status}"`);
}
if (model.cost) {
lines.push("");
lines.push(`[cost]`);
lines.push(`input = ${model.cost.input}`);
lines.push(`output = ${model.cost.output}`);
if (model.cost.cache_read !== undefined) {
lines.push(`cache_read = ${model.cost.cache_read}`);
}
}
lines.push("");
lines.push(`[limit]`);
lines.push(`context = ${formatNumber(model.limit.context)}`);
lines.push(`output = ${formatNumber(model.limit.output)}`);
lines.push("");
lines.push(`[modalities]`);
lines.push(`input = [${model.modalities.input.map((m) => `"${m}"`).join(", ")}]`);
lines.push(`output = [${model.modalities.output.map((m) => `"${m}"`).join(", ")}]`);
if (model.interleaved !== undefined) {
lines.push("");
if (model.interleaved === true) {
lines.push(`interleaved = true`);
} else if (typeof model.interleaved === "object") {
lines.push(`[interleaved]`);
lines.push(`field = "${model.interleaved.field}"`);
}
}
return lines.join("\n") + "\n";
}
// ── Change detection ─────────────────────────────────────────────────
function detectChanges(
existing: ExistingModel | null,
merged: MergedModel,
): Changes[] {
if (!existing) return [];
const changes: Changes[] = [];
const EPSILON = 0.001;
const shouldSkipZero = (field: string, oldVal: unknown, newVal: unknown): boolean => {
if (!Object.values(SkipZeroFields).includes(field as SkipZeroFields)) {
return false;
}
return (typeof oldVal === "number" && oldVal === 0) || (typeof newVal === "number" && newVal === 0);
};
const formatValue = (val: unknown): string => {
if (typeof val === "number") return formatNumber(val);
if (Array.isArray(val)) return `[${val.join(", ")}]`;
if (val === undefined) return "(none)";
return String(val);
};
const isMaterialPriceDiff = (oldPrice: unknown, newPrice: unknown): boolean => {
if (oldPrice === 0 && newPrice === undefined) return false;
if (oldPrice !== undefined && newPrice !== undefined) {
return Math.abs((oldPrice as number) - (newPrice as number)) > EPSILON;
}
return oldPrice !== newPrice;
};
const compare = (field: string, oldVal: unknown, newVal: unknown) => {
if (shouldSkipZero(field, oldVal, newVal)) return;
const isDiff = field.startsWith("cost.")
? isMaterialPriceDiff(oldVal, newVal)
: JSON.stringify(oldVal) !== JSON.stringify(newVal);
if (isDiff) {
changes.push({
field,
oldValue: formatValue(oldVal),
newValue: formatValue(newVal),
});
}
};
compare("name", existing.name, merged.name);
compare("family", existing.family, merged.family);
compare("attachment", existing.attachment, merged.attachment);
compare("reasoning", existing.reasoning, merged.reasoning);
compare("tool_call", existing.tool_call, merged.tool_call);
compare("structured_output", existing.structured_output, merged.structured_output);
compare("open_weights", existing.open_weights, merged.open_weights);
compare("release_date", existing.release_date, merged.release_date);
compare("cost.input", existing.cost?.input, merged.cost?.input);
compare("cost.output", existing.cost?.output, merged.cost?.output);
compare("cost.cache_read", existing.cost?.cache_read, merged.cost?.cache_read);
compare("limit.context", existing.limit?.context, merged.limit.context);
compare("limit.output", existing.limit?.output, merged.limit.output);
compare("modalities.input", existing.modalities?.input, merged.modalities.input);
compare("modalities.output", existing.modalities?.output, merged.modalities.output);
return changes;
}
// ── Main ─────────────────────────────────────────────────────────────
async function main() {
const args = process.argv.slice(2);
const dryRun = args.includes("--dry-run");
const newOnly = args.includes("--new-only");
const keepOrphans = args.includes("--keep-orphans");
const modelsDir = path.join(
import.meta.dirname,
"..",
"..",
"..",
"providers",
"chutes",
"models",
);
console.log(`${dryRun ? "[DRY RUN] " : ""}${newOnly ? "[NEW ONLY] " : ""}${keepOrphans ? "[KEEP ORPHANS] " : ""}Fetching Chutes models from API...`);
const res = await fetch(API_ENDPOINT);
if (!res.ok) {
console.error(`Failed to fetch API: ${res.status} ${res.statusText}`);
process.exit(1);
}
const json = await res.json();
const parsed = ChutesResponse.safeParse(json);
if (!parsed.success) {
console.error("Invalid API response:", parsed.error.errors);
process.exit(1);
}
const apiModels = parsed.data.data;
// Scan existing TOML files
const existingFiles = new Set<string>();
try {
for await (const file of new Bun.Glob("**/*.toml").scan({
cwd: modelsDir,
absolute: false,
})) {
existingFiles.add(file);
}
} catch {
}
console.log(`Found ${apiModels.length} models in API, ${existingFiles.size} existing files\n`);
const apiModelIds = new Set<string>();
let created = 0;
let updated = 0;
let unchanged = 0;
for (const apiModel of apiModels) {
const relativePath = `${apiModel.id}.toml`;
const filePath = path.join(modelsDir, relativePath);
const dirPath = path.dirname(filePath);
apiModelIds.add(relativePath);
const existing = await loadExistingModel(filePath);
const merged = mergeModel(apiModel, existing);
const tomlContent = formatToml(merged);
if (existing === null) {
created++;
if (dryRun) {
console.log(`[DRY RUN] Would create: ${relativePath}`);
console.log(` name = "${merged.name}"`);
if (merged.family) {
console.log(` family = "${merged.family}" (inferred)`);
}
console.log("");
} else {
await mkdir(dirPath, { recursive: true });
await Bun.write(filePath, tomlContent);
console.log(`Created: ${relativePath}`);
}
} else {
if (newOnly) {
unchanged++;
continue;
}
const changes = detectChanges(existing, merged);
const existingContent = await Bun.file(filePath).text();
const formatChanged = existingContent !== tomlContent;
if (changes.length > 0 || formatChanged) {
updated++;
if (dryRun) {
console.log(`[DRY RUN] Would update: ${relativePath}`);
} else {
await mkdir(dirPath, { recursive: true });
await Bun.write(filePath, tomlContent);
console.log(`Updated: ${relativePath}`);
}
for (const change of changes) {
console.log(` ${change.field}: ${change.oldValue}${change.newValue}`);
}
if (changes.length === 0 && formatChanged) {
console.log(` (format-only change)`);
}
console.log("");
} else {
unchanged++;
}
}
}
// Handle orphaned files (on disk but not in API)
const orphaned: string[] = [];
for (const file of existingFiles) {
if (!apiModelIds.has(file)) {
orphaned.push(file);
const orphanPath = path.join(modelsDir, file);
if (keepOrphans) {
console.log(`Orphaned (kept): ${file}`);
} else if (dryRun) {
console.log(`[DRY RUN] Would delete: ${file}`);
} else {
await Bun.file(orphanPath).delete();
console.log(`Deleted: ${file}`);
// Clean up empty parent directories
const parentDir = path.dirname(orphanPath);
try {
const remaining = [];
for await (const entry of new Bun.Glob("*").scan({ cwd: parentDir })) {
remaining.push(entry);
}
if (remaining.length === 0) {
const { rmdir } = await import("node:fs/promises");
await rmdir(parentDir);
console.log(` Removed empty directory: ${path.basename(parentDir)}/`);
}
} catch {
// Directory not empty or other error, ignore
}
}
}
}
console.log("");
if (dryRun) {
console.log(
`Summary: ${created} would be created, ${updated} would be updated, ${unchanged} unchanged, ${orphaned.length} would be deleted`,
);
} else if (keepOrphans) {
console.log(
`Summary: ${created} created, ${updated} updated, ${unchanged} unchanged, ${orphaned.length} orphaned (kept)`,
);
} else {
console.log(
`Summary: ${created} created, ${updated} updated, ${unchanged} unchanged, ${orphaned.length} deleted`,
);
}
}
await main();
+287
View File
@@ -0,0 +1,287 @@
#!/usr/bin/env bun
/**
* Generates Databricks model TOML files from the Foundation Model API endpoint.
*
* Each Databricks endpoint exposes a model from another provider (Anthropic,
* OpenAI, Google, etc.), so the generated TOML uses [extends] to inherit
* canonical metadata from that upstream provider's TOML in models.dev.
*
* Usage:
* DATABRICKS_HOST=<host> DATABRICKS_TOKEN=<pat> bun run databricks:generate
* bun run databricks:generate --workspace <host> --token <pat>
*
* Flags:
* --dry-run: Preview changes without writing files
* --new-only: Only create new models, skip updating existing ones
*/
import { z } from "zod";
import path from "node:path";
import { mkdir, readFile } from "node:fs/promises";
import { existsSync } from "node:fs";
const args = process.argv.slice(2);
const flag = (name: string) => {
const i = args.indexOf(`--${name}`);
return i !== -1 ? args[i + 1] : undefined;
};
const dryRun = args.includes("--dry-run");
const newOnly = args.includes("--new-only");
const host = flag("workspace") ?? process.env.DATABRICKS_HOST;
const token = flag("token") ?? process.env.DATABRICKS_TOKEN;
if (!host || !token) {
console.error(
"Usage: DATABRICKS_HOST=<host> DATABRICKS_TOKEN=<pat> bun run databricks:generate",
);
process.exit(1);
}
const workspace = host.replace(/^https?:\/\//, "").replace(/\/$/, "");
const PROVIDERS_DIR = path.join(import.meta.dirname, "..", "..", "..", "providers");
const MODELS_DIR = path.join(PROVIDERS_DIR, "databricks", "models");
// ---------------------------------------------------------------------------
// API schemas
// ---------------------------------------------------------------------------
const FoundationModel = z
.object({
ai_gateway_v2_supported: z.boolean().optional(),
api_types: z.array(z.string()).optional(),
})
.passthrough();
const ServedEntity = z
.object({
foundation_model: FoundationModel.optional(),
})
.passthrough();
const Endpoint = z
.object({
name: z.string(),
config: z
.object({
served_entities: z.array(ServedEntity).optional(),
})
.passthrough()
.optional(),
})
.passthrough();
const FoundationModelsResponse = z
.object({
endpoints: z.array(Endpoint),
})
.passthrough();
// ---------------------------------------------------------------------------
// Canonical resolution: map a Databricks endpoint name to a models.dev entry
// ---------------------------------------------------------------------------
const PREFIX_TO_PROVIDER: [string, string][] = [
["claude-", "anthropic"],
["gpt-", "openai"],
["gemini-", "google"],
["mistral-", "mistral"],
["mixtral-", "mistral"],
];
type Resolution =
| { type: "extends"; from: string }
| { type: "inline"; content: string }
| null;
async function resolveCanonical(endpointName: string): Promise<Resolution> {
const bare = endpointName.replace(/^databricks-/, "");
// Models in provider subdirectories (e.g. openrouter/openai/gpt-oss-*)
// can't use [extends] (schema requires provider/model format), so inline.
if (bare.startsWith("gpt-oss-")) {
const p = path.join(PROVIDERS_DIR, "openrouter", "models", "openai", `${bare}.toml`);
if (existsSync(p)) {
return { type: "inline", content: await readFile(p, "utf8") };
}
}
// Meta Llama: "meta-llama-3-3-70b-instruct" → "llama-3.3-70b-instruct"
if (bare.startsWith("meta-llama-") || bare.startsWith("llama-")) {
const llamaId = bare
.replace(/^meta-llama-/, "llama-")
.replace(/^(llama-\d+)-(\d+)-/, "$1.$2-");
const p = path.join(PROVIDERS_DIR, "llama", "models", `${llamaId}.toml`);
if (existsSync(p)) return { type: "extends", from: `llama/${llamaId}` };
}
for (const [prefix, provider] of PREFIX_TO_PROVIDER) {
if (!bare.startsWith(prefix)) continue;
const exact = path.join(PROVIDERS_DIR, provider, "models", `${bare}.toml`);
if (existsSync(exact)) return { type: "extends", from: `${provider}/${bare}` };
// Try with hyphens-as-dots in version (e.g. gpt-5-4 → gpt-5.4)
const dotted = bare.replace(/^((?:[a-z]+-)+\d+)-(\d)/, "$1.$2");
if (dotted !== bare) {
const dottedExact = path.join(PROVIDERS_DIR, provider, "models", `${dotted}.toml`);
if (existsSync(dottedExact)) return { type: "extends", from: `${provider}/${dotted}` };
}
// Fuzzy: longest filename that shares a prefix with bare or its dotted form
const candidates = [bare, ...(dotted !== bare ? [dotted] : [])];
const files: string[] = [];
try {
for await (const f of new Bun.Glob("*.toml").scan({
cwd: path.join(PROVIDERS_DIR, provider, "models"),
})) {
files.push(f);
}
} catch {
// provider directory may not exist
}
const match = files
.map((f) => f.replace(/\.toml$/, ""))
.filter((id) => candidates.some((c) => id.startsWith(c) || c.startsWith(id)))
.sort((a, b) => b.length - a.length)[0];
if (match) return { type: "extends", from: `${provider}/${match}` };
}
return null;
}
function formatToml(resolution: Resolution, endpointName: string): string {
if (resolution?.type === "extends") {
return `[extends]\nfrom = "${resolution.from}"\n`;
}
if (resolution?.type === "inline") {
return resolution.content;
}
return `# TODO: fill in details for ${endpointName}\nname = "${endpointName}"\n`;
}
// ---------------------------------------------------------------------------
// Main
// ---------------------------------------------------------------------------
const IGNORE_PREFIXES = [
"databricks-llama-",
"databricks-meta-llama-",
"databricks-qwen",
"databricks-gemma-",
];
async function main() {
console.log(
`${dryRun ? "[DRY RUN] " : ""}${newOnly ? "[NEW ONLY] " : ""}Fetching Databricks foundation-models...`,
);
const url = `https://${workspace}/api/2.0/serving-endpoints:foundation-models`;
const res = await fetch(url, {
headers: { Authorization: `Bearer ${token}` },
});
if (!res.ok) {
console.error(`Failed to fetch API: ${res.status} ${res.statusText}`);
console.error(await res.text().catch(() => ""));
process.exit(1);
}
const json = await res.json();
const parsed = FoundationModelsResponse.safeParse(json);
if (!parsed.success) {
console.error("Invalid API response:", parsed.error.errors);
process.exit(1);
}
const endpoints = parsed.data.endpoints.filter(
(e) =>
!IGNORE_PREFIXES.some((p) => e.name.startsWith(p)) &&
e.config?.served_entities?.some(
(se) =>
se.foundation_model?.ai_gateway_v2_supported === true &&
se.foundation_model?.api_types?.includes("mlflow/v1/chat/completions"),
),
);
const existingFiles = new Set<string>();
try {
for await (const f of new Bun.Glob("*.toml").scan({ cwd: MODELS_DIR })) {
existingFiles.add(f);
}
} catch {
// directory may not exist yet
}
console.log(
`Found ${endpoints.length} models in API, ${existingFiles.size} existing files\n`,
);
const apiModelIds = new Set<string>();
let created = 0;
let updated = 0;
let unchanged = 0;
for (const ep of endpoints) {
const filename = `${ep.name}.toml`;
apiModelIds.add(filename);
const filePath = path.join(MODELS_DIR, filename);
const resolution = await resolveCanonical(ep.name);
const newContent = formatToml(resolution, ep.name);
const tag = resolution?.type === "extends" ? `extends ${resolution.from}` : resolution?.type ?? "stub";
const existed = existsSync(filePath);
if (!existed) {
created++;
if (dryRun) {
console.log(`[DRY RUN] Would create: ${filename}${tag}`);
} else {
await mkdir(MODELS_DIR, { recursive: true });
await Bun.write(filePath, newContent);
console.log(`Created: ${filename}${tag}`);
}
continue;
}
if (newOnly) {
unchanged++;
continue;
}
const existingContent = await readFile(filePath, "utf8");
if (existingContent === newContent) {
unchanged++;
continue;
}
updated++;
if (dryRun) {
console.log(`[DRY RUN] Would update: ${filename}${tag}`);
} else {
await Bun.write(filePath, newContent);
console.log(`Updated: ${filename}${tag}`);
}
}
const orphaned: string[] = [];
for (const file of existingFiles) {
if (!apiModelIds.has(file)) {
orphaned.push(file);
console.log(`Warning: Orphaned file (not in API): ${file}`);
}
}
console.log("");
if (dryRun) {
console.log(
`Summary: ${created} would be created, ${updated} would be updated, ${unchanged} unchanged, ${orphaned.length} orphaned`,
);
} else {
console.log(
`Summary: ${created} created, ${updated} updated, ${unchanged} unchanged, ${orphaned.length} orphaned`,
);
}
}
await main();
@@ -0,0 +1,732 @@
#!/usr/bin/env bun
/**
* Generates DigitalOcean model TOML files from two public APIs:
*
* - https://api.digitalocean.com/v2/gen-ai/models (model metadata, lifecycle, modalities, limits)
* - https://www.digitalocean.com/api/static-content/v1/products (pricing, including >200k tiers)
*
* The v2 models API requires a DigitalOcean personal access token or model access key,
* read from the DIGITALOCEAN_API_TOKEN environment variable (or --api-key flag).
* The static-content pricing API is public and requires no auth.
*
* Cache pricing (cache_read, cache_write) is NOT available from any DO API and is
* preserved from existing TOML files when present.
*
* Fields the APIs cannot provide (preserved from existing TOMLs, never overwritten):
* family, knowledge, open_weights, interleaved, attachment, release_date,
* cache_read, cache_write
*
* Flags:
* --dry-run Preview changes without writing files
* --new-only Only create new models, skip updating existing ones
* --api-key=<key> DigitalOcean API key (overrides DIGITALOCEAN_API_TOKEN env var)
*/
import { z } from "zod";
import path from "node:path";
import { mkdir } from "node:fs/promises";
import { ModelFamilyValues } from "../src/family.js";
const MODELS_API = "https://api.digitalocean.com/v2/gen-ai/models";
const PRICING_API = "https://www.digitalocean.com/api/static-content/v1/products";
// ---------------------------------------------------------------------------
// v2 models API schema
// ---------------------------------------------------------------------------
const DoModel = z
.object({
id: z.string(),
name: z.string(),
lifecycle_status: z.string(),
type: z.string().optional(),
thinking: z.boolean().optional(),
context_window: z.union([z.number(), z.string()]).optional(),
modalities: z
.object({
input: z.array(z.string()).optional(),
output: z.array(z.string()).optional(),
})
.optional(),
settings: z
.array(
z.object({
name: z.string(),
max: z.number().optional(),
default_value: z.number().optional(),
}),
)
.optional(),
created_at: z.string().optional(),
})
.passthrough();
const DoModelsResponse = z
.object({
models: z.array(DoModel),
})
.passthrough();
// ---------------------------------------------------------------------------
// static-content pricing API schema
// ---------------------------------------------------------------------------
const PricingEntry = z
.object({
name: z.string(),
slug: z.string(),
model: z.string(),
prompt_tokens: z.string().optional(), // "≤200k" | ">200k" | undefined
price: z.object({ rate: z.number() }),
})
.passthrough();
const StaticContentResponse = z
.object({
gradient: z.object({
models: z.array(PricingEntry),
}),
})
.passthrough();
// ---------------------------------------------------------------------------
// Derived pricing map
// ---------------------------------------------------------------------------
interface ModelPricing {
input: number;
output: number;
inputOver200k?: number;
outputOver200k?: number;
}
// Map marketing names from /v1/products to API model IDs from /v2/gen-ai/models.
// The pricing API uses display names, not the machine IDs, so this table is the
// join key. Add entries here when DO adds new models with tiered pricing.
const PRICING_NAME_MAP: Record<string, string> = {
// Anthropic
"claude sonnet 4.6": "anthropic-claude-4.6-sonnet",
"claude sonnet 4.5": "anthropic-claude-4.5-sonnet",
"claude sonnet 4": "anthropic-claude-sonnet-4",
"claude haiku 4.5": "anthropic-claude-haiku-4.5",
"claude opus 4.6": "anthropic-claude-opus-4.6",
"claude opus 4.5": "anthropic-claude-opus-4.5",
"claude opus 4.1": "anthropic-claude-4.1-opus",
"claude opus 4": "anthropic-claude-opus-4",
// OpenAI
"gpt-5.4": "openai-gpt-5.4",
"gpt-5.4 mini": "openai-gpt-5.4-mini",
"gpt-5.4 nano": "openai-gpt-5.4-nano",
"gpt-5.4 pro": "openai-gpt-5.4-pro",
"gpt-5.3-codex": "openai-gpt-5.3-codex",
"gpt-5.2": "openai-gpt-5.2",
"gpt-5.2 pro": "openai-gpt-5.2-pro",
"gpt-5.1-codex-max": "openai-gpt-5.1-codex-max",
"gpt-5": "openai-gpt-5",
"gpt-5 mini": "openai-gpt-5-mini",
"gpt-5 nano": "openai-gpt-5-nano",
"gpt-4.1": "openai-gpt-4.1",
"gpt image 1": "openai-gpt-image-1",
"gpt image 1.5": "openai-gpt-image-1.5",
"gpt-oss-120b": "openai-gpt-oss-120b",
"gpt-oss-20b": "openai-gpt-oss-20b",
"gpt-4o": "openai-gpt-4o",
"gpt-4o mini": "openai-gpt-4o-mini",
"o1": "openai-o1",
"o3-mini": "openai-o3-mini",
// DeepSeek
"deepseek r1 distill llama 70b": "deepseek-r1-distill-llama-70b",
// Llama
"llama 3.3 70b": "llama3.3-70b-instruct",
// DO-hosted
"qwen3-32b": "alibaba-qwen3-32b",
"minimax m2.5 (public preview)": "minimax-m2.5",
"kimi k2.5": "kimi-k2.5",
"nvidia nemotron 3 super 120b (public preview)": "nvidia-nemotron-3-super-120b",
"glm 5": "glm-5",
};
function normalizeDisplayName(raw: string): string {
// Strip " Input Tokens" / " Output Tokens" suffix and lowercase
return raw
.replace(/\s+(input|output)\s+tokens$/i, "")
.trim()
.toLowerCase();
}
function buildPricingMap(entries: z.infer<typeof PricingEntry>[]): Map<string, ModelPricing> {
const map = new Map<string, ModelPricing>();
for (const entry of entries) {
const displayName = normalizeDisplayName(entry.name);
const modelId = PRICING_NAME_MAP[displayName];
if (!modelId) continue;
const isInput = entry.name.toLowerCase().includes("input tokens");
const isOver200k = entry.prompt_tokens === ">200k";
// Round to avoid float noise (e.g. 0.9900000000000001)
const rate = Math.round(entry.price.rate * 10000) / 10000;
const existing = map.get(modelId) ?? ({} as ModelPricing);
if (isInput && isOver200k) existing.inputOver200k = rate;
else if (!isInput && isOver200k) existing.outputOver200k = rate;
else if (isInput) existing.input = rate;
else existing.output = rate;
map.set(modelId, existing);
}
return map;
}
// ---------------------------------------------------------------------------
// Existing TOML shape (fields we read and may preserve)
// ---------------------------------------------------------------------------
interface ExistingModel {
name?: string;
family?: string;
attachment?: boolean;
reasoning?: boolean;
tool_call?: boolean;
structured_output?: boolean;
temperature?: boolean;
knowledge?: string;
release_date?: string;
last_updated?: string;
open_weights?: boolean;
interleaved?: boolean | { field: string };
status?: string;
cost?: {
input?: number;
output?: number;
cache_read?: number;
cache_write?: number;
context_over_200k?: {
input?: number;
output?: number;
cache_read?: number;
cache_write?: number;
context_min?: number;
};
tiers?: Array<{
tier: {
type?: "context";
size: number;
};
input?: number;
output?: number;
cache_read?: number;
cache_write?: number;
}>;
};
limit?: {
context?: number;
input?: number;
output?: number;
};
modalities?: {
input?: string[];
output?: string[];
};
}
async function loadExisting(filePath: string): Promise<ExistingModel | null> {
const file = Bun.file(filePath);
if (!(await file.exists())) return null;
try {
const mod = await import(filePath, { with: { type: "toml" } });
return mod.default as ExistingModel;
} catch (e) {
console.warn(`Warning: failed to parse ${filePath}:`, e);
return null;
}
}
// ---------------------------------------------------------------------------
// Merged model shape (what we write)
// ---------------------------------------------------------------------------
interface MergedModel {
name: string;
family?: string;
attachment: boolean;
reasoning: boolean;
tool_call: boolean;
structured_output?: boolean;
temperature: boolean;
knowledge?: string;
release_date: string;
last_updated: string;
open_weights: boolean;
interleaved?: boolean | { field: string };
status?: string;
cost?: {
input: number;
output: number;
cache_read?: number;
cache_write?: number;
context_over_200k?: {
input: number;
output: number;
cache_read?: number;
cache_write?: number;
context_min?: number;
};
};
limit: {
context: number;
output: number;
};
modalities: {
input: string[];
output: string[];
};
}
// ---------------------------------------------------------------------------
// Helpers
// ---------------------------------------------------------------------------
const VALID_INPUT_MODALITIES = new Set(["text", "audio", "image", "video", "pdf"]);
const VALID_OUTPUT_MODALITIES = new Set(["text", "audio", "image", "video", "pdf"]);
function filterInputModalities(raw: string[]): string[] {
return raw.filter((m) => VALID_INPUT_MODALITIES.has(m));
}
function filterOutputModalities(raw: string[]): string[] {
// "code" is not a valid modality in the schema — map to "text"
return [...new Set(raw.map((m) => (m === "code" ? "text" : m)).filter((m) => VALID_OUTPUT_MODALITIES.has(m)))];
}
function getTodayDate(): string {
return new Date().toISOString().slice(0, 10);
}
function formatNumber(n: number): string {
return n >= 1000 ? n.toString().replace(/\B(?=(\d{3})+(?!\d))/g, "_") : n.toString();
}
function inferFamily(modelId: string, modelName: string): string | undefined {
const sorted = [...ModelFamilyValues].sort((a, b) => b.length - a.length);
const targets = [modelId.toLowerCase(), modelName.toLowerCase()];
for (const family of sorted) {
const f = family.toLowerCase();
for (const t of targets) {
if (t.includes(f)) return family;
}
}
return undefined;
}
function getExistingLongContextCost(existing: ExistingModel | null) {
const tier = existing?.cost?.tiers?.find(
(tier) =>
(tier.tier.type === undefined || tier.tier.type === "context") &&
tier.tier.size >= 200_000,
);
if (tier) {
return {
...tier,
context_min: tier.tier.size,
};
}
return existing?.cost?.context_over_200k === undefined
? undefined
: {
...existing.cost.context_over_200k,
context_min: 200_000,
};
}
function getLongContextMin(cost: { context_min?: number }) {
return cost.context_min ?? 200_000;
}
function formatInlineNumber(n: number): string {
return n >= 1000 ? n.toString().replace(/\B(?=(\d{3})+(?!\d))/g, "_") : n.toString();
}
// ---------------------------------------------------------------------------
// Merge API data with existing TOML
// ---------------------------------------------------------------------------
function mergeModel(
apiModel: z.infer<typeof DoModel>,
pricing: ModelPricing | undefined,
existing: ExistingModel | null,
): MergedModel {
const rawInput = apiModel.modalities?.input ?? [];
const rawOutput = apiModel.modalities?.output ?? [];
const inputMods = filterInputModalities(rawInput.length > 0 ? rawInput : existing?.modalities?.input ?? ["text"]);
const outputMods = filterOutputModalities(rawOutput.length > 0 ? rawOutput : existing?.modalities?.output ?? ["text"]);
const maxTokensSetting = apiModel.settings?.find((s) => s.name === "max_tokens");
const maxTokens = maxTokensSetting?.max ?? existing?.limit?.output ?? 0;
const rawContext = apiModel.context_window;
const contextWindow =
rawContext !== undefined
? typeof rawContext === "string"
? parseInt(rawContext, 10)
: rawContext
: (existing?.limit?.context ?? 0);
const isDeprecated = apiModel.lifecycle_status === "end_of_life";
// Fields preserved from existing TOML (APIs don't provide these)
const family = existing?.family ?? inferFamily(apiModel.id, apiModel.name);
const knowledge = existing?.knowledge;
const openWeights = existing?.open_weights ?? false;
const interleaved = existing?.interleaved;
const attachment = existing?.attachment ?? inputMods.some((m) => m !== "text");
// reasoning: trust existing if set, else use API thinking flag as a hint
// (thinking flag is unreliable for non-LLM models so gate on output modality)
const isTextOutput = outputMods.includes("text") && !outputMods.includes("image") && !outputMods.includes("video");
const reasoning = existing?.reasoning ?? (isTextOutput && (apiModel.thinking ?? false));
// tool_call: no API signal, preserve existing or default true for text models
const toolCall = existing?.tool_call ?? isTextOutput;
// temperature: no API signal, preserve or default true
const temperature = existing?.temperature ?? true;
// structured_output: no API signal, preserve only
const structuredOutput = existing?.structured_output;
const releaseDate = existing?.release_date ?? apiModel.created_at?.slice(0, 10) ?? getTodayDate();
const merged: MergedModel = {
name: apiModel.name,
family,
attachment,
reasoning,
tool_call: toolCall,
temperature,
release_date: releaseDate,
last_updated: getTodayDate(),
open_weights: openWeights,
...(structuredOutput !== undefined && { structured_output: structuredOutput }),
...(knowledge && { knowledge }),
...(interleaved !== undefined && { interleaved }),
...(isDeprecated && { status: "deprecated" }),
limit: { context: contextWindow, output: maxTokens },
modalities: { input: inputMods, output: outputMods },
};
// Pricing: static-content API is the sole source of truth for prices.
// The v2 models API pricing is intentionally ignored. If a model has no
// entry in the static-content API, preserve existing TOML prices.
const inputPrice = pricing?.input ?? existing?.cost?.input;
const outputPrice = pricing?.output ?? existing?.cost?.output;
if (inputPrice !== undefined && outputPrice !== undefined) {
merged.cost = {
input: inputPrice,
output: outputPrice,
// Always preserve cache pricing — not available from any DO API
...(existing?.cost?.cache_read !== undefined && { cache_read: existing.cost.cache_read }),
...(existing?.cost?.cache_write !== undefined && { cache_write: existing.cost.cache_write }),
};
// Context-tiered pricing (>200k) from the static-content API
const existingLongContextCost = getExistingLongContextCost(existing);
if (pricing?.inputOver200k !== undefined && pricing?.outputOver200k !== undefined) {
merged.cost.context_over_200k = {
input: pricing.inputOver200k,
output: pricing.outputOver200k,
context_min: existingLongContextCost?.context_min ?? 200_000,
...(existingLongContextCost?.cache_read !== undefined && {
cache_read: existingLongContextCost.cache_read,
}),
...(existingLongContextCost?.cache_write !== undefined && {
cache_write: existingLongContextCost.cache_write,
}),
};
} else if (existingLongContextCost) {
// Preserve manually-entered tiered pricing if API has no data
merged.cost.context_over_200k = {
input: existingLongContextCost.input ?? inputPrice,
output: existingLongContextCost.output ?? outputPrice,
context_min: existingLongContextCost.context_min,
...(existingLongContextCost.cache_read !== undefined && {
cache_read: existingLongContextCost.cache_read,
}),
...(existingLongContextCost.cache_write !== undefined && {
cache_write: existingLongContextCost.cache_write,
}),
};
}
}
return merged;
}
// ---------------------------------------------------------------------------
// TOML serialiser
// ---------------------------------------------------------------------------
function formatToml(model: MergedModel): string {
const lines: string[] = [];
lines.push(`name = "${model.name.replace(/"/g, '\\"')}"`);
if (model.family) lines.push(`family = "${model.family}"`);
lines.push(`release_date = "${model.release_date}"`);
lines.push(`last_updated = "${model.last_updated}"`);
lines.push(`attachment = ${model.attachment}`);
lines.push(`reasoning = ${model.reasoning}`);
lines.push(`temperature = ${model.temperature}`);
lines.push(`tool_call = ${model.tool_call}`);
if (model.structured_output !== undefined) lines.push(`structured_output = ${model.structured_output}`);
if (model.knowledge) lines.push(`knowledge = "${model.knowledge}"`);
lines.push(`open_weights = ${model.open_weights}`);
if (model.status) lines.push(`status = "${model.status}"`);
if (model.interleaved !== undefined) {
lines.push("");
if (model.interleaved === true) {
lines.push(`interleaved = true`);
} else if (typeof model.interleaved === "object") {
lines.push(`[interleaved]`);
lines.push(`field = "${model.interleaved.field}"`);
}
}
if (model.cost) {
lines.push("");
lines.push(`[cost]`);
lines.push(`input = ${model.cost.input}`);
lines.push(`output = ${model.cost.output}`);
if (model.cost.cache_read !== undefined) lines.push(`cache_read = ${model.cost.cache_read}`);
if (model.cost.cache_write !== undefined) lines.push(`cache_write = ${model.cost.cache_write}`);
if (model.cost.context_over_200k) {
lines.push("");
lines.push(`[[cost.tiers]]`);
lines.push(`tier = { size = ${formatInlineNumber(getLongContextMin(model.cost.context_over_200k))} }`);
lines.push(`input = ${model.cost.context_over_200k.input}`);
lines.push(`output = ${model.cost.context_over_200k.output}`);
if (model.cost.context_over_200k.cache_read !== undefined)
lines.push(`cache_read = ${model.cost.context_over_200k.cache_read}`);
if (model.cost.context_over_200k.cache_write !== undefined)
lines.push(`cache_write = ${model.cost.context_over_200k.cache_write}`);
}
}
lines.push("");
lines.push(`[limit]`);
lines.push(`context = ${formatNumber(model.limit.context)}`);
lines.push(`output = ${formatNumber(model.limit.output)}`);
lines.push("");
lines.push(`[modalities]`);
lines.push(`input = [${model.modalities.input.map((m) => `"${m}"`).join(", ")}]`);
lines.push(`output = [${model.modalities.output.map((m) => `"${m}"`).join(", ")}]`);
return lines.join("\n") + "\n";
}
// ---------------------------------------------------------------------------
// Change detection
// ---------------------------------------------------------------------------
interface Change {
field: string;
oldValue: string;
newValue: string;
}
function formatValue(val: unknown): string {
if (val === undefined) return "(none)";
if (Array.isArray(val)) return `[${val.join(", ")}]`;
if (typeof val === "number") return formatNumber(val);
return String(val);
}
function detectChanges(existing: ExistingModel | null, merged: MergedModel): Change[] {
if (!existing) return [];
const changes: Change[] = [];
const EPSILON = 0.001;
const compare = (field: string, oldVal: unknown, newVal: unknown) => {
if (oldVal === undefined && newVal === undefined) return;
const isDiff = field.startsWith("cost.")
? Math.abs((oldVal as number ?? 0) - (newVal as number ?? 0)) > EPSILON
: JSON.stringify(oldVal) !== JSON.stringify(newVal);
if (isDiff) changes.push({ field, oldValue: formatValue(oldVal), newValue: formatValue(newVal) });
};
compare("name", existing.name, merged.name);
compare("reasoning", existing.reasoning, merged.reasoning);
compare("tool_call", existing.tool_call, merged.tool_call);
compare("attachment", existing.attachment, merged.attachment);
compare("status", existing.status, merged.status);
compare("cost.input", existing.cost?.input, merged.cost?.input);
compare("cost.output", existing.cost?.output, merged.cost?.output);
const existingLongContextCost = getExistingLongContextCost(existing);
compare("cost.context_over_200k.input", existingLongContextCost?.input, merged.cost?.context_over_200k?.input);
compare("cost.context_over_200k.output", existingLongContextCost?.output, merged.cost?.context_over_200k?.output);
compare("limit.context", existing.limit?.context, merged.limit.context);
compare("limit.output", existing.limit?.output, merged.limit.output);
compare("modalities.input", existing.modalities?.input, merged.modalities.input);
compare("modalities.output", existing.modalities?.output, merged.modalities.output);
return changes;
}
// ---------------------------------------------------------------------------
// Main
// ---------------------------------------------------------------------------
async function main() {
const args = process.argv.slice(2);
const dryRun = args.includes("--dry-run");
const newOnly = args.includes("--new-only");
// Resolve API key
const apiKeyArg = args.find((a) => a.startsWith("--api-key"));
const apiKey =
(apiKeyArg?.includes("=") ? apiKeyArg.split("=")[1] : args[args.indexOf(apiKeyArg!) + 1]) ??
process.env.DIGITALOCEAN_API_TOKEN;
if (!apiKey) {
console.error("Error: DIGITALOCEAN_API_TOKEN is required (or pass --api-key=<key>)");
console.error("Get one from: https://cloud.digitalocean.com/account/api/tokens");
process.exit(1);
}
const modelsDir = path.join(import.meta.dirname, "..", "..", "..", "providers", "digitalocean", "models");
const prefix = dryRun ? "[DRY RUN] " : "";
console.log(`${prefix}Fetching DigitalOcean models from API...`);
// Fetch both APIs in parallel
const [modelsRes, pricingRes] = await Promise.all([
fetch(MODELS_API, { headers: { Authorization: `Bearer ${apiKey}`, "Content-Type": "application/json" } }),
fetch(PRICING_API, { headers: { "User-Agent": "models.dev/digitalocean-sync" } }),
]);
if (!modelsRes.ok) {
console.error(`Failed to fetch models API: ${modelsRes.status} ${modelsRes.statusText}`);
if (modelsRes.status === 401 || modelsRes.status === 403)
console.error("Check your DIGITALOCEAN_API_TOKEN has read access.");
process.exit(1);
}
if (!pricingRes.ok) {
console.error(`Failed to fetch pricing API: ${pricingRes.status} ${pricingRes.statusText}`);
process.exit(1);
}
const modelsParsed = DoModelsResponse.safeParse(await modelsRes.json());
if (!modelsParsed.success) {
console.error("Unexpected models API response:", modelsParsed.error.errors);
process.exit(1);
}
const pricingParsed = StaticContentResponse.safeParse(await pricingRes.json());
if (!pricingParsed.success) {
console.error("Unexpected pricing API response:", pricingParsed.error.errors);
process.exit(1);
}
const apiModels = modelsParsed.data.models;
const pricingMap = buildPricingMap(pricingParsed.data.gradient.models);
// Collect existing TOML filenames for orphan detection
const existingFiles = new Set<string>();
for await (const file of new Bun.Glob("**/*.toml").scan({ cwd: modelsDir, absolute: false })) {
existingFiles.add(file);
}
console.log(`Found ${apiModels.length} models in API, ${existingFiles.size} existing TOML files\n`);
const apiModelFiles = new Set<string>();
let created = 0;
let updated = 0;
let unchanged = 0;
for (const apiModel of apiModels) {
// Skip non-text models that opencode can't use: image, video, audio, embedding, reranking
const outputMods = filterOutputModalities(apiModel.modalities?.output ?? []);
const isTextModel = outputMods.includes("text");
const isEmbedding = apiModel.type === "embedding";
const isReranking = apiModel.type === "reranking";
if (!isTextModel || isEmbedding || isReranking) continue;
// Model IDs may contain slashes (e.g. fal-ai/flux/schnell) — use as subpath
const relativePath = `${apiModel.id}.toml`;
const filePath = path.join(modelsDir, relativePath);
const dirPath = path.dirname(filePath);
apiModelFiles.add(relativePath);
const existing = await loadExisting(filePath);
const pricing = pricingMap.get(apiModel.id);
const merged = mergeModel(apiModel, pricing, existing);
const toml = formatToml(merged);
if (existing === null) {
created++;
if (dryRun) {
console.log(`[DRY RUN] Would create: ${relativePath}`);
console.log(` name = "${merged.name}"`);
if (pricing) console.log(` pricing: $${merged.cost?.input}/$${merged.cost?.output} per M tokens`);
if (merged.family) console.log(` family = "${merged.family}" (inferred)`);
console.log("");
} else {
await mkdir(dirPath, { recursive: true });
await Bun.write(filePath, toml);
console.log(`Created: ${relativePath}`);
}
continue;
}
if (newOnly) {
unchanged++;
continue;
}
const changes = detectChanges(existing, merged);
if (changes.length > 0) {
updated++;
if (dryRun) {
console.log(`[DRY RUN] Would update: ${relativePath}`);
} else {
await mkdir(dirPath, { recursive: true });
await Bun.write(filePath, toml);
console.log(`Updated: ${relativePath}`);
}
for (const c of changes) console.log(` ${c.field}: ${c.oldValue}${c.newValue}`);
console.log("");
} else {
unchanged++;
}
}
// Orphan detection: files in the TOML directory but not in the API
const orphaned: string[] = [];
for (const file of existingFiles) {
if (!apiModelFiles.has(file)) {
orphaned.push(file);
console.log(`Warning: orphaned file (not in API): ${file}`);
}
}
console.log("");
if (dryRun) {
console.log(
`Summary: ${created} would be created, ${updated} would be updated, ${unchanged} unchanged, ${orphaned.length} orphaned`,
);
} else {
console.log(`Summary: ${created} created, ${updated} updated, ${unchanged} unchanged, ${orphaned.length} orphaned`);
}
}
await main();
+45 -8
View File
@@ -162,7 +162,18 @@ interface ExistingModel {
output?: number;
cache_read?: number;
cache_write?: number;
context_min?: number;
};
tiers?: Array<{
tier: {
type?: "context";
size: number;
};
input?: number;
output?: number;
cache_read?: number;
cache_write?: number;
}>;
};
limit?: {
context?: number;
@@ -195,6 +206,30 @@ async function loadExistingModel(filePath: string): Promise<ExistingModel | null
}
}
function getExistingLongContextMin(existing: ExistingModel | null) {
return (
existing?.cost?.tiers?.find(
(tier) =>
(tier.tier.type === undefined || tier.tier.type === "context") &&
tier.tier.size >= 200_000,
)?.tier.size ?? 200_000
);
}
function getExistingLongContextCost(existing: ExistingModel | null) {
return (
existing?.cost?.tiers?.find(
(tier) =>
(tier.tier.type === undefined || tier.tier.type === "context") &&
tier.tier.size >= 200_000,
) ?? existing?.cost?.context_over_200k
);
}
function getLongContextMin(cost: { context_min?: number }) {
return cost.context_min ?? 200_000;
}
interface MergedModel {
name: string;
family?: string;
@@ -219,6 +254,7 @@ interface MergedModel {
output: number;
cache_read?: number;
cache_write?: number;
context_min?: number;
};
};
limit: {
@@ -241,9 +277,7 @@ function mergeModel(
const contextTokens = spec.availableContextTokens;
const outputTokens = spec.maxCompletionTokens ?? Math.floor(contextTokens / 4);
const openWeights = spec.modelSource
? spec.modelSource.toLowerCase().includes("huggingface")
: spec.privacy === "private";
const openWeights = spec.modelSource?.toLowerCase().includes("huggingface") ?? false;
const inputModalities = buildInputModalities(caps);
@@ -294,6 +328,7 @@ function mergeModel(
merged.cost.context_over_200k = {
input: spec.pricing.extended.input.usd,
output: spec.pricing.extended.output.usd,
context_min: spec.pricing.extended.context_token_threshold,
...(spec.pricing.extended.cache_input && { cache_read: spec.pricing.extended.cache_input.usd }),
...(spec.pricing.extended.cache_write && { cache_write: spec.pricing.extended.cache_write.usd }),
};
@@ -368,7 +403,8 @@ function formatToml(model: MergedModel): string {
if (model.cost.context_over_200k) {
lines.push("");
lines.push(`[cost.context_over_200k]`);
lines.push(`[[cost.tiers]]`);
lines.push(`tier = { size = ${formatNumber(getLongContextMin(model.cost.context_over_200k))} }`);
lines.push(`input = ${model.cost.context_over_200k.input}`);
lines.push(`output = ${model.cost.context_over_200k.output}`);
if (model.cost.context_over_200k.cache_read !== undefined) {
@@ -440,10 +476,11 @@ function detectChanges(
compare("cost.output", existing.cost?.output, merged.cost?.output);
compare("cost.cache_read", existing.cost?.cache_read, merged.cost?.cache_read);
compare("cost.cache_write", existing.cost?.cache_write, merged.cost?.cache_write);
compare("cost.context_over_200k.input", existing.cost?.context_over_200k?.input, merged.cost?.context_over_200k?.input);
compare("cost.context_over_200k.output", existing.cost?.context_over_200k?.output, merged.cost?.context_over_200k?.output);
compare("cost.context_over_200k.cache_read", existing.cost?.context_over_200k?.cache_read, merged.cost?.context_over_200k?.cache_read);
compare("cost.context_over_200k.cache_write", existing.cost?.context_over_200k?.cache_write, merged.cost?.context_over_200k?.cache_write);
const existingLongContextCost = getExistingLongContextCost(existing);
compare("cost.context_over_200k.input", existingLongContextCost?.input, merged.cost?.context_over_200k?.input);
compare("cost.context_over_200k.output", existingLongContextCost?.output, merged.cost?.context_over_200k?.output);
compare("cost.context_over_200k.cache_read", existingLongContextCost?.cache_read, merged.cost?.context_over_200k?.cache_read);
compare("cost.context_over_200k.cache_write", existingLongContextCost?.cache_write, merged.cost?.context_over_200k?.cache_write);
compare("limit.context", existing.limit?.context, merged.limit.context);
compare("limit.output", existing.limit?.output, merged.limit.output);
compare("modalities.input", existing.modalities?.input, merged.modalities.input);
+19 -2
View File
@@ -20,10 +20,12 @@ enum ModelType {
Embedding = "embedding",
Image = "image",
Video = "video",
Reranking = "reranking",
}
enum SkipZeroFields {
LimitContext = "limit.context",
LimitInput = "limit.input",
LimitOutput = "limit.output",
}
@@ -82,6 +84,7 @@ interface ExistingModel {
};
limit?: {
context?: number;
input?: number;
output?: number;
};
modalities?: {
@@ -112,6 +115,7 @@ interface MergedModel {
};
limit: {
context: number;
input?: number;
output: number;
};
modalities: {
@@ -232,6 +236,10 @@ async function loadExistingModel(filePath: string): Promise<ExistingModel | null
}
}
function isOpenAIModel(modelId: string): boolean {
return modelId.startsWith("openai/");
}
function mergeModel(
apiModel: z.infer<typeof VercelModel>,
existing: ExistingModel | null,
@@ -281,6 +289,7 @@ function mergeModel(
...(status && { status }),
limit: {
context: contextLimit,
...(isOpenAIModel(apiModel.id) && contextLimit > outputLimit && { input: contextLimit - outputLimit }),
output: outputLimit,
},
modalities: {
@@ -362,6 +371,9 @@ function formatToml(model: MergedModel): string {
lines.push("");
lines.push(`[limit]`);
lines.push(`context = ${formatNumber(model.limit.context)}`);
if (model.limit.input !== undefined) {
lines.push(`input = ${formatNumber(model.limit.input)}`);
}
lines.push(`output = ${formatNumber(model.limit.output)}`);
lines.push("");
@@ -435,6 +447,7 @@ function detectChanges(
compare("cost.cache_read", existing.cost?.cache_read, merged.cost?.cache_read);
compare("cost.cache_write", existing.cost?.cache_write, merged.cost?.cache_write);
compare("limit.context", existing.limit?.context, merged.limit.context);
compare("limit.input", existing.limit?.input, merged.limit.input);
compare("limit.output", existing.limit?.output, merged.limit.output);
compare("modalities.input", existing.modalities?.input, merged.modalities.input);
@@ -493,8 +506,12 @@ async function main() {
let unchanged = 0;
for (const apiModel of apiModels) {
// Skip these since OpenCode does not support image / video generation yet
if (apiModel.type === ModelType.Image || apiModel.type === ModelType.Video) {
// Skip these since OpenCode does not support image / video / reranking yet
if (
apiModel.type === ModelType.Image ||
apiModel.type === ModelType.Video ||
apiModel.type === ModelType.Reranking
) {
continue;
}
+525
View File
@@ -0,0 +1,525 @@
#!/usr/bin/env bun
import path from "node:path";
import { mkdir } from "node:fs/promises";
import { z } from "zod";
import { ModelFamilyValues } from "../src/family.js";
const API_ENDPOINT = "https://trace.wandb.ai/inference/analysis/artificialanalysis/models";
const Pricing = z
.object({
prompt: z.string().optional(),
completion: z.string().optional(),
image: z.string().optional(),
request: z.string().optional(),
input_cache_reads: z.string().optional(),
input_cache_writes: z.string().optional(),
})
.passthrough();
const WandbModel = z
.object({
id: z.string(),
name: z.string(),
created: z.number(),
input_modalities: z.array(z.string()),
output_modalities: z.array(z.string()),
context_length: z.number(),
max_output_length: z.number(),
pricing: Pricing.optional(),
supported_sampling_parameters: z.array(z.string()).default([]),
supported_features: z.array(z.string()).default([]),
})
.passthrough();
const WandbResponse = z
.object({
data: z.array(WandbModel),
})
.strict();
interface ExistingModel {
name?: string;
family?: string;
attachment?: boolean;
reasoning?: boolean;
tool_call?: boolean;
structured_output?: boolean;
temperature?: boolean;
knowledge?: string;
release_date?: string;
last_updated?: string;
open_weights?: boolean;
interleaved?: boolean | { field: string };
status?: string;
cost?: {
input?: number;
output?: number;
cache_read?: number;
cache_write?: number;
};
limit?: {
context?: number;
input?: number;
output?: number;
};
modalities?: {
input?: string[];
output?: string[];
};
}
interface MergedModel {
name: string;
family?: string;
attachment: boolean;
reasoning: boolean;
tool_call: boolean;
structured_output?: boolean;
temperature: boolean;
knowledge?: string;
release_date: string;
last_updated: string;
open_weights: boolean;
interleaved?: boolean | { field: string };
status?: string;
cost?: {
input: number;
output: number;
cache_read?: number;
cache_write?: number;
};
limit: {
context: number;
output: number;
};
modalities: {
input: Array<"text" | "audio" | "image" | "video" | "pdf">;
output: Array<"text" | "audio" | "image" | "video" | "pdf">;
};
}
interface Changes {
field: string;
oldValue: string;
newValue: string;
}
type SupportedModality = "text" | "audio" | "image" | "video" | "pdf";
const modalityMap: Record<string, SupportedModality | undefined> = {
text: "text",
image: "image",
audio: "audio",
video: "video",
pdf: "pdf",
file: "pdf",
files: "pdf",
};
const openWeightsPrefixes = new Set([
"deepseek-ai/",
"meta-llama/",
"microsoft/",
"MiniMaxAI/",
"moonshotai/",
"nvidia/",
"OpenPipe/",
"Qwen/",
"zai-org/",
]);
function timestampToDate(timestamp: number): string {
return new Date(timestamp * 1000).toISOString().slice(0, 10);
}
function getTodayDate(): string {
return new Date().toISOString().slice(0, 10);
}
function formatNumber(n: number): string {
if (n >= 1000) {
return n.toString().replace(/\B(?=(\d{3})+(?!\d))/g, "_");
}
return n.toString();
}
function formatDecimal(n: number): string {
return Number(n.toFixed(6)).toString();
}
function priceToPerMillion(value: string): number {
return Number((parseFloat(value) * 1_000_000).toFixed(6));
}
function isSubstring(target: string, family: string): boolean {
return target.toLowerCase().includes(family.toLowerCase());
}
function matchesFamily(target: string, family: string): boolean {
const targetLower = target.toLowerCase();
const familyLower = family.toLowerCase();
let familyIdx = 0;
for (let i = 0; i < targetLower.length && familyIdx < familyLower.length; i++) {
if (targetLower[i] === familyLower[familyIdx]) {
familyIdx++;
}
}
return familyIdx === familyLower.length;
}
function inferFamily(modelId: string, modelName: string): string | undefined {
const sortedFamilies = [...ModelFamilyValues].sort((a, b) => b.length - a.length);
for (const family of sortedFamilies) {
if (isSubstring(modelId, family) || isSubstring(modelName, family)) {
return family;
}
}
for (const family of sortedFamilies) {
if (matchesFamily(modelId, family) || matchesFamily(modelName, family)) {
return family;
}
}
return undefined;
}
function normalizeName(apiModel: z.infer<typeof WandbModel>): string {
const stripped = apiModel.name.replace(/^[^:]+:\s*/, "").trim();
return stripped || path.basename(apiModel.id);
}
function inferReasoning(apiModel: z.infer<typeof WandbModel>): boolean {
const text = `${apiModel.id} ${apiModel.name}`.toLowerCase();
return text.includes("thinking") || /\br1\b/.test(text) || text.includes("reasoning");
}
function inferOpenWeights(modelId: string): boolean {
for (const prefix of openWeightsPrefixes) {
if (modelId.startsWith(prefix)) {
return true;
}
}
return false;
}
function normalizeModalities(values: string[]): SupportedModality[] {
const normalized = values
.map((value) => modalityMap[value.toLowerCase()])
.filter((value): value is SupportedModality => value !== undefined);
return [...new Set(normalized)];
}
async function loadExistingModel(filePath: string): Promise<ExistingModel | null> {
try {
const file = Bun.file(filePath);
if (!(await file.exists())) {
return null;
}
const toml = await import(filePath, { with: { type: "toml" } }).then((mod) => mod.default);
return toml as ExistingModel;
} catch (cause) {
console.warn(`Warning: Failed to parse existing file ${filePath}:`, cause);
return null;
}
}
function mergeModel(
apiModel: z.infer<typeof WandbModel>,
existing: ExistingModel | null,
): MergedModel {
const featureSet = new Set(apiModel.supported_features);
const samplingSet = new Set(apiModel.supported_sampling_parameters);
const inputModalities = normalizeModalities(apiModel.input_modalities);
const outputModalities = normalizeModalities(apiModel.output_modalities);
const merged: MergedModel = {
name: existing?.name ?? normalizeName(apiModel),
family: existing?.family ?? inferFamily(apiModel.id, apiModel.name),
attachment: existing?.attachment ?? inputModalities.some((m) => m !== "text"),
reasoning: existing?.reasoning ?? inferReasoning(apiModel),
tool_call: existing?.tool_call ?? featureSet.has("tools"),
temperature: existing?.temperature ?? samplingSet.has("temperature"),
release_date: existing?.release_date ?? timestampToDate(apiModel.created),
last_updated: getTodayDate(),
open_weights: existing?.open_weights ?? inferOpenWeights(apiModel.id),
...(existing?.structured_output !== undefined
? { structured_output: existing.structured_output }
: featureSet.has("structured_outputs")
? { structured_output: true }
: {}),
...(existing?.knowledge ? { knowledge: existing.knowledge } : {}),
...(existing?.interleaved !== undefined ? { interleaved: existing.interleaved } : {}),
...(existing?.status ? { status: existing.status } : {}),
limit: {
context: apiModel.context_length > 0 ? apiModel.context_length : (existing?.limit?.context ?? 0),
output: apiModel.max_output_length > 0
? apiModel.max_output_length
: (existing?.limit?.output ?? 0),
},
modalities: {
input: inputModalities.length > 0
? inputModalities
: ((existing?.modalities?.input as SupportedModality[] | undefined) ?? ["text"]),
output: outputModalities.length > 0
? outputModalities
: ((existing?.modalities?.output as SupportedModality[] | undefined) ?? ["text"]),
},
};
const prompt = apiModel.pricing?.prompt;
const completion = apiModel.pricing?.completion;
const cacheRead = apiModel.pricing?.input_cache_reads;
const cacheWrite = apiModel.pricing?.input_cache_writes;
if (prompt && completion) {
merged.cost = {
input: priceToPerMillion(prompt),
output: priceToPerMillion(completion),
...(cacheRead && parseFloat(cacheRead) > 0
? { cache_read: priceToPerMillion(cacheRead) }
: {}),
...(cacheWrite && parseFloat(cacheWrite) > 0
? { cache_write: priceToPerMillion(cacheWrite) }
: {}),
};
} else if (existing?.cost?.input !== undefined && existing.cost.output !== undefined) {
merged.cost = {
input: existing.cost.input,
output: existing.cost.output,
...(existing.cost.cache_read !== undefined ? { cache_read: existing.cost.cache_read } : {}),
...(existing.cost.cache_write !== undefined ? { cache_write: existing.cost.cache_write } : {}),
};
}
return merged;
}
function formatToml(model: MergedModel): string {
const lines: string[] = [];
lines.push(`name = "${model.name.replace(/"/g, '\\"')}"`);
if (model.family) {
lines.push(`family = "${model.family}"`);
}
lines.push(`release_date = "${model.release_date}"`);
lines.push(`last_updated = "${model.last_updated}"`);
lines.push(`attachment = ${model.attachment}`);
lines.push(`reasoning = ${model.reasoning}`);
if (model.structured_output !== undefined) {
lines.push(`structured_output = ${model.structured_output}`);
}
lines.push(`temperature = ${model.temperature}`);
lines.push(`tool_call = ${model.tool_call}`);
if (model.knowledge) {
lines.push(`knowledge = "${model.knowledge}"`);
}
lines.push(`open_weights = ${model.open_weights}`);
if (model.status) {
lines.push(`status = "${model.status}"`);
}
if (model.interleaved !== undefined) {
lines.push("");
if (model.interleaved === true) {
lines.push("interleaved = true");
} else {
lines.push("[interleaved]");
lines.push(`field = "${model.interleaved.field}"`);
}
}
if (model.cost) {
lines.push("");
lines.push("[cost]");
lines.push(`input = ${formatDecimal(model.cost.input)}`);
lines.push(`output = ${formatDecimal(model.cost.output)}`);
if (model.cost.cache_read !== undefined) {
lines.push(`cache_read = ${formatDecimal(model.cost.cache_read)}`);
}
if (model.cost.cache_write !== undefined) {
lines.push(`cache_write = ${formatDecimal(model.cost.cache_write)}`);
}
}
lines.push("");
lines.push("[limit]");
lines.push(`context = ${formatNumber(model.limit.context)}`);
lines.push(`output = ${formatNumber(model.limit.output)}`);
lines.push("");
lines.push("[modalities]");
lines.push(`input = [${model.modalities.input.map((m) => `"${m}"`).join(", ")}]`);
lines.push(`output = [${model.modalities.output.map((m) => `"${m}"`).join(", ")}]`);
return `${lines.join("\n")}\n`;
}
function detectChanges(existing: ExistingModel | null, merged: MergedModel): Changes[] {
if (!existing) {
return [];
}
const changes: Changes[] = [];
const epsilon = 0.001;
const formatValue = (value: unknown): string => {
if (typeof value === "number") return formatNumber(value);
if (Array.isArray(value)) return `[${value.join(", ")}]`;
if (value === undefined) return "(none)";
return String(value);
};
const compare = (field: string, oldValue: unknown, newValue: unknown) => {
const changed = field.startsWith("cost.")
? (
oldValue === undefined && newValue === undefined
? false
: oldValue === undefined || newValue === undefined
? true
: Math.abs((oldValue as number) - (newValue as number)) > epsilon
)
: JSON.stringify(oldValue) !== JSON.stringify(newValue);
if (changed) {
changes.push({
field,
oldValue: formatValue(oldValue),
newValue: formatValue(newValue),
});
}
};
compare("name", existing.name, merged.name);
compare("family", existing.family, merged.family);
compare("release_date", existing.release_date, merged.release_date);
compare("attachment", existing.attachment, merged.attachment);
compare("reasoning", existing.reasoning, merged.reasoning);
compare("structured_output", existing.structured_output, merged.structured_output);
compare("temperature", existing.temperature, merged.temperature);
compare("tool_call", existing.tool_call, merged.tool_call);
compare("open_weights", existing.open_weights, merged.open_weights);
compare("cost.input", existing.cost?.input, merged.cost?.input);
compare("cost.output", existing.cost?.output, merged.cost?.output);
compare("cost.cache_read", existing.cost?.cache_read, merged.cost?.cache_read);
compare("cost.cache_write", existing.cost?.cache_write, merged.cost?.cache_write);
compare("limit.context", existing.limit?.context, merged.limit.context);
compare("limit.output", existing.limit?.output, merged.limit.output);
compare("modalities.input", existing.modalities?.input, merged.modalities.input);
compare("modalities.output", existing.modalities?.output, merged.modalities.output);
return changes;
}
async function main() {
const args = process.argv.slice(2);
const dryRun = args.includes("--dry-run");
const newOnly = args.includes("--new-only");
const modelsDir = path.join(import.meta.dirname, "..", "..", "..", "providers", "wandb", "models");
console.log(`${dryRun ? "[DRY RUN] " : ""}${newOnly ? "[NEW ONLY] " : ""}Fetching WandB models from API...`);
const res = await fetch(API_ENDPOINT);
if (!res.ok) {
console.error(`Failed to fetch API: ${res.status} ${res.statusText}`);
process.exit(1);
}
const json = await res.json();
const parsed = WandbResponse.safeParse(json);
if (!parsed.success) {
console.error("Invalid API response:", parsed.error.errors);
process.exit(1);
}
const apiModels = parsed.data.data;
const existingFiles = new Set<string>();
for await (const file of new Bun.Glob("**/*.toml").scan({ cwd: modelsDir, absolute: false })) {
existingFiles.add(file);
}
console.log(`Found ${apiModels.length} models in API, ${existingFiles.size} existing files\n`);
const apiModelIds = new Set<string>();
let created = 0;
let updated = 0;
let unchanged = 0;
for (const apiModel of apiModels) {
const relativePath = `${apiModel.id}.toml`;
const filePath = path.join(modelsDir, relativePath);
const dirPath = path.dirname(filePath);
apiModelIds.add(relativePath);
const existing = await loadExistingModel(filePath);
const merged = mergeModel(apiModel, existing);
const tomlContent = formatToml(merged);
if (existing === null) {
created++;
if (dryRun) {
console.log(`[DRY RUN] Would create: ${relativePath}`);
console.log(` name = "${merged.name}"`);
if (merged.family) {
console.log(` family = "${merged.family}"`);
}
console.log("");
} else {
await mkdir(dirPath, { recursive: true });
await Bun.write(filePath, tomlContent);
console.log(`Created: ${relativePath}`);
}
continue;
}
if (newOnly) {
unchanged++;
continue;
}
const changes = detectChanges(existing, merged);
if (changes.length === 0) {
unchanged++;
continue;
}
updated++;
if (dryRun) {
console.log(`[DRY RUN] Would update: ${relativePath}`);
} else {
await mkdir(dirPath, { recursive: true });
await Bun.write(filePath, tomlContent);
console.log(`Updated: ${relativePath}`);
}
for (const change of changes) {
console.log(` ${change.field}: ${change.oldValue}${change.newValue}`);
}
console.log("");
}
const orphaned = [...existingFiles].filter((file) => !apiModelIds.has(file));
for (const file of orphaned) {
console.log(`Warning: Orphaned file (not in API): ${file}`);
}
console.log("");
console.log(
dryRun
? `Summary: ${created} would be created, ${updated} would be updated, ${unchanged} unchanged, ${orphaned.length} orphaned`
: `Summary: ${created} created, ${updated} updated, ${unchanged} unchanged, ${orphaned.length} orphaned`,
);
}
await main();
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bun
import { main } from "../src/sync/index.js";
await main();
+44 -3
View File
@@ -14,6 +14,7 @@ export const ModelFamilyValues = [
"gpt-mini",
"gpt-nano",
"gpt-oss",
"gpt-image",
// OpenAI o-series (reasoning models)
"o",
@@ -46,16 +47,25 @@ export const ModelFamilyValues = [
// Alibaba Qwen
"qwen",
"qwen3.5",
"qwen3.6",
"qwen3.7-max",
"qwen-free",
// DeepSeek
"deepseek",
"deepseek-thinking",
"deepseek-flash",
"deepseek-flash-free",
"deepseek-flash-think",
// Microsoft Phi
"phi",
// Moonshot Kimi
"kimi",
"kimi-k2.5",
"kimi-k2.6",
"kimi-free",
"kimi-thinking",
@@ -73,6 +83,7 @@ export const ModelFamilyValues = [
// xAI Grok
"grok",
"grok-build",
"grok-vision",
"grok-beta",
@@ -96,6 +107,7 @@ export const ModelFamilyValues = [
// NVIDIA Nemotron
"nemotron",
"nemotron-free",
// AWS Titan
"titan",
@@ -103,11 +115,16 @@ export const ModelFamilyValues = [
// MiniMax
"minimax",
"minimax-m2.5",
"minimax-m2.7",
"minimax-free",
// Hunyuan
"hunyuan",
// Hy
"Hy",
// Yi
"yi",
@@ -193,6 +210,15 @@ export const ModelFamilyValues = [
// Mimo
"mimo",
"mimo-pro",
"mimo-omni",
"mimo-v2-pro",
"mimo-v2-omni",
"mimo-v2.5-pro",
"mimo-v2.5",
"mimo-pro-free",
"mimo-omni-free",
"mimo-flash-free",
// Clarifai
"mm-poly",
@@ -265,9 +291,15 @@ export const ModelFamilyValues = [
// RNJ
"rnj",
// Tecent Hy
"hy3",
"hy3-free",
// Ling & Ring (InclusionAI)
"ling",
"ling-flash-free",
"ring",
"ring-1t-free",
// Kat Coder
"kat-coder",
@@ -284,9 +316,6 @@ export const ModelFamilyValues = [
// Parakeet
"parakeet",
// MiMo
"mimo-flash-free",
// NeMo
"nemoretriever",
@@ -371,6 +400,18 @@ export const ModelFamilyValues = [
// Writer
"palmyra",
// ALLaM
"allam",
// Canopy Labs
"canopylabs",
// Groq
"groq",
// Elephant
"elephant",
] as const;
export const ModelFamily = z.enum(ModelFamilyValues);
+160 -4
View File
@@ -1,9 +1,31 @@
import path from "path";
import { mergeDeep } from "remeda";
import { z } from "zod";
import { Provider, Model } from "./schema.js";
import { Provider, Model, AuthoredModel, AuthoredModelShape } from "./schema.js";
const ExtendsModel = AuthoredModelShape
.partial()
.extend({
extends: z
.object({
from: z
.string()
.regex(/^[^/]+\/[^/]+$/, "Must be in provider/model format"),
omit: z.array(z.string()).optional(),
})
.strict(),
})
.strict();
export async function generate(directory: string) {
const result = {} as Record<string, Provider>;
const result: Record<string, Provider> = {};
const extendsModels: Array<{
providerID: string;
modelID: string;
modelPath: string;
model: z.infer<typeof ExtendsModel>;
}> = [];
for await (const providerPath of new Bun.Glob("*/provider.toml").scan({
cwd: directory,
absolute: true,
@@ -35,15 +57,149 @@ export async function generate(directory: string) {
},
}).then((mod) => mod.default);
toml.id = modelID;
const model = Model.safeParse(toml);
if (toml.extends !== undefined) {
const model = ExtendsModel.safeParse(toml);
if (!model.success) {
model.error.cause = { modelPath, toml };
throw model.error;
}
extendsModels.push({
providerID,
modelID,
modelPath,
model: model.data,
});
continue;
}
const model = AuthoredModel.safeParse(toml);
if (!model.success) {
model.error.cause = { modelPath, toml };
throw model.error;
}
provider.data.models[modelID] = model.data;
provider.data.models[modelID] = normalizeModelCost(model.data);
}
result[providerID] = provider.data;
}
for (const pendingModel of extendsModels) {
const [providerID, modelID] = pendingModel.model.extends.from.split("/");
const baseModel = result[providerID]?.models[modelID];
if (baseModel === undefined) {
throw new Error(`Unable to resolve extends.from: ${pendingModel.model.extends.from}`, {
cause: { modelPath: pendingModel.modelPath, toml: pendingModel.model },
});
}
const { extends: extendsConfig, ...overrides } = pendingModel.model;
const merged: Record<string, unknown> = structuredClone(
mergeDeep(baseModel, overrides),
);
for (const omit of extendsConfig.omit ?? []) {
const parts = omit.split(".");
const parents: Array<{
value: Record<string, unknown>;
key: string;
}> = [];
let current = merged;
for (const part of parts.slice(0, -1)) {
const next = current[part];
if (
next === undefined ||
next === null ||
typeof next !== "object" ||
Array.isArray(next)
) {
throw new Error(`Unable to omit missing path: ${omit}`, {
cause: { modelPath: pendingModel.modelPath, toml: pendingModel.model },
});
}
parents.push({ value: current, key: part });
current = next as Record<string, unknown>;
}
const lastPart = parts.at(-1);
if (lastPart === undefined || !(lastPart in current)) {
throw new Error(`Unable to omit missing path: ${omit}`, {
cause: { modelPath: pendingModel.modelPath, toml: pendingModel.model },
});
}
delete current[lastPart];
for (let index = parents.length - 1; index >= 0; index--) {
const parent = parents[index];
const value = parent?.value[parent.key];
if (
value === null ||
value === undefined ||
typeof value !== "object" ||
Array.isArray(value) ||
Object.keys(value).length > 0
) {
break;
}
delete parent.value[parent.key];
}
}
const model = Model.safeParse(normalizeCost(merged));
if (!model.success) {
model.error.cause = { modelPath: pendingModel.modelPath, toml: merged };
throw model.error;
}
result[pendingModel.providerID]!.models[pendingModel.modelID] = model.data;
}
return result;
}
function normalizeModelCost(model: z.infer<typeof AuthoredModel>): Model {
return normalizeCost(model) as Model;
}
function normalizeCost(model: Record<string, unknown>) {
const cost = model.cost;
if (cost === undefined || cost === null || typeof cost !== "object" || Array.isArray(cost)) {
return model;
}
const tiers = (cost as { tiers?: unknown }).tiers;
if (!Array.isArray(tiers)) {
return model;
}
if (tiers.length !== 1) {
return model;
}
const contextOver200k = tiers.find((tier) => {
if (tier === null || typeof tier !== "object" || Array.isArray(tier)) return false;
const tierConfig = (tier as { tier?: unknown }).tier;
if (tierConfig === null || typeof tierConfig !== "object" || Array.isArray(tierConfig)) return false;
const type = (tierConfig as { type?: unknown }).type;
const size = (tierConfig as { size?: unknown }).size;
// context_over_200k is a legacy compatibility field. It intentionally
// includes higher thresholds; cost.tiers carries the exact threshold.
return (
(type === undefined || type === "context") &&
typeof size === "number" &&
size >= 200_000
);
});
if (contextOver200k === undefined) {
return model;
}
const { tier: _tier, ...legacyCost } = contextOver200k as Record<string, unknown>;
return {
...model,
cost: {
...(cost as Record<string, unknown>),
context_over_200k: legacyCost,
},
};
}
+175 -79
View File
@@ -2,91 +2,183 @@ import { z } from "zod";
import { ModelFamily } from "./family";
const Cost = z.object({
input: z.number().min(0, "Input price cannot be negative"),
output: z.number().min(0, "Output price cannot be negative"),
reasoning: z.number().min(0, "Input price cannot be negative").optional(),
cache_read: z
.number()
.min(0, "Cache read price cannot be negative")
type JsonValue =
| string
| number
| boolean
| null
| { [key: string]: JsonValue }
| JsonValue[];
const JsonValue: z.ZodType<JsonValue> = z.lazy(() =>
z.union([
z.string(),
z.number(),
z.boolean(),
z.null(),
z.array(JsonValue),
z.record(JsonValue),
]),
);
const Cost = z
.object({
input: z.number().min(0, "Input price cannot be negative"),
output: z.number().min(0, "Output price cannot be negative"),
reasoning: z
.number()
.min(0, "Reasoning price cannot be negative")
.optional(),
cache_read: z
.number()
.min(0, "Cache read price cannot be negative")
.optional(),
cache_write: z
.number()
.min(0, "Cache write price cannot be negative")
.optional(),
input_audio: z
.number()
.min(0, "Audio input price cannot be negative")
.optional(),
output_audio: z
.number()
.min(0, "Audio output price cannot be negative")
.optional(),
});
const CostTier = Cost.extend({
tier: z
.object({
type: z.literal("context").default("context"),
size: z.number().int().min(0, "Context tier size cannot be negative"),
})
.strict(),
}).strict();
const AuthoredCost = Cost.extend({
context_over_200k: z.never().optional(),
tiers: z.array(CostTier).optional(),
});
const OutputCost = Cost.extend({
context_over_200k: Cost.optional(),
tiers: z.array(CostTier).optional(),
});
const ModelBase = z.object({
id: z.string(),
name: z.string().min(1, "Model name cannot be empty"),
family: ModelFamily.optional(),
attachment: z.boolean(),
reasoning: z.boolean(),
tool_call: z.boolean(),
interleaved: z
.union([
z.literal(true),
z
.object({
field: z.enum(["reasoning_content", "reasoning_details"]),
})
.strict(),
])
.optional(),
cache_write: z
.number()
.min(0, "Cache write price cannot be negative")
structured_output: z.boolean().optional(),
temperature: z.boolean().optional(),
knowledge: z
.string()
.regex(/^\d{4}-\d{2}(-\d{2})?$/, {
message: "Must be in YYYY-MM or YYYY-MM-DD format",
})
.optional(),
input_audio: z
.number()
.min(0, "Audio input price cannot be negative")
release_date: z.string().regex(/^\d{4}-\d{2}(-\d{2})?$/, {
message: "Must be in YYYY-MM or YYYY-MM-DD format",
}),
last_updated: z.string().regex(/^\d{4}-\d{2}(-\d{2})?$/, {
message: "Must be in YYYY-MM or YYYY-MM-DD format",
}),
modalities: z.object({
input: z.array(z.enum(["text", "audio", "image", "video", "pdf"])),
output: z.array(z.enum(["text", "audio", "image", "video", "pdf"])),
}),
open_weights: z.boolean(),
limit: z.object({
context: z.number().min(0, "Context window must be positive"),
input: z.number().min(0, "Input tokens must be positive").optional(),
output: z.number().min(0, "Output tokens must be positive"),
}),
status: z.enum(["alpha", "beta", "deprecated"]).optional(),
experimental: z
.object({
modes: z
.record(
z.object({
cost: Cost.optional(),
provider: z
.object({
body: z.record(JsonValue).optional(),
headers: z.record(z.string()).optional(),
})
.optional(),
}),
)
.optional(),
})
.optional(),
output_audio: z
.number()
.min(0, "Audio output price cannot be negative")
provider: z
.object({
npm: z.string().optional(),
api: z.string().optional(),
shape: z.enum(["responses", "completions"]).optional(),
body: z.record(JsonValue).optional(),
headers: z.record(z.string()).optional(),
})
.optional(),
});
export const Model = z
function refineModel<T extends z.ZodTypeAny>(schema: T) {
return schema
.refine(
(data) => {
return !(data.reasoning === false && data.cost?.reasoning !== undefined);
},
{
message: "Cannot set cost.reasoning when reasoning is false",
path: ["cost", "reasoning"],
},
)
.refine(
(data) => {
const tiers = data.cost?.tiers;
if (tiers === undefined) return true;
const sizes = tiers.map((tier: { tier: { size: number } }) => tier.tier.size);
return new Set(sizes).size === sizes.length;
},
{
message: "Cost context tiers must not have duplicate sizes",
path: ["cost", "tiers"],
},
);
}
export const ModelShape = z
.object({
id: z.string(),
name: z.string().min(1, "Model name cannot be empty"),
family: ModelFamily.optional(),
attachment: z.boolean(),
reasoning: z.boolean(),
tool_call: z.boolean(),
interleaved: z
.union([
z.literal(true),
z
.object({
field: z.enum(["reasoning_content", "reasoning_details"]),
})
.strict(),
])
.optional(),
structured_output: z.boolean().optional(),
temperature: z.boolean().optional(),
knowledge: z
.string()
.regex(/^\d{4}-\d{2}(-\d{2})?$/, {
message: "Must be in YYYY-MM or YYYY-MM-DD format",
})
.optional(),
release_date: z.string().regex(/^\d{4}-\d{2}(-\d{2})?$/, {
message: "Must be in YYYY-MM or YYYY-MM-DD format",
}),
last_updated: z.string().regex(/^\d{4}-\d{2}(-\d{2})?$/, {
message: "Must be in YYYY-MM or YYYY-MM-DD format",
}),
modalities: z.object({
input: z.array(z.enum(["text", "audio", "image", "video", "pdf"])),
output: z.array(z.enum(["text", "audio", "image", "video", "pdf"])),
}),
open_weights: z.boolean(),
cost: Cost.extend({
context_over_200k: Cost.optional(),
}).optional(),
limit: z.object({
context: z.number().min(0, "Context window must be positive"),
input: z.number().min(0, "Input tokens must be positive").optional(),
output: z.number().min(0, "Output tokens must be positive"),
}),
status: z.enum(["alpha", "beta", "deprecated"]).optional(),
provider: z
.object({
npm: z.string().optional(),
api: z.string().optional(),
shape: z.enum(["responses", "completions"]).optional(),
})
.optional(),
...ModelBase.shape,
cost: OutputCost.optional(),
})
.strict()
.refine(
(data) => {
return !(data.reasoning === false && data.cost?.reasoning !== undefined);
},
{
message: "Cannot set cost.reasoning when reasoning is false",
path: ["cost", "reasoning"],
},
);
.strict();
export const AuthoredModelShape = z
.object({
...ModelBase.shape,
cost: AuthoredCost.optional(),
})
.strict();
export const Model = refineModel(ModelShape);
export const AuthoredModel = refineModel(AuthoredModelShape);
export type Model = z.infer<typeof Model>;
@@ -112,6 +204,7 @@ export const Provider = z
const isOpenAIcompatible = data.npm === "@ai-sdk/openai-compatible";
const isOpenrouter = data.npm === "@openrouter/ai-sdk-provider";
const isAnthropic = data.npm === "@ai-sdk/anthropic";
const isKiro = data.npm === "kiro-acp-ai-provider";
const hasApi = data.api !== undefined;
return (
@@ -123,17 +216,20 @@ export const Provider = z
isAnthropic ||
// openai: api optional (always allowed)
isOpenAI ||
// kiro: api optional (always allowed)
isKiro ||
// all others: must NOT have api
(!isOpenAI &&
!isOpenAIcompatible &&
!isOpenrouter &&
!isAnthropic &&
!isKiro &&
!hasApi)
);
},
{
message:
"'api' is required for openai-compatible and openrouter, optional for anthropic and openai, forbidden otherwise",
"'api' is required for openai-compatible and openrouter, optional for anthropic, openai, and kiro, forbidden otherwise",
path: ["api"],
},
);
+481
View File
@@ -0,0 +1,481 @@
import path from "node:path";
import { mkdir, readdir, rm } from "node:fs/promises";
import { z } from "zod";
import { AuthoredModel, AuthoredModelShape } from "../schema.js";
import { cloudflareWorkersAi } from "./providers/cloudflare-workers-ai.js";
import { google } from "./providers/google.js";
import { openrouter } from "./providers/openrouter.js";
import { xai } from "./providers/xai.js";
const ExtendsConfig = z
.object({
from: z.string(),
omit: z.array(z.string()).optional(),
})
.strict();
const ExistingExtendsConfig = z
.object({
from: z.string(),
omit: z.array(z.string()).optional(),
})
.passthrough();
const ExistingModel = AuthoredModelShape.partial()
.extend({
extends: ExistingExtendsConfig.optional(),
})
.strict();
const SyncedExtendsModel = AuthoredModelShape.partial()
.extend({
id: z.string(),
extends: ExtendsConfig,
})
.strict();
const SyncedAuthoredModel = z.union([AuthoredModel, SyncedExtendsModel]);
export type ExistingModel = z.infer<typeof ExistingModel>;
export type SyncedFullModel = Omit<z.infer<typeof AuthoredModelShape>, "id">;
export type SyncedExtendsModel = Omit<z.infer<typeof SyncedExtendsModel>, "id">;
export type SyncedModel = SyncedFullModel | SyncedExtendsModel;
export interface SyncProvider<SourceModel> {
id: string;
name: string;
modelsDir: string;
skipCreates?: boolean;
sourceID?(model: SourceModel): string;
skippedNotice?(ids: string[]): string[];
fetchModels(): Promise<unknown>;
parseModels(raw: unknown): SourceModel[];
translateModel(
model: SourceModel,
context: { existing(id: string): ExistingModel | undefined },
): { id: string; model: SyncedModel } | undefined;
}
export interface SyncResult {
id: string;
name: string;
status: "changed" | "unchanged";
created: number;
updated: number;
deleted: number;
unchanged: number;
notices: string[];
files: Array<{ status: "created" | "updated" | "deleted"; path: string }>;
}
export const providers: {
"cloudflare-workers-ai": SyncProvider<any>;
google: SyncProvider<any>;
openrouter: SyncProvider<any>;
xai: SyncProvider<any>;
} = {
"cloudflare-workers-ai": cloudflareWorkersAi,
google,
openrouter,
xai,
};
export const groups = {
aggregators: ["openrouter"],
cloudflare: ["cloudflare-workers-ai"],
direct: ["google", "xai"],
} as const;
type ProviderID = keyof typeof providers;
interface SyncOptions {
dryRun?: boolean;
newOnly?: boolean;
}
export async function syncProviderByID(id: ProviderID, options: SyncOptions = {}) {
return syncProvider(providers[id], options);
}
export async function syncProvider<SourceModel>(
provider: SyncProvider<SourceModel>,
options: SyncOptions = {},
): Promise<SyncResult> {
console.log(`\nSyncing ${provider.name}...`);
const existing = await readExisting(provider.modelsDir);
const sourceModels = provider.parseModels(await provider.fetchModels());
const desired = new Map<string, { model: z.infer<typeof SyncedAuthoredModel>; content: string }>();
const skippedRemote: string[] = [];
for (const sourceModel of sourceModels) {
const translated = provider.translateModel(sourceModel, {
existing(id) {
return existing.get(`${id}.toml`)?.toml;
},
});
if (translated === undefined) {
if (provider.skipCreates) skippedRemote.push(provider.sourceID?.(sourceModel) ?? "unknown");
continue;
}
const relativePath = `${translated.id}.toml`;
if (provider.skipCreates && !existing.has(relativePath)) {
skippedRemote.push(translated.id);
continue;
}
if (desired.has(relativePath)) {
throw new Error(`Duplicate synced model path: ${provider.id}/${relativePath}`);
}
const parsed = SyncedAuthoredModel.safeParse({
id: translated.id,
...translated.model,
});
if (!parsed.success) {
parsed.error.cause = { provider: provider.id, path: relativePath };
throw parsed.error;
}
desired.set(relativePath, {
model: parsed.data,
content: formatToml(parsed.data),
});
}
const files: SyncResult["files"] = [];
let unchanged = 0;
for (const [relativePath, file] of desired) {
const filePath = path.join(provider.modelsDir, relativePath);
const current = existing.get(relativePath);
if (current === undefined) {
files.push({ status: "created", path: filePath });
if (options.dryRun) {
console.log(`Would create ${relativePath}`);
} else {
await mkdir(path.dirname(filePath), { recursive: true });
await Bun.write(filePath, file.content);
}
continue;
}
if (!sameModel(relativePath, current.toml, file.model)) {
if (options.newOnly) {
unchanged++;
continue;
}
files.push({ status: "updated", path: filePath });
if (options.dryRun) {
console.log(`Would update ${relativePath}`);
} else {
if (current.symlink) await rm(filePath, { force: true });
await Bun.write(filePath, file.content);
}
} else {
unchanged++;
}
}
for (const relativePath of existing.keys()) {
if (desired.has(relativePath)) continue;
if (options.newOnly) {
console.log(`Skipping removal in new-only mode: ${relativePath}`);
unchanged++;
continue;
}
const filePath = path.join(provider.modelsDir, relativePath);
files.push({ status: "deleted", path: filePath });
if (options.dryRun) {
console.log(`Would remove ${relativePath}`);
} else {
await rm(filePath, { force: true });
}
}
const result = summarize(provider, files, unchanged, provider.skippedNotice?.(skippedRemote) ?? []);
console.log(
`${options.dryRun ? "Dry run: " : ""}${result.created} created, ${result.updated} updated, ${result.deleted} removed, ${result.unchanged} unchanged`,
);
return result;
}
export async function syncTargets(target: string, options: SyncOptions = {}) {
const ids = target in groups
? groups[target as keyof typeof groups]
: target in providers
? [target as ProviderID]
: undefined;
if (ids === undefined) {
throw new Error(`Unknown sync target: ${target}`);
}
const results: SyncResult[] = [];
for (const id of ids) {
results.push(await syncProviderByID(id as ProviderID, options));
}
return results;
}
export function syncProviderMatrix() {
return {
include: Object.values(providers).map((provider) => ({
provider: provider.id,
name: provider.name,
})),
};
}
async function readExisting(modelsDir: string) {
const existing = new Map<string, { text: string; toml: ExistingModel; symlink: boolean }>();
for (const { file, symlink } of await tomlFiles(modelsDir)) {
const text = await Bun.file(path.join(modelsDir, file)).text();
const parsed = ExistingModel.safeParse(Bun.TOML.parse(text));
if (!parsed.success) {
parsed.error.cause = { path: path.join(modelsDir, file) };
throw parsed.error;
}
existing.set(file, { text, toml: parsed.data, symlink });
}
return existing;
}
async function tomlFiles(root: string, dir = "") {
const result: Array<{ file: string; symlink: boolean }> = [];
for (const entry of await readdir(path.join(root, dir), { withFileTypes: true })) {
const file = path.join(dir, entry.name);
if (entry.isDirectory()) {
result.push(...await tomlFiles(root, file));
} else if (entry.name.endsWith(".toml") && (entry.isFile() || entry.isSymbolicLink())) {
result.push({ file, symlink: entry.isSymbolicLink() });
}
}
return result;
}
function summarize(
provider: { id: string; name: string },
files: SyncResult["files"],
unchanged: number,
notices: string[],
): SyncResult {
return {
id: provider.id,
name: provider.name,
status: files.length > 0 ? "changed" : "unchanged",
created: files.filter((file) => file.status === "created").length,
updated: files.filter((file) => file.status === "updated").length,
deleted: files.filter((file) => file.status === "deleted").length,
unchanged,
notices,
files,
};
}
function sameModel(
relativePath: string,
current: ExistingModel,
desired: z.infer<typeof SyncedAuthoredModel>,
) {
const parsed = SyncedAuthoredModel.safeParse({
id: relativePath.slice(0, -5),
...current,
});
return parsed.success && stable(parsed.data) === stable(desired);
}
function stable(value: unknown): string {
if (Array.isArray(value)) {
const items = value.map(stable);
const ordered = value.every((item) => item === null || typeof item !== "object")
? items.sort()
: items;
return `[${ordered.join(",")}]`;
}
if (value !== null && typeof value === "object") {
return `{${Object.entries(value)
.filter(([, item]) => item !== undefined)
.sort(([a], [b]) => a.localeCompare(b))
.map(([key, item]) => `${JSON.stringify(key)}:${stable(item)}`)
.join(",")}}`;
}
return JSON.stringify(value);
}
async function writeReport(target: string, results: SyncResult[]) {
await mkdir(".sync", { recursive: true });
const lines = [
`Updates model TOMLs for the \`${target}\` sync target.`,
"",
"| Provider | Status | Created | Updated | Deleted |",
"| --- | --- | ---: | ---: | ---: |",
];
for (const result of results) {
lines.push(
`| ${result.name} | ${result.status} | ${result.created} | ${result.updated} | ${result.deleted} |`,
);
}
for (const result of results.filter((item) => item.files.length > 0)) {
lines.push("", `<details><summary>${result.name} changed files</summary>`, "");
for (const file of result.files) {
lines.push(`- ${file.status}: \`${file.path}\``);
}
lines.push("", "</details>");
}
const noticeResults = results.filter((item) => item.notices.length > 0);
if (noticeResults.length > 0) {
lines.push("", "## Notices");
for (const result of noticeResults) {
lines.push("", `### ${result.name}`);
for (const notice of result.notices) {
lines.push(`- ${notice}`);
}
}
}
lines.push("", "This PR was created automatically by the daily model sync workflow.");
await Bun.write(".sync/model-sync-report.md", `${lines.join("\n")}\n`);
}
function quote(value: string) {
return `"${value.replaceAll("\\", "\\\\").replaceAll('"', '\\"')}"`;
}
function formatInteger(n: number) {
return String(n).replace(/\B(?=(\d{3})+(?!\d))/g, "_");
}
function formatNumber(n: number) {
return Number.isInteger(n) ? formatInteger(n) : String(n);
}
function formatToml(model: z.infer<typeof SyncedAuthoredModel>) {
const lines: string[] = [];
const extendsLines: string[] = [];
if ("extends" in model) {
extendsLines.push("[extends]");
extendsLines.push(`from = ${quote(model.extends.from)}`);
if (model.extends.omit !== undefined) {
extendsLines.push(`omit = [${model.extends.omit.map(quote).join(", ")}]`);
}
}
if (model.name !== undefined) lines.push(`name = ${quote(model.name)}`);
if (model.family !== undefined) lines.push(`family = ${quote(model.family)}`);
if (model.release_date !== undefined) lines.push(`release_date = ${quote(model.release_date)}`);
if (model.last_updated !== undefined) lines.push(`last_updated = ${quote(model.last_updated)}`);
if (model.attachment !== undefined) lines.push(`attachment = ${model.attachment}`);
if (model.reasoning !== undefined) lines.push(`reasoning = ${model.reasoning}`);
if (model.temperature !== undefined) lines.push(`temperature = ${model.temperature}`);
if (model.tool_call !== undefined) lines.push(`tool_call = ${model.tool_call}`);
if (model.structured_output !== undefined) {
lines.push(`structured_output = ${model.structured_output}`);
}
if (model.knowledge !== undefined) lines.push(`knowledge = ${quote(model.knowledge)}`);
if (model.open_weights !== undefined) lines.push(`open_weights = ${model.open_weights}`);
if (model.status !== undefined) lines.push(`status = ${quote(model.status)}`);
if (extendsLines.length > 0) {
if (lines.length > 0) lines.push("");
lines.push(...extendsLines);
}
if (model.interleaved !== undefined) {
lines.push("");
if (model.interleaved === true) {
lines.push("interleaved = true");
} else {
lines.push("[interleaved]");
lines.push(`field = ${quote(model.interleaved.field)}`);
}
}
if (model.cost !== undefined) {
lines.push("", "[cost]");
lines.push(`input = ${formatNumber(model.cost.input)}`);
lines.push(`output = ${formatNumber(model.cost.output)}`);
if (model.cost.reasoning !== undefined) {
lines.push(`reasoning = ${formatNumber(model.cost.reasoning)}`);
}
if (model.cost.cache_read !== undefined) {
lines.push(`cache_read = ${formatNumber(model.cost.cache_read)}`);
}
if (model.cost.cache_write !== undefined) {
lines.push(`cache_write = ${formatNumber(model.cost.cache_write)}`);
}
if (model.cost.input_audio !== undefined) {
lines.push(`input_audio = ${formatNumber(model.cost.input_audio)}`);
}
if (model.cost.output_audio !== undefined) {
lines.push(`output_audio = ${formatNumber(model.cost.output_audio)}`);
}
for (const tier of model.cost.tiers ?? []) {
lines.push("", "[[cost.tiers]]");
lines.push(`tier = { size = ${formatInteger(tier.tier.size)} }`);
lines.push(`input = ${formatNumber(tier.input)}`);
lines.push(`output = ${formatNumber(tier.output)}`);
if (tier.reasoning !== undefined) lines.push(`reasoning = ${formatNumber(tier.reasoning)}`);
if (tier.cache_read !== undefined) lines.push(`cache_read = ${formatNumber(tier.cache_read)}`);
if (tier.cache_write !== undefined) lines.push(`cache_write = ${formatNumber(tier.cache_write)}`);
}
}
if (model.limit !== undefined) {
lines.push("", "[limit]");
if (model.limit.context !== undefined) lines.push(`context = ${formatInteger(model.limit.context)}`);
if (model.limit.input !== undefined) lines.push(`input = ${formatInteger(model.limit.input)}`);
if (model.limit.output !== undefined) lines.push(`output = ${formatInteger(model.limit.output)}`);
}
if (model.modalities !== undefined) {
lines.push("", "[modalities]");
if (model.modalities.input !== undefined) {
lines.push(`input = [${model.modalities.input.map(quote).join(", ")}]`);
}
if (model.modalities.output !== undefined) {
lines.push(`output = [${model.modalities.output.map(quote).join(", ")}]`);
}
}
return `${lines.join("\n")}\n`;
}
export async function main(args = process.argv.slice(2)) {
if (args.includes("--list-providers")) {
console.log(JSON.stringify(syncProviderMatrix()));
return;
}
const target = args.find((arg) => !arg.startsWith("-")) ?? "aggregators";
const results = await syncTargets(target, {
dryRun: args.includes("--dry-run"),
newOnly: args.includes("--new-only"),
});
await writeReport(target, results);
console.log("\nSync summary");
for (const result of results) {
console.log(
`${result.name}: ${result.created} created, ${result.updated} updated, ${result.deleted} deleted`,
);
}
}
if (import.meta.main) await main();
@@ -0,0 +1,176 @@
import { z } from "zod";
import type { ExistingModel, SyncProvider } from "../index.js";
import {
buildOpenRouterModel,
OpenRouterModel,
OpenRouterResponse,
} from "./openrouter.js";
const API_BASE = "https://api.cloudflare.com/client/v4/accounts";
const CloudflareOpenRouterResponse = z.object({
result: z.union([OpenRouterResponse, z.array(OpenRouterModel)]).optional(),
result_info: z.object({
page: z.number().optional(),
total_pages: z.number().optional(),
}).passthrough().optional(),
}).passthrough();
const CloudflareModel = z.object({
id: z.string(),
name: z.string(),
created: z.number(),
hugging_face_id: z.string().nullable().optional(),
context_length: z.number(),
max_output_length: z.number().nullable().optional(),
input_modalities: z.array(z.string()).optional(),
output_modalities: z.array(z.string()).optional(),
pricing: z.object({
prompt: z.string(),
completion: z.string(),
internal_reasoning: z.string().optional(),
input_cache_read: z.string().optional(),
input_cache_write: z.string().optional(),
}),
supported_features: z.array(z.string()).optional(),
supported_sampling_parameters: z.array(z.string()).optional(),
}).passthrough();
const CloudflareResponse = z.object({
data: z.array(CloudflareModel),
}).passthrough();
type CloudflareModel = z.infer<typeof CloudflareModel>;
export const cloudflareWorkersAi = {
id: "cloudflare-workers-ai",
name: "Cloudflare Workers AI",
modelsDir: "providers/cloudflare-workers-ai/models",
async fetchModels() {
const accountID = process.env.CLOUDFLARE_WORKERS_AI_SYNC_ACCOUNT_ID;
const token = process.env.CLOUDFLARE_WORKERS_AI_SYNC_API_TOKEN;
if (accountID === undefined || token === undefined) {
throw new Error(
"Cloudflare Workers AI sync requires CLOUDFLARE_WORKERS_AI_SYNC_ACCOUNT_ID and CLOUDFLARE_WORKERS_AI_SYNC_API_TOKEN",
);
}
const first = await fetchPage(accountID, token, 1);
const models = parseCloudflareModels(first);
const pageInfo = CloudflareOpenRouterResponse.safeParse(first).success
? CloudflareOpenRouterResponse.parse(first).result_info
: undefined;
for (let page = 2; page <= (pageInfo?.total_pages ?? 1); page++) {
models.push(...parseCloudflareModels(await fetchPage(accountID, token, page)));
}
return { data: models };
},
parseModels(raw) {
return parseCloudflareModels(raw);
},
translateModel(model, context) {
const normalized = normalizeModel(model);
const id = normalized.id.replace(/^workers-ai\//, "");
return {
id,
model: buildWorkersAiModel(normalized, context.existing(id)),
};
},
} satisfies SyncProvider<CloudflareModel>;
function buildWorkersAiModel(model: z.infer<typeof OpenRouterModel>, existing: ExistingModel | undefined) {
const synced = buildOpenRouterModel(model, existing);
return {
...synced,
name: existing?.name ?? synced.name,
release_date: existing?.release_date ?? synced.release_date,
last_updated: existing?.last_updated ?? synced.last_updated,
limit: {
...synced.limit,
output: existing?.limit?.output ?? synced.limit.output,
},
};
}
async function fetchPage(accountID: string, token: string, page: number) {
const url = new URL(`${API_BASE}/${accountID}/ai/models/search`);
url.searchParams.set("format", "openrouter");
url.searchParams.set("per_page", "1000");
url.searchParams.set("page", String(page));
const response = await fetch(url, {
headers: { Authorization: `Bearer ${token}` },
});
if (!response.ok) {
throw new Error(
`Cloudflare Workers AI models request failed: ${response.status} ${response.statusText}${await responseDetails(response)}`,
);
}
return response.json();
}
function parseCloudflareModels(raw: unknown) {
const cloudflare = CloudflareResponse.safeParse(raw);
if (cloudflare.success) return cloudflare.data.data;
const direct = OpenRouterResponse.safeParse(raw);
if (direct.success) return direct.data.data;
const wrapped = CloudflareOpenRouterResponse.parse(raw);
if (wrapped.result === undefined) {
throw new Error("Cloudflare Workers AI response did not include model data");
}
return Array.isArray(wrapped.result) ? wrapped.result : wrapped.result.data;
}
function normalizeModel(model: CloudflareModel) {
if ("architecture" in model && "top_provider" in model && "supported_parameters" in model) {
return OpenRouterModel.parse(model);
}
return OpenRouterModel.parse({
id: model.id.startsWith("@cf/") ? model.id : `@cf/${model.id.replace(/^@cf\//, "")}`,
name: model.name,
created: model.created,
hugging_face_id: model.hugging_face_id ?? null,
knowledge_cutoff: null,
context_length: model.context_length,
architecture: {
input_modalities: model.input_modalities ?? ["text"],
output_modalities: model.output_modalities ?? ["text"],
},
pricing: model.pricing,
top_provider: {
context_length: model.context_length,
max_completion_tokens: model.max_output_length ?? null,
},
supported_parameters: [
...model.supported_sampling_parameters ?? [],
...model.supported_features ?? [],
],
});
}
async function responseDetails(response: Response) {
const text = await response.text();
if (text.length === 0) return "";
try {
const body = z.object({
errors: z.array(z.object({
code: z.union([z.string(), z.number()]).optional(),
message: z.string().optional(),
}).passthrough()).optional(),
}).passthrough().parse(JSON.parse(text));
const details = body.errors
?.map((error) => [error.code, error.message].filter(Boolean).join(": "))
.filter((message) => message.length > 0)
.join("; ");
return details === undefined || details.length === 0 ? "" : ` (${details})`;
} catch {
return "";
}
}
+138
View File
@@ -0,0 +1,138 @@
import { z } from "zod";
import type { ExistingModel, SyncProvider, SyncedModel } from "../index.js";
const API_ENDPOINT = "https://generativelanguage.googleapis.com/v1beta/models";
const GoogleModel = z.object({
name: z.string(),
baseModelId: z.string().optional(),
version: z.string().optional(),
displayName: z.string().optional(),
description: z.string().optional(),
inputTokenLimit: z.number().int().nonnegative(),
outputTokenLimit: z.number().int().nonnegative(),
supportedGenerationMethods: z.array(z.string()).optional(),
temperature: z.number().optional(),
topP: z.number().optional(),
topK: z.number().optional(),
maxTemperature: z.number().optional(),
thinking: z.boolean().optional(),
}).passthrough();
const GoogleResponse = z.object({
models: z.array(GoogleModel).optional(),
nextPageToken: z.string().optional(),
}).passthrough();
type GoogleModel = z.infer<typeof GoogleModel>;
export const google = {
id: "google",
name: "Google",
modelsDir: "providers/google/models",
skipCreates: true,
sourceID(model) {
return model.name.replace(/^models\//, "");
},
skippedNotice(ids) {
if (ids.length === 0) return [];
return [
`${ids.length} Google models returned by the API were not created because the Models API does not provide authoritative modalities, pricing, knowledge cutoff, release date, tool calling, or structured output metadata. Existing models are still updated from API-authoritative fields.`,
`Skipped remote IDs: ${ids.map((id) => `\`${id}\``).join(", ")}`,
];
},
async fetchModels() {
const key = process.env.GOOGLE_API_KEY
?? process.env.GEMINI_API_KEY
?? process.env.GOOGLE_GENERATIVE_AI_API_KEY;
if (key === undefined) {
throw new Error("Google sync requires GOOGLE_API_KEY, GEMINI_API_KEY, or GOOGLE_GENERATIVE_AI_API_KEY");
}
const models: GoogleModel[] = [];
let pageToken: string | undefined;
do {
const url = new URL(API_ENDPOINT);
url.searchParams.set("key", key);
url.searchParams.set("pageSize", "1000");
if (pageToken !== undefined) url.searchParams.set("pageToken", pageToken);
const response = await fetch(url);
if (!response.ok) {
throw new Error(`Google models request failed: ${response.status} ${response.statusText}`);
}
const page = GoogleResponse.parse(await response.json());
models.push(...page.models ?? []);
pageToken = page.nextPageToken;
} while (pageToken !== undefined);
return { models };
},
parseModels(raw) {
return GoogleResponse.parse(raw).models ?? [];
},
translateModel(model, context) {
const id = model.name.replace(/^models\//, "");
const existing = context.existing(id);
if (existing === undefined) return undefined;
return {
id,
model: buildModel(model, existing),
};
},
} satisfies SyncProvider<GoogleModel>;
function buildModel(model: GoogleModel, existing: ExistingModel): SyncedModel {
const name = existing.name;
const releaseDate = existing.release_date;
const lastUpdated = existing.last_updated;
const attachment = existing.attachment;
const reasoning = existing.reasoning;
const toolCall = existing.tool_call;
const openWeights = existing.open_weights;
const limit = existing.limit;
const modalities = existing.modalities;
if (
name === undefined
|| releaseDate === undefined
|| lastUpdated === undefined
|| attachment === undefined
|| reasoning === undefined
|| toolCall === undefined
|| openWeights === undefined
|| limit === undefined
|| modalities === undefined
) {
throw new Error(`Google model ${model.name} has incomplete local TOML metadata required for sync`);
}
return {
name: model.displayName ?? name,
family: existing.family,
release_date: releaseDate,
last_updated: lastUpdated,
attachment,
reasoning: model.thinking ?? reasoning,
temperature: model.temperature !== undefined || model.maxTemperature !== undefined
? true
: existing.temperature,
tool_call: toolCall,
structured_output: existing.structured_output,
knowledge: existing.knowledge,
open_weights: openWeights,
status: existing.status,
interleaved: existing.interleaved,
cost: existing.cost,
limit: {
input: limit.input,
context: model.inputTokenLimit,
output: model.outputTokenLimit,
},
modalities,
};
}
@@ -0,0 +1,334 @@
import { z } from "zod";
import { readFileSync, readdirSync } from "node:fs";
import path from "node:path";
import { ModelFamilyValues } from "../../family.js";
import type { ExistingModel, SyncProvider, SyncedFullModel, SyncedModel } from "../index.js";
const API_ENDPOINT = "https://openrouter.ai/api/v1/models";
const PROVIDERS_DIR = path.join(import.meta.dirname, "..", "..", "..", "..", "..", "providers");
const modelFilesByProvider = new Map<string, Set<string>>();
const canonicalTomlByModel = new Map<string, Record<string, unknown>>();
const CANONICAL_PROVIDER_PREFIXES = {
anthropic: "anthropic",
cohere: "cohere",
deepseek: "deepseek",
google: "google",
meta: "llama",
"meta-llama": "llama",
minimax: "minimax",
mistralai: "mistral",
moonshotai: "moonshotai",
openai: "openai",
"x-ai": "xai",
xai: "xai",
xiaomi: "xiaomi",
zai: "zai",
"z-ai": "zai",
} as const;
export const OpenRouterModel = z.object({
id: z.string(),
name: z.string(),
created: z.number(),
hugging_face_id: z.string().nullable(),
knowledge_cutoff: z.string().nullable(),
context_length: z.number(),
architecture: z.object({
input_modalities: z.array(z.string()),
output_modalities: z.array(z.string()),
}),
pricing: z.object({
prompt: z.string(),
completion: z.string(),
internal_reasoning: z.string().optional(),
input_cache_read: z.string().optional(),
input_cache_write: z.string().optional(),
}),
top_provider: z.object({
context_length: z.number().nullable(),
max_completion_tokens: z.number().nullable(),
}),
supported_parameters: z.array(z.string()),
});
export const OpenRouterResponse = z.object({
data: z.array(OpenRouterModel),
}).passthrough();
export type OpenRouterModel = z.infer<typeof OpenRouterModel>;
export const openrouter = {
id: "openrouter",
name: "OpenRouter",
modelsDir: "providers/openrouter/models",
async fetchModels() {
const headers = process.env.OPENROUTER_API_KEY
? { Authorization: `Bearer ${process.env.OPENROUTER_API_KEY}` }
: undefined;
const response = await fetch(API_ENDPOINT, { headers });
if (!response.ok) {
throw new Error(`OpenRouter request failed: ${response.status} ${response.statusText}`);
}
return response.json();
},
parseModels(raw) {
return OpenRouterResponse.parse(raw).data;
},
translateModel(model, context) {
return {
id: model.id,
model: buildOpenRouterModel(model, context.existing(model.id)),
};
},
} satisfies SyncProvider<OpenRouterModel>;
function dateFromTimestamp(timestamp: number) {
return new Date(timestamp * 1000).toISOString().slice(0, 10);
}
function price(value: string | undefined) {
if (value === undefined) return undefined;
const number = Number(value);
return Number.isFinite(number) && number >= 0
? Math.round(number * 1_000_000_000_000) / 1_000_000
: undefined;
}
type Modality = "text" | "audio" | "image" | "video" | "pdf";
function modalities(values: string[], fallback: Modality[]): Modality[] {
const allowed = new Set<Modality>(["text", "audio", "image", "video", "pdf"]);
const result = values
.map((value) => value.toLowerCase())
.map((value) => value === "file" ? "pdf" : value)
.filter((value): value is Modality => allowed.has(value as Modality));
return [...new Set(result.length > 0 ? result : fallback)];
}
function inferFamily(model: OpenRouterModel, name: string) {
const target = `${model.id} ${name}`.toLowerCase();
return [...ModelFamilyValues]
.sort((a, b) => b.length - a.length)
.find((family) => {
const value = family.toLowerCase().replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
if (family === "o") {
return new RegExp(`(^|[^a-z0-9])${value}(?=\\d|$|[^a-z0-9])`).test(target);
}
return new RegExp(`(^|[^a-z0-9])${value}(?=$|[^a-z0-9])`).test(target);
});
}
export function buildOpenRouterModel(model: OpenRouterModel, existing: ExistingModel | undefined): SyncedModel {
const params = new Set(model.supported_parameters);
const name = model.name.replace(/^[^:]+:\s+/, "");
const input = modalities(model.architecture.input_modalities, ["text"]);
const output = modalities(model.architecture.output_modalities, ["text"]);
const prompt = price(model.pricing.prompt);
const completion = price(model.pricing.completion);
const reasoning = params.has("reasoning") || params.has("include_reasoning");
const context = model.top_provider.context_length ?? model.context_length;
const family = inferFamily(model, name);
const releaseDate = dateFromTimestamp(model.created);
const familyValue = existing?.family === "o" && family !== "o"
? family
: (existing?.family ?? family);
const attachment = input.some((value) => value !== "text");
const toolCall = params.has("tools") || params.has("tool_choice");
const structuredOutput = params.has("structured_outputs");
const knowledge = model.knowledge_cutoff?.slice(0, 10) ?? existing?.knowledge;
const openWeights = Boolean(model.hugging_face_id);
const cost = prompt !== undefined && completion !== undefined
? {
input: prompt,
output: completion,
reasoning: reasoning ? price(model.pricing.internal_reasoning) : undefined,
cache_read: price(model.pricing.input_cache_read),
cache_write: price(model.pricing.input_cache_write),
tiers: existing?.cost?.tiers,
}
: existing?.cost;
const limit = {
context,
input: existing?.limit?.input,
output: model.top_provider.max_completion_tokens ?? existing?.limit?.output ?? context,
};
const canonical = resolveCanonicalModel(model.id);
if (canonical !== undefined) {
return {
extends: {
from: canonical.from,
omit: canonicalOmit(canonical.provider, canonical.modelID, cost, limit),
},
...canonicalRuntimeOverrides(canonical.provider, canonical.modelID, {
name: model.id.endsWith(":free") ? name : undefined,
attachment,
reasoning,
}),
temperature: params.has("temperature"),
tool_call: toolCall,
structured_output: structuredOutput,
status: existing?.status,
interleaved: existing?.interleaved,
cost,
limit,
modalities: { input, output },
};
}
return {
name,
family: familyValue,
release_date: releaseDate,
last_updated: releaseDate,
attachment,
reasoning,
temperature: params.has("temperature"),
tool_call: toolCall,
structured_output: structuredOutput,
knowledge,
open_weights: openWeights,
status: existing?.status,
interleaved: existing?.interleaved,
cost,
limit,
modalities: { input, output },
} satisfies SyncedFullModel;
}
function resolveCanonicalModel(openrouterID: string) {
const [prefix, ...modelParts] = openrouterID.split("/");
if (prefix === undefined || modelParts.length === 0) return undefined;
if (openrouterID.startsWith("~/") || prefix.startsWith("~")) return undefined;
const provider = CANONICAL_PROVIDER_PREFIXES[prefix as keyof typeof CANONICAL_PROVIDER_PREFIXES];
if (provider === undefined) return undefined;
const modelID = modelParts.join("/").replace(/:free$/, "");
const candidates = canonicalCandidates(provider, modelID);
const match = candidates.find((candidate) => {
return canonicalModelExists(provider, candidate);
});
return match === undefined
? undefined
: {
from: `${provider}/${match}`,
provider,
modelID: match,
};
}
function canonicalModelExists(provider: string, modelID: string) {
let files = modelFilesByProvider.get(provider);
if (files === undefined) {
try {
files = new Set(readdirSync(path.join(PROVIDERS_DIR, provider, "models")));
} catch {
files = new Set();
}
modelFilesByProvider.set(provider, files);
}
return files.has(`${modelID}.toml`);
}
function canonicalOmit(
provider: string,
modelID: string,
cost: SyncedFullModel["cost"],
limit: SyncedFullModel["limit"],
) {
const toml = canonicalToml(provider, modelID);
const omit = ["provider", "experimental"].filter((key) => toml[key] !== undefined);
const baseCost = toml.cost;
if (baseCost !== undefined && baseCost !== null && typeof baseCost === "object" && !Array.isArray(baseCost)) {
if (cost === undefined) {
omit.push("cost");
} else {
for (const key of ["reasoning", "cache_read", "cache_write", "input_audio", "output_audio", "tiers"] as const) {
if ((baseCost as Record<string, unknown>)[key] !== undefined && cost[key] === undefined) {
omit.push(`cost.${key}`);
}
}
if (hasLegacyContextOver200k(baseCost) && cost.tiers === undefined) {
omit.push("cost.context_over_200k");
}
}
}
const baseLimit = toml.limit;
if (
baseLimit !== undefined &&
baseLimit !== null &&
typeof baseLimit === "object" &&
!Array.isArray(baseLimit) &&
(baseLimit as Record<string, unknown>).input !== undefined &&
limit.input === undefined
) {
omit.push("limit.input");
}
return omit.length > 0 ? omit : undefined;
}
function hasLegacyContextOver200k(cost: object) {
const tiers = (cost as { tiers?: unknown }).tiers;
if (!Array.isArray(tiers) || tiers.length !== 1) return false;
const tier = tiers[0];
if (tier === null || typeof tier !== "object" || Array.isArray(tier)) return false;
const tierConfig = (tier as { tier?: unknown }).tier;
if (tierConfig === null || typeof tierConfig !== "object" || Array.isArray(tierConfig)) return false;
const size = (tierConfig as { size?: unknown }).size;
return typeof size === "number" && size >= 200_000;
}
function canonicalRuntimeOverrides(
provider: string,
modelID: string,
values: Pick<SyncedFullModel, "name" | "attachment" | "reasoning">,
) {
const toml = canonicalToml(provider, modelID);
return Object.fromEntries(
Object.entries(values).filter(([key, value]) => value !== undefined && toml[key] !== value),
);
}
function canonicalToml(provider: string, modelID: string) {
const key = `${provider}/${modelID}`;
let toml = canonicalTomlByModel.get(key);
if (toml === undefined) {
const filePath = path.join(PROVIDERS_DIR, provider, "models", `${modelID}.toml`);
toml = Bun.TOML.parse(readFileSync(filePath, "utf8")) as Record<string, unknown>;
canonicalTomlByModel.set(key, toml);
}
return toml;
}
function canonicalCandidates(provider: string, modelID: string) {
const candidates = [modelID];
if (provider === "anthropic") {
candidates.push(modelID.replace(/(claude-(?:opus|sonnet|haiku)-\d+)\.(\d+)/, "$1-$2"));
candidates.push(modelID.replace(/^claude-3\.5-/, "claude-3-5-"));
}
if (provider === "llama") {
candidates.push(modelID.replace(/^llama-(\d+)-(\d+)/, "llama-$1.$2"));
candidates.push(modelID.replace(/^llama-(4)-(maverick|scout)$/, "llama-$1-$2-17b"));
}
if (provider === "mistral") {
candidates.push(modelID.replace(/-latest$/, ""));
}
if (provider === "minimax") {
candidates.push(modelID.replace(/^minimax-m/, "MiniMax-M"));
}
return [...new Set(candidates)];
}
+211
View File
@@ -0,0 +1,211 @@
import { z } from "zod";
import type { ExistingModel, SyncProvider, SyncedModel } from "../index.js";
const API_BASE = "https://api.x.ai/v1";
const XAIModel = z.object({
id: z.string(),
canonical_id: z.string().optional(),
created: z.number().int().nonnegative(),
aliases: z.array(z.string()).optional(),
input_modalities: z.array(z.string()).optional(),
output_modalities: z.array(z.string()).optional(),
prompt_text_token_price: z.number().int().nonnegative().optional(),
cached_prompt_text_token_price: z.number().int().nonnegative().optional(),
completion_text_token_price: z.number().int().nonnegative().optional(),
max_prompt_length: z.number().int().nonnegative().optional(),
}).passthrough();
const XAIModelList = z.object({
models: z.array(XAIModel),
}).passthrough();
const XAIResponse = z.object({
models: z.array(XAIModel),
});
const XAIAPIKey = z.object({
acls: z.array(z.string()),
}).passthrough();
type XAIModel = z.infer<typeof XAIModel>;
export const xai = {
id: "xai",
name: "xAI",
modelsDir: "providers/xai/models",
skipCreates: true,
sourceID(model) {
return model.id;
},
skippedNotice(ids) {
if (ids.length === 0) return [];
return [
`${ids.length} xAI models returned by the API were not created because the Models API does not provide enough authoritative metadata for the catalog, especially output token limits and some feature/capability flags. Existing models are still updated from API-authoritative fields.`,
`Skipped remote IDs: ${ids.map((id) => `\`${id}\``).join(", ")}`,
];
},
async fetchModels() {
const key = process.env.XAI_API_KEY;
if (key === undefined) throw new Error("xAI sync requires XAI_API_KEY");
await assertFullModelAccess(key);
const models = await Promise.all([
fetchTypedModels(key, "language-models"),
fetchTypedModels(key, "image-generation-models"),
fetchTypedModels(key, "video-generation-models"),
]);
return { models: models.flat() };
},
parseModels(raw) {
const models = XAIResponse.parse(raw).models;
const seen = new Set<string>();
const expanded: XAIModel[] = [];
for (const model of models) {
if (!seen.has(model.id)) {
seen.add(model.id);
expanded.push(model);
}
}
for (const model of models) {
for (const alias of model.aliases ?? []) {
if (seen.has(alias)) continue;
seen.add(alias);
expanded.push({ ...model, id: alias, canonical_id: model.id });
}
}
return expanded;
},
translateModel(model, context) {
const existing = context.existing(model.id);
if (existing === undefined) return undefined;
return {
id: model.id,
model: buildModel(model, existing),
};
},
} satisfies SyncProvider<XAIModel>;
async function assertFullModelAccess(key: string) {
const response = await fetch(`${API_BASE}/api-key`, {
headers: { Authorization: `Bearer ${key}` },
});
if (!response.ok) {
throw new Error(`xAI API key metadata request failed: ${response.status} ${response.statusText}`);
}
const apiKey = XAIAPIKey.parse(await response.json());
if (!apiKey.acls.includes("api-key:model:*")) {
throw new Error("xAI sync requires XAI_API_KEY to include api-key:model:* so the model list is not ACL-filtered");
}
}
async function fetchTypedModels(key: string, endpoint: string) {
const response = await fetch(`${API_BASE}/${endpoint}`, {
headers: { Authorization: `Bearer ${key}` },
});
if (!response.ok) {
throw new Error(`xAI ${endpoint} request failed: ${response.status} ${response.statusText}`);
}
return XAIModelList.parse(await response.json()).models;
}
function dateFromTimestamp(timestamp: number) {
return new Date(timestamp * 1000).toISOString().slice(0, 10);
}
type Modality = "text" | "audio" | "image" | "video" | "pdf";
function modalities(values: string[] | undefined, fallback: Modality[]) {
const allowed = new Set<Modality>(["text", "audio", "image", "video", "pdf"]);
const result = (values ?? [])
.map((value) => value.toLowerCase())
.filter((value): value is Modality => allowed.has(value as Modality));
if (result.includes("image")) result.push("pdf");
return [...new Set(result.length > 0 ? result : fallback)];
}
function tokenPrice(value: number | undefined) {
if (value === undefined) return undefined;
return value / 10_000;
}
function preservedCostTiers(existing: ExistingModel) {
// The xAI models API exposes base pricing only; long-context tiers are curated from xAI docs/console.
return existing.cost?.tiers;
}
function cost(model: XAIModel, existing: ExistingModel) {
const input = tokenPrice(model.prompt_text_token_price);
const output = tokenPrice(model.completion_text_token_price);
if (input === undefined || output === undefined) return existing.cost;
return {
input,
output,
reasoning: existing.cost?.reasoning,
cache_read: tokenPrice(model.cached_prompt_text_token_price),
cache_write: existing.cost?.cache_write,
input_audio: existing.cost?.input_audio,
output_audio: existing.cost?.output_audio,
tiers: preservedCostTiers(existing),
};
}
function buildModel(model: XAIModel, existing: ExistingModel): SyncedModel {
const name = existing.name;
const attachment = existing.attachment;
const reasoning = existing.reasoning;
const toolCall = existing.tool_call;
const openWeights = existing.open_weights;
const limit = existing.limit;
const releaseDate = existing.release_date;
const lastUpdated = existing.last_updated;
if (
name === undefined
|| attachment === undefined
|| reasoning === undefined
|| toolCall === undefined
|| openWeights === undefined
|| limit === undefined
|| (model.canonical_id !== undefined && releaseDate === undefined)
|| (model.canonical_id !== undefined && lastUpdated === undefined)
) {
throw new Error(`xAI model ${model.id} has incomplete local TOML metadata required for sync`);
}
const input = modalities(model.input_modalities, existing.modalities?.input ?? ["text"]);
const output = modalities(model.output_modalities, existing.modalities?.output ?? ["text"]);
const created = dateFromTimestamp(model.created);
return {
name,
family: existing.family,
release_date: model.canonical_id === undefined ? created : releaseDate!,
last_updated: model.canonical_id === undefined ? created : lastUpdated!,
attachment: input.some((value) => value !== "text"),
reasoning,
temperature: existing.temperature,
tool_call: toolCall,
structured_output: existing.structured_output,
knowledge: existing.knowledge,
open_weights: openWeights,
status: existing.status,
interleaved: existing.interleaved,
cost: cost(model, existing),
limit: {
input: limit.input,
context: model.max_prompt_length ?? limit.context,
output: limit.output,
},
modalities: { input, output },
};
}
+1
View File
@@ -6,6 +6,7 @@
"build": "./script/build.ts"
},
"dependencies": {
"@tanstack/virtual-core": "^3.14.0",
"hono": "^4.8.0",
"models.dev": "workspace:*"
},
+32 -8
View File
@@ -41,6 +41,8 @@
html,
body {
height: 100%;
overflow: hidden;
font-family: 'Rubik', sans-serif;
line-height: 1.6;
color: var(--color-text);
@@ -205,12 +207,17 @@ header {
}
}
.table-viewport {
height: calc(100svh - var(--header-height));
margin-top: var(--header-height);
overflow: auto;
}
table {
border-collapse: separate;
border-spacing: 0;
font-size: 0.875rem;
width: 100%;
margin-top: var(--header-height);
}
thead,
@@ -218,7 +225,7 @@ tbody {}
table thead th {
position: sticky;
top: var(--header-height);
top: 0;
border-top: 1px solid var(--color-border);
border-bottom: 1px solid var(--color-border);
font-size: 0.75rem;
@@ -264,6 +271,7 @@ td {
text-align: left;
border-bottom: 1px solid var(--color-border);
white-space: nowrap;
height: 48px;
}
tbody {
@@ -319,21 +327,37 @@ tbody {
gap: 0.375rem;
}
.provider-cell span:first-child {
flex: 0 0 auto;
}
.provider-cell img,
.provider-cell svg {
flex: 0 0 auto;
display: block;
width: 1rem;
height: 1rem;
color: var(--color-text-secondary);
}
.virtual-spacer td {
height: inherit;
padding: 0;
border-bottom: 0;
}
.empty-row td {
padding: 2rem 0.75rem;
color: var(--color-text-tertiary);
}
.empty-row div {
position: sticky;
left: 0;
width: calc(100vw - 1.5rem);
text-align: center;
}
.model-id-cell {
display: flex;
align-items: center;
justify-content: space-between;
justify-content: flex-start;
gap: 0.375rem;
}
@@ -546,4 +570,4 @@ dialog {
}
}
}
}
+295 -94
View File
@@ -1,7 +1,53 @@
import {
Virtualizer,
elementScroll,
observeElementOffset,
observeElementRect,
} from "@tanstack/virtual-core";
import {
type TableRow,
renderRow,
escapeHtml,
booleanText,
knowledgeText,
weightsText,
} from "./shared.js";
declare global {
interface Window {
__TABLE_DATA__: TableRow[];
}
}
interface VirtualizedRow extends TableRow {
key: string;
searchText: string;
sortValues: Array<string | number | undefined>;
}
type SortDirection = "asc" | "desc";
const ESTIMATED_ROW_HEIGHT = 48;
const VIRTUAL_OVERSCAN = 5;
const modal = document.getElementById("modal") as HTMLDialogElement;
const modalClose = document.getElementById("close")!;
const help = document.getElementById("help")!;
const search = document.getElementById("search")! as HTMLInputElement;
const viewport = document.getElementById("table-viewport") as HTMLElement;
const tbody = document.getElementById(
"models-table-body"
) as HTMLTableSectionElement;
const headers = Array.from(document.querySelectorAll("th.sortable"));
const columnCount = document.querySelectorAll("thead th").length;
let isLoaded = false;
let allRows: VirtualizedRow[] = [];
let visibleRows: VirtualizedRow[] = [];
let currentSort: { column: number; direction: SortDirection } = {
column: -1,
direction: "asc",
};
/////////////////////////
// URL State Management
@@ -31,10 +77,7 @@ function getColumnNameForURL(headerEl: Element): string {
}
function getColumnIndexByUrlName(name: string): number {
const headers = document.querySelectorAll("th.sortable");
return Array.from(headers).findIndex(
(header) => getColumnNameForURL(header) === name
);
return headers.findIndex((header) => getColumnNameForURL(header) === name);
}
/////////////////////////
@@ -63,32 +106,181 @@ modal.addEventListener("click", (e) => {
});
////////////////////
// Handle Sorting
// Row Data
////////////////////
let currentSort = { column: -1, direction: "asc" };
function lockColumnWidths() {
const ths = document.querySelectorAll("#models-table thead th");
const widths = Array.from(ths).map((th) => th.getBoundingClientRect().width);
function sortTable(column: number, direction: "asc" | "desc") {
const header = document.querySelectorAll("th.sortable")[column];
const columnType = header.getAttribute("data-type");
if (!columnType) return;
const measurementRow = tbody.querySelector('tr[aria-hidden="true"]');
if (measurementRow) measurementRow.remove();
// update state
currentSort = { column, direction };
updateQueryParams({
sort: getColumnNameForURL(header),
order: direction,
});
const table = document.getElementById("models-table")!;
table.style.tableLayout = "fixed";
// sort rows
const tbody = document.querySelector("table tbody")!;
const rows = Array.from(
tbody.querySelectorAll("tr")
) as HTMLTableRowElement[];
rows.sort((a, b) => {
const aValue = getCellValue(a.cells[column], columnType);
const bValue = getCellValue(b.cells[column], columnType);
const colgroup = document.createElement("colgroup");
for (const width of widths) {
const col = document.createElement("col");
col.style.width = `${width}px`;
colgroup.appendChild(col);
}
table.insertBefore(colgroup, table.firstChild);
}
function prepareRow(row: TableRow): VirtualizedRow {
const sortValues: VirtualizedRow["sortValues"] = [
row.providerName,
row.modelName,
row.family,
row.providerId,
row.modelId,
booleanText(row.toolCall),
booleanText(row.reasoning),
row.input.length,
row.output.length,
row.inputCost,
row.outputCost,
row.reasoningCost,
row.cacheReadCost,
row.cacheWriteCost,
row.audioInputCost,
row.audioOutputCost,
row.contextLimit,
row.inputLimit,
row.outputLimit,
row.structuredOutput === undefined
? undefined
: booleanText(row.structuredOutput),
booleanText(row.temperature),
weightsText(row.openWeights),
row.knowledge ? knowledgeText(row.knowledge) : undefined,
row.releaseDate,
row.lastUpdated,
];
const searchableValues = [
row.providerName,
row.modelName,
row.family ?? "",
row.providerId,
row.modelId,
row.releaseDate,
row.lastUpdated,
];
return {
...row,
key: `${row.providerId}/${row.modelId}`,
searchText: searchableValues.join(" ").toLowerCase(),
sortValues,
};
}
////////////////////
// Virtual Table
////////////////////
function getVirtualizerOptions(count: number) {
return {
count,
getScrollElement: () => viewport,
estimateSize: () => ESTIMATED_ROW_HEIGHT,
getItemKey: (index: number) => visibleRows[index]?.key ?? index,
initialRect: {
width: viewport.clientWidth || window.innerWidth,
height: viewport.clientHeight || window.innerHeight,
},
overscan: VIRTUAL_OVERSCAN,
observeElementRect,
observeElementOffset,
scrollToFn: elementScroll,
onChange: () => renderVirtualRows(),
};
}
const virtualizer = new Virtualizer<HTMLElement, HTMLTableRowElement>(
getVirtualizerOptions(0)
);
const cleanupVirtualizer = virtualizer._didMount();
virtualizer._willUpdate();
window.addEventListener("pagehide", () => cleanupVirtualizer());
function renderStatusRow(message: string) {
tbody.innerHTML = `<tr class="empty-row"><td colspan="${columnCount}"><div>${escapeHtml(
message
)}</div></td></tr>`;
}
function setVirtualizerCount(count: number, resetScroll: boolean) {
virtualizer.setOptions(getVirtualizerOptions(count));
virtualizer._willUpdate();
if (resetScroll) virtualizer.scrollToOffset(0);
renderVirtualRows();
}
function renderVirtualRows() {
if (!isLoaded) return;
if (visibleRows.length === 0) {
renderStatusRow("No models found");
return;
}
const virtualRows = virtualizer.getVirtualItems();
if (virtualRows.length === 0) return;
const firstRow = virtualRows[0]!;
const lastRow = virtualRows[virtualRows.length - 1]!;
const paddingTop = firstRow.start;
const paddingBottom = Math.max(virtualizer.getTotalSize() - lastRow.end, 0);
const html: string[] = [];
if (paddingTop > 0) html.push(renderSpacerRow(paddingTop));
for (const virtualRow of virtualRows) {
const row = visibleRows[virtualRow.index];
if (row) html.push(renderRow(row, virtualRow.index));
}
if (paddingBottom > 0) html.push(renderSpacerRow(paddingBottom));
tbody.innerHTML = html.join("");
tbody.querySelectorAll<HTMLTableRowElement>("tr[data-index]").forEach((row) =>
virtualizer.measureElement(row)
);
}
function renderSpacerRow(height: number) {
return `<tr class="virtual-spacer" style="height: ${height}px"><td colspan="${columnCount}"></td></tr>`;
}
function applyRows(resetScroll = true) {
if (!isLoaded) return;
visibleRows = getRowsForDisplay();
setVirtualizerCount(visibleRows.length, resetScroll);
}
////////////////////
// Sorting
////////////////////
function getRowsForDisplay() {
const terms = search.value
.toLowerCase()
.split(",")
.map((term) => term.trim())
.filter(Boolean);
const filteredRows =
terms.length === 0
? allRows
: allRows.filter((row) =>
terms.some((term) => row.searchText.includes(term))
);
if (currentSort.column === -1) return filteredRows;
const columnType = headers[currentSort.column]?.getAttribute("data-type");
if (!columnType) return filteredRows;
return [...filteredRows].sort((a, b) => {
const aValue = a.sortValues[currentSort.column];
const bValue = b.sortValues[currentSort.column];
// Handle undefined values - always sort to bottom
if (aValue === undefined && bValue === undefined) return 0;
if (aValue === undefined) return 1;
if (bValue === undefined) return -1;
@@ -96,45 +288,48 @@ function sortTable(column: number, direction: "asc" | "desc") {
let comparison = 0;
if (columnType === "number" || columnType === "modalities") {
comparison = (aValue as number) - (bValue as number);
} else if (columnType === "boolean") {
comparison = (aValue as string).localeCompare(bValue as string);
} else {
comparison = (aValue as string).localeCompare(bValue as string);
comparison = String(aValue).localeCompare(String(bValue));
}
return direction === "asc" ? comparison : -comparison;
return currentSort.direction === "asc" ? comparison : -comparison;
});
rows.forEach((row) => tbody.appendChild(row));
}
// update sort indicators
const headers = document.querySelectorAll("th.sortable");
function sortTable(
column: number,
direction: SortDirection,
updateURL = true
) {
const header = headers[column];
if (!header?.getAttribute("data-type")) return;
currentSort = { column, direction };
if (updateURL) {
updateQueryParams({
sort: getColumnNameForURL(header),
order: direction,
});
}
updateSortIndicators();
applyRows();
}
function updateSortIndicators() {
headers.forEach((header, i) => {
const indicator = header.querySelector(".sort-indicator")!;
if (i === column) {
indicator.textContent = direction === "asc" ? "↑" : "↓";
} else {
indicator.textContent = "";
}
indicator.textContent =
i === currentSort.column
? currentSort.direction === "asc"
? "↑"
: ""
: "";
});
}
function getCellValue(
cell: HTMLTableCellElement,
type: string
): string | number | undefined {
if (type === "modalities")
return cell.querySelectorAll(".modality-icon").length;
const text = cell.textContent?.trim() || "";
if (text === "-") return;
if (type === "number") return parseFloat(text.replace(/[$,]/g, "")) || 0;
return text;
}
document.querySelectorAll("th.sortable").forEach((header) => {
headers.forEach((header, column) => {
header.addEventListener("click", () => {
const column = Array.from(header.parentElement!.children).indexOf(header);
const direction =
currentSort.column === column && currentSort.direction === "asc"
? "desc"
@@ -144,34 +339,19 @@ document.querySelectorAll("th.sortable").forEach((header) => {
});
///////////////////
// Handle Search
// Search
///////////////////
function filterTable(value: string) {
const lowerCaseValues = value.toLowerCase().split(",").filter(str => str.trim() !== "");
const rows = document.querySelectorAll(
"table tbody tr"
) as NodeListOf<HTMLTableRowElement>;
rows.forEach((row) => {
const cellTexts = Array.from(row.cells).map((cell) =>
cell.textContent!.toLowerCase()
);
const isVisible = lowerCaseValues.length === 0 ||
lowerCaseValues.some((lowerCaseValue) => cellTexts.some((text) => text.includes(lowerCaseValue)));
row.style.display = isVisible ? "" : "none";
});
updateQueryParams({ search: value || null });
}
search.addEventListener("input", () => {
filterTable(search.value);
updateQueryParams({ search: search.value || null });
applyRows();
});
document.addEventListener("keydown", (e) => {
if ((e.metaKey || e.ctrlKey) && e.key === "k") {
const key = e.key.toLowerCase();
if ((e.metaKey || e.ctrlKey) && (key === "k" || key === "f")) {
e.preventDefault();
search.focus();
search.select();
}
});
@@ -185,10 +365,7 @@ search.addEventListener("keydown", (e) => {
///////////////////////////////////
// Handle Copy model ID function
///////////////////////////////////
(window as any).copyModelId = async (
button: HTMLButtonElement,
modelId: string
) => {
async function copyModelId(button: HTMLButtonElement, modelId: string) {
try {
if (navigator.clipboard) {
await navigator.clipboard.writeText(modelId);
@@ -209,7 +386,18 @@ search.addEventListener("keydown", (e) => {
} catch (err) {
console.error("Failed to copy text: ", err);
}
};
}
document.addEventListener("click", (event) => {
if (!(event.target instanceof Element)) return;
const button = event.target.closest<HTMLButtonElement>(
".copy-button[data-model-id]"
);
if (!button) return;
const modelId = button.dataset.modelId;
if (modelId) void copyModelId(button, modelId);
});
///////////////////////////////////
// Initialize State from URL
@@ -217,24 +405,37 @@ search.addEventListener("keydown", (e) => {
function initializeFromURL() {
const params = getQueryParams();
(() => {
const searchQuery = params.get("search");
if (!searchQuery) return;
search.value = searchQuery;
filterTable(searchQuery);
})();
(() => {
const columnName = params.get("sort");
if (!columnName) return;
search.value = params.get("search") ?? "";
currentSort = { column: -1, direction: "asc" };
const columnName = params.get("sort");
if (columnName) {
const columnIndex = getColumnIndexByUrlName(columnName);
if (columnIndex === -1) return;
if (columnIndex !== -1) {
currentSort = {
column: columnIndex,
direction: params.get("order") === "desc" ? "desc" : "asc",
};
}
}
const direction = (params.get("order") as "asc" | "desc") || "asc";
sortTable(columnIndex, direction);
})();
updateSortIndicators();
applyRows(false);
}
document.addEventListener("DOMContentLoaded", initializeFromURL);
function loadRows() {
try {
allRows = window.__TABLE_DATA__.map(prepareRow);
lockColumnWidths();
isLoaded = true;
initializeFromURL();
} catch (error) {
console.error(error);
isLoaded = true;
visibleRows = [];
renderStatusRow("Failed to load model data");
}
}
loadRows();
window.addEventListener("popstate", initializeFromURL);
+55 -232
View File
@@ -6,6 +6,7 @@ import { Fragment } from "hono/jsx";
import { renderToString } from "hono/jsx/dom/server";
import { existsSync } from "fs";
import path from "path";
import { type TableRow, renderRow, getLargestRow } from "./shared.js";
export const Providers = await generate(
path.join(import.meta.dir, "..", "..", "..", "providers")
@@ -38,7 +39,6 @@ const loadProviderSvg = async (providerId: string): Promise<string | null> => {
const file = Bun.file(providerLogoPath);
return await file.text();
}
//
// Fall back to default logo
if (existsSync(defaultLogoPath)) {
const file = Bun.file(defaultLogoPath);
@@ -62,122 +62,47 @@ for (const [providerId] of Object.entries(Providers)) {
}
}
function renderProviderLogo(providerId: string) {
const svgContent = providerLogos.get(providerId) || "";
export const INITIAL_ROW_COUNT = 50;
return <span dangerouslySetInnerHTML={{ __html: svgContent }} />;
}
export const TableRows: TableRow[] = Object.entries(Providers)
.sort(([, providerA], [, providerB]) =>
providerA.name.localeCompare(providerB.name)
)
.flatMap(([providerId, provider]) =>
Object.entries(provider.models)
.filter(([, model]) => model.status !== "alpha")
.sort(([, modelA], [, modelB]) => modelA.name.localeCompare(modelB.name))
.map(([modelId, model]) => ({
providerId,
providerName: provider.name,
providerLogoSvg: providerLogos.get(providerId) || "",
modelId,
modelName: model.name,
family: model.family,
toolCall: model.tool_call,
reasoning: model.reasoning,
input: model.modalities.input,
output: model.modalities.output,
inputCost: model.cost?.input,
outputCost: model.cost?.output,
reasoningCost: model.cost?.reasoning,
cacheReadCost: model.cost?.cache_read,
cacheWriteCost: model.cost?.cache_write,
audioInputCost: model.cost?.input_audio,
audioOutputCost: model.cost?.output_audio,
contextLimit: model.limit.context,
inputLimit: model.limit.input,
outputLimit: model.limit.output,
structuredOutput: model.structured_output,
temperature: model.temperature ?? false,
openWeights: model.open_weights,
knowledge: model.knowledge,
releaseDate: model.release_date,
lastUpdated: model.last_updated,
}))
);
const getModalityIcon = (modality: string) => {
switch (modality) {
case "text":
return (
<span class="modality-icon" data-tooltip="Text">
<svg
xmlns="http://www.w3.org/2000/svg"
width="16"
height="16"
viewBox="0 0 24 24"
fill="none"
stroke="currentColor"
stroke-width="2"
stroke-linecap="round"
stroke-linejoin="round"
>
<polyline points="4,7 4,4 20,4 20,7"></polyline>
<line x1="9" y1="20" x2="15" y2="20"></line>
<line x1="12" y1="4" x2="12" y2="20"></line>
</svg>
</span>
);
case "image":
return (
<span class="modality-icon" data-tooltip="Image">
<svg
xmlns="http://www.w3.org/2000/svg"
width="16"
height="16"
viewBox="0 0 24 24"
fill="none"
stroke="currentColor"
stroke-width="2"
stroke-linecap="round"
stroke-linejoin="round"
>
<rect width="18" height="18" x="3" y="3" rx="2" ry="2"></rect>
<circle cx="9" cy="9" r="2"></circle>
<path d="m21 15-3.086-3.086a2 2 0 0 0-2.828 0L6 21"></path>
</svg>
</span>
);
case "audio":
return (
<span class="modality-icon" data-tooltip="Audio">
<svg
xmlns="http://www.w3.org/2000/svg"
width="16"
height="16"
viewBox="0 0 24 24"
fill="none"
stroke="currentColor"
stroke-width="2"
stroke-linecap="round"
stroke-linejoin="round"
>
<polygon points="11 5 6 9 2 9 2 15 6 15 11 19 11 5"></polygon>
<path d="m19.07 4.93a10 10 0 0 1 0 14.14M15.54 8.46a5 5 0 0 1 0 7.07"></path>
</svg>
</span>
);
case "video":
return (
<span class="modality-icon" data-tooltip="Video">
<svg
xmlns="http://www.w3.org/2000/svg"
width="16"
height="16"
viewBox="0 0 24 24"
fill="none"
stroke="currentColor"
stroke-width="2"
stroke-linecap="round"
stroke-linejoin="round"
>
<path d="m22 8-6 4 6 4V8Z"></path>
<rect width="14" height="12" x="2" y="6" rx="2" ry="2"></rect>
</svg>
</span>
);
case "pdf":
return (
<span class="modality-icon" data-tooltip="PDF">
<svg
xmlns="http://www.w3.org/2000/svg"
width="16"
height="16"
viewBox="0 0 24 24"
fill="none"
stroke="currentColor"
stroke-width="2"
stroke-linecap="round"
stroke-linejoin="round"
>
<path d="M14 2H6a2 2 0 0 0-2 2v16a2 2 0 0 0 2 2h12a2 2 0 0 0 2-2V8z"></path>
<polyline points="14,2 14,8 20,8"></polyline>
<line x1="16" y1="13" x2="8" y2="13"></line>
<line x1="16" y1="17" x2="8" y2="17"></line>
<polyline points="10,9 9,9 8,9"></polyline>
</svg>
</span>
);
default:
return null;
}
};
const renderCost = (cost?: number) => {
return cost === undefined ? "-" : `$${cost.toFixed(2)}`;
};
const largestRow = getLargestRow(TableRows);
export const Rendered = renderToString(
<Fragment>
@@ -213,7 +138,8 @@ export const Rendered = renderToString(
<button id="help">How to use</button>
</div>
</header>
<table>
<div id="table-viewport" class="table-viewport">
<table id="models-table">
<thead>
<tr>
<th class="sortable" data-type="text">
@@ -342,120 +268,12 @@ export const Rendered = renderToString(
</th>
</tr>
</thead>
<tbody>
{Object.entries(Providers)
.sort(([, providerA], [, providerB]) =>
providerA.name.localeCompare(providerB.name)
)
.flatMap(([providerId, provider]) =>
Object.entries(provider.models)
.filter(([, model]) => model.status !== "alpha")
.sort(([, modelA], [, modelB]) =>
modelA.name.localeCompare(modelB.name)
)
.map(([modelId, model]) => (
<tr key={`${providerId}-${modelId}`}>
<td>
<div class="provider-cell">
{renderProviderLogo(providerId)}
<span>{provider.name}</span>
</div>
</td>
<td>{model.name}</td>
<td>{model.family ?? "-"}</td>
<td>{providerId}</td>
<td>
<div class="model-id-cell">
<span class="model-id-text">{modelId}</span>
<button
class="copy-button"
onclick={`copyModelId(this, '${modelId}')`}
>
<svg
class="copy-icon"
xmlns="http://www.w3.org/2000/svg"
width="14"
height="14"
viewBox="0 0 24 24"
fill="none"
stroke="currentColor"
stroke-width="2"
stroke-linecap="round"
stroke-linejoin="round"
>
<rect
width="14"
height="14"
x="8"
y="8"
rx="2"
ry="2"
/>
<path d="m4 16c-1.1 0-2-.9-2-2V4c0-1.1.9-2 2-2h10c1.1 0 2 .9 2 2" />
</svg>
<svg
class="check-icon"
xmlns="http://www.w3.org/2000/svg"
width="14"
height="14"
viewBox="0 0 24 24"
fill="none"
stroke="currentColor"
stroke-width="2"
stroke-linecap="round"
stroke-linejoin="round"
style="display: none;"
>
<polyline points="20,6 9,17 4,12" />
</svg>
</button>
</div>
</td>
<td>{model.tool_call ? "Yes" : "No"}</td>
<td>{model.reasoning ? "Yes" : "No"}</td>
<td>
<div class="modalities">
{model.modalities.input.map((modality) =>
getModalityIcon(modality)
)}
</div>
</td>
<td>
<div class="modalities">
{model.modalities.output.map((modality) =>
getModalityIcon(modality)
)}
</div>
</td>
<td>{renderCost(model.cost?.input)}</td>
<td>{renderCost(model.cost?.output)}</td>
<td>{renderCost(model.cost?.reasoning)}</td>
<td>{renderCost(model.cost?.cache_read)}</td>
<td>{renderCost(model.cost?.cache_write)}</td>
<td>{renderCost(model.cost?.input_audio)}</td>
<td>{renderCost(model.cost?.output_audio)}</td>
<td>{model.limit.context.toLocaleString()}</td>
<td>{model.limit.input?.toLocaleString() ?? "-"}</td>
<td>{model.limit.output.toLocaleString()}</td>
<td>
{model.structured_output === undefined
? "-"
: model.structured_output
? "Yes"
: "No"}
</td>
<td>{model.temperature ? "Yes" : "No"}</td>
<td>{model.open_weights ? "Open" : "Closed"}</td>
<td>
{model.knowledge ? model.knowledge.substring(0, 7) : "-"}
</td>
<td>{model.release_date}</td>
<td>{model.last_updated}</td>
</tr>
))
)}
</tbody>
</table>
<tbody id="models-table-body" dangerouslySetInnerHTML={{
__html: TableRows.slice(0, INITIAL_ROW_COUNT).map((row, i) => renderRow(row, i)).join('')
+ renderRow(largestRow, -1).replace('<tr', '<tr style="visibility:hidden" aria-hidden="true"')
}} />
</table>
</div>
<dialog id="modal">
<div class="header">
<h2>How to use</h2>
@@ -565,10 +383,15 @@ export const Rendered = renderToString(
>
Edit on GitHub
</a>
<a href="https://sst.dev" target="_blank" rel="noopener noreferrer">
Created by SST
<a href="https://opencode.ai" target="_blank" rel="noopener noreferrer">
Created by OpenCode
</a>
</div>
</dialog>
<script
dangerouslySetInnerHTML={{
__html: `window.__TABLE_DATA__ = ${JSON.stringify(TableRows)}`,
}}
></script>
</Fragment>
);
+176
View File
@@ -0,0 +1,176 @@
export interface TableRow {
providerId: string;
providerName: string;
providerLogoSvg: string;
modelId: string;
modelName: string;
family?: string;
toolCall: boolean;
reasoning: boolean;
input: string[];
output: string[];
inputCost?: number;
outputCost?: number;
reasoningCost?: number;
cacheReadCost?: number;
cacheWriteCost?: number;
audioInputCost?: number;
audioOutputCost?: number;
contextLimit: number;
inputLimit?: number;
outputLimit: number;
structuredOutput?: boolean;
temperature: boolean;
openWeights: boolean;
knowledge?: string;
releaseDate: string;
lastUpdated: string;
}
const MODALITY_ICONS: Record<string, string> = {
text: `<svg xmlns="http://www.w3.org/2000/svg" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polyline points="4,7 4,4 20,4 20,7"></polyline><line x1="9" y1="20" x2="15" y2="20"></line><line x1="12" y1="4" x2="12" y2="20"></line></svg>`,
image: `<svg xmlns="http://www.w3.org/2000/svg" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><rect width="18" height="18" x="3" y="3" rx="2" ry="2"></rect><circle cx="9" cy="9" r="2"></circle><path d="m21 15-3.086-3.086a2 2 0 0 0-2.828 0L6 21"></path></svg>`,
audio: `<svg xmlns="http://www.w3.org/2000/svg" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><polygon points="11 5 6 9 2 9 2 15 6 15 11 19 11 5"></polygon><path d="m19.07 4.93a10 10 0 0 1 0 14.14M15.54 8.46a5 5 0 0 1 0 7.07"></path></svg>`,
video: `<svg xmlns="http://www.w3.org/2000/svg" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="m22 8-6 4 6 4V8Z"></path><rect width="14" height="12" x="2" y="6" rx="2" ry="2"></rect></svg>`,
pdf: `<svg xmlns="http://www.w3.org/2000/svg" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><path d="M14 2H6a2 2 0 0 0-2 2v16a2 2 0 0 0 2 2h12a2 2 0 0 0 2-2V8z"></path><polyline points="14,2 14,8 20,8"></polyline><line x1="16" y1="13" x2="8" y2="13"></line><line x1="16" y1="17" x2="8" y2="17"></line><polyline points="10,9 9,9 8,9"></polyline></svg>`,
};
export function escapeHtml(value: string | number) {
return String(value).replace(/[&<>'"]/g, (char) => {
switch (char) {
case "&":
return "&amp;";
case "<":
return "&lt;";
case ">":
return "&gt;";
case "'":
return "&#39;";
case '"':
return "&quot;";
default:
return char;
}
});
}
export function booleanText(value: boolean) {
return value ? "Yes" : "No";
}
export function optionalBooleanText(value?: boolean) {
return value === undefined ? "-" : booleanText(value);
}
export function formatCost(cost?: number) {
return cost === undefined ? "-" : `$${cost.toFixed(2)}`;
}
export function formatNumber(value?: number) {
return value === undefined ? "-" : value.toLocaleString();
}
export function knowledgeText(value?: string) {
return value ? value.substring(0, 7) : "-";
}
export function weightsText(value: boolean) {
return value ? "Open" : "Closed";
}
export function renderModalityIcon(modality: string) {
const label =
modality === "pdf"
? "PDF"
: modality[0]!.toUpperCase() + modality.slice(1);
const icon = MODALITY_ICONS[modality];
if (!icon) return "";
return `<span class="modality-icon" data-tooltip="${label}">${icon}</span>`;
}
export function renderModalities(modalities: string[]) {
return `<div class="modalities">${modalities
.map(renderModalityIcon)
.join("")}</div>`;
}
export function renderCopyButton(modelId: string) {
const escapedModelId = escapeHtml(modelId);
return `<button type="button" class="copy-button" data-model-id="${escapedModelId}" aria-label="Copy model ID"><svg class="copy-icon" xmlns="http://www.w3.org/2000/svg" width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round"><rect width="14" height="14" x="8" y="8" rx="2" ry="2"></rect><path d="m4 16c-1.1 0-2-.9-2-2V4c0-1.1.9-2 2-2h10c1.1 0 2 .9 2 2"></path></svg><svg class="check-icon" xmlns="http://www.w3.org/2000/svg" width="14" height="14" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" style="display: none;"><polyline points="20,6 9,17 4,12"></polyline></svg></button>`;
}
export function renderRow(row: TableRow, index: number) {
return `<tr data-index="${index}">
<td><div class="provider-cell">${row.providerLogoSvg}<span>${escapeHtml(
row.providerName
)}</span></div></td>
<td>${escapeHtml(row.modelName)}</td>
<td>${escapeHtml(row.family ?? "-")}</td>
<td>${escapeHtml(row.providerId)}</td>
<td><div class="model-id-cell"><span class="model-id-text">${escapeHtml(
row.modelId
)}</span>${renderCopyButton(row.modelId)}</div></td>
<td>${booleanText(row.toolCall)}</td>
<td>${booleanText(row.reasoning)}</td>
<td>${renderModalities(row.input)}</td>
<td>${renderModalities(row.output)}</td>
<td>${formatCost(row.inputCost)}</td>
<td>${formatCost(row.outputCost)}</td>
<td>${formatCost(row.reasoningCost)}</td>
<td>${formatCost(row.cacheReadCost)}</td>
<td>${formatCost(row.cacheWriteCost)}</td>
<td>${formatCost(row.audioInputCost)}</td>
<td>${formatCost(row.audioOutputCost)}</td>
<td>${formatNumber(row.contextLimit)}</td>
<td>${formatNumber(row.inputLimit)}</td>
<td>${formatNumber(row.outputLimit)}</td>
<td>${optionalBooleanText(row.structuredOutput)}</td>
<td>${booleanText(row.temperature)}</td>
<td>${weightsText(row.openWeights)}</td>
<td>${knowledgeText(row.knowledge)}</td>
<td>${escapeHtml(row.releaseDate)}</td>
<td>${escapeHtml(row.lastUpdated)}</td>
</tr>`;
}
export function getLargestRow(rows: TableRow[]): TableRow {
const worst: TableRow = {
providerId: "", providerName: "", providerLogoSvg: "", modelId: "", modelName: "",
toolCall: true, reasoning: true,
input: [], output: [],
contextLimit: 0, outputLimit: 0,
structuredOutput: true, temperature: true, openWeights: false,
releaseDate: "", lastUpdated: "",
};
for (const row of rows) {
if (row.providerName.length > worst.providerName.length) worst.providerName = row.providerName;
if (row.modelName.length > worst.modelName.length) worst.modelName = row.modelName;
if ((row.family ?? "").length > (worst.family ?? "").length) worst.family = row.family;
if (row.providerId.length > worst.providerId.length) worst.providerId = row.providerId;
if (row.modelId.length > worst.modelId.length) worst.modelId = row.modelId;
if ((row.knowledge ?? "").length > (worst.knowledge ?? "").length) worst.knowledge = row.knowledge;
if (row.releaseDate.length > worst.releaseDate.length) worst.releaseDate = row.releaseDate;
if (row.lastUpdated.length > worst.lastUpdated.length) worst.lastUpdated = row.lastUpdated;
if (row.input.length > worst.input.length) worst.input = row.input;
if (row.output.length > worst.output.length) worst.output = row.output;
const costWider = (a: number | undefined, b: number | undefined) =>
b !== undefined && (a === undefined || formatCost(b).length > formatCost(a).length);
if (costWider(worst.inputCost, row.inputCost)) worst.inputCost = row.inputCost;
if (costWider(worst.outputCost, row.outputCost)) worst.outputCost = row.outputCost;
if (costWider(worst.reasoningCost, row.reasoningCost)) worst.reasoningCost = row.reasoningCost;
if (costWider(worst.cacheReadCost, row.cacheReadCost)) worst.cacheReadCost = row.cacheReadCost;
if (costWider(worst.cacheWriteCost, row.cacheWriteCost)) worst.cacheWriteCost = row.cacheWriteCost;
if (costWider(worst.audioInputCost, row.audioInputCost)) worst.audioInputCost = row.audioInputCost;
if (costWider(worst.audioOutputCost, row.audioOutputCost)) worst.audioOutputCost = row.audioOutputCost;
const numWider = (a: number | undefined, b: number | undefined) =>
b !== undefined && (a === undefined || formatNumber(b).length > formatNumber(a).length);
if (numWider(worst.contextLimit as number | undefined, row.contextLimit)) worst.contextLimit = row.contextLimit;
if (numWider(worst.inputLimit, row.inputLimit)) worst.inputLimit = row.inputLimit;
if (numWider(worst.outputLimit as number | undefined, row.outputLimit)) worst.outputLimit = row.outputLimit;
}
return worst;
}
@@ -1,6 +1,6 @@
name = "Qwen: Qwen-Max "
release_date = "2024-04-03"
last_updated = "2025-01-25"
name = "MiniMax-M2.7-highspeed"
release_date = "2026-03-19"
last_updated = "2026-03-19"
attachment = false
reasoning = false
temperature = true
@@ -8,13 +8,12 @@ tool_call = true
open_weights = false
[cost]
input = 1.6
output = 6.4
cache_read = 0.32
input = 0.600
output = 4.800
[limit]
context = 32768
output = 8192
context = 204_800
output = 131_072
[modalities]
input = ["text"]
+20
View File
@@ -0,0 +1,20 @@
name = "MiniMax-M2.7"
release_date = "2026-03-19"
last_updated = "2026-03-19"
attachment = false
reasoning = false
temperature = true
tool_call = true
open_weights = false
[cost]
input = 0.300
output = 1.200
[limit]
context = 204_800
output = 131_072
[modalities]
input = ["text"]
output = ["text"]
@@ -1,19 +1,17 @@
name = "Claude Haiku 3.5"
name = "claude-3-5-haiku-20241022"
family = "claude-haiku"
release_date = "2024-10-22"
last_updated = "2024-10-22"
attachment = true
reasoning = false
temperature = true
knowledge = "2024-07"
tool_call = true
open_weights = false
knowledge = "2024-07-31"
[cost]
input = 0.80
output = 4.00
cache_read = 0.08
cache_write = 1.00
input = 0.800
output = 4.000
[limit]
context = 200_000
@@ -1,19 +1,17 @@
name = "Claude Sonnet 3.5 v2"
family = "claude-sonnet"
name = "claude-3-5-haiku-latest"
family = "claude-haiku"
release_date = "2024-10-22"
last_updated = "2024-10-22"
attachment = true
reasoning = false
temperature = true
knowledge = "2024-04"
tool_call = true
open_weights = false
knowledge = "2024-07-31"
[cost]
input = 3.00
output = 15.00
cache_read = 0.30
cache_write = 3.75
input = 0.800
output = 4.000
[limit]
context = 200_000
@@ -1,12 +1,13 @@
name = "claude-haiku-4-5-20251001"
family = "claude-haiku"
release_date = "2025-10-16"
last_updated = "2025-10-16"
attachment = true
reasoning = false
reasoning = true
temperature = true
tool_call = true
open_weights = false
knowledge = "2025-03"
knowledge = "2025-02-28"
[cost]
input = 1.000
@@ -17,5 +18,5 @@ context = 200_000
output = 64_000
[modalities]
input = ["text", "image"]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -0,0 +1,22 @@
name = "claude-haiku-4-5"
family = "claude-haiku"
release_date = "2025-10-16"
last_updated = "2025-10-16"
attachment = true
reasoning = true
temperature = true
tool_call = true
open_weights = false
knowledge = "2025-02-28"
[cost]
input = 1.000
output = 5.000
[limit]
context = 200_000
output = 64_000
[modalities]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -1,12 +1,13 @@
name = "claude-opus-4-1-20250805"
family = "claude-opus"
release_date = "2025-08-05"
last_updated = "2025-08-05"
attachment = true
reasoning = false
reasoning = true
temperature = true
tool_call = true
open_weights = false
knowledge = "2025-03"
knowledge = "2025-03-31"
[cost]
input = 15.000
@@ -17,5 +18,5 @@ context = 200_000
output = 32_000
[modalities]
input = ["text", "image"]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -1,19 +1,17 @@
name = "Claude Opus 4 (US)"
name = "claude-opus-4-20250514"
family = "claude-opus"
release_date = "2025-05-22"
last_updated = "2025-05-22"
attachment = true
reasoning = true
temperature = true
knowledge = "2024-04"
tool_call = true
open_weights = false
knowledge = "2025-03-31"
[cost]
input = 15.00
output = 75.00
cache_read = 1.50
cache_write = 18.75
input = 15.000
output = 75.000
[limit]
context = 200_000
@@ -1,12 +1,13 @@
name = "claude-opus-4-5-20251101"
family = "claude-opus"
release_date = "2025-11-25"
last_updated = "2025-11-25"
attachment = true
reasoning = false
reasoning = true
temperature = true
tool_call = true
open_weights = false
knowledge = "2025-03"
knowledge = "2025-03-31"
[cost]
input = 5.000
@@ -17,5 +18,5 @@ context = 200_000
output = 64_000
[modalities]
input = ["text", "image"]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -1,23 +1,22 @@
name = "Claude Opus 4.5"
name = "claude-opus-4-5"
family = "claude-opus"
release_date = "2025-11-25"
last_updated = "2025-11-25"
attachment = true
reasoning = true
temperature = true
tool_call = true
knowledge = "2025-03"
open_weights = false
knowledge = "2025-03-31"
[cost]
input = 5.00
output = 25.00
cache_read = 0.50
cache_write = 6.25
input = 5.000
output = 25.000
[limit]
context = 200_000
output = 32_000
output = 64_000
[modalities]
input = ["text", "image"]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -0,0 +1,21 @@
name = "claude-opus-4-6-thinking"
release_date = "2026-02-06"
last_updated = "2026-03-13"
attachment = true
reasoning = true
temperature = true
tool_call = true
open_weights = false
knowledge = "2025-05"
[cost]
input = 5.000
output = 25.000
[limit]
context = 1_000_000
output = 128_000
[modalities]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -0,0 +1,22 @@
name = "claude-opus-4-6"
family = "claude-opus"
release_date = "2026-02-06"
last_updated = "2026-03-13"
attachment = true
reasoning = true
temperature = true
tool_call = true
open_weights = false
knowledge = "2025-05-31"
[cost]
input = 5.000
output = 25.000
[limit]
context = 1_000_000
output = 128_000
[modalities]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -0,0 +1,31 @@
name = "claude-opus-4-7"
family = "claude-opus"
release_date = "2026-04-17"
last_updated = "2026-04-17"
attachment = true
reasoning = true
temperature = true
tool_call = true
open_weights = false
knowledge = "2026-01-31"
[cost]
input = 5.000
output = 25.000
cache_read = 0.500
cache_write = 6.250
[[cost.tiers]]
tier = { size = 200_000 }
input = 10.000
output = 37.500
cache_read = 1.000
cache_write = 12.500
[limit]
context = 1_000_000
output = 128_000
[modalities]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -1,19 +1,17 @@
name = "Claude Sonnet 4"
name = "claude-sonnet-4-20250514"
family = "claude-sonnet"
release_date = "2025-05-22"
last_updated = "2025-05-22"
attachment = true
reasoning = true
temperature = true
knowledge = "2024-04"
tool_call = true
open_weights = false
knowledge = "2025-03-31"
[cost]
input = 3.00
output = 15.00
cache_read = 0.30
cache_write = 3.75
input = 3.000
output = 15.000
[limit]
context = 200_000
@@ -17,5 +17,5 @@ context = 200_000
output = 64_000
[modalities]
input = ["text", "image"]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -1,12 +1,13 @@
name = "claude-sonnet-4-5-20250929"
release_date = "2025-09-29"
last_updated = "2025-09-29"
family = "claude-sonnet"
release_date = "2025-09-30"
last_updated = "2025-09-30"
attachment = true
reasoning = false
reasoning = true
temperature = true
tool_call = true
open_weights = false
knowledge = "2025-03"
knowledge = "2025-07-31"
[cost]
input = 3.000
@@ -17,5 +18,5 @@ context = 200_000
output = 64_000
[modalities]
input = ["text", "image"]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -1,19 +1,17 @@
name = "Claude Sonnet 4.5"
name = "claude-sonnet-4-5"
family = "claude-sonnet"
release_date = "2025-09-29"
last_updated = "2025-09-29"
release_date = "2025-09-30"
last_updated = "2025-09-30"
attachment = true
reasoning = true
temperature = true
tool_call = true
knowledge = "2025-07-31"
open_weights = false
knowledge = "2025-07-31"
[cost]
input = 3
output = 15
cache_read = 0.3
cache_write = 3.75
input = 3.000
output = 15.000
[limit]
context = 200_000
@@ -0,0 +1,21 @@
name = "claude-sonnet-4-6-thinking"
release_date = "2026-02-18"
last_updated = "2026-03-13"
attachment = true
reasoning = true
temperature = true
tool_call = true
open_weights = false
knowledge = "2025-08"
[cost]
input = 3.000
output = 15.000
[limit]
context = 1_000_000
output = 64_000
[modalities]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -0,0 +1,22 @@
name = "claude-sonnet-4-6"
family = "claude-sonnet"
release_date = "2026-02-18"
last_updated = "2026-03-13"
attachment = true
reasoning = true
temperature = true
tool_call = true
open_weights = false
knowledge = "2025-08-31"
[cost]
input = 3.000
output = 15.000
[limit]
context = 1_000_000
output = 64_000
[modalities]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -0,0 +1,21 @@
name = "gemini-3.1-flash-image-preview"
release_date = "2026-02-27"
last_updated = "2026-02-27"
attachment = true
reasoning = false
temperature = true
tool_call = false
open_weights = false
knowledge = "2025-01"
[cost]
input = 0.500
output = 60.000
[limit]
context = 131_072
output = 32_768
[modalities]
input = ["text", "image", "pdf"]
output = ["text", "image"]
@@ -1,25 +1,22 @@
name = "GLM 4.5"
family = "glm"
name = "glm-4.5-air"
family = "glm-air"
release_date = "2025-07-29"
last_updated = "2025-07-29"
attachment = false
reasoning = true
temperature = true
tool_call = true
knowledge = "2025-04"
open_weights = true
knowledge = "2025-04"
[cost]
input = 0.55
output = 2.19
input = 0.1143
output = 0.286
[limit]
context = 131_072
output = 131_072
output = 98_304
[modalities]
input = ["text"]
output = ["text"]
[interleaved]
field = "reasoning_content"
@@ -1,20 +1,21 @@
name = "Mistral Large (24.02)"
family = "mistral-large"
release_date = "2024-12-01"
last_updated = "2024-12-01"
name = "glm-4.5-airx"
family = "glm"
release_date = "2025-07-29"
last_updated = "2025-07-29"
attachment = false
reasoning = false
temperature = true
tool_call = true
open_weights = false
knowledge = "2025-04"
[cost]
input = 0.50
output = 1.50
input = 0.572
output = 1.714
[limit]
context = 128_000
output = 4_096
output = 16_384
[modalities]
input = ["text"]
@@ -1,20 +1,21 @@
name = "Titan Text G1 - Express"
family = "titan"
release_date = "2024-12-01"
last_updated = "2024-12-01"
name = "glm-4.5-x"
family = "glm"
release_date = "2025-07-29"
last_updated = "2025-07-29"
attachment = false
reasoning = false
temperature = true
tool_call = true
open_weights = false
knowledge = "2025-04"
[cost]
input = 0.20
output = 0.60
input = 1.143
output = 2.290
[limit]
context = 128_000
output = 4_096
output = 16_384
[modalities]
input = ["text"]
+5 -4
View File
@@ -1,19 +1,20 @@
name = "GLM-4.5"
family = "glm"
release_date = "2025-07-29"
last_updated = "2025-07-29"
attachment = false
reasoning = false
reasoning = true
temperature = true
tool_call = true
open_weights = false
knowledge = "2024-10"
open_weights = true
knowledge = "2025-04"
[cost]
input = 0.286
output = 1.142
[limit]
context = 128_000
context = 131_072
output = 98_304
[modalities]
+7 -6
View File
@@ -1,12 +1,13 @@
name = "GLM-4.5V"
release_date = "2025-07-29"
last_updated = "2025-07-29"
family = "glm"
release_date = "2025-08-12"
last_updated = "2025-08-12"
attachment = true
reasoning = false
reasoning = true
temperature = true
tool_call = true
open_weights = false
knowledge = "2024-10"
open_weights = true
knowledge = "2025-04"
[cost]
input = 0.290
@@ -17,5 +18,5 @@ context = 64_000
output = 16_384
[modalities]
input = ["text", "image"]
input = ["text", "image", "video"]
output = ["text"]
+5 -4
View File
@@ -1,19 +1,20 @@
name = "glm-4.6"
family = "glm"
release_date = "2025-09-30"
last_updated = "2025-09-30"
attachment = false
reasoning = false
reasoning = true
temperature = true
tool_call = true
open_weights = false
knowledge = "2025-03"
open_weights = true
knowledge = "2025-04"
[cost]
input = 0.286
output = 1.142
[limit]
context = 200_000
context = 204_800
output = 131_072
[modalities]
+5 -4
View File
@@ -1,12 +1,13 @@
name = "GLM-4.6V"
family = "glm"
release_date = "2025-12-08"
last_updated = "2025-12-08"
attachment = true
reasoning = false
reasoning = true
temperature = true
tool_call = true
open_weights = false
knowledge = "2025-03"
open_weights = true
knowledge = "2025-04"
[cost]
input = 0.145
@@ -17,5 +18,5 @@ context = 128_000
output = 32_768
[modalities]
input = ["text", "image"]
input = ["text", "image", "video"]
output = ["text"]
@@ -1,20 +1,20 @@
name = "GLM 4.5 Air"
family = "glm-air"
release_date = "2025-08-01"
last_updated = "2025-08-01"
name = "glm-4.7-flashx"
family = "glm-flash"
release_date = "2026-01-20"
last_updated = "2026-01-20"
attachment = false
reasoning = true
temperature = true
tool_call = true
knowledge = "2025-04"
open_weights = true
knowledge = "2025-04"
[cost]
input = 0.22
output = 0.88
input = 0.0715
output = 0.429
[limit]
context = 131_072
context = 200_000
output = 131_072
[modalities]
+8 -4
View File
@@ -1,19 +1,23 @@
name = "glm-4.7"
family = "glm"
release_date = "2025-12-22"
last_updated = "2025-12-22"
attachment = false
reasoning = false
reasoning = true
temperature = true
tool_call = true
open_weights = false
knowledge = "2025-06"
open_weights = true
knowledge = "2025-04"
[interleaved]
field = "reasoning_content"
[cost]
input = 0.286
output = 1.142
[limit]
context = 200_000
context = 204_800
output = 131_072
[modalities]
+25
View File
@@ -0,0 +1,25 @@
name = "glm-5-turbo"
family = "glm"
release_date = "2026-03-16"
last_updated = "2026-03-16"
attachment = false
reasoning = true
temperature = true
tool_call = true
structured_output = true
open_weights = false
[interleaved]
field = "reasoning_content"
[cost]
input = 0.720
output = 3.200
[limit]
context = 200_000
output = 131_072
[modalities]
input = ["text"]
output = ["text"]
+25
View File
@@ -0,0 +1,25 @@
name = "glm-5.1"
family = "glm"
release_date = "2026-04-10"
last_updated = "2026-04-10"
attachment = false
reasoning = true
temperature = true
tool_call = true
structured_output = true
open_weights = false
[interleaved]
field = "reasoning_content"
[cost]
input = 0.860
output = 3.500
[limit]
context = 200_000
output = 131_072
[modalities]
input = ["text"]
output = ["text"]
@@ -1,4 +1,4 @@
name = "GLM5"
name = "glm-5"
family = "glm"
release_date = "2026-02-12"
last_updated = "2026-02-12"
@@ -6,19 +6,18 @@ attachment = false
reasoning = true
temperature = true
tool_call = true
structured_output = true
open_weights = true
[interleaved]
field = "reasoning_content"
[cost]
input = 0.0
output = 0.0
input = 0.600
output = 2.600
[limit]
context = 202752
output = 131000
context = 204_800
output = 131_072
[modalities]
input = ["text"]
+24
View File
@@ -0,0 +1,24 @@
name = "GLM-5V-Turbo"
family = "glm"
release_date = "2026-04-02"
last_updated = "2026-04-02"
attachment = true
reasoning = true
temperature = true
tool_call = true
open_weights = false
[interleaved]
field = "reasoning_content"
[cost]
input = 0.720
output = 3.200
[limit]
context = 200_000
output = 131_072
[modalities]
input = ["text", "image", "video", "audio", "pdf"]
output = ["text"]
@@ -1,21 +1,20 @@
name = "GLM 4.6"
name = "glm-for-coding"
family = "glm"
release_date = "2025-09-30"
last_updated = "2025-09-30"
attachment = false
reasoning = false
reasoning = true
temperature = true
tool_call = true
knowledge = "2025-09"
open_weights = true
open_weights = false
[cost]
input = 0.6
output = 2.2
input = 0.086
output = 0.343
[limit]
context = 200_000
output = 200_000
output = 131_072
[modalities]
input = ["text"]
+3 -2
View File
@@ -6,6 +6,7 @@ attachment = true
reasoning = false
temperature = true
tool_call = true
structured_output = true
open_weights = false
knowledge = "2024-04"
@@ -14,9 +15,9 @@ input = 0.400
output = 1.600
[limit]
context = 1_000_000
context = 1_047_576
output = 32_768
[modalities]
input = ["text", "image"]
input = ["text", "image", "pdf"]
output = ["text"]
+2 -1
View File
@@ -6,6 +6,7 @@ attachment = true
reasoning = false
temperature = true
tool_call = true
structured_output = true
open_weights = false
knowledge = "2024-04"
@@ -14,7 +15,7 @@ input = 0.100
output = 0.400
[limit]
context = 1_000_000
context = 1_047_576
output = 32_768
[modalities]
+3 -2
View File
@@ -6,6 +6,7 @@ attachment = true
reasoning = false
temperature = true
tool_call = true
structured_output = true
open_weights = false
knowledge = "2024-04"
@@ -14,9 +15,9 @@ input = 2.000
output = 8.000
[limit]
context = 1_000_000
context = 1_047_576
output = 32_768
[modalities]
input = ["text", "image"]
input = ["text", "image", "pdf"]
output = ["text"]
+2 -1
View File
@@ -6,6 +6,7 @@ attachment = true
reasoning = false
temperature = true
tool_call = true
structured_output = true
open_weights = false
knowledge = "2023-09"
@@ -18,5 +19,5 @@ context = 128_000
output = 16_384
[modalities]
input = ["text", "image"]
input = ["text", "image", "pdf"]
output = ["text"]
+6 -3
View File
@@ -1,12 +1,14 @@
name = "gpt-5-mini"
family = "gpt-mini"
release_date = "2025-08-08"
last_updated = "2025-08-08"
attachment = true
reasoning = false
temperature = true
reasoning = true
temperature = false
tool_call = true
structured_output = true
open_weights = false
knowledge = "2024-10"
knowledge = "2024-05-30"
[cost]
input = 0.250
@@ -14,6 +16,7 @@ output = 2.000
[limit]
context = 400_000
input = 272_000
output = 128_000
[modalities]
+6 -3
View File
@@ -1,12 +1,14 @@
name = "gpt-5-pro"
family = "gpt-pro"
release_date = "2025-10-08"
last_updated = "2025-10-08"
attachment = true
reasoning = false
temperature = true
reasoning = true
temperature = false
tool_call = true
structured_output = true
open_weights = false
knowledge = "2024-10"
knowledge = "2024-09-30"
[cost]
input = 15.000
@@ -14,6 +16,7 @@ output = 120.000
[limit]
context = 400_000
input = 272_000
output = 272_000
[modalities]
@@ -1,12 +1,14 @@
name = "gpt-5.1-chat-latest"
family = "gpt-codex"
release_date = "2025-11-14"
last_updated = "2025-11-14"
attachment = true
reasoning = false
temperature = true
reasoning = true
temperature = false
tool_call = true
structured_output = true
open_weights = false
knowledge = "2024-10"
knowledge = "2024-09-30"
[cost]
input = 1.250
+6 -3
View File
@@ -1,12 +1,14 @@
name = "gpt-5.1"
family = "gpt"
release_date = "2025-11-14"
last_updated = "2025-11-14"
attachment = true
reasoning = false
temperature = true
reasoning = true
temperature = false
tool_call = true
structured_output = true
open_weights = false
knowledge = "2024-10"
knowledge = "2024-09-30"
[cost]
input = 1.250
@@ -14,6 +16,7 @@ output = 10.000
[limit]
context = 400_000
input = 272_000
output = 128_000
[modalities]
@@ -1,12 +1,14 @@
name = "gpt-5.2-chat-latest"
family = "gpt-codex"
release_date = "2025-12-12"
last_updated = "2025-12-12"
attachment = true
reasoning = false
temperature = true
reasoning = true
temperature = false
tool_call = true
structured_output = true
open_weights = false
knowledge = "2024-10"
knowledge = "2025-08-31"
[cost]
input = 1.750
+6 -3
View File
@@ -1,12 +1,14 @@
name = "gpt-5.2"
family = "gpt"
release_date = "2025-12-12"
last_updated = "2025-12-12"
attachment = true
reasoning = false
temperature = true
reasoning = true
temperature = false
tool_call = true
structured_output = true
open_weights = false
knowledge = "2024-10"
knowledge = "2025-08-31"
[cost]
input = 1.750
@@ -14,6 +16,7 @@ output = 14.000
[limit]
context = 400_000
input = 272_000
output = 128_000
[modalities]
@@ -0,0 +1,24 @@
name = "gpt-5.4-mini-2026-03-17"
family = "gpt-mini"
release_date = "2026-03-19"
last_updated = "2026-03-19"
attachment = true
reasoning = true
temperature = false
tool_call = true
structured_output = true
open_weights = false
knowledge = "2025-08-31"
[cost]
input = 0.750
output = 4.500
[limit]
context = 400_000
input = 272_000
output = 128_000
[modalities]
input = ["text", "image"]
output = ["text"]
+24
View File
@@ -0,0 +1,24 @@
name = "gpt-5.4-mini"
family = "gpt-mini"
release_date = "2026-03-19"
last_updated = "2026-03-19"
attachment = true
reasoning = true
temperature = false
tool_call = true
structured_output = true
open_weights = false
knowledge = "2025-08-31"
[cost]
input = 0.750
output = 4.500
[limit]
context = 400_000
input = 272_000
output = 128_000
[modalities]
input = ["text", "image"]
output = ["text"]
@@ -0,0 +1,24 @@
name = "gpt-5.4-nano-2026-03-17"
family = "gpt-nano"
release_date = "2026-03-19"
last_updated = "2026-03-19"
attachment = true
reasoning = true
temperature = false
tool_call = true
structured_output = true
open_weights = false
knowledge = "2025-08-31"
[cost]
input = 0.200
output = 1.250
[limit]
context = 400_000
input = 272_000
output = 128_000
[modalities]
input = ["text", "image"]
output = ["text"]
+24
View File
@@ -0,0 +1,24 @@
name = "gpt-5.4-nano"
family = "gpt-nano"
release_date = "2026-03-19"
last_updated = "2026-03-19"
attachment = true
reasoning = true
temperature = false
tool_call = true
structured_output = true
open_weights = false
knowledge = "2025-08-31"
[cost]
input = 0.200
output = 1.250
[limit]
context = 400_000
input = 272_000
output = 128_000
[modalities]
input = ["text", "image"]
output = ["text"]
+31
View File
@@ -0,0 +1,31 @@
name = "gpt-5.4-pro"
family = "gpt-pro"
release_date = "2026-03-05"
last_updated = "2026-03-05"
attachment = true
reasoning = true
temperature = false
tool_call = true
structured_output = false
open_weights = false
knowledge = "2025-08-31"
[cost]
input = 30.000
output = 180.000
cache_read = 0
cache_write = 0
[[cost.tiers]]
tier = { size = 272_000 }
input = 60.000
output = 270.000
[limit]
context = 1_050_000
input = 922_000
output = 128_000
[modalities]
input = ["text", "image"]
output = ["text"]
+31
View File
@@ -0,0 +1,31 @@
name = "gpt-5.4"
family = "gpt"
release_date = "2026-03-05"
last_updated = "2026-03-05"
attachment = true
reasoning = true
temperature = false
tool_call = true
structured_output = true
open_weights = false
knowledge = "2025-08-31"
[cost]
input = 2.500
output = 15.000
cache_read = 0.250
cache_write = 0
[[cost.tiers]]
tier = { size = 272_000 }
input = 5.000
output = 22.500
[limit]
context = 1_050_000
input = 922_000
output = 128_000
[modalities]
input = ["text", "image", "pdf"]
output = ["text"]
+6 -3
View File
@@ -1,12 +1,14 @@
name = "gpt-5"
family = "gpt"
release_date = "2025-08-08"
last_updated = "2025-08-08"
attachment = true
reasoning = false
temperature = true
reasoning = true
temperature = false
tool_call = true
structured_output = true
open_weights = false
knowledge = "2024-10"
knowledge = "2024-09-30"
[cost]
input = 1.250
@@ -14,6 +16,7 @@ output = 10.000
[limit]
context = 400_000
input = 272_000
output = 128_000
[modalities]
@@ -1,18 +1,15 @@
name = "Grok 4 Fast (Non-Reasoning)"
family = "grok"
release_date = "2025-09-19"
last_updated = "2025-09-19"
name = "grok-4.20-beta-0309-non-reasoning"
release_date = "2026-03-16"
last_updated = "2026-03-16"
attachment = true
reasoning = false
temperature = true
knowledge = "2025-07"
tool_call = true
open_weights = false
[cost]
input = 0.20
output = 0.50
cache_read = 0.05
input = 2.000
output = 6.000
[limit]
context = 2_000_000
@@ -1,6 +1,6 @@
name = "xAI: Grok 4.1 Fast"
release_date = "2025-11-19"
last_updated = "2025-11-19"
name = "grok-4.20-beta-0309-reasoning"
release_date = "2026-03-16"
last_updated = "2026-03-16"
attachment = true
reasoning = true
temperature = true
@@ -8,13 +8,12 @@ tool_call = true
open_weights = false
[cost]
input = 0.2
output = 0.5
cache_read = 0.05
input = 2.000
output = 6.000
[limit]
context = 2000000
output = 30000
context = 2_000_000
output = 30_000
[modalities]
input = ["text", "image"]
@@ -0,0 +1,20 @@
name = "grok-4.20-multi-agent-beta-0309"
release_date = "2026-03-16"
last_updated = "2026-03-16"
attachment = true
reasoning = true
temperature = true
tool_call = true
open_weights = false
[cost]
input = 2.000
output = 6.000
[limit]
context = 2_000_000
output = 30_000
[modalities]
input = ["text", "image"]
output = ["text"]
+1 -1
View File
@@ -6,7 +6,7 @@ attachment = true
reasoning = true
temperature = true
tool_call = true
knowledge = "2025-05"
knowledge = "2025-05-31"
open_weights = false
[cost]
@@ -6,7 +6,7 @@ attachment = true
reasoning = true
temperature = true
tool_call = true
knowledge = "2025-08"
knowledge = "2025-08-31"
open_weights = false
[cost]
@@ -12,6 +12,8 @@ open_weights = false
[cost]
input = 0.25
output = 1.50
cache_read = 0.025
cache_write = 1.00
[limit]
context = 1_048_576
+6
View File
@@ -0,0 +1,6 @@
<svg viewBox="0 0 220 32" xmlns="http://www.w3.org/2000/svg">
<text x="0" y="24" fill="currentColor">
<tspan font-size="24">Abliteration</tspan>
<tspan dx="8" dy="2" font-size="12" fill-opacity="0.6">.ai</tspan>
</text>
</svg>

After

Width:  |  Height:  |  Size: 239 B

@@ -0,0 +1,22 @@
name = "Abliterated Model"
release_date = "2026-01-06"
last_updated = "2026-01-06"
attachment = true
reasoning = false
tool_call = true
structured_output = false
temperature = true
open_weights = true
[cost]
input = 3.00
output = 3.00
[limit]
context = 150_000
input = 150_000
output = 8_192
[modalities]
input = ["text", "image"]
output = ["text"]
+5
View File
@@ -0,0 +1,5 @@
name = "abliteration.ai"
env = ["ABLIT_KEY"]
npm = "@ai-sdk/openai-compatible"
api = "https://api.abliteration.ai/v1"
doc = "https://docs.abliteration.ai/models"
@@ -0,0 +1,27 @@
name = "DeepSeek V4 Flash (Alibaba Cloud)"
family = "deepseek-flash"
release_date = "2026-04-24"
last_updated = "2026-04-24"
attachment = false
reasoning = true
temperature = true
tool_call = true
structured_output = true
knowledge = "2025-05"
open_weights = true
[interleaved]
field = "reasoning_content"
[cost]
input = 0.14
output = 0.28
cache_read = 0.028
[limit]
context = 1_000_000
output = 384_000
[modalities]
input = ["text"]
output = ["text"]
@@ -0,0 +1,27 @@
name = "DeepSeek V4 Pro (Alibaba Cloud)"
family = "deepseek-thinking"
release_date = "2026-04-24"
last_updated = "2026-04-24"
attachment = false
reasoning = true
temperature = true
tool_call = true
structured_output = true
knowledge = "2025-05"
open_weights = true
[interleaved]
field = "reasoning_content"
[cost]
input = 1.69
output = 3.38
cache_read = 0.13
[limit]
context = 1_000_000
output = 384_000
[modalities]
input = ["text"]
output = ["text"]
@@ -0,0 +1,27 @@
name = "GLM-5.1 (Alibaba Cloud)"
family = "glm"
release_date = "2026-03-27"
last_updated = "2026-03-27"
attachment = false
reasoning = true
temperature = true
tool_call = true
structured_output = true
open_weights = true
[interleaved]
field = "reasoning_content"
[cost]
input = 0.84
output = 3.38
cache_read = 0.169
cache_write = 1.05625
[limit]
context = 200_000
output = 128_000
[modalities]
input = ["text"]
output = ["text"]
@@ -1,28 +1,33 @@
name = "Claude Opus 4.6 Think"
name = "Claude Opus 4.6 Thinking"
family = "claude-opus"
release_date = "2026-02-05"
last_updated = "2026-02-05"
last_updated = "2026-03-13"
attachment = true
reasoning = true
temperature = true
tool_call = true
knowledge = "2025-05"
structured_output = true
knowledge = "2025-05-31"
open_weights = false
[cost]
input = 5.00
output = 25.00
cache_read = 0.30
cache_write = 3.75
[interleaved]
field = "reasoning_content"
[cost.context_over_200k]
input = 6.00
output = 22.00
cache_read = 0.60
cache_write = 7.50
[cost]
input = 5
output = 25
cache_read = 0.5
cache_write = 6.25
[[cost.tiers]]
tier = { size = 200_000 }
input = 10
output = 37.5
cache_read = 1.0
cache_write = 12.5
[limit]
context = 200_000
context = 1_000_000
output = 128_000
[modalities]
+17 -13
View File
@@ -1,28 +1,32 @@
name = "Claude Opus 4.6"
family = "claude-opus"
release_date = "2026-02-05"
last_updated = "2026-02-05"
last_updated = "2026-03-13"
attachment = true
reasoning = true
temperature = true
tool_call = true
knowledge = "2025-05"
structured_output = true
knowledge = "2025-05-31"
open_weights = false
[cost]
input = 5.00
output = 25.00
cache_read = 0.30
cache_write = 3.75
interleaved = true
[cost.context_over_200k]
input = 6.00
output = 22.00
cache_read = 0.60
cache_write = 7.50
[cost]
input = 5
output = 25
cache_read = 0.5
cache_write = 6.25
[[cost.tiers]]
tier = { size = 200_000 }
input = 10
output = 37.5
cache_read = 1.0
cache_write = 12.5
[limit]
context = 200_000
context = 1_000_000
output = 128_000
[modalities]
@@ -0,0 +1,35 @@
name = "Claude Opus 4.7 Thinking"
family = "claude-opus"
release_date = "2026-04-16"
last_updated = "2026-04-16"
attachment = true
reasoning = true
temperature = false
tool_call = true
structured_output = true
knowledge = "2026-01-31"
open_weights = false
[interleaved]
field = "reasoning_content"
[cost]
input = 5
output = 25
cache_read = 0.5
cache_write = 6.25
[[cost.tiers]]
tier = { size = 200_000 }
input = 10
output = 37.5
cache_read = 1.0
cache_write = 12.5
[limit]
context = 1_000_000
output = 128_000
[modalities]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -0,0 +1,34 @@
name = "Claude Opus 4.7"
family = "claude-opus"
release_date = "2026-04-16"
last_updated = "2026-04-16"
attachment = true
reasoning = true
temperature = false
tool_call = true
structured_output = true
knowledge = "2026-01-31"
open_weights = false
interleaved = true
[cost]
input = 5
output = 25
cache_read = 0.5
cache_write = 6.25
[[cost.tiers]]
tier = { size = 200_000 }
input = 10
output = 37.5
cache_read = 1.0
cache_write = 12.5
[limit]
context = 1_000_000
output = 128_000
[modalities]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -1,30 +1,35 @@
name = "Claude Sonnet 4.6 Think"
name = "Claude Sonnet 4.6 Thinking"
family = "claude-sonnet"
release_date = "2026-02-17"
last_updated = "2026-02-17"
last_updated = "2026-03-13"
attachment = true
reasoning = true
temperature = true
tool_call = true
knowledge = "2025-08"
structured_output = true
knowledge = "2025-08-31"
open_weights = false
[interleaved]
field = "reasoning_content"
[cost]
input = 3.00
output = 15.00
cache_read = 0.30
cache_write = 3.75
[cost.context_over_200k]
[[cost.tiers]]
tier = { size = 200_000 }
input = 6.00
output = 22.50
cache_read = 0.60
cache_write = 7.50
[limit]
context = 200_000
context = 1_000_000
output = 64_000
[modalities]
input = ["text", "image", "pdf"]
output = ["text"]
output = ["text"]
@@ -1,28 +1,32 @@
name = "Claude Sonnet 4.6"
family = "claude-sonnet"
release_date = "2026-02-17"
last_updated = "2026-02-17"
last_updated = "2026-03-13"
attachment = true
reasoning = true
temperature = true
tool_call = true
knowledge = "2025-08"
structured_output = true
knowledge = "2025-08-31"
open_weights = false
interleaved = true
[cost]
input = 3.00
output = 15.00
cache_read = 0.30
cache_write = 3.75
[cost.context_over_200k]
[[cost.tiers]]
tier = { size = 200_000 }
input = 6.00
output = 22.50
cache_read = 0.60
cache_write = 7.50
[limit]
context = 200_000
context = 1_000_000
output = 64_000
[modalities]

Some files were not shown because too many files have changed in this diff Show More