Compare commits

...

220 Commits

Author SHA1 Message Date
Aiden Cline cda892e0d7 sync github copilot limits 2026-03-19 21:54:42 -05:00
Aiden Cline 098ff4f5bf Merge pull request #1227 from Verizane/dev
add OpenRouter models for gpt-5.4 mini and gpt-5.4 nano
2026-03-19 21:23:15 -05:00
Aiden Cline 0f70b8959f Merge pull request #1234 from mchenco/dev
Add Workers AI models: kimi-k2.5, nemotron-3-120b-a12b, glm-4.7-flash
2026-03-19 15:11:19 -05:00
mchen b8e6d58e5b add workers-ai models: kimi-k2.5, nemotron-3-120b-a12b, glm-4.7-flash 2026-03-19 14:59:17 -04:00
Roman Koslowski a855001a7e apply changes from review 2026-03-19 17:20:26 +01:00
Aiden Cline ac760b2268 Merge pull request #1230 from SamizuHM/feature/zhipuai-coding-plan-add-glm-5-turbo
zhipuai-coding-plan: Add glm-5-turbo.toml and replace symlink
2026-03-19 10:42:43 -05:00
Aiden Cline d4a5ea7ae7 Merge pull request #1226 from spiffytech/dev
Add Ollama Cloud support for Minimax M2.7
2026-03-19 10:41:47 -05:00
Aiden Cline 434ed89ba2 Merge pull request #1228 from dpuyosa/minimax_m2_7
Venice: Add MiniMax M2.7 and update DeepSeek V3.2 pricing
2026-03-19 10:41:16 -05:00
Aiden Cline 6d7719a62a Merge pull request #1229 from 0b1000/dev
Xiaomi: Add MiMo-V2-Pro and MiMo-V2-Omni
2026-03-19 10:41:06 -05:00
Aiden Cline 93637039ef Merge pull request #1231 from ariane-emory/feat/feat/add-xiaomi-mimo-v2-pro-and-omni
feat: add the Xiaomi MiMo V2 Pro and Xiaomi MiMo V2 Omni models to the OpenRouter provide
2026-03-19 10:40:44 -05:00
Ariane Emory 9c95f796c0 Merge remote-tracking branch 'upstream/dev' into feat/feat/add-xiaomi-mimo-v2-pro 2026-03-19 11:22:43 -04:00
Ariane Emory e8650b6073 feat: add xiaomi mimo-v2-pro and mimo-v2-omni models to openrouter 2026-03-19 11:18:46 -04:00
SamizuHM 23c2be6ff7 feat(zhipuai-coding-plan): add glm-5-turbo.toml and replace glm-5-turbo with symlink 2026-03-19 18:09:17 +08:00
Frank 913a63dbe6 update zen models 2026-03-19 00:33:45 -04:00
0b1000 503087e99b Merge branch 'anomalyco:dev' into dev 2026-03-19 12:28:38 +08:00
0b1000 48150f09d3 Xiaomi: Add MiMo-V2-Pro and MiMo-V2-Omni 2026-03-19 12:27:00 +08:00
Aiden Cline 5fef681657 Disable tool_call in grok model configuration 2026-03-18 23:09:30 -05:00
Frank 123054ae0c update zen models 2026-03-18 20:45:44 -04:00
Frank 03060d154b update zen models 2026-03-18 20:37:47 -04:00
dpuyosa 5c9b8108e0 Update minimax-m27.toml 2026-03-19 01:02:24 +01:00
dpuyosa c8084681f9 [venice] Add MiniMax M2.7 and update DeepSeek V3.2 pricing
- Add MiniMax M2.7 model with reasoning and tool_call support
 - Update DeepSeek V3.2 pricing (input: $0.33, output: $0.48, cache: $0.16)
2026-03-19 00:58:50 +01:00
Roman Koslowski 352ab4ae1b add gpt-5.4 mini and gpt-5.4 nano 2026-03-18 22:16:55 +01:00
spiffytech cf0b416b15 Added Ollama Cloud support for Minimax M2.7 2026-03-18 16:15:07 -04:00
Aiden Cline 38339a2a90 Merge pull request #1224 from APonce911/minimax-m2.7-openrouter
add MiniMax M2.7 to OpenRouter
2026-03-18 14:10:13 -05:00
Aiden Cline ff9040bf52 Update minimax-m2.7.toml 2026-03-18 14:09:26 -05:00
Aiden Cline 3039804af4 Delete providers/opencode/models/minimax-m2.7.toml 2026-03-18 14:08:55 -05:00
Frank 7a4ad7bec8 update go models 2026-03-18 14:40:25 -04:00
airton 721cc122bc add MiniMax M2.7 to OpenRouter and OpenCode 2026-03-18 18:57:09 +01:00
Aiden Cline 0527f019af Merge pull request #1221 from sergical/fix/bedrock-claude-4-6-context-window-and-pricing
fix(amazon-bedrock): set Claude Sonnet 4.6 and Opus 4.6 context window to 1M
2026-03-18 12:17:57 -05:00
Aiden Cline c89371de50 Merge pull request #1223 from sylviezhang37/update-vercel-models-20260318-1659
Update Vercel models
2026-03-18 12:17:22 -05:00
Sylvie Zhang 6d6d4220d8 Enable open_weights in minimax-m2.7.toml 2026-03-18 10:12:37 -07:00
Sylvie Zhang 8b984eeec1 Enable open_weights in minimax-m2.7-highspeed model 2026-03-18 10:12:21 -07:00
github-actions[bot] 586027c8f1 chore(vercel): update Vercel model definitions
Auto-generated by weekly workflow from Vercel AI Gateway API.

Co-Authored-By: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-03-18 16:59:24 +00:00
Sergiy Dybskiy 343b5f87ef fix(amazon-bedrock): set Claude Sonnet 4.6 and Opus 4.6 context window to 1M
Both models support a 1M token context window natively on Bedrock via the
Converse API with no beta headers required. Verified empirically via the
AWS CLI (bedrock-runtime converse): 950K tokens succeeds, >1M returns
'prompt is too long: N tokens > 1000000 maximum'.

The AWS Bedrock pricing page confirms long context pricing for these two
models is identical to standard pricing (no surcharge), so the
[cost.context_over_200k] section is removed as it was incorrect.
2026-03-18 12:20:38 -04:00
Aiden Cline 955b773ee5 Merge pull request #1218 from pomidornijfrukt/azure/5.4-mini-nano
Add GPT-5.4 Mini and Nano models for Azure providers
2026-03-18 10:31:13 -05:00
Aiden Cline 98559071f0 Merge pull request #1217 from cgilly2fast/dev
chore(firmware): update base url and docs url
2026-03-18 10:30:44 -05:00
eCube-cachy 0660308816 add: GPT-5.4 Mini and Nano model configurations for Azure providers 2026-03-18 15:17:52 +02:00
Jack 380f9dd8eb Merge pull request #1216 from no1wudi/dev
Add MiniMax M2.7 and M2.7-highspeed models to 4 official providers
2026-03-18 16:29:59 +08:00
Jack 1cfdab1b18 update MiniMax-M2.7 cache_read to 0.06 2026-03-18 16:27:53 +08:00
Colby Gilbert 75a981f957 chore(firmware): update base url and docs url 2026-03-18 00:41:05 -07:00
Huang Qi 7fadbcadc8 Add MiniMax M2.7 and M2.7-highspeed models to 4 official providers 2026-03-18 15:21:06 +08:00
Frank 38f9092292 update zen models 2026-03-18 02:30:18 -04:00
Aiden Cline 92149b9eaa rm nonexistant github model 2026-03-17 21:41:51 -05:00
Aiden Cline b614f0e69c Merge pull request #1214 from luisrudge/dev
Add GPT-5.4 mini and nano to GitHub Copilot provider
2026-03-17 20:13:46 -05:00
Luís Rudge 67d6dac5c5 Add GPT-5.4 mini and nano to GitHub Copilot provider 2026-03-17 18:44:38 -06:00
Aiden Cline 7d3cc61a48 Merge pull request #1207 from PedroACosta/feat/add-dinference-provider
feat(providers): add dinference provider
2026-03-17 14:51:31 -05:00
Aiden Cline f02ea6c4d2 Merge pull request #1115 from skywalker512/feat/add-tencent-coding-plan
feat: add Tencent Coding Plan provider
2026-03-17 14:51:19 -05:00
Aiden Cline 0cb50eeece Merge pull request #1208 from scwgoire/march-update
Scaleway 26-03 model updates
2026-03-17 14:48:12 -05:00
Aiden Cline 878311d2e0 Merge pull request #1210 from dm-cohere/dm/fix-update-cohere-model-capabilities
fix(models): update cohere model capabilities
2026-03-17 14:32:27 -05:00
Aiden Cline a0e89f65d6 Merge pull request #1206 from 0b1000/dev
Rename minimax-m2.5.toml to MiniMax-M2.5.toml
2026-03-17 14:32:19 -05:00
Aiden Cline 74099b7c9c Merge pull request #1213 from smrdotgg/add-openai-gpt-5-4-mini-and-nano
Add OpenAI GPT-5.4 mini and nano
2026-03-17 14:31:24 -05:00
Aiden Cline ec522435c3 Merge pull request #1211 from sylviezhang37/update-vercel-models-20260317-1807
Update Vercel models
2026-03-17 14:30:24 -05:00
smr d839cd37d4 Add OpenAI GPT-5.4 mini and nano
Capture the newly released mini and nano model metadata so models.dev reflects OpenAI's latest GPT-5.4 lineup with current pricing, limits, and knowledge cutoff.
2026-03-17 22:09:13 +03:00
github-actions[bot] ecb6ef7f93 chore(vercel): update Vercel model definitions
Auto-generated by weekly workflow from Vercel AI Gateway API.

Co-Authored-By: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-03-17 18:07:06 +00:00
Deirdre Meehan 8b4d341054 fix: cohere models on non-cohere providers 2026-03-17 16:51:24 +00:00
Deirdre Meehan 62f4a28308 fix: cohere provider models 2026-03-17 16:44:55 +00:00
Pedro 2fb8ef0dc8 feat(providers): add dinference provider 2026-03-17 14:13:36 +01:00
Gregoire de Turckheim 96968e2bf8 feat: Scaleway 26-03 model updates 2026-03-17 12:15:52 +01:00
0b1000 ee9d7879ce Rename minimax-m2.5.toml to MiniMax-M2.5.toml 2026-03-17 14:50:35 +08:00
Frank 71283512a6 update zen models 2026-03-17 02:21:13 -04:00
Frank cd4afd7e7c update zen models 2026-03-17 02:19:17 -04:00
Aiden Cline 1239d0190b Merge pull request #1204 from cyberofficial/vultr
VULTR: Updated Vultr model pricing to reflect current serverless inference rates
2026-03-16 16:10:39 -05:00
Aiden Cline 491bf6ccba Merge pull request #1202 from RaviTharuma/fix/chutes-pricing-update-2026-03
fix(chutes): update pricing and limits from live API
2026-03-16 16:10:25 -05:00
Cyber Official c993d0c121 Updated Vultr model pricing to reflect current serverless inference rates
Updated Vultr model pricing to reflect current serverless inference rates

This commit updates the cost configuration for all Vultr models to align with their latest pricing tiers:

**Cost Reductions:**
- DeepSeek-R1-Distill-Qwen-32B: Input $0.55→$0.30, Output $2.75→$0.30 (73% reduction)
- NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4: Input $0.55→$0.20, Output $2.75→$0.80 (64% input, 71% output reduction)
- Qwen2.5-Coder-32B-Instruct: Input $0.55→$0.20, Output $2.75→$0.60 (64% input, 78% output reduction)
- gpt-oss-120b: Input $0.55→$0.15, Output $2.75→$0.60 (73% input, 78% output reduction)
- MiniMax-M2.5: Input $0.55→$0.30, Output $2.75→$1.20 (45% input, 56% output reduction)

**Cost Adjustments:**
- DeepSeek-R1-Distill-Llama-70B: Input $0.55→$2.00, Output $2.75→$2.00 (significant increase)
- DeepSeek-V3.2: Output $2.75→$1.65 (40% reduction)
- Llama-3.1-Nemotron-Ultra-253B-v1: Output $2.75→$1.80 (35% reduction)
- GLM-5-FP8: Input $0.55→$0.85, Output $2.75→$3.10 (55% input, 13% output increase)
2026-03-16 13:58:28 -04:00
Aiden Cline e55c39a83d Merge pull request #1141 from sk0x0y/feature/nanogpt-confirmed-suffix2-fixes
fix(nano-gpt): rename confirmed 2-suffix model ids
2026-03-16 10:57:44 -05:00
Aiden Cline 6dea000e25 Merge pull request #1148 from sk0x0y/feature/nanogpt-bundled-confirmed-suffix2-fixes
fix(nano-gpt): rename bundled confirmed 2-suffix model ids
2026-03-16 10:57:30 -05:00
Aiden Cline 3f7a757b3f Merge pull request #1142 from sk0x0y/feature/nanogpt-more-confirmed-suffix2-fixes
fix(nano-gpt): rename more confirmed 2-suffix model ids
2026-03-16 10:56:36 -05:00
Aiden Cline c693fd71e2 Merge pull request #1194 from cyberofficial/vultr
Update Vultr model list with 10 new models and updated pricing
2026-03-16 10:55:19 -05:00
Aiden Cline 54e04e288a Merge pull request #1198 from amritbanerjee/add-glm-5-turbo
Add GLM-5-Turbo model support
2026-03-16 10:47:08 -05:00
Aiden Cline 462a179eee Merge pull request #1203 from jerome-benoit/fix/sap-ai-core-model-specs
fix(sap-ai-core): align model specs with official sources
2026-03-16 10:46:40 -05:00
Aiden Cline 95db59034d Merge pull request #1201 from dpuyosa/venice-new-models
Venice: Add new provider models
2026-03-16 10:45:59 -05:00
Aiden Cline 74dcc74e32 Merge pull request #1200 from dpuyosa/venice/pricing-update
Venice: Update model pricing for 7 models
2026-03-16 10:45:47 -05:00
Jérôme Benoit 57975f5f25 fix(sap-ai-core): align model specs with official sources 2026-03-16 13:59:06 +01:00
Ravi Tharuma ad7b063747 fix(chutes): update pricing and limits from live API
Synced 6 Chutes model definitions against the live API at
https://llm.chutes.ai/v1/models (queried 2026-03-16).

Models updated:
- deepseek-ai/DeepSeek-V3.2-TEE: cost 0.25/0.38→0.28/0.42, cache 0.125→0.14, context 163840→131072
- zai-org/GLM-5-TEE: cost 0.75/2.5→0.95/3.15, added cache_read 0.475
- zai-org/GLM-4.6-TEE: cost 0.35/1.5→0.4/1.7, added cache_read 0.2
- zai-org/GLM-4.6V: added cache_read 0.15
- MiniMaxAI/MiniMax-M2.5-TEE: cost 0.15/0.6→0.3/1.1, added cache_read 0.15
- Qwen/Qwen3.5-397B-A17B-TEE: cost 0.3/1.2→0.39/2.34, cache 0.15→0.195
2026-03-16 11:42:29 +01:00
dpuyosa f76e9f0551 [venice] Add new provider models
- Add mistral-small-3.2-24b-instruct, qwen3-5-9b, venice-uncensored-role-play, zai-org-glm-4.6
2026-03-16 09:37:41 +01:00
dpuyosa d70a49b36f [venice] Update model pricing for 7 models
- Remove context_over_200k pricing from Claude models
- Update Grok cache_read pricing from 0.5 to 0.25
- Update Kimi, MiniMax input/output pricing
2026-03-16 09:05:37 +01:00
amrit 3487135f9f Add GLM-5-Turbo model support 2026-03-16 12:14:50 +11:00
Aiden Cline 458a66c766 Merge pull request #1197 from kesku/update-perplexity-agent-models
Update Perplexity Agent API models
2026-03-15 10:59:23 -05:00
Frank d3a84dc7ec update zen models 2026-03-15 10:59:52 -04:00
Kesku ae61b25583 update perplexity-agent: add gpt-5.4 & nemotron, remove gemini-3-pro 2026-03-15 03:46:50 +00:00
Aiden Cline 74be576eda Merge pull request #1178 from Sewer56/change-synthetic-endpoint
Add OpenAI and Anthropic compatible endpoints
2026-03-14 20:55:30 -05:00
Aiden Cline 164df2cda0 Merge pull request #1191 from Alcatraz-Zhang/update/kilo-models
Sync Kilo model definitions with latest gateway catalog
2026-03-14 20:54:45 -05:00
Cyber Official 2cd7908369 Update Vultr model list with 10 new models and updated pricing
- Updated pricing to $0.55/M input tokens, $2.75/M output tokens
- Updated context limits to safe floor values from official testing
- Added accurate output token limits from official model documentation
- Added 5 new models: MiniMax M2.5, DeepSeek V3.2, GLM-5 FP8, Llama 3.1 Nemotron Ultra 253B, NVIDIA Nemotron 3 Super 120B A12B NVFP4
- Updated existing models: DeepSeek R1 Distill variants, GPT OSS 120B, Kimi K2.5, Qwen2.5 Coder 32B

Model specifications:
- MiniMax M2.5: 196K context, 4,096 output
- Qwen2.5-Coder-32B: 15K context, 256 output (notable low default)
- DeepSeek R1 Distill Llama 70B: 130K context, 4,096 output
- DeepSeek R1 Distill Qwen 32B: 130K context, 4,096 output
- DeepSeek V3.2: 163K context, 4,096 output
- Kimi K2.5: 261K context, 32,768 output (high output limit)
- GPT OSS 120B: 130K context, 8,192 output
- GLM-5 FP8: 202K context, 131,072 output (exceptionally high)
- Llama 3.1 Nemotron Ultra 253B: 32K context, 4,096 output
- NVIDIA Nemotron 3 Super 120B A12B NVFP4: 260K context, 8,192 output

All models set to text-only (no vision support) as confirmed.
2026-03-14 19:47:01 -04:00
Alcatraz-Zhang cc667340f5 Sync Kilo model definitions with latest gateway catalog
Refresh the Kilo provider catalog so models.dev matches the current gateway inventory, pricing, and availability.
2026-03-15 04:35:38 +08:00
Sewer56 f2cfc1435d Changed: Synthetic to use newer openai endpoint 2026-03-14 17:09:44 +00:00
Aiden Cline 35bb8cca47 Merge pull request #1172 from bigfluffycookie/add-deepinfra-llama-models
Add deepinfra llama models
2026-03-14 10:55:13 -05:00
Aiden Cline 3468a410e1 Merge pull request #1177 from ar27111994/dev
Add Grok 4.1 Fast configurations for reasoning and non-reasoning
2026-03-14 10:54:57 -05:00
Aiden Cline b1b5e3c5cd Merge pull request #1174 from dacbd/patch-1
fix(wandb): fix k2.5 settings
2026-03-14 10:54:35 -05:00
Aiden Cline 97f03ec672 Merge pull request #1175 from dacbd/patch-2
chore(docs): add note for manual testing with opencode
2026-03-14 10:54:22 -05:00
BigFluffyCookie 9b516924aa Add limit output for llama models 2026-03-14 11:49:57 +01:00
Ahmed Rehan 929a39600b feat(models): add Grok 4.1 Fast (Reasoning and Non-Reasoning) configurations 2026-03-14 14:27:24 +05:00
Daniel Barnes a87d8bb8cc chore(docs): add note for manual testing with opencode 2026-03-14 13:42:57 +09:00
Daniel Barnes 574139eb49 fix(wandb): fix k2.5 settings 2026-03-14 13:07:16 +09:00
Aiden Cline 1e3bc38b31 Merge pull request #1137 from mcowger/mcowger/correct-gemini-flash-lite-pricing
Fix incorrect pricing for gemini-3.1-flash-lite-preview
2026-03-13 18:41:41 -05:00
Aiden Cline 8916fe9874 Merge pull request #1171 from stephenkuhn214/dev
fix(amazon-bedrock): Remove deprecated and add missing models
2026-03-13 18:26:25 -05:00
BigFluffyCookie 5d956b41a6 Rename llama models to remove "Meta" prefix 2026-03-13 23:15:27 +01:00
BigFluffyCookie 42a7a14f69 Add Meta Llama models to DeepInfra provider 2026-03-13 22:53:39 +01:00
Stephen Kuhn f24ee000d7 fix(amazon-bedrock): update and add models
- Remove 19 deprecated/EOL models
- Add 7 new models: DeepSeek V3.2, Llama 3.1 405B, Magistral Small 1.2, Ministral 3 3B, Mistral Large 3, Pixtral Large, NVIDIA Nemotron Nano 3 30B
- Fix Devstral 2 123B: correct name, family, and open_weights
- Set accurate Bedrock launch dates for all new models
2026-03-13 16:02:04 -04:00
Aiden Cline 7196b1fb2c Merge pull request #1170 from anomalyco/revert-1166-fix/update-gpt53-codex-spark-preview
Revert "fix(openai): rename gpt-5.3-codex-spark to gpt-5.3-codex-spark-preview"
2026-03-13 14:31:33 -05:00
Aiden Cline f6c0d5a29d Revert "fix(openai): rename gpt-5.3-codex-spark to gpt-5.3-codex-spark-preview" 2026-03-13 14:30:58 -05:00
Aiden Cline ee63449aa5 sonnet 4.6 and opus 4.6 1M context 2026-03-13 14:27:55 -05:00
Aiden Cline 92aa44ec00 Merge pull request #1166 from rluisr/fix/update-gpt53-codex-spark-preview
fix(openai): rename gpt-5.3-codex-spark to gpt-5.3-codex-spark-preview
2026-03-13 14:18:41 -05:00
Aiden Cline 477284535c Rename model from 'GPT-5.3 Codex Spark Preview' to 'GPT-5.3 Codex Spark' 2026-03-13 14:17:44 -05:00
Aiden Cline 304233bdda Merge pull request #1169 from mdrxy/mdrxy/anthropic-token-limits
Update Claude 4.6 context/pricing
2026-03-13 14:13:40 -05:00
Aiden Cline 25d782ee2c Reduce context limit from 1,000,000 to 200,000 2026-03-13 14:13:30 -05:00
Aiden Cline 0f63393d51 Update context limit in claude-opus-4-6.toml 2026-03-13 14:12:56 -05:00
rluisr e780eefce2 fix(openai): rename gpt-5.3-codex-spark to gpt-5.3-codex-spark-preview
The OpenAI API expects model ID 'gpt-5.3-codex-spark-preview', not
'gpt-5.3-codex-spark'. Rename model files in both openai and opencode
providers so the generated model ID matches the actual API.
2026-03-14 03:59:03 +09:00
Aiden Cline a79585fa83 Merge pull request #1163 from micuintus/feature/Kimi2.5-fast
feat(nebius): add Kimi-K2.5-fast model
2026-03-13 13:14:38 -05:00
Aiden Cline 00801f74f2 Merge pull request #1164 from butyess/dev
Openrouter models: gemini 3.1 flash lite preview, grok 4.20 beta models.
2026-03-13 13:14:22 -05:00
Aiden Cline 185f6731ee Merge pull request #1162 from dpuyosa/feature/venice-grok-4-20-beta
Venice: Add Grok 4.20 Beta models
2026-03-13 12:53:28 -05:00
Aiden Cline d291b0575c Merge pull request #1167 from sylviezhang37/update-vercel-models-20260313-1639
Update Vercel models
2026-03-13 12:53:11 -05:00
Mason Daugherty 382d9f3e7d Update Claude 4.6 context/pricing 2026-03-13 13:53:04 -04:00
Aiden Cline e64f5fe963 Merge pull request #1168 from mdrxy/mdrxy/update-baseten
Update Baseten models
2026-03-13 12:51:56 -05:00
Mason Daugherty ea57ddfe7e Update Baseten models 2026-03-13 13:48:41 -04:00
github-actions[bot] 29463d7fa8 chore(vercel): update Vercel model definitions
Auto-generated by weekly workflow from Vercel AI Gateway API.

Co-Authored-By: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-03-13 16:39:30 +00:00
Jack bcc8db49ee Merge pull request #1165 from anomalyco/chore/openrouter-alpha-reasoning-details-20260313
feat(openrouter): add interleaved reasoning details for alpha models
2026-03-13 22:25:25 +08:00
Jack c8521d70f3 feat(openrouter): add interleaved reasoning details for alpha models 2026-03-13 22:20:54 +08:00
Federico Masi 490cd249e4 Openrouter models: gemini 3.1 flash lite preview, grok 4.20 beta models. 2026-03-13 15:11:12 +01:00
Michael Voigt dbc636f5f3 feat(nebius): add Kimi-K2.5-fast model 2026-03-13 12:35:13 +01:00
Michael Voigt 9a32f671a1 fix(nebius): lowercase model ID for Nemotron-3-Super-120B-A12B
The filename must match the API casing (lowercase) to avoid 'model does not exist' errors.
2026-03-13 12:35:08 +01:00
dpuyosa 856d925eda [venice] Add Grok 4.20 Beta models
- Add Grok 4.20 Beta model configuration (2M context, 128K output)
- Add Grok 4.20 Multi-Agent Beta model configuration
2026-03-13 10:48:38 +01:00
Aiden Cline 066a425917 Merge pull request #1158 from micuintus/feature/Nebius_Nemotron-3-Super-120b-a12b
feat(nebius): Add support for Nemotron-3-Super-120B-A12B
2026-03-12 22:20:20 -05:00
Aiden Cline 6df7f20cdc Merge pull request #1156 from dsingal0/dev
added nemotron super on baseten
2026-03-12 22:20:06 -05:00
Aiden Cline 78bb47b90e Merge pull request #1151 from dacbd/dacbd
fix(wandb): update models
2026-03-12 22:19:43 -05:00
Aiden Cline c121d86419 Merge pull request #1160 from kreatoo/dev
feat: add zai-org/glm-4.7 and zai-org/glm-4.7-flash to NanoGPT
2026-03-12 22:11:18 -05:00
Aiden Cline ab148eeb14 Merge pull request #1161 from Grin1024/dev
Add Claude Opus 4.6 and Sonnet 4.6 models to RequestY provider
2026-03-12 22:11:07 -05:00
lihui 49d196d326 Add Claude Opus 4.6 and Sonnet 4.6 models to RequestY provider 2026-03-13 09:00:54 +08:00
Kreato 8899b390ef feat: add zai-org/glm-4.7 and zai-org/glm-4.7-flash to NanoGPT 2026-03-13 00:27:09 +03:00
Michael Voigt 5217f62ddf fix(nebius): Follow context updates for Kimi 2.5 and GLM-5 2026-03-12 20:22:48 +01:00
Michael Voigt 55eaff9af1 feat(nebius): Add support for Nemotron-3-Super-120B-A12B 2026-03-12 20:22:21 +01:00
Dhruv Singal 7557c06ac0 update output length 2026-03-12 09:41:25 -07:00
Dhruv Singal e85d820121 fix input output 2026-03-12 08:29:01 -07:00
Dhruv Singal 499d3a39ef remove cache pricing 2026-03-12 08:21:22 -07:00
Dhruv Singal b9b38d6e33 added nemotron super on baseten 2026-03-12 08:18:46 -07:00
Aiden Cline ca24ac14fa Merge pull request #1153 from dpuyosa/dev
Venice: Update model output token limits
2026-03-12 10:08:46 -05:00
Aiden Cline 822546fc67 Merge pull request #1155 from spiffytech/dev
Add Ollama Cloud support for Nemotron 3 Super
2026-03-12 10:08:31 -05:00
Aiden Cline 4555195b71 Merge pull request #1152 from v1gnesh/dev
Update grok-4.20 model defs
2026-03-12 10:08:15 -05:00
spiffytech 5eae8effc6 Added Ollama Cloud support for Nemotron 3 Super 2026-03-12 09:28:47 -04:00
dpuyosa c1801aef87 [venice] Normalize model output token limits
- Update output limits to standard values across all models
2026-03-12 10:08:39 +01:00
v1gnesh 5e6464b272 Update grok-4.20-beta-reasoning 2026-03-12 10:27:40 +05:30
v1gnesh e1a4f23332 Update grok-4.20-beta-non-reasoning 2026-03-12 10:26:03 +05:30
v1gnesh 753e1f9f0c grok-multi-agent-beta update 2026-03-12 10:23:57 +05:30
Daniel Barnes 123ecd2ba5 docs url 2026-03-12 13:27:56 +09:00
Daniel Barnes f15cda9fcb remove old 2026-03-12 13:26:08 +09:00
Daniel Barnes 0205debbd3 fix values 2026-03-12 13:22:29 +09:00
Daniel Barnes 0059766509 number formating 2026-03-12 13:17:22 +09:00
Daniel Barnes be81b02916 additional model files 2026-03-12 13:02:17 +09:00
Daniel Barnes 2dab141166 initial script & model updates 2026-03-12 13:01:35 +09:00
Aiden Cline 45aa49af25 tweak: azure kimi k2.5 2026-03-11 22:35:20 -05:00
Aiden Cline 781fad3ad4 Merge pull request #1150 from cau1k/5.4-family
feat(azure): add 5.4/pro families
2026-03-11 22:14:08 -05:00
cau1k 99d2ffcfdd feat(azure): add 5.4/pro families 2026-03-11 20:59:11 -04:00
Aiden Cline 381d7cc19d Merge pull request #1149 from ariane-emory/fear/add-march-or-stealth-models
Add OpenRouter stealth models: Hunter Alpha and Healer Alpha
2026-03-11 18:07:50 -05:00
Ariane Emory 7482e22458 Fix family field to use 'alpha' for stealth models 2026-03-11 18:49:32 -04:00
Ariane Emory f5e6a402e6 Add OpenRouter stealth models: Hunter Alpha and Healer Alpha 2026-03-11 18:41:58 -04:00
Aiden Cline 9265852852 tweak: adjust some gh limits to align better w/ api 2026-03-11 15:23:44 -05:00
Aiden Cline dc98a32996 Merge pull request #1018 from Sewer56/add-synthetic-missing-models
Update synthetic.new models: promote MiniMax-M2.5, add GLM-4.7-Flash
2026-03-11 14:55:50 -05:00
Aiden Cline 56c39ae0f6 Merge pull request #1140 from sk0x0y/feature/nanogpt-thudm-id-fixes
fix(nano-gpt): rename THUDM 2 ids to canonical THUDM ids
2026-03-11 14:55:07 -05:00
Aiden Cline b1f43a7595 Merge pull request #1147 from msadiks/fix/alibaba-coding-minimax
fix: alibaba-coding-plan MiniMax-M2.5 context window
2026-03-11 14:54:37 -05:00
Matt Cowger fed8bcae19 Merge branch 'dev' into mcowger/correct-gemini-flash-lite-pricing 2026-03-11 12:23:42 -07:00
sk0x0y fb95150d02 fix(nano-gpt): rename VongolaChouko model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:20:39 +09:00
sk0x0y a7c9a240b4 fix(nano-gpt): rename Steelskull model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:20:39 +09:00
sk0x0y 6432a4a3e6 fix(nano-gpt): rename Sao10K model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:20:38 +09:00
sk0x0y f2e4a249fe fix(nano-gpt): rename NeverSleep model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:20:38 +09:00
sk0x0y 8667a6eed8 fix(nano-gpt): rename MarinaraSpaghetti model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:19:56 +09:00
sk0x0y 429554397a fix(nano-gpt): rename LatitudeGames model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:19:56 +09:00
sk0x0y a64e6ad0ac fix(nano-gpt): rename LLM360 model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:19:56 +09:00
sk0x0y d68d79888c fix(nano-gpt): rename Infermatic model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:19:56 +09:00
sk0x0y 6c52905c6a fix(nano-gpt): rename Gryphe model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:19:55 +09:00
sk0x0y 62410b8f26 fix(nano-gpt): rename GalrionSoftworks model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:19:55 +09:00
sk0x0y 50ce68ccab fix(nano-gpt): rename Envoid model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:19:20 +09:00
sk0x0y d1c6a6b873 fix(nano-gpt): rename EVA-UNIT-01 model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-12 04:19:20 +09:00
Frank 7193b068a5 update zen models 2026-03-11 13:52:50 -04:00
Sadik 79a8a06bd7 fix MiniMax-M2.5 context window 2026-03-11 20:50:33 +03:00
Aiden Cline b60c03e11c Merge pull request #1139 from zainhas/dev
[Together AI] add prompt caching pricing for MiniMax m2.5
2026-03-11 12:31:56 -05:00
Aiden Cline 15cf98d57b Merge pull request #1146 from gotjoshua/patch-1
Rename step-3-5-flash.toml to step-3.5-flash.toml
2026-03-11 12:31:39 -05:00
Aiden Cline b2ee6c407b Merge pull request #1144 from micuintus/feature/update-nebius-changes
Feat: update Nebius changes
2026-03-11 12:31:29 -05:00
gotjoshua 96a14a06e7 Rename step-3-5-flash.toml to step-3.5-flash.toml
on nvidia it is 3.5 not 3-5
2026-03-11 11:41:36 +00:00
Michael Voigt adc358606d fix(nebius): update model context limits per API 2026-03-11 11:33:14 +01:00
Michael Voigt 63d52adf6f feat(nebius): add GLM-5 model 2026-03-11 11:33:14 +01:00
sk0x0y 9a31387766 fix(nano-gpt): rename Salesforce model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 18:17:41 +09:00
sk0x0y 735157b837 fix(nano-gpt): rename ReadyArt model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 18:17:41 +09:00
sk0x0y d75b46fb37 fix(nano-gpt): rename Doctor-Shotgun model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 18:17:41 +09:00
sk0x0y cc555f8482 fix(nano-gpt): rename CrucibleLab model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 18:17:13 +09:00
sk0x0y 7fbbcf2b49 fix(nano-gpt): rename MiniMaxAI model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 16:04:28 +09:00
sk0x0y 14c8ec8ca5 fix(nano-gpt): rename Tongyi-Zhiwen model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 16:04:28 +09:00
sk0x0y 72568bbdb3 fix(nano-gpt): rename Alibaba-NLP model id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 16:03:57 +09:00
sk0x0y c2225b715f fix(nano-gpt): rename THUDM GLM-Z1 rumination id
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 15:49:20 +09:00
sk0x0y 7ce25e3742 fix(nano-gpt): rename THUDM GLM-Z1 model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 15:49:20 +09:00
sk0x0y 427868604b fix(nano-gpt): rename THUDM GLM-4 model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 15:49:20 +09:00
Zain Hasan 247cd801a8 add prompt caching pricing for MiniMax m2.5 2026-03-10 22:54:42 -07:00
Aiden Cline 1aa2ee22b1 Merge pull request #1134 from sk0x0y/feature/nanogpt-catalog-fixes
fix(nano-gpt): correct TEE path ids and add missing canonical entries
2026-03-10 22:02:52 -05:00
Aiden Cline 0f57233eff Merge pull request #1105 from sylviezhang37/add-vercel-input-context-and-new-models
feat(vercel): add input context calculation + new models
2026-03-10 22:01:52 -05:00
Aiden Cline 73a78eebfc Merge pull request #1138 from mugnimaestra/feat/add-glm-5-turbo-chutes
feat: add GLM-5-Turbo to Chutes provider listings
2026-03-10 22:01:08 -05:00
Sylvie Zhang 3a6789b819 Merge branch 'dev' into add-vercel-input-context-and-new-models 2026-03-10 17:44:14 -07:00
Sylvie Zhang f7c505e140 remove context from gemini models 2026-03-10 17:43:08 -07:00
Sylvie Zhang 6bb36806d6 only calc input context for openai models 2026-03-10 17:40:46 -07:00
Sylvie Zhang 20a404eb88 revert non openai changes 2026-03-10 17:38:46 -07:00
Muhammad Mugni Hadi 65ecb5cd4a feat: add GLM-5-Turbo to Chutes provider listings 2026-03-11 05:26:11 +07:00
Matt Cowger 56062a9129 Fix incorrect pricing 2026-03-10 14:57:44 -07:00
Aiden Cline d3d9c580d4 Merge pull request #1135 from gitpush-gitpaid/fix/gpt-5-4-pdf-input-modalities
Added PDF to input modalities for GPT-5.4
2026-03-10 13:53:42 -05:00
gitpush-gitpaid ef98d8a9cb Updated GPT-5.4 PDF input modalities 2026-03-10 13:59:29 -04:00
sk0x0y 9d17752b88 fix(nano-gpt): add missing GLM 5 thinking model
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 00:56:22 +09:00
sk0x0y b5a838fe8b fix(nano-gpt): add missing TEE qwen3.5 model
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 00:56:22 +09:00
sk0x0y aa1ac39ee6 fix(nano-gpt): rename TEE gemma and minimax ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 00:56:22 +09:00
sk0x0y 4bc17ccf96 fix(nano-gpt): rename TEE oss and llama ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 00:56:22 +09:00
sk0x0y 08c1899bfe fix(nano-gpt): rename TEE deepseek model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 00:56:02 +09:00
sk0x0y ad50e4a5ed fix(nano-gpt): rename TEE qwen model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 00:56:02 +09:00
sk0x0y 730915a123 fix(nano-gpt): rename TEE kimi model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 00:56:02 +09:00
sk0x0y 6f12d18cb8 fix(nano-gpt): rename TEE glm model ids
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-11 00:56:02 +09:00
Aiden Cline bd8774db99 Merge pull request #1132 from sk0x0y/feature/nanogpt-model-sync
feat(nano-gpt): add text and image models
2026-03-10 10:31:50 -05:00
Aiden Cline 88fbea52a4 Merge pull request #1133 from anomalyco/fix-model
fix: bedrock devstral
2026-03-10 10:31:08 -05:00
sk0x0y 898b3c18b7 feat(nano-gpt): add image models
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-10 22:13:48 +09:00
sk0x0y 6316e543ef feat(nano-gpt): add text models
Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-opencode)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
2026-03-10 22:13:48 +09:00
Sewer56 7a02946620 Update synthetic models: promote MiniMax-M2.5, add GLM-4.7-Flash, remove deprecated Qwen3.5 2026-03-08 22:56:31 +00:00
skywalker512 236af40da3 feat: add Tencent Coding Plan provider
Add support for Tencent Coding Plan with 8 models:
- Auto (tc-code-latest)
- Hunyuan 2.0 Instruct
- Hunyuan 2.0 Think
- Hunyuan-T1
- Hunyuan-TurboS
- MiniMax-M2.5
- Kimi-K2.5
- GLM-5

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-08 15:34:36 +08:00
Sylvie Zhang 26465319d6 Merge branch 'dev' into add-vercel-input-context-and-new-models 2026-03-07 11:39:53 -08:00
Sylvie Zhang 7a11ef241d update more models 2026-03-06 08:24:49 -08:00
Sylvie Zhang 145862315d add input calculation + new models 2026-03-06 08:07:11 -08:00
Sewer56 0428299773 Added: Qwen3.5-397B natively supports image, MM2.5 No Image as it was a mistake. 2026-02-25 08:10:42 +00:00
Sewer56 eee3303df0 Add missing synthetic.new models
Add configuration for hf:Qwen/Qwen3.5-397B-A17B and hf:MiniMaxAI/MiniMax-M2.5
to the synthetic provider, based on API specs from synthetic.new.

Note: API reports image support but these models may not natively support
images (likely rerouted/proxied through vision-capable infrastructure).
2026-02-24 09:19:57 +00:00
613 changed files with 6019 additions and 1528 deletions
+11
View File
@@ -199,6 +199,17 @@ $ bun run dev
And it'll open the frontend at http://localhost:3000
### Manual testing with opencode
You can manually check provider changes with opencode by:
```bash
$ bun install
$ cd packages/web
$ bun run build
$ OPENCODE_MODELS_PATH="dist/_api.json" opencode
```
### Questions?
Open an issue if you need help or have questions about contributing.
+2 -1
View File
@@ -18,7 +18,8 @@
"validate": "bun ./packages/core/script/validate.ts",
"helicone:generate": "bun ./packages/core/script/generate-helicone.ts",
"venice:generate": "bun ./packages/core/script/generate-venice.ts",
"vercel:generate": "bun ./packages/core/script/generate-vercel.ts"
"vercel:generate": "bun ./packages/core/script/generate-vercel.ts",
"wandb:generate": "bun ./packages/core/script/generate-wandb.ts"
},
"dependencies": {
"@cloudflare/workers-types": "^4.20250801.0",
+12
View File
@@ -24,6 +24,7 @@ enum ModelType {
enum SkipZeroFields {
LimitContext = "limit.context",
LimitInput = "limit.input",
LimitOutput = "limit.output",
}
@@ -82,6 +83,7 @@ interface ExistingModel {
};
limit?: {
context?: number;
input?: number;
output?: number;
};
modalities?: {
@@ -112,6 +114,7 @@ interface MergedModel {
};
limit: {
context: number;
input?: number;
output: number;
};
modalities: {
@@ -232,6 +235,10 @@ async function loadExistingModel(filePath: string): Promise<ExistingModel | null
}
}
function isOpenAIModel(modelId: string): boolean {
return modelId.startsWith("openai/");
}
function mergeModel(
apiModel: z.infer<typeof VercelModel>,
existing: ExistingModel | null,
@@ -281,6 +288,7 @@ function mergeModel(
...(status && { status }),
limit: {
context: contextLimit,
...(isOpenAIModel(apiModel.id) && contextLimit > outputLimit && { input: contextLimit - outputLimit }),
output: outputLimit,
},
modalities: {
@@ -362,6 +370,9 @@ function formatToml(model: MergedModel): string {
lines.push("");
lines.push(`[limit]`);
lines.push(`context = ${formatNumber(model.limit.context)}`);
if (model.limit.input !== undefined) {
lines.push(`input = ${formatNumber(model.limit.input)}`);
}
lines.push(`output = ${formatNumber(model.limit.output)}`);
lines.push("");
@@ -435,6 +446,7 @@ function detectChanges(
compare("cost.cache_read", existing.cost?.cache_read, merged.cost?.cache_read);
compare("cost.cache_write", existing.cost?.cache_write, merged.cost?.cache_write);
compare("limit.context", existing.limit?.context, merged.limit.context);
compare("limit.input", existing.limit?.input, merged.limit.input);
compare("limit.output", existing.limit?.output, merged.limit.output);
compare("modalities.input", existing.modalities?.input, merged.modalities.input);
+525
View File
@@ -0,0 +1,525 @@
#!/usr/bin/env bun
import path from "node:path";
import { mkdir } from "node:fs/promises";
import { z } from "zod";
import { ModelFamilyValues } from "../src/family.js";
const API_ENDPOINT = "https://trace.wandb.ai/inference/analysis/artificialanalysis/models";
const Pricing = z
.object({
prompt: z.string().optional(),
completion: z.string().optional(),
image: z.string().optional(),
request: z.string().optional(),
input_cache_reads: z.string().optional(),
input_cache_writes: z.string().optional(),
})
.passthrough();
const WandbModel = z
.object({
id: z.string(),
name: z.string(),
created: z.number(),
input_modalities: z.array(z.string()),
output_modalities: z.array(z.string()),
context_length: z.number(),
max_output_length: z.number(),
pricing: Pricing.optional(),
supported_sampling_parameters: z.array(z.string()).default([]),
supported_features: z.array(z.string()).default([]),
})
.passthrough();
const WandbResponse = z
.object({
data: z.array(WandbModel),
})
.strict();
interface ExistingModel {
name?: string;
family?: string;
attachment?: boolean;
reasoning?: boolean;
tool_call?: boolean;
structured_output?: boolean;
temperature?: boolean;
knowledge?: string;
release_date?: string;
last_updated?: string;
open_weights?: boolean;
interleaved?: boolean | { field: string };
status?: string;
cost?: {
input?: number;
output?: number;
cache_read?: number;
cache_write?: number;
};
limit?: {
context?: number;
input?: number;
output?: number;
};
modalities?: {
input?: string[];
output?: string[];
};
}
interface MergedModel {
name: string;
family?: string;
attachment: boolean;
reasoning: boolean;
tool_call: boolean;
structured_output?: boolean;
temperature: boolean;
knowledge?: string;
release_date: string;
last_updated: string;
open_weights: boolean;
interleaved?: boolean | { field: string };
status?: string;
cost?: {
input: number;
output: number;
cache_read?: number;
cache_write?: number;
};
limit: {
context: number;
output: number;
};
modalities: {
input: Array<"text" | "audio" | "image" | "video" | "pdf">;
output: Array<"text" | "audio" | "image" | "video" | "pdf">;
};
}
interface Changes {
field: string;
oldValue: string;
newValue: string;
}
type SupportedModality = "text" | "audio" | "image" | "video" | "pdf";
const modalityMap: Record<string, SupportedModality | undefined> = {
text: "text",
image: "image",
audio: "audio",
video: "video",
pdf: "pdf",
file: "pdf",
files: "pdf",
};
const openWeightsPrefixes = new Set([
"deepseek-ai/",
"meta-llama/",
"microsoft/",
"MiniMaxAI/",
"moonshotai/",
"nvidia/",
"OpenPipe/",
"Qwen/",
"zai-org/",
]);
function timestampToDate(timestamp: number): string {
return new Date(timestamp * 1000).toISOString().slice(0, 10);
}
function getTodayDate(): string {
return new Date().toISOString().slice(0, 10);
}
function formatNumber(n: number): string {
if (n >= 1000) {
return n.toString().replace(/\B(?=(\d{3})+(?!\d))/g, "_");
}
return n.toString();
}
function formatDecimal(n: number): string {
return Number(n.toFixed(6)).toString();
}
function priceToPerMillion(value: string): number {
return Number((parseFloat(value) * 1_000_000).toFixed(6));
}
function isSubstring(target: string, family: string): boolean {
return target.toLowerCase().includes(family.toLowerCase());
}
function matchesFamily(target: string, family: string): boolean {
const targetLower = target.toLowerCase();
const familyLower = family.toLowerCase();
let familyIdx = 0;
for (let i = 0; i < targetLower.length && familyIdx < familyLower.length; i++) {
if (targetLower[i] === familyLower[familyIdx]) {
familyIdx++;
}
}
return familyIdx === familyLower.length;
}
function inferFamily(modelId: string, modelName: string): string | undefined {
const sortedFamilies = [...ModelFamilyValues].sort((a, b) => b.length - a.length);
for (const family of sortedFamilies) {
if (isSubstring(modelId, family) || isSubstring(modelName, family)) {
return family;
}
}
for (const family of sortedFamilies) {
if (matchesFamily(modelId, family) || matchesFamily(modelName, family)) {
return family;
}
}
return undefined;
}
function normalizeName(apiModel: z.infer<typeof WandbModel>): string {
const stripped = apiModel.name.replace(/^[^:]+:\s*/, "").trim();
return stripped || path.basename(apiModel.id);
}
function inferReasoning(apiModel: z.infer<typeof WandbModel>): boolean {
const text = `${apiModel.id} ${apiModel.name}`.toLowerCase();
return text.includes("thinking") || /\br1\b/.test(text) || text.includes("reasoning");
}
function inferOpenWeights(modelId: string): boolean {
for (const prefix of openWeightsPrefixes) {
if (modelId.startsWith(prefix)) {
return true;
}
}
return false;
}
function normalizeModalities(values: string[]): SupportedModality[] {
const normalized = values
.map((value) => modalityMap[value.toLowerCase()])
.filter((value): value is SupportedModality => value !== undefined);
return [...new Set(normalized)];
}
async function loadExistingModel(filePath: string): Promise<ExistingModel | null> {
try {
const file = Bun.file(filePath);
if (!(await file.exists())) {
return null;
}
const toml = await import(filePath, { with: { type: "toml" } }).then((mod) => mod.default);
return toml as ExistingModel;
} catch (cause) {
console.warn(`Warning: Failed to parse existing file ${filePath}:`, cause);
return null;
}
}
function mergeModel(
apiModel: z.infer<typeof WandbModel>,
existing: ExistingModel | null,
): MergedModel {
const featureSet = new Set(apiModel.supported_features);
const samplingSet = new Set(apiModel.supported_sampling_parameters);
const inputModalities = normalizeModalities(apiModel.input_modalities);
const outputModalities = normalizeModalities(apiModel.output_modalities);
const merged: MergedModel = {
name: existing?.name ?? normalizeName(apiModel),
family: existing?.family ?? inferFamily(apiModel.id, apiModel.name),
attachment: existing?.attachment ?? inputModalities.some((m) => m !== "text"),
reasoning: existing?.reasoning ?? inferReasoning(apiModel),
tool_call: existing?.tool_call ?? featureSet.has("tools"),
temperature: existing?.temperature ?? samplingSet.has("temperature"),
release_date: existing?.release_date ?? timestampToDate(apiModel.created),
last_updated: getTodayDate(),
open_weights: existing?.open_weights ?? inferOpenWeights(apiModel.id),
...(existing?.structured_output !== undefined
? { structured_output: existing.structured_output }
: featureSet.has("structured_outputs")
? { structured_output: true }
: {}),
...(existing?.knowledge ? { knowledge: existing.knowledge } : {}),
...(existing?.interleaved !== undefined ? { interleaved: existing.interleaved } : {}),
...(existing?.status ? { status: existing.status } : {}),
limit: {
context: apiModel.context_length > 0 ? apiModel.context_length : (existing?.limit?.context ?? 0),
output: apiModel.max_output_length > 0
? apiModel.max_output_length
: (existing?.limit?.output ?? 0),
},
modalities: {
input: inputModalities.length > 0
? inputModalities
: ((existing?.modalities?.input as SupportedModality[] | undefined) ?? ["text"]),
output: outputModalities.length > 0
? outputModalities
: ((existing?.modalities?.output as SupportedModality[] | undefined) ?? ["text"]),
},
};
const prompt = apiModel.pricing?.prompt;
const completion = apiModel.pricing?.completion;
const cacheRead = apiModel.pricing?.input_cache_reads;
const cacheWrite = apiModel.pricing?.input_cache_writes;
if (prompt && completion) {
merged.cost = {
input: priceToPerMillion(prompt),
output: priceToPerMillion(completion),
...(cacheRead && parseFloat(cacheRead) > 0
? { cache_read: priceToPerMillion(cacheRead) }
: {}),
...(cacheWrite && parseFloat(cacheWrite) > 0
? { cache_write: priceToPerMillion(cacheWrite) }
: {}),
};
} else if (existing?.cost?.input !== undefined && existing.cost.output !== undefined) {
merged.cost = {
input: existing.cost.input,
output: existing.cost.output,
...(existing.cost.cache_read !== undefined ? { cache_read: existing.cost.cache_read } : {}),
...(existing.cost.cache_write !== undefined ? { cache_write: existing.cost.cache_write } : {}),
};
}
return merged;
}
function formatToml(model: MergedModel): string {
const lines: string[] = [];
lines.push(`name = "${model.name.replace(/"/g, '\\"')}"`);
if (model.family) {
lines.push(`family = "${model.family}"`);
}
lines.push(`release_date = "${model.release_date}"`);
lines.push(`last_updated = "${model.last_updated}"`);
lines.push(`attachment = ${model.attachment}`);
lines.push(`reasoning = ${model.reasoning}`);
if (model.structured_output !== undefined) {
lines.push(`structured_output = ${model.structured_output}`);
}
lines.push(`temperature = ${model.temperature}`);
lines.push(`tool_call = ${model.tool_call}`);
if (model.knowledge) {
lines.push(`knowledge = "${model.knowledge}"`);
}
lines.push(`open_weights = ${model.open_weights}`);
if (model.status) {
lines.push(`status = "${model.status}"`);
}
if (model.interleaved !== undefined) {
lines.push("");
if (model.interleaved === true) {
lines.push("interleaved = true");
} else {
lines.push("[interleaved]");
lines.push(`field = "${model.interleaved.field}"`);
}
}
if (model.cost) {
lines.push("");
lines.push("[cost]");
lines.push(`input = ${formatDecimal(model.cost.input)}`);
lines.push(`output = ${formatDecimal(model.cost.output)}`);
if (model.cost.cache_read !== undefined) {
lines.push(`cache_read = ${formatDecimal(model.cost.cache_read)}`);
}
if (model.cost.cache_write !== undefined) {
lines.push(`cache_write = ${formatDecimal(model.cost.cache_write)}`);
}
}
lines.push("");
lines.push("[limit]");
lines.push(`context = ${formatNumber(model.limit.context)}`);
lines.push(`output = ${formatNumber(model.limit.output)}`);
lines.push("");
lines.push("[modalities]");
lines.push(`input = [${model.modalities.input.map((m) => `"${m}"`).join(", ")}]`);
lines.push(`output = [${model.modalities.output.map((m) => `"${m}"`).join(", ")}]`);
return `${lines.join("\n")}\n`;
}
function detectChanges(existing: ExistingModel | null, merged: MergedModel): Changes[] {
if (!existing) {
return [];
}
const changes: Changes[] = [];
const epsilon = 0.001;
const formatValue = (value: unknown): string => {
if (typeof value === "number") return formatNumber(value);
if (Array.isArray(value)) return `[${value.join(", ")}]`;
if (value === undefined) return "(none)";
return String(value);
};
const compare = (field: string, oldValue: unknown, newValue: unknown) => {
const changed = field.startsWith("cost.")
? (
oldValue === undefined && newValue === undefined
? false
: oldValue === undefined || newValue === undefined
? true
: Math.abs((oldValue as number) - (newValue as number)) > epsilon
)
: JSON.stringify(oldValue) !== JSON.stringify(newValue);
if (changed) {
changes.push({
field,
oldValue: formatValue(oldValue),
newValue: formatValue(newValue),
});
}
};
compare("name", existing.name, merged.name);
compare("family", existing.family, merged.family);
compare("release_date", existing.release_date, merged.release_date);
compare("attachment", existing.attachment, merged.attachment);
compare("reasoning", existing.reasoning, merged.reasoning);
compare("structured_output", existing.structured_output, merged.structured_output);
compare("temperature", existing.temperature, merged.temperature);
compare("tool_call", existing.tool_call, merged.tool_call);
compare("open_weights", existing.open_weights, merged.open_weights);
compare("cost.input", existing.cost?.input, merged.cost?.input);
compare("cost.output", existing.cost?.output, merged.cost?.output);
compare("cost.cache_read", existing.cost?.cache_read, merged.cost?.cache_read);
compare("cost.cache_write", existing.cost?.cache_write, merged.cost?.cache_write);
compare("limit.context", existing.limit?.context, merged.limit.context);
compare("limit.output", existing.limit?.output, merged.limit.output);
compare("modalities.input", existing.modalities?.input, merged.modalities.input);
compare("modalities.output", existing.modalities?.output, merged.modalities.output);
return changes;
}
async function main() {
const args = process.argv.slice(2);
const dryRun = args.includes("--dry-run");
const newOnly = args.includes("--new-only");
const modelsDir = path.join(import.meta.dirname, "..", "..", "..", "providers", "wandb", "models");
console.log(`${dryRun ? "[DRY RUN] " : ""}${newOnly ? "[NEW ONLY] " : ""}Fetching WandB models from API...`);
const res = await fetch(API_ENDPOINT);
if (!res.ok) {
console.error(`Failed to fetch API: ${res.status} ${res.statusText}`);
process.exit(1);
}
const json = await res.json();
const parsed = WandbResponse.safeParse(json);
if (!parsed.success) {
console.error("Invalid API response:", parsed.error.errors);
process.exit(1);
}
const apiModels = parsed.data.data;
const existingFiles = new Set<string>();
for await (const file of new Bun.Glob("**/*.toml").scan({ cwd: modelsDir, absolute: false })) {
existingFiles.add(file);
}
console.log(`Found ${apiModels.length} models in API, ${existingFiles.size} existing files\n`);
const apiModelIds = new Set<string>();
let created = 0;
let updated = 0;
let unchanged = 0;
for (const apiModel of apiModels) {
const relativePath = `${apiModel.id}.toml`;
const filePath = path.join(modelsDir, relativePath);
const dirPath = path.dirname(filePath);
apiModelIds.add(relativePath);
const existing = await loadExistingModel(filePath);
const merged = mergeModel(apiModel, existing);
const tomlContent = formatToml(merged);
if (existing === null) {
created++;
if (dryRun) {
console.log(`[DRY RUN] Would create: ${relativePath}`);
console.log(` name = "${merged.name}"`);
if (merged.family) {
console.log(` family = "${merged.family}"`);
}
console.log("");
} else {
await mkdir(dirPath, { recursive: true });
await Bun.write(filePath, tomlContent);
console.log(`Created: ${relativePath}`);
}
continue;
}
if (newOnly) {
unchanged++;
continue;
}
const changes = detectChanges(existing, merged);
if (changes.length === 0) {
unchanged++;
continue;
}
updated++;
if (dryRun) {
console.log(`[DRY RUN] Would update: ${relativePath}`);
} else {
await mkdir(dirPath, { recursive: true });
await Bun.write(filePath, tomlContent);
console.log(`Updated: ${relativePath}`);
}
for (const change of changes) {
console.log(` ${change.field}: ${change.oldValue}${change.newValue}`);
}
console.log("");
}
const orphaned = [...existingFiles].filter((file) => !apiModelIds.has(file));
for (const file of orphaned) {
console.log(`Warning: Orphaned file (not in API): ${file}`);
}
console.log("");
console.log(
dryRun
? `Summary: ${created} would be created, ${updated} would be updated, ${unchanged} unchanged, ${orphaned.length} orphaned`
: `Summary: ${created} created, ${updated} updated, ${unchanged} unchanged, ${orphaned.length} orphaned`,
);
}
await main();
+6 -3
View File
@@ -96,6 +96,7 @@ export const ModelFamilyValues = [
// NVIDIA Nemotron
"nemotron",
"nemotron-free",
// AWS Titan
"titan",
@@ -103,6 +104,8 @@ export const ModelFamilyValues = [
// MiniMax
"minimax",
"minimax-m2.5",
"minimax-m2.7",
"minimax-free",
// Hunyuan
@@ -193,6 +196,9 @@ export const ModelFamilyValues = [
// Mimo
"mimo",
"mimo-pro-free",
"mimo-omni-free",
"mimo-flash-free",
// Clarifai
"mm-poly",
@@ -284,9 +290,6 @@ export const ModelFamilyValues = [
// Parakeet
"parakeet",
// MiMo
"mimo-flash-free",
// NeMo
"nemoretriever",
@@ -12,6 +12,8 @@ open_weights = false
[cost]
input = 0.25
output = 1.50
cache_read = 0.025
cache_write = 1.00
[limit]
context = 1_048_576
@@ -18,7 +18,8 @@ cache_read = 0
cache_write = 0
[limit]
context = 16_608
context = 196_608
input = 196_601
output = 24_576
[modalities]
@@ -1,22 +0,0 @@
name = "Jamba 1.5 Large"
family = "jamba"
release_date = "2024-08-15"
last_updated = "2024-08-15"
attachment = false
reasoning = false
temperature = true
knowledge = "2024-08"
tool_call = true
open_weights = true
[cost]
input = 2.00
output = 8.00
[limit]
context = 256_000
output = 4_096
[modalities]
input = ["text"]
output = ["text"]
@@ -1,22 +0,0 @@
name = "Jamba 1.5 Mini"
family = "jamba"
release_date = "2024-08-15"
last_updated = "2024-08-15"
attachment = false
reasoning = false
temperature = true
knowledge = "2024-08"
tool_call = true
open_weights = true
[cost]
input = 0.20
output = 0.40
[limit]
context = 256_000
output = 4_096
[modalities]
input = ["text"]
output = ["text"]
@@ -1,21 +0,0 @@
name = "Titan Text G1 - Express"
family = "titan"
release_date = "2024-12-01"
last_updated = "2024-12-01"
attachment = false
reasoning = false
temperature = true
tool_call = true
open_weights = false
[cost]
input = 0.20
output = 0.60
[limit]
context = 128_000
output = 4_096
[modalities]
input = ["text"]
output = ["text"]
@@ -1,21 +0,0 @@
name = "Titan Text G1 - Express"
family = "titan"
release_date = "2024-12-01"
last_updated = "2024-12-01"
attachment = false
reasoning = false
temperature = true
tool_call = true
open_weights = false
[cost]
input = 0.20
output = 0.60
[limit]
context = 128_000
output = 4_096
[modalities]
input = ["text"]
output = ["text"]
@@ -1,22 +0,0 @@
name = "Claude Opus 3"
family = "claude-opus"
release_date = "2024-02-29"
last_updated = "2024-02-29"
attachment = true
reasoning = false
temperature = true
knowledge = "2023-08"
tool_call = true
open_weights = false
[cost]
input = 15.00
output = 75.00
[limit]
context = 200_000
output = 4_096
[modalities]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -1,22 +0,0 @@
name = "Claude Sonnet 3"
family = "claude-sonnet"
release_date = "2024-03-04"
last_updated = "2024-03-04"
attachment = true
reasoning = false
temperature = true
knowledge = "2023-08"
tool_call = true
open_weights = false
[cost]
input = 3.00
output = 15.00
[limit]
context = 200_000
output = 4_096
[modalities]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -1,22 +0,0 @@
name = "Claude Instant"
family = "claude"
release_date = "2023-03-01"
last_updated = "2023-03-01"
attachment = false
reasoning = false
temperature = true
knowledge = "2023-08"
tool_call = false
open_weights = false
[cost]
input = 0.80
output = 2.40
[limit]
context = 100_000
output = 4_096
[modalities]
input = ["text"]
output = ["text"]
@@ -1,7 +1,7 @@
name = "Claude Opus 4.6"
family = "claude-opus"
release_date = "2026-02-05"
last_updated = "2026-02-05"
last_updated = "2026-03-18"
attachment = true
reasoning = true
temperature = true
@@ -15,14 +15,8 @@ output = 25.00
cache_read = 0.50
cache_write = 6.25
[cost.context_over_200k]
input = 10.00
output = 37.50
cache_read = 1.00
cache_write = 12.50
[limit]
context = 200_000
context = 1_000_000
output = 128_000
[modalities]
@@ -1,7 +1,7 @@
name = "Claude Sonnet 4.6"
family = "claude-sonnet"
release_date = "2026-02-17"
last_updated = "2026-02-17"
last_updated = "2026-03-18"
attachment = true
reasoning = true
temperature = true
@@ -15,14 +15,8 @@ output = 15.00
cache_read = 0.30
cache_write = 3.75
[cost.context_over_200k]
input = 6.00
output = 22.50
cache_read = 0.60
cache_write = 7.50
[limit]
context = 200_000
context = 1_000_000
output = 64_000
[modalities]
@@ -1,22 +0,0 @@
name = "Claude 2"
family = "claude"
release_date = "2023-07-11"
last_updated = "2023-07-11"
attachment = false
reasoning = false
temperature = true
knowledge = "2023-08"
tool_call = false
open_weights = false
[cost]
input = 8.00
output = 24.00
[limit]
context = 100_000
output = 4_096
[modalities]
input = ["text"]
output = ["text"]
@@ -1,22 +0,0 @@
name = "Claude 2.1"
family = "claude"
release_date = "2023-11-21"
last_updated = "2023-11-21"
attachment = false
reasoning = false
temperature = true
knowledge = "2023-08"
tool_call = false
open_weights = false
[cost]
input = 8.00
output = 24.00
[limit]
context = 200_000
output = 4_096
[modalities]
input = ["text"]
output = ["text"]
@@ -1,22 +0,0 @@
name = "Command Light"
family = "command-light"
release_date = "2023-11-01"
last_updated = "2023-11-01"
attachment = false
reasoning = false
temperature = true
knowledge = "2023-08"
tool_call = false
open_weights = true
[cost]
input = 0.30
output = 0.60
[limit]
context = 4_096
output = 4_096
[modalities]
input = ["text"]
output = ["text"]
@@ -1,22 +0,0 @@
name = "Command"
family = "command"
release_date = "2023-11-01"
last_updated = "2023-11-01"
attachment = false
reasoning = false
temperature = true
knowledge = "2023-08"
tool_call = false
open_weights = true
[cost]
input = 1.50
output = 2.00
[limit]
context = 4_096
output = 4_096
[modalities]
input = ["text"]
output = ["text"]
@@ -1,7 +1,7 @@
name = "DeepSeek-V3.2"
family = "deepseek"
release_date = "2026-02-15"
last_updated = "2026-02-15"
release_date = "2026-02-06"
last_updated = "2026-02-06"
attachment = false
reasoning = true
temperature = true
@@ -1,7 +1,7 @@
name = "Claude Opus 4.6 (EU)"
family = "claude-opus"
release_date = "2026-02-05"
last_updated = "2026-02-05"
last_updated = "2026-03-18"
attachment = true
reasoning = true
temperature = true
@@ -15,14 +15,8 @@ output = 25.00
cache_read = 0.50
cache_write = 6.25
[cost.context_over_200k]
input = 10.00
output = 37.50
cache_read = 1.00
cache_write = 12.50
[limit]
context = 200_000
context = 1_000_000
output = 128_000
[modalities]
@@ -1,7 +1,7 @@
name = "Claude Sonnet 4.6 (EU)"
family = "claude-sonnet"
release_date = "2026-02-17"
last_updated = "2026-02-17"
last_updated = "2026-03-18"
attachment = true
reasoning = true
temperature = true
@@ -15,14 +15,8 @@ output = 15.00
cache_read = 0.30
cache_write = 3.75
[cost.context_over_200k]
input = 6.00
output = 22.50
cache_read = 0.60
cache_write = 7.50
[limit]
context = 200_000
context = 1_000_000
output = 64_000
[modalities]
@@ -1,7 +1,7 @@
name = "Claude Opus 4.6 (Global)"
family = "claude-opus"
release_date = "2026-02-05"
last_updated = "2026-02-05"
last_updated = "2026-03-18"
attachment = true
reasoning = true
temperature = true
@@ -15,14 +15,8 @@ output = 25.00
cache_read = 0.50
cache_write = 6.25
[cost.context_over_200k]
input = 10.00
output = 37.50
cache_read = 1.00
cache_write = 12.50
[limit]
context = 200_000
context = 1_000_000
output = 128_000
[modalities]
@@ -1,7 +1,7 @@
name = "Claude Sonnet 4.6 (Global)"
family = "claude-sonnet"
release_date = "2026-02-17"
last_updated = "2026-02-17"
last_updated = "2026-03-18"
attachment = true
reasoning = true
temperature = true
@@ -15,14 +15,8 @@ output = 15.00
cache_read = 0.30
cache_write = 3.75
[cost.context_over_200k]
input = 6.00
output = 22.50
cache_read = 0.60
cache_write = 7.50
[limit]
context = 200_000
context = 1_000_000
output = 64_000
[modalities]
@@ -1,4 +1,4 @@
name = "Llama 3 70B Instruct"
name = "Llama 3.1 405B Instruct"
family = "llama"
release_date = "2024-07-23"
last_updated = "2024-07-23"
@@ -6,16 +6,16 @@ attachment = false
reasoning = false
temperature = true
knowledge = "2023-12"
tool_call = false
tool_call = true
open_weights = true
[cost]
input = 2.65
output = 3.50
input = 2.40
output = 2.40
[limit]
context = 8_192
output = 2_048
context = 128_000
output = 4_096
[modalities]
input = ["text"]
@@ -1,12 +1,12 @@
name = "Devstral 2 135B"
family = "mistral"
name = "Devstral 2 123B"
family = "devstral"
release_date = "2026-02-17"
last_updated = "2026-02-17"
attachment = false
reasoning = false
temperature = true
tool_call = true
open_weights = false
open_weights = true
[cost]
input = 0.40
@@ -0,0 +1,21 @@
name = "Magistral Small 1.2"
family = "magistral"
release_date = "2025-12-02"
last_updated = "2025-12-02"
attachment = false
reasoning = true
temperature = true
tool_call = true
open_weights = true
[cost]
input = 0.50
output = 1.50
[limit]
context = 128_000
output = 40_000
[modalities]
input = ["text", "image"]
output = ["text"]
@@ -0,0 +1,21 @@
name = "Ministral 3 3B"
family = "ministral"
release_date = "2025-12-02"
last_updated = "2025-12-02"
attachment = false
reasoning = false
temperature = true
tool_call = true
open_weights = true
[cost]
input = 0.10
output = 0.10
[limit]
context = 256_000
output = 8_192
[modalities]
input = ["text", "image"]
output = ["text"]
@@ -1,21 +0,0 @@
name = "Mistral Large (24.02)"
family = "mistral-large"
release_date = "2024-12-01"
last_updated = "2024-12-01"
attachment = false
reasoning = false
temperature = true
tool_call = true
open_weights = false
[cost]
input = 0.50
output = 1.50
[limit]
context = 128_000
output = 4_096
[modalities]
input = ["text"]
output = ["text"]
@@ -0,0 +1,21 @@
name = "Mistral Large 3"
family = "mistral"
release_date = "2025-12-02"
last_updated = "2025-12-02"
attachment = false
reasoning = false
temperature = true
tool_call = true
open_weights = true
[cost]
input = 0.50
output = 1.50
[limit]
context = 256_000
output = 8_192
[modalities]
input = ["text", "image"]
output = ["text"]
@@ -1,22 +0,0 @@
name = "Mixtral-8x7B-Instruct-v0.1"
family = "mixtral"
release_date = "2025-04-01"
last_updated = "2025-04-01"
attachment = false
reasoning = false
tool_call = false
structured_output = true
temperature = true
open_weights = true
[cost]
input = 0.70
output = 0.70
[limit]
context = 32_000
output = 32_000
[modalities]
input = ["text"]
output = ["text"]
@@ -0,0 +1,21 @@
name = "Pixtral Large (25.02)"
family = "mistral"
release_date = "2025-04-08"
last_updated = "2025-04-08"
attachment = false
reasoning = false
temperature = true
tool_call = true
open_weights = false
[cost]
input = 2.00
output = 6.00
[limit]
context = 128_000
output = 8_192
[modalities]
input = ["text", "image"]
output = ["text"]
@@ -1,17 +1,16 @@
name = "Command R"
family = "command-r"
release_date = "2024-03-11"
last_updated = "2024-03-11"
name = "NVIDIA Nemotron Nano 3 30B"
family = "nemotron"
release_date = "2025-12-23"
last_updated = "2025-12-23"
attachment = false
reasoning = false
reasoning = true
temperature = true
knowledge = "2024-04"
tool_call = true
open_weights = true
[cost]
input = 0.50
output = 1.50
input = 0.06
output = 0.24
[limit]
context = 128_000
@@ -1,7 +1,7 @@
name = "Claude Opus 4.6 (US)"
family = "claude-opus"
release_date = "2026-02-05"
last_updated = "2026-02-05"
last_updated = "2026-03-18"
attachment = true
reasoning = true
temperature = true
@@ -15,14 +15,8 @@ output = 25.00
cache_read = 0.50
cache_write = 6.25
[cost.context_over_200k]
input = 10.00
output = 37.50
cache_read = 1.00
cache_write = 12.50
[limit]
context = 200_000
context = 1_000_000
output = 128_000
[modalities]
@@ -1,7 +1,7 @@
name = "Claude Sonnet 4.6 (US)"
family = "claude-sonnet"
release_date = "2026-02-17"
last_updated = "2026-02-17"
last_updated = "2026-03-18"
attachment = true
reasoning = true
temperature = true
@@ -15,14 +15,8 @@ output = 15.00
cache_read = 0.30
cache_write = 3.75
[cost.context_over_200k]
input = 6.00
output = 22.50
cache_read = 0.60
cache_write = 7.50
[limit]
context = 200_000
context = 1_000_000
output = 64_000
[modalities]
@@ -1,7 +1,7 @@
name = "Claude Opus 4.6"
family = "claude-opus"
release_date = "2026-02-05"
last_updated = "2026-02-05"
last_updated = "2026-03-13"
attachment = true
reasoning = true
temperature = true
@@ -15,14 +15,8 @@ output = 25.00
cache_read = 0.50
cache_write = 6.25
[cost.context_over_200k]
input = 10.00
output = 37.50
cache_read = 1.00
cache_write = 12.50
[limit]
context = 200_000
context = 1_000_000
output = 128_000
[modalities]
@@ -1,7 +1,7 @@
name = "Claude Sonnet 4.6"
family = "claude-sonnet"
release_date = "2026-02-17"
last_updated = "2026-02-17"
last_updated = "2026-03-13"
attachment = true
reasoning = true
temperature = true
@@ -15,14 +15,8 @@ output = 15.00
cache_read = 0.30
cache_write = 3.75
[cost.context_over_200k]
input = 6.00
output = 22.50
cache_read = 0.60
cache_write = 7.50
[limit]
context = 200_000
context = 1_000_000
output = 64_000
[modalities]
@@ -0,0 +1,25 @@
name = "GPT-5.4 Mini"
family = "gpt-mini"
release_date = "2026-03-17"
last_updated = "2026-03-17"
attachment = true
reasoning = true
temperature = false
knowledge = "2025-08-31"
tool_call = true
structured_output = true
open_weights = false
[cost]
input = 0.75
output = 4.50
cache_read = 0.075
[limit]
context = 400_000
input = 272_000
output = 128_000
[modalities]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -0,0 +1,25 @@
name = "GPT-5.4 Nano"
family = "gpt-nano"
release_date = "2026-03-17"
last_updated = "2026-03-17"
attachment = true
reasoning = true
temperature = false
knowledge = "2025-08-31"
tool_call = true
structured_output = true
open_weights = false
[cost]
input = 0.20
output = 1.25
cache_read = 0.02
[limit]
context = 400_000
input = 272_000
output = 128_000
[modalities]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -3,7 +3,7 @@ family = "command-r"
release_date = "2024-08-30"
last_updated = "2024-08-30"
attachment = false
reasoning = true
reasoning = false
temperature = true
knowledge = "2024-06-01"
tool_call = true
@@ -3,7 +3,7 @@ family = "command-r"
release_date = "2024-08-30"
last_updated = "2024-08-30"
attachment = false
reasoning = true
reasoning = false
temperature = true
knowledge = "2024-06-01"
tool_call = true
+25
View File
@@ -0,0 +1,25 @@
name = "GPT-5.4 Mini"
family = "gpt-mini"
release_date = "2026-03-17"
last_updated = "2026-03-17"
attachment = true
reasoning = true
temperature = false
knowledge = "2025-08-31"
tool_call = true
structured_output = true
open_weights = false
[cost]
input = 0.75
output = 4.50
cache_read = 0.075
[limit]
context = 400_000
input = 272_000
output = 128_000
[modalities]
input = ["text", "image", "pdf"]
output = ["text"]
+25
View File
@@ -0,0 +1,25 @@
name = "GPT-5.4 Nano"
family = "gpt-nano"
release_date = "2026-03-17"
last_updated = "2026-03-17"
attachment = true
reasoning = true
temperature = false
knowledge = "2025-08-31"
tool_call = true
structured_output = true
open_weights = false
[cost]
input = 0.20
output = 1.25
cache_read = 0.02
[limit]
context = 400_000
input = 272_000
output = 128_000
[modalities]
input = ["text", "image", "pdf"]
output = ["text"]
+1
View File
@@ -1,4 +1,5 @@
name = "GPT-5.4 Pro"
family = "gpt-pro"
release_date = "2026-03-05"
last_updated = "2026-03-05"
attachment = true
+2 -1
View File
@@ -1,4 +1,5 @@
name = "GPT-5.4"
family = "gpt"
release_date = "2026-03-05"
last_updated = "2026-03-05"
attachment = true
@@ -20,5 +21,5 @@ input = 272_000
output = 128_000
[modalities]
input = ["text", "image"]
input = ["text", "image", "pdf"]
output = ["text"]
@@ -0,0 +1,24 @@
name = "Grok 4.1 Fast (Non-Reasoning)"
family = "grok"
release_date = "2025-06-27"
last_updated = "2025-06-27"
attachment = true
reasoning = false
temperature = true
tool_call = true
open_weights = false
status = "beta"
[cost]
input = 0.20
output = 0.50
cache_read = 0.05
[limit]
context = 128_000
input = 128_000
output = 8_192
[modalities]
input = ["text", "image"]
output = ["text"]
@@ -0,0 +1,24 @@
name = "Grok 4.1 Fast (Reasoning)"
family = "grok"
release_date = "2025-06-27"
last_updated = "2025-06-27"
attachment = true
reasoning = true
temperature = true
tool_call = true
open_weights = false
status = "beta"
[cost]
input = 0.20
output = 0.50
cache_read = 0.05
[limit]
context = 128_000
input = 128_000
output = 8_192
[modalities]
input = ["text", "image"]
output = ["text"]
+2
View File
@@ -25,3 +25,5 @@ output = ["text"]
[provider]
shape = "completions"
npm = "@ai-sdk/openai-compatible"
api = "https://${AZURE_RESOURCE_NAME}.services.ai.azure.com/models"
@@ -1,22 +1,22 @@
name = "DeepSeek-V3-0324"
name = "DeepSeek V3 0324"
family = "deepseek"
release_date = "2025-03-24"
last_updated = "2025-03-24"
attachment = false
reasoning = false
temperature = true
knowledge = "2024-10"
knowledge = "2024-12"
tool_call = true
open_weights = true
[cost]
input = 1.14
output = 2.75
[limit]
context = 161_000
output = 8192
context = 164_000
output = 131_000
[modalities]
input = ["text"]
output = ["text"]
[cost]
input = 0.77
output = 0.77
@@ -0,0 +1,21 @@
name = "DeepSeek V3.1"
family = "deepseek"
release_date = "2025-08-25"
last_updated = "2025-08-25"
attachment = false
reasoning = true
temperature = true
tool_call = true
open_weights = true
[limit]
context = 164_000
output = 131_000
[modalities]
input = ["text"]
output = ["text"]
[cost]
input = 0.50
output = 1.50
@@ -1,9 +1,10 @@
name = "DeepSeek V3.2"
family = "deepseek"
status = "deprecated"
release_date = "2025-12-01"
last_updated = "2025-12-01"
last_updated = "2026-03-06"
attachment = false
reasoning = false
reasoning = true
temperature = true
knowledge = "2025-10"
tool_call = true
@@ -1,7 +1,8 @@
name = "Kimi K2 Instruct 0905"
family = "kimi"
status = "deprecated"
release_date = "2025-09-05"
last_updated = "2025-09-05"
last_updated = "2026-03-06"
attachment = false
reasoning = false
temperature = true
@@ -1,7 +1,8 @@
name = "Kimi K2 Thinking"
family = "kimi-thinking"
status = "deprecated"
release_date = "2025-11-06"
last_updated = "2025-11-06"
last_updated = "2026-03-06"
attachment = false
reasoning = true
temperature = true
@@ -0,0 +1,25 @@
name = "Nemotron 3 Super"
family = "nemotron"
release_date = "2026-03-11"
last_updated = "2026-03-11"
attachment = false
reasoning = true
temperature = true
tool_call = true
knowledge = "2026-02"
open_weights = true
[interleaved]
field = "reasoning_content"
[cost]
input = 0.30
output = 0.75
[limit]
context = 262144
output = 32678
[modalities]
input = ["text"]
output = ["text"]
@@ -1,20 +1,22 @@
name = "OpenAI: gpt-oss-120b (exacto)"
name = "GPT OSS 120B"
family = "gpt-oss"
release_date = "2025-08-05"
last_updated = "2025-08-05"
attachment = false
reasoning = true
temperature = true
knowledge = "2025-08"
tool_call = true
open_weights = true
[cost]
input = 0.039
output = 0.19
[limit]
context = 131072
output = 26215
context = 128_000
output = 128_000
[modalities]
input = ["text"]
output = ["text"]
[cost]
input = 0.10
output = 0.50
@@ -10,8 +10,9 @@ structured_output = true
open_weights = true
[cost]
input = 0.15
output = 0.6
input = 0.30
output = 1.10
cache_read = 0.15
[limit]
context = 196_608
@@ -10,9 +10,9 @@ structured_output = true
open_weights = true
[cost]
input = 0.30
output = 1.20
cache_read = 0.15
input = 0.39
output = 2.34
cache_read = 0.195
[limit]
context = 262_144
@@ -10,12 +10,12 @@ structured_output = true
open_weights = true
[cost]
input = 0.25
output = 0.38
cache_read = 0.125
input = 0.28
output = 0.42
cache_read = 0.14
[limit]
context = 163_840
context = 131_072
output = 65_536
[modalities]
@@ -10,8 +10,9 @@ structured_output = true
open_weights = true
[cost]
input = 0.35
output = 1.50
input = 0.40
output = 1.70
cache_read = 0.20
[limit]
context = 202_752
@@ -12,6 +12,7 @@ open_weights = true
[cost]
input = 0.30
output = 0.90
cache_read = 0.15
[limit]
context = 131_072
@@ -10,8 +10,9 @@ structured_output = true
open_weights = true
[cost]
input = 0.750
output = 2.50
input = 0.95
output = 3.15
cache_read = 0.475
[limit]
context = 202_752
@@ -0,0 +1,26 @@
name = "GLM 5 Turbo"
family = "glm"
release_date = "2026-03-11"
last_updated = "2026-03-11"
attachment = false
reasoning = true
temperature = true
tool_call = true
structured_output = true
open_weights = true
[cost]
input = 0.49
output = 1.96
cache_read = 0.245
[limit]
context = 202_752
output = 65_535
[modalities]
input = ["text"]
output = ["text"]
[interleaved]
field = "reasoning_content"
@@ -0,0 +1,27 @@
name = "Kimi K2.5"
family = "kimi"
release_date = "2026-01-27"
last_updated = "2026-01-27"
attachment = true
reasoning = true
structured_output = true
temperature = true
tool_call = true
knowledge = "2025-01"
open_weights = true
[interleaved]
field = "reasoning_content"
[cost]
input = 0.60
output = 3.00
cache_read = 0.10
[limit]
context = 256_000
output = 256_000
[modalities]
input = ["text", "image"]
output = ["text"]
@@ -0,0 +1,24 @@
name = "Nemotron 3 Super 120B"
family = "nemotron"
release_date = "2026-03-11"
last_updated = "2026-03-11"
attachment = false
reasoning = true
temperature = true
tool_call = true
open_weights = true
[interleaved]
field = "reasoning_content"
[cost]
input = 0.50
output = 1.50
[limit]
context = 256_000
output = 256_000
[modalities]
input = ["text"]
output = ["text"]
@@ -0,0 +1,22 @@
name = "GLM-4.7-Flash"
family = "glm-flash"
release_date = "2026-01-19"
last_updated = "2026-01-19"
attachment = false
reasoning = true
temperature = true
tool_call = true
knowledge = "2025-04"
open_weights = true
[cost]
input = 0.06
output = 0.40
[limit]
context = 131_072
output = 131_072
[modalities]
input = ["text"]
output = ["text"]
@@ -0,0 +1,27 @@
name = "Kimi K2.5"
family = "kimi"
release_date = "2026-01-27"
last_updated = "2026-01-27"
attachment = true
reasoning = true
structured_output = true
temperature = true
tool_call = true
knowledge = "2025-01"
open_weights = true
[interleaved]
field = "reasoning_content"
[cost]
input = 0.60
output = 3.00
cache_read = 0.10
[limit]
context = 256_000
output = 256_000
[modalities]
input = ["text", "image"]
output = ["text"]
@@ -0,0 +1,24 @@
name = "Nemotron 3 Super 120B"
family = "nemotron"
release_date = "2026-03-11"
last_updated = "2026-03-11"
attachment = false
reasoning = true
temperature = true
tool_call = true
open_weights = true
[interleaved]
field = "reasoning_content"
[cost]
input = 0.50
output = 1.50
[limit]
context = 256_000
output = 256_000
[modalities]
input = ["text"]
output = ["text"]
@@ -4,7 +4,7 @@ last_updated = "2024-10-24"
attachment = false
reasoning = false
temperature = true
tool_call = true
tool_call = false
open_weights = true
[limit]
@@ -4,7 +4,7 @@ last_updated = "2024-10-24"
attachment = false
reasoning = false
temperature = true
tool_call = true
tool_call = false
open_weights = true
[limit]
@@ -4,7 +4,7 @@ last_updated = "2025-05-14"
attachment = true
reasoning = false
temperature = true
tool_call = true
tool_call = false
open_weights = true
[limit]
@@ -4,7 +4,7 @@ last_updated = "2025-05-14"
attachment = true
reasoning = false
temperature = true
tool_call = true
tool_call = false
open_weights = true
[limit]
@@ -3,7 +3,7 @@ family = "command-a"
release_date = "2025-03-13"
last_updated = "2025-03-13"
attachment = false
reasoning = true
reasoning = false
temperature = true
knowledge = "2024-06-01"
tool_call = true
@@ -3,7 +3,7 @@ family = "command-r"
release_date = "2024-08-30"
last_updated = "2024-08-30"
attachment = false
reasoning = true
reasoning = false
temperature = true
knowledge = "2024-06-01"
tool_call = true
@@ -3,7 +3,7 @@ family = "command-r"
release_date = "2024-08-30"
last_updated = "2024-08-30"
attachment = false
reasoning = true
reasoning = false
temperature = true
knowledge = "2024-06-01"
tool_call = true
@@ -1,21 +1,19 @@
name = "Llama 3 8B Instruct"
name = "Llama 3.1 70B Turbo"
family = "llama"
release_date = "2024-07-23"
last_updated = "2024-07-23"
attachment = false
reasoning = false
temperature = true
knowledge = "2023-03"
tool_call = false
tool_call = true
open_weights = true
[cost]
input = 0.30
output = 0.60
input = 0.40
output = 0.40
[limit]
context = 8_192
output = 2_048
context = 131_072
output = 16_384
[modalities]
input = ["text"]
@@ -0,0 +1,20 @@
name = "Llama 3.1 70B"
family = "llama"
release_date = "2024-07-23"
last_updated = "2024-07-23"
attachment = false
reasoning = false
tool_call = true
open_weights = true
[cost]
input = 0.40
output = 0.40
[limit]
context = 131_072
output = 16_384
[modalities]
input = ["text"]
output = ["text"]
@@ -0,0 +1,20 @@
name = "Llama 3.1 8B Turbo"
family = "llama"
release_date = "2024-07-23"
last_updated = "2024-07-23"
attachment = false
reasoning = false
tool_call = true
open_weights = true
[cost]
input = 0.02
output = 0.03
[limit]
context = 131_072
output = 16_384
[modalities]
input = ["text"]
output = ["text"]
@@ -0,0 +1,20 @@
name = "Llama 3.1 8B"
family = "llama"
release_date = "2024-07-23"
last_updated = "2024-07-23"
attachment = false
reasoning = false
tool_call = true
open_weights = true
[cost]
input = 0.02
output = 0.05
[limit]
context = 131_072
output = 16_384
[modalities]
input = ["text"]
output = ["text"]
@@ -1,19 +1,19 @@
name = "Meta: Llama 3.3 70B Instruct (free)"
name = "Llama 3.3 70B Turbo"
family = "llama"
release_date = "2024-12-06"
last_updated = "2024-12-06"
attachment = false
reasoning = false
temperature = true
tool_call = true
open_weights = true
[cost]
input = 0
output = 0
input = 0.10
output = 0.32
[limit]
context = 128000
output = 128000
context = 131_072
output = 16_384
[modalities]
input = ["text"]
@@ -0,0 +1,20 @@
name = "Llama 4 Maverick 17B FP8"
family = "llama"
release_date = "2025-04-05"
last_updated = "2025-04-05"
attachment = false
reasoning = false
tool_call = true
open_weights = true
[cost]
input = 0.15
output = 0.60
[limit]
context = 1_000_000
output = 16_384
[modalities]
input = ["text", "image"]
output = ["text"]
@@ -0,0 +1,20 @@
name = "Llama 4 Scout 17B"
family = "llama"
release_date = "2025-04-05"
last_updated = "2025-04-05"
attachment = false
reasoning = false
tool_call = true
open_weights = true
[cost]
input = 0.08
output = 0.30
[limit]
context = 10_000_000
output = 16_384
[modalities]
input = ["text", "image"]
output = ["text"]
+1
View File
@@ -0,0 +1 @@
<svg xmlns="http://www.w3.org/2000/svg" width="24" height="24" fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="2" class="lucide lucide-terminal h-5 w-5 text-primary"><path d="m4 17 6-6-6-6M12 19h8"/></svg>

After

Width:  |  Height:  |  Size: 253 B

+25
View File
@@ -0,0 +1,25 @@
name = "GLM-4.7"
family = "glm"
release_date = "2025-12"
last_updated = "2025-12"
attachment = false
reasoning = true
temperature = true
tool_call = true
knowledge = "2025-04"
open_weights = true
[interleaved]
field = "reasoning_content"
[cost]
input = 0.45
output = 1.65
[limit]
context = 200_000
output = 128_000
[modalities]
input = ["text"]
output = ["text"]
+25
View File
@@ -0,0 +1,25 @@
name = "GLM-5"
family = "glm"
release_date = "2026-02"
last_updated = "2026-02"
attachment = false
reasoning = true
temperature = true
tool_call = true
knowledge = "2025-04"
open_weights = true
[interleaved]
field = "reasoning_content"
[cost]
input = 0.75
output = 2.40
[limit]
context = 200_000
output = 128_000
[modalities]
input = ["text"]
output = ["text"]
@@ -1,6 +1,6 @@
name = "MoonshotAI: Kimi K2 0905 (exacto)"
release_date = "2025-09-05"
last_updated = "2025-09-05"
name = "GPT OSS 120B"
release_date = "2025-08"
last_updated = "2025-08"
attachment = false
reasoning = false
temperature = true
@@ -8,12 +8,12 @@ tool_call = true
open_weights = true
[cost]
input = 0.6
output = 2.5
input = 0.0675
output = 0.27
[limit]
context = 262144
output = 52429
context = 131_072
output = 32_768
[modalities]
input = ["text"]
+5
View File
@@ -0,0 +1,5 @@
name = "DInference"
env = ["DINFERENCE_API_KEY"]
npm = "@ai-sdk/openai-compatible"
doc = "https://dinference.com"
api = "https://api.dinference.com/v1"
+2 -2
View File
@@ -2,5 +2,5 @@ name = "Firmware"
# Token for Firmware's OpenAI-compatible proxy
env = ["FIRMWARE_API_KEY"]
npm = "@ai-sdk/openai-compatible"
api = "https://app.firmware.ai/api/v1"
doc = "https://docs.firmware.ai"
api = "https://app.frogbot.ai/api/v1"
doc = "https://docs.frogbot.ai"
@@ -14,8 +14,9 @@ input = 0
output = 0
[limit]
context = 128_000
context = 144_000
output = 32_000
input = 128_000
[modalities]
input = ["text", "image"]
@@ -14,8 +14,9 @@ input = 0
output = 0
[limit]
context = 128_000
context = 160_000
output = 32_000
input = 128_000
[modalities]
input = ["text", "image"]
@@ -14,8 +14,9 @@ input = 0
output = 0
[limit]
context = 128_000
context = 144_000
output = 64_000
input = 128_000
[modalities]
input = ["text", "image"]
@@ -14,8 +14,9 @@ input = 0
output = 0
[limit]
context = 128_000
context = 144_000
output = 32_000
input = 128_000
[modalities]
input = ["text", "image"]
@@ -13,8 +13,9 @@ input = 0
output = 0
[limit]
context = 128_000
context = 200_000
output = 32_000
input = 128_000
[modalities]
input = ["text", "image"]
@@ -14,8 +14,9 @@ input = 0
output = 0
[limit]
context = 128_000
context = 216_000
output = 16_000
input = 128_000
[modalities]
input = ["text", "image"]
@@ -16,6 +16,7 @@ output = 0
[limit]
context = 128_000
output = 64_000
input = 128_000
[modalities]
input = ["text", "image", "audio", "video"]
@@ -17,6 +17,7 @@ output = 0
[limit]
context = 128_000
output = 64_000
input = 128_000
[modalities]
input = ["text", "image", "audio", "video"]
@@ -17,6 +17,7 @@ output = 0
[limit]
context = 128_000
output = 64_000
input = 128_000
[modalities]
input = ["text", "image", "audio", "video"]
@@ -17,6 +17,7 @@ output = 0
[limit]
context = 128_000
output = 64_000
input = 128_000
[modalities]
input = ["text", "image"]
+2 -1
View File
@@ -14,8 +14,9 @@ input = 0
output = 0
[limit]
context = 64_000
context = 128_000
output = 16_384
input = 64_000
[modalities]
input = ["text", "image"]

Some files were not shown because too many files have changed in this diff Show More