chenxiao5580-cmd 95cf7bc77c fix(modelis): declare reasoning_options per model from measurements (#3951)
* fix(modelis): declare reasoning_options per model from measurements

Follow-up to #3932. That PR landed with the same six-value effort list on
all nine models; the review bot was right that this is over-broad, and
re-measuring showed it is also incomplete.

Measured one control at a time against the live endpoint:

- effort kept only where the levels measurably change reasoning
  (Claude x3, Gemini x2). Dropped on both DeepSeek and both Qwen models,
  which accept every value and return 200 but do not change behaviour.
- toggle added where both states are caller-reachable. The mechanism
  differs by family: reasoning.enabled for Claude/Gemini/Qwen, and
  reasoning_effort "none" for DeepSeek, which ignores reasoning.enabled.
- budget_tokens added where reasoning_tokens tracks the requested budget
  (Gemini x2, Qwen x2). No min/max, since no boundary was probed.
- claude-fable-5 and gemini-2.5-pro reject disabling with a 400, so
  neither declares a toggle.

Also drops the header comment that claimed all six effort values were
reflected in reasoning_tokens: that holds for five models, not nine.

Costs are unchanged and re-verified against the live pricing endpoint.

* fix(modelis): move wire-path comments to a leading header block

Review finding: every declared control needs its exact request syntax in a
leading top-of-file comment, not an inline one next to the option.

I had put them inline because Modelis has no sync module, so nothing would
strip mid-file comments today. That was the wrong call: the sync rewrites
provider TOMLs by parsing and re-serializing them and keeps only a leading
header, so an inline comment is one sync module away from vanishing with
nobody noticing.

Each file now opens with the wire path for every control it declares.

* fix(modelis): narrow effort values to measured separable levels

Review finding: the six-value lists were the gateway's global accept-set
minus none, not per-model truth.

Re-measured at three task difficulties, asking which ADJACENT levels are
actually distinguishable (sample ranges that do not overlap):

- minimal collapses into low on every Claude model at every difficulty
  -> dropped from all three, as the lab baseline predicted.
- xhigh never rises above high on opus, sonnet or gemini-2.5-flash
  -> dropped there; kept on fable, where it does separate.
- gemini-2.5-flash keeps minimal: 37 vs 107 with zero scatter across
  three repeats.
- claude-fable-5 returns 145 reasoning tokens at reasoning_effort none,
  so it has no off switch at all and declares neither toggle nor none.

Per-file: opus/sonnet/gemini-2.5-pro low|medium|high|max, fable
low|medium|high|xhigh|max, gemini-2.5-flash minimal|low|medium|high|max.

DeepSeek and Qwen still declare no effort list: repeats at one setting
scatter up to 5x and the ordering inverts at medium on both DeepSeek
models. Numbers are in the PR discussion.

* fix(modelis): effort-none authored as effort; restore lab-baseline levels

Review findings:

1. Off via reasoning_effort "none" must be authored as effort with none
   in values, not as toggle. Both DeepSeek files had a toggle declaration
   whose own wire comment named the effort parameter -- self-contradicting.
   They now declare effort = [none, high, max] per the peer set.
   Qwen keeps toggle because there the mechanism really is a separate
   field: reasoning.enabled false -> 0, while reasoning_effort none
   leaves those models reasoning unchanged.

2. Dropping a level because adjacent reasoning_tokens ranges overlapped
   was the wrong test -- a level can differ in latency or quality without
   differing in thinking tokens. Reverted to the lab/peer baseline and
   restored xhigh on claude-opus-4-8.

minimal stays dropped on the Claude models: it is absent from the lab
baseline and returned output identical to low at every difficulty tested.
2026-08-02 10:57:31 -05:00
2026-07-03 14:10:05 +02:00
2026-07-06 13:13:11 +02:00
2026-06-02 14:24:55 -04:00

Models.dev logo


Models.dev is a comprehensive open-source database of AI model specifications, pricing, and capabilities.

There's no single database with information about all the available AI models. We started Models.dev as a community-contributed project to address this. We also use it internally in opencode.

API

You can access this data through an API.

curl https://models.dev/api.json

Use the Model ID field to do a lookup on any model; it's the identifier used by AI SDK.

Provider-agnostic model metadata is available separately:

curl https://models.dev/models.json

Use this for facts about the model itself, independent of where it is served. If you need both provider endpoints and model-only metadata in one response:

curl https://models.dev/catalog.json

Logos

Provider logos are available as SVG files:

curl https://models.dev/logos/{provider}.svg

Replace {provider} with the Provider ID (e.g., anthropic, openai, google). If we don't have a provider's logo, a default logo is served instead.

Contributing

The data is stored in the repo as TOML files; organized by provider and model. The logo is stored as an SVG. This is used to generate this page and power the API.

We need your help keeping the data up to date.

Adding Model Metadata

Model-only facts live in models/, using the same path-style IDs as provider models. For example, models/openai/gpt-5.toml defines metadata for the underlying GPT-5 model, while providers/openai/models/gpt-5.toml defines OpenAI-specific serving details such as pricing.

Use model metadata for provider-agnostic facts:

  • name, family, release_date, last_updated, knowledge
  • attachment, reasoning, tool_call, structured_output, temperature
  • [limit] defaults like context, input, and output token limits
  • [modalities] defaults
  • open_weights, license, links, weights, and benchmarks

Example:

name = "GPT-5"
family = "gpt"
release_date = "2025-08-07"
last_updated = "2025-08-07"
attachment = true
reasoning = true
temperature = false
tool_call = true
structured_output = true
open_weights = false

[limit]
context = 400_000
input = 272_000
output = 128_000

[modalities]
input = ["text", "image"]
output = ["text"]

[[benchmarks]]
name = "Benchmark Name"
score = 72.5
metric = "accuracy"
source = "https://example.com/results"

[[weights]]
label = "Model weights"
url = "https://huggingface.co/example/model"
format = "safetensors"

Provider TOMLs can inherit these facts with base_model and then keep only provider-specific fields or overrides:

base_model = "openai/gpt-5"

[cost]
input = 1.25
output = 10.00
cache_read = 0.125

[limit]
context = 200_000 # optional provider override
output = 32_000

Provider fields win over model metadata during generation. Use this when the underlying model is the same but a provider serves it with different context limits, modalities, features, or pricing.

Adding a New Provider Model

To add a new model, start by checking if the provider already exists in the providers/ directory. If not, then:

1. Create a Provider

If the provider isn't already in providers/:

  1. Create a new folder in providers/ with the provider's ID. For example, providers/newprovider/.

  2. Add a provider.toml with the provider details:

    name = "Provider Name"
    npm = "@ai-sdk/provider" # AI SDK Package name
    env = ["PROVIDER_API_KEY"] # Environment Variable keys used for auth
    doc = "https://example.com/docs/models" # Link to provider's documentation
    

    If the provider doesnt publish an npm package but exposes an OpenAI-compatible endpoint, set the npm field accordingly and include the base URL:

    npm = "@ai-sdk/openai-compatible" # Use OpenAI-compatible SDK
    api = "https://api.example.com/v1" # Required with openai-compatible
    

2. Add a Logo (required for new providers)

To add a logo for the provider:

  1. Add a logo.svg file to the provider's directory (e.g., providers/newprovider/logo.svg)
  2. Use SVG format with no fixed size or colors - use currentColor for fills/strokes

Example SVG structure:

<svg xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24" fill="currentColor">
  <!-- Logo paths here -->
</svg>

3. Add a Model Definition

Create a new TOML file in the provider's models/ directory where the filename is the model ID.

If the model ID contains /, use subfolders. For example, for the model ID openai/gpt-5, create a folder openai/ and place a file named gpt-5.toml inside it.

name = "Model Display Name"
attachment = true           # or false - supports file attachments
reasoning = false           # or true - supports reasoning / chain-of-thought
tool_call = true            # or false - supports tool calling
structured_output = true    # or false - supports a dedicated structured output feature
temperature = true          # or false - supports temperature control
knowledge = "2024-04"       # Knowledge-cutoff date
release_date = "2025-02-19" # First public release date
last_updated = "2025-02-19" # Most recent update date
open_weights = true         # or false  - models trained weights are publicly available

[cost]
input = 3.00                # Cost per million input tokens (USD)
output = 15.00              # Cost per million output tokens (USD)
reasoning = 15.00           # Cost per million reasoning tokens (USD)
cache_read = 0.30           # Cost per million cached read tokens (USD)
cache_write = 3.75          # Cost per million cached write tokens (USD)
input_audio = 1.00          # Cost per million audio input tokens (USD)
output_audio = 10.00        # Cost per million audio output tokens (USD)

[limit]
context = 400_000           # Maximum context window (tokens)
input = 272_000             # Maximum input tokens
output = 8_192              # Maximum output tokens

[modalities]
input = ["text", "image"]   # Supported input modalities
output = ["text"]           # Supported output modalities

[interleaved]
field = "reasoning_content" # Name of the interleaved field "reasoning_content" or "reasoning_details"

3a. Reuse Model Metadata with base_model

For wrapper providers that mirror an existing model, prefer referencing the model-only metadata instead of duplicating provider-agnostic fields.

Use base_model when the provider serves the same underlying model and only provider-specific fields differ.

base_model = "anthropic/claude-opus-4-6"
# Match lab/peer controls for this model (not a stripped L/M/H guess)
reasoning_options = [
  { type = "effort", values = ["low", "medium", "high", "max"] },
  { type = "budget_tokens", min = 1_024 },
]

[cost]
input = 5.00
output = 25.00

Rules:

  • base_model must point to a TOML file in models/ using <provider>/<model-id>.
  • Override-only: after base_model, write only provider-specific fields and values that differ from the base. Do not restate the same description, structured_output, modalities, tool_call, dates, etc.
  • You may override any top-level model field when the provider actually differs.
  • If you override a nested table like [cost], [limit], or [modalities], include the full values needed for that table (arrays/primitives replace; plain objects deep-merge).
  • base_model_omit is optional and removes inherited model metadata fields after local overrides are merged. Use dot-path strings, for example base_model_omit = ["limit.input"].
  • Provider-specific fields (cost, reasoning_options, interleaved, status, provider, experimental) belong on the provider model when needed.
  • id still comes from the filename; do not add it to the TOML.

Reasoning options (short): classify first-party lab vs multi-model relay (not by npm). Copy the underlying models controls from the lab entry and same-surface peers — often low/medium/high on GPT-style relays, but DeepSeek V4 is toggle+high/max, etc. Do not use [] from uncertainty on relays. Full policy: AGENTS.md.

Use base_model when the wrapper model is materially the same as the source model and only differs by provider-specific pricing, limits, modalities, provider request shape, or lifecycle flags.

Sync and generator scripts should preserve existing base_model / base_model_omit fields when updating provider TOMLs. Do not use legacy [extends] tables.

4. Submit a Pull Request

  1. Fork this repo
  2. Create a new branch with your changes
  3. Add your provider and/or model files
  4. Open a PR with a clear description

Validation

There's a GitHub Action that will automatically validate your submission against our schema to ensure:

  • All required fields are present
  • Data types are correct
  • Values are within acceptable ranges
  • TOML syntax is valid

When moving existing provider fields into model metadata, compare generated output before and after the change:

bun run compare:migrations

This prints a diff for each changed model TOML so you can confirm the generated JSON only changed where you intended.

Schema Reference

Models must conform to the following schema, as defined in packages/core/src/schema.ts.

Provider Schema:

  • name: String - Display name of the provider
  • npm: String - AI SDK Package name
  • env: String[] - Environment variable keys used for auth
  • doc: String - Link to the provider's documentation
  • api (optional): String - OpenAI-compatible API endpoint. Required only when using @ai-sdk/openai-compatible as the npm package

Model Schema:

  • name: String — Display name of the model
  • attachment: Boolean — Supports file attachments
  • reasoning: Boolean — Supports reasoning / chain-of-thought
  • tool_call: Boolean - Supports tool calling
  • structured_output (optional): Boolean — Supports structured output feature
  • temperature (optional): Boolean — Supports temperature control
  • knowledge (optional): String — Knowledge-cutoff date in YYYY-MM or YYYY-MM-DD format
  • release_date: String — First public release date in YYYY-MM or YYYY-MM-DD
  • last_updated: String — Most recent update date in YYYY-MM or YYYY-MM-DD
  • open_weights: Boolean - Indicate the model's trained weights are publicly available
  • interleaved (optional): Boolean or Object — Supports interleaved reasoning. Use true for general support or an object with field to specify the format
  • interleaved.field: String — Name of the interleaved field ("reasoning_content" or "reasoning_details")
  • cost.input: Number — Cost per million input tokens (USD)
  • cost.output: Number — Cost per million output tokens (USD)
  • cost.reasoning (optional): Number — Cost per million reasoning tokens (USD)
  • cost.cache_read (optional): Number — Cost per million cached read tokens (USD)
  • cost.cache_write (optional): Number — Cost per million cached write tokens (USD)
  • cost.input_audio (optional): Number — Cost per million audio input tokens, if billed separately (USD)
  • cost.output_audio (optional): Number — Cost per million audio output tokens, if billed separately (USD)
  • limit.context: Number — Maximum context window (tokens)
  • limit.input: Number — Maximum input tokens
  • limit.output: Number — Maximum output tokens
  • modalities.input: Array of strings — Supported input modalities (e.g., ["text", "image", "audio", "video", "pdf"])
  • modalities.output: Array of strings — Supported output modalities (e.g., ["text"])
  • status (optional): String — Supported status:
    • alpha - Indicate the model is in alpha testing
    • beta - Indicate the model is in beta testing
    • deprecated - Indicate the model is no longer served by the provider's public API

Examples

See existing providers in the providers/ directory for reference:

  • providers/anthropic/ - Anthropic Claude models
  • providers/openai/ - OpenAI GPT models
  • providers/google/ - Google Gemini models

Working on frontend

Make sure you have Bun installed.

$ bun install
$ cd packages/web
$ bun run dev

And it'll open the frontend at http://localhost:3000

Manual testing with opencode

You can manually check provider changes with opencode by:

$ bun install
$ cd packages/web
$ bun run build
$ OPENCODE_MODELS_PATH="dist/_api.json" opencode

Questions?

Open an issue if you need help or have questions about contributing.


Models.dev is created by the maintainers of SST.

Join our community Discord | YouTube | X.com

S
Description
开源 AI 模型数据库|GitHub 镜像 6.6k · 🍴 1.5k
https://github.com/anomalyco/models.dev Readme MIT 20 MiB
Languages
TypeScript 95%
CSS 2.4%
Shell 2.4%
HTML 0.2%