发布

  • [issue-6976] [BE] fix: discount Gemini/Google cached tokens in cost calculation (#6980)

    frostbyte_neo 发布于 2026-06-10 09:13:24 +00:00

    Register the raw LiteLLM providers vertex_ai-language-models and gemini in PROVIDERS_CACHE_COST_CALCULATOR and add a Google cache handler so Gemini context-cache tokens are billed at the cache-read rate instead of the full input rate. Fixes #6976.

    Co-authored-by: Andres Cruz andresc@comet.com

    下载附件