Skip to content

fix(core): read cache stats when streaming through providers like OpenRouter - #82

Merged
devproje merged 1 commit into
masterfrom
fix/provider-cache-usage
Sep 24, 2026
Merged

devproje merged 1 commit into
masterfrom
fix/provider-cache-usage

Conversation

@devproje

Copy link
Copy Markdown
Owner

What this changes

ChatCompletionAccumulator.AddChunk in the openai-go SDK only copies CompletionTokens, PromptTokens, and TotalTokens field by field when merging stream chunks — it never carries over PromptTokensDetails or the chunk's raw JSON. That made accumulator.Usage.PromptTokensDetails.CachedTokens always read 0, so every cache-hit ratio we recorded for a session was 0 regardless of provider.

Anthropic models routed through OpenRouter compound this: their usage payload reports cache hits via a top-level cache_read_input_tokens field, which the OpenAI-shaped PromptTokensDetails struct has no slot for even once the accumulator bug is worked around, so nothing ever populated it for Claude.

chatStreamRound now inspects each chunk's Usage as it arrives (before the accumulator drops the detail), preferring the OpenAI field and falling back to cache_read_input_tokens via gjson when it's empty, and returns the resolved cached-token count alongside the accumulator so SendChatMessage can pass the real number to SessionUsageSave.

How it was verified

make test-race passes. Added core/chat_cache_usage_test.go covering the OpenAI field, the Anthropic fallback field, and the no-cache-data case.

  • make test-race passes
  • New behaviour has a test, or there is nothing to test

Checklist

  • Follows docs/CONVENTION.md: no comments, no :=, one var block per function in first-use order with err last, callees before callers, and main unconditionally last
  • New .go files carry the two-line SPDX header
  • Documentation updated if behaviour a user can see has changed
  • My contribution is licensed GPL-3.0-only, matching the project

ChatCompletionAccumulator.AddChunk only copies CompletionTokens,
PromptTokens, and TotalTokens field by field; it never carries over
PromptTokensDetails or the chunk's raw JSON. That made
accumulator.Usage.PromptTokensDetails.CachedTokens always read 0, so
every cache-hit ratio we recorded for a session was 0 regardless of
provider. Anthropic models routed through OpenRouter compound this:
their usage payload reports cache hits as a top-level
cache_read_input_tokens field, which the OpenAI-shaped
PromptTokensDetails struct has no slot for even when the accumulator
bug is fixed, so nothing ever populated it for Claude.

chatStreamRound now inspects each chunk's Usage as it arrives (before
the accumulator drops the detail) and falls back to
cache_read_input_tokens via gjson when the OpenAI field is empty, then
returns the resolved cached-token count alongside the accumulator so
SendChatMessage can pass the real number to SessionUsageSave.
@devproje
devproje merged commit aecbb99 into master Sep 24, 2026
3 checks passed
@devproje
devproje deleted the fix/provider-cache-usage branch September 24, 2026 10:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant