fix(core): read cache stats when streaming through providers like OpenRouter - #82
Merged
Merged
Conversation
ChatCompletionAccumulator.AddChunk only copies CompletionTokens, PromptTokens, and TotalTokens field by field; it never carries over PromptTokensDetails or the chunk's raw JSON. That made accumulator.Usage.PromptTokensDetails.CachedTokens always read 0, so every cache-hit ratio we recorded for a session was 0 regardless of provider. Anthropic models routed through OpenRouter compound this: their usage payload reports cache hits as a top-level cache_read_input_tokens field, which the OpenAI-shaped PromptTokensDetails struct has no slot for even when the accumulator bug is fixed, so nothing ever populated it for Claude. chatStreamRound now inspects each chunk's Usage as it arrives (before the accumulator drops the detail) and falls back to cache_read_input_tokens via gjson when the OpenAI field is empty, then returns the resolved cached-token count alongside the accumulator so SendChatMessage can pass the real number to SessionUsageSave.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
ChatCompletionAccumulator.AddChunkin the openai-go SDK only copiesCompletionTokens,PromptTokens, andTotalTokensfield by field when merging stream chunks — it never carries overPromptTokensDetailsor the chunk's raw JSON. That madeaccumulator.Usage.PromptTokensDetails.CachedTokensalways read 0, so every cache-hit ratio we recorded for a session was 0 regardless of provider.Anthropic models routed through OpenRouter compound this: their usage payload reports cache hits via a top-level
cache_read_input_tokensfield, which the OpenAI-shapedPromptTokensDetailsstruct has no slot for even once the accumulator bug is worked around, so nothing ever populated it for Claude.chatStreamRoundnow inspects each chunk'sUsageas it arrives (before the accumulator drops the detail), preferring the OpenAI field and falling back tocache_read_input_tokensviagjsonwhen it's empty, and returns the resolved cached-token count alongside the accumulator soSendChatMessagecan pass the real number toSessionUsageSave.How it was verified
make test-racepasses. Addedcore/chat_cache_usage_test.gocovering the OpenAI field, the Anthropic fallback field, and the no-cache-data case.make test-racepassesChecklist
:=, onevarblock per function in first-use order witherrlast, callees before callers, andmainunconditionally last.gofiles carry the two-line SPDX headerGPL-3.0-only, matching the project