Skip to content

fix(pricing): one rate lookup for every path; no double normalization on reload - #642

Merged
lis186 merged 1 commit into
mainfrom
fix/pricing-single-source-pr
Sep 28, 2026
Merged

lis186 merged 1 commit into
mainfrom
fix/pricing-single-source-pr

Conversation

@lis186

@lis186 lis186 commented Sep 28, 2026

Copy link
Copy Markdown
Owner

摘要

統一 ccxray 各條路徑的計價。

  • 原本的問題: 同一筆 Codex(gpt-6-astra)用量,proxy、CLI 匯入、伺服器重新載入會算出三個金額:$0.18609、$0.0558、$0.1405。重新載入時連 token 數都被扣錯。
  • 原因: 三個問題疊在一起。
    • 程式裡有兩套價格查詢,匯入器和 cost-worker 那套只看內建價格表。
    • 匯入的 Codex 用量在重新載入時被重複扣掉快取 token。
    • 啟動時重新計價和價格表載入之間有競態。
  • 修正後:
    • 每條路徑都經過同一個 lookupRates,同一筆用量得到同一個金額和同一個可信度標示。
    • 查不到價格的模型一律標 unknown,不再用替代價格估一個可能錯三倍的數字。
    • 讀取時會重算 fallback 和 unknown 的紀錄,已經是 exact 的不動。

這是 agentflow 整合規格(#638)A1a 的前置條件。

Details

  • Root causes (reproduced on 685321f, with an isolated CCXRAY_HOME/HOME, a copy of the same pricing-cache.json, and network blocked):
    1. Two rate lookups. Proxy, reload, cold load and rebuild-index used pricing.calculateCost (built-ins plus the LiteLLM cache). The importer and cost-worker used default-rates.calculateCostSimple: built-ins only, with claude-sonnet-4 rates as a fallback for anything unknown. gpt-6-astra exists only in the cache, so imports were priced at about 1/3 of the real rate.
    2. Double normalisation. The Codex importer writes cache-exclusive usage but set no _ccxrayUsageNormalized marker. restore.js and cold load (routes/api.js) then subtracted cached tokens a second time: input 17,695 → 10,655 and 225 → 0, repriced to $0.1405.
    3. Startup race. Restore started before fetchPricing() finished, so reload repricing depended on how far the table had loaded.
  • Fix:
    • lookupRates(model, provider) is the only rate lookup. Its table is built by one buildRateTable (built-ins, the pricing cache, xai/ key mirroring, lag overrides) and loads the cache synchronously on first use. There is no startup wait and no race.
    • Every path uses it: proxy HTTP (Anthropic and OpenAI), WebSocket, Claude and Codex import (import --once / --target-transcript), the cost-worker child, rebuild-index, restore and cold load.
    • A model with no rate is cost: null / unknown on every path, shown as a lower bound per ADR 0017. The forward.js live-path invariant comment is updated to match.
    • The importer marks Codex usage as normalized. Reads treat existing imported OpenAI rows as already normalized, which also repairs rows written before this change.
    • fallback and unknown rows are repriced on read when a rate is now available; exact rows are untouched. A PRICING_REV stamp makes sessions.json rebuild once on the next start.
  • Test isolation: pricing now reads the package-relative pricing-cache.json. test/default-rates.test.js and test/grok-wire.test.js point the cache at a missing path, and docs/testing.md documents this.

Verification

  • CCXRAY_HOME=$(mktemp -d) CCXRAY_EXPORT_DISABLE=1 npm test: 2683/2683 (run again independently before opening this PR). The real ~/.ccxray was untouched during the run.
  • New test/pricing-single-source.test.js (31 cases) covers every path listed above, plus repricing, the revision stamp and the no-cache case.
    • Fail-on-old: the new file imports a test-only hook, so it fails wholesale on the old code, which is not behavioural evidence. With a no-op shim for that hook on 685321f, the 26 behaviour-changing cases fail and the invariance checks pass. With the fix, all pass.
    • The 3 regression tests added after the first acceptance review fail 3/3 on the pre-review commit.
  • End-to-end check on real data: the same Codex rollout (gpt-6-astra, 2 turns), imported with import --target-transcript against the same pricing cache.
    • Old code: $0.055827 and $0.0081228, fallback.
    • New code: $0.18609 and $0.027076, exact, matching what the proxy recorded for the same turns.
    • Token counts are unchanged.
  • Claude figures unchanged: the re-verification sessions give identical results on old and new code across proxy, import and reload (claude -p worker $0.3473501 via proxy; interactive host $0.7334226 via proxy).
  • Independent review (agentflow acceptance by codex, a different model family):
    • The first pass FAILED. The first-use table lacked xai/ key mirroring, so import and proxy could price xai/-prefixed models differently. This is fixed here with a shared table builder and regression tests.
    • The second pass: Outcome, Minimality and Conformance all PASS.

Not verified:

  • sessions.json totals refresh on the next rebuild, not immediately, when the pricing cache first learns a model; per-entry views reprice at once.
  • F4 (minor, follow-up): with an expired cache and no network, the Usage page's model list falls back to built-ins. Per-path pricing is unaffected.
  • CLI coverage: CLI import --once / --target-transcript are covered in tests through their shared import function; the real CLI was exercised by the end-to-end check above.

🤖 Generated with Claude Code

… on reload

The same Codex turn priced three ways: the proxy used the LiteLLM table
($0.18609 exact), the importer and cost-worker used only DEFAULT_PRICING and
substituted claude-sonnet-4 rates for unknown models ($0.055827 fallback), and
reload/cold-load subtracted cached tokens a second time from imported Codex
usage that lacked the normalized marker ($0.1405 in total).

- lookupRates (default-rates.js) is the only rate lookup, over one table built
  by one pure builder, buildRateTable (xai/ mirror, DEFAULT_PRICING floor, lag
  overrides). The table is read synchronously from pricing-cache.json on first
  use and replaced when fetchPricing lands; buildPricingTable is a thin
  wrapper around the same builder. calculateCost and calculateCostSimple both
  use it, and so do the proxy, WebSocket, importer, import --once and
  --target-transcript, cost-worker, rebuild-index, restore and cold-load.
  Every path passes the same provider key.
- A model the table does not match is cost null / unknown on every path; the
  substitute-rate fallback is removed.
- The Codex importer marks its cache-exclusive usage as normalized; reload and
  cold-load treat existing imported openai lines the same way.
- Reload and cold-load reprice fallback/unknown costs with the current lookup
  (exact, prefix and legacy numeric costs are untouched; nothing is rewritten
  on disk), and session aggregates carry PRICING_REV so sessions.json rebuilds
  once.
- Tests pin pricing-cache.json (docs/testing.md), isolate their homes, and
  cover every path with fail-on-685321f / pass-on-fix evidence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@lis186
lis186 merged commit 15fbbff into main Sep 28, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant