Skip to content

fix: subagent import, Claude/Codex proxy cwd, 1h cache pricing, and indexed effort - #637

Merged
lis186 merged 7 commits into
mainfrom
fix/agentflow-attribution-accuracy
Sep 27, 2026
Merged

lis186 merged 7 commits into
mainfrom
fix/agentflow-attribution-accuracy

Conversation

@lis186

@lis186 lis186 commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

摘要

修正 ccxray 在六個地方少算或算錯用量,另加一項強化,共 7 個 commit,每個都附回歸測試:

  • 匯入漏掉 Claude 子代理 transcript:只靠匯入時,成本少算約 37%。
  • Claude Code 2.1.283 與 Codex 經 proxy 時 cwd 是 null:影響所有新版 Claude Code 流量的專案歸屬。
  • 1 小時快取寫入誤用 5 分鐘價格:Claude Code 的主對話都用 1 小時快取,所以成本少算。
  • 匯入 id 只有 10 毫秒精度:同一個 10 毫秒內寫入的紀錄會被誤判為重複而丟掉。
  • effort、thinking token、turn 耗時沒有寫進索引:_req.json 14 天後就會被刪,事後無法補。

修正後,實驗 session 的匯入與 proxy 兩條路徑都算出 $0.4030201,與 Claude Code 自己回報的成本完全一致。起因是設計 agentflow 的每個 Ask 成本報表,但除了 effort 欄位以外,其他修正都會影響一般使用者。

Details

  • Subagent transcripts were never imported (b79473a). collectJsonlFiles listed only projects/<slug>/*.jsonl and skipped <sid>/subagents/agent-*.jsonl. Subagent turns now import under the parent session with isSubagent, agentKey/agentLabel taken from the .meta.json agentType, plus new subagentId and subagentToolUseId fields. The deployment-identity fields agentId/agentType (feat(index): optional identity fields + local date/tz + config-dir extraction #504) are left alone. Dedup goes through msg.id, the responseId (ADR 0012).
  • Import ids used 10 ms resolution (d651552), so two turns written in the same 10 ms were treated as duplicates and one was dropped. Ids now keep full milliseconds. Dedup asks "is this turn already indexed" (by responseId, or by a same-session entry under the new id or the legacy 10 ms id), so a rescan does not rewrite existing imports. A genuinely different turn gets a -1/-2 suffix instead of being dropped.
  • Proxy cwd was null on Claude Code 2.1.283 (b2e0d22). 2.1.283 moved the environment block into messages[1] with role: system, and later system messages can be plain strings. extractCwd now scans messages[0] and every system-role message. As a last resort it reads safeguards[].classifier_context.live_cwd. It never scans other user messages, so pasted text cannot decide the project. This affected project attribution for all proxied traffic from current Claude Code.
  • 1-hour cache writes were priced at the 5-minute rate (3e41b3a). When usage.cache_creation has the TTL split, 5-minute writes are now priced at 1.25× input and 1-hour writes at 2× input. The 1-hour rate uses LiteLLM's cache_creation_input_token_cost_above_1hr when available. Both proxy (calculateCost) and import/cost-worker (calculateCostSimple) paths are fixed. A non-numeric or negative split falls back to the old calculation.
  • Proxy cwd was null for Codex (e131be2). Codex sends cwd in an input[] user message that starts with <environment_context><cwd>…, and none of the existing sources read it. It is now the last fallback. Only a block at the start of the text counts, because Codex memory-consolidation prompts quote older rollouts' environment_context, and taking those would attribute the request to the wrong project.
  • Effort, thinking tokens and turn duration were not indexed (0fa5cbe). New index fields effort, thinkingTokens and turnDurationMs are written at ingest time, so they survive the 14-day _req.json retention.
    • Import sources: Claude perTurnEffort or effort, output_tokens_details.thinking_tokens, and turn_duration rows; Codex turn_context.effort.
    • Proxy sources: Claude output_config.effort; Codex reasoning.effort, then metadata.reasoning_effort, then the same session's latest prewarm.
    • This commit also fixes the WebSocket path, which used to drop reasoning from response.create.
  • Prewarm map hardening (96c7c53). The Codex prewarm → effort map now expires entries after 30 idle minutes (sliding, because Codex sends only one prewarm per session). It refuses effort/model values over 128 chars instead of truncating them, since a truncated model name would misprice. The existing 500-entry cap now evicts least-recently-used.

These are one PR rather than six because they share one goal (correct per-session attribution and cost across import and proxy) and several touch the same functions. Each commit stands alone if you'd rather split it.

Verification

  • CCXRAY_HOME=$(mktemp -d) CCXRAY_EXPORT_DISABLE=1 npm test passes: 2631/2631.
  • Evidence the change does what it claims:
    • Fail-on-old / pass-on-new: the 9 changed test files (309 tests) run against origin/main code fail 49; against this branch all 309 pass.
    • Real-traffic check: a Claude Code 2.1.283 claude -p --effort low session with one Task subagent was recorded through an isolated proxy and also imported from its transcript.
      • Cost: Claude Code's own total_cost_usd is $0.4030201. Before this branch, import gave $0.2043 (subagents missing) and proxy gave $0.3230 (1-hour cache mispriced). Now both paths give $0.4030201.
      • Rows: the import went from 7 to 11 (7 main + 4 subagent).
      • Effort: low on this session, and medium on a second interactive session.
    • Proxy cwd: proxied Claude requests with cwd went from 0/24 to 22/24. The other 2 are title-generation requests, which carry no working directory.
    • Codex cwd: a captured Codex exec request (ChatGPT login, WebSocket) now yields its cwd.
    • Browser smoke: an isolated --port 5602 smoke run loads the dashboard with no console errors, and /_api/entries returns the imported rows.
  • Independent review (agentflow cross-check/acceptance, Claude-family reviewer): Outcome, Minimality and Conformance all PASS. Security scan: no critical or high findings.

Not verified:

  • Dashboard lists: with a home populated only by import --target-transcript, the dashboard shows 0 projects and 0 sessions. origin/main code shows the same with the same data, so this is not a regression, but it means the dashboard rendering of imported subagent turns and the new fields was not checked by eye.
  • Successful Codex turn through the proxy: not observed, because the account hit its usage limit. Codex usage/elapsed recording on a completed WebSocket turn was not re-checked.
  • Proxy-side thinking tokens and turn duration: not recorded. Proxy records effort only; those two come from import.
  • Grok: untouched.
  • Deferred proposals, not in this PR:
    • .meta.json agentType allowlist
    • Codex uses only the first user message for environment_context
    • extractConfigDir scans only messages[0]
    • exposing the 1-hour rate in /_api/pricing

License

  • I agree to license this contribution under PolyForm Noncommercial 1.0.0
    and grant the maintainer the right to relicense it under different terms in
    the future, including commercial licenses
    (see CONTRIBUTING.md) — maintainer-authored; left for the maintainer to tick.
  • My commits are signed off (git commit -s): maintainer-authored, matching existing main history.

🤖 Generated with Claude Code

lis186 and others added 7 commits September 27, 2026 02:45
…ssion

Claude Code writes Task-tool subagent transcripts to
<slug>/<sid>/subagents/agent-*.jsonl, which the importer never read, so an
import-only session lost every subagent turn (37% of the exp8 cost).

Subagent turns now import under the parent sessionId with isSubagent,
agentKey/agentLabel from the sidecar .meta.json agentType, and two new
add-only index fields: subagentId (transcript agentId) and
subagentToolUseId (meta toolUseId). The existing agentId/agentType fields
stay reserved for #504 deployment identity. responseId is unchanged, so
ADR 0012 merge folds them with proxy copies. Targeted import of a parent
transcript also picks up its subagent files, each path-checked.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…urns

tsToId kept only 10 ms of precision and pushImportedEntry dropped any turn
whose id already existed anywhere in the index, so two sessions (or a main
turn and its subagent) writing within the same 10 ms silently lost one
turn; targeted import threw an identity-collision error instead.

Ids now keep all three millisecond digits. Dedup asks whether the same
logical turn was already imported (imported responseId for Claude; same
session holding the new id, its legacy 10 ms id, or a suffixed form for
Codex), so rescans stay idempotent across the id-format change. A
genuinely different turn that shares an id gets a deterministic -N suffix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude Code 2.1.283 sends context_management requests with no top-level
system; the environment block ("Primary working directory: …") moved to a
role:"system" message at messages[1], and later ones fold to a plain
string. extractCwd only scanned messages[0]'s block array, so every proxy
Claude line recorded cwd: null.

extractCwd now scans messages[0] plus role:"system" messages (string or
block content) and falls back to safeguards[].classifier_context.live_cwd
(interactive sessions). Other user messages are never scanned, so quoted
text cannot misattribute a project. exp8 proxy requests: 0/24 -> 22/24
(the two nulls are title-generation requests with no cwd signal).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…minute

Anthropic bills 5-minute cache writes at 1.25x input and 1-hour writes at
2x input. ccxray priced every cache write at the single cache_create
(5-minute) rate, so a Claude Code session that caches its main context
for an hour under-reported cost: exp8 session 33539803 showed $0.323
against Claude Code's total_cost_usd of $0.4030201.

calculateCost and calculateCostSimple now price
usage.cache_creation.ephemeral_5m/1h_input_tokens separately when the
split is present (flat counter otherwise). The 1-hour rate comes from
LiteLLM's cache_creation_input_token_cost_above_1hr when listed, else it
is derived as 2x input, never the 5-minute rate, so old pricing-cache.json
rows and DEFAULT_PRICING work unchanged. The importer keeps the split in
the stored usage. Session 33539803 now sums to $0.4030201 on both the
import and proxy paths; cost-worker gains the split for free.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codex 0.157 carries the working directory only in an input[] user message
that starts with <environment_context><cwd>…</cwd>…; getCodexCwd read
metadata, workspaces, the turn-metadata header, and instructions, so a
Codex process started outside the ccxray launcher (e.g. an agentflow
worker) recorded cwd: null.

getCodexCwd now falls back, after every existing source, to that block
(<cwd>, else the first <workspace_roots><root>). Only a text that starts
with the tag counts: memory-consolidation requests quote past rollouts'
environment_context inside a larger prompt, and that cwd belongs to
another session.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rebuilding per-task cost and time needs effort, thinking tokens and turn
duration, but none reached the index, and _req.json bodies are pruned
after the retention window, so they must be captured at write time.

Three add-only index fields (omitted when null, summarized, fill-if-empty
on merge):
- effort: Claude import perTurnEffort/effort; Codex import turn_context
  effort; Claude proxy output_config.effort; Codex proxy
  reasoning.effort -> metadata.reasoning_effort -> the latest prewarm
  metadata for the same session id. Prewarm and main turn arrive on
  separate WebSocket connections, so a bounded (500) session map carries
  it; the prewarm entry also takes metadata.model when it has none.
- thinkingTokens: Claude import usage.output_tokens_details.thinking_tokens.
- turnDurationMs: Claude import system/turn_duration durationMs, attached
  to the last assistant turn before it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Codex prewarm map carried a session's reasoning effort and model to
later turns on other WebSocket connections, but an entry never expired
(a prewarm could be applied days later to the same session id) and its
values had no length limit, only an entry-count cap.

Entries now expire after 30 idle minutes; each hit refreshes the entry,
so a long active session keeps its effort while an idle one is
forgotten. Effort or model values longer than 128 characters are not
stored (dropped rather than truncated, since a cut model name would
misprice). The 500-entry cap stays, now evicting least recently used.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@lis186
lis186 merged commit fa6db61 into main Sep 27, 2026
3 checks passed
lis186 added a commit that referenced this pull request Sep 27, 2026
Claude Code 2.1.283 still sends a top-level `system` on context_management
requests, but the env block with "Primary working directory:" moved to a
role:'system' message. extractCwd returned null as soon as `system` lacked
the env line, so the live proxy recorded cwd: null on every Claude entry and
session. #637 was validated against _req.json files, which store `system`
only as sysHash, so the saved bodies took the context_management branch and
looked fixed.

extractCwd now falls through to the existing messages / safeguards scan when
a context_management request's `system` has no env line. A `system` that
names the directory still wins. role:'system' messages are read before
messages[0], because on 2.1.283 messages[0] embeds CLAUDE.md and other
<system-reminder> text that can quote an env line of its own.

Subagent classification is unchanged: without context_management a
cwd-less `system` still yields null, so title-generation requests and
subagent kickoffs keep "no cwd"; isAnthropicSubagent and isLikelySubagent
already return early on context_management. A title-generation entry takes
the session's already-known cwd, as before.

New e2e test drives the real proxy path with a mock upstream and checks
index.ndjson (main turn cwd, title-gen isSubagent and inherited cwd) and
sessions.json after shutdown. Old code: 0/3 e2e and the new unit tests fail;
new code: all pass.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
lis186 added a commit that referenced this pull request Sep 28, 2026
Transcript-first per-Ask usage reports for agentflow notebooks (phase A,
no agentflow change), and two tool-neutral upstream proposals (phase B:
worker run ledger, close-time usage reporter). Records the 2026-09-26
evidence behind PR #637.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
lis186 added a commit that referenced this pull request Sep 28, 2026
* docs: ccxray × agentflow integration spec (draft)

Transcript-first per-Ask usage reports for agentflow notebooks (phase A,
no agentflow change), and two tool-neutral upstream proposals (phase B:
worker run ledger, close-time usage reporter). Records the 2026-09-26
evidence behind PR #637.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(agentflow): dashboard hides imported turns by design; adapter reads the index

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(agentflow): revise spec after adversarial review (GPT-6 Astra, Fable 5.1)

- Ask windows: Reply stamp is the end (close commit only a cross-check);
  start is the host user message matching the Ask body; empty scaffolds,
  idle gaps, overlaps, retried closes and reopened rounds defined
- External workers join host-authored dispatch records first; clone path
  plus time bound; B1 becomes 'make the dispatch record a contract'
- Codex has no responseId: one source per Codex session until a tested key
- Reports default to the ccxray data dir; --beside uses the git common dir
  exclude and is refused in stream worktrees (agentflow cleanup rejects
  unknown ignored files)
- Claude transcripts default to 30-day retention: snapshots move into A1;
  settled is decided from on-disk data with source fingerprints
- B2 block is writer-owned like the stamp, generated once, reused on retry
- Corrected stamp format, close semantics, archive naming, workspace-dir;
  added privacy rules and background-agent rows

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(agentflow): record re-verification on merged main e1b882c

- Interactive hosts: ~12% proxy-only traffic (prompt suggestions, titles);
  transcript-only host rows are lower bounds and labelled as such
- Claude proxy+import merge verified at read time; index holds both rows,
  hideImported is presence-based; Codex double-counting verified
- Dispatch join verified 18/18; match on recorded cwd, effort from the
  top-level field, dedupe usage by message.id
- Ask start by first bullet, window tail after the Reply stamp, setup row,
  host session from the matched message
- New A1 prerequisites: live proxy cwd still null on CC 2.1.283, pricing
  differs across proxy/import/reload paths, prompt-suggestion executor kind

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(agentflow): put the proxy in the MVP; split A1 into A1a/A1b

The re-verification showed transcript-only reporting misses about 12% of an
interactive host (prompt suggestions and title generation never reach the
transcript), so the MVP now reads merged proxy + import data and labels each
Ask complete / partial proxy coverage / transcripts only. Records that proxy
index lines and their imported twins are pruned after 14 days, splits A1 into
a terminal-only A1a and a persistence A1b behind a pricing prerequisite, and
updates the #640 cwd gap. Revised after a threeways review.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* docs(agentflow): record owner decisions on completeness and A1a acceptance

Three completeness labels and the per-turn-twin rule for complete are
decided; A1a does not require exact executor time. A-001..A-008 never went
through the proxy, so A1a acceptance expects transcripts only for all of
them and the complete/partial paths need a later proxied Ask or fixtures;
listed as a known gap. Notes that runner truncation of the captured report
does not affect the dispatch fields attribution uses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant