qwen-code/packages
Shaojin Wen 4f79036a22
feat(review): a cost ledger from the records already on disk (#8471)
* feat(review): a cost ledger from the records already on disk

"0.21.3 was fine, 0.21.4 got slow" was settled only by replaying a whole
review under a telemetry exporter and hand-aggregating the output — hours
of forensics to find a repair round that had silently doubled a run
(measured: a +93/-48 PR at high effort cost 523 model calls and 37.8M
input tokens, 9.7M of them redelivering prompts the agents had already
acted on). The usage data was on disk the whole time: every chat and
subagent transcript event carries usageMetadata.

qwen review cost-ledger --plan <plan> aggregates those records — the same
files check-coverage trusts for delivery, found via the same
environment-exported location, floored at the plan's mtime so a review
started an hour into a session does not bill that hour — into per-stream
totals: the main loop and each agent, with input / cached / output /
thinking counts and wall time. Step 8 pastes the printed block into the
saved report, so the next slowness question is a diff of two archives
instead of an excavation.

Informational by construction: an incomputable ledger prints why and
exits 0. Validated against the measured run above — 521 calls, 37.7M
input (93% cached), 849k output, 107 min — matching the telemetry-side
aggregation, minus the two side-query calls that are not the review's.

* test: register cost-ledger in the subcommand registry and its demand message

* fix(review): honest cost-ledger output and a safe archive write (#8471)

Address the review of the cost ledger: report output tokens once
(thinking is a subset of candidates, not a sibling), keep the --out
write inside the exit-0 contract and mkdir its parent, name a missing
plan as the plan, read each transcript once, compare timestamps as
instants, fold relaunched agents into marked rows, and archive the
full ledger next to the Step 8 report.

* fix(review): cost ledger — honest failures, validated plan, shared records (#8471)

Address the second review round: a missing or faulted chat transcript
now says "cost-ledger unavailable" instead of rendering agents-only
totals as the whole cost (the plan proves the main loop ran), and a
subagent dir that fails listing with anything but ENOENT does the
same. The --plan file is validated as a plan report before its mtime
alone sets the billing window. Output derives from totalTokenCount −
promptTokenCount when present, correct under both usage conventions.
Chunk agents label "chunk N" via the shared CHUNK_RE instead of the
malformed "agent chunk N of M"; the transcript listing is one helper
shared with the coverage gate; glued JSONL records are recovered via
parseLineTolerant; totals reuse the rows' accumulator; folded (×N)
rows rank by combined total; stale agent files are skipped by mtime
without being opened. Rendered block gains pluralization, "agent
runs: N", and a B tier; SKILL.md states the ledger's bounded window.

* fix(review): close the cost-ledger audit — refusals, labels, pinned math

Address the remaining review threads on the cost ledger:

- Refuse agents-only totals when the chat file exists but holds no
  above-floor records: a degraded recorder leaves the file present and
  empty while agents run — the same infrastructure fact as an
  unreadable transcript, and exactly the output the missing-file
  refusal exists to prevent.
- Read the agent label from the first user record, not a raw 64KB head
  slice: a fork's agent_bootstrap record precedes the launch prompt,
  quotes other agents' identity lines, and can outgrow any fixed
  window.
- Distinguish parallel invariant agents by their owned file, so
  per-file runs stop folding into a phantom (xN) relaunch row.
- Coerce negative provider counts to zero: the agent path records
  usage uncoerced, and summed negatives rendered >100% cached shares.
- Accept degraded diff-less Step 1 reports: validate diffLines +
  chunks, the pair every plan report carries, instead of
  check-coverage's stricter contract that refused them.
- Pin every branch the second round proved unobservable: the exit
  code on all handler paths, total - prompt under both usage
  conventions, per-agent fault tolerance, the wall-minutes
  conversion, array-shaped usage, the mtime pre-filter and the
  event-level floor, human() rounding, per-condition plan validation,
  sort order against a lexical readdir, the zero-event skip, the
  --out per-stream archive contract, error messages naming their
  paths, truncation membership and folded-run counting, and the
  assistant-type filter.

Every new assertion was mutation-probed: each mutant the review named
now turns the suite red.

* review: pipeline stages keep their own ledger rows

The (×N) fold keyed on the label alone, and three legitimate multi-launch
shapes shared one: a reverse-audit chunk auditor is launched with the same
'chunk N of M' identity as the Step 3B territory finder (five audit rounds
folded into the finder's row — one agent where six pipeline stages ran),
and repeat rounds of the findings roles carry their round OUTSIDE the
backticks (every round folded as a phantom relaunch). labelOf now reads
the stage from the audit brief's record key in the launch (audit chunk N
(round K)) and the round from the identity LINE — never the whole launch,
whose folded findings can quote a budget disclosure's own '(round N)' —
so rounds are rows and only true relaunches and same-round verify shards
fold. The (×N) comment now says what the marker means: N runs under one
label.

* fix(cli): annotate cost-ledger test helper to restore strict build (#8471)

* fix(cli): anchor cost-ledger labels and harden broken-usage defenses (#8471)

---------

Co-authored-by: verify <verify@local>
Co-authored-by: qwen-code-dev-bot <qwen-code-dev@service.alibaba.com>
2026-08-05 02:18:06 +00:00
..
acp-bridge feat(serve): add a required external tool guard provider (#8125) 2026-08-04 14:37:13 +00:00
audio-capture chore(release): v0.21.5 (#8505) 2026-08-04 02:30:44 +00:00
channels chore(release): v0.21.5 (#8505) 2026-08-04 02:30:44 +00:00
chrome-extension feat(browser-ext): add alpha readiness diagnostics (#6739) 2026-08-04 03:23:47 +00:00
cli feat(review): a cost ledger from the records already on disk (#8471) 2026-08-05 02:18:06 +00:00
core feat(review): a cost ledger from the records already on disk (#8471) 2026-08-05 02:18:06 +00:00
cua-driver fix(mcp): add opt-in model payload filtering (#7413) 2026-07-21 09:49:04 +00:00
desktop feat(core): Normalize tool-call terminal telemetry (#8176) 2026-07-31 11:51:14 +00:00
desktop-shell fix(desktop): codesign ripgrep and node binaries before tauri build (#8518) 2026-08-04 14:41:36 +00:00
mobile-mcp chore(release): v0.20.1 (#7461) 2026-07-22 02:00:22 +00:00
sdk-java fix(sdk-java): accept ERROR terminal in teardown E2E to fix flaky race (#8354) 2026-08-02 03:34:47 +00:00
sdk-python fix(sdk-python): validate max_tool_calls and max_subagent_depth as integers (#7548) 2026-07-23 08:45:13 +00:00
sdk-typescript feat(channels): add Web Shell management support for GitHub and GitLab (#8310) 2026-08-02 12:29:10 +00:00
vscode-ide-companion perf(core): clear tool results to a low watermark to preserve prompt cache (#8464) 2026-08-04 19:00:11 +00:00
web-shell feat(web-shell): bind plan approval to its Todo revision (#8393) 2026-08-04 14:25:50 +00:00
web-templates chore(release): v0.21.5 (#8505) 2026-08-04 02:30:44 +00:00
webui chore(release): v0.21.5 (#8505) 2026-08-04 02:30:44 +00:00
zed-extension chore(deps): upgrade ink 6.2.3 → 7.0.2 + bump Node engine to 22 (#3860) 2026-05-11 17:29:50 +08:00