qwen-code/packages/cli
Shaojin Wen 4f79036a22
feat(review): a cost ledger from the records already on disk (#8471)
* feat(review): a cost ledger from the records already on disk

"0.21.3 was fine, 0.21.4 got slow" was settled only by replaying a whole
review under a telemetry exporter and hand-aggregating the output — hours
of forensics to find a repair round that had silently doubled a run
(measured: a +93/-48 PR at high effort cost 523 model calls and 37.8M
input tokens, 9.7M of them redelivering prompts the agents had already
acted on). The usage data was on disk the whole time: every chat and
subagent transcript event carries usageMetadata.

qwen review cost-ledger --plan <plan> aggregates those records — the same
files check-coverage trusts for delivery, found via the same
environment-exported location, floored at the plan's mtime so a review
started an hour into a session does not bill that hour — into per-stream
totals: the main loop and each agent, with input / cached / output /
thinking counts and wall time. Step 8 pastes the printed block into the
saved report, so the next slowness question is a diff of two archives
instead of an excavation.

Informational by construction: an incomputable ledger prints why and
exits 0. Validated against the measured run above — 521 calls, 37.7M
input (93% cached), 849k output, 107 min — matching the telemetry-side
aggregation, minus the two side-query calls that are not the review's.

* test: register cost-ledger in the subcommand registry and its demand message

* fix(review): honest cost-ledger output and a safe archive write (#8471)

Address the review of the cost ledger: report output tokens once
(thinking is a subset of candidates, not a sibling), keep the --out
write inside the exit-0 contract and mkdir its parent, name a missing
plan as the plan, read each transcript once, compare timestamps as
instants, fold relaunched agents into marked rows, and archive the
full ledger next to the Step 8 report.

* fix(review): cost ledger — honest failures, validated plan, shared records (#8471)

Address the second review round: a missing or faulted chat transcript
now says "cost-ledger unavailable" instead of rendering agents-only
totals as the whole cost (the plan proves the main loop ran), and a
subagent dir that fails listing with anything but ENOENT does the
same. The --plan file is validated as a plan report before its mtime
alone sets the billing window. Output derives from totalTokenCount −
promptTokenCount when present, correct under both usage conventions.
Chunk agents label "chunk N" via the shared CHUNK_RE instead of the
malformed "agent chunk N of M"; the transcript listing is one helper
shared with the coverage gate; glued JSONL records are recovered via
parseLineTolerant; totals reuse the rows' accumulator; folded (×N)
rows rank by combined total; stale agent files are skipped by mtime
without being opened. Rendered block gains pluralization, "agent
runs: N", and a B tier; SKILL.md states the ledger's bounded window.

* fix(review): close the cost-ledger audit — refusals, labels, pinned math

Address the remaining review threads on the cost ledger:

- Refuse agents-only totals when the chat file exists but holds no
  above-floor records: a degraded recorder leaves the file present and
  empty while agents run — the same infrastructure fact as an
  unreadable transcript, and exactly the output the missing-file
  refusal exists to prevent.
- Read the agent label from the first user record, not a raw 64KB head
  slice: a fork's agent_bootstrap record precedes the launch prompt,
  quotes other agents' identity lines, and can outgrow any fixed
  window.
- Distinguish parallel invariant agents by their owned file, so
  per-file runs stop folding into a phantom (xN) relaunch row.
- Coerce negative provider counts to zero: the agent path records
  usage uncoerced, and summed negatives rendered >100% cached shares.
- Accept degraded diff-less Step 1 reports: validate diffLines +
  chunks, the pair every plan report carries, instead of
  check-coverage's stricter contract that refused them.
- Pin every branch the second round proved unobservable: the exit
  code on all handler paths, total - prompt under both usage
  conventions, per-agent fault tolerance, the wall-minutes
  conversion, array-shaped usage, the mtime pre-filter and the
  event-level floor, human() rounding, per-condition plan validation,
  sort order against a lexical readdir, the zero-event skip, the
  --out per-stream archive contract, error messages naming their
  paths, truncation membership and folded-run counting, and the
  assistant-type filter.

Every new assertion was mutation-probed: each mutant the review named
now turns the suite red.

* review: pipeline stages keep their own ledger rows

The (×N) fold keyed on the label alone, and three legitimate multi-launch
shapes shared one: a reverse-audit chunk auditor is launched with the same
'chunk N of M' identity as the Step 3B territory finder (five audit rounds
folded into the finder's row — one agent where six pipeline stages ran),
and repeat rounds of the findings roles carry their round OUTSIDE the
backticks (every round folded as a phantom relaunch). labelOf now reads
the stage from the audit brief's record key in the launch (audit chunk N
(round K)) and the round from the identity LINE — never the whole launch,
whose folded findings can quote a budget disclosure's own '(round N)' —
so rounds are rows and only true relaunches and same-round verify shards
fold. The (×N) comment now says what the marker means: N runs under one
label.

* fix(cli): annotate cost-ledger test helper to restore strict build (#8471)

* fix(cli): anchor cost-ledger labels and harden broken-usage defenses (#8471)

---------

Co-authored-by: verify <verify@local>
Co-authored-by: qwen-code-dev-bot <qwen-code-dev@service.alibaba.com>
2026-08-05 02:18:06 +00:00
..
src feat(review): a cost ledger from the records already on disk (#8471) 2026-08-05 02:18:06 +00:00
.gitignore feat(core): add opt-in built-in web_search backed by the DashScope Responses API (#7215) 2026-07-21 10:59:36 +00:00
index.ts fix(cli): add bootstrap fast paths (#6188) 2026-07-02 22:28:11 +00:00
package.json chore(release): v0.21.5 (#8505) 2026-08-04 02:30:44 +00:00
test-setup.ts feat(serve): persist dynamic workspace registrations (#6716) 2026-07-11 16:49:40 +00:00
tsconfig.json feat(serve): resolve and report the daemon memory budget (#8245) 2026-08-02 04:19:31 +00:00
vitest.config.ts feat(serve): add a required external tool guard provider (#8125) 2026-08-04 14:37:13 +00:00