Re-instrumenting this exact head over the full corpus traced the
original 12.5%/46% figures to a replication gap: the earlier
prevCumulativeTotal guard (codex.ts) already discards any token_count
event whose running total exactly repeats the previous kept one, so a
drop at the seenKeys dedup site is always a byte-identical replay of
already-counted tokens -- never a real loss. The condition BUG-2's fix
guarded against does not occur in real Codex output.
Removes taskDedupedTokens, its increment at the dedup site, the
active-time scaling at task_complete, and the now-unexercisable unit
test (its fixture forces a dedup collision that is not also a
cumulative-total repeat, a state the real writer never produces). The
dedup site keeps a short comment recording why no rescaling is needed,
so the next investigator doesn't retrace this.
Keeps: the reasoning double-count fix, the harness-startup exclusion
(BUG-1, confirmed to the decimal), the shared mergeToolIntervals
helper and permissive duration parsing (BUG-8), and the legend rename.
The real-corpus table is unchanged -- BUG-2 measured zero effect on it
before this revert too.
Extends #1079 (reasoning double-count) with three more findings from an
exactness pass over the same throughput path:
- BUG-1: task_started fires before Codex assembles the request, so the
gap to the first request-context event (turn_context, world_state,
event_msg/user_message, or response_item/message) was counted as
active model time. The active window now starts at that event instead.
- BUG-2: a token_count event dropped by fork-replay dedup lost its
tokens from the numerator while the task's real duration still spanned
it in the denominator, understating Tok/s for a partially (not fully)
deduped task. The active window is now scaled down by the dropped
tokens' proportional share.
- BUG-8: the tool-interval clip/merge/cap logic was duplicated between
providers/codex.ts and codex-throughput.ts and had already drifted
(task_complete only read a plain-number duration, unlike
mcp_tool_call_end). Now one shared mergeToolIntervals, and
task_complete's duration parses the same permissive forms.
Cost and every token count remain byte-identical. activeGeneratedTokens/
activeDurationMs/toolWaitMs are stored verbatim in both Codex caches, so
none of this self-heals -- but the CODEX_CACHE_VERSION 13 bump already
shipped for #1079 covers the same fields, so no further bump is needed.
Dashboard's per-model column keeps the "Tok/s" header (zero width slack
at the standard layout, verified against a real test); the legend now
spells out "Effective Tok/s" with a decode-speed disclaimer.
The same reasoning-inclusion bug #1075/#1078 fixed for cost also affected
the throughput display: activeGeneratedTokens/taskGeneratedTokens in the
Codex parser and generatedTokens in the live codex-tps reader summed
outputTokens + reasoningTokens, but reasoning is already inside output.
All three sites now route through the billableOutputTokens('codex', ...)
helper #1078 introduced, so the throughput numerator can never drift from
the billed one.
Cost and every token count are unchanged (verified byte-identical on a
real 51,753-call corpus); Tok/s drops 20-42% depending on how
reasoning-heavy the model is.
activeGeneratedTokens/activeDurationMs/toolWaitMs are stored verbatim in
both the Codex result cache and the session cache rather than re-derived
on read, so neither self-heals: CODEX_CACHE_VERSION moves 11 -> 13 (12 is
claimed by feat/core-extraction's own port of this feature) and
PROVIDER_PARSE_VERSIONS.codex gains a codex-tps-v1 suffix, forcing Codex
sessions to re-parse once.
Closes#1079
Mixed-version binaries were clobbering the unsuffixed Codex, Cursor,
and Antigravity result files and each re-parsed the whole corpus.
Write *-results.v<n>.json like the daily cache. Leave the unsuffixed
file for older binaries; adopt a matching-version copy once.
The COST_CHANGED_BY_DESIGN carve-out in compare.mjs left codex cost entirely
unasserted after #1075/#1078. It now requires the upgraded cost to be strictly
lower than baseline and within 25% of it, and the row verdict says "repriced"
instead of the misleading "identical (cost N% drift)".
grok.ts's comment on the reasoning/output split still claimed provider-side
splitting was the repo's only mechanism; it now also names
billableOutputTokens/REASONING_INCLUDED_IN_OUTPUT (models.ts), which is the
other half since #1078. usage-aggregator.ts's "folds reasoning into output"
comment was true pre-#1078 but is backwards for codex now (reasoning is
already inside output, not added to it).
Test exemplars for the "reasoning is additive" case used hermes, whose
upstream is OpenAI-shaped and may not stay a safe example; swapped to gemini,
which documents "thoughts" as genuinely separate output.
CHANGELOG's #1075 entry gets one line noting days whose codex transcripts
have aged out keep their pre-fix totals via the daily-cache never-lose guard,
matching the disclosure already given for #1040.
The pricing cache written to disk had no schema version, so a cache written
by a pre-#1078 binary lacked cacheWriteCostIsExplicit on every entry. Reading
it back resolved the missing key to undefined (falsy), silently reintroducing
the surcharge-fabrication bug #1078 killed for up to CACHE_TTL_MS after an
upgrade. loadCachedPricing now rejects any cache whose version doesn't match
the current schema instead of reading it verbatim.
codexCredits() still accepted an optional reasoningTokens param that added it
to output - the exact double-count #1078 removed from every real caller. The
only caller never passed it; deleted it so it can't be reintroduced by
accident.
parser.ts's activeGeneratedTokens fallback went through billableOutputTokens
in #1078, but codex is the only caller of activeDurationMs/activeGeneratedTokens
and always sets both together, so the fallback branch is unreachable for it.
Reverted to reduce diff noise.
Reasoning tokens are a subset of output_tokens for OpenAI models, not an
extra bucket: on a 1,396-rollout corpus all 134,316 token_count events
carrying a total satisfy input + output == total. CodeBurn added
reasoning_output_tokens on top when pricing a codex call, in the
cache-rehydration re-price, and in the models/audit display sums. That
overstated codex cost by $166.03 (3.5%) and displayed output tokens by
34.6% on that corpus. Both cost sites and the display sums now go through
one shared billableOutputTokens() so a cold parse and a warm read cannot
drift apart.
cache_write_input_tokens (codex PR #33454) was never read and
cacheCreationInputTokens was hardcoded to 0. It is now carved out of the
uncached-input bucket and clamped to it, but routed to the cache-write
bucket ONLY when the pricing source publishes an explicit cache-write rate.
buildCosts fabricates 1.25x input when a source omits one, which is correct
for Anthropic and would have invented a surcharge OpenAI never charged on
gpt-5.5 / 5.4 / 5.3-codex / gpt-5. ModelCosts now carries
cacheWriteCostIsExplicit so that distinction survives getModelCosts.
A cost change invalidates persisted output: codex-results.json v10 -> v11
(stores costUSD verbatim), the codex parse version moves (the token-bucket
change does not self-heal on read), and the daily cache goes 20 -> 23 (21 is
claimed by the #946 landing branch and 22 by PR #1056). The upgrade-path
corpus asserts codex tokens and calls exactly and reports the repricing.
Closes#1075
Extra High MERGE AFTER FIX on 7e51413. JSON-report headlines
come from durable.data; daily[] is the durable.days consumer.
Oracle helper and assertions unchanged.
#1070 is hygiene, not a leak. User docs no longer name
deriveTraceId or list session ids under What is NOT sent.
Code comment matches the wire key. Class test pins every
attribution span: no ai.session_id, join is shared traceId.
Attribution spans already share deriveTraceId(sessionId) with
usage spans. Emitting the raw session id was redundant and
undid the pre-attribution wire rule that session id is hash
input only. Receivers upsert session spans by that keyed
traceId. Usage spans never had the field. Docs match.
Extra High MERGE AFTER FIX on 1fa284f. The parity helper is an
independent per-call oracle for durable.days, not a mirror of a
live dailyMap fallback that no longer exists. Helper and test
unchanged.
The process suite required the loser of a stale-lock contest to
be timed-out. On a slow runner the winner's unlink-guard /
create-successor gap is missing+missing, which the lock honestly
reports as completed-by-other (#904).
Exactly one owner still publishes. The loser may be timed-out,
completed-by-other, or unavailable — never parsed. Do not delay
the clean-release path by treating missing+missing as wait-out.
The visible prefix that decides title-vs-id order must match the
title cap. A UTF-16 slice can split an emoji and treat two distinct
titles as the same truncated series.
Hourly session series still opened with a truncated UUID, so a
monorepo's legend was six identical hex fragments. Sessions
report already prefers title; the chart did not.
Lead with the cleaned title when that prefix is unique in the
visible ~24-glyph budget. Fall back to short-id-first when
titles collide or share a long prefix. Untitled sessions stay
project fallback. No full session ids on the happy path.
Extra High MERGE AFTER FIX: source-scraping the setup file treated a
comment containing 'HERMES_HOME' as isolation. The lists now live in a
side-effect-free module that applyIsolation and the declaration test
both import. Comment-only sabotage fails with hermes:HERMES_HOME.
HERMES_HOME and eight sibling PROVIDER_ENV_VARS data-dir overrides were
fingerprinted for cache invalidation but never CLEARED by the vitest
setup file, so a Hermes-shell laptop parsed real sessions in fixtures.
A static guard fails closed when the map grows another undeclared home.
Maintainer review on #1054: two-journal collision is unreachable
through discovery. Keying by timestamp+journal collapsed a 3-leg
stampless journal onto one row, and the :n strip re-sent ledger
keys. Keep occurrence keys. Take lastEventTimestamp before
sessionStartTime for the call timestamp only. No migration.
Maintainer review on #1049: empty-state copy under --provider codex
named detectors that scanSessions never ran. Say the session-scan
detectors do not cover that provider yet. Drop the dead TUI provider
threading; optimize view is still gated to all|claude.
Mirrors the pricing guard: CODEBURN_FX_NO_FETCH short-circuits getExchangeRate
after the cache read and before the fetch, returning the same USD-equivalent
rate an unreachable network already yields. Tests that seed their own
exchange-rate.json (serve-stdio's EUR case) keep working unchanged.
loadPricing() fetched the live LiteLLM table during tests and live data wins
over the bundled snapshot, so DeepSeek v4 assertions went red when upstream
dropped the off-peak discount. CODEBURN_PRICING_SNAPSHOT_ONLY skips the fetch;
env-isolation sets it for the whole suite.
Extra High HOLD on #1054: timestamp-only keys still collided across
journals, and the provider-wide rollup-ts-v1 bump dropped present
OTel sources, erasing conversations already pruned from the DB.
Identity is now (session, model, timestamp, journal basename).
Legacy :n keys migrate on that JSONL file only.
#1040 fixed nested provenance.model. The compact Buffer path still
took the first cwd/session_id/originator/name anywhere in the payload.
Same depth-1 window for every session_meta string field. Bump parse
fingerprint and CODEX_CACHE_VERSION so present sources re-parse.
Two journals for one session both emitted :1 and the shared seenKeys
set dropped the second file's first leg. Identity is now
(session, model, shutdown timestamp). Bump the Copilot parse
fingerprint so present sources re-parse; durable orphans stay.
Extra High on #1049: keep the shipped Claude Code empty-state and
TUI header strings; resolve agent nouns from Provider.displayName;
cover all five cross-provider detector labels.
It was the only MiMo row still rendering as its raw slug next to
"MiMo v2.5" and "MiMo v2.5 Pro". SORTED_SHORT_NAMES is longest-first, so
the two v2.5 entries keep their own labels.
githubOwnerRepoFromRoot re-read .git/config once per session, so every
session in the same repo paid for the same two syscalls. Memoize it per repo
root for the life of the process.
The prLinks plumbing from a provider call through the session cache into the
session summary had no test above the provider boundary; add one that runs
the real parseAllSessions pipeline against a temp HERMES_HOME. It fails
against the pre-change parser.
The hand-written KNOWN_NAMESPACES set dropped pricing for vendor prefixes
LiteLLM itself indexes: x-ai/, nousresearch/, zhipu/, litellm_proxy/ and
openai_like/ all priced on main and went unpriced here. Derive the set from
the loaded pricing keys instead, so a vendor the catalog knows is never lost
to a stale list, and keep only the spellings no catalog carries as explicit
extras: the routing wrappers, the client-side kimi/ and mimo/ prefixes, and
the litellm_proxy/ + openai_like/ routes. Local runners are excluded on
purpose, so an unlisted ollama tag cannot strip down to a priced cloud row.
xiaomi/ stops being a routing wrapper: it is the namespace LiteLLM prices
MiMo under, and BUILTIN_ALIASES maps the bare MiMo ids INTO it, so peeling
pulled against the alias. It stays known via the derived set.
Also: a user price override for a bare id now wins over the catalog row a
routed spelling of it would otherwise hit, and a namespaced GLM-5.3 is no
longer LABELLED GLM-5.2 by the sibling alias it prices through.
Tests assert the allowlist through a price override on a synthetic id, so
they cannot rot with the snapshot; the glm-5.4 assertion that pinned on the
snapshot NOT carrying a model is dropped.
The budget window comes from computePeriodFromResetDay, which builds an
anniversary period from plan.resetDay (1-28, settable per plan with
`codeburn plan set --reset-day`). "Calendar-month budget" and "Next
calendar reset" are therefore wrong for anyone who moved the reset day,
which is the same class of inaccuracy this change set exists to remove.
Say "budget" and "Next budget reset" instead, and use one wording across
the TUI and the desktop cards.
Both TUI lines truncate end-first at the terminal width. The headline had
grown past the point where an 80-column terminal still showed the
percentage, so it drops "vs ... /mo" for "/ $300.00 budget", and the
status line drops the clause repeating "budget" from the headline. At 80
columns the longest label (custom plans carry their provider) now fits
the percentage, and the status line still shows the projection.
The codex parse version and CODEX_CACHE_VERSION bumps in #1040 make codex
sessions re-parse, but the daily cache has no per-provider invalidation, so
every day already finalized keeps its old per-model rows - and usage-aggregator
serves every day before today from that cache, with ten-year retention. Raise
DAILY_CACHE_VERSION and MIN_SUPPORTED_VERSION to 20 so history re-derives once
off the warm session cache.
The re-derivation test now seeds v19, the last shipped version, so it models the
real 19 -> 20 path, and the upgrade-path check expects daily-cache.v20.json.
Measured on a real 110-day cache: no day lost value, none disappeared, 100 came
back identical, and 9 grok days rose by $19.80 in total from rollups an earlier
parse change had left stale. Every codex model row was unchanged - that corpus
predates the provenance field the fix corrects.
The `mimo-v2-flash -> xiaomi/mimo-v2-flash` alias shipped before this
branch and already cycled through display-name resolution, so
getShortModelName threw RangeError on every real MiMo v2 Flash session.
The new cycle-safe resolver fixes it, but nothing pinned the ids that
actually crashed in production: cover the four spellings found in a real
session cache, including the unnamespaced `mimo/mimo-v2-flash`.
Add the base `mimo-v2.5` display name so the row reads next to
"MiMo v2.5 Pro" instead of showing a raw slug; SORTED_SHORT_NAMES is
longest-first, so the Pro tier still wins its own entry.
CI typecheck failed: sessions-report maps getShortModelName, and
the Extra High cycle Set was a second parameter. Array.map fed
the index as `seen`.
Cycle tracking stays on an internal helper. Display and alias
behavior unchanged. No second Extra High.
Extra High MERGE AFTER FIX on d3f86f5. mimo-v2.5 aliased to
xiaomi/mimo-v2.5 then last-segment recursed forever. Looking up
SHORT_NAMES before resolveAlias also froze user remaps of known
ids (gpt-4o still displayed as GPT-4o).
Follow user aliases first. Break strip→alias→leaf cycles.
Do not invent a Kimi rate. Do not paper over this with a
mimo-v2.5 SHORT_NAMES row.
Hermes and token-plan sessions store mimo-v2.5-pro. The snapshot
row is xiaomi/mimo-v2.5-pro. Same class as the existing
mimo-v2-flash alias. No invented rate.
Looking up the display name on the stripped leaf before following a
pricing alias, so cline-pass/mimo-v2.5-pro cannot recurse
strip → alias → last-segment forever.