Reconcile with #914 (which independently fixed the same date-sensitive
fixtures): keep this branch's equivalent date fixes for parser.test.ts and
cli-durable-totals.test.ts, and drop the global vitest retry #914 added now
that this branch fixes the flakes at the root (fs.rm retries, longer timeouts,
serial cache-lock CI step). The targeted cache-lock local retries stay.
Hermes, Cursor, OpenCode and copilot OTel sources all live in SQLite
databases that their agents keep open in WAL mode for the life of the
process. Committed writes park in <db>-wal until a checkpoint, so the
main file's stat can sit hours or days behind the newest committed data.
fingerprintFile only statted the main file, which broke two ways:
- The date-range mtime pre-filter in parseProviderSources read the stale
mtime as "nothing in range" and skipped the source entirely. Every
Hermes session committed after the last checkpoint vanished from
reports: the today-parse skipped the db (mtime < local midnight) while
the backfill only keeps days through yesterday. Exactly the "17
sessions in the DB, 14 reported, the 3 from today missing" report in
issue #913.
- reconcileFile saw an unchanged fingerprint between checkpoints and kept
serving stale cached turns for sessions that had since grown.
Fold the -wal sibling into the fingerprint: newest mtime wins and sizes
add, so both WAL growth and a checkpoint (db grows, wal truncates) move
the fingerprint. -shm is deliberately ignored (it mutates on reads).
Bare SQLite paths get the fold only when the extension says database, so
JSONL transcript fingerprints (offset-based append detection) are
untouched.
Refs #913
parser.test.ts (a)/(f): createJsonlSession stamped events at a fixed 2026-05-01
that aged past the 90-day retention window, pruning to zero; date them relative
to now. cli-durable-totals: seedLiveTodaySession stamped noon, which is in the
future on a pre-noon run so the provider-scoped today slice (ends at now)
dropped it while the all path (ends at range end) kept it; seed a past-today
time. cache-refresh-lock and other integration tests starve under a saturated
parallel run and fail closed; add a small global retry and raise the two most
load-sensitive lock tests. Test-only; no production code changed.
When the project directory IS the home directory, countSkills pushed both
~/.claude/skills and <project>/.claude/skills - the same path - and counted
every skill twice, and scanMemoryFiles read ~/.claude/CLAUDE.md twice,
inflating the context-budget estimate. Dedupe both by resolved path.
Mutation-checked: a single home skill counts 2 before the fix, 1 after.
cacheKey fingerprinted only project count + api-call sum, so two datasets
agreeing on those two numbers collided onto one cached OptimizeResult, and
a cost/token change that left call count unchanged (e.g. a re-price) served
stale findings within the 60s TTL - reachable in the long-lived menubar.
Fold total cost, savings and proxied cost (scaled to micro-dollars) into
the key. Exported cacheKey and mutation-checked: the old key collides two
same-shape datasets and a re-price; the new one separates both, while an
identical dataset still keys identically.
parseChatFile estimated input tokens from pendingUserMessage - the last
human turn sliced to 500 chars - while output summed every bot char, so a
multi-turn session or any prompt over 500 chars undercounted input tokens
and therefore costUSD severalfold. Accumulate every human turn's full
length (inputChars), matching the modern-execution path; keep the 500
slice for the display userMessage only. Mutation-checked: a 2400-char
prompt reports 125 tokens before, 600 after.
The shared base computation guarded hours >= 2 but its h < 2 branch
still subtracted five minutes past midnight, escaping into yesterday
during the first five minutes of UTC hours 0 and 1 and zeroing every
'today' assertion - which is exactly when runs #6 landed. Midnight
clamp replaces the guard at all four sites.
server.close() only stops new connections; an in-flight fire-and-forget
cache save can land a file mid-recursive-rm, surfacing as ENOTEMPTY on
slower runners (run #5). fs.rm's built-in retries absorb the window.
Every case spawns the real CLI and does genuine multi-provider parse
work; run #3 showed a second sibling crossing the 5s default on the
shared runner. File-level cap replaces the earlier single-test one.
- cli-durable-totals: the live fixture session was stamped at noon today,
so every before-noon run saw it in the future; the provider-filtered
path drops future instants while the all-provider path keeps the whole
day, failing the parity assertion. Relative-and-clamped timestamps,
the same fix project-filter-durable-totals got in 1596220.
- parser (copilot, 2 cases): the fixture's fixed 2026-05-01 dates crossed
copilot's durable 90-day age-out on 2026-07-30, so the first parse
pruned the freshly-cached session. Relative timestamps.
- parser-incremental-append: unlink-then-create let ext4 hand the freed
inode straight back, breaking the new-inode premise. The replacement
is now created beside the original and renamed over it.
- parser-proxy-pricing: normalizeProxyPath folds case only on darwin and
win32, deliberately; the test now asserts the platform-correct
behavior on both kinds of filesystem instead of hardcoding macOS.
- cli-status-menubar: the config-source filter case does real multi-parse
work and needs more than the 5s default on shared runners; 30s cap.
The rollup fallback was gated on the post-dedup emitted counter, so a session
whose per-message calls were all suppressed by the shared dedup (a duplicated
session directory reusing a session_id) fell through to the metadata.usage
rollup and double-counted its cost. Gate on a hadMetrics flag set before the
dedup check instead.
The Cline CLI (npm `cline`, 3.x) stores sessions as
<sessions>/<id>/<id>.json + <id>.messages.json. The existing `cline`
provider only discovers tasks/<id>/ui_messages.json, so every CLI session
was silently reported as $0.00 — no warning, not even under --verbose.
Adds `cline-cli` as its own provider rather than a third root on `cline`,
leaving the shared Cline-family parser (Roo Code, KiloCode, IBM Bob)
untouched. It mirrors the CLI's own root resolution
(CLINE_SESSION_DATA_DIR -> CLINE_DATA_DIR -> CLINE_DIR -> ~/.cline),
implements probeRoots() so `doctor` can tell "not installed" from "wrong
override", emits one call per assistant message's `metrics` block, and
falls back to the session rollup when a session carries none. The
fallback reads `usage`, not `aggregateUsage`, which folds in spawned
subagents that are themselves separate session directories.
Two supporting changes, both required for CLI costs to report correctly:
- parser.ts re-priced cline-cli calls from tokens because the provider
was not on the reported-cost allowlist, inflating a real 12-session
local sample from $1.11 to $3.92.
- session-cache.ts gains the matching PROVIDER_ENV_VARS entry (so a
changed override invalidates) and a `reported-cost-v1` parse version
(so sessions cached before the allowlist fix re-parse once instead of
being re-priced forever).
Cost is treated as metered only when actually present and non-negative,
so a metered $0 stays reported while a missing or negative cost falls
back to token pricing — applied identically on the per-message and
rollup paths. Timestamps promote a seconds-resolution value rather than
silently landing in 1970, matching the guard kiro.ts uses.
CLINE_DIR / CLINE_DATA_DIR / CLINE_SESSION_DATA_DIR are added to the test
env-isolation list so a developer's real sessions cannot bleed into
fixtures.
The VS Code variant discovery bug reported alongside this in #874 is
deliberately NOT fixed here — it shipped in #882.
Verified against 18 real local sessions: 142 calls, 4,934,762 input /
224,561 output tokens, and a cost matching the CLI's own metered total to
the cent. `codeburn doctor` reports "Cline CLI OK".
Refs: #874
Several model ids price correctly but had no SHORT_NAMES entry, so the
By Model panel rendered the raw slug next to properly named siblings:
`gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna`, `grok-4.5`,
`qwen3.7-max`, `minimax-m3` and `mimo-v2.5-pro`.
All display-only; no dollar amounts move.
Notes on the less obvious ones:
- The GPT-5.6 variants are listed individually rather than as a bare
`gpt-5.6`. A base entry would swallow every future `gpt-5.6-*` through
the prefix match and hide the variant behind a sibling's label, which
is exactly what getShortModelName's version-boundary rule prevents. An
unlisted variant still falls through to its raw id, and there is a test
pinning that.
- `grok-4.5` is the model the Grok Build harness runs and reports as
`current_model_id`, so it takes the model's own name. Ids that really
are `grok-build*` keep the "Grok Build" label, also covered by a test.
- ClinePass routes models as `cline-pass/<slug>`. No new prefix handling
was needed: getShortModelName's path fallback already strips the
prefix and re-resolves the bare slug, the same way it handles
`accounts/fireworks/models/<slug>`.
- MiniMax M3 is mapped under both the lowercase OpenRouter slug and the
capitalized spelling sessions report, since SHORT_NAMES matching is
case-sensitive (the case-insensitive index covers pricing only).
`mimo-v2.5-pro` remains unpriced upstream; this only gives it a name.
deferRow now sums only the servers no applied mcp-remove / mcp-project-scope
record already measures, so a defer row and an MCP row can never claim the
same server's schema tokens over the same post-apply sessions and
totalRealizedTokens stays a disjoint sum. Conservative by design: the defer
row drops a claimed server for its whole window, and when every server is
claimed it reports not measurable instead of guessing.
deferredSessions === 0 now reports the new 'pending' status instead of
asserting 'reverted': the note already said the cause was ambiguous, and
--json consumers could not tell a not-yet-restarted client from a genuine
revert. The table renders it as 'not yet in effect'.
The provider list rebuilt its own day set straight from the daily cache,
unioning unfiltered historical days with an already-filtered today. Per-provider
costs therefore counted every carried day whole while today honoured the name
filters, so the list could not be reconciled with the headline or the By Project
panel. #864 fixed the headline and left this deliberately untouched.
Reuse durable.days, which is the same union the headline is built from, already
narrowed by range, day selection and project filter. That is what the comment
above the buildDurablePeriod call already promised this section would do.
Providers whose entire spend is excluded do not vanish from the list: the
installed-but-zero backfill below still adds them at cost 0.
Codex driving a Kimi backend records the model as kimi/k3[1m] (provider prefix
plus a [1m] context tag). getCanonicalName stripped the prefix but not the tag,
so it matched no alias and priced to $0 - Kimi-via-codex spend was silently
reported as free, and the menubar showed no Codex segment. Strip a trailing
[...] context tag so kimi/k3[1m] -> k3 -> kimi-k3, repairing both cost and the
display name.
Follow-up to #881. Structural discovery admits third-party rollouts whose
schema is unverified. Two unchecked JSON.parse fields still reached string ops
on the parse path: an unparseable timestamp threw RangeError out of the
fork-cutoff Date math, and a non-string model threw TypeError from calculateCost
(.replace). Either sank that session's usage to zero. Skip the fork cutoff for
an unparseable timestamp, and only adopt a string model (falling back to a real
model otherwise) so the session is counted instead of silently reading zero.
Follow-up to #880. Make fetchReleaseAsset generic over a consume callback
that runs inside the retry loop, so a socket dropped mid-download is retried
and its partial file removed rather than aborting the install and leaving a
truncated zip. Clamp a non-finite attempt budget, drain non-ok bodies, and
carry the original error as cause. The checksum comparison stays outside the
retry, so an integrity mismatch still aborts immediately and never re-downloads.
Codex session discovery required `payload.originator` to start with
"codex" (case-insensitive). `originator` is a free-form client identity
string, not a format marker: any tool driving `codex app-server` writes
structurally identical rollouts under ~/.codex/sessions with its own
value ("t3code_desktop", "JetBrains.IntelliJ IDEA", ...). Those sessions
were silently dropped from every report, and each past fix only admitted
one more spelling.
Gate on structure instead: a first line that parses as JSON, has
type === "session_meta", and carries a plain-object payload. Foreign and
malformed files are still rejected. Directory ownership decides the
provider — codex.ts is the only provider that reads ~/.codex, and the
walk only visits rollout-*.jsonl under the strict YYYY/MM/DD path or
archived_sessions/ — so no double counting is possible. `originator` is
still parsed onto the meta entry; nothing downstream reads it.
Bump the daily cache to v16. Historical days are served from that cache
(usage-aggregator only recomputes today) and retention is ten years, so
without a bump an upgrading user with a warm cache keeps the pre-fix
rollups forever: discovery reruns, so the session COUNT moves, while
cost and calls stay frozen — a self-contradicting report that reads as
"fixed". Measured on a fixture with two same-day rollouts, one
codex-cli and one t3code_desktop:
pristine main, fresh cache cost 4.55 calls 1 sessions 1
this branch, main's warm cache cost 4.55 calls 1 sessions 2 (was)
this branch, main's warm cache cost 18.2 calls 2 sessions 2 (now)
this branch, fresh cache cost 18.2 calls 2 sessions 2 (truth)
CODEX_CACHE_VERSION and PROVIDER_PARSE_VERSIONS.codex deliberately stay
put: both caches are keyed per file path and are written only after a
successful parse, so a file rejected at discovery has no entry to
invalidate. Verified on the fixture above — main's codex-results.json
and session-cache.v7.json hold only the first-party rollout, and reusing
them unchanged still yields the correct total.
Harden `payload.cwd` while admitting unverified clients. It is declared
`string` but comes straight off JSON.parse, and a number/object/array
threw "cwd.replace is not a function" out of sanitizeProject; the throw
escaped discoverSessions into safeDiscoverSessions, which returns [] for
the WHOLE provider, so one malformed file made every Codex report read
zero. Guarded in discovery (falls back to the `unknown` project) and on
the parse side, where a non-string cwd would otherwise ride into
projectPath/workingDirectory and reach the parser's path helpers.
Closes#873, closes#626.
Cline discovery only looked at the stable VS Code globalStorage root, so
tasks created in VS Code Insiders or VSCodium were never found. The
singular getVSCodeGlobalStoragePath helper returns paths[0], and because
the provider always passed a concrete overrideDir, the 3-variant fallback
inside discoverClineTasks was never reached - unlike the Roo Code and
KiloCode siblings, which pass overrideDir straight through.
Build the default roots from getVSCodeGlobalStoragePaths (stable,
Insiders, VSCodium) plus the ~/.cline/data root and hand them to
discoverClineTasks in one call. The existing dedupe by task id still
collapses a task id seen in more than one root, so totals cannot inflate.
The configuredDirs override used by tests and createClineProvider(dirs)
is unchanged.
buildDurablePeriod derived the today slice of the multi-day, all-provider
headline from the unsliced whole-range parse, so a turn spanning local midnight
kept its category and turn count anchored on its yesterday start. The per-call
cost and calls bucketed onto today correctly, but By Activity and the JSON
daily turn count lost the post-midnight half — categories summed to only the
pre-midnight cost while the headline, By Model and By Project were right.
Slice the today parse with filterProjectsByDays first, which re-anchors the
straddling turn to its surviving today calls, so today's category cost lands on
today. Category cost is the sum of the slice's own calls, so day-N + day-N+1
still equals the whole-range total (no over-count); the per-day turn-count
split matches the cache side and the documented per-day semantics.
Adds a regression test in the straddling-turn conservation suite
(mutation-checked: fails on the pre-fix code). Also fills in the CHANGELOG
Unreleased entries for the batch (#853, #856, #872, #846/#859, #866/#867, #833).
`codeburn menubar` aborted the install on the first bad response from
GitHub release-asset delivery. A transient HTTP 500 on the checksum
fetch (issue #876) killed an otherwise healthy install even though the
asset was published correctly and the next request succeeded.
Retry the zip and checksum downloads up to 3 times with a short
exponential backoff (0.5s, 1s) on 5xx responses and network-level
errors. 4xx is never retried: 404/410 still falls through to the
release-API discovery path unchanged, and a 403/429 rate limit cannot
clear inside the backoff window so it surfaces immediately with its
retry-after hint. The checksum comparison stays outside the retry loop
so a genuine digest mismatch still aborts on the first look.
Thrown errors now name the requested URL so a failure is actionable.
Retry parameters and the fetch/sleep/log seams are injectable, matching
the options pattern in src/sync/push.ts and src/cache-refresh-lock.ts.
The tests added in #864 seeded today's live session at a fixed wall-clock
hour (12:00 local). The periods they build end at `new Date()`, and the
suite runs under TZ=UTC, so for any run before 12:00 UTC that timestamp
is in the FUTURE and the range filter correctly drops it. The live half
of the cache/live union then contributes nothing, and the one assertion
that needs a non-zero headline — the unattributed-cost footnote — fell
into renderOverview's "No usage found" early return and went red. Half
of every day was a failing window; a769b50 fixed the start-of-month
flake but this one survived it.
Verified by bisecting the fixture on the unpatched test: moving the
seeded hour from 12:00 to 01:00 (past, at a 03:38 UTC run) turns the
same 12 tests green, so the timestamp's position relative to `now` is
the whole cause.
- Seed the session a few minutes BEFORE now, clamped to today's
midnight, so it is always both inside today and already in the past.
- Stop the footnote test depending on the live parse at all: seed a
second, attributable cached day so the headline is non-zero from the
cache alone. The test now exercises the footnote instead of the
fixture's timing.
Claude is scanned via scanProjectDirs instead of parseProviderSources, and
that call had no provider-filter guard. On a --provider <other> run
discoverAllSessions correctly returns no claude sources, so claudeDirs is
empty, but scanProjectDirs still ran: its orphan pass reads the whole cached
claude section and treats every file as no-longer-discovered, re-injecting
PR-bearing entries (and in read-only mode every cached entry) into the result.
The headline stayed correct because it comes from the provider-sliced daily
cache, so only the live-parse panels were wrong. By Model then listed
Anthropic models under --provider cursor while the total showed cursor alone.
Guard the scan with claudeInScope, mirroring the guard the durable-orphan
loop already applies. Deliberately not a claudeDirs.length check: when claude
is in scope but every transcript has been pruned, the orphan pass is what
keeps PR-attributed spend from vanishing.
Review rounds 2-3 + self-review on --attribution:
Credential egress (round 2):
- normalizeRemoteUrl: scp userinfo expressed as an optional regex group
let backtracking re-parse a credential prefix as host:path
(x-access-token:ghp_...@host/repo -> token in git.repo). Userinfo is
now split off at the first @ BEFORE any host matching.
- Positive validation (allow-list) as the final gate on EVERY branch:
host must be hostname-shaped, every path segment repo-shaped, total
identity <= 200 chars. Kills transport-helper remotes (ext:: leaks
local SSH key paths, codecommit:: leaks AWS profile names), residual
@, spaces/colons, and unbounded strings.
- sanitizePrLinks: links are rebuilt from origin + pathname — userinfo,
query strings, and fragments are dropped instead of passed through;
collapsed duplicates dedupe.
Attribution correctness (round 3 + self-review):
- Double-count fix with precise retraction semantics: when a commit
migrates to a later-parsed tighter-window session, the loser re-emits
git.commit_count=0. Empty records are emitted ONLY on a true loss in
THIS computation (lostCandidacy) — a commit that merely aged out of
the --since range was lost to nobody, and retracting it would
permanently zero a still-correct server-side count. The sync layer
additionally requires a prior ledgered state for the session.
- Session dedup key includes project + both window timestamps, so
ongoing sessions re-emit with corrected span times.
- Span end times clamped like the usage builder (never 0, never
earlier than start + 1ms).
- CLI mirrors the usage path on attribution push failures instead of
claiming success.
- Identity normalization: case-insensitive .git strip, doubled path
slashes collapse.
AI-Origin: human
Review findings on the --attribution PR:
- Privacy: sessions whose project path no longer resolves inherited the
cwd-fallback repo identity, egressing whatever (possibly confidential)
repo the user pushes from and falsely attributing its commits.
buildRepoGroups now tracks per-session identity provenance; the
attribution path excludes fallback sessions from commit attribution
entirely (no repo, no commits, PR links only) — they also can no
longer steal a commit from a genuine session's window.
- Privacy: Windows drive-letter paths (C:/..., C:\..., drive-relative)
parsed as scp-like remotes, emitting local filesystem paths as repo
identities. normalizeRemoteUrl rejects drive letters and
single-character hosts (dotless intranet hosts still accepted).
- Hardening: PR links are shape-checked before sending (https,
/org/repo/pull/N path, <=256 chars, max 20 per session) — upstream
parsers only truthiness-check them.
- Safety valve: MAX_ATTRIBUTION_PER_PUSH (10k) caps a first
--since all --attribution push; dry-run reports the cap.
- Tests: adversarial normalize corpus, cwd-fallback egress repro,
commit-stealing prevention, PR-link sanitization, and CLI-level tests
(mock IdP + collector): dry-run sends nothing to the traces endpoint,
flag-off emits no attribution span names on the wire.
- Docs: reconciled the 'never sent' wording with reality (PR links ride
even when repo is null; device_id/methodology/timestamps disclosed).
CHANGELOG Unreleased entry added.
AI-Origin: human
Expose the yield session-to-commit correlation through codeburn sync so
backends can join AI usage to git activity without local git hooks.
- yield: export normalizeRemoteUrl (host/org/repo; credentials, ports,
and .git stripped) and computeAttributionRecords, which reuses the
exact repo-grouping + tightest-window attribution from computeYield
(extracted into a shared buildRepoGroups) and joins in the normalized
origin remote and session prLinks.
- otlp: two new span types sharing the session traceId —
codeburn.session.attribution (git.repo, git.pr_links, git.commit_count)
and codeburn.commit (git.sha, git.in_main, git.was_reverted). Resource
attribute codeburn.attribution_methodology=timestamp-window marks the
attribution as inferred.
- push: generic send core reused by usage and attribution batches. Dedup
keys encode mutable state (inMain/wasReverted), so a state transition
re-sends the updated fact while identical states dedupe via the
existing sent-ledger.
- cli: opt-in --attribution flag on sync push (dry-run aware); commits
in repos with no network remote are never sent.
AI-Origin: human
Two issues on top of the --project/--exclude durable-headline fix:
- sanitizeProjects dropped any project whose key is an Object.prototype member
name (constructor, valueOf, __proto__, ...). A project key is a directory
basename, so such a name is legitimate, and dropping it left the day's
per-project split summing to less than the day cost — so the sliced,
project-filtered headline silently lost that project's spend with no footnote.
The keys are written via setOwn (defineProperty), so keeping them is
pollution-safe; only the redundant `name in Object.prototype` guard is removed.
Regression test added (mutation-checked: fails without the guard removed).
- The new project-filter tests seeded a carried day 10 days ago but ranged over
the calendar month, so within the first 10 days of a month that day fell out
of range and the tests went red. Replaced with a fixed 20-day window that
always spans the seeded day.