mirror of
https://github.com/AgentSeal/codeburn.git
synced 2026-08-29 18:33:01 +00:00
The Copilot CLI and the GitHub Copilot desktop app both write ~/.copilot/session-store.db unconditionally; its assistant_usage_events table holds one row per API request. Until now input/cache tokens for these surfaces came only from the session.shutdown rollups in events.jsonl, which are written only on clean shutdown (a crash loses the whole leg's input/cache accounting) and lump each session leg into one per-model total. The rollup also RESETS its counters at in-session compaction (traced on a clean single-process 107-request session whose sole rollup covered exactly its five post-compaction requests), so even cleanly-closed long sessions were truncated; on a long-history machine the store recovered ~35% of real Copilot spend lost to crashes and compaction resets. The DB rows are per-request, crash-proof, and carry real timestamps. The store's input_tokens is cache-INCLUSIVE (input + cache_read + cache_write), the same convention as the shutdown rollups — verified against each row's token_details_json and by reconciling per-session sums against the CLI's own footers and rollups across two machines (1,380+ rows, 8 models, CLI 1.0.70–1.0.79, schema_version 6): every divergence was a rollup gap. Emitted calls mirror the shutdown-call contract: input/cache/reasoning only, output 0 — per-turn output stays owned by the events.jsonl assistant.message calls. Rollup-vs-store precedence is RECONCILED at serve time, per (session, model), and only there. Both representations always parse and cache; parseProviderSources aggregates the cached calls and, wherever store rows exist for a (session, model), drops the rollup calls and serves the rows plus per-leg RESIDUAL calls: each rollup leg subtracts only the rows in its own interval — rows commit strictly before their leg's shutdown line, so a leg at time T covers exactly the rows in (previous leg's T, T] — and any remainder (per token component, floored at zero) serves once at that leg's own timestamp. A store missing requests a leg covered — adopted mid-session, rows pruned before ever being read — therefore still serves that tail exactly once ON THAT LEG'S DAY, a crash-tail row the rollup never saw can never cancel it, and a complete store serves pure per-request granularity with every residual retired to zero. The decision reads only cached contents, never discovery: deleting or resetting the store changes nothing served, so finalized daily history can never flip on an absence epoch; cached rows of a deleted store remain the record until the 90-day orphan age-out (which exempts still-discovered paths). The serve set is the one coherent snapshot — nothing a writer does between discovery and a parse can change what one pass sees — and read-time precedence heals persisted duplication (stale epochs, runtimes without node:sqlite, restored files) instead of preserving it, following the buildDurablePeriod pattern. Store rows and rollups carry supplementary accounting weight. A rollup (or its residual) is aggregate accounting, never a request: zero api-call/model-call/turn weight, tokens and cost fully retained. A store row is one real request, but when it pairs with a served per-turn call it is supplementary too; rows pair with same-model per-turn calls by timestamp adjacency (monotone matching, tight 2-minute window — the two are written at the same completion moment, and a wide window would let a crash-only row pair against a neighbor whose own row is missing), computed once over the FULL serve set so a date-range boundary that separates a row from its call cannot double the request across adjacent day queries. Only the unpaired rows — store-only requests, exactly where crash-lost requests sit — count. Supplementary-only turns fold into the nearest behavioral turn within 30 minutes; with no behavioral turn to fold into they stay separate weightless turns, each on its own day, with apiCalls 0 — and the session emission gate admits usage-bearing zero-call sessions. The weight propagates into the daily cache: aggregateProjectsIntoDays applies the same rule to every calls counter and category-turn count it seals, so v19 history and live summaries can never disagree about what was a request. A changed source whose read defers on the busy shape (locked, EACCES, corrupt mid-replace — discovery still emits the source; only true absence or a schema mismatch reads as absent) now marks session hydration incomplete, so the daily backfill holds its watermark instead of finalizing a day the deferred rows never reached; an unchanged unreadable store defers nothing. The verdict travels with its result — the 180s memo and the serve burst-reuse restore the hydration verdict their cached data was parsed under, so a memoized partial parse cannot inherit a later parse's complete — and a discovered source whose FINGERPRINT cannot be read (EACCES on a present file) defers instead of silently skipping, while a genuinely deleted file stays a silent skip. Copilot reasoning tokens are no longer double-billed at the report layer: they are a subset of the output the per-turn calls already price, and copilot joins claude in the reasoning-inside-output case of the query-time cost recompute. Store dedup keys are content-discriminated — copilot-store:<sid>:<rowId>:<fnv1a64(created_at|tokens|model)> — because AUTOINCREMENT prevents id reuse only within one database lifetime: a same-path DB reset reusing row ids now mints new keys instead of the durable union swallowing the new usage, while a byte-identical re-insert still collapses (64-bit: 32-bit FNV collisions between plausible token tuples are constructible). Every call of a session serves under one project label resolved at serve time — the session-state-derived label when the serve set knows it, else the store rows' own — so neither rows cached before events.jsonl existed nor an events.jsonl orphaned by a session-state prune can split the session across two grouping keys. CODEBURN_COPILOT_SESSION_STORE_DB is read but deliberately NOT fingerprinted, per the #927 ruling (any copilot fingerprint change drops cached entries whose path still exists, destroying pruned history only the cache holds); the read is allowlisted in the #927 guard, and serve-time reconciliation makes repointing safe without a fingerprint — the new store's rows parse on sight and the old path's entries persist as durable orphans. The copilot parse version appends session-store-v2 and the daily cache bumps v17 → v19: per-day attribution, call counts and costs all change against pre-store builds. 19, not 18: an earlier pushed head of this PR already claimed v18 under different accounting, and the carry-forward would adopt those days as finalized without re-deriving them. Verified by A/B on snapshots of two real stores, a live SIGKILL crash test (row present, no rollup, tokens recovered exactly), live resumes whose warm-cache deltas matched new rows to the token, upgrade-healing at 4,800-session scale, and serve-level regressions pinning every maintainer finding from six review rounds: the rows-then-shutdown race, stale-cache healing, age-out exemption, absence-epoch identity, progressive row landing with residual retirement, behavioral weight across all four pinned scenarios, the hydration fence, project unification in both directions, the same-path reset, mixed coverage (crash tail vs covered-leg gap), multi-leg residual day attribution, range-invariant pairing, memo-scoped hydration verdicts, and the fingerprint-failure fence.
208 lines
7.2 KiB
TypeScript
208 lines
7.2 KiB
TypeScript
import { describe, expect, it } from 'vitest'
|
|
|
|
import { aggregateByBranch, aggregateSessions, attributeSessionPrSpend, renderJson, renderTable } from '../src/sessions-report.js'
|
|
import type { ClassifiedTurn, ParsedApiCall, ProjectSummary, SessionSummary } from '../src/types.js'
|
|
|
|
function makeProject(): ProjectSummary {
|
|
const turn: ClassifiedTurn = {
|
|
userMessage: 'build it',
|
|
timestamp: '2026-07-10T10:00:00.000Z',
|
|
sessionId: 'session-1',
|
|
category: 'feature',
|
|
retries: 0,
|
|
hasEdits: true,
|
|
assistantCalls: [{
|
|
provider: 'claude',
|
|
model: 'claude-sonnet-4-5',
|
|
usage: {
|
|
inputTokens: 100,
|
|
outputTokens: 20,
|
|
cacheCreationInputTokens: 30,
|
|
cacheReadInputTokens: 400,
|
|
cachedInputTokens: 0,
|
|
reasoningTokens: 0,
|
|
webSearchRequests: 0,
|
|
},
|
|
costUSD: 0.12,
|
|
tools: [],
|
|
mcpTools: [],
|
|
skills: [],
|
|
subagentTypes: [],
|
|
hasAgentSpawn: false,
|
|
hasPlanMode: false,
|
|
speed: 'standard',
|
|
timestamp: '2026-07-10T10:01:00.000Z',
|
|
bashCommands: [],
|
|
deduplicationKey: 'call-1',
|
|
}],
|
|
}
|
|
const session: SessionSummary = {
|
|
sessionId: 'session-1',
|
|
project: 'codeburn',
|
|
firstTimestamp: '2026-07-10T10:00:00.000Z',
|
|
lastTimestamp: '2026-07-10T10:05:00.000Z',
|
|
totalCostUSD: 0.12,
|
|
totalSavingsUSD: 0.03,
|
|
totalInputTokens: 100,
|
|
totalOutputTokens: 20,
|
|
totalReasoningTokens: 0,
|
|
totalCacheReadTokens: 400,
|
|
totalCacheWriteTokens: 30,
|
|
apiCalls: 1,
|
|
turns: [turn],
|
|
modelBreakdown: {
|
|
'claude-sonnet-4-5': { calls: 1, costUSD: 0.12, tokens: turn.assistantCalls[0]!.usage, savingsUSD: 0.03 },
|
|
},
|
|
toolBreakdown: {},
|
|
mcpBreakdown: {},
|
|
bashBreakdown: {},
|
|
categoryBreakdown: {} as SessionSummary['categoryBreakdown'],
|
|
skillBreakdown: {},
|
|
subagentBreakdown: {},
|
|
}
|
|
return {
|
|
project: 'codeburn',
|
|
projectPath: '/tmp/codeburn',
|
|
sessions: [session],
|
|
totalCostUSD: 0.12,
|
|
totalSavingsUSD: 0.03,
|
|
totalApiCalls: 1,
|
|
totalProxiedCostUSD: 0,
|
|
}
|
|
}
|
|
|
|
describe('sessions JSON emitter', () => {
|
|
it('flattens SessionSummary fields into the exact JSON row shape', () => {
|
|
const rows = aggregateSessions([makeProject()])
|
|
const parsed = JSON.parse(renderJson(rows))
|
|
|
|
expect(parsed).toEqual([{
|
|
sessionId: 'session-1',
|
|
title: '',
|
|
project: 'codeburn',
|
|
provider: 'claude',
|
|
models: ['claude-sonnet-4-5'],
|
|
cost: 0.12,
|
|
savingsUSD: 0.03,
|
|
calls: 1,
|
|
turns: 1,
|
|
inputTokens: 100,
|
|
outputTokens: 20,
|
|
cacheReadTokens: 400,
|
|
cacheWriteTokens: 30,
|
|
startedAt: '2026-07-10T10:00:00.000Z',
|
|
endedAt: '2026-07-10T10:05:00.000Z',
|
|
durationMs: 300_000,
|
|
}])
|
|
})
|
|
|
|
it('renders a simple table', () => {
|
|
const output = renderTable(aggregateSessions([makeProject()]), { terminalWidth: 120 })
|
|
expect(output).toContain('Session')
|
|
expect(output).toContain('session-1')
|
|
expect(output).toContain('Sonnet 4.5')
|
|
expect(output).toContain('1 sessions')
|
|
})
|
|
|
|
it('hides home-directory slugs and renders compact agent/worktree names', () => {
|
|
const rows = aggregateSessions([makeProject()])
|
|
rows[0]!.sessionId = 'agent-a72cc958e305e4957'
|
|
rows[0]!.project = '-Users-torukmakto-Projects-eywa-eywa--claude-worktrees-issue-131'
|
|
const output = renderTable(rows, { terminalWidth: 160 })
|
|
|
|
expect(output).toContain('Agent a72cc958')
|
|
expect(output).toContain('eywa · issue-131')
|
|
expect(output).not.toContain('Users-torukmakto')
|
|
})
|
|
|
|
it('normalizes Claude dot-worktree slugs without exposing the home directory', () => {
|
|
const rows = aggregateSessions([makeProject()])
|
|
rows[0]!.project = 'Users-torukmakto-codeburn-.claude-worktrees-agent-a213f7c77871f483f'
|
|
const output = renderTable(rows, { terminalWidth: 120 })
|
|
|
|
expect(output).toContain('codeburn · agent a213f7c7')
|
|
expect(output).not.toContain('Users-torukmakto')
|
|
expect(output).not.toContain('.claude-worktrees')
|
|
})
|
|
|
|
it('uses captured titles, sorts newest first, and fits the requested width', () => {
|
|
const rows = aggregateSessions([makeProject()])
|
|
const older = { ...rows[0]!, sessionId: 'old', title: 'Older task', startedAt: '2026-07-01T10:00:00.000Z' }
|
|
const newer = {
|
|
...rows[0]!,
|
|
sessionId: 'new',
|
|
title: 'Review the authentication migration without exposing filesystem details',
|
|
project: '-Users-private-Projects-codeburn',
|
|
startedAt: '2026-07-20T10:00:00.000Z',
|
|
}
|
|
const output = renderTable([older, newer], { terminalWidth: 80 })
|
|
const lines = output.split('\n')
|
|
|
|
expect(output.indexOf('Review the')).toBeLessThan(output.indexOf('Older task'))
|
|
expect(output).not.toContain('Users-private')
|
|
expect(Math.max(...lines.slice(0, -1).map(line => line.length))).toBeLessThanOrEqual(80)
|
|
})
|
|
})
|
|
|
|
// A copilot serve set pairs some calls with an already-counted per-turn call
|
|
// (shutdown rollups, residuals, store rows). Their tokens and cost are real and
|
|
// must survive into every spend surface, but they carry no behavioral weight, so
|
|
// no user-visible calls/turns counter may count them.
|
|
function copilotCall(overrides: Partial<ParsedApiCall> & { deduplicationKey: string }): ParsedApiCall {
|
|
return { ...makeProject().sessions[0]!.turns[0]!.assistantCalls[0]!, provider: 'copilot', ...overrides }
|
|
}
|
|
|
|
function makeSupplementaryProject(): ProjectSummary {
|
|
const project = makeProject()
|
|
const session = project.sessions[0]!
|
|
const mixed = session.turns[0]!
|
|
mixed.gitBranch = 'feat/copilot'
|
|
mixed.assistantCalls = [
|
|
copilotCall({ deduplicationKey: 'behavioral-1', costUSD: 0.10 }),
|
|
copilotCall({ deduplicationKey: 'supp-1', costUSD: 0.02, supplementaryAccounting: true }),
|
|
]
|
|
session.turns.push({
|
|
...mixed,
|
|
gitBranch: undefined,
|
|
timestamp: '2026-07-10T10:02:00.000Z',
|
|
assistantCalls: [copilotCall({ deduplicationKey: 'supp-2', costUSD: 0.05, supplementaryAccounting: true })],
|
|
})
|
|
session.apiCalls = 1
|
|
session.totalCostUSD = 0.17
|
|
return project
|
|
}
|
|
|
|
describe('supplementary accounting weight', () => {
|
|
it('counts only turns with a behavioral call in the session rows', () => {
|
|
const rows = aggregateSessions([makeSupplementaryProject()])
|
|
expect(rows[0]!.turns).toBe(1)
|
|
expect(rows[0]!.calls).toBe(1)
|
|
expect(rows[0]!.cost).toBeCloseTo(0.17)
|
|
})
|
|
|
|
it('attributes supplementary spend to a PR while counting only behavioral calls', () => {
|
|
const url = 'https://github.com/acme/app/pull/7'
|
|
const { perUrl } = attributeSessionPrSpend({
|
|
turns: [
|
|
{ prRefs: [url], assistantCalls: [{ costUSD: 0.10 }, { costUSD: 0.02, supplementaryAccounting: true }] },
|
|
// Supplementary-only turn: no calls, but its cost still belongs to the PR.
|
|
{ assistantCalls: [{ costUSD: 0.05, supplementaryAccounting: true }] },
|
|
],
|
|
totalCostUSD: 0.17,
|
|
apiCalls: 1,
|
|
totalSavingsUSD: 0,
|
|
})
|
|
|
|
const pr = perUrl.get(url)!
|
|
expect(pr.calls).toBe(1)
|
|
expect(pr.cost).toBeCloseTo(0.17)
|
|
})
|
|
|
|
it('attributes supplementary spend to a branch while counting only behavioral calls', () => {
|
|
const rows = aggregateByBranch([makeSupplementaryProject()])
|
|
expect(rows).toHaveLength(1)
|
|
expect(rows[0]!.branch).toBe('feat/copilot')
|
|
expect(rows[0]!.calls).toBe(1)
|
|
expect(rows[0]!.cost).toBeCloseTo(0.17)
|
|
})
|
|
})
|