# CodeBurn Architecture A map of the codebase. Read this once before opening a non-trivial PR. ## Three Surfaces CodeBurn is one Node.js CLI plus two GUI clients that shell out to it. ``` +----------------------+ +-----------------+ | mac/ (Swift) | ---> | | +----------------------+ | src/cli.ts | | gnome/ (JavaScript) | ---> | (the CLI) | +----------------------+ | | | status | | --format | | menubar-json | +-----------------+ | v +----------------------------+ | session files on disk | | (JSONL, SQLite, protobuf) | +----------------------------+ ``` The macOS menubar (`mac/`) and the GNOME extension (`gnome/`) both invoke `codeburn status --format menubar-json --period
` and parse the JSON. They do not share code with the CLI; they only depend on its output contract.
## CLI (`src/`)
`src/cli.ts` is the Commander.js entry point. The bin field in `package.json` points at `dist/cli.js`. Twelve commands are registered:
| Command | Line | Purpose |
|---|---|---|
| `report` | 274 | Default. Interactive Ink TUI dashboard. |
| `status` | 358 | Compact text status, plus `--format menubar-json` for clients. |
| `today` | 524 | Today-only view of `report`. |
| `month` | 542 | Month-only view of `report`. |
| `export` | 560 | CSV or JSON dump of usage data. |
| `menubar` | 621 | Downloads and launches the macOS menubar bundle. |
| `currency` | 636 | Sets display currency. |
| `model-alias` | 687 | Maps an unknown model name to a known one for pricing. |
| `plan` | 737 | Configures a subscription plan for overage tracking. |
| `optimize` | 857 | Runs all 14 waste detectors. |
| `compare` | 870 | Compares two models side by side. |
| `yield` | 882 | Tracks which sessions shipped to main vs. were reverted (experimental). |
### Pipeline
```
provider.discoverSessions()
|
v
provider.createSessionParser(source, seenKeys)
|
v yields ParsedProviderCall (see src/providers/types.ts)
|
v
src/parser.ts: parseAllSessions()
|
v aggregates into ProjectSummary[]
|
v
src/daily-cache.ts: aggregate per day, persist
|
v
output formatter (Ink TUI, JSON, or menubar-json)
```
`src/parser.ts` is the central aggregator. Public exports: `parseAllSessions`, `filterProjectsByName`, `extractMcpInventory`. It owns the dedup `Set` (`seenKeys`) that is passed into every provider parser so a turn that surfaces in two providers (Claude logs vs. Cursor mirror, for instance) is counted once.
### Parallel Cold Parse
A cold parse spends most of its time on work that is per-file and pure: reading a
session JSONL or a Codex rollout, decoding it, and turning each line into a
journal entry. `src/parse-workers.ts` moves that onto `worker_threads` when the
pending workload is big enough to pay for them. Each worker runs the same
per-file function the serial path runs — `parseClaudeFileFull` for a Claude
session, `parseCodexFileFull` for a Codex rollout — against an empty dedup set,
and ships the result back as a JSON string together with every dedup key it
claimed. The parent installs results in the same order the serial loop would, and
everything with cross-file state (the dedup sets, canonical project paths, spawn
links, PR correlation, the Codex result cache) stays on the main thread. A file
whose keys were already claimed by an earlier file, or whose worker failed, is
re-parsed in-process — so the output is identical to the serial path either way.
That overlap check is what makes a forked Codex rollout safe: it replays its
parent's token_count history under the parent's key namespace, collides, and is
re-parsed against the real dedup set.
A Codex worker never touches `src/codex-cache.ts`: it returns the cache entry it
would have written and the parent writes it, in install order, so
`flushCodexCache` publishes exactly what a serial parse would. Only whole-file
parses go off-thread; the append/incremental paths (a Claude append, a Codex
byte-offset resume) are untouched and stay in-process. The decision is made per
provider — the Claude scan and the provider loop run one after the other, so at
most one pool is alive — and the pool is terminated when its scan ends, so the
resident `serve` child never accumulates threads.
The pool is off by default for anything that is not a large cold parse:
| Gate | Serial when |
|---|---|
| Pending bytes | under 200 MB behind the pending whole-file parses |
| Cores | `availableParallelism() <= 2` |
| Memory | under 4 GB available |
Otherwise the worker count is
`min(cores - 1, min(0.25 * available, 2 GB) / perWorker, max(pendingFiles / 50, pendingBytes / 200 MB))`.
Files and bytes each earn threads on their own, so a few hundred multi-hundred-MB
Codex rollouts parallelize as well as a few thousand small Claude transcripts. The
gate is bytes only, deliberately: 250 pending files holding under a megabyte
between them spawn threads that make the run ~5% slower, and a file count only
starts paying for itself around 400.
`perWorker` is the per-thread memory budget, derived per parse as
`clamp(256 MB, 2 x (pendingBytes / pendingFiles) + 128 MB, 1 GB)`. A flat figure
was wrong in both directions: small Claude transcripts peak well under 256 MB,
while a 260 MB Codex rollout peaks near 430 MB in its worker and scales linearly
with the pool. The budget also covers the parent, which buffers up to `pool.size`
finished results while it installs one.
"Available" is `process.availableMemory()`, falling back to `os.totalmem()`. It is
deliberately not `os.freemem()`: on macOS that counts free pages rather than
available memory and reads as a few hundred MB on an idle 128 GB machine, so a
gate built on it switches the feature on and off between runs. On Linux outside a
memory-limited cgroup, `availableMemory()` reports free memory and can still
under-report on a busy host — which fails safe, to fewer threads or none.
`CODEBURN_PARSE_WORKERS` overrides the decision and skips every gate above:
`0` forces the serial parse, `N` forces N workers (capped at the core count).
`CODEBURN_VERBOSE=1` prints the resolved worker count and the reason for it.
### Cache Layers
Three caches under `~/.cache/codeburn/` (override with `CODEBURN_CACHE_DIR`):
| File | Owner | Invalidation |
|---|---|---|
| `codex-results.json` | `src/codex-cache.ts` | `mtimeMs + sizeBytes` per Codex `.jsonl`. |
| `cursor-results.json` | `src/cursor-cache.ts` | `mtimeMs + sizeBytes` of the Cursor SQLite db. |
| `daily-cache.json` | `src/daily-cache.ts` | Tracks `lastComputedDate`; new days are backfilled, old days are reused. |
All three use atomic write (temp file + `rename`) and write with mode `0o600`. All three carry a numeric `version` field; bumping it forces a recompute next run.
### Optimize Detectors
`src/optimize.ts` exports 14 detectors. Each returns a `WasteFinding | null`. They are composed by `runOptimize()` which collects findings, ranks them by impact, and returns them with `WasteAction` objects (paste-to-CLAUDE.md, paste-to-session-opener, prompt-now, edit shell config).
| Detector | Line | What it catches |
|---|---|---|
| `detectJunkReads` | 428 | Reads into `node_modules`, `.git`, `dist`, etc. |
| `detectDuplicateReads` | 477 | Re-reads of the same file in a session. |
| `detectMcpToolCoverage` | 795 | MCP servers with many tools but low usage. |
| `detectUnusedMcp` | 855 | MCP servers configured but never invoked. |
| `detectBloatedClaudeMd` | 944 | `CLAUDE.md` files past a healthy size. |
| `detectLowReadEditRatio` | 987 | Edit-heavy sessions with too few prior reads. |
| `detectCacheBloat` | 1048 | High `cache_creation_input_tokens`. |
| `detectGhostAgents` | 1124 | Defined but never-invoked Claude agents. |
| `detectGhostSkills` | 1154 | Defined but never-invoked skills. |
| `detectGhostCommands` | 1184 | Defined but never-invoked slash commands. |
| `detectBashBloat` | 1228 | Shell output limit set above the recommended 15K chars. |
| `detectLowWorthSessions` | 1405 | Sessions with cost but no edits or git delivery. |
| `detectContextBloat` | 1512 | Input:output token ratio above 25:1. |
| `detectSessionOutliers` | 1558 | Sessions costing more than 2x the project average. |
### Output Formats
| Command | `--format` choices | Default |
|---|---|---|
| `report`, `today`, `month` | `tui`, `json` | `tui` |
| `status` | `terminal`, `menubar-json`, `json` | `terminal` |
| `export` | `csv`, `json` | `csv` |
| `plan` | `text`, `json` | `text` |
The macOS menubar and GNOME extension consume `menubar-json`. `src/menubar-json.ts` defines the contract; `tests/menubar-json.test.ts` pins it.
## Providers (`src/providers/`)
Every provider implements the `Provider` interface in `src/providers/types.ts`:
```ts
type Provider = {
name: string
displayName: string
modelDisplayName(model: string): string
toolDisplayName(rawTool: string): string
discoverSessions(): Promise