mirror of
https://github.com/razzant/ouroboros.git
synced 2026-10-03 04:07:04 +00:00
The growth guard was a per-process memo keyed by (inode, device, size), and a committed pass REPLACES the file. So at the instant one process folds, every other process's memo stops matching and re-enters a full pass on its next reservation, and a process that has just started has no memo at all while the residue that cannot fold — group rows, retained idempotent and review-attributed rows, terminal rows younger than the 48-hour fold horizon — holds the file above the trigger for good. Each of those passes rewrites the whole monetary authority under the held money lock and copies the entire live file into a new archive segment. Measured on the owner's live install: 100 passes archiving 2.37 GB in three days for 4.88 MB of live-file gain, with 93 of 99 consecutive passes entered after less than 1 MB of growth. A guard that must hold across processes cannot live inside one of them. The pass already stamps what it needs into line 1: the baseline header records source_size_bytes, the size that pass read. maybe_compact_usage_ledger_locked now declines while the file is below that plus the existing USAGE_LEDGER_COMPACT_RETRY_GROWTH_BYTES, read lock-free through the existing _live_baseline_header (one readline; appends never touch line 1 and the swap is atomic). No new file, no header field, no constant, no lock ordering. The per-process memo stays, because an abort changes no bytes and so leaves the ledger's own floor naming the window that let that pass in: only the memo can throttle the retry. It is also the whole guard on a ledger that states no floor — nothing has folded here yet, the recorded size is not a positive count, or the leading row cannot be read at all. That last case is contained exactly where it was: the floor neither raises into the caller's reservation nor declines in place of the pass, so the corruption still surfaces from the normal ledger read. Disclosed cost: the stamp names the PRE-pass size, so after a high-gain fold the next pass waits until the file outgrows what the last one read. The 20 MB USAGE_LEDGER_WARN_BYTES tripwire and the startup note on state/usage_attempts.jsonl now name that as a third cause of growth, because a decline happens before the pass and leaves no typed event. test_a_committed_fold_arms_the_same_growth_guard_as_an_abort kept its claim but had simulated growth by narrowing the retry window to one byte; it now appends real chains, since the ledger's floor is the size the last pass read.
818 lines
55 KiB
Markdown
818 lines
55 KiB
Markdown
# Usage-ledger compaction and monetary authority
|
||
|
||
Terminal usage history is compacted into a stamped baseline block while the
|
||
exact original ledger bytes remain in an append-only archive. The ledger is
|
||
the monetary authority, so compaction preserves exact money, attribution,
|
||
in-flight attempts and historical joins.
|
||
|
||
## 1. Problem
|
||
|
||
`state/usage_attempts.jsonl` is the append-only monetary authority
|
||
(`ouroboros/usage_ledger.py`). Every reservation re-reads it under the
|
||
cross-process monetary lock; a ~20 MB ledger costs ~0.5 s per full re-read
|
||
under that lock. `USAGE_LEDGER_WARN_BYTES` in
|
||
`ouroboros/context_budget.py` warns at that measured degradation point.
|
||
The in-process warm caches bound the *steady-state* cost,
|
||
but every cold read (process start, refold on any doubt) still replays the
|
||
whole file. Without compaction the live file grows without bound: each
|
||
physical attempt appends a 2–4 row lifecycle chain that remains after it is terminal.
|
||
|
||
## 2. Compacted representation
|
||
|
||
Fold the terminal history into a **stamped baseline block** at the head of the
|
||
ledger and move the raw pre-compaction bytes, verbatim, into an append-only
|
||
archive segment. Nothing is deleted: the archive holds every original row
|
||
(with its original `seq`); the live file holds an exact, structurally
|
||
validated aggregate plus every row that is still live.
|
||
|
||
## 3. What folds, what never folds
|
||
|
||
| rows | disposition | why |
|
||
|---|---|---|
|
||
| `kind="attempt"`, final state `settled` / `released`, no abandonment marker, and **no review attribution** (`review_skill`/`review_wave_id`/`review_slot_id` all empty) | folded (their whole seq chain) after the fold horizon | ordinary terminal rows are immutable; aggregation-complete under §5 |
|
||
| `kind="attempt"`, final state `reserved` / `dispatched` / `unresolved`, or an administrative `settled` row with `settle_reason="abandoned"` | **retained verbatim** | their next valid transition or late receipt still has to join the exact `attempt_id` in the live replay |
|
||
| `usage_baseline` / `usage_baseline_group` from a previous compaction | re-folded (header replaced, groups merged by key, exact-decimal sums added) | baselines must not accumulate per epoch; historical unknown-cost groups remain valid aggregates, never recreated individual attempts |
|
||
| `kind="subscription_session"`, `"external_unmetered"` | **retained** | their `attempt_id` is deterministically re-derived from a stable external id and re-asserted on replay: `_append_single_settled_row` dedups and conflict-checks against the LIVE replay. Folding them would turn an idempotent replay into a silent double charge. Disclosed residual: these rows keep growing (slowly — one row per delegated run / external dispatch). |
|
||
| `kind="legacy_*"` | **retained** | same idempotency argument: `ensure_legacy_imported` dedups candidate rows against live `attempt_id`s if the completion watermark is ever lost mid-history. Bounded one-time set. |
|
||
| attempts with review attribution | **retained** | `skill_review_usage` projects historical waves per-attempt (`attempt_ids`, `attempts` lists) for durable review receipts; folding would erase that projection. Disclosed residual (skill-review waves only; ordinary task/review traffic carries no `review_*` attribution). |
|
||
| unknown future kinds | **retained** | fail-safe default: fold only what this design proves aggregation-complete |
|
||
| quarantine file | untouched | quarantined bytes are already out of the replay; `integrity_degraded` stays path-derived and unaffected |
|
||
|
||
An administrative abandonment uses the existing `settled` state with
|
||
`settle_reason="abandoned"`, `cost_usd=None` and `cost_final=false`. The reservation
|
||
bound remains accounted without becoming confirmed spend; closing ownership does
|
||
not make an unknown price final. `usage_ledger.is_abandoned_settlement` is the shared
|
||
structural predicate, and validation rejects an abandonment marker that claims a
|
||
known price or final cost.
|
||
|
||
An unresolved attempt can become this administrative settlement or accept one real
|
||
settlement tagged `late_receipt`; an abandoned settlement can accept that same late
|
||
receipt. Either may instead become `released` only on the existing positive
|
||
`before_dispatch_failed:` proof that generation never started. After the real
|
||
settlement or release the row is immutable again. Identical actual-receipt retries
|
||
are no-op at the accounting writer; conflicting receipts cannot append. The full
|
||
validator and incremental `LedgerResumeState.late_receipt_ids` preserve the same
|
||
eligibility, so a warm read grants no extra transition and loses no owed receipt.
|
||
Compaction preserves the whole eligible chain, not just its last row. Existing
|
||
baseline groups keep their recorded sums and weights; the archive is not migrated
|
||
to invent individually correctable attempts.
|
||
|
||
## 4. Baseline block shape
|
||
|
||
Subscription-session replay binds the stable external session id, route and all
|
||
ownership/review-attribution fields. Its observed model is disclosure, not a new
|
||
physical session or a pricing input: a later observation can differ or become
|
||
unknown while replay returns the original ledger row byte-identically. Existing
|
||
legacy model labels are historical evidence, not newly confirmed observations;
|
||
new custody/result observations remain separate. External-unmetered replay keeps
|
||
its own existing model identity check.
|
||
|
||
The compacted file is `[header row] [group rows …] [retained rows …]`.
|
||
|
||
**Header** (`kind="usage_baseline"`, exactly one, always seq 1 when present):
|
||
|
||
```json
|
||
{"kind":"usage_baseline","attempt_id":"baseline-<hex12>","state":"settled",
|
||
"seq":1,"ts":"…","baseline_id":"<hex12>","compaction_epoch":N,
|
||
"archive_rel":"archive/usage_ledger/segment_….jsonl",
|
||
"source_sha256":"…","source_size_bytes":B,"source_row_count":R,
|
||
"source_first_seq":1,"source_last_seq":K,
|
||
"folded_row_count":F,"folded_attempt_count":A,
|
||
"group_count":G,"retained_row_count":T}
|
||
```
|
||
|
||
The header is a pure stamp: `_summary`/`_breakdown_bucket` skip it entirely
|
||
(it contributes no money, no counts, no tokens). `source_sha256` is the
|
||
SHA-256 of the archived segment's exact bytes — together with the segment's
|
||
own embedded previous header this forms a tamper-evident hash chain over the
|
||
whole history.
|
||
|
||
**Group rows** (`kind="usage_baseline_group"`): one per attribution tuple
|
||
|
||
```
|
||
key = (state, model, provider, category, source,
|
||
task_id, root_task_id, parent_task_id,
|
||
prompt_cache_ttl, cost_known, cost_final, pricing_known, bound_known)
|
||
```
|
||
|
||
carrying: the key fields verbatim; `folded_attempt_count` (int ≥ 1);
|
||
`cost_usd` / `reservation_upper_bound_usd` as **exact-decimal JSON strings**
|
||
(§5); token sums (`prompt_tokens`, `completion_tokens`, `cached_tokens`,
|
||
`cache_write_tokens`) as ints, absent when no folded row reported the field;
|
||
`root_limit_usd` = min over the group's known values (else absent);
|
||
`baseline_id` joining the header; empty `review_*` attribution.
|
||
|
||
Why per-group rows rather than a single global aggregate:
|
||
budget enforcement is **per-root** (`reserve_attempt` filters finals by
|
||
`root_task_id`; `usage_projection` takes `min` of row `root_limit_usd`), and
|
||
`usage_breakdown` groups by model/provider/category/task/root. A single global
|
||
row would silently zero per-root accounting — a budget bypass. One stamped
|
||
baseline *block* whose group rows preserve the full attribution tuple keeps
|
||
every existing aggregation exact while remaining one atomic, stamped unit.
|
||
|
||
## 5. Monetary exactness rule (the fixed rule)
|
||
|
||
Monetary equality is defined **on decimals, never on float accumulation**:
|
||
|
||
- The compactor parses the source segment with `parse_float=Decimal` and sums
|
||
each group's `cost_usd` / `reservation_upper_bound_usd` as exact `Decimal`s
|
||
of the literals actually stored in the file.
|
||
- Those sums run in an **explicit** decimal context (`_exact_money`:
|
||
`prec = MONEY_PRECISION = 60`, `Inexact` trapped), never the ambient one.
|
||
The interpreter default keeps 28 significant digits, which silently rounds
|
||
a large-magnitude sum — and because the group row and the self-check that
|
||
approves it are computed the same way, a rounded total verifies against
|
||
itself while the float projection stays equal on both sides. Sixty digits
|
||
is far past any real ledger; past even that, the trap turns the loss into
|
||
an abort instead of an approximation. Decimal *construction* is
|
||
context-free by language rule, so stored literals are always captured
|
||
exactly; only arithmetic needs the context.
|
||
- Group sums are stored as exact-decimal **JSON strings** (`"cost_usd":
|
||
"12.3456789"`, `format(dec, "f")`, no exponent). `_number()` — the single
|
||
row-level monetary parser used by validation, `_summary`, and every
|
||
projection — already accepts numeric strings, so no reader changes shape.
|
||
A float re-serialization would round the exact sum to the nearest double;
|
||
the string keeps the invariant byte-checkable forever, including across
|
||
re-compactions (group merges sum the decimal strings exactly).
|
||
- Retained rows are re-serialized from the standard float parse (shortest-repr
|
||
round-trip). The compactor VERIFIES, per retained row, that the re-parsed
|
||
`Decimal` view of the new line equals the `Decimal` view of the original
|
||
line on every field (only `seq` / `pre_compaction_seq` may differ); every
|
||
literal our writers ever emit is shortest-repr so this holds identically,
|
||
and a foreign non-canonical literal (one that is not double-round-trippable)
|
||
triggers **abort, not approximation**.
|
||
- Before committing, the compactor replays the candidate bytes through the
|
||
PRODUCTION aggregation (`_final_rows` → `_summary`, per-root summaries with
|
||
`min` limits, `_breakdown_bucket` global and per model/provider/category/
|
||
task/root axis) and requires the **non-money** rendered dicts to equal the
|
||
corresponding source projection. The separate `decimal_totals` check
|
||
requires exact money equality. Either mismatch aborts the compaction and
|
||
leaves the ledger byte-identical; float accumulation is not an additional
|
||
equality gate.
|
||
|
||
The stored decimal money and non-money projections are preserved before the
|
||
swap. Existing readers still accumulate and round floats, so their displayed
|
||
money can differ across a fold even when the exact ledger sums are identical.
|
||
|
||
## 6. seq policy
|
||
|
||
The live file keeps the substrate's strongest integrity property: **dense
|
||
`seq` from 1** (the validator's density check is what detects mid-file loss
|
||
and tampering). A gap-tolerant validator would have weakened that authority,
|
||
so instead the compacted file starts a fresh dense epoch:
|
||
|
||
- baseline header = seq 1, groups 2..G+1, retained rows follow densely in
|
||
their original relative order;
|
||
- every retained row keeps its original seq as `pre_compaction_seq`;
|
||
- the header records `source_first_seq`/`source_last_seq`, and the archive
|
||
segment holds every original row with its original `seq` untouched.
|
||
|
||
Monotonicity and density are preserved; the original seq
|
||
values are never lost (archive + `pre_compaction_seq`). Nothing durable
|
||
references ledger rows by `seq` (cross-references are `attempt_id`s); resume
|
||
fingerprints are invalidated structurally by the inode change (§8).
|
||
|
||
Validator additions (`_validate_records`): baseline rows are legal only as
|
||
the leading block of a full-file validation (a baseline row in an appended
|
||
tail, or after any non-baseline row, is corrupt); exactly one header, first;
|
||
group rows must carry the header's `baseline_id` and a positive
|
||
`folded_attempt_count`; group state ∈ {settled, unresolved, released}, including
|
||
historical unresolved groups. Per-attempt transitions follow the late-receipt
|
||
contract in §3; numeric checks accept monetary strings through `_number`.
|
||
|
||
**The stamp's own provenance is validated, not trusted.** The header is the
|
||
only ledger row that points at bytes outside its file, so the substrate — the
|
||
thing that decides what a well-formed row IS — checks it:
|
||
|
||
- `compaction_epoch` is a positive int; `source_sha256` is 64 hex;
|
||
- `archive_rel` is *bounded* (`usage_ledger.valid_archive_rel`): relative,
|
||
forward-slash, exactly `archive/usage_ledger/<name>`, no `..`, no absolute
|
||
path, no drive letter, no separator inside the name. A tampered header can
|
||
therefore never aim a reader at a file elsewhere on the host;
|
||
- the counts must **close**: `folded_row_count + retained_row_count ==
|
||
source_row_count`, `source_first_seq == 1`, `source_last_seq ==
|
||
source_row_count`, and `source_size_bytes ≥ 1`;
|
||
- the block must agree with its header: the number of group rows equals
|
||
`group_count` and their `folded_attempt_count`s sum to the header's. Once
|
||
the first non-baseline row closes the block, a later group row — even one
|
||
carrying the real `baseline_id` — is corrupt. That is the money-injection
|
||
shape the position rule exists to refuse;
|
||
- `pre_compaction_seq` is checked as the provenance claim it is: legal only
|
||
under a leading header, strictly increasing, only up to the first
|
||
non-baseline row that carries none (post-compaction appends never carry
|
||
it, and a re-compaction rewrites the whole file), and **inside the source
|
||
range the header declares** (`source_first_seq`..`source_last_seq`) — a
|
||
retained row may not claim to come from bytes nobody archived.
|
||
|
||
## 7. Aggregation contract (`_usage_rows`)
|
||
|
||
`_summary` and `_physical_call_count`/`_breakdown_bucket` are
|
||
baseline-aware in the narrowest way:
|
||
|
||
- `usage_baseline` header: skipped (no money, no counts);
|
||
- `usage_baseline_group`: every **count** increment (`attempt_counts`,
|
||
`unknown_unmetered`, `non_final_rows`, physical calls, `prompt_cache_ttls`)
|
||
uses `weight = folded_attempt_count`; every **sum** adds the row's carried
|
||
aggregate once. For all existing kinds `weight == 1` and the code path is
|
||
unchanged.
|
||
|
||
The group key (§4) makes each group homogeneous in every branch predicate
|
||
`_summary` evaluates per row (`cost is None`, `cost_final`,
|
||
`pricing_known is False`, `bound is None`, state), so the per-group branch is
|
||
exactly the per-row branch taken `weight` times with the sums pre-added.
|
||
|
||
## 8. Concurrency, crash-safety, caches
|
||
|
||
- Compaction runs **only under the same monetary lock** (`_locked(root)`) as
|
||
every read-check-append transaction, invoked from `reserve_attempt`'s
|
||
locked section before its ledger read (§9). No second lock, no new lock
|
||
ordering.
|
||
- **The lock has two tiers, chosen by a capability predicate, and the pass
|
||
runs on one of them only.** `platform_layer.kernel_file_locks_enforced`
|
||
decides once per lock directory by kernel-locking a scratch file there:
|
||
only the kernel's own "this filesystem cannot" selects the **name tier**;
|
||
every other answer is the **enforced tier**. The tier is never chosen by a
|
||
refusal on a live acquisition. The refusals that mean the filesystem takes
|
||
no kernel locks at all are exactly `EOPNOTSUPP`/`ENOTSUP`/`ENOSYS` — on
|
||
Windows `ERROR_INVALID_FUNCTION` and `ERROR_NOT_SUPPORTED`, which
|
||
`_win32_lock_error` maps onto `ENOSYS` and `EOPNOTSUPP`, because CPython's
|
||
own winerror→errno table lands both on `EINVAL` and would leave the name
|
||
tier structurally unreachable there: a Windows volume without byte-range
|
||
locks would fail EVERY monetary append closed instead of degrading to it.
|
||
`ENOLCK` ("no locks available" — a filesystem without a lock daemon, or an
|
||
exhausted kernel lock table) selects the **name tier** too. Making this
|
||
failure close every caller would stop locked state writes, task-result and
|
||
custody updates, and model dispatch on a lockd-less filesystem, although
|
||
those non-monetary consumers can use the name protocol. The probe records
|
||
the errno beside its verdict, and each caller can refuse that tier by errno
|
||
(`acquire_exclusive_file_lock(refuse_name_tier_errnos=…)`).
|
||
Only the monetary lock does: `usage_ledger._named_lock` names `ENOLCK`, so on
|
||
such an install every monetary write refuses typed (`UsageAccountingError` —
|
||
no lock, no append, no pass; money never runs the name protocol where locks
|
||
might merely be unavailable) while every other lock keeps the name protocol
|
||
it always ran there; moving `state/` onto a filesystem that locks is the
|
||
repair. The verdict is cached per process per directory under one module
|
||
lock, so racing threads share one probe and one answer rather than two
|
||
probes that could disagree — except a directory where the scratch probe
|
||
cannot be created, which answers enforced for that call and is probed again
|
||
next time (not cached).
|
||
**Windows takes the enforced tier on a byte range beyond the stamp.**
|
||
Windows byte-range locks are mandatory: locking the whole file prevents a
|
||
contender from reading the owner's stamp, so it cannot judge the hold and
|
||
times out. Falling back to the name tier would expose another constraint:
|
||
the contender's open handle can prevent the owner's release unlink through
|
||
a sharing violation (no `FILE_SHARE_DELETE`). `_unlink_lock_path` retries
|
||
that transient refusal for a bounded window. The enforced-tier hold is one
|
||
byte at
|
||
`platform_layer._WIN32_LOCK_OFFSET` (`0x7FFFFFFF00000000`, length 1 — the
|
||
common Win32 idiom; a lock beyond end-of-file is legal there and no lock
|
||
file's one-line stamp can reach that far), so the bytes a contender reads,
|
||
[0, 512), lie outside every range this protocol locks. The capability probe
|
||
then runs on Windows exactly as on POSIX (scratch file, lock, unlock;
|
||
ERROR_LOCK_VIOLATION = held, ERROR_INVALID_FUNCTION/ERROR_NOT_SUPPORTED = no
|
||
byte-range locks on this volume → the name tier), and the compaction pass
|
||
runs there. Windows eviction on this tier takes the SAME non-blocking kernel
|
||
lock on the judged descriptor as POSIX — a creator stalled between its create
|
||
and its lock is judged by the kernel, not by age alone — but it cannot unlink
|
||
what it holds open, so it releases that hold, closes its probe and only then
|
||
re-checks the identity and unlinks. It therefore does NOT give the POSIX
|
||
guarantee below («of two racing reclaimers at most one can evict», held by
|
||
the kernel across the whole judge → re-check → unlink span): two Windows
|
||
reclaimers may both re-check, and what keeps the shape exclusive across that
|
||
gap is Windows itself — the winner's freshly won lock is held open by its
|
||
owner, so the loser's unlink is REFUSED rather than obeyed (a sharing
|
||
violation the eviction path deliberately does not retry; the poll loop
|
||
re-judges). Release order there is unlock → close → unlink (a handle closed
|
||
with an outstanding lock leaves the release undefined; the unlink still
|
||
retries a contender's transient refusal). What Windows still does not have on
|
||
this tier is named elsewhere in this section and unchanged: no directory
|
||
fsync, and no old-inode witness across `os.replace`, so a charge landed in
|
||
the swap's last syscall is lost silently there. Windows execution is checked
|
||
by the CI matrix; host-side pins (the range constant and its two wrappers, an emulated LockFileEx that
|
||
refuses the same range, the delete-semantics simulator) stand in for the
|
||
mechanism, never for the platform.
|
||
*Enforced tier* (POSIX `fcntl.flock`, Windows `LockFileEx` —
|
||
both held on the lock fd): every other hold of this lock is milliseconds; a compaction
|
||
pass over a multi-megabyte ledger can legitimately exceed its 90 s
|
||
staleness window, and a lock evicted purely by age would put a second
|
||
writer on the same authority. So the acquisition takes a **kernel lock on
|
||
the lock fd** in addition to the O_EXCL name. The refusals that mean HELD
|
||
BY SOMEONE — stand down and re-contend — are exactly `EAGAIN`/`EWOULDBLOCK`
|
||
(on Windows `ERROR_LOCK_VIOLATION` alone, mapped onto the first); every
|
||
other refusal, `EACCES` included, **fails closed** — no descriptor, our own
|
||
file removed, a stale lock never evicted without the held flock, never a
|
||
silent fall back to the name protocol — `_named_lock` acquires
|
||
**owner-aware** (a live owner PID is never evicted on age), and the hold
|
||
yields a **heartbeat**
|
||
(`platform_layer.refresh_exclusive_file_lock`), REQUIRED by both entry
|
||
points rather than defaulted — an omitted heartbeat would turn every proof
|
||
below into a no-op and swap the authority unproven, so a caller that drops
|
||
it gets a TypeError — and the pass beats at every checkpoint — inside the long span (both row walks of the candidate build
|
||
and each verification stage), then **immediately before each snapshot
|
||
re-check, including the final one inside the atomic replace** (writing and
|
||
fsyncing the candidate temp can take arbitrarily long, so the proof of
|
||
ownership runs once the temp is durable, right before the rename — ownership
|
||
first, then the snapshot): a re-check answered while the lock already
|
||
belongs to someone else is a meaningless answer, so the loss must abort the
|
||
pass before the answer is even asked. The lock primitives are
|
||
ownership-exact and kernel-guarded: a stale eviction unlinks only while
|
||
HOLDING the flock on the very fd it judged abandoned (path re-checked
|
||
against it), so of two racing reclaimers at most one can evict; on POSIX a
|
||
release unlinks before its close, under the still-held flock; the heartbeat
|
||
answers OWNERSHIP rather than success, including for a lock file replaced
|
||
atomically (the path never absent), and answers `False` — never a renewal —
|
||
when its own descriptor's identity cannot be read or the kernel refuses the
|
||
`utime`. Windows cannot unlink an open file: its eviction takes the same
|
||
probe lock and then re-checks the path after releasing and closing that
|
||
probe, AND its release re-checks after its own close, a freshly won lock —
|
||
held open by its owner — being undeletable there, which is what keeps that
|
||
shape exclusive across the gap the kernel does not cover (above). Disclosed, both tiers: a contention
|
||
answer (`EAGAIN`) on the creator's OWN fresh file — a foreign flock holder
|
||
that never unlinks it — leaves that file on the path, stamped with the
|
||
creator's live pid; the creator re-contends against it until its timeout,
|
||
and owner-aware acquirers never age it out while this process lives. No
|
||
in-protocol holder produces that shape (an evictor's flock unlinks what it
|
||
judged), so it is theoretical on the enforced tier. The acquisition is identity-checked too: a creator
|
||
stalled between its O_EXCL create and its kernel lock (SIGSTOP, a suspend,
|
||
a debugger, clock skew) can be judged abandoned and evicted, and its lock
|
||
then lands on an inode the path no longer names — not a hold: a won lock is
|
||
returned only while the path still names its descriptor, else the creator
|
||
closes it and re-contends, and the owner pid is written BEFORE the lock so
|
||
an owner-aware reclaimer never judges a live creator's fresh file empty
|
||
(every caller of the primitive gets this, the age-only non-monetary locks
|
||
included: at most one descriptor ever answers, the cost is a re-contention).
|
||
A descriptor whose OWN identity cannot be read — `fstat` answering ESTALE
|
||
or EIO on the very filesystems this tier exists for — proves neither
|
||
ownership nor eviction, so it is not a hold either: comparing two unreadable
|
||
identities raw made them equal and returned a descriptor for an unlinked
|
||
inode. It fails closed, and the file we stamped with our LIVE pid is removed
|
||
with it when its bytes are still exactly the ones we wrote — left behind, no
|
||
owner-aware reclaimer could ever evict it and the lock wedges for good.
|
||
**Residual, disclosed:** the owner-aware rule asks
|
||
`pid_is_alive(owner_pid)`, and a recycled PID reads as alive whoever now
|
||
owns it: `kill(0)` succeeds for the same user and answers `EPERM` for
|
||
another user's process; both mean alive. A lock whose owner died and whose
|
||
PID was reused is never reclaimed by age while the impostor lives (`pid_is_alive` is the
|
||
ONE liveness primitive, shared by every consumer — custody settlements,
|
||
claim reclaims, staging reaps — so a pid recycled onto another user's
|
||
process reads alive everywhere and those defer while the impostor lives,
|
||
exactly as they do for a same-uid recycle): the wedge begins once
|
||
the dead owner's file is older than the 90 s staleness window
|
||
(`usage_ledger._locked`, `stale_sec=90.0` — a literal there, not a
|
||
`config.py` constant) and ends when the impostor exits, even though the
|
||
enforced tier's probe flock would settle it at once — deliberately not
|
||
consulted while the pid reads alive, because a mixed-tier name-tier holder
|
||
holds no flock; `state/usage_attempts.lock` needs a hand repair meanwhile
|
||
(`docs/PERSISTENCE.md` says so in its row).
|
||
On this tier we do not claim a pass can never be robbed. The bounded claim:
|
||
a concurrent holder can exist only after the lock file is removed by an
|
||
actor OUTSIDE the lock protocol — a hand repair, a foreign helper, a
|
||
name-tier process of a mixed-tier install — because in-protocol eviction
|
||
is impossible under the held flock and the heartbeat-fresh mtime; such a
|
||
robbery is caught at the next proof, and ownership is proven at every
|
||
checkpoint, immediately before every rename attempt and once more AFTER
|
||
the in-swap snapshot look, so the irreducible residual is the interval
|
||
between that last proof and the rename syscall: a charge the robber lands
|
||
inside it is erased by the swap. A post-swap re-read cannot detect that
|
||
loss: it compares the new inode against the candidate, while the archive
|
||
contains the pre-row snapshot. On POSIX the swap therefore holds the old
|
||
inode open across the rename (the only witness left) and reads whatever
|
||
landed beyond the proven snapshot's length AFTER the fact: those bytes go
|
||
to `state/usage_attempts.quarantine.jsonl` (`raw_base64`, the shape a torn
|
||
tail takes, which flips `integrity_degraded`) and the pass raises
|
||
`UsageLedgerCorrupt` instead of returning a receipt — the charge is still
|
||
gone from the live file (never re-appended: `seq` belongs to the live
|
||
file), but it is preserved, flagged and typed. Detected by size: a
|
||
same-size in-place rewrite inside that one syscall is not a landed charge
|
||
and is not seen. Windows cannot hold the destination open through
|
||
`os.replace`, so there the loss stays silent, disclosed. A `False`
|
||
heartbeat — or one that cannot be answered at all — aborts the pass,
|
||
leaving the ledger byte-identical.
|
||
*Name tier* (a filesystem the kernel says cannot lock — one answering
|
||
`EOPNOTSUPP`/`ENOSYS` to `flock`; a lockd-less NFS answers `ENOLCK` and is
|
||
fail-closed, above): the O_EXCL name protocol runs alone with re-check-then-unlink
|
||
eviction. There is no kernel exclusion, so a heartbeat there is an identity
|
||
check, not a proof. The pass therefore **refuses to run at all** on the
|
||
name tier (`usage_compaction.NAME_TIER_REFUSAL`: logged, and written ONCE
|
||
per process per data root as a typed `usage_ledger_compaction_refused`
|
||
event to `logs/events.jsonl`; the ledger stays uncompacted and the 20 MB
|
||
tripwire names this tier as a cause), while ordinary monetary appends
|
||
continue under the name protocol as a disclosed best effort: on that tier
|
||
the appends are only as exclusive as the name protocol, which is exactly
|
||
why the whole-file rewrite is withheld. Residual, disclosed: the tier is
|
||
decided per process per lock directory, so a lockd dying mid-run can leave
|
||
one process on each tier until restart — the name-tier process never
|
||
compacts, but it evicts by NAME with no kernel hold, so in that mixed mode
|
||
the two-writer class returns for the enforced-tier process's heartbeat-less
|
||
APPENDS (its compaction pass is still guarded by its proofs).
|
||
- **The swap re-proves its snapshot — before, at the last instant, and
|
||
after.** Because the swap replaces the WHOLE file, the pass re-reads the
|
||
live ledger under the same held lock before the archive write and again
|
||
before the rename, and the atomic writer evaluates one more re-proof
|
||
**inside the swap** — after the candidate temp bytes are durable,
|
||
immediately before `os.replace`, and again before EVERY retried attempt
|
||
(`utils.replace_atomic` retries a Windows sharing violation with pauses,
|
||
and a proof taken before a refused attempt is stale by the next one), the
|
||
last instant each replace can still be refused, with the ownership
|
||
heartbeat asked FIRST, the snapshot compare only under a proven hold, and
|
||
the heartbeat asked AGAIN after that compare — so a hold lost while the
|
||
temp was written, between two attempts, or during the compare itself
|
||
refuses the replace, a row that lands after the outer re-check or between
|
||
attempts survives too, and the only interval left between the last proof
|
||
and the rename is the syscall, not the milliseconds of a full-file
|
||
compare. The cost of any lost race is one skipped pass, never a charge.
|
||
(Belt to the lock's braces: it also covers a lock broken by a hand repair
|
||
or an older build — up to that last proof; the bounded claim above names
|
||
the residual.) The compare→replace window is closed structurally as well: EVERY
|
||
writer of this ledger appends under this same owner-aware lock — kernel-held
|
||
on the enforced tier, the only tier the pass runs on — and there is no
|
||
unlocked fallback: an acquisition that times out, or that the kernel
|
||
refuses, raises `UsageAccountingError` and writes nothing. After the rename the pass
|
||
re-reads what landed and refuses to report a receipt for bytes that are not
|
||
there.
|
||
- Commit order: (1) build + fully verify the candidate in memory (§5, §6);
|
||
(2) write the archive segment — exact source bytes — via O_EXCL write and
|
||
`fsync` the file **and every directory in the chain, on every pass**: the
|
||
segment's parent, `archive/`, and the data root, since syncing a directory
|
||
persists only the entries IT holds. Unconditional, not only where this pass
|
||
created a level: an earlier pass may have created one and died before its
|
||
fsync, so on a retry the directories exist while their durability does not.
|
||
The archive path must also BE this data root's own — a symlink at
|
||
`archive/` or `archive/usage_ledger` aborts the pass, because history must
|
||
never be written through a link — and on POSIX that bound is enforced **at
|
||
the write itself**, not only at a preceding check: the chain
|
||
root→`archive/`→`usage_ledger` is opened `O_DIRECTORY|O_NOFOLLOW`
|
||
handle-to-handle (`dir_fd`), the segment is created `O_NOFOLLOW` relative
|
||
to the held handle, and file plus every directory HANDLE are fsync'd — so
|
||
the chain proven durable is the one the bytes actually landed in, and a
|
||
link planted between check and write is an abort, not a redirect. On POSIX
|
||
a directory-fsync failure **aborts the pass** (an unsynced archive
|
||
directory plus a completed swap is exactly the crash that loses history);
|
||
Windows has no `dir_fd`/`O_DIRECTORY` and no directory handle to fsync, so
|
||
it keeps the path-based chain as a disclosed best effort chosen by the
|
||
platform predicate, never by swallowing `OSError`; (3) atomically replace
|
||
the live ledger (`_write_bytes_atomic_fsync`), which re-proves the snapshot
|
||
one last time before its `os.replace` (the swap bullet above); (4) emit the
|
||
`usage_ledger_compacted` event. A crash before (3) leaves the ledger
|
||
byte-identical (an orphaned archive segment is harmless and disclosed); a
|
||
crash during (3) leaves either the old or the new file — both valid.
|
||
- Cache coherence is structural, not cooperative: the swap changes the inode,
|
||
so every resume fingerprint (`_LEDGER_READ_CACHE`, `_ROWS_MEMO`) refuses to
|
||
warm-resume and refolds from the new file. No cache is asked to remember to
|
||
invalidate.
|
||
|
||
## 9. Trigger policy (config SSOT, no env knob)
|
||
|
||
- `ouroboros/config.py`: `USAGE_LEDGER_COMPACT_BYTES = 8_000_000` (compact at
|
||
~0.2 s-per-replay scale, well under the 20 MB measured-degradation WARN) and
|
||
`USAGE_LEDGER_COMPACT_RETRY_GROWTH_BYTES = 1_000_000`. Constants, not env
|
||
handles.
|
||
- `reserve_attempt` calls `maybe_compact_usage_ledger_locked(root)` at the top
|
||
of its locked section: an `os.stat` fast-path (~µs) below the threshold;
|
||
above it, one compaction pass on exactly the path whose lock-hold the file
|
||
size degrades. Every failure inside compaction is contained and never fails
|
||
the reservation: a policy abort (`_Abort`) is logged AND records the typed
|
||
`usage_ledger_compaction_skipped` event naming its reason, once per process
|
||
per cause. Disclosed: the two snapshot-integrity abandons (a row that landed
|
||
under the lock between the proven snapshot and the archive write, or between
|
||
that write and the swap) log a warning and return without the typed event,
|
||
because they report a lost race the next pass simply repeats rather than a
|
||
cause an operator has to diagnose. A structurally corrupt ledger still fails
|
||
in the normal read path with the normal error.
|
||
- Thrash guard, in two parts, and a pass runs only past both. The threshold
|
||
alone is no brake: the unfoldable residue (group rows, retained idempotent
|
||
and review-attributed rows, terminal rows younger than the fold horizon)
|
||
only grows, so once it reaches the trigger every reservation would run a
|
||
full rewrite of the authority under the held lock and copy the whole live
|
||
file into a new archive segment for a gain of a few kilobytes. Profitability
|
||
is not the question the guard asks; a pass is worth its cost only after real
|
||
growth.
|
||
- **The floor is the ledger's own, so it holds across processes.** The
|
||
header the last committed pass stamped into line 1 records
|
||
`source_size_bytes`, the size that pass READ, and no pass runs until the
|
||
live file reaches it plus `…_RETRY_GROWTH_BYTES`. Read lock-free by
|
||
`_live_baseline_header` (one `readline`; appends never touch line 1 and
|
||
the swap is atomic), so every process answers from the same number. This
|
||
is what a per-process memo cannot do: a committed pass REPLACES the file,
|
||
so at that instant every other process's memo stops matching and re-enters
|
||
a pass, and a freshly started process has no memo at all while the residue
|
||
holds the file above the trigger for good. Measured on the owner's live
|
||
install before this floor: 100 passes archiving 2.37 GB in three days for
|
||
4.88 MB of live-file gain, 93 of 99 consecutive passes entered after less
|
||
than 1 MB of growth. **Disclosed cost:** the stamp names the PRE-pass
|
||
size, so after a high-gain fold the next pass waits until the file
|
||
outgrows what the last one read — deliberately conservative, and the 20 MB
|
||
tripwire below is what watches it. No floor is claimed when the ledger
|
||
states none: no stamp (nothing has folded here yet), a leading row that
|
||
cannot be read at all (the caller's own read reports that corruption — the
|
||
guard neither raises it early nor stands in for the pass), or a recorded
|
||
size that is not a positive count.
|
||
- **The per-process memo remains, for the pass that changed nothing.** It
|
||
holds the last attempted (inode, size) and arms after ANY pass —
|
||
unprofitable (nothing foldable / no shrink / verify-abort) or committed,
|
||
where it takes the COMPACTED size and the new inode. An abort leaves the
|
||
bytes untouched, so the ledger's own floor still names the window that let
|
||
that pass in; only the memo can throttle the retry. It is also the whole
|
||
guard on a ledger that states no floor.
|
||
- `USAGE_LEDGER_WARN_BYTES` (20 MB) stays as the regression tripwire above the
|
||
mechanism, exactly like the rotation-bounded log warns: it fires if
|
||
compaction is broken, if the unfoldable residue itself reaches 20 MB, or if
|
||
the floor above is still holding a ledger that has not outgrown what its
|
||
last pass read. The startup note on `state/usage_attempts.jsonl`
|
||
(`agent_startup_checks._hot_store_thresholds`) names all three, because the
|
||
last one is declined before the pass and so leaves no typed event.
|
||
|
||
## 10. History readers: model-send reconciliation and audits
|
||
|
||
The reverse reconciliation in
|
||
[Model-send observability](MODEL_SEND_OBSERVABILITY.md#33-reverse-direction-audit-model_send-only)
|
||
(`ouroboros/model_send_seal.py`) joins model-send seals to usage-accounting
|
||
attempts. A folded attempt is absent from the live replay, so the sweep uses
|
||
the archive-aware join:
|
||
|
||
- `usage_compaction.archived_attempt_ids(root)` — the `attempt_id` set of
|
||
every archived segment, walked through the tamper-evident header chain
|
||
(live header → segment; segment's own embedded header → older segment, …).
|
||
Each hop is verified, not trusted:
|
||
- the reference is bounded twice — by the substrate's textual rule (§6) and
|
||
by the RESOLVED path having the archive directory as its parent, so a
|
||
symlink planted in the archive cannot import a foreign file's ids;
|
||
- the segment must match the naming header's `source_sha256`,
|
||
`source_size_bytes` and `source_row_count`, and must itself **validate as
|
||
a ledger** (`_validate_records`: dense seq, legal transitions,
|
||
well-formed rows). A tamperer who also repairs the hash still has to
|
||
produce a structurally valid former generation;
|
||
- **the chain's epochs must step down by exactly one and end at epoch 1
|
||
with a header-less segment.** The source of epoch N is exactly the file
|
||
epoch N-1 produced, so this holds by construction — and it is what makes
|
||
re-pointing a live header at an older *genuine* segment (correct hash,
|
||
correct counts) corruption rather than a legitimately shorter history;
|
||
- **the archive anchors the live file.** The step-down rule alone only
|
||
proves the chain BELOW the header, and `compaction_epoch` is as mutable as
|
||
the rest of the row, so a forgery that repoints AND lowers the epoch walks
|
||
a valid, short chain. The generations such a forgery orphans are still on
|
||
disk: no segment may carry a generation newer than the live file's own,
|
||
read from each segment's own embedded header rather than from its name —
|
||
a live file with NO stamp being the same question with the floor at zero,
|
||
so a stamp that was REMOVED is a corruption verdict rather than an archive
|
||
nobody consulted. The one legal newer case is an uncommitted orphan of the
|
||
CURRENT generation (a lost snapshot race, a crash before the swap), and it
|
||
is recognised by what it IS: the pass writes that segment BEFORE the swap,
|
||
so it is the byte-for-byte copy of the live file at that instant, and the
|
||
live file only grows behind it — the orphan's bytes are still a PREFIX of
|
||
it, and every id it holds is live, read from the descriptor the entry was
|
||
classified through (one open per entry, never a re-open of the name, which
|
||
could name a different file by then). Matching only its leading row against
|
||
the live header was not that proof: the newest segment IS the previous
|
||
generation's whole file, so a live file restored from a backup taken just
|
||
after that compaction matched the row while being a strict SUBSET of what
|
||
the segment held, and the attempts that compaction folded — which exist
|
||
nowhere else — were reported absent by the join. Silent absence is the one
|
||
verdict this surface may never reach, so the prefix is the test and a
|
||
restored generation is corruption. The anchor scan runs in the directory the chain
|
||
was walked in — on POSIX through the `O_DIRECTORY|O_NOFOLLOW` dir-fd
|
||
held from the moment the handles are opened, after the live header is
|
||
read, for the rest of the question, entries opened relative to it — so a
|
||
directory swapped after the walk cannot hide a generation from it (a
|
||
directory swapped BEFORE the handles open is the same power as deleting
|
||
the newer segments: disclosed, out of the anchor's reach), and an entry
|
||
the scan cannot list, open or read is `UsageLedgerCorrupt`: the scan did
|
||
not complete, so the question is UNKNOWN — the data root's own handle
|
||
included, since a bare `OSError` from THAT one open (an unreadable root,
|
||
fd exhaustion) would escape the sweep's UNKNOWN mapping entirely. An entry
|
||
that OPENS but is not a regular file — a stray `backup/` directory, a
|
||
FIFO, a device — is no segment: segments are regular files by
|
||
construction, no generation lives there, and it is skipped; one the
|
||
kernel refuses to open at all (a UNIX socket, `ENXIO`) is corruption like
|
||
any other unopenable entry — on the dir-fd shape; without a held handle
|
||
(Windows, no `O_DIRECTORY`) the `stat` before the open classifies a socket
|
||
as not regular and skips it. Where a dir-fd is held that classification is
|
||
an `fstat` after an `O_NONBLOCK` open; where none is (Windows, or any os
|
||
without `O_DIRECTORY`) it is a `stat` BEFORE the open, because there the
|
||
open is the step a directory refuses and a writer-less FIFO blocks on. A
|
||
first row that reads but does not parse is a torn segment from a crashed
|
||
write: no evidence of any generation, left to the walk. Every path
|
||
inspection the reader makes is typed the same way: `pathlib`
|
||
re-raises every `OSError` but `ENOENT`/`ENOTDIR`/`EBADF`/`ELOOP` from
|
||
`is_symlink`/`is_dir`, so the symlink bounds on both archive levels and
|
||
on the named segment (an `archive/usage_ledger` readable but not
|
||
searchable — the shape a `chmod -R 600 data/` hardening produces —
|
||
refused the segment's own `lstat`) raise `UsageLedgerCorrupt`, never a
|
||
bare `OSError`; and on a STAMP-LESS live file the question ends early
|
||
only on the kernel's exact "no archive directory" (`ENOENT`) — a regular
|
||
file standing where the directory belongs, or an archive that cannot be
|
||
inspected, is UNKNOWN (typed), never a silent empty answer — and every
|
||
case that reads anything then goes through the same typed root open.
|
||
Disclosed, stamp-less anchor: with the floor at epoch zero every parsable
|
||
regular file in `archive/usage_ledger/` that is not a byte-prefix of a
|
||
stamp-less live file is `generation newer` — a fresh ledger started
|
||
beside a surviving archive (a reset that deleted the ledger alone) leaves
|
||
every history question a `generation newer` corruption verdict until the
|
||
fresh ledger has compacted as many times as the old one had — after which
|
||
the surviving segments fall under the epoch floor and are silently ignored,
|
||
their ids absent from every history question — and an operator's stray `notes.jsonl` there is skipped on a
|
||
stamped ledger but is corruption on a stamp-less one; move or delete the
|
||
archive with the ledger, or keep both. Disclosed, lock-free readers: a
|
||
compaction that commits between a question's live-header read and its
|
||
anchor scan makes that ONE question UNKNOWN (`generation newer`: the
|
||
freshly committed segment is prefix-tested against the new live file);
|
||
the next question walks the new chain;
|
||
- the archive directory must be **this data root's own**: neither
|
||
`archive/` nor `archive/usage_ledger` may be a symlink, the resolved
|
||
directory must be exactly the resolved root's archive path, and no segment
|
||
may be a link. Otherwise "the segment resolves next to the archive
|
||
directory" is a tautology — both sides resolve through the same link. On
|
||
POSIX the reader enforces this **at the open**: the segment is opened
|
||
`O_NOFOLLOW` through the same `O_DIRECTORY|O_NOFOLLOW` `dir_fd` handles
|
||
the writer uses — opened once, after the live header read, and held for
|
||
the rest of the question, chain walk and epoch anchor alike — and is
|
||
fingerprinted and read from that fd — so a link planted after any
|
||
path-based look is an open error, never a read through it (a
|
||
byte-identical copy behind a link hashes perfectly; only refusing the
|
||
traversal itself defends); a named segment that opens but is not a
|
||
regular file, or whose read fails, is typed `UsageLedgerCorrupt`, never a
|
||
bare `OSError` the sweep's UNKNOWN mapping would miss. Windows keeps the
|
||
path-based checks, best effort, by the platform predicate;
|
||
- per-segment results are cached by path, but the hit requires the file's
|
||
fingerprint (inode/device/size/mtime_ns) to match, so a segment deleted
|
||
or rewritten after a warm read surfaces as `UsageLedgerCorrupt` instead
|
||
of keeping an audit's "logged" verdict alive. A fingerprint is not proof
|
||
of identity, though: an in-place same-size rewrite within the filesystem's
|
||
timestamp granularity keeps it. So a hit ALSO requires a file whose mtime
|
||
has settled (> 2 s old) and an entry younger than 60 s — a cached read is
|
||
evidence with a shelf life. **Residual, disclosed:** a rewrite that
|
||
restores `mtime_ns` exactly can still be answered from a warm entry for up
|
||
to that minute. Closing it would mean re-hashing every segment on every
|
||
question, which is the quadratic cost this cache exists to avoid; the
|
||
chain hash still binds every segment the answer depends on, and any
|
||
process that starts, or asks a minute later, re-hashes;
|
||
- the union over a whole chain is cached by the chain's identity
|
||
((`archive_rel`, sha) per hop). A bulk seal audit selects its manifests
|
||
before reading live attempts, then lazily obtains ONE validated archived-id
|
||
set for that pass. It retains an UNKNOWN archive answer for the rest of the
|
||
reverse pass without suppressing independent forward checks. The next pass
|
||
reads anew with the existing fingerprint, hash and TTL checks; TTL expiry
|
||
inside a batch does not multiply the history read by its seal count.
|
||
Honest about the rest of each archive query's cost: the epoch
|
||
anchor runs on every question, ahead of that cache — one directory listing
|
||
plus a bounded first-row read per entry the chain did not walk, and a full
|
||
read only of an entry that claims a newer generation. The union cache is
|
||
bounded (the chain key changes at every compaction and only the newest
|
||
chain can be asked again), so a long-lived process does not keep one
|
||
archived-id set per epoch it has ever seen.
|
||
- `usage_compaction.usage_attempt_recorded(root, attempt_id, live_ids)` —
|
||
membership in live replay ∪ archive.
|
||
- A live leading row that cannot be **read at all** (bad UTF-8, bad JSON, not
|
||
an object) raises `UsageLedgerCorrupt` rather than returning "no baseline
|
||
header". `None` from `_live_baseline_header` means a readable first row
|
||
that is not a stamp — which is either "no compaction has happened" or "one
|
||
has and its stamp is gone". That function cannot tell them apart and no
|
||
longer claims to: the archive does, because the epoch anchor runs on a
|
||
stamp-less file too, and the wrong answer there turns a folded attempt into
|
||
a reported orphan seal.
|
||
|
||
The model-send reconciliation contract is that the
|
||
reverse sweep's "no attempt row" verdict (`orphan_seal`) must consult this
|
||
union, not the live replay alone; an unreadable/mismatched segment is the
|
||
sweep's existing UNKNOWN → skip-pass case (fail-soft, the API raises typed
|
||
`UsageLedgerCorrupt`). The baseline header in the live file is the structural
|
||
signal that the live replay is not the full per-attempt history.
|
||
|
||
`legacy-import` needs no such lookup: its rows are never folded (§3), so its
|
||
existing live-replay dedup keeps working under a lost watermark.
|
||
|
||
## 11. Module placement
|
||
|
||
`ouroboros/usage_compaction.py` (domain D16) owns fold policy +
|
||
archive/verify/swap + history readers. It imports FROM `usage_ledger`
|
||
(substrate) and `_usage_rows` (aggregation leaf); `usage_accounting` calls
|
||
INTO it from `reserve_attempt`. The substrate stays policy-free (it learns
|
||
only the baseline row kinds' validation), the one-way seam
|
||
`usage_ledger ← usage_accounting` is unchanged, and the compactor — which must
|
||
know the aggregation semantics — lives beside the aggregation, not inside the
|
||
byte authority.
|
||
|
||
## 12. Invariants (pinned by tests/test_usage_compaction.py — the pass side,
|
||
invariants 1-4, 6, 7, 9, 10 — and tests/test_usage_compaction_archive.py —
|
||
the archive reader, invariants 5 and 8; shared fixtures in
|
||
tests/fixtures_usage_compaction.py)
|
||
|
||
1. **Byte-exact money**: decimal sums of `cost_usd` /
|
||
`reservation_upper_bound_usd` over finals are identical before/after, and
|
||
the NON-money projection of `usage_projection` (global + per-root incl.
|
||
limits) and of `usage_breakdown` (all axes) renders equal dicts: state
|
||
counts and folded weights, physical calls, token sums, finality,
|
||
subscription sessions, per-root limits, every axis shape. The float dollars
|
||
those renders carry are deliberately NOT part of that equality.
|
||
Readers round money at six places, so one history summed per row and summed
|
||
per group can land on either side of that boundary. The regression fixture
|
||
in `tests/test_usage_compaction_fingerprint.py` observes a 1e-6 USD shift,
|
||
a ten-thousandth of a cent: `settled_usd` 2.467588 becomes 2.467589 while
|
||
the exact decimal sum is unchanged. This observation is not a universal
|
||
float-drift bound enforced by the compactor.
|
||
Comparing the float view as well would abort correct folds forever instead
|
||
of protecting a cent; the money itself stays exact by the decimal check
|
||
that runs beside it. A sum needing more than the ambient 28 digits
|
||
(10²⁸ + 1) keeps its last digit — pinned by an oracle summing in its own,
|
||
wider context.
|
||
2. **Correctable attempts never fold**: reserved/dispatched/unresolved and
|
||
administratively abandoned chains survive verbatim (modulo seq). The exact
|
||
attempt accepts its valid terminal transition after compaction, while an actual
|
||
late settlement/release becomes foldable under the ordinary rules. Historical
|
||
baseline groups remain aggregates (`tests/test_usage_abandoned_ledger.py`).
|
||
3. **Crash-safety**: a failure injected at the ledger rename ITSELF leaves a
|
||
byte-identical, valid, further-usable ledger, with the archive segment
|
||
already on disk holding the exact source bytes; the archive directory
|
||
chain is fsync'd before the swap — including on the RETRY after a pass that
|
||
died on that fsync — and a POSIX directory-fsync failure aborts with the
|
||
ledger untouched. The live file's own half is pinned the same way: the
|
||
candidate temp is fsync'd BEFORE the rename (without it the renamed inode
|
||
can hold unwritten data — neither the old ledger nor the approved new one)
|
||
and the ledger's directory after it.
|
||
4. **Budget limits are preserved**: root/global enforcement thresholds are
|
||
unchanged across compaction.
|
||
5. **Model-send join survives**: every pre-compaction `attempt_id` remains
|
||
resolvable through live ∪ archive, across chained compactions; a tampered
|
||
segment, a re-hashed but structurally broken segment, a deleted segment
|
||
behind a warm cache, a same-size rewrite once the cache window closes, an
|
||
out-of-archive reference, a symlinked archive directory, an epoch-skipping
|
||
chain, a rollback the archive out-anchors — a previous generation restored
|
||
over the live file, whether or not the stamp came with it — a directory
|
||
swapped between the chain walk and the anchor scan (a look-alike carrying
|
||
the newest NAME with the forged live header included), an archive entry the
|
||
anchor cannot open (the data root's own handle included), a named segment
|
||
that is not a regular file, an unreadable leading row, and a path
|
||
inspection the kernel refuses (the symlink bound on either archive level
|
||
or on the named segment; the stamp-less archive check, which ends the
|
||
question early only on an exact `ENOENT`) are each typed corruption
|
||
(UNKNOWN/skip), never silent absence and never a bare `OSError`; an
|
||
uncommitted orphan segment of the live generation is NOT corruption (its
|
||
bytes are still a prefix of the live file), a stray directory or FIFO in
|
||
the archive — an entry that opens but is not a regular file — is no
|
||
segment and is skipped on both shapes of the scan (the FIFO cannot hang
|
||
the question; an entry the kernel refuses to open at all, a UNIX socket,
|
||
is corruption on the dir-fd shape and skipped by the path shape's
|
||
stat-before-open), and the chain union is built once per chain and cached
|
||
within a bound.
|
||
6. **Idempotency survives**: subscription/external replays after compaction
|
||
dedup (no double charge) and still conflict-check; legacy import stays
|
||
correct with and without its watermark.
|
||
7. **Trigger policy**: no compaction below threshold; the growth guard holds
|
||
ACROSS processes — the floor a pass stamps into the ledger throttles every
|
||
other process and every freshly started one, and the per-process memo still
|
||
throttles a pass that aborted, a ledger that states no floor, and one whose
|
||
leading row cannot be read (which stays the normal read's error to report);
|
||
verify-abort leaves the ledger untouched; the pass is entered with the
|
||
monetary lock demonstrably HELD *and* with that lock's heartbeat wired
|
||
through — the parameter is required, so a caller that drops it raises
|
||
instead of proving nothing.
|
||
8. **Structure**: baseline rows only at head — a header is refused for its
|
||
POSITION while the identical block validates at the head, and a group row
|
||
cannot rejoin a closed block; header counts must close and agree with the
|
||
block; `pre_compaction_seq` is unique, increasing, and legal only under a
|
||
stamp; quarantine/`integrity_degraded` semantics unchanged.
|
||
9. **No pass loses a concurrent charge**: a row appended between snapshot and
|
||
swap aborts the pass and survives byte-for-byte, with money equal to
|
||
before-plus-that-row — including a row that lands between the pre-swap
|
||
re-check and the rename, refused by the re-proof inside the atomic writer
|
||
at the last instant; the lock's tier is chosen by the capability
|
||
predicate — on the enforced tier it is kernel-held (flock / LockFileEx on
|
||
the lock fd) and a kernel refusal that is not contention fails the
|
||
acquisition closed rather than degrading to the name protocol, on the
|
||
name tier the pass refuses to run while appends continue (a typed
|
||
`usage_ledger_compaction_refused` event, once per process per data root)
|
||
— owner-aware,
|
||
and heartbeaten through the long span, with ownership proven immediately
|
||
before each snapshot re-check — including inside the atomic writer, once
|
||
the temp is durable and right before EVERY rename attempt, retries
|
||
included, ownership first (a loss as the commit section is entered aborts
|
||
before the pre-archive look and writes no orphan; a loss at the archive
|
||
aborts before the re-check is even asked; a loss while the temp is
|
||
written, or between two rename attempts, refuses the replace — as does a
|
||
charge that lands between them, and a hold lost the instant the in-swap
|
||
look answered True refuses it too, so the last proof precedes the rename
|
||
by one syscall); a pass that loses the hold abandons its work at its next
|
||
proof instead of swapping (bounded, not absolute: a charge landed by an
|
||
out-of-protocol holder inside that one syscall is erased by the rename —
|
||
on POSIX its bytes are read back from the old inode held open across the
|
||
swap, quarantined with `integrity_degraded` raised, and the pass raises
|
||
typed instead of returning a receipt; on Windows the loss is silent,
|
||
disclosed); of two reclaimers
|
||
racing over a stale lock at most one can evict (eviction only under the
|
||
held flock on the judged fd); a creator evicted while still lock-less
|
||
never returns a descriptor, nor does one whose own identity cannot be read
|
||
— and that one leaves no live-pid stamp behind either; a heartbeat after an
|
||
atomic replacement of the lock file answers False; a writer that cannot take the lock refuses in
|
||
typed form rather than appending without it; and the swap reports a
|
||
receipt only for bytes it re-read at the path.
|
||
10. **Provenance is bounded**: `pre_compaction_seq` names a row inside the
|
||
header's declared source range; the archive directory is the data root's
|
||
own, through no symlink — enforced on POSIX at the open/create itself via
|
||
`O_NOFOLLOW` `dir_fd` handles for both writer and reader, so a link
|
||
planted after any check can neither receive nor serve history; the epoch
|
||
anchor is scanned through the very handle the chain was walked in,
|
||
entries opened relative to it, so a directory swapped in between cannot
|
||
hide a generation; and no archived segment carries a generation newer
|
||
than the live file's own (bar an uncommitted orphan of it, proven by
|
||
still being a prefix of that file), whether or not the live file carries
|
||
a stamp.
|