Fix CyberGym late delivery and terminal accounting

Recover final measured results below historically held bounds without refunding or repeating paid work. Normalize coherent terminal finality for failure paths, retain custody and append-only evidence, and correct the documented benchmark treatment.
This commit is contained in:
Ouroboros 2026-09-18 02:35:37 +03:00
parent d655de7f46
commit 484f6617e7
6 changed files with 156 additions and 37 deletions

View file

@ -217,10 +217,9 @@ first-turn fallback allowance does not authorize cross-family substitution.
No model price is hardcoded in this adapter. Cost is read from the exact
provider route and usage record. A missing or `null` cost is `cost unknown`,
not zero. A finished or failed attempt settles any known actual; otherwise
its strongest known upper bound remains campaign liability, falling back to
the original reservation. An unknown bound blocks further paid dispatch.
Only an explicit settled/released transition removes that liability.
not zero. Known final cost, measured non-final cost, and a finished attempt
with no cost evidence have separate settlement rules (§10); all preserve
campaign liability. An unknown bound blocks further paid dispatch.
## 5. No-swarm and tool policy
@ -477,24 +476,10 @@ finalization grace (30 min) instead of cancelling it, and a cancellation
custody poll that observes such a frame keeps custody for the same grace.
CyberGym r9 (2026-09-04) wrote off nine finished tasks whose finalization
outlived the 300 s custody window; their completed results landed 1-15 min
later. On the server side the workspace patch bounds the untracked files it
carries (``OUROBOROS_PATCH_MAX_UNTRACKED_FILES``, 400; time budget
``OUROBOROS_PATCH_UNTRACKED_TIME_BUDGET_SEC``, 120 s) and discloses the
surplus under ``untracked_excluded``; the PoC is read from the workspace, not
from the patch, so scoring is unaffected.
Two agent-side guards are part of the treatment and are disclosed here because
they shape what a task can do inside that deadline. A single assistant turn
executes at most ``OUROBOROS_MAX_TOOL_CALLS_PER_TURN`` tool calls (32); the
surplus is pruned from the turn, the model is told what was discarded, and the
task's event log carries a ``tool_call_burst_truncated`` checkpoint (a degenerate
1113-call turn otherwise overflowed the context window and ended the task).
The blocking post-task cognition chain (consolidation, summary, reflection) is
skipped when the task's ``deadline_at`` is nearer than
``OUROBOROS_POST_TASK_COGNITION_MIN_REMAINING_SEC`` (900 s); the PoC is already
final at that point, and the skip is recorded as ``post_task_cognition_skipped``.
Without the guard the supervisor's deadline kill landed mid-reflection, and a task
whose work had finished in time was lost or published cost-non-final.
later. Workspace patch capture discloses its existing per-file exclusions
under ``untracked_excluded``. Eligible binary or larger-than-5-MiB untracked
outputs travel as separate file artifacts instead of Git patch blobs. The
PoC is read from the workspace independently of the patch.
The summary always names the metric, numerator, denominator, task-data hash,
source order, model identity, provider distribution, effort, and whether the
@ -535,6 +520,13 @@ with requested tasks having neither rows nor checkpoints is
`reconcile_incomplete`. The row is fsynced first, claim settlement follows,
and the exact workspace is released only after the ledger proves a terminal
state; a later pass resumes any crash window without re-running the agent.
When a transport failure already settled at its held bound without measured
cost, a later final measured amount at or below that bound may supersede the
failure. The late row retains its measured ``cost_usd`` and separately
discloses the unchanged claim as ``ledger_accounted_usd``. Recovery proves
the bound from the original no-cost row and pre-settlement claim history,
including historical untagged settlements. It refunds nothing and refuses
an amount above the held bound or a conflicting previously measured cost.
Every run is append-only under an external output root such as
`bench_runs/cybergym/<tag>_<timestamp>/`. Large image/binary caches use the
@ -611,10 +603,12 @@ owner-authorized full run the launcher applies the explicit runtime tree cap
runtime tree cap and the latter is the separate campaign-ledger reservation.
Both values are visible without conflating their roles, and paid invocations
must state the runtime cap explicitly. A new claim still requires a finite
per-task estimate. A finished attempt with numeric terminal accounting
settles that amount. Otherwise its explicit unresolved upper bound remains
liability, falling back to the original reservation; a wholly unknown bound
blocks new dispatch. A nullable provider cost is never interpreted as zero.
per-task estimate. A finished attempt with known final, non-estimated cost
settles that amount. Measured but non-final or unattested cost remains
unresolved at its known bound. A finished attempt with no cost evidence
settles at its current held liability, preserving the campaign projection
without claiming that amount as its measured invoice. A wholly unknown
bound blocks new dispatch. A nullable provider cost is never interpreted as zero.
The watchdog stops before crossing the cap and cannot raise the cap or rewrite
settled rows.

View file

@ -670,6 +670,9 @@ def _terminal_gateway_accounting(payload: Mapping[str, Any] | None) -> dict[str,
"non_final_rows", "unknown_unmetered", "reserved_usd", "unresolved_upper_bound_usd",
))):
projected["cost_final"] = False
# Coherent explicit finality supplies the absent estimate flag on every path.
if projected.get("cost_final") is True and not estimated_present:
projected["cost_estimated"] = False
return projected

View file

@ -43,6 +43,9 @@ from typing import Any
from devtools.benchmarks.cybergym.cybergym_adapter import (
_TERMINAL_GATEWAY_STATUSES,
_active_attempt_liability,
_finished_attempt_actual_usd,
_terminal_gateway_accounting,
DEFAULT_LEVEL,
BudgetLedger,
CyberGymError,
@ -636,13 +639,18 @@ def _append_row_if_unrecorded(run_dir: pathlib.Path, row: Mapping[str, Any]) ->
return True
def _settled_attempt_cost_usd(ledger: BudgetLedger, attempt_id: str) -> float:
"""Return the durable settled amount for an attempt, or refuse ambiguity."""
def _settled_attempt_accounting(
ledger: BudgetLedger, attempt_id: str, prior_row: Mapping[str, Any],
) -> tuple[float, bool]:
"""Return settled amount and proof it held a no-evidence attempt's bound."""
events = ledger.events()
latest: Mapping[str, Any] | None = None
for event in ledger.events():
latest_index = -1
for index, event in enumerate(events):
if str(event.get("attempt_id") or "") == str(attempt_id):
latest = event
latest_index = index
kind = (
str(latest.get("event", latest.get("kind", "")) or "").lower()
if latest is not None
@ -656,7 +664,18 @@ def _settled_attempt_cost_usd(ledger: BudgetLedger, attempt_id: str) -> float:
raise LedgerError("settled attempt has no finite cost") from exc
if not math.isfinite(cost):
raise LedgerError("settled attempt has no finite cost")
return cost
# Historical no-evidence settlements have no basis tag. Join the original
# outcome with the exact liability it held, using the settlement owner's
# existing projection; a previously measured amount must still agree.
held_bound = False
if not _terminal_gateway_accounting(prior_row.get("runtime_result")) and all(
prior_row.get(key) is None
for key in ("cost_usd", "cost_upper_bound_usd", "unresolved_upper_bound_usd")
):
held_bound = round(cost, 6) == round(
_active_attempt_liability(events[:latest_index], attempt_id), 6,
)
return cost, held_bound
def _record_reconcile_pass(manifest: dict[str, Any], report: Mapping[str, Any]) -> None:
@ -1086,13 +1105,20 @@ def reconcile_main(args: argparse.Namespace) -> int:
if not math.isfinite(delivered_cost):
raise LedgerError("late terminal row has no finite cost")
if claim_state_at_entry == "settled":
# An already-settled claim is terminal accounting:
# the redelivered row must agree with it exactly.
settled_cost = _settled_attempt_cost_usd(ledger, attempt_id)
if round(settled_cost, 6) != round(delivered_cost, 6):
settled_cost, held_bound = _settled_attempt_accounting(
ledger, attempt_id, recorded_rows[attempt_key],
)
if round(settled_cost, 6) != round(delivered_cost, 6) and not (
held_bound and _finished_attempt_actual_usd(row) is not None
and round(delivered_cost, 6) <= round(settled_cost, 6)
):
raise LedgerError(
"late terminal cost disagrees with settled claim"
)
if held_bound:
# Keep the paid claim unchanged; disclose its
# conservative amount beside the measured cost.
row["ledger_accounted_usd"] = settled_cost
checkpoint_value = json.loads(
checkpoint.read_text(encoding="utf-8")
)

View file

@ -706,9 +706,10 @@ def test_terminal_telemetry_failure_preserves_settled_cost(tmp_path, monkeypatch
assert rows[0]["status"] == "infra_failed"
assert rows[0]["lifecycle"] == "post_gateway_evaluation_failed"
assert rows[0]["cost_usd"] == pytest.approx(0.25)
assert rows[0]["cost_final"] is True and rows[0]["cost_estimated"] is False
projection = BudgetLedger(config.run_root / "claims.jsonl", cap_usd=2).projection()
assert projection.settled_usd == 0
assert projection.unresolved_upper_bound_usd == pytest.approx(0.25)
assert projection.settled_usd == pytest.approx(0.25)
assert projection.unresolved_upper_bound_usd == 0
def test_missing_marker_with_failed_execution_stays_infra(tmp_path, monkeypatch):

View file

@ -7,6 +7,7 @@ Docker daemon, upstream package, or provider credential is used.
from __future__ import annotations
import hashlib
import json
import pathlib
from types import SimpleNamespace
@ -19,9 +20,11 @@ from devtools.benchmarks.cybergym.cybergym_adapter import (
CyberGymError,
append_cybergym_result,
campaign_execution_lock,
run_campaign,
task_slug,
)
from devtools.benchmarks.cybergym.cybergym_reconcile import reconcile_main
from devtools.benchmarks.cybergym.cybergym_wire import GatewayTransportError
from devtools.benchmarks.cybergym.cybergym_result_index import (
append_cybergym_late_result,
effective_task_rows,
@ -437,6 +440,92 @@ def test_reconcile_supersedes_settled_transport_failure_without_resettling(
assert fake.released == [(task_id, attempt_id)]
@pytest.mark.parametrize("prior_cost,actual,final,admitted", [
(None, 7.31, True, True), (None, 20.0, True, True), (None, 0.0, True, True),
(None, 25.0, True, False), (None, None, True, False),
(None, 7.31, False, False), (20.0, 7.31, True, False),
])
def test_campaign_late_result_preserves_historical_held_cost(
prior_cost, actual, final, admitted, tmp_path, monkeypatch,
):
"""A lost paid attempt can deliver later without refunding its held bound."""
task_id = "arvo:1"
root = _write_run(tmp_path / "run", [task_id])
dispatched = []
def original_attempt(task, task_dir):
attempt = task.metadata["attempt_id"]
dispatched.append(attempt)
checkpoint = root / "checkpoints" / task_slug(task_id) / attempt / "gateway_checkpoint.json"
checkpoint.parent.mkdir(parents=True, exist_ok=True)
checkpoint.write_text(json.dumps({
"gateway_task_id": "gateway-" + attempt, "status": "submitted",
}), encoding="utf-8")
if prior_cost is None:
raise GatewayTransportError("bounded gateway poll exhausted")
return {"status": "infra_failed", "lifecycle": "executor_failed",
"cost_usd": prior_cost, "cost_final": True, "cost_estimated": False}
rows = run_campaign([task_id], run_root=root, executor=original_attempt,
estimated_cost_usd=20.0, budget_cap_usd=3000.0)
attempt, = dispatched
assert rows[0]["cost_usd"] == prior_cost
ledger = BudgetLedger(root / "claims.jsonl", cap_usd=3000.0)
assert ledger.attempt_state(attempt) == "settled"
assert ledger.projection().settled_usd == 20.0
ledger_before = (root / "claims.jsonl").read_bytes()
# Historical events carry no settlement-basis tag; recovery must use them.
assert not any("basis" in event for event in ledger.events())
poc = b"late official final PoC"
digest = hashlib.sha256(poc).hexdigest()
outcome = {
"status": "completed", "lifecycle": "official_verified",
"observed_effort": "high", "cost_usd": actual,
"cost_final": final, "cost_estimated": False, "final_poc_sha256": digest,
"trials": [{"trial_id": "final", "is_final": True, "poc_hash": digest,
"vul_exit_code": 1, "fix_exit_code": 0}],
"runtime_result": {"task_id": "gateway-" + attempt, "status": "completed",
"cost_usd": actual, "cost_final": final, "cost_estimated": False},
}
class LateExecutor(_FakeExecutor):
def reconcile_task(self, spec, task_dir, attempt_id, checkpoint):
task_dir.mkdir(parents=True, exist_ok=True)
(task_dir / "final.poc").write_bytes(poc)
return super().reconcile_task(spec, task_dir, attempt_id, checkpoint)
def release_reconciled_workspace(self, spec, attempt_id):
# Both paired rows, including the disclosure, precede cleanup.
assert _read_rows(root / task_slug(task_id)) == _read_rows(root)
assert _read_rows(root)[-1]["ledger_accounted_usd"] == 20.0
return super().release_reconciled_workspace(spec, attempt_id)
fake = LateExecutor(outcome)
_install_fake_executor(monkeypatch, fake)
assert reconcile_main(_reconcile_args(root)) == (0 if admitted else 2)
if admitted:
delivered = _read_rows(root)[-1]
assert delivered["official_success"] is True
assert delivered["row_role"] == "late_delivery"
assert delivered["cost_usd"] == actual
assert delivered["cost_final"] is True
assert delivered["ledger_accounted_usd"] == 20.0
assert fake.released == [(task_id, attempt)]
# A crash-torn task-local pair is repaired without repeating delivery.
task_index = root / task_slug(task_id) / "result_index.jsonl"
task_index.write_text(json.dumps(rows[0]) + "\n", encoding="utf-8")
assert reconcile_main(_reconcile_args(root)) == 0
assert _read_rows(root / task_slug(task_id)) == _read_rows(root)
assert len(_read_rows(root)) == 2
assert fake.reconciled == [(task_id, attempt)]
else:
assert reconcile_main(_reconcile_args(root)) == 2
assert fake.released == []
assert len(_read_rows(root)) == 1
assert (root / "claims.jsonl").read_bytes() == ledger_before
assert dispatched == [attempt]
def test_reconcile_supersedes_unresolved_transport_failure_and_settles_measured(
tmp_path, monkeypatch,
):

View file

@ -163,6 +163,9 @@ def test_explicit_legacy_estimate_is_not_repaired_by_new_proof(value):
frame = _frame()
frame["cost_estimated"] = value
assert _accept(frame) is None
projected = _terminal_gateway_accounting(frame)
assert projected["cost_estimated"] is True
assert projected["cost_final"] is False
@pytest.mark.parametrize("phase", ["pending_once", "running", "", None])
@ -283,7 +286,10 @@ def test_closed_snapshot_preserves_explicit_finality(marker):
projected = _terminal_gateway_accounting(frame)
assert projected["cost_usd"] == pytest.approx(1.05)
assert projected.get("cost_final") is (True if marker is True else False if marker != "absent" else None)
assert "cost_estimated" not in projected
if marker is True:
assert projected["cost_estimated"] is False
else:
assert "cost_estimated" not in projected
def test_closed_snapshot_cannot_override_explicit_partiality():