mirror of
https://github.com/razzant/ouroboros.git
synced 2026-10-03 04:07:04 +00:00
Fix CyberGym late delivery and terminal accounting
Recover final measured results below historically held bounds without refunding or repeating paid work. Normalize coherent terminal finality for failure paths, retain custody and append-only evidence, and correct the documented benchmark treatment.
This commit is contained in:
parent
d655de7f46
commit
484f6617e7
6 changed files with 156 additions and 37 deletions
|
|
@ -217,10 +217,9 @@ first-turn fallback allowance does not authorize cross-family substitution.
|
|||
|
||||
No model price is hardcoded in this adapter. Cost is read from the exact
|
||||
provider route and usage record. A missing or `null` cost is `cost unknown`,
|
||||
not zero. A finished or failed attempt settles any known actual; otherwise
|
||||
its strongest known upper bound remains campaign liability, falling back to
|
||||
the original reservation. An unknown bound blocks further paid dispatch.
|
||||
Only an explicit settled/released transition removes that liability.
|
||||
not zero. Known final cost, measured non-final cost, and a finished attempt
|
||||
with no cost evidence have separate settlement rules (§10); all preserve
|
||||
campaign liability. An unknown bound blocks further paid dispatch.
|
||||
|
||||
## 5. No-swarm and tool policy
|
||||
|
||||
|
|
@ -477,24 +476,10 @@ finalization grace (30 min) instead of cancelling it, and a cancellation
|
|||
custody poll that observes such a frame keeps custody for the same grace.
|
||||
CyberGym r9 (2026-09-04) wrote off nine finished tasks whose finalization
|
||||
outlived the 300 s custody window; their completed results landed 1-15 min
|
||||
later. On the server side the workspace patch bounds the untracked files it
|
||||
carries (``OUROBOROS_PATCH_MAX_UNTRACKED_FILES``, 400; time budget
|
||||
``OUROBOROS_PATCH_UNTRACKED_TIME_BUDGET_SEC``, 120 s) and discloses the
|
||||
surplus under ``untracked_excluded``; the PoC is read from the workspace, not
|
||||
from the patch, so scoring is unaffected.
|
||||
|
||||
Two agent-side guards are part of the treatment and are disclosed here because
|
||||
they shape what a task can do inside that deadline. A single assistant turn
|
||||
executes at most ``OUROBOROS_MAX_TOOL_CALLS_PER_TURN`` tool calls (32); the
|
||||
surplus is pruned from the turn, the model is told what was discarded, and the
|
||||
task's event log carries a ``tool_call_burst_truncated`` checkpoint (a degenerate
|
||||
1113-call turn otherwise overflowed the context window and ended the task).
|
||||
The blocking post-task cognition chain (consolidation, summary, reflection) is
|
||||
skipped when the task's ``deadline_at`` is nearer than
|
||||
``OUROBOROS_POST_TASK_COGNITION_MIN_REMAINING_SEC`` (900 s); the PoC is already
|
||||
final at that point, and the skip is recorded as ``post_task_cognition_skipped``.
|
||||
Without the guard the supervisor's deadline kill landed mid-reflection, and a task
|
||||
whose work had finished in time was lost or published cost-non-final.
|
||||
later. Workspace patch capture discloses its existing per-file exclusions
|
||||
under ``untracked_excluded``. Eligible binary or larger-than-5-MiB untracked
|
||||
outputs travel as separate file artifacts instead of Git patch blobs. The
|
||||
PoC is read from the workspace independently of the patch.
|
||||
|
||||
The summary always names the metric, numerator, denominator, task-data hash,
|
||||
source order, model identity, provider distribution, effort, and whether the
|
||||
|
|
@ -535,6 +520,13 @@ with requested tasks having neither rows nor checkpoints is
|
|||
`reconcile_incomplete`. The row is fsynced first, claim settlement follows,
|
||||
and the exact workspace is released only after the ledger proves a terminal
|
||||
state; a later pass resumes any crash window without re-running the agent.
|
||||
When a transport failure already settled at its held bound without measured
|
||||
cost, a later final measured amount at or below that bound may supersede the
|
||||
failure. The late row retains its measured ``cost_usd`` and separately
|
||||
discloses the unchanged claim as ``ledger_accounted_usd``. Recovery proves
|
||||
the bound from the original no-cost row and pre-settlement claim history,
|
||||
including historical untagged settlements. It refunds nothing and refuses
|
||||
an amount above the held bound or a conflicting previously measured cost.
|
||||
|
||||
Every run is append-only under an external output root such as
|
||||
`bench_runs/cybergym/<tag>_<timestamp>/`. Large image/binary caches use the
|
||||
|
|
@ -611,10 +603,12 @@ owner-authorized full run the launcher applies the explicit runtime tree cap
|
|||
runtime tree cap and the latter is the separate campaign-ledger reservation.
|
||||
Both values are visible without conflating their roles, and paid invocations
|
||||
must state the runtime cap explicitly. A new claim still requires a finite
|
||||
per-task estimate. A finished attempt with numeric terminal accounting
|
||||
settles that amount. Otherwise its explicit unresolved upper bound remains
|
||||
liability, falling back to the original reservation; a wholly unknown bound
|
||||
blocks new dispatch. A nullable provider cost is never interpreted as zero.
|
||||
per-task estimate. A finished attempt with known final, non-estimated cost
|
||||
settles that amount. Measured but non-final or unattested cost remains
|
||||
unresolved at its known bound. A finished attempt with no cost evidence
|
||||
settles at its current held liability, preserving the campaign projection
|
||||
without claiming that amount as its measured invoice. A wholly unknown
|
||||
bound blocks new dispatch. A nullable provider cost is never interpreted as zero.
|
||||
The watchdog stops before crossing the cap and cannot raise the cap or rewrite
|
||||
settled rows.
|
||||
|
||||
|
|
|
|||
|
|
@ -670,6 +670,9 @@ def _terminal_gateway_accounting(payload: Mapping[str, Any] | None) -> dict[str,
|
|||
"non_final_rows", "unknown_unmetered", "reserved_usd", "unresolved_upper_bound_usd",
|
||||
))):
|
||||
projected["cost_final"] = False
|
||||
# Coherent explicit finality supplies the absent estimate flag on every path.
|
||||
if projected.get("cost_final") is True and not estimated_present:
|
||||
projected["cost_estimated"] = False
|
||||
return projected
|
||||
|
||||
|
||||
|
|
|
|||
|
|
@ -43,6 +43,9 @@ from typing import Any
|
|||
|
||||
from devtools.benchmarks.cybergym.cybergym_adapter import (
|
||||
_TERMINAL_GATEWAY_STATUSES,
|
||||
_active_attempt_liability,
|
||||
_finished_attempt_actual_usd,
|
||||
_terminal_gateway_accounting,
|
||||
DEFAULT_LEVEL,
|
||||
BudgetLedger,
|
||||
CyberGymError,
|
||||
|
|
@ -636,13 +639,18 @@ def _append_row_if_unrecorded(run_dir: pathlib.Path, row: Mapping[str, Any]) ->
|
|||
return True
|
||||
|
||||
|
||||
def _settled_attempt_cost_usd(ledger: BudgetLedger, attempt_id: str) -> float:
|
||||
"""Return the durable settled amount for an attempt, or refuse ambiguity."""
|
||||
def _settled_attempt_accounting(
|
||||
ledger: BudgetLedger, attempt_id: str, prior_row: Mapping[str, Any],
|
||||
) -> tuple[float, bool]:
|
||||
"""Return settled amount and proof it held a no-evidence attempt's bound."""
|
||||
|
||||
events = ledger.events()
|
||||
latest: Mapping[str, Any] | None = None
|
||||
for event in ledger.events():
|
||||
latest_index = -1
|
||||
for index, event in enumerate(events):
|
||||
if str(event.get("attempt_id") or "") == str(attempt_id):
|
||||
latest = event
|
||||
latest_index = index
|
||||
kind = (
|
||||
str(latest.get("event", latest.get("kind", "")) or "").lower()
|
||||
if latest is not None
|
||||
|
|
@ -656,7 +664,18 @@ def _settled_attempt_cost_usd(ledger: BudgetLedger, attempt_id: str) -> float:
|
|||
raise LedgerError("settled attempt has no finite cost") from exc
|
||||
if not math.isfinite(cost):
|
||||
raise LedgerError("settled attempt has no finite cost")
|
||||
return cost
|
||||
# Historical no-evidence settlements have no basis tag. Join the original
|
||||
# outcome with the exact liability it held, using the settlement owner's
|
||||
# existing projection; a previously measured amount must still agree.
|
||||
held_bound = False
|
||||
if not _terminal_gateway_accounting(prior_row.get("runtime_result")) and all(
|
||||
prior_row.get(key) is None
|
||||
for key in ("cost_usd", "cost_upper_bound_usd", "unresolved_upper_bound_usd")
|
||||
):
|
||||
held_bound = round(cost, 6) == round(
|
||||
_active_attempt_liability(events[:latest_index], attempt_id), 6,
|
||||
)
|
||||
return cost, held_bound
|
||||
|
||||
|
||||
def _record_reconcile_pass(manifest: dict[str, Any], report: Mapping[str, Any]) -> None:
|
||||
|
|
@ -1086,13 +1105,20 @@ def reconcile_main(args: argparse.Namespace) -> int:
|
|||
if not math.isfinite(delivered_cost):
|
||||
raise LedgerError("late terminal row has no finite cost")
|
||||
if claim_state_at_entry == "settled":
|
||||
# An already-settled claim is terminal accounting:
|
||||
# the redelivered row must agree with it exactly.
|
||||
settled_cost = _settled_attempt_cost_usd(ledger, attempt_id)
|
||||
if round(settled_cost, 6) != round(delivered_cost, 6):
|
||||
settled_cost, held_bound = _settled_attempt_accounting(
|
||||
ledger, attempt_id, recorded_rows[attempt_key],
|
||||
)
|
||||
if round(settled_cost, 6) != round(delivered_cost, 6) and not (
|
||||
held_bound and _finished_attempt_actual_usd(row) is not None
|
||||
and round(delivered_cost, 6) <= round(settled_cost, 6)
|
||||
):
|
||||
raise LedgerError(
|
||||
"late terminal cost disagrees with settled claim"
|
||||
)
|
||||
if held_bound:
|
||||
# Keep the paid claim unchanged; disclose its
|
||||
# conservative amount beside the measured cost.
|
||||
row["ledger_accounted_usd"] = settled_cost
|
||||
checkpoint_value = json.loads(
|
||||
checkpoint.read_text(encoding="utf-8")
|
||||
)
|
||||
|
|
|
|||
|
|
@ -706,9 +706,10 @@ def test_terminal_telemetry_failure_preserves_settled_cost(tmp_path, monkeypatch
|
|||
assert rows[0]["status"] == "infra_failed"
|
||||
assert rows[0]["lifecycle"] == "post_gateway_evaluation_failed"
|
||||
assert rows[0]["cost_usd"] == pytest.approx(0.25)
|
||||
assert rows[0]["cost_final"] is True and rows[0]["cost_estimated"] is False
|
||||
projection = BudgetLedger(config.run_root / "claims.jsonl", cap_usd=2).projection()
|
||||
assert projection.settled_usd == 0
|
||||
assert projection.unresolved_upper_bound_usd == pytest.approx(0.25)
|
||||
assert projection.settled_usd == pytest.approx(0.25)
|
||||
assert projection.unresolved_upper_bound_usd == 0
|
||||
|
||||
|
||||
def test_missing_marker_with_failed_execution_stays_infra(tmp_path, monkeypatch):
|
||||
|
|
|
|||
|
|
@ -7,6 +7,7 @@ Docker daemon, upstream package, or provider credential is used.
|
|||
|
||||
from __future__ import annotations
|
||||
|
||||
import hashlib
|
||||
import json
|
||||
import pathlib
|
||||
from types import SimpleNamespace
|
||||
|
|
@ -19,9 +20,11 @@ from devtools.benchmarks.cybergym.cybergym_adapter import (
|
|||
CyberGymError,
|
||||
append_cybergym_result,
|
||||
campaign_execution_lock,
|
||||
run_campaign,
|
||||
task_slug,
|
||||
)
|
||||
from devtools.benchmarks.cybergym.cybergym_reconcile import reconcile_main
|
||||
from devtools.benchmarks.cybergym.cybergym_wire import GatewayTransportError
|
||||
from devtools.benchmarks.cybergym.cybergym_result_index import (
|
||||
append_cybergym_late_result,
|
||||
effective_task_rows,
|
||||
|
|
@ -437,6 +440,92 @@ def test_reconcile_supersedes_settled_transport_failure_without_resettling(
|
|||
assert fake.released == [(task_id, attempt_id)]
|
||||
|
||||
|
||||
@pytest.mark.parametrize("prior_cost,actual,final,admitted", [
|
||||
(None, 7.31, True, True), (None, 20.0, True, True), (None, 0.0, True, True),
|
||||
(None, 25.0, True, False), (None, None, True, False),
|
||||
(None, 7.31, False, False), (20.0, 7.31, True, False),
|
||||
])
|
||||
def test_campaign_late_result_preserves_historical_held_cost(
|
||||
prior_cost, actual, final, admitted, tmp_path, monkeypatch,
|
||||
):
|
||||
"""A lost paid attempt can deliver later without refunding its held bound."""
|
||||
task_id = "arvo:1"
|
||||
root = _write_run(tmp_path / "run", [task_id])
|
||||
dispatched = []
|
||||
|
||||
def original_attempt(task, task_dir):
|
||||
attempt = task.metadata["attempt_id"]
|
||||
dispatched.append(attempt)
|
||||
checkpoint = root / "checkpoints" / task_slug(task_id) / attempt / "gateway_checkpoint.json"
|
||||
checkpoint.parent.mkdir(parents=True, exist_ok=True)
|
||||
checkpoint.write_text(json.dumps({
|
||||
"gateway_task_id": "gateway-" + attempt, "status": "submitted",
|
||||
}), encoding="utf-8")
|
||||
if prior_cost is None:
|
||||
raise GatewayTransportError("bounded gateway poll exhausted")
|
||||
return {"status": "infra_failed", "lifecycle": "executor_failed",
|
||||
"cost_usd": prior_cost, "cost_final": True, "cost_estimated": False}
|
||||
|
||||
rows = run_campaign([task_id], run_root=root, executor=original_attempt,
|
||||
estimated_cost_usd=20.0, budget_cap_usd=3000.0)
|
||||
attempt, = dispatched
|
||||
assert rows[0]["cost_usd"] == prior_cost
|
||||
ledger = BudgetLedger(root / "claims.jsonl", cap_usd=3000.0)
|
||||
assert ledger.attempt_state(attempt) == "settled"
|
||||
assert ledger.projection().settled_usd == 20.0
|
||||
ledger_before = (root / "claims.jsonl").read_bytes()
|
||||
# Historical events carry no settlement-basis tag; recovery must use them.
|
||||
assert not any("basis" in event for event in ledger.events())
|
||||
poc = b"late official final PoC"
|
||||
digest = hashlib.sha256(poc).hexdigest()
|
||||
outcome = {
|
||||
"status": "completed", "lifecycle": "official_verified",
|
||||
"observed_effort": "high", "cost_usd": actual,
|
||||
"cost_final": final, "cost_estimated": False, "final_poc_sha256": digest,
|
||||
"trials": [{"trial_id": "final", "is_final": True, "poc_hash": digest,
|
||||
"vul_exit_code": 1, "fix_exit_code": 0}],
|
||||
"runtime_result": {"task_id": "gateway-" + attempt, "status": "completed",
|
||||
"cost_usd": actual, "cost_final": final, "cost_estimated": False},
|
||||
}
|
||||
|
||||
class LateExecutor(_FakeExecutor):
|
||||
def reconcile_task(self, spec, task_dir, attempt_id, checkpoint):
|
||||
task_dir.mkdir(parents=True, exist_ok=True)
|
||||
(task_dir / "final.poc").write_bytes(poc)
|
||||
return super().reconcile_task(spec, task_dir, attempt_id, checkpoint)
|
||||
|
||||
def release_reconciled_workspace(self, spec, attempt_id):
|
||||
# Both paired rows, including the disclosure, precede cleanup.
|
||||
assert _read_rows(root / task_slug(task_id)) == _read_rows(root)
|
||||
assert _read_rows(root)[-1]["ledger_accounted_usd"] == 20.0
|
||||
return super().release_reconciled_workspace(spec, attempt_id)
|
||||
|
||||
fake = LateExecutor(outcome)
|
||||
_install_fake_executor(monkeypatch, fake)
|
||||
assert reconcile_main(_reconcile_args(root)) == (0 if admitted else 2)
|
||||
if admitted:
|
||||
delivered = _read_rows(root)[-1]
|
||||
assert delivered["official_success"] is True
|
||||
assert delivered["row_role"] == "late_delivery"
|
||||
assert delivered["cost_usd"] == actual
|
||||
assert delivered["cost_final"] is True
|
||||
assert delivered["ledger_accounted_usd"] == 20.0
|
||||
assert fake.released == [(task_id, attempt)]
|
||||
# A crash-torn task-local pair is repaired without repeating delivery.
|
||||
task_index = root / task_slug(task_id) / "result_index.jsonl"
|
||||
task_index.write_text(json.dumps(rows[0]) + "\n", encoding="utf-8")
|
||||
assert reconcile_main(_reconcile_args(root)) == 0
|
||||
assert _read_rows(root / task_slug(task_id)) == _read_rows(root)
|
||||
assert len(_read_rows(root)) == 2
|
||||
assert fake.reconciled == [(task_id, attempt)]
|
||||
else:
|
||||
assert reconcile_main(_reconcile_args(root)) == 2
|
||||
assert fake.released == []
|
||||
assert len(_read_rows(root)) == 1
|
||||
assert (root / "claims.jsonl").read_bytes() == ledger_before
|
||||
assert dispatched == [attempt]
|
||||
|
||||
|
||||
def test_reconcile_supersedes_unresolved_transport_failure_and_settles_measured(
|
||||
tmp_path, monkeypatch,
|
||||
):
|
||||
|
|
|
|||
|
|
@ -163,6 +163,9 @@ def test_explicit_legacy_estimate_is_not_repaired_by_new_proof(value):
|
|||
frame = _frame()
|
||||
frame["cost_estimated"] = value
|
||||
assert _accept(frame) is None
|
||||
projected = _terminal_gateway_accounting(frame)
|
||||
assert projected["cost_estimated"] is True
|
||||
assert projected["cost_final"] is False
|
||||
|
||||
|
||||
@pytest.mark.parametrize("phase", ["pending_once", "running", "", None])
|
||||
|
|
@ -283,7 +286,10 @@ def test_closed_snapshot_preserves_explicit_finality(marker):
|
|||
projected = _terminal_gateway_accounting(frame)
|
||||
assert projected["cost_usd"] == pytest.approx(1.05)
|
||||
assert projected.get("cost_final") is (True if marker is True else False if marker != "absent" else None)
|
||||
assert "cost_estimated" not in projected
|
||||
if marker is True:
|
||||
assert projected["cost_estimated"] is False
|
||||
else:
|
||||
assert "cost_estimated" not in projected
|
||||
|
||||
|
||||
def test_closed_snapshot_cannot_override_explicit_partiality():
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue