ouroboros/devtools/benchmarks/continual_learning
Ouroboros 62f87cc94c Merge upstream ouroboros db6d7cf8 into the v7 line: absorb PR #609 net-resilience and PR #614 update letter
Second absorption of the frozen upstream line (23ab428f..db6d7cf8: 89
commits, 47 files) on top of the rc.10 hotfix tip, by the F2 rules (S1
upstream body in the owning leaf, S2 hand-merge, S3 only with proof;
retired 7.0 surfaces never return):

- PR #609 net-resilience: interactive transport-wait episodes bounded by
  the task idle timeout with the typed task_incident/toast_once pair, a
  bounded paid repeat after a typed post-dispatch transport death with a
  round-keyed record that fences every other send, the shutdown-aware
  supervisor crash counter and bounded lifespan join, Darwin keepalive
  tuning. loop_transport/transport_custody/net_transport/loop_llm_call
  land verbatim (same shapes on both sides); the loop.py deltas are
  relocated into the v7 leaves (loop_round_limits, loop_model_call,
  loop_delivery, loop_forced_finalization, loop_nudges, loop_messages,
  loop_budget) with bodies AST-equal to upstream modulo the call-time
  handles; _emit_overflow_retry_skipped stays a public helper (the v7
  facade contract) and upstream's nested _skipped delegates to it.
- PR #614 update letter: ouroboros/update_letter.py and its web module
  land verbatim; the new OUROBOROS_UPDATE_LETTER_TIMEOUT_SEC key and
  get_update_letter_timeout_sec live in their v7 owners
  (settings_defaults.py, runtime_limits.py, re-exported by config.py);
  _supervisor_stop lives in server_process.py beside the restart events;
  docs/PERSISTENCE.md gains the state/update_letter.json row and the
  inventory pin moves to 286.
- Tests: the relocated run_llm_loop tests take the emit_progress
  incident keyword (every one-argument progress fake in tests/ swept, a
  gap upstream itself left in test_tree_cost_ceiling); the official-update
  runtime-section test lands in tests/test_context.py; _MOVED_OWNERS
  registers the relocated getter.
- Docs: ARCHITECTURE/DEVELOPMENT hunks land on the upstream text; the two
  legacy timeout rows upstream's context still carries stay retired (7.0).
- Size ratchet regenerated; the band rationale for tests/test_update_letter.py
  is carried verbatim from upstream.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-04 22:53:15 +00:00
..
operator_patches Fix benchmark actor provenance 2026-08-20 02:32:32 +03:00
__init__.py feat: v6.55.0 bench devtools alignment — scaffold defaults, ProgramBench e2e, CLB launcher, OSWorld 2.0 2026-07-03 22:49:01 +03:00
METHODOLOGY.md Migrate benchmark subagent profiles 2026-08-20 02:32:32 +03:00
README.md feat(v6.81.0): benchmark provenance becomes a gate, and the artefacts stop lying 2026-07-26 03:40:44 +00:00
run_clb.py Migrate benchmark subagent profiles 2026-08-20 02:32:32 +03:00
RUNBOOK.md Benchmark disclosures: repeats live inside the call's attempt budget 2026-09-04 14:17:57 +00:00
settings_base.json v7next F3.1-D ABI-10: remove the reviewer comma-list migration read (owner 5.4=A) 2026-08-31 16:54:17 +00:00

CL-Bench (continual-learning-bench) launcher

Launcher wrapper for running Ouroboros on CL-Bench (continual-learning-bench.com, arXiv 2606.05661) — a benchmark that feeds a system a strictly sequential stream of task instances and measures continual learning: a memoryless stateless baseline vs a stateful rollout where one persistent system carries state across the whole ordered sequence. For Ouroboros the carried state is its native memory (scratchpad / knowledge), which is the point of running this bench: runs go against a live agent with memory ON, evolution OFF by default.

See METHODOLOGY.md for what the benchmark measures, official scoring, our scaffold disclosures, the known failure taxonomy, and honest limits. See RUNBOOK.md for the field-tested at-scale operating recipe (VM memory sizing, mandatory companion daemons, one-seed vs leaderboard-submission flow, known loss classes).

External runner REQUIRED (not vendored)

This directory contains only the launcher. The benchmark itself — tasks, scoring, schedules, and the in-runner Ouroboros adapter — lives in the external continual-learning-bench repository and must be obtained separately:

  • The runner repo provides run_benchmark.py, src/, schedules/, and the official analysis scripts (scripts/analyze_final_results.py, scripts/generate_leaderboard.py).
  • The Ouroboros adapter is the src/systems/ouroboros/ subtree of that repo (system.py, _launcher.py, _docker_launcher.py, run_clbench_bridge_agent.py, clbench_step_shim.py). The v6.71.1 reference campaign pinned it at commit 3ec3761 (adds a network-outage hold that pauses the scope clock instead of forfeiting questions, a format-repair round that re-emits a prose-final answer as typed JSON in the same container, an OUROBOROS_REVIEW_MAX_PASSES override, and extra-overrides passthrough); a full copy ships in the run handoff bundle (clbench-671-full-2026-07-21.tar.gz, under bench-config/external-adapters/ouroboros/) together with the adapter's own METHODOLOGY. If the adapter is missing from your checkout, restore it from that bundle.
  • The adapter needs a dedicated Ouroboros clone (--ouroboros-clone, never the LIVE repo — $OUROBOROS_REPO_DIR; a pinned seed IS allowed to be the checkout you launched from) whose devtools/benchmarks/common/server_runner.py it imports, plus (docker path) the clbench-ouroboros:dev image and the runner-side clbench_skill/remote_work skill source (renamed from clbench_remote 2026-07-09 — bench-identifying tells stripped from the agent-visible surface; same tools get_observation/submit_action).
  • Run under a Python that has the runner's deps (litellm/pydantic/...) AND the Ouroboros deps; the launcher defaults to <runner>/.venv/bin/python when it exists.

The launcher fails loudly when any of these are absent; --dry-run works without them.

v6.56.0 operator patches for the pinned adapter

See operator_patches/README.md — three host/runtime incompatibilities of the pinned external adapter with v6.56.0 (safety-lowering owner-guard, host-loopback shim unreachability from Linux containers, rootful-daemon uid mismatch) and the unified diffs that port it.

Strictly sequential — parallelism is forbidden

Two separate knobs, deliberately kept apart:

  1. Cross-task order is the benchmark. The stateful rollout MUST process instances strictly one-by-one (--instance-workers 1, the default; the launcher refuses more unless --allow-parallel-baseline is passed, and even then fan-out applies only to the per-instance-independent stateless baseline arm — each parallel worker boots its own container, so watch disk: the runner's default schedules carry max_workers: 12 and that has filled a disk before).
  2. max_workers (settings template) is the agent's internal worker pool — subagent decomposition WITHIN one task. It does not (and must not) create cross-task parallelism; it is disclosed as a scaffold parameter in the manifest. The validated at-scale value is 3 (not the earlier template's 4/10): pool size drives engine-container RSS, and larger pools OOM the Docker VM when several containers run concurrently — see RUNBOOK.md for the sizing formula and the failure signature.

Launch

cd repo
# Dry run: manifest + rendered settings + exact planned runner argv, no spend.
python devtools/benchmarks/continual_learning/run_clb.py \
  --runner-path ~/continual-learning-bench \
  --ouroboros-clone ~/ouroboros-bench-src \
  --dry-run

# Standard "clean path" (per-action, runner owns the loop) — DB domain, 1-seed smoke:
python devtools/benchmarks/continual_learning/run_clb.py \
  --runner-path ~/continual-learning-bench \
  --ouroboros-clone ~/ouroboros-bench-src \
  --path standard --domain database_exploration --runs 1

# Bridge path (whole-question; the shape of the reference full40 2026-07-01 run):
python devtools/benchmarks/continual_learning/run_clb.py \
  --runner-path ~/continual-learning-bench \
  --ouroboros-clone ~/ouroboros-bench-src \
  --path bridge --domain database_exploration \
  --phases stateless,stateful_noevo --num-instances 40

# Re-normalize an existing run dir:
python devtools/benchmarks/continual_learning/run_clb.py --collect-only bench_runs/continual_learning/<run>

Provider keys are taken from the environment first (OPENROUTER_API_KEY, ...), falling back to the live data/settings.json; they travel via the child environment only and are never written into run artifacts (the rendered _run_settings.json has all secret-shaped values blanked).

Run layout (append-only, one fresh dir per launch)

bench_runs/continual_learning/continual_learning_<stamp>_<pid>/
  run_manifest.json     # provenance + template-fidelity report + disclosures
  _run_settings.json    # rendered settings template, secrets blanked
  traces/               # bridge path: <condition>/<domain>/q###/{prompt.txt,task_outcome.json,...}
  runner_state/         # adapter sidecar ledgers + isolated engine run roots
  results.json          # normalized per-condition means + memory/evolution effects
  result_index.jsonl    # denominator-preserving per-instance ledger

Official scoring authority stays with the external runner (normalized_reward_mean via its analysis scripts); results.json is an audit sidecar of raw per-instance rewards.

Template fidelity (read the warnings)

The pinned external adapter enforces model slots, effort, worker pool, evolution flag, budget, and OpenRouter provider routing through its own interface; the launcher feeds those from settings_base.json. Three template knobs are declared but not forwarded by the pinned adapter — OUROBOROS_SAFETY_MODE=light and OUROBOROS_REVIEW_ENFORCEMENT=blocking (docker path drops both) and CLBENCH_SOLVE_DISABLED_TOOLS (incl. claude_code_edit) — the launcher exports them, records the gap in the manifest, and prints loud warnings. See METHODOLOGY.md "Template fidelity" for the two-line adapter patch that closes this.