mirror of
https://github.com/razzant/ouroboros.git
synced 2026-08-04 16:19:50 +00:00
Commit 3/3 of the bench post-mortem sprint (plan smooth-skipping-charm; commits 1-2 = bf8bde7 v6.54.3, ef786bb v6.54.4). Minor bump per plan. Shared bench-template scaffold defaults, disclosed in the benchmarks index and per-bench METHODOLOGY files: OUROBOROS_MAX_WORKERS=4 (same-model decomposition slots within one task, never best-of-N; TB 2->4, GAIA --max-workers default 1->4 with explicit 1 as the strict-baseline ablation and the quality-profile silent 1->5 bump removed, SWE-pro 5->4), OUROBOROS_SAFETY_MODE=light (disposable jails; deterministic guards stay), RUNTIME_MODE=pro for container benches (GAIA deliberately stays light: its solver runs unisolated against a live repo), and claude_code_edit disabled in every bench solve task (single-model harness measurement). TB raises _DEADLINE_SAFETY_SEC 30->105 from measured finalization overhead, mirrors the registry _WEB_TOOLS denylist (youtube_transcript drift fixed, pinned by a sync test), and the README pins why --all-model keeps review single-model. ProgramBench e2e runner ported from the colleague's tree, adapted to our APIs: metadata.budget_profile onto the v6.54.4 contract (footer keys dropped, reserve pct unit-fixed 0.15->15), solve-model preflight through migrate_model_value (fail-fast on direct-route legacy ids), atomic per-instance checkpoints with honored reattach (client poll timeout leaves the executor alive; the next run skips seed/start and reattaches), explicit payload-status terminal detection, denominator-preserving ledgers (skipped_existing rows; skipped = successful for exit code), source-only submissions (both root binaries excluded by name), idempotent workspace normalization shared by the prepare-only flow. New continual_learning/ launcher wraps the EXTERNAL CL-Bench runner (nothing vendored; fails loudly with obtain instructions), strict-sequential guard (parallel opt-in is exclusively-stateless-bridge only), standard-path pointer rows keep the ledger denominator, fidelity gaps of the pinned external adapter recorded and warned. OSWorld aligned to the official 2.0 protocol (pinned OSWorld-V2@c261cb57, 500-step default, show_result.py layout, official traj rows), bridge-level final_answer capture with VM-state-only prompting, preflight verifies the target server's effective scaffold settings via /api/settings (env cannot configure a --url server) and refuses the live desktop URL without an explicit flag. LifelongAgentBench documented as blocked (its 100% run was a no-gold-oracle artifact). Carriers synced to 6.55.0. Review: 6 triad+scope rounds to a fully clean verdict (16 confirmed criticals fixed, each with a regression test); gemini/opus clean in all rounds. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| benchmarks | ||
| __init__.py | ||
| README.md | ||
Ouroboros Devtools
devtools/ contains operator-side and benchmark support code that should be
versioned with Ouroboros without becoming part of the runtime core.
Rules:
- Generated logs, datasets, run outputs, Docker layers, and secrets do not live here.
- Default benchmark outputs go under
/Users/anton/Ouroboros/bench_runs/. - Runtime modules must not import
devtools. - This is not an immune-system bypass: touched files are reviewed normally.
- Promote code out of
devtoolsonly through a separate reviewed runtime plan.