ouroboros/devtools
Ouroboros 8e847e7982 fix(e2e_live): spend from the product ledger; skill fixture declares its runtime; landing skip flags recorded and checked
Three stand corrections from the deep read of rc.15 run3.

Money: lane_spend summed the llm_usage telemetry rows, which the product
itself calls compatibility telemetry (pricing.py); the money authority is
the settled physical-attempt ledger state/usage_attempts.jsonl. run3 read
114.81 by telemetry and 141.63 by the ledger — the skill-review triad, the
advisory pre-review through the agent SDK and the post-task synthesis
write no telemetry row — so the run cap admitted on a short count. The
reader now sums settled, cost-final ledger rows (cost_usd), counting an
unpriced one as unknown; the docstring's claim that the telemetry carried
every row is gone.

Skill fixture: the stand's reference SKILL.md declares type extension with
no runtime, which CREATING_SKILLS.md says is required for extensions; the
model copied the fixture byte for byte in all three lanes and the terra
reviewer failed manifest_schema (a hard-critical item) in two of them. The
fixture now declares runtime: python3.

Landing path: DEVELOPMENT.md describes the SM1 landing as commit_reviewed
with no skip flags and a real advisory row, but the stand only checked
that some advisory run was real. run2 SM1_a1 and run3 SM1_a1 landed
through review_rebuttal + skip_advisory_review=True and still passed.
commit_refusal_facts now records the landing call's truthy skip_* args as
landing_skip_flags and run_sm1 checks landed_without_skip_flags, so a
bypass landing is a typed failed check with the facts beside it.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
2026-09-06 11:08:02 +00:00
..
benchmarks fix(e2e_live): reserve the evolution root only for the scenario that absorbs; typed idle reasons relative to the wait 2026-09-06 01:31:11 +00:00
e2e_live fix(e2e_live): spend from the product ledger; skill fixture declares its runtime; landing skip flags recorded and checked 2026-09-06 11:08:02 +00:00
__init__.py feat(devtools-benchmarks): add official benchmark harnesses and workspace executor 2026-06-06 12:03:30 +03:00
measure_review_pack.py devtools: measure the FULL scope input with the real assembler (P4) 2026-09-05 09:52:35 +00:00
README.md feat(devtools-benchmarks): add official benchmark harnesses and workspace executor 2026-06-06 12:03:30 +03:00

Ouroboros Devtools

devtools/ contains operator-side and benchmark support code that should be versioned with Ouroboros without becoming part of the runtime core.

Rules:

  • Generated logs, datasets, run outputs, Docker layers, and secrets do not live here.
  • Default benchmark outputs go under /Users/anton/Ouroboros/bench_runs/.
  • Runtime modules must not import devtools.
  • This is not an immune-system bypass: touched files are reviewed normally.
  • Promote code out of devtools only through a separate reviewed runtime plan.