mirror of
https://github.com/razzant/ouroboros.git
synced 2026-10-03 04:07:04 +00:00
Motivated by benchmark forensics (terminal-bench/SWE-bench/GAIA): most trial deaths were harness infrastructure, not agent inability. All three sub-blocks are general-purpose robustness for normal users. 1a. Per-class same-model transient retry (loop_llm_call.py): - Transient classes (finish_reason=null / empty-response shapes, provider_transient 429/5xx/overloaded) retry the SAME model with a larger budget: transient_retry_max() (OUROBOROS_TRANSIENT_RETRY_MAX, SSOT default 6 in SETTINGS_DEFAULTS + apply_settings_to_env, floored at the caller budget), exponential backoff capped 60s. - Backoff sleeps are deadline-bounded (task_metadata.deadline_at threaded as deadline_ts through the main loop, budget-limit and round-limit wrap-up calls); stopping emits a durable llm_retry_deadline_exhausted event from BOTH transient paths. - Permanent classes (auth/quota/bad_request/request_too_large) fail fast unchanged. NO cross-model fallback is introduced: single-model setups (all slots one model, empty fallback) die only after the real budget, and the failure text reports actual attempts used. 1b. Encrypted-reasoning strip-retry (llm.py): - _is_openrouter_signature_error also matches "encrypted reasoning", "encrypted content for item" (observed gpt-5 shape "...for item rs_..."), "reasoning item", "reasoning_details" - reusing the existing one-shot roundtrip-metadata strip-and-retry on the same model. The allow_fallbacks pin is untouched. 1c. Compaction robustness (context_compaction.py + loop.py): - Per-batch isolation: a failed batch leaves only its own rounds raw; the old whole-pass try/except discarded every successful summary. - Per-round degradation: a missing summary leaves that round raw instead of the all-or-nothing completeness ValueError. - Structured emit_round_summaries tool protocol (tool_choice=required, reliable round_id keying) with text-protocol fallback for local light models or prose answers; spend from failed batches is accounted (_BatchSummaryError carries usage, including across fallback failures). - Warning protection scans the first two non-empty lines (autocorrect notes can prefix the marker); SHELL_EXIT_ERROR rounds are deliberately compactable - trial-and-error history must compact, with the first error line preserved by summarizer instruction. - Emergency compaction adapts keep_recent to min(50, max(6, spans//2), max(1, spans-1)) so oversized transcripts with few huge rounds actually compact instead of no-opping. Review of record: triad+scope rounds 1-4 via run_external_review.py; round 4 blocked=False with zero criticals (scope fable-5 responded, 851,542 real tokens). Remaining advisory (param count on two pre-existing over-limit signatures) is documented pre-existing debt; context-object consolidation is out of block scope. Carriers: VERSION, pyproject.toml, web/package.json, api_types.js GATEWAY_CONTRACT_VERSION, README badge+history (oldest minor row trimmed per P9 cap), ARCHITECTURE.md header + retry/compaction docs.
22 lines
919 B
Python
22 lines
919 B
Python
"""Regression test for PR-B: retry backoff (#15).
|
|
|
|
(#14 — not billing provider-glitch empties — was moved to CONSULT-BUGS.md: doing
|
|
it correctly requires deciding whether the durable usage SSOT in events.jsonl
|
|
should exclude finish_reason=null responses, a provider-billing semantics call
|
|
left to the maintainer.)
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import pathlib
|
|
|
|
REPO = pathlib.Path(__file__).resolve().parents[1]
|
|
|
|
|
|
def test_backoff_doubled_with_cap():
|
|
"""Backoff stays exponential (x4 base) and per-class capped: 30s for
|
|
generic retryable errors, 60s for transient provider classes (v6.28.0)."""
|
|
src = (REPO / "ouroboros" / "loop_llm_call.py").read_text(encoding="utf-8")
|
|
assert "2.0 ** attempt * 4" in src # doubled per-attempt backoff
|
|
assert "min(2 ** attempt * 2, 30)" not in src # old value gone
|
|
assert "_TRANSIENT_BACKOFF_CAP_SEC if is_transient else 30.0" in src
|