Merge pull request #611 from razzant/claude/model-defaults-2026-09

Shipped defaults: Main and the first triad reviewer on gemini-3.8-flash,
deep self-review on plain sol, the Antigravity preset on gemini-3.8-flash.

Co-authored-by: Ouroboros <311266734+ouroboros-agent@users.noreply.github.com>
This commit is contained in:
Anton Razzhigaev 2026-09-04 09:13:23 +00:00
commit 23ab428f51
16 changed files with 61 additions and 50 deletions

View file

@ -568,7 +568,7 @@ Every host renders one served page: `GET /onboarding` returns `onboarding_templa
Completion is one HTTP conversation on every host — `POST /api/onboarding/complete` (`gateway/onboarding.py`; its docstring carries the numbered step order) — so providers and runtime mode cannot be left half-saved. Its single write, re-proved under the settings lock, persists settings, next-boot runtime mode, the fresh-install `OUROBOROS_SAFETY_MODE=light` default no other endpoint may author, the one-shot preset marker, and the durable completion fact; only then does the supervisor start. Install time means three proofs together (`gateway/onboarding.py::preset_eligible`): onboarding never completed here, no preset generation applied, and no `settings.json` yet — three, because "no startup-ready provider" alone is a state an old install reaches whenever its key stops working. `GET /api/onboarding` normalizes what it displays but never persists: a read that created `settings.json` would silently disqualify the install-time latches. A completion failing AFTER bytes reach disk reports that it saved plus the failing stage; a 2xx is not a completion — only the exact success envelope is, because the saved runtime mode and restart receipt live in that body. On a boot-pinned change the framed overlay shows its restart card and a plain browser tab shows the saved-but-restart-required screen (the launcher recycle: above). The overlay sandbox grants popup permission because the agent sign-in link is the step's primary action.
`ouroboros/subscription_install_presets.py` is a pure compiler with two sibling products — task actors and reviewer slots, not a reviewer/subagent matrix; its docstring carries the emission rules, the ratified Claude/Codex/Cursor reviewer subset, and the Agy-neutrality contract. Normalized API/local settings alone suffice for API-only and local-only installs; a declared subscription makes `gateway/onboarding.py` read exactly one live Claudexor snapshot before the settings transaction — the authority for supported harnesses, exact model ids (Agy's automatic row: `gemini-3.7-flash-high`), and enabled accounts. Durable seat facts (kind/enabled/verified), not the hour's quota reading, decide the once-only preset: a spent Claude window at save time must not yield a Codex-only preset permanently. Task-actor rows stay unpinned and grow linearly, never as a powerset; on the fresh-install path reviewer slots are `subagent_id` references into the roster the preset ships (a seat matching no task actor mints a `review-<harness>` roster row), while an owner-configured roster stays validate-only. Missing required exact discovery is a typed pre-write refusal; no partial preset persists.
`ouroboros/subscription_install_presets.py` is a pure compiler with two sibling products — task actors and reviewer slots, not a reviewer/subagent matrix; its docstring carries the emission rules, the ratified Claude/Codex/Cursor reviewer subset, and the Agy-neutrality contract. Normalized API/local settings alone suffice for API-only and local-only installs; a declared subscription makes `gateway/onboarding.py` read exactly one live Claudexor snapshot before the settings transaction — the authority for supported harnesses, exact model ids (Agy's automatic row: `gemini-3.8-flash-high`), and enabled accounts. Durable seat facts (kind/enabled/verified), not the hour's quota reading, decide the once-only preset: a spent Claude window at save time must not yield a Codex-only preset permanently. Task-actor rows stay unpinned and grow linearly, never as a powerset; on the fresh-install path reviewer slots are `subagent_id` references into the roster the preset ships (a seat matching no task actor mints a `review-<harness>` roster row), while an owner-configured roster stays validate-only. Missing required exact discovery is a typed pre-write refusal; no partial preset persists.
`POST /api/onboarding/subagents/preview` runs the same compiler against the open draft without persisting; the editable result becomes `OUROBOROS_SUBAGENTS`, and an owner-edited draft is validated and preserved, not regenerated. Completion persists actors, reviewer disposition, preset receipt, and completion facts through the one-write document-lock/fingerprint boundary; network discovery happens before the lock. A daemon failure keeps the wizard open with an explicit finish-without-agent-defaults path; a read or failed refresh never rewrites saved intent.
@ -1216,7 +1216,7 @@ A task with `type=deep_self_review` bypasses the ordinary tool loop and calls `d
- **Native inspection episode** (a configured-subagent `api_chat` row): `NativeToolRoundReviewExecutor` over the repository root with `policy["native_data_root"]` = the real runtime root, so the reviewer reads the repository with the host's read-only tools (the runtime root is its readable data plane) while the memory whitelist reaches it inline byte-exact. The route-owned task carries the role prompt, up to seven memory files inline byte-exact with every whitelisted entry's disposition stated (`inlined` / `missing` / `empty` / `oversized` / `read_error` — memory coverage is by inlining and disclosure, never by receipts), BIBLE.md as a MANDATORY full read (with its size, for chunked reads) and ARCHITECTURE/DEVELOPMENT/CHECKLISTS as navigation maps read on demand; the output contract is the report shape (free markdown, most critical first, a one-line model-side coverage header). Bounds are the shared ones: the window-derived transcript bound with the landing notice, the owner deadline together with the slot's logical window (the task's absolute ceiling narrowed by the deadline), and the paid ledger; exhaustion delivers the collected draft marked `native_incomplete`. After the episode the host derives BIBLE.md coverage from the executed repository-root `read_file` receipts, matched on the reader's `opened_path` and `opened_root` (never on the model's spelling; a data-plane read never counts): `read` only when the merged extents of measured receipts cover the whole file (a single result is capped, so a full read is multi-chunk by construction), otherwise `partial(fraction)`, `missing` (nothing of the file delivered at a repository root) or `unobserved` (receipts capped below the call count, or a matching receipt without an extent — full coverage is proven by measured receipts alone); the edge rules (the raw-spelling fallback for an extent-less receipt, `..` shapes, zero-line deliveries) are `_native_read_coverage`'s docstring. Anything but `read` is disclosed in the header and as a typed `capability_delta`, never as a refusal (R8).
- **Delegated session** (an `agent_session` row): `AgentSessionReviewExecutor` with the same task and contract (no output schema — a report is prose); the session reads the repository itself, the memory files reach it inline, its reads are not host-observed and coverage is recorded `unobserved`.
Every delivered report is prefixed by the host provenance header — `<!-- deep-review provenance: delivery=…, model=…, memory=n/7[, memory_missing=…, memory_empty=…, memory_oversized=…, memory_read_error=…], coverage=…, incomplete=…, attestation=…[, rounds/tool_calls/receipts/end_reason/transcript/landing on the native episode only] -->` plus one human-readable line, every comment value and every external value on the human line sanitized and bounded through `_header_value` (one value per memory disposition, each at most the seven whitelisted basenames), the fact set built per delivery (a session carries no round or receipt facts and its attestation is `unobserved` by construction) — so a reader, and the next task's context (which quotes `memory/deep_review.md`), can tell a packed report from a retrieved one and a complete one from an incomplete one (R9). `incomplete` is derived from the facts each delivery actually carries: the native episode's typed `native_incomplete`; the packed call's provider stop marker — the OpenAI-compatible normalizer's `response_finish_reason == "length"` or, on the direct-Anthropic lane (the shipped `anthropic::` deep default, which sets no usage finish reason), the message's `stop_reason == "max_tokens"` — as `output_reserve` (the report hit the 100K output reserve); a session's completeness is `unobserved`. The packed pack's memory dispositions ride its OMITTED section and `state/deep_self_review_context.json`'s sibling usage fact `deep_review_memory`. The packed path records normal usage evidence and writes the coverage manifest to `state/deep_self_review_context.json`. The retrieving deliveries record the row's «Выполняется как» execution on every outcome (responded, empty response, executor exception) and persist their prompt and response through `persist_call`; the packed delivery persists its request and response through `chat_observed` and records its execution only once the call has returned (responded, or an empty response as an error row) — a transport exception on the packed call is the typed `deep_self_review_error` failure without an execution record. `agent.py` stores the report at `memory/deep_review.md` ONLY when the review delivered: every failure returns typed usage (`execution_status=infra_failed` with `deep_self_review_unavailable` / `deep_self_review_error`), lands in the task result and a typed `task_error` event, and leaves the previous report in place. Availability is route-aware (`deep_review_route`): the packed row keeps the ≥1M floor — a route whose Capability Evidence (`reviewer_window.resolve_reviewer_window`, the shared resolver) confirms a sub-1M window is refused typed, because the pack IS this delivery's guarantee and is never silently shrunk to a smaller window (a native or session `deep_review` row serves such a route); an UNKNOWN window keeps the documented full-window assumption every packet surface shares and is disclosed in the header (`window=assumed_1000000`); the `openai::` route is trusted only without `OPENAI_BASE_URL`, and the OpenRouter `-pro` slug is rewritten to the direct default — a native row needs its routed model's credentials, a session row a healthy delegated route (`subagents.route_health`); the `request_deep_self_review` tool and the agent read the ROW, not the model key. On every delivery the model has no mutating tools and does not run plan, task-acceptance, or commit reviewers; its report is durable diagnostic memory under BIBLE authority, not implementation or publication authority.
Every delivered report is prefixed by the host provenance header — `<!-- deep-review provenance: delivery=…, model=…, memory=n/7[, memory_missing=…, memory_empty=…, memory_oversized=…, memory_read_error=…], coverage=…, incomplete=…, attestation=…[, rounds/tool_calls/receipts/end_reason/transcript/landing on the native episode only] -->` plus one human-readable line, every comment value and every external value on the human line sanitized and bounded through `_header_value` (one value per memory disposition, each at most the seven whitelisted basenames), the fact set built per delivery (a session carries no round or receipt facts and its attestation is `unobserved` by construction) — so a reader, and the next task's context (which quotes `memory/deep_review.md`), can tell a packed report from a retrieved one and a complete one from an incomplete one (R9). `incomplete` is derived from the facts each delivery actually carries: the native episode's typed `native_incomplete`; the packed call's provider stop marker — the OpenAI-compatible normalizer's `response_finish_reason == "length"` or, on the direct-Anthropic lane (the shipped `anthropic::` deep default, which sets no usage finish reason), the message's `stop_reason == "max_tokens"` — as `output_reserve` (the report hit the 100K output reserve); a session's completeness is `unobserved`. The packed pack's memory dispositions ride its OMITTED section and `state/deep_self_review_context.json`'s sibling usage fact `deep_review_memory`. The packed path records normal usage evidence and writes the coverage manifest to `state/deep_self_review_context.json`. The retrieving deliveries record the row's «Выполняется как» execution on every outcome (responded, empty response, executor exception) and persist their prompt and response through `persist_call`; the packed delivery persists its request and response through `chat_observed` and records its execution only once the call has returned (responded, or an empty response as an error row) — a transport exception on the packed call is the typed `deep_self_review_error` failure without an execution record. `agent.py` stores the report at `memory/deep_review.md` ONLY when the review delivered: every failure returns typed usage (`execution_status=infra_failed` with `deep_self_review_unavailable` / `deep_self_review_error`), lands in the task result and a typed `task_error` event, and leaves the previous report in place. Availability is route-aware (`deep_review_route`): the packed row keeps the ≥1M floor — a route whose Capability Evidence (`reviewer_window.resolve_reviewer_window`, the shared resolver) confirms a sub-1M window is refused typed, because the pack IS this delivery's guarantee and is never silently shrunk to a smaller window (a native or session `deep_review` row serves such a route); an UNKNOWN window keeps the documented full-window assumption every packet surface shares and is disclosed in the header (`window=assumed_1000000`); the `openai::` route is trusted only without `OPENAI_BASE_URL`, and an owner-pinned OpenRouter `-pro` slug is rewritten to the direct default, plain Sol being the shipped default on both routes — a native row needs its routed model's credentials, a session row a healthy delegated route (`subagents.route_health`); the `request_deep_self_review` tool and the agent read the ROW, not the model key. On every delivery the model has no mutating tools and does not run plan, task-acceptance, or commit reviewers; its report is durable diagnostic memory under BIBLE authority, not implementation or publication authority.
Rationale. The review ran without tools for as long as it did because its guarantee was the PACK: one ≥1M-context call whose Atlas assembles fail-closed — a required artifact that does not fit refuses the whole review instead of shrinking it — so the reviewer provably held BIBLE.md, the protected runtime and the whole memory whitelist at once, and no smaller-window model could impersonate that coverage. That guarantee is also the packed delivery's limit: it exists only on a ≥1M route (sub-1M installs had no deep review at all), the reviewer cannot follow a call chain out of the files the greedy budget admitted, and "the pack contained X" was silently read as "the reviewer considered X". Retrieving deliveries are admitted now because they buy two things the pack never could — a deep review on any install with a payable route or a subscription session, and host-OBSERVED reads on the native episode (the receipts say which files were actually opened) — while the guarantee they lack stays honest instead of hidden: the mandatory BIBLE.md read is checked after the fact and a miss is disclosed, never assumed; a delegated session's reads are `unobserved` and say so; the packed delivery remains the default and the only one whose coverage is assembled rather than retrieved. Whole-repository retrieval in one episode is not a goal: the four canonical docs alone (≈1.06M chars) exceed any transcript bound, so the retrieving task is a targeted survey through the navigation maps, not a promise of full coverage.
@ -1335,7 +1335,7 @@ A registry of `config.SETTINGS_DEFAULTS` (exact defaults stay canonical in `conf
| OUROBOROS_MANAGED_UPDATE_FETCH_TIMEOUT_SEC | 300 | Managed-update fetch ceiling |
| OUROBOROS_RESCUE_GIT_TIMEOUT_SEC | 300 | Per-process ceiling on rescue Git commands |
| OUROBOROS_TRUST_NONLOCAL_BIND_WITHOUT_PASSWORD | unset | Env-only: `1` permits saving a non-loopback bind without a password |
| OUROBOROS_MODEL | google/gemini-3.7-flash | Main model |
| OUROBOROS_MODEL | google/gemini-3.8-flash | Main model |
| OUROBOROS_MODEL_HEAVY | "" | Legacy slot: readable for migration/history only, out of active routing |
| OUROBOROS_MODEL_LIGHT | openai/gpt-5.6-luna | Light model |
| OUROBOROS_MODEL_VISION | "" | Vision model (empty inherits) |
@ -1351,7 +1351,7 @@ A registry of `config.SETTINGS_DEFAULTS` (exact defaults stay canonical in `conf
| OUROBOROS_FALLBACK_COOLDOWN_SEC | 120 | Cooldown window |
| OUROBOROS_FALLBACK_ATTEMPTS_PER_MODEL | 1 | Attempts per model in the fallback walk |
| OUROBOROS_REVIEW_NATIVE_MAX_TRANSCRIPT_CHARS | 900000 | Owner CEILING (chars) on the native review inspection episode transcript; the effective bound is the reviewer window's calibrated capacity, never above this. No round cap exists (`OUROBOROS_REVIEW_NATIVE_MAX_ROUNDS` is retired): exhaustion is a typed fail-closed refusal for verdict shapes and a disclosed incomplete product for the report shape, never a silent truncation |
| OUROBOROS_MODEL_DEEP_SELF_REVIEW | openai/gpt-5.6-sol-pro | Deep self-review model key — the invisible migration source and fallback for the optional `deep_review` reviewer row: with no row saved, `deep_review_slot()` synthesizes the packed api row from it (the historical delivery); a saved row wins and the key is not read — so the provider-default migrations of this key (`server_runtime.py`) reach only installs that still synthesize from it; a row-configured install keeps its row, by design. Not a Settings UI field any more — the row lives in Agents → Review lanes |
| OUROBOROS_MODEL_DEEP_SELF_REVIEW | openai/gpt-5.6-sol | Deep self-review model key — the invisible migration source and fallback for the optional `deep_review` reviewer row: with no row saved, `deep_review_slot()` synthesizes the packed api row from it (the historical delivery); a saved row wins and the key is not read — so the provider-default migrations of this key (`server_runtime.py`) reach only installs that still synthesize from it; a row-configured install keeps its row, by design. Not a Settings UI field any more — the row lives in Agents → Review lanes |
| OUROBOROS_MAX_WORKERS | 10 | Worker-pool size |
| OUROBOROS_MAX_ACTIVE_SUBAGENTS_PER_ROOT | 6 | Live-subagent cap per root (hard cap 500 ids; depth hard cap 10) |
| OUROBOROS_MAX_SUBAGENT_DEPTH | 3 | Subagent tree depth |
@ -1386,7 +1386,7 @@ A registry of `config.SETTINGS_DEFAULTS` (exact defaults stay canonical in `conf
| OUROBOROS_OBSERVABILITY_KEEP_RAW | unset | Env-only: truthy enables raw observability payload persistence |
| OUROBOROS_GENERATIVE_PROBE | 1 (on) | Generative-write probe toggle |
| OUROBOROS_GENERATIVE_PROBE_CHARS | 5000000 | Generative-probe size companion |
| OUROBOROS_REVIEW_MODELS | google/gemini-3.7-flash,openai/gpt-5.6-terra,anthropic/claude-opus-5 | Legacy triad reviewer roster shared by commit/plan/task/skill review when `OUROBOROS_REVIEWER_SLOTS` is absent (duplicate model IDs are independent slots); with the structured setting present, every review surface (commit/plan/skill review and task acceptance) reads the exact per-row delivery and this key is only a runtime projection of the panel's api model ids for legacy consumers (the external review script's key ordering, benchmark manifests) — never a second write and never a review input |
| OUROBOROS_REVIEW_MODELS | google/gemini-3.8-flash,openai/gpt-5.6-terra,anthropic/claude-opus-5 | Legacy triad reviewer roster shared by commit/plan/task/skill review when `OUROBOROS_REVIEWER_SLOTS` is absent (duplicate model IDs are independent slots); with the structured setting present, every review surface (commit/plan/skill review and task acceptance) reads the exact per-row delivery and this key is only a runtime projection of the panel's api model ids for legacy consumers (the external review script's key ordering, benchmark manifests) — never a second write and never a review input |
| OUROBOROS_REVIEWER_SLOTS | (empty) | (6.1) Structured reviewer-slot SSOT (`reviewer_slot_config.py`): JSON `{triad[], scope[], advisory, deep_review?}`; each row is EITHER an inline route `{slot_id, route:{kind: api_chat\|agent_session, target_id}, effort}` OR a roster reference `{slot_id, subagent_id, effort}` — mutually exclusive (a row naming both refuses typed; the reference materializes route/effort from the Available-subagents roster at load time, an explicit row `effort` winning over the roster row's) — with a STABLE owner-assigned slot_id (never an array index); an `agent_session` route may add the optional `route.profile_id` credential pin (empty = account rotation; a pin on an `api_chat` route refuses typed); the optional `deep_review` singleton carries the same row keys minus `slot_id` (fixed `deep_review_slot_1`) and, absent, is synthesized as the packed api row from `OUROBOROS_MODEL_DEEP_SELF_REVIEW`. Empty = read the legacy comma keys + phase-5 route envs as the migration source. Malformed value refuses typed at save AND at review time on every surface, task acceptance included (owner R3); env-apply logs and leaves legacy keys unprojected. The save that FIRST gives the triad a retrieving row (agent session or configured-subagent native inspection) returns the one-time R12 migration disclosure in the save response's `warnings` — the rows by id and target, and the measured API packet-panel cost it replaces (≈12 s / ≈$0.07 per model row per task, median of the 2026-09-01 OSWorld traces; ≈75 s / ≈$0.82 for a three-row panel on ProgramBench) against minutes of subscription window per task for a session row; a later save that keeps a retrieving triad is silent, and so — by design — is the reverse transition back to a packet-only triad: R12 is a one-time migration notice, not a routing monitor. The onboarding ladder footnote states the same numbers. |
| OUROBOROS_SUBSCRIPTION_PRESET_VERSION | (empty) | One-shot install-preset marker; endpoint-authored, DISK-ONLY (`ENDPOINT_AUTHORED_SETTINGS`) — its absence authorizes nothing, which is why install time is proved by three facts (§2) |
| OUROBOROS_SUBAGENT_PRESET_RECEIPT | (empty) | Install-preset receipt; endpoint-authored, disk-only |

View file

@ -332,18 +332,22 @@ def provider_credential_plan(
# provider profiles gives onboarding, runtime defaults, and tests one vocabulary
# instead of repeating model ids across those surfaces.
OPENROUTER_DEFAULTS = {
"main": "google/gemini-3.7-flash",
"main": "google/gemini-3.8-flash",
"heavy": "",
"light": "openai/gpt-5.6-luna",
"vision": "",
"consciousness": "",
"fallback": "openai/gpt-5.6-luna",
"deep_self_review": "openai/gpt-5.6-sol-pro",
# Plain Sol, not the `-pro` routing slug: a pro-mode call bills the prompt
# three to four times over (parallel test-time compute), so the density
# witness it records can never admit a repository-sized pack, and every
# packed review costs that multiple. Same id the direct-OpenAI slot ships.
"deep_self_review": "openai/gpt-5.6-sol",
}
OPENROUTER_REVIEW_DEFAULTS = {
"triad": (
"google/gemini-3.7-flash",
"google/gemini-3.8-flash",
"openai/gpt-5.6-terra",
"anthropic/claude-opus-5",
),
@ -366,15 +370,13 @@ OPENAI_DIRECT_DEFAULTS = {
# Cloud.ru and GigaChat are documented BELOW that floor, so filling their slot
# would advertise a deep review that is doomed to overflow its real route.
#
# DELIBERATELY plain Sol, NOT the OpenRouter default's `-pro`: that suffix is an
# OpenRouter slug, not an OpenAI model id. Live-probed 2026-07-29 against
# api.openai.com: `gpt-5.6-sol-pro` on /v1/chat/completions -> 404; the pro
# Plain Sol, the same model the OpenRouter default names. A `-pro` suffix is an
# OpenRouter routing slug, not an OpenAI model id: live-probed 2026-07-29 against
# api.openai.com, `gpt-5.6-sol-pro` on /v1/chat/completions -> 404; the pro
# reasoning mode exists only on /v1/responses as `reasoning.mode="pro"` (200),
# and passing `reasoning` to /v1/chat/completions -> 400 "Unknown parameter".
# Every LLM call in llm.py is a chat.completions call, so a direct-OpenAI
# install runs deep review on plain Sol — an owner-accepted capability
# difference from the OpenRouter default, disclosed in README/ARCHITECTURE
# rather than papered over with a slug that does not exist.
# Every LLM call in llm.py is a chat.completions call, so an owner's pinned
# `-pro` slug lands here too (deep_self_review.deep_review_route).
"deep_self_review": "openai::gpt-5.6-sol",
}

View file

@ -54,9 +54,9 @@ _DIRECT_PROVIDER_LEGACY_DEFAULTS = {
"OUROBOROS_MODEL_FALLBACKS": {
"anthropic/claude-sonnet-4.6", "openai/gpt-5.4-mini", "openai::gpt-5.4-mini",
},
# The SHIPPED OpenRouter deep default and its migrated direct spelling both
# name a router slug that does not exist on api.openai.com (404) — a direct
# install must land on the real model, not on the -pro id.
# Prior shipped OpenRouter deep defaults and their migrated direct spellings
# all name router slugs that do not exist on api.openai.com (404) — a direct
# install must land on the real model, not on a -pro id.
"OUROBOROS_MODEL_DEEP_SELF_REVIEW": {
"openai/gpt-5.6-sol-pro", "openai::gpt-5.6-sol-pro",
"openai/gpt-5.5-pro", "openai::gpt-5.5-pro",
@ -91,7 +91,7 @@ _LEGACY_GEMINI_3_FLASH_PREVIEW = "google/gemini-" + "3-flash-preview"
for _legacy_defaults in _DIRECT_PROVIDER_LEGACY_DEFAULTS.values():
for _slot in ("OUROBOROS_MODEL", "OUROBOROS_MODEL_HEAVY", "OUROBOROS_MODEL_LIGHT"):
_legacy_defaults[_slot].add(_LEGACY_GEMINI_31_FLASH_LITE)
# Outgoing SHIPPED OpenRouter defaults (through v6.104), applied for EVERY
# Outgoing SHIPPED OpenRouter defaults, applied for EVERY
# exclusive-direct provider (incl. cloudru/gigachat/minimax/deepseek, which have no per-provider
# legacy table): before each defaults refresh a stored copy of the shipped default matched the
# `current in {"", default}` check because SETTINGS_DEFAULTS still carried it;
@ -101,6 +101,7 @@ for _legacy_defaults in _DIRECT_PROVIDER_LEGACY_DEFAULTS.values():
_PRIOR_SHIPPED_SLOT_DEFAULTS = {
"OUROBOROS_MODEL": {
"google/gemini-3.5-flash",
"google/gemini-3.7-flash",
"x-ai/grok-4.5",
},
"OUROBOROS_MODEL_HEAVY": {"google/gemini-3.5-flash"},
@ -109,9 +110,13 @@ _PRIOR_SHIPPED_SLOT_DEFAULTS = {
"google/gemini-3.6-flash",
},
"OUROBOROS_MODEL_FALLBACKS": {"anthropic/claude-sonnet-4.6"},
# v6.81's shipped deep-review value: an upgraded direct-provider install still
# carries it, and it is just as unreachable without an OpenRouter credential.
"OUROBOROS_MODEL_DEEP_SELF_REVIEW": {"openai/gpt-5.5-pro", "openai::gpt-5.5-pro"},
# Prior shipped deep-review values (v6.81's gpt-5.5-pro, then the gpt-5.6-sol-pro
# routing slug): an upgraded direct-provider install still carries one, and each
# is just as unreachable without an OpenRouter credential.
"OUROBOROS_MODEL_DEEP_SELF_REVIEW": {
"openai/gpt-5.5-pro", "openai::gpt-5.5-pro",
"openai/gpt-5.6-sol-pro", "openai::gpt-5.6-sol-pro",
},
}
# Heavy is no longer an active role, but its bounded migration reader still
# needs to distinguish an owner's custom value from values Ouroboros itself

View file

@ -221,7 +221,7 @@ _SUBSCRIPTION_FIELDS = _rows(("id", "payloadKey", "label", "note"), (
("skip-subscription-presets", SKIP_SUBSCRIPTION_PRESETS_FIELD, "Finish without agent defaults", "Completes onboarding without moving reviewers and subagents onto the connected subscriptions. Everything stays editable in Settings afterwards."),
))
_MODEL_SUGGESTIONS = list(dict.fromkeys(("google/gemini-3.7-flash", "x-ai/grok-4.6", "openai/gpt-5.6-terra", "openai/gpt-5.6-sol", "openai/gpt-5.6-luna", "openai::gpt-5.6-terra", "openai::gpt-5.6-sol", "openai::gpt-5.6-luna", "anthropic/claude-sonnet-5", "anthropic/claude-opus-5", "anthropic::claude-sonnet-5", "anthropic::claude-opus-5", "anthropic::claude-opus-4-6", "deepseek/deepseek-v4-pro", "deepseek::deepseek-v4-pro", "deepseek::deepseek-v4-flash", "openai-compatible::meta-llama/compatible", "cloudru::zai-org/GLM-4.7", "minimax::MiniMax-M3", "minimax::MiniMax-M2.7")))
_MODEL_SUGGESTIONS = list(dict.fromkeys(("google/gemini-3.8-flash", "x-ai/grok-4.6", "openai/gpt-5.6-terra", "openai/gpt-5.6-sol", "openai/gpt-5.6-luna", "openai::gpt-5.6-terra", "openai::gpt-5.6-sol", "openai::gpt-5.6-luna", "anthropic/claude-sonnet-5", "anthropic/claude-opus-5", "anthropic::claude-sonnet-5", "anthropic::claude-opus-5", "anthropic::claude-opus-4-6", "deepseek/deepseek-v4-pro", "deepseek::deepseek-v4-pro", "deepseek::deepseek-v4-flash", "openai-compatible::meta-llama/compatible", "cloudru::zai-org/GLM-4.7", "minimax::MiniMax-M3", "minimax::MiniMax-M2.7")))
def _string(value: Any) -> str:

View file

@ -86,7 +86,7 @@ _MODEL_ALIASES: Dict[str, Dict[str, Tuple[str, ...]]] = {
# agy (Antigravity) spells effort inside the id like cursor. Flash High is
# the automatic task actor; Pro remains an ordinary manual editor choice.
HARNESS_AGY: {
"gemini-3.7-flash": ("gemini-3.7-flash-{effort}",),
"gemini-3.8-flash": ("gemini-3.8-flash-{effort}",),
"gemini-3.1-pro": ("gemini-3.1-pro-{effort}",),
},
}
@ -218,7 +218,7 @@ _TASK_POLICIES: Dict[str, SurfacePolicy] = {
HARNESS_CLAUDE: _surface("opus-5", "medium"),
HARNESS_CODEX: _surface("gpt-5.6-sol", "medium"),
HARNESS_CURSOR: _surface("grok-4.6", "high"),
HARNESS_AGY: _surface("gemini-3.7-flash", "high"),
HARNESS_AGY: _surface("gemini-3.8-flash", "high"),
}
_POLICY_HARNESSES = (HARNESS_CLAUDE, HARNESS_CODEX, HARNESS_CURSOR)

View file

@ -44,7 +44,7 @@ class ProviderCanary:
_OPENROUTER_CANARIES = (
ProviderCanary(
"openrouter_gemini", "google/gemini-3.7-flash", "openrouter",
"openrouter_gemini", "google/gemini-3.8-flash", "openrouter",
"OPENROUTER_API_KEY", True, "medium",
),
ProviderCanary(

View file

@ -784,7 +784,7 @@ class TestOmissionSectionBound:
def test_direct_openai_deep_review_sends_a_real_openai_model_id():
"""PHYSICAL-PAYLOAD proof, not a defaults-table assertion.
The OpenRouter default is the slug `openai/gpt-5.6-sol-pro`. That `-pro`
An owner may pin the slug `openai/gpt-5.6-sol-pro`. That `-pro`
suffix is an OpenRouter routing slug, NOT an OpenAI model id: live-probed
2026-07-29, `gpt-5.6-sol-pro` on api.openai.com /v1/chat/completions returns
404, while pro reasoning exists only on /v1/responses as

View file

@ -208,9 +208,9 @@ def test_antigravity_install_succeeds_without_inventing_reviewer_seats(onboardin
"harnesses": [{
"id": "agy", "status": "ok", "enabled": True,
"models": [
{"id": "gemini-3.7-flash-low"},
{"id": "gemini-3.7-flash-medium"},
{"id": "gemini-3.7-flash-high"},
{"id": "gemini-3.8-flash-low"},
{"id": "gemini-3.8-flash-medium"},
{"id": "gemini-3.8-flash-high"},
],
}],
"profiles": {
@ -229,7 +229,7 @@ def test_antigravity_install_succeeds_without_inventing_reviewer_seats(onboardin
items = json.loads(saved[SUBAGENTS_SETTING])["items"]
assert items[0]["route"] == {
"kind": "agent_session",
"target_id": "agy=gemini-3.7-flash-high",
"target_id": "agy=gemini-3.8-flash-high",
"credential_profile_id": "",
}
assert saved["OUROBOROS_REVIEWER_SLOTS"] == ""
@ -1377,7 +1377,7 @@ def test_a_connected_agy_account_composes_with_core_reviewers(onboarding):
"models": [{"id": "claude-opus-5"}, {"id": "claude-sonnet-5"},
{"id": "claude-fable-5"}, {"id": "claude-opus-4-6"}]},
{"id": "agy", "status": "unavailable", "enabled": True,
"models": [{"id": "gemini-3.7-flash-high"}, {"id": "gemini-3.1-pro-high"}]},
"models": [{"id": "gemini-3.8-flash-high"}, {"id": "gemini-3.1-pro-high"}]},
],
"profiles": {
# next_up kind="none": the engine returns none whenever the
@ -1405,7 +1405,7 @@ def test_a_connected_agy_account_composes_with_core_reviewers(onboarding):
row["route"]["target_id"]
for row in json.loads(saved[SUBAGENTS_SETTING])["items"]
]
assert "agy=gemini-3.7-flash-high" in actor_targets
assert "agy=gemini-3.8-flash-high" in actor_targets
reviewer = json.loads(saved["OUROBOROS_REVIEWER_SLOTS"])
roster = {
row["subagent_id"]: row["route"]["target_id"]

View file

@ -171,7 +171,7 @@ def test_provider_alarm_output_sanitizes_token_shaped_evidence(capsys):
def test_exact_provider_canary_matrix_logical_turns_and_attempt_bound():
matrix = provider_canary_matrix()
assert [(row.canary_id, row.model) for row in matrix] == [
("openrouter_gemini", "google/gemini-3.7-flash"),
("openrouter_gemini", "google/gemini-3.8-flash"),
("openrouter_opus", "anthropic/claude-opus-5"),
("openrouter_gpt", "openai/gpt-5.6-luna"),
("openrouter_grok", "x-ai/grok-4.6"),

View file

@ -352,7 +352,7 @@ def test_explicit_settings_mapping_is_authoritative(monkeypatch, model, settings
"CLAUDE_CODE_MODEL": "claude-name-only",
}, "anthropic::light"),
("cloudru", {}, "cloudru::zai-org/GLM-4.7"),
("openrouter", {"CLAUDE_AGENT_SDK_MODEL": "opus"}, "google/gemini-3.7-flash"),
("openrouter", {"CLAUDE_AGENT_SDK_MODEL": "opus"}, "google/gemini-3.8-flash"),
("openai-compatible", {
"OUROBOROS_MODEL_FALLBACKS": "openai::other,openai-compatible::chosen",
}, "openai-compatible::chosen"),
@ -381,7 +381,7 @@ def test_catalog_is_used_only_for_compatible_discovery(monkeypatch):
"openrouter", {"OPENROUTER_API_KEY": "key"},
) == {"ok": True}
assert discoveries == []
assert probes == ["google/gemini-3.7-flash"]
assert probes == ["google/gemini-3.8-flash"]
assert provider_api._run_provider_test_with_settings("openai-compatible", {
"OPENAI_COMPATIBLE_BASE_URL": "https://compat.example/v1",
}) == {"ok": True}

View file

@ -40,7 +40,7 @@ def test_settings_defaults_include_phase2_keys():
assert SETTINGS_DEFAULTS["OUROBOROS_RUNTIME_MODE"] == "advanced"
assert SETTINGS_DEFAULTS["OUROBOROS_SKILLS_REPO_PATH"] == ""
assert SETTINGS_DEFAULTS["OUROBOROS_MODEL"] == "google/gemini-3.7-flash"
assert SETTINGS_DEFAULTS["OUROBOROS_MODEL"] == "google/gemini-3.8-flash"
# Empty role slots inherit Main; Light and Fallback use Luna explicitly.
assert SETTINGS_DEFAULTS["OUROBOROS_MODEL_HEAVY"] == ""
assert SETTINGS_DEFAULTS["OUROBOROS_MODEL_VISION"] == ""
@ -49,7 +49,7 @@ def test_settings_defaults_include_phase2_keys():
assert SETTINGS_DEFAULTS["OUROBOROS_MODEL_FALLBACKS"] == "openai/gpt-5.6-luna"
assert (
SETTINGS_DEFAULTS["OUROBOROS_MODEL_DEEP_SELF_REVIEW"]
== "openai/gpt-5.6-sol-pro"
== "openai/gpt-5.6-sol"
)
assert SETTINGS_DEFAULTS["TOTAL_BUDGET"] == 200.0
assert SETTINGS_DEFAULTS["OUROBOROS_PER_TASK_COST_USD"] == 50.0

View file

@ -424,11 +424,11 @@ def test_apply_runtime_provider_defaults_keeps_new_triad_on_openrouter():
assert not changed
assert changed_keys == []
assert normalized["OUROBOROS_MODEL"] == "google/gemini-3.7-flash"
assert normalized["OUROBOROS_MODEL"] == "google/gemini-3.8-flash"
assert normalized["OUROBOROS_MODEL_LIGHT"] == "openai/gpt-5.6-luna"
assert normalized["OUROBOROS_MODEL_FALLBACKS"] == "openai/gpt-5.6-luna"
assert normalized["OUROBOROS_REVIEW_MODELS"] == (
"google/gemini-3.7-flash,openai/gpt-5.6-terra,anthropic/claude-opus-5"
"google/gemini-3.8-flash,openai/gpt-5.6-terra,anthropic/claude-opus-5"
)
assert normalized["OUROBOROS_SCOPE_REVIEW_MODEL"] == "openai/gpt-5.6-terra"
assert normalized["OUROBOROS_SCOPE_REVIEW_MODELS"] == "openai/gpt-5.6-terra"

View file

@ -87,7 +87,7 @@ def test_review_models_default_in_config():
assert val # non-empty
models = [m.strip() for m in val.split(",") if m.strip()]
assert models == [
"google/gemini-3.7-flash",
"google/gemini-3.8-flash",
"openai/gpt-5.6-terra",
"anthropic/claude-opus-5",
]

View file

@ -58,9 +58,9 @@ LIVE_MODELS = {
"claude-fable-5-thinking-xhigh", "claude-sonnet-5-medium",
),
"agy": (
"gemini-3.7-flash-low",
"gemini-3.7-flash-medium",
"gemini-3.7-flash-high",
"gemini-3.8-flash-low",
"gemini-3.8-flash-medium",
"gemini-3.8-flash-high",
),
}
@ -246,7 +246,7 @@ def test_antigravity_compiles_task_actor_without_changing_core_reviewer_bytes(co
row["route"]["target_id"]
for row in json.loads(preset.available_subagents)["items"]
]
assert "agy=gemini-3.7-flash-high" in actor_routes
assert "agy=gemini-3.8-flash-high" in actor_routes
core = tuple(harness for harness in connected if harness in CORE_HARNESSES)
if not core:
assert preset.reviewer_slots == ""
@ -351,9 +351,13 @@ def test_compiler_reads_no_settings_and_carries_no_transport(monkeypatch):
# Verbatim from the Antigravity CLI the Claudexor 3.5.0 agy adapter pins
# (AGY_KNOWN_MODELS, verified against agy 1.1.13). Fourteen ids; effort rides
# inside the slug, and gemini-3.1-pro exists ONLY at high/low.
# (AGY_KNOWN_MODELS, verified against agy 1.1.13, plus the gemini-3.8-flash
# triple the shipped preset now targets — assumed to be published by the vendor
# CLI on the owner's decision, not read from an installed agy on this host).
# Seventeen ids; effort rides inside the slug, and gemini-3.1-pro exists ONLY
# at high/low.
AGY_LIVE_MODELS = (
"gemini-3.8-flash-high", "gemini-3.8-flash-medium", "gemini-3.8-flash-low",
"gemini-3.7-flash-high", "gemini-3.7-flash-medium", "gemini-3.7-flash-low",
"gemini-3.6-flash-high", "gemini-3.6-flash-medium", "gemini-3.6-flash-low",
"gemini-3.5-flash-high", "gemini-3.5-flash-medium", "gemini-3.5-flash-low",
@ -397,7 +401,7 @@ def test_agy_alias_table_spells_effort_inside_the_id():
aliases = _MODEL_ALIASES[HARNESS_AGY]
# Every alias candidate formats to an id the pinned vendor CLI really
# publishes, so automatic and manually selected rows resolve exact ids.
assert aliases["gemini-3.7-flash"][0].format(effort="high") in AGY_LIVE_MODELS
assert aliases["gemini-3.8-flash"][0].format(effort="high") in AGY_LIVE_MODELS
assert aliases["gemini-3.1-pro"][0].format(effort="high") in AGY_LIVE_MODELS
assert aliases["gemini-3.1-pro"][0].format(effort="low") in AGY_LIVE_MODELS
# Documented trap for the future dictation: pro has no -medium slug.

View file

@ -318,7 +318,7 @@ function collectSecretValue(id, body) {
// Fallback picker pills mirror config defaults plus useful direct-provider ids.
const SETTINGS_FALLBACK_MODELS = [
'google/gemini-3.7-flash',
'google/gemini-3.8-flash',
'x-ai/grok-4.6',
'openai/gpt-5.6-terra',
'openai/gpt-5.6-sol',

View file

@ -20,7 +20,7 @@ const SETTINGS_TABS = [
// Guard markers: renderTabStrip emits behavior/advanced tabs at runtime.
const MODEL_CARDS = [
['Main', 'Primary reasoning model.', 's-model', 's-local-main', 'google/gemini-3.7-flash'],
['Main', 'Primary reasoning model.', 's-model', 's-local-main', 'google/gemini-3.8-flash'],
['Light', 'Fast summaries, lightweight internal work, reflections, and the default Fast scout. Empty uses Main.', 's-model-light', 's-local-light', 'openai/gpt-5.6-luna'],
['Vision', 'Caption and VLM lane. Empty uses Main.', 's-model-vision', '', ''],
['Consciousness', 'High-horizon background consciousness. Empty uses Main.', 's-model-consciousness', 's-local-consciousness', ''],