ouroboros/devtools/benchmarks/common
Ouroboros b463bb3d93 Add the Z.ai (GLM) direct provider with effort projection at the send boundary
zai:: joins the direct providers exactly the way deepseek:: did: prefix and
credential registry, ZAI_API_KEY plus a ZAI_PLAN endpoint selector
(provider_models.resolve_zai_base_url: empty/payg = api.z.ai/api/paas/v4,
coding = the Coding Plan endpoint), the routing target, live catalog fetch,
provider Test, settings card, onboarding contract, review-fallback roles,
single-provider startup and review detection, secret masking, benchmark
env hygiene, and docs.

Reasoning effort now reaches Z.ai. The provider serves an ABSENT
reasoning_effort at its maximum tier, so every call on the old generic
compatible route was billed at max regardless of the configured effort.
The canonical scale is projected onto Z.ai's own low/high/max enum
(ZAI_REASONING_EFFORT_ALIASES: none/minimal -> low, medium -> high,
xhigh/ultra -> max), disclosed as reasoning_effort_clamped when the tier
changes; GLM-5.3 rejects every other value and cannot disable thinking
(HTTP 400 code 1210), and forced tool_choice works with thinking on, so
there is no DeepSeek-style suppression arm. The projection is keyed on the
provider id the owner configured, never on a model name: a GLM served
from an owner's own OpenAI-compatible endpoint keeps today's behavior.

The provider port is the contributor's own work from the closed PR #1194,
narrowed to Z.ai (the DashScope and Moonshot lanes were not measured and
stay out). 07-configuration gains two settings rows and one route
paragraph (budget 37300 -> 38400), 02-naming records the dated Z.ai
probe in the external-fact inventory, and the onboarding bootstrap
fixture and data-layout inventory are regenerated.

Co-authored-by: josephsteuerjr <josephsteuerjr@gmail.com>
2026-09-25 17:41:54 +03:00
..
__init__.py feat(devtools-benchmarks): add official benchmark harnesses and workspace executor 2026-06-06 12:03:30 +03:00
launcher_audit.py Merge frozen release repairs into development 2026-09-18 13:51:38 +03:00
manifests.py Merge current ouroboros into CyberGym recovery 2026-09-18 00:41:54 +03:00
model_slots.py feat: integrate subscription accounts into model setup and execution 2026-09-07 10:54:51 +00:00
official_commands.py feat: v6.55.0 bench devtools alignment — scaffold defaults, ProgramBench e2e, CLB launcher, OSWorld 2.0 2026-07-03 22:49:01 +03:00
result_index.py fix(benchmarks): retain official evaluator evidence independently of execution 2026-09-25 03:21:08 +03:00
run_roots.py fix(bench): preserve Windows campaign paths and exclusive execution 2026-09-18 12:09:21 +03:00
secrets.py Add the Z.ai (GLM) direct provider with effort projection at the send boundary 2026-09-25 17:41:54 +03:00
server_runner.py Add the Z.ai (GLM) direct provider with effort projection at the send boundary 2026-09-25 17:41:54 +03:00
subprocesses.py feat(devtools-benchmarks): add official benchmark harnesses and workspace executor 2026-06-06 12:03:30 +03:00