Commit graph

17 commits

Author SHA1 Message Date
Peter Steinberger
7cbaa6b36f
refactor(scripts): deslop shared tooling flows (#162824)
Consolidate script lifecycle and helper ownership while preserving CLI, generated-output, cleanup, and guard contracts. Reject inherited object keys in model-matrix option admission with the existing unknown-argument diagnostic.

Validation: 90 full sibling suites; original-fails parser regression; byte-identical native outputs; changed checks; SDK surface and API comparison; both zero-cycle checks; independent isolated review.
2026-10-01 21:30:19 +00:00
Peter Steinberger
2291fe823e
refactor(scripts): deslop tooling scripts second pass (#161169)
Share repeated tooling parsing, projections, and fixture transforms while preserving command and generated-output contracts. Repair the OpenGrep help range so bootstrap code no longer replaces documented usage.
2026-09-29 12:05:18 +00:00
Peter Steinberger
567d7d393c
refactor(scripts): deslop root scripts a–k (#159199)
Consolidate repeated CLI parsing, scanning, benchmark and release projections through their existing owners. Preserve script and CI contracts.

Fix malformed installer version diagnostics, flat ClawHub artifact sealing, completed polling sleep listener retention, and metadata output symlink confinement.
2026-09-27 01:08:19 -07:00
Peter Steinberger
dfb2cee6d8
chore: compare Code Mode against disabled on complex workloads (#155613)
* test: add paired Code Mode workload evaluation

* test: preserve initial evidence assertion during build rechecks

* test: align Code Mode evidence checks with paired admission

* ci: refresh merge proof after upstream pairing repair

Refresh the merge ref after main fixed the observed SSH-pairing fixture race in f87f23c483. No source changes.
2026-09-22 04:24:26 -07:00
Peter Steinberger
09e06495a3
feat: run Code Mode on Node or isolated QuickJS (#154522)
* feat: add selectable Code Mode executors

Default enabled Code Mode to trusted Node execution and move QuickJS into a bundled executor plugin. Preserve typed discovery, JavaScript-only execution, tool authorization, continuation ownership, and explicit legacy QuickJS selections. Add a web settings selector and document both security boundaries.

* refactor: finish Code Mode executor source cutover

Remove the retired core worker copies and regenerate config documentation for the requested QuickJS plugin. The 21 added plugin paths are standard plugin management fields; core and channel counts stay unchanged.

* refactor: narrow the Code Mode plugin contract

Keep only the executor and guest protocol exports consumed by the QuickJS plugin, budget that generic contract, and load the public plugin artifact through the existing runtime test boundary.

* refactor: align Code Mode workers with current runtime boundaries

Use the worker-side task server, keep asynchronous cleanup ownership explicit, and model the real Promise contracts in lifecycle fixtures. Regenerate the requested plugin config surface and budget the exact 35 public executor exports.

* refactor: keep executor implementation types private

* chore: regenerate code mode config baseline

* fix: satisfy Code Mode executor integration contracts

* test: cover QuickJS plugin metadata and exact settings titles

* test: retain QuickJS integration in the agent runtime suite

* test: keep Code Mode validation within lint and type-shard contracts
2026-09-21 17:33:04 -07:00
Peter Steinberger
1cdc5bd4f9
fix: avoid reusing builds from reverted source edits (#149866)
* fix: retain producer provenance when reusing build artifacts

* test: record clean producer facts in artifact fixtures

* docs: clarify provenance requirements for frozen runtimes
2026-09-16 02:45:49 -07:00
Vincent Koc
9ccba21473
feat(qa): retain evidence across repeated scenarios and retries (#147756)
* feat(qa): retain invocation evidence and explicit proof ownership

* test(qa): require evidence in maturity attestation guard

* refactor(qa): document evidence assertion invariants

* fix(qa): preserve typed scenario grouping and nested evidence

* test(qa): use canonical direct conversation fixture

* refactor(qa): keep internal evidence definitions private

* refactor(qa): split evidence validation and command planning

* test(qa): remove unused planning mock import

* refactor(qa): preserve matrix evidence with explicit copying

* fix(qa): preserve invocation ownership across retained child evidence

* fix(qa): preserve immutable publication and artifact bases

* fix(qa): preserve continuation artifacts and selected outcomes

* fix(qa): retry observed outcomes while retaining selected evidence

* fix(qa): frame source evidence and preserve artifact test boundaries

* fix(qa): separate diagnostics from scenario proof

* fix(qa): evaluate rowless retry proof by selection

* fix(qa): retain unresolved maturity evidence

* fix(qa): preserve child coverage and repeated Docker lanes

* refactor(qa): keep evidence schemas independent of consumers

* test(qa): isolate matrix interruption from the test runner

* test(qa): pass the signal name to the interruption listener

* test(plugins): align Copilot auth-choice registry expectations (#148981)

(cherry picked from commit 2b8c9f1ad6)
2026-09-15 21:30:41 +08:00
Vincent Koc
bcfe3383b8
fix(qa): avoid runtime startup for evidence scripts (#148153)
Some checks are pending
Native App Locale Refresh / Refresh native ja-JP (push) Blocked by required conditions
Native App Locale Refresh / Refresh native ko (push) Blocked by required conditions
Native App Locale Refresh / Refresh native nl (push) Blocked by required conditions
Native App Locale Refresh / Refresh native pl (push) Blocked by required conditions
Native App Locale Refresh / Refresh native pt-BR (push) Blocked by required conditions
Native App Locale Refresh / Refresh native ru (push) Blocked by required conditions
Native App Locale Refresh / Refresh native sv (push) Blocked by required conditions
Native App Locale Refresh / Refresh native th (push) Blocked by required conditions
Native App Locale Refresh / Refresh native tr (push) Blocked by required conditions
Native App Locale Refresh / Refresh native uk (push) Blocked by required conditions
Native App Locale Refresh / Refresh native vi (push) Blocked by required conditions
Native App Locale Refresh / Refresh native zh-CN (push) Blocked by required conditions
Native App Locale Refresh / Refresh native zh-TW (push) Blocked by required conditions
Native App Locale Refresh / Commit native locale refresh (push) Blocked by required conditions
OpenClaw Stable Main Closeout / Resolve stable release closeout inputs (push) Waiting to run
OpenClaw Stable Main Closeout / Verify stable main closeout (push) Blocked by required conditions
Plugin Init Scaffold Validation / Validate provider scaffold (push) Waiting to run
Plugin NPM Release / preview_plugins_npm (push) Waiting to run
Plugin NPM Release / Validate release publish approval (push) Blocked by required conditions
Plugin NPM Release / preview_plugin_pack (push) Blocked by required conditions
Plugin NPM Release / Preflight plugin npm package () (push) Blocked by required conditions
Plugin NPM Release / Seal prepared plugin npm release (push) Blocked by required conditions
Plugin NPM Release / Trusted publisher OIDC exchange (push) Blocked by required conditions
Plugin NPM Release / publish_plugins_npm (push) Blocked by required conditions
Plugin NPM Release / verify_plugins_npm (push) Blocked by required conditions
Vitest Cache Warm / warm (linux) (push) Waiting to run
Vitest Cache Warm / warm (macos) (push) Waiting to run
Workflow Sanity / no-tabs (push) Waiting to run
Workflow Sanity / actionlint (push) Waiting to run
Workflow Sanity / generated-doc-baselines (push) Waiting to run
2026-09-14 07:27:51 -07:00
Peter Steinberger
520442c337
test(agents): add live Code Mode comparison workloads (#147537)
Extend the existing model matrix with isolated Gateway tasks, fixed-workload build comparisons, and separate task and interview evidence. Preserve failed trials, exact source identities, task effects, logical cell completion, and observed preview coverage.

Validation: 187 focused tests, changed-file and targeted type/lint/docs checks, exact-candidate runtime build, actual OpenAI smoke, read-only replay of 24 original trials plus the final smoke, and fresh P2 review. Original scores and interview qualifications remain preserved.
2026-09-13 16:32:28 -07:00
RoboClaw
d8af304112
improve(qa): extend Code Mode evaluation scenarios (#140913)
Extend the existing Code Mode evaluation harness with opt-in large-result reduction, independent-read composition and dependent-chain tasks. Preserve default matrix size and distinguish missing telemetry from observed zeros.

Focused tests, negative controls, dry-run evidence, changed checks, independent review and exact-head CI pass. No live model performance result is claimed.

Worked on by:
- @Takhoffman

Co-authored-by: Takhoffman <781889+Takhoffman@users.noreply.github.com>
2026-09-06 23:27:26 -07:00
Peter Steinberger
eeec0d4b2a
fix(cli): keep setup and agent help startup lazy (#133272)
* fix(cli): keep command actions out of help registration

Defer onboarding rejection and agent finalization imports until actions execute. Split agent-exec input/config preparation and result projection into actual production owners while preserving helper bodies and test contracts. Production LOC stays neutral.

* refactor(cli): complete agent exec owner migration
2026-08-30 04:46:54 -07:00
Vincent Koc
63f265c0dd
refactor(scripts): share required option arguments (#133095)
* refactor(scripts): share required option arguments

* test(scripts): copy option helper into OpenGrep fixture
2026-08-30 14:17:08 +08:00
Peter Steinberger
c70aee247e
refactor(scripts): migrate JavaScript tools to TypeScript (#121005)
* refactor(scripts): migrate JavaScript tools to TypeScript

* fix(ci): keep changed-scope preflight zero-install

* fix(ci): preserve zero-install script owners

* fix(ci): complete script migration follow-through

* fix(release): keep stable closeout zero-install

* fix(scripts): preserve standalone execution boundaries

* fix(scripts): repair standalone loader boundaries

* fix(scripts): normalize gateway observation ids

* fix(scripts): keep Docker packager standalone

* test(scripts): preserve rebase cleanup helpers

* test(sessions): use tracked temp directory
2026-08-09 07:21:35 -07:00
Peter Steinberger
6ec84bc3bd
fix(models): expose custom max and ultra reasoning tiers (#115690)
* fix(models): honor custom advanced reasoning levels

* fix(ci): restore code mode matrix boundaries
2026-07-29 03:33:30 -04:00
Vincent Koc
9daa60961c
fix(agents): serialize sandbox provisioning (#115645)
* fix(agents): serialize sandbox provisioning

* test(docker): add sandbox browser sidecar e2e

* fix(ci): register sandbox browser e2e entrypoint

* fix(ci): restore code mode matrix checks
2026-07-29 14:50:15 +08:00
Dallin Romney
7a8df4920d
fix(ci): type code mode matrix catch callback (#115674) 2026-07-29 14:24:57 +08:00
Vincent Koc
1c36054494
feat(agents): add code mode model acceptance matrix (#115305)
* feat(agents): add code mode model acceptance matrix

* fix(qa): emit canonical Code Mode matrix evidence

* fix(qa): reserve code mode matrix output safely

* fix(qa): require fresh matrix builds

* fix(qa): protect matrix evidence paths
2026-07-29 14:12:20 +08:00