Share repeated tooling parsing, projections, and fixture transforms while preserving command and generated-output contracts. Repair the OpenGrep help range so bootstrap code no longer replaces documented usage.
* feat: add selectable Code Mode executors
Default enabled Code Mode to trusted Node execution and move QuickJS into a bundled executor plugin. Preserve typed discovery, JavaScript-only execution, tool authorization, continuation ownership, and explicit legacy QuickJS selections. Add a web settings selector and document both security boundaries.
* refactor: finish Code Mode executor source cutover
Remove the retired core worker copies and regenerate config documentation for the requested QuickJS plugin. The 21 added plugin paths are standard plugin management fields; core and channel counts stay unchanged.
* refactor: narrow the Code Mode plugin contract
Keep only the executor and guest protocol exports consumed by the QuickJS plugin, budget that generic contract, and load the public plugin artifact through the existing runtime test boundary.
* refactor: align Code Mode workers with current runtime boundaries
Use the worker-side task server, keep asynchronous cleanup ownership explicit, and model the real Promise contracts in lifecycle fixtures. Regenerate the requested plugin config surface and budget the exact 35 public executor exports.
* refactor: keep executor implementation types private
* chore: regenerate code mode config baseline
* fix: satisfy Code Mode executor integration contracts
* test: cover QuickJS plugin metadata and exact settings titles
* test: retain QuickJS integration in the agent runtime suite
* test: keep Code Mode validation within lint and type-shard contracts
Extend the existing model matrix with isolated Gateway tasks, fixed-workload build comparisons, and separate task and interview evidence. Preserve failed trials, exact source identities, task effects, logical cell completion, and observed preview coverage.
Validation: 187 focused tests, changed-file and targeted type/lint/docs checks, exact-candidate runtime build, actual OpenAI smoke, read-only replay of 24 original trials plus the final smoke, and fresh P2 review. Original scores and interview qualifications remain preserved.
Extend the existing Code Mode evaluation harness with opt-in large-result reduction, independent-read composition and dependent-chain tasks. Preserve default matrix size and distinguish missing telemetry from observed zeros.
Focused tests, negative controls, dry-run evidence, changed checks, independent review and exact-head CI pass. No live model performance result is claimed.
Worked on by:
- @Takhoffman
Co-authored-by: Takhoffman <781889+Takhoffman@users.noreply.github.com>
* fix(cli): keep command actions out of help registration
Defer onboarding rejection and agent finalization imports until actions execute. Split agent-exec input/config preparation and result projection into actual production owners while preserving helper bodies and test contracts. Production LOC stays neutral.
* refactor(cli): complete agent exec owner migration