mirror of
https://github.com/QwenLM/qwen-code.git
synced 2026-08-04 05:40:58 +00:00
* feat(triage): sponsored /verify for external PRs on ephemeral runners
/verify used to require the PR AUTHOR to hold write access, because the
lane executes the author's code on the persistent ECS pool — which
excluded external contributors' PRs entirely (16 of 38 open PRs
eligible when measured on 2026-07-28), and those are the PRs where
sandboxed evidence is worth the most. The practical alternative was a
maintainer running the verification on their own machine, which is
strictly worse: personal credentials, SSH keys, and no isolation at
all. The risk was not avoided, only relocated.
The author's permission now ROUTES instead of gating, using the same
conditional runs-on pattern the authorize/triage jobs already use for
fork PRs:
- write-access author -> the persistent ECS pool, unchanged;
- external author -> a SPONSORED run on an ephemeral GitHub-hosted VM
(destroyed after the job), free on a public repository.
The commenter gate is unchanged: only a write-access maintainer can
trigger either lane, so runs stay maintainer-metered. Their comment is
the sponsorship, and authorize snapshots the head OID at that moment;
the resolve step refuses to run any other commit, so code pushed while
the job queues is never executed on the sponsor's approval. On the ECS
lane the execution-time re-check keeps its old meaning (the author's
write permission is the control there, and it can be revoked while
queued); its refusal now names the sponsored alternative.
Before anything executes on the hosted lane, a pre-execution risk
screen reads the PR diff as data: mechanical checks (added npm
lifecycle scripts, dependencies resolving outside registry.npmjs.org,
package-manager config changes, long opaque single-line content) and a
model screen, every arm failing CLOSED — an unfetchable diff, missing
screen config, a model error, and an unparseable reply each refuse the
run. Reasons posted to the PR are fixed strings chosen by the
workflow, never diff text, so a crafted diff cannot steer the bot's
comment. The screen is defense in depth, not the control: what bounds
a malicious PR is the ephemeral VM. The model can be fooled;
the machine being destroyed cannot.
The ECS kill switch no longer strands the hosted lane (an ephemeral
runner does not depend on the pool), and the authorize-side denial
notice for external authors is gone — the case it explained no longer
denies. The 2b-bis skill guidance now offers the sponsored run instead
of declaring the lanes unavailable, with an explicit warning that an
external author's report is adversarial input to be read like the
fork's own CI logs.
Mutation-verified 8/8, each with landing proof: dropping the commenter
gate, inverting the routing, dropping the OID snapshot, letting a
permission-API failure route instead of deny, dropping the head-moved
check, dropping the lifecycle-script check, tolerating an unparseable
model reply, and pinning the hosted lane back to ECS each turn a test
red. The screen tests run the real resolve-step text against a stubbed
gh and a live local HTTP server standing in for the model endpoint —
in a separate process, because spawnSync blocks the test runner's
event loop and an in-process server would deadlock into the fetch's
120 s abort instead of testing anything.
actionlint category diff vs main: -3 SC2016:info (the removed denial
step), +3 SC2129:style (new GITHUB_OUTPUT groups); both non-gating.
* docs(triage): persistence, not lateral movement, is the ECS threat
The maintainer confirmed the ECS runners are network-isolated, so the
comments justifying the hosted routing now name the decisive risk
correctly: machines are reused and the next run arrives carrying
credentials that planted artifacts could harvest. Lateral movement was
never the load-bearing argument; persistence is, and it is the one
threat a destroyed VM removes without depending on sweep completeness.
* feat(triage): run external /verify on ECS behind a workspace wipe
Maintainer decision (2026-07-29): external contributors' PRs run on the
persistent ECS pool like everyone else, not on ephemeral GitHub-hosted
VMs. The pool is network-isolated, and the realistic alternative — a
maintainer verifying on a personal machine with real credentials and no
isolation — is worse than a hardened shared runner.
That removes machine destruction from the design, so the author's
permission now selects CONTROLS rather than a machine:
trusted -> unchanged: write permission is the control, re-checked at
execution time, head may move freely.
external -> pinned head OID (the sponsoring comment approves the code
that existed then), pre-execution risk screen, and a full
workspace wipe.
The wipe is what replaces destruction. The residual ECS risk for
untrusted code is PERSISTENCE: machines are reused, and the next run
starts by handling credentials (CI_BOT_PAT in the resolve step, the
model key in the proxy). The existing sweep is an allowlist of vectors
we have thought of, and an adversarial author gets unlimited attempts to
find one we have not — so external runs delete the workspace outright
before checkout instead of sweeping it. Deny-by-default costs one cold
clone and removes hooks, worktree config, planted dotfiles, staged
artifacts and out-of-tree symlinks in one move, known or not. Scope is
stated in the step: RUNNER_TEMP and /tmp also persist, and the steps
reading them already rm -rf their own subtrees; the container is
recreated per job, so /home/node and its npm cache are not cross-run
state.
The screen's status changes with the machine: on ephemeral hardware it
was a cost-saver, here it is load-bearing — but still not sufficient
alone, and the comment says so, naming the wipe and the sweep as what
actually bounds a screened-but-hostile diff.
Mutation-verified 6/6: removing the wipe, moving it after checkout,
applying it to every run, tolerating a non-empty workspace, dropping the
suspicious-path guard, and classifying an external author as trusted
each turn a test red.
Also fixes a live hazard in the guard's own test. It passed real paths
(/, /usr, /root) to a live `rm -rf`, which is safe only while the guard
exists — so it detonates precisely when the guard is removed, which is
when it is meant to protect you. It did: a mutation run that deleted the
guard spent six minutes attempting to delete / before being killed.
macOS permissions absorbed it and nothing was lost, but that outcome was
luck, not design. `rm` is now stubbed to a recorder on PATH, so the
destructive primitive cannot fire from the test suite under any edit,
and the assertion is on the decision — with the guard gone the recorder
shows the attempted delete and the test fails, having deleted nothing.
* fix(triage): close two screen bypasses and a vacuous routing assertion
Review round on #7985 found four issues; all four were real.
1. npm lifecycle alternation stopped at the install lifecycle, but npm
runs pre/post hooks for ANY `npm run <script>`. This job runs
`npm run build` and the agent runs the PR's suites, so a `prebuild`
or `pretest` executes exactly like a `postinstall` and walked past
the mechanical screen. Added build and test to the alternation.
2. The off-registry check excluded any line CONTAINING
`registry.npmjs.org`, so a lookalike host — registry.npmjs.org.evil.com,
or an @-userinfo form — read as npmjs and passed. The tarball's own
postinstall lives inside the tarball, never in the diff, so this arm
was the only mechanical defense against it. The exclusion is now
anchored to scheme + host + '/'.
3. A routing assertion read `.runner`, a key from the earlier
ephemeral-VM design that `gate()` never returns, so it was
`expect(undefined).toBeUndefined()` — green no matter how badly the
trust level leaked into /triage or /tmux. Asserts the real keys now,
on both non-verify paths.
4. Three screen arms had no scenario at all: package-manager config,
the long-opaque-line awk, and the unconfigured-model fail-closed
path. All three are covered, each with a negative control so the arm
is not merely refusing everything — the genuine registry URL and a
700-char lockfile integrity line must still pass.
Mutation-verified 6/6, each failure naming the payload it let through:
reverting the alternation (`prebuild was not flagged`), un-anchoring the
registry match (`registry.npmjs.org.evil.com was not flagged`), breaking
the package-manager grep (`.yarnrc was not flagged`), raising the awk
threshold, letting the unconfigured path fall through to clear, and
leaking verify_trust into /triage — which is exactly what the vacuous
assertion could never see.
93/93 tests; prettier, eslint and shellcheck clean.
* fix(triage): close post-run persistence gap and screen bypasses (#7985)
Maintainer review found the wipe only ran BEFORE external code, leaving
the next pool job to meet the allowlist sweep alone. Add a post-run wipe
with `if: always()` so a cancelled or timed-out external job still
cleans up.
The model screen silently truncated at 200 KB — a fail-open in a
fail-closed step. Refuse oversized diffs mechanically before the model
call.
The screen prompt still described an "ephemeral, credential-free
sandbox" from the deleted ephemeral-runner design, calibrating the
model for leniency the persistent pool does not warrant. Describe the
actual environment. Also fix the stale "hosted lane" comment.
Three cheap mechanical-screen hardenings from the review: drop the
root anchor on the package-manager config grep (a nested .npmrc is
just as dangerous), match `rename to` headers (a rename-only hunk
emits no `+++ b/` line), and relax the opaque-line awk from
"no space in the line" to "no space-free field >= 600" so a single
space anywhere is no longer a bypass.
94/94 tests; build, typecheck, lint clean.
* test(triage): cover the unfetchable-diff fail-closed screen guard (#7985)
---------
Co-authored-by: wenshao <wenshao@example.com>
Co-authored-by: qwen-code-ci-bot <qwen-code-ci-bot@users.noreply.github.com>
Co-authored-by: qwen-code-dev-bot <qwen-code-dev@service.alibaba.com>
|
||
|---|---|---|
| .. | ||
| installation | ||
| lib | ||
| tests | ||
| acp-http-smoke.mjs | ||
| audit-runtime-critical.js | ||
| benchmark-api-latency.mjs | ||
| build-hosted-installation-assets.js | ||
| build-standalone-release.js | ||
| build.js | ||
| build_package.js | ||
| build_sandbox.js | ||
| build_vscode_companion.js | ||
| check-build-status.js | ||
| check-desktop-isolation.js | ||
| check-i18n.ts | ||
| check-lockfile.js | ||
| check-serve-fast-path-bundle.js | ||
| clean-package-build-artifacts.js | ||
| clean.js | ||
| cli-entry.js | ||
| copy_bundle_assets.js | ||
| copy_files.js | ||
| create-standalone-package.js | ||
| create_alias.sh | ||
| daemon-dev.js | ||
| desktop-openwork-sync.ts | ||
| dev.js | ||
| esbuild-shims.js | ||
| generate-changelog.js | ||
| generate-git-commit-info.js | ||
| generate-release-notes.js | ||
| generate-settings-schema.ts | ||
| get-release-version.js | ||
| lint.js | ||
| local_telemetry.js | ||
| measure-flicker.mjs | ||
| pre-commit.js | ||
| prepare-package.js | ||
| prepare.js | ||
| release-script-utils.js | ||
| run-java-daemon-sdk-e2e.ts | ||
| sandbox_command.js | ||
| sdk-node-exporter-stub.js | ||
| sign-release.sh | ||
| start.js | ||
| sync-computer-use-schemas.ts | ||
| telemetry.js | ||
| telemetry_gcp.js | ||
| telemetry_utils.js | ||
| test-rewind-e2e.sh | ||
| test-windows-paths.js | ||
| unused-keys-only-in-locales.json | ||
| upload-aliyun-oss-assets.js | ||
| verify-installation-release.js | ||
| version.js | ||
| workspaces.js | ||