mirror of
https://github.com/QwenLM/qwen-code.git
synced 2026-07-30 19:35:18 +00:00
|
Some checks are pending
E2E Tests / E2E Test (Linux) - sandbox:docker - shard 1/3 (push) Waiting to run
E2E Tests / E2E Test (Linux) - sandbox:docker - shard 2/3 (push) Waiting to run
E2E Tests / E2E Test (Linux) - sandbox:docker - shard 3/3 (push) Waiting to run
E2E Tests / E2E Test (Linux) - sandbox:none - shard 1/3 (push) Waiting to run
E2E Tests / E2E Test (Linux) - sandbox:none - shard 2/3 (push) Waiting to run
E2E Tests / E2E Test (Linux) - sandbox:none - shard 3/3 (push) Waiting to run
E2E Tests / E2E Test - macOS - shard 1/2 (push) Waiting to run
E2E Tests / E2E Test - macOS - shard 2/2 (push) Waiting to run
E2E Tests / channel-plugin E2E (nightly) (push) Waiting to run
E2E Tests / cron-interactive E2E (nightly) (push) Waiting to run
E2E Tests / web-shell Browser Regression (push) Waiting to run
SDK Java / ubuntu-latest / Java 11 (push) Waiting to run
SDK Java / ubuntu-latest / Java 17 (push) Waiting to run
SDK Java / macos-latest / Java 21 (push) Waiting to run
SDK Java / ubuntu-latest / Java 21 (push) Waiting to run
SDK Java / windows-latest / Java 21 (push) Waiting to run
SDK Java / Real daemon E2E / Java 11 (push) Waiting to run
Every gate compose-review enforces proves the agents READ the diff — coverage proves the transcripts, anchors prove the quotes — but none proves the review could tell good code from bad. Dogfooded: a weak-model run drafted nothing from its entire roster on a diff where stronger same-condition runs found a verified blocking Critical, and compose printed a bare, confident "Verdict: Approve". A reader takes that as evidence of quality when it is only absence of signal — the most dangerous shape a review verdict can have. Disclose it, deterministically. When the composed event is APPROVE and the plan's srcDiffLines — the same field the review topology is chosen from, with tests, docs and generated files excluded by construction — exceeds 100, the verdict line names the shape: Verdict: Approve — low signal: none of the 11 review agents reported a finding on a non-trivial diff (155 source diff lines) The event never moves on it: a cap would punish every genuinely clean diff, and nothing was in fact found. Docs-only and typo-class diffs keep their bare Approve — there, finding nothing is the expected outcome. The floor sits well past the typo-fix class (a tiny edit scattered one line per hunk stays under it) and at a fifth of the smallest diff the topology gate calls big. Both numbers in the line are the run's own: the roster the plan required (all on record at APPROVE, or coverage would have capped) and the plan's source-line count. Co-authored-by: verify <verify@local> |
||
|---|---|---|
| .. | ||
| acp-bridge | ||
| audio-capture | ||
| channels | ||
| chrome-extension | ||
| cli | ||
| core | ||
| cua-driver | ||
| desktop | ||
| mobile-mcp | ||
| sdk-java | ||
| sdk-python | ||
| sdk-typescript | ||
| vscode-ide-companion | ||
| web-shell | ||
| web-templates | ||
| webui | ||
| zed-extension | ||