Summary: - Add the `zed-eval` Python CLI for Modal/Harbor/Pier benchmark orchestration, including content-addressed remote builds, run/suite management, reporting, rejudge, baseline, and cleanup workflows. - Extend `eval-cli` for remote evals with provider/model overrides and step/tool-call metrics in `result.json`. - Add install/source-run helper scripts so `zed-eval` can be installed or run from the checkout without manually setting `PYTHONPATH`. - Harden the remote harness wrappers around exit-code preservation, archive extraction, custom secret wiring, and Harbor/Pier option parity, with regression coverage. Testing: - Using the CLI for two weeks - `PYTHONPATH=crates/eval_cli python3 -m compileall -q crates/eval_cli/zed_eval` - `uv run --project crates/eval_cli/zed_eval python -m unittest discover -s crates/eval_cli/zed_eval/tests` - `bash -n crates/eval_cli/script/install-zed-eval crates/eval_cli/script/zed-eval` - `cargo check -p eval_cli` - `cargo fmt --package eval_cli -- --check` - `cargo test -p eval_cli --no-run` - `./script/clippy -p eval_cli` Release Notes: - N/A
2.3 KiB
eval-cli
Headless Rust binary for running Zed's agent in evaluation and benchmark environments. It is designed for containerized harnesses such as Harbor and Pier, where the repository is already checked out and model API keys are provided via environment variables.
eval-cli uses the same NativeAgent + AcpThread pipeline as the production
Zed editor: a full agentic loop with tool calls, subagents, and retries, without
a GUI.
This directory also contains zed_eval/, the Python zed-eval package used to
build this binary, launch remote benchmark runs on Modal/Harbor/Pier, and fetch
results. For normal benchmark orchestration, start with
zed_eval/README.md.
Building
Native, for local testing on the same OS
cargo build --release -p eval_cli
Linux x86_64, for Harbor/Pier sandboxes
Harbor and Pier containers run Linux x86_64. From the repository root, use the Docker-based build script:
crates/eval_cli/script/build-linux
This produces target/eval-cli, an x86_64 Linux ELF binary. You can also
specify a custom output path:
crates/eval_cli/script/build-linux --output ~/bin/eval-cli-linux
Standalone usage
eval-cli \
--workdir /testbed \
--model anthropic/claude-sonnet-4-6 \
--instruction "Fix the bug described in..." \
--timeout 600 \
--output-dir /logs/agent
eval-cli reads provider API keys from environment variables such as
ANTHROPIC_API_KEY and OPENAI_API_KEY. It writes result.json, thread.md,
and thread.json to the output directory.
Exit codes
| Code | Meaning |
|---|---|
| 0 | Agent finished |
| 1 | Error, such as model/auth/runtime failure |
| 2 | Timeout |
| 3 | Interrupted by SIGTERM or SIGINT |
Running benchmarks
Most benchmark runs should use the Python zed-eval CLI instead of invoking
eval-cli directly. From the repository root:
crates/eval_cli/script/install-zed-eval
zed-eval doctor --create-volume
zed-eval run rf --from local --n-tasks 2
For one-off source runs without installing the tool globally, use
crates/eval_cli/script/zed-eval <args>.
See zed_eval/README.md for supported benchmarks, remote
builds, Modal setup, reporting, rejudging, baselines, and Harbor/Pier installed
agent usage.