zed/crates/eval_cli/README.md
Anant Goel 10f501d700
eval_cli: Add remote benchmark orchestration (#59802)
Summary:

- Add the `zed-eval` Python CLI for Modal/Harbor/Pier benchmark
orchestration, including content-addressed remote builds, run/suite
management, reporting, rejudge, baseline, and cleanup workflows.
- Extend `eval-cli` for remote evals with provider/model overrides and
step/tool-call metrics in `result.json`.
- Add install/source-run helper scripts so `zed-eval` can be installed
or run from the checkout without manually setting `PYTHONPATH`.
- Harden the remote harness wrappers around exit-code preservation,
archive extraction, custom secret wiring, and Harbor/Pier option parity,
with regression coverage.

Testing:

- Using the CLI for two weeks
- `PYTHONPATH=crates/eval_cli python3 -m compileall -q
crates/eval_cli/zed_eval`
- `uv run --project crates/eval_cli/zed_eval python -m unittest discover
-s crates/eval_cli/zed_eval/tests`
- `bash -n crates/eval_cli/script/install-zed-eval
crates/eval_cli/script/zed-eval`
- `cargo check -p eval_cli`
- `cargo fmt --package eval_cli -- --check`
- `cargo test -p eval_cli --no-run`
- `./script/clippy -p eval_cli`

Release Notes:

- N/A
2026-06-24 15:32:41 +00:00

2.3 KiB

eval-cli

Headless Rust binary for running Zed's agent in evaluation and benchmark environments. It is designed for containerized harnesses such as Harbor and Pier, where the repository is already checked out and model API keys are provided via environment variables.

eval-cli uses the same NativeAgent + AcpThread pipeline as the production Zed editor: a full agentic loop with tool calls, subagents, and retries, without a GUI.

This directory also contains zed_eval/, the Python zed-eval package used to build this binary, launch remote benchmark runs on Modal/Harbor/Pier, and fetch results. For normal benchmark orchestration, start with zed_eval/README.md.

Building

Native, for local testing on the same OS

cargo build --release -p eval_cli

Linux x86_64, for Harbor/Pier sandboxes

Harbor and Pier containers run Linux x86_64. From the repository root, use the Docker-based build script:

crates/eval_cli/script/build-linux

This produces target/eval-cli, an x86_64 Linux ELF binary. You can also specify a custom output path:

crates/eval_cli/script/build-linux --output ~/bin/eval-cli-linux

Standalone usage

eval-cli \
  --workdir /testbed \
  --model anthropic/claude-sonnet-4-6 \
  --instruction "Fix the bug described in..." \
  --timeout 600 \
  --output-dir /logs/agent

eval-cli reads provider API keys from environment variables such as ANTHROPIC_API_KEY and OPENAI_API_KEY. It writes result.json, thread.md, and thread.json to the output directory.

Exit codes

Code Meaning
0 Agent finished
1 Error, such as model/auth/runtime failure
2 Timeout
3 Interrupted by SIGTERM or SIGINT

Running benchmarks

Most benchmark runs should use the Python zed-eval CLI instead of invoking eval-cli directly. From the repository root:

crates/eval_cli/script/install-zed-eval
zed-eval doctor --create-volume
zed-eval run rf --from local --n-tasks 2

For one-off source runs without installing the tool globally, use crates/eval_cli/script/zed-eval <args>.

See zed_eval/README.md for supported benchmarks, remote builds, Modal setup, reporting, rejudging, baselines, and Harbor/Pier installed agent usage.