feat(moe): --io-two-wave, publish a layer's read batch in two waves (#128)

A cold layer's batch became visible to the I/O lanes only after every
miss took its page commits - up to three vm_commit syscalls per cold
expert of bookkeeping sitting in front of the first byte of I/O, which
is the latency-to-first-slice the sidecar refutation (PR #90) identified
as the binding constraint. (#118)

With the flag on, only the first present projection - the one
mul_mat_id blocks on first - is committed up front; its jobs publish
and wake the lanes immediately, and the remaining projections are
committed and appended while the lanes already read.

The drain protocol grew the one thing this needs:

- io_drain copies each job out under the lock, so jobs_ growing (and
  possibly reallocating) mid-batch cannot leave a worker holding a
  dangling reference;
- the worker wait predicate admits next_idx_ < batch_njobs_, so a
  worker that drained wave one and left comes back for a batch that
  grew in the SAME generation - the gen comparison alone never would;
- a wave-two commit failure goes fatal and wakes the ready waiters,
  because wave one already published flags this batch will never flip.

Batch completion cannot fire between the waves: the only thread that
waits on done_cnt_ == batch_njobs_ is the eval thread, and it is the
one appending wave two.

Overlap + LRU cache only (validate() enforces both); recorded in the
CSV preamble as io_two_wave. Default off: the win is bounded by the
commit cost per cold layer, and the failure mode of a drain-protocol
bug is a hang the host cannot reproduce - so a new gate (G4d) holds
two-wave output byte-identical to serial streaming, and the flag stays
off until the on-device A/B (#120) says the win is real.
This commit is contained in:
Helldez 2026-07-28 15:46:52 +02:00 • committed by GitHub
parent cc4aed1304
commit ad13038ce0
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
11 changed files with 172 additions and 37 deletions

View file

@ -167,7 +167,7 @@ prints just the summary lines.
# model=<file> arch=<arch> n_layer=<n> n_expert=<n> n_expert_used=<k> threads=<n>
n_ctx=<n> n_ubatch=<n> chatml=<0|1>
# moe_stream=<0|1> cache_mb=<n> cache_auto=<0|1> cache_floor_mb=<n> cache_ceil_mb=<n>
force_cache=<0|1> load_all=<0|1> io_threads=<n> o_direct=<0|1> overlap=<0|1> prefetch=<n>
force_cache=<0|1> load_all=<0|1> io_threads=<n> o_direct=<0|1> overlap=<0|1> io_two_wave=<0|1> prefetch=<n>
predict_prefetch=<0|1> predict_log=<0|1> predict_spec_max=<n> prefetch_sync=<0|1>
dense_weights=<mmap|warm|anon|ahwb> drop_cold_frac=<f> drop_renorm=<0|1> drop_prefill=<0|1>
# temp=<f> top_k=<n> top_p=<f> seed=<u> compute_trace_layers=<n>