mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
feat(moe): --io-two-wave, publish a layer's read batch in two waves (#128)
A cold layer's batch became visible to the I/O lanes only after every miss took its page commits - up to three vm_commit syscalls per cold expert of bookkeeping sitting in front of the first byte of I/O, which is the latency-to-first-slice the sidecar refutation (PR #90) identified as the binding constraint. (#118) With the flag on, only the first present projection - the one mul_mat_id blocks on first - is committed up front; its jobs publish and wake the lanes immediately, and the remaining projections are committed and appended while the lanes already read. The drain protocol grew the one thing this needs: - io_drain copies each job out under the lock, so jobs_ growing (and possibly reallocating) mid-batch cannot leave a worker holding a dangling reference; - the worker wait predicate admits next_idx_ < batch_njobs_, so a worker that drained wave one and left comes back for a batch that grew in the SAME generation - the gen comparison alone never would; - a wave-two commit failure goes fatal and wakes the ready waiters, because wave one already published flags this batch will never flip. Batch completion cannot fire between the waves: the only thread that waits on done_cnt_ == batch_njobs_ is the eval thread, and it is the one appending wave two. Overlap + LRU cache only (validate() enforces both); recorded in the CSV preamble as io_two_wave. Default off: the win is bounded by the commit cost per cold layer, and the failure mode of a drain-protocol bug is a hang the host cannot reproduce - so a new gate (G4d) holds two-wave output byte-identical to serial streaming, and the flag stays off until the on-device A/B (#120) says the win is real.
This commit is contained in:
parent
cc4aed1304
commit
ad13038ce0
11 changed files with 172 additions and 37 deletions
|
|
@ -167,7 +167,7 @@ prints just the summary lines.
|
|||
# model=<file> arch=<arch> n_layer=<n> n_expert=<n> n_expert_used=<k> threads=<n>
|
||||
n_ctx=<n> n_ubatch=<n> chatml=<0|1>
|
||||
# moe_stream=<0|1> cache_mb=<n> cache_auto=<0|1> cache_floor_mb=<n> cache_ceil_mb=<n>
|
||||
force_cache=<0|1> load_all=<0|1> io_threads=<n> o_direct=<0|1> overlap=<0|1> prefetch=<n>
|
||||
force_cache=<0|1> load_all=<0|1> io_threads=<n> o_direct=<0|1> overlap=<0|1> io_two_wave=<0|1> prefetch=<n>
|
||||
predict_prefetch=<0|1> predict_log=<0|1> predict_spec_max=<n> prefetch_sync=<0|1>
|
||||
dense_weights=<mmap|warm|anon|ahwb> drop_cold_frac=<f> drop_renorm=<0|1> drop_prefill=<0|1>
|
||||
# temp=<f> top_k=<n> top_p=<f> seed=<u> compute_trace_layers=<n>
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue