mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
* feat(moe): stream split multi-shard ggufs natively + DeepSeek V4 Flash recipe Hugging Face rejects single files above 50 GB, so every large model ships as -00001-of-0000N.gguf shards; until now the streamer assumed one file, forcing a merge with double the disk. gguf_offsets now fans the first shard out to the whole set and resolves every tensor to (shard, offset); the expert streamer and the dense loader open one positioned reader per shard and route each read by the tensor's shard index. Pass the first shard, exactly as llama.cpp takes it; a missing sibling fails the load with the shard named. Add the deepseek4 recipe row: V3.2-style routing (256 routed experts, a per-expert bias like lfm2moe, an always-on shared expert that stays resident) over the standard split expert suffixes. The V4 compressed-attention machinery is dense-side llama.cpp code, invisible to the streaming seam. The byte-identity gates gain a 4-shard qwen3moe fixture (metadata-only first shard, the layout large quants actually use); make-tiny-moe.py learns --split-max-tensors. All gates pass, split included. * fix(moe): cache auto must budget for the anon dense conversion The auto budget read MemAvailable while the dense weights were still reclaimable page cache, then dense-weights=anon converted them into buffers the kernel cannot take back: the same bytes planned twice. Latent since the anon policy shipped (dense sets were 2-3 GiB and explicit budgets were the benched path); DeepSeek V4 Flash's 6.5 GiB dense set turned it into a device-taking overcommit on first load. The budget now deducts the pending conversion and says so in the log. * fix(moe): review pass on the multi-shard path Three defects the split rewrite introduced, none of which the gates could see: - The shard index rode in an int8_t, so a model past 127 shards wrapped to a negative index into the reader vector. The bounds check could never catch it: it validated the untruncated value. Widened to int16_t, which covers the whole -%05d-of-%05d filename space. - DenseWeights::warm() reused one flag as both the inner loop condition and the partial-warm report, so the first shard that failed to open silently skipped the warm-up of every later shard. Per-shard condition, sticky report. - The dense readers stayed allocated for the session after read_anonymous had copied and rebound every tensor: fds and a per-lane bounce buffer per shard, sitting next to a cache counting every MiB. Released at the end of init. Also: the streaming banner read O_DIRECT off shard 0, which under the small-first-shard layout is metadata only and too short to verify, so it could claim a mode the shards carrying experts had not got. It now reports the weakest of the readers. * build: the engine version says 0.19.0, like the changelog does The version is declared in CMakeLists.txt and reported by `--version` and by the run-parameter preamble of every metrics CSV, so a committed benchmark file names the engine that produced it. This release section landed while the number stayed at 0.18.0, which would have stamped the wrong engine on every CSV this branch produces, defeating the one purpose the string has.
3.3 KiB
3.3 KiB
Limitations and prior art
Prior art
BigMoeOnEdge is an engineering package, not a new technique. The ideas it combines:
- AirLLM — layer-by-layer streaming of >RAM models from disk.
- Apple, "LLM in a flash" — flash-aware weight streaming, windowing, sparsity-driven loading.
- FlexGen — offloading and I/O-bound throughput scheduling for large models.
- PowerInfer / EdgeMoE — hot/cold expert locality and expert-granularity residency on the edge.
The contribution here is a clean, modular, llama.cpp-native implementation of
expert-selective streaming that stays lossless and runs on the public API — no fork for the
serial path, and only a single ~25-line hook (with an explicit sunset) for the optional
--overlap feature. See seam.md § 3.
Limitations
- One setting makes output non-reproducible. Every other knob is deterministic given a
configuration:
--n-expert-usedchanges the output, but changes it the same way on every run.--drop-cold-expertsdecides per routing from live cache state, so the same prompt and the same flags can decode differently run to run, and the byte-identity gates cannot cover its output — only its machinery. Off by default in the CLI. - n=1 only. The expert sparsity exists only for single-token decode, so streaming is incompatible with speculative decoding or batching. Prefill streams the union of the prompt's routed experts (still far below the full bank, but larger than one token's).
- CPU experts. Streamed experts are computed on CPU; the rebind targets host memory. GPU offload of the streamed experts is not supported (the dense parts can still use the GPU). Decode is flash-I/O-bound anyway, so this is rarely the bottleneck.
- Shared experts stay resident. Architectures with an always-on shared expert (e.g.
gemma4,deepseek4) stream the routed experts but keep the shared expert — and any dense layers — resident (in the page cache, or in the engine's own buffers under--dense-weights anon), so the streamed fraction (and the memory saving) is smaller than for a purely routed model likeqwen3moe. The same applies to architectures whose first blocks are dense by design (lfm2moehas aleading_dense_block_count): those blocks name no expert tensors, so they are never streamed. - Streaming does not help a model that fits. The engine's reason to exist is a model larger than RAM. Registering an architecture says the layout streams losslessly, not that streaming is the fast way to run every model using it — a small MoE that fits in memory is faster loaded resident, and the registry rows are about coverage, not a recommendation.
- Repack must stay off. Loading uses
use_extra_bufts=false; you cannot combine streaming with weight repacking. - Windows throughput. The cache's reserve-then-commit-per-slice path is heavier on Windows than the POSIX lazy-commit path. The gates run on Windows; the throughput targets are stated for Android/Linux.
- Depends on a ggml scheduling behaviour (documented in seam.md) that is not a stability-guaranteed contract. Re-verified by the gates on each submodule bump.
Not goals
- Distributing a model across devices (a different axis).
- Beating a model that already fits in RAM — if it fits, run it resident.