BigMoeOnEdge/third_party
Helldez 927e2d3b31
feat(moe): Qwen3.8-Flash-Next support (#172)
Qwen3.8-Flash-Next (qwen4exp): 125B total, ~6B active, 512 routed experts at
top-10 plus one shared, 48 hybrid gated-delta SSM / sparse attention layers,
and a 51B n-gram embedding table. One registry row streams the experts; a
dense-policy guard keeps the n-gram table (larger than any phone's RAM)
mmap'd under every mode so pinned and anonymous dense weights survive load.
Runs on the 12 GB test phone at ~2 tok/s with pinned dense weights, compute-
bound, and sits in the app catalog as a three-shard download. Submodule
pinned to upstream master b10666, the first with the architecture merged,
with the expert-ready hook on top. README hero clip, changelog and docs
updated. App 0.22.0 (versionCode 37).
2026-08-28 10:07:43 +02:00
..
llama.cpp@0e8c83e512 feat(moe): Qwen3.8-Flash-Next support (#172) 2026-08-28 10:07:43 +02:00