mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
Qwen3.8-Flash-Next (qwen4exp): 125B total, ~6B active, 512 routed experts at top-10 plus one shared, 48 hybrid gated-delta SSM / sparse attention layers, and a 51B n-gram embedding table. One registry row streams the experts; a dense-policy guard keeps the n-gram table (larger than any phone's RAM) mmap'd under every mode so pinned and anonymous dense weights survive load. Runs on the 12 GB test phone at ~2 tok/s with pinned dense weights, compute- bound, and sits in the app catalog as a three-shard download. Submodule pinned to upstream master b10666, the first with the architecture merged, with the expert-ready hook on top. README hero clip, changelog and docs updated. App 0.22.0 (versionCode 37). |
||
|---|---|---|
| .. | ||
| llama.cpp@0e8c83e512 | ||