feat(moe): Qwen3.8-Flash-Next support (#172)

Qwen3.8-Flash-Next (qwen4exp): 125B total, ~6B active, 512 routed experts at
top-10 plus one shared, 48 hybrid gated-delta SSM / sparse attention layers,
and a 51B n-gram embedding table. One registry row streams the experts; a
dense-policy guard keeps the n-gram table (larger than any phone's RAM)
mmap'd under every mode so pinned and anonymous dense weights survive load.
Runs on the 12 GB test phone at ~2 tok/s with pinned dense weights, compute-
bound, and sits in the app catalog as a three-shard download. Submodule
pinned to upstream master b10666, the first with the architecture merged,
with the expert-ready hook on top. README hero clip, changelog and docs
updated. App 0.22.0 (versionCode 37).
This commit is contained in:
Helldez 2026-08-28 10:07:43 +02:00 • committed by GitHub
parent f9371b4ddb
commit 927e2d3b31
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
13 changed files with 488 additions and 14 deletions

View file

@ -95,7 +95,7 @@ check). Nothing below needs a storage permission except the last option.
adb shell mv /data/local/tmp/shardllm /data/local/tmp/bmoe
```
### Sharded models (gpt-oss-120b, DeepSeek V4 Flash)
### Sharded models (gpt-oss-120b, DeepSeek V4 Flash, Qwen3.8-Flash-Next)
Models above Hugging Face's 50 GB per-file limit ship as several shard files
(`-00001-of-0000N.gguf`). The engine streams a split set natively, so these download in-app
@ -111,7 +111,13 @@ the first one:
adb push DeepSeek-V4-Flash-0731-UD-IQ2_M-0000*-of-00003.gguf /data/local/tmp/bmoe/
```
Mind the space: DeepSeek V4 Flash UD-IQ2_M is ~91 GB on disk.
Mind the space: DeepSeek V4 Flash UD-IQ2_M is ~91 GB on disk, Qwen3.8-Flash-Next UD-IQ3_XXS ~82 GB.
Qwen3.8-Flash-Next needs **Dense weights = Pinned (dma-buf)** in Settings. Its dense side is 4.3 GB
and every token walks it: with Anon the kernel swaps it to zram and single tokens stall for 10-20 s,
with Mmap it is refaulted from flash every token. Its 51B n-gram table is held back automatically and
stays mmap'd whatever the setting; the engine says so on stderr at load. Keep the expert cache at
1000-1500 MiB on a 12 GB phone: the pinned dense set leaves no room for more.
## Expected numbers

View file

@ -51,8 +51,8 @@ android {
applicationId 'io.bigmoeonedge.example'
minSdk 29
targetSdk 34
versionCode 36
versionName '0.21.0'
versionCode 37
versionName '0.22.0'
buildConfigField 'String', 'GIT_SHA', "\"${gitSha}\""
ndk {
// The engine ships as prebuilt arm64 binaries staged by build-android.ps1.

View file

@ -153,6 +153,37 @@ object ModelCatalog {
),
),
),
Entry(
title = "Qwen3.8-Flash-Next",
quant = "UD-IQ3_XXS",
fileName = "Qwen3.8-Flash-Next-UD-IQ3_XXS-00001-of-00003.gguf",
approxBytes = 81_961_823_936L,
url = null,
// The dense side is 4.3 GB and is walked every token, so on a 12 GB phone this one
// runs only with the dense weights pinned (dma-buf); anon swaps, mmap refaults. The
// 51B n-gram table stays mmap'd on its own — see docs/android-memory.md.
blurb = "6B active of 125B, the Qwen4 preview. ~82 GB on disk; needs Pinned dense weights.",
shards = listOf(
Shard(
"Qwen3.8-Flash-Next-UD-IQ3_XXS-00001-of-00003.gguf",
"https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/resolve/main/UD-IQ3_XXS/" +
"Qwen3.8-Flash-Next-UD-IQ3_XXS-00001-of-00003.gguf?download=true",
10_946_624L,
),
Shard(
"Qwen3.8-Flash-Next-UD-IQ3_XXS-00002-of-00003.gguf",
"https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/resolve/main/UD-IQ3_XXS/" +
"Qwen3.8-Flash-Next-UD-IQ3_XXS-00002-of-00003.gguf?download=true",
49_567_921_344L,
),
Shard(
"Qwen3.8-Flash-Next-UD-IQ3_XXS-00003-of-00003.gguf",
"https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/resolve/main/UD-IQ3_XXS/" +
"Qwen3.8-Flash-Next-UD-IQ3_XXS-00003-of-00003.gguf?download=true",
32_382_955_968L,
),
),
),
)
/**