mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
feat(moe): Qwen3.8-Flash-Next support (#172)
Qwen3.8-Flash-Next (qwen4exp): 125B total, ~6B active, 512 routed experts at top-10 plus one shared, 48 hybrid gated-delta SSM / sparse attention layers, and a 51B n-gram embedding table. One registry row streams the experts; a dense-policy guard keeps the n-gram table (larger than any phone's RAM) mmap'd under every mode so pinned and anonymous dense weights survive load. Runs on the 12 GB test phone at ~2 tok/s with pinned dense weights, compute- bound, and sits in the app catalog as a three-shard download. Submodule pinned to upstream master b10666, the first with the architecture merged, with the expert-ready hook on top. README hero clip, changelog and docs updated. App 0.22.0 (versionCode 37).
This commit is contained in:
parent
f9371b4ddb
commit
927e2d3b31
13 changed files with 488 additions and 14 deletions
|
|
@ -95,7 +95,7 @@ check). Nothing below needs a storage permission except the last option.
|
|||
adb shell mv /data/local/tmp/shardllm /data/local/tmp/bmoe
|
||||
```
|
||||
|
||||
### Sharded models (gpt-oss-120b, DeepSeek V4 Flash)
|
||||
### Sharded models (gpt-oss-120b, DeepSeek V4 Flash, Qwen3.8-Flash-Next)
|
||||
|
||||
Models above Hugging Face's 50 GB per-file limit ship as several shard files
|
||||
(`-00001-of-0000N.gguf`). The engine streams a split set natively, so these download in-app
|
||||
|
|
@ -111,7 +111,13 @@ the first one:
|
|||
adb push DeepSeek-V4-Flash-0731-UD-IQ2_M-0000*-of-00003.gguf /data/local/tmp/bmoe/
|
||||
```
|
||||
|
||||
Mind the space: DeepSeek V4 Flash UD-IQ2_M is ~91 GB on disk.
|
||||
Mind the space: DeepSeek V4 Flash UD-IQ2_M is ~91 GB on disk, Qwen3.8-Flash-Next UD-IQ3_XXS ~82 GB.
|
||||
|
||||
Qwen3.8-Flash-Next needs **Dense weights = Pinned (dma-buf)** in Settings. Its dense side is 4.3 GB
|
||||
and every token walks it: with Anon the kernel swaps it to zram and single tokens stall for 10-20 s,
|
||||
with Mmap it is refaulted from flash every token. Its 51B n-gram table is held back automatically and
|
||||
stays mmap'd whatever the setting; the engine says so on stderr at load. Keep the expert cache at
|
||||
1000-1500 MiB on a 12 GB phone: the pinned dense set leaves no room for more.
|
||||
|
||||
## Expected numbers
|
||||
|
||||
|
|
|
|||
|
|
@ -51,8 +51,8 @@ android {
|
|||
applicationId 'io.bigmoeonedge.example'
|
||||
minSdk 29
|
||||
targetSdk 34
|
||||
versionCode 36
|
||||
versionName '0.21.0'
|
||||
versionCode 37
|
||||
versionName '0.22.0'
|
||||
buildConfigField 'String', 'GIT_SHA', "\"${gitSha}\""
|
||||
ndk {
|
||||
// The engine ships as prebuilt arm64 binaries staged by build-android.ps1.
|
||||
|
|
|
|||
|
|
@ -153,6 +153,37 @@ object ModelCatalog {
|
|||
),
|
||||
),
|
||||
),
|
||||
Entry(
|
||||
title = "Qwen3.8-Flash-Next",
|
||||
quant = "UD-IQ3_XXS",
|
||||
fileName = "Qwen3.8-Flash-Next-UD-IQ3_XXS-00001-of-00003.gguf",
|
||||
approxBytes = 81_961_823_936L,
|
||||
url = null,
|
||||
// The dense side is 4.3 GB and is walked every token, so on a 12 GB phone this one
|
||||
// runs only with the dense weights pinned (dma-buf); anon swaps, mmap refaults. The
|
||||
// 51B n-gram table stays mmap'd on its own — see docs/android-memory.md.
|
||||
blurb = "6B active of 125B, the Qwen4 preview. ~82 GB on disk; needs Pinned dense weights.",
|
||||
shards = listOf(
|
||||
Shard(
|
||||
"Qwen3.8-Flash-Next-UD-IQ3_XXS-00001-of-00003.gguf",
|
||||
"https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/resolve/main/UD-IQ3_XXS/" +
|
||||
"Qwen3.8-Flash-Next-UD-IQ3_XXS-00001-of-00003.gguf?download=true",
|
||||
10_946_624L,
|
||||
),
|
||||
Shard(
|
||||
"Qwen3.8-Flash-Next-UD-IQ3_XXS-00002-of-00003.gguf",
|
||||
"https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/resolve/main/UD-IQ3_XXS/" +
|
||||
"Qwen3.8-Flash-Next-UD-IQ3_XXS-00002-of-00003.gguf?download=true",
|
||||
49_567_921_344L,
|
||||
),
|
||||
Shard(
|
||||
"Qwen3.8-Flash-Next-UD-IQ3_XXS-00003-of-00003.gguf",
|
||||
"https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF/resolve/main/UD-IQ3_XXS/" +
|
||||
"Qwen3.8-Flash-Next-UD-IQ3_XXS-00003-of-00003.gguf?download=true",
|
||||
32_382_955_968L,
|
||||
),
|
||||
),
|
||||
),
|
||||
)
|
||||
|
||||
/**
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue