GgufHeader.arch() read general.architecture out of the header "to pick the right chat turn format and to gate arch-specific prompt switches", and nothing ever called it. Neither job is the app's to do. --chatml initialises the chat templates from the model itself, so the format comes from the gguf that declares it and ChatML is only llama.cpp's fallback for a model that ships none; the one arch-conditional decision the app does act on, whether thinking can be turned off, is measured by the engine at open() and arrives as think_ctl on BMOE_READY. Wiring the probe up would therefore have meant choosing behaviour from a model name, which is what session.h asks this codebase not to do. Deleting it takes parseArch and the magic/version reader with it — parseIsMoe reads the magic inline, so isMoe(), the single entry point in use, is unaffected. The --chatml comment in AppSettings said "(gemma / chatml)", copied from the CLI help, which reads as though the app picks between the two. Reworded to say what the flag does and to note the name is historical. Kotlin compiles clean, which is also the check that no caller was left behind. |
||
|---|---|---|
| .. | ||
| app | ||
| gradle/wrapper | ||
| build.gradle | ||
| gradle.properties | ||
| gradlew | ||
| gradlew.bat | ||
| README.md | ||
| settings.gradle | ||
BigMoeOnEdge — Android example
A minimal chat app that validates the throughput claim on a real phone: pick a pushed
.gguf, type a prompt, and watch the answer stream in while a live panel shows tok/s and
the per-token compute-vs-flash-I/O split and cache hit rate.
It runs the engine as the bmoe-cli binary (shipped as libbmoe-cli.so) via
ProcessBuilder from a foreground service — no JNI. This is the same pattern used by the
research harness and keeps the app a thin driver over the CLI.
Build
-
Cross-compile and stage the engine binaries (needs the Android NDK):
pwsh ../../scripts/build-android.ps1This fills
app/src/main/jniLibs/arm64-v8a/withlibbmoe-cli.soand thelibllama/libggmlshared libraries. -
Build and install the APK. Open this folder in Android Studio, or use the committed Gradle wrapper directly. The app has two distribution flavors (see below); build the one you want:
./gradlew assembleDevDebug adb install app/build/outputs/apk/dev/debug/app-dev-debug.apkPublished sideload builds are release-signed with a stable key instead, so an update installs over the previous one rather than being refused. That needs a
keystore.propertiesnext toapp/(gitignored — it points at the keystore and holds its passwords); without it the release build falls back to debug signing.
Flavors
Two build flavors differ only in how a model reaches the device:
- dev — sideloaded (this is what CI attaches to releases). Keeps all-files access, so it
can also read a model adb-pushed to shared storage. Application id
….example.dev. - play — Play-Store-compliant. No broad storage permission: models come only through the
in-app downloader or the file picker.
./gradlew assemblePlayDebug.
Getting a model onto the device
The picker lists every MoE .gguf it finds (dense models are filtered out by a gguf-header
check). Nothing below needs a storage permission except the last option.
-
Built-in catalog (both flavors) — the "Get a model" card offers the models this engine is measured on, each a single tap: Qwen3-30B-A3B-Q4_K_M (~18.6 GB, the reference model), Qwen3.6-35B-A3B-Q4_K_M (~22.3 GB, a hybrid attention/SSM MoE, comfortably past device RAM) and Gemma-4-26B-A4B-it-Q4_K_M (~17 GB). Downloads run in a foreground worker, survive the app being killed, resume an interrupted transfer instead of restarting, and appear in the picker when done.
-
Any other model — under Other model, paste a direct gguf URL (e.g. a Hugging Face
…/resolve/main/model.gguflink), or pick a.ggufalready on the device to import it.In-app downloads and picker imports both land in the app's internal storage (
filesDir, a real f2fs/ext4 volume), so the streamed expert reads use O_DIRECT at full speed. Only models read from the emulated external dirs (adb-pushed to/sdcard/Download) fall back to buffered I/O. A download needs free space equal to the model size — no temporary second copy. -
adb push (dev flavor only — needs all-files access, which the dev build requests):
adb push Qwen3-30B-A3B-Q4_K_M.gguf /sdcard/Download/ # /data/local/tmp/bmoe avoids duplicating a model too big to copy, and is on a real # filesystem where O_DIRECT works (the emulated dirs fall back to buffered I/O) adb push Qwen3-30B-A3B-Q4_K_M.gguf /data/local/tmp/bmoe/This directory was named
shardllmbefore v0.8.0. To keep models already pushed there:adb shell mv /data/local/tmp/shardllm /data/local/tmp/bmoe
gpt-oss-120b
Listed in the catalog but not downloadable in-app: Hugging Face ships the Q4_K_M quant as two shards (the 50 GB per-file limit), and expert streaming reads tensors by byte offset from a single file. Merge the shards on a PC, then transfer the result:
llama-gguf-split --merge gpt-oss-120b-Q4_K_M-00001-of-00002.gguf gpt-oss-120b-Q4_K_M.gguf
adb push gpt-oss-120b-Q4_K_M.gguf /data/local/tmp/bmoe/ # or import it with the file picker
Expected numbers
On a phone with UFS 4.x storage and ~12 GB RAM, streaming Qwen3-30B-A3B-Q4_K_M with the
expert cache at 4000 MiB, 4 I/O lanes and 4 compute threads, decode settles around
0.55–0.6 s/token (~1.8 tok/s) — a model ~1.7× the device RAM, lossless. That 4000 MiB is a
sweep point from the benchmark protocol, not the app default: the app ships a fixed 2000 MiB
expert cache. See ../../docs/benchmark-method.md for the full procedure and the cache/thread
sweep.