Getting a first model meant knowing which MoE to look for and pasting a raw Hugging Face link. The Get-a-model card now offers the models this engine is measured on — Qwen3-30B-A3B-Q4_K_M and Gemma-4-26B-A4B-it-Q4_K_M — as a single tap each, with size, free-space check before enqueuing, and per-entry progress. gpt-oss-120b is listed but not downloadable in-app: Hugging Face ships that quant as two shards, and expert streaming reads tensors by byte offset from one file. The entry carries the PC-side merge recipe instead; once the merged file reaches the device it is recognized like any other. Downloads are keyed by filename and seeded from DownloadManager, so a multi-GB transfer started before the app was killed is picked back up rather than left running unseen. Arbitrary URLs and the file picker stay, under Other model. The catalog is UI convenience only — nothing about these models reaches the engine, which still discovers every architecture at runtime. |
||
|---|---|---|
| .. | ||
| app | ||
| gradle/wrapper | ||
| build.gradle | ||
| gradle.properties | ||
| gradlew | ||
| gradlew.bat | ||
| README.md | ||
| settings.gradle | ||
BigMoeOnEdge — Android example
A minimal chat app that validates the throughput claim on a real phone: pick a pushed
.gguf, type a prompt, and watch the answer stream in while a live panel shows tok/s and
the per-token compute-vs-flash-I/O split and cache hit rate.
It runs the engine as the bmoe-cli binary (shipped as libbmoe-cli.so) via
ProcessBuilder from a foreground service — no JNI. This is the same pattern used by the
research harness and keeps the app a thin driver over the CLI.
Build
-
Cross-compile and stage the engine binaries (needs the Android NDK):
pwsh ../../scripts/build-android.ps1This fills
app/src/main/jniLibs/arm64-v8a/withlibbmoe-cli.soand thelibllama/libggmlshared libraries. -
Build and install the APK. Open this folder in Android Studio, or use the committed Gradle wrapper directly. The app has two distribution flavors (see below); build the one you want:
./gradlew assembleDevDebug adb install app/build/outputs/apk/dev/debug/app-dev-debug.apk
Flavors
Two build flavors differ only in how a model reaches the device:
- dev — sideloaded (this is what CI attaches to releases). Keeps all-files access, so it
can also read a model adb-pushed to shared storage. Application id
….example.dev. - play — Play-Store-compliant. No broad storage permission: models come only through the
in-app downloader or the file picker.
./gradlew assemblePlayDebug.
Getting a model onto the device
Any of these — the picker lists every MoE .gguf it finds (dense models are filtered out by
a gguf-header check):
-
In-app URL download (both flavors). Tap Add → Download and paste a direct gguf URL (e.g. a Hugging Face
…/resolve/main/model.gguflink). It downloads in the background to the app's files dir — no permission needed — and appears in the picker when done. -
In-app file picker (both flavors). Tap Add → Pick file and choose a
.ggufalready on the device; it is imported into the app's files dir. -
adb push (dev flavor only — needs all-files access, which the dev build requests):
adb push Qwen3-30B-A3B-Q4_K_M.gguf /sdcard/Download/ # or the app's own dir, or /data/local/tmp/shardllm for a model too big to duplicate
Expected numbers
On a phone with UFS 4.x storage and ~12 GB RAM, streaming Qwen3-30B-A3B-Q4_K_M with the
expert cache at 4000 MiB, 4 I/O lanes and 4 compute threads, decode settles around
0.55–0.6 s/token (~1.8 tok/s) — a model ~1.7× the device RAM, lossless. See
../../docs/benchmark-method.md for the full procedure and the cache/thread sweep.