BigMoeOnEdge/examples/android
Helldez 2ededf1d7a feat(android): built-in model catalog with one-tap downloads
Getting a first model meant knowing which MoE to look for and pasting a raw
Hugging Face link. The Get-a-model card now offers the models this engine is
measured on — Qwen3-30B-A3B-Q4_K_M and Gemma-4-26B-A4B-it-Q4_K_M — as a single
tap each, with size, free-space check before enqueuing, and per-entry progress.

gpt-oss-120b is listed but not downloadable in-app: Hugging Face ships that
quant as two shards, and expert streaming reads tensors by byte offset from one
file. The entry carries the PC-side merge recipe instead; once the merged file
reaches the device it is recognized like any other.

Downloads are keyed by filename and seeded from DownloadManager, so a multi-GB
transfer started before the app was killed is picked back up rather than left
running unseen. Arbitrary URLs and the file picker stay, under Other model.

The catalog is UI convenience only — nothing about these models reaches the
engine, which still discovers every architecture at runtime.
2026-07-17 12:43:23 +02:00
..
app feat(android): built-in model catalog with one-tap downloads 2026-07-17 12:43:23 +02:00
gradle/wrapper build(android): commit Gradle wrapper (8.9) for reproducible APK builds 2026-07-11 08:37:55 +02:00
build.gradle feat(android): example chat app with live streaming telemetry 2026-07-10 18:18:24 +02:00
gradle.properties feat(android): example chat app with live streaming telemetry 2026-07-10 18:18:24 +02:00
gradlew build(android): commit Gradle wrapper (8.9) for reproducible APK builds 2026-07-11 08:37:55 +02:00
gradlew.bat build(android): commit Gradle wrapper (8.9) for reproducible APK builds 2026-07-11 08:37:55 +02:00
README.md feat(android): in-app model acquisition, no broad storage permission 2026-07-11 17:03:01 +02:00
settings.gradle feat(android): example chat app with live streaming telemetry 2026-07-10 18:18:24 +02:00

BigMoeOnEdge — Android example

A minimal chat app that validates the throughput claim on a real phone: pick a pushed .gguf, type a prompt, and watch the answer stream in while a live panel shows tok/s and the per-token compute-vs-flash-I/O split and cache hit rate.

It runs the engine as the bmoe-cli binary (shipped as libbmoe-cli.so) via ProcessBuilder from a foreground service — no JNI. This is the same pattern used by the research harness and keeps the app a thin driver over the CLI.

Build

  1. Cross-compile and stage the engine binaries (needs the Android NDK):

    pwsh ../../scripts/build-android.ps1
    

    This fills app/src/main/jniLibs/arm64-v8a/ with libbmoe-cli.so and the libllama/libggml shared libraries.

  2. Build and install the APK. Open this folder in Android Studio, or use the committed Gradle wrapper directly. The app has two distribution flavors (see below); build the one you want:

    ./gradlew assembleDevDebug
    adb install app/build/outputs/apk/dev/debug/app-dev-debug.apk
    

Flavors

Two build flavors differ only in how a model reaches the device:

  • dev — sideloaded (this is what CI attaches to releases). Keeps all-files access, so it can also read a model adb-pushed to shared storage. Application id …​.example.dev.
  • play — Play-Store-compliant. No broad storage permission: models come only through the in-app downloader or the file picker. ./gradlew assemblePlayDebug.

Getting a model onto the device

Any of these — the picker lists every MoE .gguf it finds (dense models are filtered out by a gguf-header check):

  1. In-app URL download (both flavors). Tap Add → Download and paste a direct gguf URL (e.g. a Hugging Face …/resolve/main/model.gguf link). It downloads in the background to the app's files dir — no permission needed — and appears in the picker when done.

  2. In-app file picker (both flavors). Tap Add → Pick file and choose a .gguf already on the device; it is imported into the app's files dir.

  3. adb push (dev flavor only — needs all-files access, which the dev build requests):

    adb push Qwen3-30B-A3B-Q4_K_M.gguf /sdcard/Download/
    # or the app's own dir, or /data/local/tmp/shardllm for a model too big to duplicate
    

Expected numbers

On a phone with UFS 4.x storage and ~12 GB RAM, streaming Qwen3-30B-A3B-Q4_K_M with the expert cache at 4000 MiB, 4 I/O lanes and 4 compute threads, decode settles around 0.55–0.6 s/token (~1.8 tok/s) — a model ~1.7× the device RAM, lossless. See ../../docs/benchmark-method.md for the full procedure and the cache/thread sweep.