BigMoeOnEdge/examples/android
Helldez 5db3399897 docs(android): drop the reference to the retired dense pins
The teardown comment listed "unlock the dense pins" among the ordered steps.
That step belonged to --lock-dense, which is retired: RLIMIT_MEMLOCK is 65536
bytes hard on the target, so the flag could pin 0.003% of a 2 GB cache and its
premise does not survive measurement. See docs/android-memory.md.

The teardown is unchanged — only the comment described a step that never ran.
2026-07-15 20:19:21 +02:00
..
app docs(android): drop the reference to the retired dense pins 2026-07-15 20:19:21 +02:00
gradle/wrapper build(android): commit Gradle wrapper (8.9) for reproducible APK builds 2026-07-11 08:37:55 +02:00
build.gradle feat(android): example chat app with live streaming telemetry 2026-07-10 18:18:24 +02:00
gradle.properties feat(android): example chat app with live streaming telemetry 2026-07-10 18:18:24 +02:00
gradlew build(android): commit Gradle wrapper (8.9) for reproducible APK builds 2026-07-11 08:37:55 +02:00
gradlew.bat build(android): commit Gradle wrapper (8.9) for reproducible APK builds 2026-07-11 08:37:55 +02:00
README.md feat(android): in-app model acquisition, no broad storage permission 2026-07-11 17:03:01 +02:00
settings.gradle feat(android): example chat app with live streaming telemetry 2026-07-10 18:18:24 +02:00

BigMoeOnEdge — Android example

A minimal chat app that validates the throughput claim on a real phone: pick a pushed .gguf, type a prompt, and watch the answer stream in while a live panel shows tok/s and the per-token compute-vs-flash-I/O split and cache hit rate.

It runs the engine as the bmoe-cli binary (shipped as libbmoe-cli.so) via ProcessBuilder from a foreground service — no JNI. This is the same pattern used by the research harness and keeps the app a thin driver over the CLI.

Build

  1. Cross-compile and stage the engine binaries (needs the Android NDK):

    pwsh ../../scripts/build-android.ps1
    

    This fills app/src/main/jniLibs/arm64-v8a/ with libbmoe-cli.so and the libllama/libggml shared libraries.

  2. Build and install the APK. Open this folder in Android Studio, or use the committed Gradle wrapper directly. The app has two distribution flavors (see below); build the one you want:

    ./gradlew assembleDevDebug
    adb install app/build/outputs/apk/dev/debug/app-dev-debug.apk
    

Flavors

Two build flavors differ only in how a model reaches the device:

  • dev — sideloaded (this is what CI attaches to releases). Keeps all-files access, so it can also read a model adb-pushed to shared storage. Application id …​.example.dev.
  • play — Play-Store-compliant. No broad storage permission: models come only through the in-app downloader or the file picker. ./gradlew assemblePlayDebug.

Getting a model onto the device

Any of these — the picker lists every MoE .gguf it finds (dense models are filtered out by a gguf-header check):

  1. In-app URL download (both flavors). Tap Add → Download and paste a direct gguf URL (e.g. a Hugging Face …/resolve/main/model.gguf link). It downloads in the background to the app's files dir — no permission needed — and appears in the picker when done.

  2. In-app file picker (both flavors). Tap Add → Pick file and choose a .gguf already on the device; it is imported into the app's files dir.

  3. adb push (dev flavor only — needs all-files access, which the dev build requests):

    adb push Qwen3-30B-A3B-Q4_K_M.gguf /sdcard/Download/
    # or the app's own dir, or /data/local/tmp/shardllm for a model too big to duplicate
    

Expected numbers

On a phone with UFS 4.x storage and ~12 GB RAM, streaming Qwen3-30B-A3B-Q4_K_M with the expert cache at 4000 MiB, 4 I/O lanes and 4 compute threads, decode settles around 0.55–0.6 s/token (~1.8 tok/s) — a model ~1.7× the device RAM, lossless. See ../../docs/benchmark-method.md for the full procedure and the cache/thread sweep.