BigMoeOnEdge/examples/android
Helldez bea5a0b99e
feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#95)
* feat(moe): --drop-cold-experts, spend quality only where it buys I/O

Turbo top-k drops the tail of a routing whether or not those experts were
already in RAM. A resident expert costs no flash read, so that trade pays
quality for nothing on the ~80% of decode routings that are cache hits.

This skips a routed expert only when it is a cache MISS and the router
weighted it below frac x (1/top-k). Replayed over the committed route
traces at frac 1.0, decode phase, that avoids 66% of flash reads for 9.5%
of the router's weight mass, where --n-expert-used 5 avoids 23% for a
comparable 10.6% -- about 3x the reads at the same quality cost.

Implementation. The decision needs the FINAL router weights, which arrive
several nodes after the topk where the streamer normally loads, so
load_layer() is deferred to the terminal node of the layer's weight chain.
Which node that is depends on the model's gating, so the hook learns it
from the graph rather than carrying an architecture table; if it fails to
arrive the hook forgets it and re-learns rather than re-betting. A dropped
slot has its weight zeroed and its expert id repointed at the routing's
top-weighted expert: an unread expert can sit in reserved-but-uncommitted
VM and mul_mat_id would touch it anyway, so the kernel is given memory
that is certainly resident and multiplies it by exactly zero.

Requires the LRU cache -- with --cache-mb 0 residency reads all-miss and
the policy would silently degenerate into an unconditional weight cut.
Prefill is excluded by default. The top expert is always pinned, so no
routing can be emptied at any threshold.

Gates: G8a/G8a' prove the deferral and the learned terminal node are
transparent (byte-identical output, zero drops, at a threshold below any
producible weight); G8b that full strength against a constantly-evicting
cache never reaches an unloaded slot; G8c that at top-k 1 dropping is a
no-op, pinning both the top-expert guarantee and the threshold tracking
the effective top-k.

Three existing metrics shift meaning under dropping and the docs now say
so: cache_hit_pct rises without the cache serving more (a dropped routing
is a miss that is never looked up), and token/layer_demand measure what
was staged rather than routed. prefetch.md's "cannot change output" is
scoped, limitations.md gains the non-reproducibility entry, and
benchmark-method.md warns that reversing the run order cannot distinguish
a moved drop rate from a contaminated cell.

Off by default in the CLI and in the app. The output is not reproducible
-- what gets dropped depends on what the cache held -- so it carries no
rows in the README tables, and switching it on by default waits on a
published on-device A/B rather than on the replay argument alone.

* feat(app): default cache-aware dropping to 75%, measured on device

Qwen3.6-35B-A3B (top-k 8 of 256), in-app, cache 3000, one variable
changed: 2.549 tok/s off, 3.938 at F=0.75 (+55%), 4.702 at F=1.0 (+84%),
with flash reads falling 248 -> 163 -> 48 GiB. Per-token bootstrap
intervals separate every pair except off vs 0.50, which overlaps -- at
half the uniform share the policy drops 2.7% of routings and buys
nothing, which doubles as a negative control that the machinery is free
when it does not fire.

Run order was 1.0, off, 0.5, 0.75, so the two fastest cells are the first
and the LAST; thermal drift would have made the last the worst. The
mechanism orders by threshold even though the run order does not.

The replay turned out conservative rather than optimistic. It is
documented as an upper bound because it cannot model the cache changing
in response to dropping: at F=0.75 it was accurate (37% predicted, 34%
measured), at F=1.0 it understated (66% predicted, 81% measured). Avoided
reads free cache capacity, which raises the hit rate, which leaves fewer
misses to drop.

75% rather than 100% is deliberate: it takes the larger part of the win
for half the discarded routings (14% against 28%). Quality is still
unquantified -- no perplexity number and no side-by-side exists -- so the
conservative end of a measured range is the defensible default. The CLI
stays off; the byte-identity gates need a deterministic default.

Also records cache_hit_pct rising 67.8 -> 90.7% as the documented
accounting artefact rather than the cache serving more, and majflt/token
as dominated by each run's starting memory state, not by the threshold.
2026-07-22 17:21:55 +02:00
..
app feat(moe): --drop-cold-experts — spend quality only where it buys I/O (#95) 2026-07-22 17:21:55 +02:00
gradle/wrapper build(android): commit Gradle wrapper (8.9) for reproducible APK builds 2026-07-11 08:37:55 +02:00
build.gradle feat(android): example chat app with live streaming telemetry 2026-07-10 18:18:24 +02:00
gradle.properties feat(android): example chat app with live streaming telemetry 2026-07-10 18:18:24 +02:00
gradlew build(android): commit Gradle wrapper (8.9) for reproducible APK builds 2026-07-11 08:37:55 +02:00
gradlew.bat build(android): commit Gradle wrapper (8.9) for reproducible APK builds 2026-07-11 08:37:55 +02:00
README.md docs(android): list Qwen3.6 in the catalog, and separate sweep from default 2026-07-19 11:27:44 +02:00
settings.gradle feat(android): example chat app with live streaming telemetry 2026-07-10 18:18:24 +02:00

BigMoeOnEdge — Android example

A minimal chat app that validates the throughput claim on a real phone: pick a pushed .gguf, type a prompt, and watch the answer stream in while a live panel shows tok/s and the per-token compute-vs-flash-I/O split and cache hit rate.

It runs the engine as the bmoe-cli binary (shipped as libbmoe-cli.so) via ProcessBuilder from a foreground service — no JNI. This is the same pattern used by the research harness and keeps the app a thin driver over the CLI.

Build

  1. Cross-compile and stage the engine binaries (needs the Android NDK):

    pwsh ../../scripts/build-android.ps1
    

    This fills app/src/main/jniLibs/arm64-v8a/ with libbmoe-cli.so and the libllama/libggml shared libraries.

  2. Build and install the APK. Open this folder in Android Studio, or use the committed Gradle wrapper directly. The app has two distribution flavors (see below); build the one you want:

    ./gradlew assembleDevDebug
    adb install app/build/outputs/apk/dev/debug/app-dev-debug.apk
    

    Published sideload builds are release-signed with a stable key instead, so an update installs over the previous one rather than being refused. That needs a keystore.properties next to app/ (gitignored — it points at the keystore and holds its passwords); without it the release build falls back to debug signing.

Flavors

Two build flavors differ only in how a model reaches the device:

  • dev — sideloaded (this is what CI attaches to releases). Keeps all-files access, so it can also read a model adb-pushed to shared storage. Application id …​.example.dev.
  • play — Play-Store-compliant. No broad storage permission: models come only through the in-app downloader or the file picker. ./gradlew assemblePlayDebug.

Getting a model onto the device

The picker lists every MoE .gguf it finds (dense models are filtered out by a gguf-header check). Nothing below needs a storage permission except the last option.

  1. Built-in catalog (both flavors) — the "Get a model" card offers the models this engine is measured on, each a single tap: Qwen3-30B-A3B-Q4_K_M (~18.6 GB, the reference model), Qwen3.6-35B-A3B-Q4_K_M (~22.3 GB, a hybrid attention/SSM MoE, comfortably past device RAM) and Gemma-4-26B-A4B-it-Q4_K_M (~17 GB). Downloads run in a foreground worker, survive the app being killed, resume an interrupted transfer instead of restarting, and appear in the picker when done.

  2. Any other model — under Other model, paste a direct gguf URL (e.g. a Hugging Face …/resolve/main/model.gguf link), or pick a .gguf already on the device to import it.

    In-app downloads and picker imports both land in the app's internal storage (filesDir, a real f2fs/ext4 volume), so the streamed expert reads use O_DIRECT at full speed. Only models read from the emulated external dirs (adb-pushed to /sdcard/Download) fall back to buffered I/O. A download needs free space equal to the model size — no temporary second copy.

  3. adb push (dev flavor only — needs all-files access, which the dev build requests):

    adb push Qwen3-30B-A3B-Q4_K_M.gguf /sdcard/Download/
    # /data/local/tmp/bmoe avoids duplicating a model too big to copy, and is on a real
    # filesystem where O_DIRECT works (the emulated dirs fall back to buffered I/O)
    adb push Qwen3-30B-A3B-Q4_K_M.gguf /data/local/tmp/bmoe/
    

    This directory was named shardllm before v0.8.0. To keep models already pushed there:

    adb shell mv /data/local/tmp/shardllm /data/local/tmp/bmoe
    

gpt-oss-120b

Listed in the catalog but not downloadable in-app: Hugging Face ships the Q4_K_M quant as two shards (the 50 GB per-file limit), and expert streaming reads tensors by byte offset from a single file. Merge the shards on a PC, then transfer the result:

llama-gguf-split --merge gpt-oss-120b-Q4_K_M-00001-of-00002.gguf gpt-oss-120b-Q4_K_M.gguf
adb push gpt-oss-120b-Q4_K_M.gguf /data/local/tmp/bmoe/   # or import it with the file picker

Expected numbers

On a phone with UFS 4.x storage and ~12 GB RAM, streaming Qwen3-30B-A3B-Q4_K_M with the expert cache at 4000 MiB, 4 I/O lanes and 4 compute threads, decode settles around 0.55–0.6 s/token (~1.8 tok/s) — a model ~1.7× the device RAM, lossless. That 4000 MiB is a sweep point from the benchmark protocol, not the app default: the app ships a fixed 2000 MiB expert cache. See ../../docs/benchmark-method.md for the full procedure and the cache/thread sweep.