feat(moe): Nemotron 3.5 (nemotron_h_moe) and Ornith 1.5 support (#203)

nemotron_h_moe is the third expert layout: gate-less. Each expert is
up, ReLU^2, down, so the registry row names ffn_up_exps and
ffn_down_exps and leaves the tail slot empty, as the fused gemma4 row
does. The Mamba2/attention blocks, the shared expert, the optional
latent projections and the MTP block all stay on the resident side of
the seam. No llama.cpp change and no submodule bump: the pinned tree
already builds nemotron_h_moe.

make-tiny-moe.py learns the whole shape in miniature (hybrid stack,
latent projections, biased sigmoid router, shared expert, a trailing
MTP block that is never loaded), and it runs as a third byte-identity
gate. Every identity gate passes on it.

The architecture never puts two MoE blocks next to each other, so the
forward predictors (predict-prefetch, route-ahead, the stale half of
predict-log) have no next layer to target. The gate reads that from the
file and reports those checks N/A instead of failing or passing them
vacuously; an unreadable file keeps them strict.

Ornith-1.5-35B-A3B is qwen35moe and needs no engine change. Both models
join the Android catalog at Q4_K_M; neither has device numbers yet.

Also releases 0.25.0: versionCode 40, versionName 0.25.0, dated changelog.
This commit is contained in:
Raffaele 2026-09-28 09:34:42 +02:00 • committed by GitHub
parent 74ba18f53d
commit 50d0a38338
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
13 changed files with 266 additions and 24 deletions

View file

@ -65,7 +65,10 @@ check). Nothing below needs a storage permission except the last option.
1. **Built-in catalog** (both flavors) — the "Get a model" card offers the models this engine
is measured on, each a single tap: **Qwen3-30B-A3B-Q4_K_M** (~18.6 GB, the reference model),
**Qwen3.6-35B-A3B-Q4_K_M** (~22.3 GB, a hybrid attention/SSM MoE, comfortably past device RAM)
and **Gemma-4-26B-A4B-it-Q4_K_M** (~17 GB). Downloads run in a foreground worker, survive the
and **Gemma-4-26B-A4B-it-Q4_K_M** (~17 GB). Two newer entries stream through the same path
but have no device numbers yet: **Ornith-1.5-35B-A3B-Q4_K_M** (~21.9 GB, the Qwen3.5 MoE
architecture) and **Nemotron-3.5-Lightning-30B-A3B-Q4_K_M** (~25.5 GB, a hybrid
Mamba2/attention MoE with gate-less experts). Downloads run in a foreground worker, survive the
app being killed, resume an interrupted transfer instead of restarting, and appear in the
picker when done.
2. **Any other model** — under **Other model**, paste a direct gguf URL (e.g. a Hugging Face

View file

@ -51,8 +51,8 @@ android {
applicationId 'io.bigmoeonedge.example'
minSdk 29
targetSdk 34
versionCode 39
versionName '0.24.0'
versionCode 40
versionName '0.25.0'
buildConfigField 'String', 'GIT_SHA', "\"${gitSha}\""
ndk {
// The engine ships as prebuilt arm64 binaries staged by build-android.ps1.

View file

@ -91,6 +91,24 @@ object ModelCatalog {
"Qwen_Qwen3.6-35B-A3B-Q4_K_M.gguf?download=true",
blurb = "3B active of 35B. Hybrid attention/SSM, comfortably past RAM.",
),
Entry(
title = "Ornith-1.5-35B-A3B",
quant = "Q4_K_M",
fileName = "Ornith-1.5-35B-A3B-Q4_K_M.gguf",
approxBytes = 21_864_081_056L,
url = "https://huggingface.co/bartowski/Ornith-1.5-35B-A3B-GGUF/resolve/main/" +
"Ornith-1.5-35B-A3B-Q4_K_M.gguf?download=true",
blurb = "3B active of 35B. Qwen3.5 MoE tuned for coding and agents.",
),
Entry(
title = "Nemotron-3.5-Lightning-30B-A3B",
quant = "Q4_K_M",
fileName = "NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q4_K_M.gguf",
approxBytes = 25_477_403_616L,
url = "https://huggingface.co/bartowski/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF/resolve/main/" +
"NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q4_K_M.gguf?download=true",
blurb = "3B active of 30B. Hybrid Mamba2/attention, gate-less experts.",
),
Entry(
title = "Gemma-4-26B-A4B-it",
quant = "Q4_K_M",