BigMoeOnEdge/examples
Raffaele 50d0a38338
feat(moe): Nemotron 3.5 (nemotron_h_moe) and Ornith 1.5 support (#203)
nemotron_h_moe is the third expert layout: gate-less. Each expert is
up, ReLU^2, down, so the registry row names ffn_up_exps and
ffn_down_exps and leaves the tail slot empty, as the fused gemma4 row
does. The Mamba2/attention blocks, the shared expert, the optional
latent projections and the MTP block all stay on the resident side of
the seam. No llama.cpp change and no submodule bump: the pinned tree
already builds nemotron_h_moe.

make-tiny-moe.py learns the whole shape in miniature (hybrid stack,
latent projections, biased sigmoid router, shared expert, a trailing
MTP block that is never loaded), and it runs as a third byte-identity
gate. Every identity gate passes on it.

The architecture never puts two MoE blocks next to each other, so the
forward predictors (predict-prefetch, route-ahead, the stale half of
predict-log) have no next layer to target. The gate reads that from the
file and reports those checks N/A instead of failing or passing them
vacuously; an unreadable file keeps them strict.

Ornith-1.5-35B-A3B is qwen35moe and needs no engine change. Both models
join the Android catalog at Q4_K_M; neither has device numbers yet.

Also releases 0.25.0: versionCode 40, versionName 0.25.0, dated changelog.
2026-09-28 09:34:42 +02:00
..
android feat(moe): Nemotron 3.5 (nemotron_h_moe) and Ornith 1.5 support (#203) 2026-09-28 09:34:42 +02:00