nemotron_h_moe is the third expert layout: gate-less. Each expert is up, ReLU^2, down, so the registry row names ffn_up_exps and ffn_down_exps and leaves the tail slot empty, as the fused gemma4 row does. The Mamba2/attention blocks, the shared expert, the optional latent projections and the MTP block all stay on the resident side of the seam. No llama.cpp change and no submodule bump: the pinned tree already builds nemotron_h_moe. make-tiny-moe.py learns the whole shape in miniature (hybrid stack, latent projections, biased sigmoid router, shared expert, a trailing MTP block that is never loaded), and it runs as a third byte-identity gate. Every identity gate passes on it. The architecture never puts two MoE blocks next to each other, so the forward predictors (predict-prefetch, route-ahead, the stale half of predict-log) have no next layer to target. The gate reads that from the file and reports those checks N/A instead of failing or passing them vacuously; an unreadable file keeps them strict. Ornith-1.5-35B-A3B is qwen35moe and needs no engine change. Both models join the Android catalog at Q4_K_M; neither has device numbers yet. Also releases 0.25.0: versionCode 40, versionName 0.25.0, dated changelog.
2.6 KiB
Adding a MoE architecture
Most MoE models in llama.cpp are built by the same build_moe_ffn helper and expose the
identical routing node (ffn_moe_topk) and expert tensors
(ffn_{gate,up,down}_exps). For those, adding support is one row.
1. Add a recipe
In core/src/moe/arch_registry.cpp:
static const MoeRecipe k_recipes[] = {
{ "qwen3moe", { "ffn_gate_exps", "ffn_up_exps", "ffn_down_exps" } },
{ "your_arch", { "ffn_gate_exps", "ffn_up_exps", "ffn_down_exps" } }, // <-- new
};
arch is the gguf general.architecture string. exps_suffix lists the layer's expert
weight tensors — the names without the blk.<il>. prefix and .weight suffix. The common
split layout names three ({gate, up, down}); leave a trailing slot nullptr for layouts
with fewer (see the fused case below). The engine discovers the expert count and per-expert
stride at runtime, so a recipe is only these names.
2. Run the gates
Generate a tiny model for your architecture and run its gate:
# scripts/make-tiny-moe.py --arch <arch> emits a synthetic model for that layout;
# tests/CMakeLists.txt wires one generate-fixture + gate per arch it knows.
cd build && ctest -R moe_gates --output-on-failure
If G1 (streamed == resident) passes, the architecture streams losslessly. For a layout
make-tiny-moe.py does not yet emit, either teach it that arch (preferred — permanent CI
coverage) or validate the recipe against a real model of that architecture.
When one row is not enough
Some models pack gate and up into a single tensor (ffn_gate_up_exps) instead of two —
gemma4 (Gemma 4 MoE) is one. This is still a single row: put the fused suffix first and
leave the tail nullptr, because to the streamer a fused gate_up is just an expert tensor
with a larger per-expert stride, discovered at runtime like any other.
{ "gemma4", { "ffn_gate_up_exps", "ffn_down_exps", nullptr } },
Some models have no gate projection at all: each expert is up, an activation, then down.
nemotron_h_moe (Nemotron 3 / 3.5 MoE, ReLU²) is one. The row names the two tensors it has and
leaves the tail nullptr, exactly like the fused case; the slots carry no meaning to the engine,
which streams whatever the row names.
{ "nemotron_h_moe", { "ffn_up_exps", "ffn_down_exps", nullptr } },
Models with shared/always-on experts (a dense expert applied to every token, as in
gemma4, DeepSeek and some Qwen variants) work, but the shared expert stays resident and
reduces the streaming saving proportionally. Note it in the model's entry when you add one.