BigMoeOnEdge/docs/adding-a-model.md
Raffaele 50d0a38338
feat(moe): Nemotron 3.5 (nemotron_h_moe) and Ornith 1.5 support (#203)
nemotron_h_moe is the third expert layout: gate-less. Each expert is
up, ReLU^2, down, so the registry row names ffn_up_exps and
ffn_down_exps and leaves the tail slot empty, as the fused gemma4 row
does. The Mamba2/attention blocks, the shared expert, the optional
latent projections and the MTP block all stay on the resident side of
the seam. No llama.cpp change and no submodule bump: the pinned tree
already builds nemotron_h_moe.

make-tiny-moe.py learns the whole shape in miniature (hybrid stack,
latent projections, biased sigmoid router, shared expert, a trailing
MTP block that is never loaded), and it runs as a third byte-identity
gate. Every identity gate passes on it.

The architecture never puts two MoE blocks next to each other, so the
forward predictors (predict-prefetch, route-ahead, the stale half of
predict-log) have no next layer to target. The gate reads that from the
file and reports those checks N/A instead of failing or passing them
vacuously; an unreadable file keeps them strict.

Ornith-1.5-35B-A3B is qwen35moe and needs no engine change. Both models
join the Android catalog at Q4_K_M; neither has device numbers yet.

Also releases 0.25.0: versionCode 40, versionName 0.25.0, dated changelog.
2026-09-28 09:34:42 +02:00

2.6 KiB

Adding a MoE architecture

Most MoE models in llama.cpp are built by the same build_moe_ffn helper and expose the identical routing node (ffn_moe_topk) and expert tensors (ffn_{gate,up,down}_exps). For those, adding support is one row.

1. Add a recipe

In core/src/moe/arch_registry.cpp:

static const MoeRecipe k_recipes[] = {
    { "qwen3moe", { "ffn_gate_exps", "ffn_up_exps", "ffn_down_exps" } },
    { "your_arch", { "ffn_gate_exps", "ffn_up_exps", "ffn_down_exps" } },  // <-- new
};

arch is the gguf general.architecture string. exps_suffix lists the layer's expert weight tensors — the names without the blk.<il>. prefix and .weight suffix. The common split layout names three ({gate, up, down}); leave a trailing slot nullptr for layouts with fewer (see the fused case below). The engine discovers the expert count and per-expert stride at runtime, so a recipe is only these names.

2. Run the gates

Generate a tiny model for your architecture and run its gate:

# scripts/make-tiny-moe.py --arch <arch> emits a synthetic model for that layout;
# tests/CMakeLists.txt wires one generate-fixture + gate per arch it knows.
cd build && ctest -R moe_gates --output-on-failure

If G1 (streamed == resident) passes, the architecture streams losslessly. For a layout make-tiny-moe.py does not yet emit, either teach it that arch (preferred — permanent CI coverage) or validate the recipe against a real model of that architecture.

When one row is not enough

Some models pack gate and up into a single tensor (ffn_gate_up_exps) instead of two — gemma4 (Gemma 4 MoE) is one. This is still a single row: put the fused suffix first and leave the tail nullptr, because to the streamer a fused gate_up is just an expert tensor with a larger per-expert stride, discovered at runtime like any other.

    { "gemma4", { "ffn_gate_up_exps", "ffn_down_exps", nullptr } },

Some models have no gate projection at all: each expert is up, an activation, then down. nemotron_h_moe (Nemotron 3 / 3.5 MoE, ReLU²) is one. The row names the two tensors it has and leaves the tail nullptr, exactly like the fused case; the slots carry no meaning to the engine, which streams whatever the row names.

    { "nemotron_h_moe", { "ffn_up_exps", "ffn_down_exps", nullptr } },

Models with shared/always-on experts (a dense expert applied to every token, as in gemma4, DeepSeek and some Qwen variants) work, but the shared expert stays resident and reduces the streaming saving proportionally. Note it in the model's entry when you add one.