mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
nemotron_h_moe is the third expert layout: gate-less. Each expert is up, ReLU^2, down, so the registry row names ffn_up_exps and ffn_down_exps and leaves the tail slot empty, as the fused gemma4 row does. The Mamba2/attention blocks, the shared expert, the optional latent projections and the MTP block all stay on the resident side of the seam. No llama.cpp change and no submodule bump: the pinned tree already builds nemotron_h_moe. make-tiny-moe.py learns the whole shape in miniature (hybrid stack, latent projections, biased sigmoid router, shared expert, a trailing MTP block that is never loaded), and it runs as a third byte-identity gate. Every identity gate passes on it. The architecture never puts two MoE blocks next to each other, so the forward predictors (predict-prefetch, route-ahead, the stale half of predict-log) have no next layer to target. The gate reads that from the file and reports those checks N/A instead of failing or passing them vacuously; an unreadable file keeps them strict. Ornith-1.5-35B-A3B is qwen35moe and needs no engine change. Both models join the Android catalog at Q4_K_M; neither has device numbers yet. Also releases 0.25.0: versionCode 40, versionName 0.25.0, dated changelog. |
||
|---|---|---|
| .. | ||
| include/bmoe | ||
| src | ||
| CMakeLists.txt | ||