mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
The serial streaming path stays public-API only and builds against stock upstream. Intra-layer overlap needs one per-expert readiness hook in the CPU mul_mat_id kernel, carried as a 1-commit fork branch (bmoe/expert-ready-hook) on Helldez/llama.cpp. CMake detects the hook in ggml-cpu.h and defines BMOE_HAVE_EXPERT_READY_HOOK; without it --overlap fails with a clear runtime error. Dropped once upstream ships an equivalent hook. |
||
|---|---|---|
| .. | ||
| llama.cpp@5236140e51 | ||