BigMoeOnEdge/core
gjjkbssg 0f193d3c05
feat(moe): warn when the expert cache budget sits below one token's cycle (#167)
Below one token's worst-case routed bytes, global LRU evicts precisely
what it is about to read: the hit rate is 0% while the run still pays
the cache's management time and its RAM. The only protection so far was
cache_min_mb, a fixed floor that happens to sit above the cycle for the
shipped models at their default top-k and stops holding the moment
--n-expert-used widens the routing without touching the budget.

The cycle is priced at init from the model's shape alone — every bound
layer's entry_bytes times min(top_k, n_expert), at the top-k the run
actually applies — and recorded as cache_cycle_mb in the metrics
preamble next to the budget it should be judged against, so a committed
CSV answers on its own whether its cache could ever have hit. The engine
prints one stderr line when the resolved budget falls under it. It warns
rather than refuses: the budget is legal and the output byte-identical,
and the engine states the fact and leaves the choice — the argument that
--cache-mb 0 is strictly better below the cliff stays in
docs/cache-sizing.md, not in the engine's output.

Docs updated in the same commit: cache-sizing.md (the guard section,
past tense), roadmap.md (the item moves from still-worth-doing to
shipped) and telemetry.md (the new preamble key).
2026-08-25 12:48:22 +02:00
..
include/bmoe feat(moe): warn when the expert cache budget sits below one token's cycle (#167) 2026-08-25 12:48:22 +02:00
src feat(moe): warn when the expert cache budget sits below one token's cycle (#167) 2026-08-25 12:48:22 +02:00
CMakeLists.txt feat(engine): self-speculative decoding — the model's own MTP head, or n-gram lookup (#134) 2026-08-02 00:09:39 +02:00