mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 11:35:50 +00:00
Experimental, off by default. Before a decode routing is committed, every expert already in the LRU cache gets its score raised by L times the token's score range and the top-k is taken again, so a near-tie goes to the expert already in RAM (Skliar et al., arXiv:2412.00099). The same number of experts runs; fewer are read from flash. Scores are read from the tensor the graph itself sorted, exact for any gating function. Desktop, Qwen3.6-35B Q4_K_M at L=0.15: 258 to 119 MiB of flash per token, 2.37 to 3.84 tok/s, perplexity +1 to 4 %, tinyMMLU 88 to 84/100, HumanEval-50 42 = 42. The on-device A/B is still owed, hence experimental. Also: --ppl / --ppl-step / --ppl-list / --ppl-choices (teacher-forced perplexity, one token per decode so cache-dependent policies are priced where they act), scripts/tinymmlu-bench.py, scripts/humaneval-bench.py, gates G8d/G8e, app switch "Prefer cached experts" under Experimental, docs/cache-aware-substitution.md. |
||
|---|---|---|
| .. | ||
| bench-analyze.py | ||
| bench-lib.ps1 | ||
| bench-matrix-rework.ps1 | ||
| bench-matrix.ps1 | ||
| bench-pr23-c2000.ps1 | ||
| bench-pr23-summary.py | ||
| bench-prefetch.ps1 | ||
| bench-report.sh | ||
| bench-run.sh | ||
| bench-warmonly.sh | ||
| build-android.ps1 | ||
| build-host.sh | ||
| decode-analyze.py | ||
| gptoss-matrix.sh | ||
| gptoss-mmap.sh | ||
| humaneval-bench.py | ||
| inspect-hf-moe-release.py | ||
| make-tiny-moe.py | ||
| route-analyze.py | ||
| route-drop-replay.py | ||
| route-replay.py | ||
| route-viewer.py | ||
| tinymmlu-bench.py | ||
| trace_io.py | ||