mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 03:25:42 +00:00
docs: point the layer-lfu record at its tag, not a branch that is being deleted
The branch carrying the experiment is going away; a tag preserves the implementation without leaving a stale branch behind, and keeps the verdict note's reference resolvable.
This commit is contained in:
parent
c8ca0b01d2
commit
08b001d2ba
2 changed files with 8 additions and 8 deletions
|
|
@ -4,10 +4,10 @@ A per-layer budget partition with least-frequently-used eviction inside a layer,
|
||||||
offline replay put it 3.5-5 points of hit rate above global LRU at shipped budgets on top-6
|
offline replay put it 3.5-5 points of hit rate above global LRU at shipped budgets on top-6
|
||||||
models. On device it delivers that hit rate and is **~30 % slower**.
|
models. On device it delivers that hit rate and is **~30 % slower**.
|
||||||
|
|
||||||
> **The implementation is not in `main`.** It lives on the branch
|
> **The implementation is not in `main`.** It is preserved under the tag
|
||||||
> [`feat/expert-cache-layer-lfu`](https://github.com/Helldez/BigMoeOnEdge/tree/feat/expert-cache-layer-lfu)
|
> [`experiment/layer-lfu`](https://github.com/Helldez/BigMoeOnEdge/tree/experiment/layer-lfu)
|
||||||
> as a `--cache-policy lru|layer-lfu` flag (default `lru`), byte-identity gated, and was left
|
> as a `--cache-policy lru|layer-lfu` flag (default `lru`), byte-identity gated (G8, G4d), and was
|
||||||
> there rather than merged: a measured regression does not belong in the engine, but the numbers
|
> deliberately not merged: a measured regression does not belong in the engine, but the numbers
|
||||||
> below are worth keeping. Read this as a closed experiment, not as an available option.
|
> below are worth keeping. Read this as a closed experiment, not as an available option.
|
||||||
|
|
||||||
## The A/B
|
## The A/B
|
||||||
|
|
@ -58,7 +58,7 @@ major fault. The cache buys 2 points of expert hit rate and hands back the dense
|
||||||
## Conclusion
|
## Conclusion
|
||||||
|
|
||||||
**Do not ship layer-lfu as a throughput feature.** The engine keeps global LRU, which is what
|
**Do not ship layer-lfu as a throughput feature.** The engine keeps global LRU, which is what
|
||||||
every published measurement was taken against; the implementation stays on its branch.
|
every published measurement was taken against; the implementation is preserved under that tag.
|
||||||
|
|
||||||
The lesson generalises: the replay models *which* entries a policy keeps, and it models that
|
The lesson generalises: the replay models *which* entries a policy keeps, and it models that
|
||||||
correctly, but it cannot model *what keeping them costs*. A hit rate is not a proxy for
|
correctly, but it cannot model *what keeping them costs*. A hit rate is not a proxy for
|
||||||
|
|
|
||||||
|
|
@ -60,11 +60,11 @@ Whether a smarter eviction policy could raise throughput is **answered, and the
|
||||||
Offline replay of the route traces through Bélády, LRU, LFU, random and a per-layer partition
|
Offline replay of the route traces through Bélády, LRU, LFU, random and a per-layer partition
|
||||||
(`scripts/route-replay.py`, validated against the recorded hit rates to the decimal) shows the
|
(`scripts/route-replay.py`, validated against the recorded hit rates to the decimal) shows the
|
||||||
offline optimum is 11-23 points above LRU, but **no online policy recovers more than ~5**. The
|
offline optimum is 11-23 points above LRU, but **no online policy recovers more than ~5**. The
|
||||||
best candidate, a per-layer budget partition with frequency eviction, was implemented on the
|
best candidate, a per-layer budget partition with frequency eviction, was implemented under the
|
||||||
branch `feat/expert-cache-layer-lfu` and measured: it delivers the predicted hit-rate gain
|
tag `experiment/layer-lfu` and measured: it delivers the predicted hit-rate gain
|
||||||
(+2.0 points, −7 % flash reads) and is **~30 % slower**, because a hard per-layer cap removes the
|
(+2.0 points, −7 % flash reads) and is **~30 % slower**, because a hard per-layer cap removes the
|
||||||
cache's ability to self-balance and the resulting `MADV_DONTNEED` churn gets paid for by the
|
cache's ability to self-balance and the resulting `MADV_DONTNEED` churn gets paid for by the
|
||||||
kernel reclaiming the dense weights (majflt/token 6 → 2370). It was left on the branch rather
|
kernel reclaiming the dense weights (majflt/token 6 → 2370). It is preserved under that tag rather
|
||||||
than merged — a measured regression does not belong in the engine.
|
than merged — a measured regression does not belong in the engine.
|
||||||
|
|
||||||
The transferable lesson: a hit-rate curve is not a throughput argument. Any future policy has to
|
The transferable lesson: a hit-rate curve is not a throughput argument. Any future policy has to
|
||||||
|
|
|
||||||
Loading…
Add table
Add a link
Reference in a new issue