mirror of
https://github.com/Helldez/BigMoeOnEdge.git
synced 2026-10-03 11:35:50 +00:00
* feat(tools): add bmoe-membench, a read-bandwidth probe for pinned memory The dense weights are the one part of the model that must stay resident, and no lever documented in docs/android-memory.md can keep them there: RLIMIT_MEMLOCK is capped at 64 KiB by the vendor, the cgroup knobs are v2-only, and MADV_COLD only redirects reclaim rather than preventing it. One allocation an unprivileged app can make is exempt by construction — a dma-buf, whose pages stay pinned because a device may DMA from them at any time. Userspace reaches it through AHardwareBuffer. Whether that is usable hinges on a property nobody publishes: gralloc decides per allocation whether a buffer is CPU-cacheable, and uncached memory loses the cache line, the prefetcher and most memory-level parallelism. A dense matmul streams weights, so it would pay that in full. On a device whose flash serves 1.3-2.5 GB/s, uncached pinned weights would read SLOWER than refaulting them off storage — so the idea is gated on a bandwidth measurement, not on a design argument. bmoe-membench measures exactly that: the same sequential read kernel over an anonymous mapping (what --dense-weights anon already allocates) and over a locked AHardwareBuffer BLOB, reporting the ratio. The two outcomes are an order of magnitude apart, so the tool is built to be unambiguous rather than precise. --probe-max additionally reports the largest BLOB the driver will hand over, since no limit is documented and the dense working set is multiple GiB. It links nothing beyond the stdlib, so it cross-compiles in seconds and also builds on the host, where the anon row still runs. No engine, CLI or app code is touched. * test(memory): measure reclaim-exempt memory — bandwidth gate passes, size capped at 2047 MiB A locked AHardwareBuffer BLOB reads at exactly anonymous-memory speed: within 0.5% on both CPU clusters single-threaded (30.4k and 45.2k MiB/s) and at 4 threads (59.0k), and the CPU_READ_OFTEN hint makes no difference to it. The way this idea could have died on arrival was an uncached mapping — gralloc chooses cacheability per allocation, and uncached memory would read at or below the flash bandwidth it is meant to save, making pinned dense weights slower than refaulting them. It does not. The gate passes. The binding constraint turned out to be size, in a place the documentation does not point at. Allocation succeeds up to the 4 GiB format cap implied by the 32-bit BLOB width, but AHardwareBuffer_lock — which is what yields a usable CPU pointer — refuses at exactly 2048 MiB with EINVAL while 2047 MiB succeeds. A ceiling landing precisely on 2^31 is a signed 32-bit type in the lock path, not memory exhaustion. The usable unit is therefore 2047 MiB and anything larger must span several buffers, which this engine can do since dense_weights already allocates per tensor. Two corrections to the tool, both cases of it producing a confident wrong number: - --probe-max probed allocation only, and so reported roughly double the usable size. It now locks every candidate, and refines to 1 MiB so a ceiling sitting on a power of two is visible as evidence about its cause. - --modes ran anon first whatever order was asked for, and a single pass turned out to rank CPU clusters rather than allocators: the first, un-pinned run reported a clean 0.67x that inverted when the modes were swapped, because the scheduler's choice of core moves this number 1.5x. Modes now run in the order given, --repeat interleaves them, and the header says to pin with taskset. What this does not show is that pinning helps. Reclaim-exempt memory does not create memory: under a >RAM model the RAM the dense weights stop yielding has to come from the expert cache or from the page cache feeding the stream, which is the same trade that refuted the bulk restore and the per-layer LFU cap. Nor did any run here put the device under the pressure a >RAM decode creates, so reclaim-exemption itself remains an inference from how dma-buf works. Recorded as an open, unproven lever. docs/android-memory.md gains the dma-buf row its lever table was missing, the roadmap records the open question, and the raw runs are committed alongside the analysis. |
||
|---|---|---|
| .. | ||
| CMakeLists.txt | ||
| iobench.cpp | ||
| membench.cpp | ||