ruvector

mirror of https://github.com/ruvnet/RuVector.git synced 2026-05-22 19:56:25 +00:00

History

rUv 04b26c8d69 feat: Add PowerInfer-style sparse inference engine with precision lanes (#106 ) ## Summary - Add PowerInfer-style sparse inference engine with precision lanes - Add memory module with QuantizedWeights and NeuronCache - Fix compilation and test issues - Demonstrated 2.9-8.7x speedup at typical sparsity levels - Published to crates.io as ruvector-sparse-inference v0.1.30 ## Key Features - Low-rank predictor using P·Q matrix factorization for fast neuron selection - Sparse FFN kernels that only compute active neurons - SIMD optimization for AVX2, SSE4.1, NEON, and WASM SIMD - GGUF parser with full quantization support (Q4_0 through Q6_K) - Precision lanes (3/5/7-bit layered quantization) - π integration for low-precision systems 🤖 Generated with [Claude Code](https://claude.com/claude-code)	2026-01-04 23:40:31 -05:00
..
ARCHITECTURE.md	feat: Add PowerInfer-style sparse inference engine with precision lanes (#106 )	2026-01-04 23:40:31 -05:00

rUv 04b26c8d69 feat: Add PowerInfer-style sparse inference engine with precision lanes (#106 )

## Summary
- Add PowerInfer-style sparse inference engine with precision lanes
- Add memory module with QuantizedWeights and NeuronCache
- Fix compilation and test issues
- Demonstrated 2.9-8.7x speedup at typical sparsity levels
- Published to crates.io as ruvector-sparse-inference v0.1.30

## Key Features
- Low-rank predictor using P·Q matrix factorization for fast neuron selection
- Sparse FFN kernels that only compute active neurons
- SIMD optimization for AVX2, SSE4.1, NEON, and WASM SIMD
- GGUF parser with full quantization support (Q4_0 through Q6_K)
- Precision lanes (3/5/7-bit layered quantization)
- π integration for low-precision systems

🤖 Generated with [Claude Code](https://claude.com/claude-code)

2026-01-04 23:40:31 -05:00

ARCHITECTURE.md

feat: Add PowerInfer-style sparse inference engine with precision lanes (#106 )

2026-01-04 23:40:31 -05:00