mirror of
https://github.com/ruvnet/RuVector.git
synced 2026-08-07 07:43:52 +00:00
feat(acrp): adaptive compression & retrieval plane — unified codec interface, search policies, adaptive escalation (ADR-297)
ADR-297 lays out the plane: one EncodedVector interface for every
representation, storage/search precision separation, per-query automatic
precision selection, memory tiers, topology-aware bit allocation, drift
detection, provenance, honest benchmarking, and a three-policy product
surface — gated by an ablation acceptance test (adaptive must beat uniform
Turbo4 by >= 30% memory at <= 0.5pp recall loss and <= 10% P95).
Phase B lands here:
- ruvector-core::encoding — CodecKind + VectorCodec/EncodedQuery traits;
Fp32, Fp16 (in-crate RNE binary16, no half dep), Int8, and Turbo4 plane
codecs; codec_for() registry; VectorProvenance schema (model id, codec
version, rotation seed, source hash, migration lineage).
ruvector-turboquant is now an unconditional core dependency (dep-free,
WASM-safe) so the codec plane exists on every build.
- SearchPolicy { Quality, Balanced, MaxCompression } on
QuantizationConfig::Turbo4 — users pick an outcome, not an algorithm.
- Turbo4HnswIndex adaptive escalation: relative kept/dropped score margin
triggers widened re-search (2-3x ef, 2x rescore pool), stopping when
top-k membership stabilizes; MaxCompression never escalates; telemetry
via adaptive_stats() targets a 5-15% escalation budget.
- clippy --all-targets clean for both crates (fixes CI identity_op).
Co-Authored-By: claude-flow <ruv@ruv.net>
Claude-Session: https://claude.ai/code/session_01XFWB9PKwsZYk5FbjBRY6mk
This commit is contained in:
parent
2315700e17
commit
4cdc1be09e
10 changed files with 1213 additions and 37 deletions
|
|
@ -14,7 +14,9 @@ description = "High-performance Rust vector database core with HNSW indexing"
|
|||
redb = { workspace = true, optional = true }
|
||||
memmap2 = { workspace = true, optional = true }
|
||||
hnsw_rs = { workspace = true, optional = true }
|
||||
ruvector-turboquant = { version = "2.3.0", path = "../ruvector-turboquant", optional = true }
|
||||
# Turbo4 codec + unified encoding plane (ADR-296/297). Dependency-free and
|
||||
# WASM-safe, so it is unconditional: the codec plane exists on every build.
|
||||
ruvector-turboquant = { version = "2.3.0", path = "../ruvector-turboquant" }
|
||||
simsimd = { workspace = true, optional = true }
|
||||
rayon = { workspace = true, optional = true }
|
||||
crossbeam = { workspace = true, optional = true }
|
||||
|
|
@ -121,7 +123,7 @@ simd = ["simsimd"] # SIMD acceleration (not available in WASM)
|
|||
simd-avx512 = []
|
||||
parallel = ["rayon", "crossbeam"] # Parallel processing (not available in WASM)
|
||||
storage = ["redb", "memmap2"] # File-based storage (not available in WASM)
|
||||
hnsw = ["hnsw_rs", "dep:ruvector-turboquant"] # HNSW indexing (not available in WASM due to mmap dependency)
|
||||
hnsw = ["hnsw_rs"] # HNSW indexing (not available in WASM due to mmap dependency)
|
||||
memory-only = [] # Pure in-memory storage for WASM
|
||||
uuid-support = [] # Deprecated: uuid is now always included
|
||||
real-embeddings = [] # Feature flag for embedding provider API (use ApiEmbedding for production)
|
||||
|
|
|
|||
198
crates/ruvector-core/src/encoding.rs
Normal file
198
crates/ruvector-core/src/encoding.rs
Normal file
|
|
@ -0,0 +1,198 @@
|
|||
//! Unified vector-representation plane (ADR-297 §1).
|
||||
//!
|
||||
//! One interface — [`VectorCodec`] / [`EncodedQuery`] — for every way
|
||||
//! RuVector can hold a vector: FP32, FP16, Int8, Turbo4 (and, from their own
|
||||
//! crates, PQ and RaBitQ). Storage, indexes, snapshots, WASM, and bindings
|
||||
//! consume `&dyn VectorCodec` and opaque blobs; they never name a concrete
|
||||
//! codec. This is what lets precision become a per-vector, per-query
|
||||
//! *decision* instead of a compile-time choice.
|
||||
//!
|
||||
//! Also home to [`VectorProvenance`] (ADR-297 §7): the metadata every stored
|
||||
//! vector must carry so that migrations are auditable and snapshots stay
|
||||
//! deterministic and independently verifiable.
|
||||
|
||||
pub mod codecs;
|
||||
|
||||
use crate::error::{Result, RuvectorError};
|
||||
use crate::types::DistanceMetric;
|
||||
use serde::{Deserialize, Serialize};
|
||||
|
||||
/// Every representation the plane knows about. `Pq` and `RaBitQ1` are
|
||||
/// reserved here (stable blob-header / provenance values) but their codecs
|
||||
/// live in `ruvector-pq-search` / `ruvector-rabitq` (ADR-297 phase C).
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize)]
|
||||
pub enum CodecKind {
|
||||
/// 4 bytes/dim, exact.
|
||||
Fp32,
|
||||
/// 2 bytes/dim, ~1e-3 relative error. Hot-tier / verification codec.
|
||||
Fp16,
|
||||
/// 1 byte/dim + 8 B header (per-vector min/scale uniform quantization).
|
||||
Int8,
|
||||
/// 0.5 bytes/dim + 8 B constants (ADR-296 Lloyd-Max rotated codes).
|
||||
Turbo4,
|
||||
/// Product quantization (external crate; reserved).
|
||||
Pq,
|
||||
/// 1-bit RaBitQ (external crate; reserved).
|
||||
RaBitQ1,
|
||||
}
|
||||
|
||||
impl CodecKind {
|
||||
/// Bytes per encoded vector at dimension `dim` (None for reserved kinds
|
||||
/// whose layout is owned by external crates).
|
||||
pub fn encoded_len(&self, dim: usize) -> Option<usize> {
|
||||
match self {
|
||||
CodecKind::Fp32 => Some(dim * 4),
|
||||
CodecKind::Fp16 => Some(dim * 2),
|
||||
CodecKind::Int8 => Some(dim + 8),
|
||||
CodecKind::Turbo4 => Some(dim / 2 + 8),
|
||||
CodecKind::Pq | CodecKind::RaBitQ1 => None,
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// A vector codec: encodes f32 vectors into opaque blobs and scores blobs
|
||||
/// without the caller knowing the representation.
|
||||
pub trait VectorCodec: Send + Sync {
|
||||
/// Which representation this codec produces.
|
||||
fn kind(&self) -> CodecKind;
|
||||
/// Dimensionality of vectors this codec accepts.
|
||||
fn dim(&self) -> usize;
|
||||
/// Bytes per encoded vector.
|
||||
fn encoded_len(&self) -> usize;
|
||||
/// Encode a vector. The blob is self-sufficient for scoring.
|
||||
fn encode(&self, v: &[f32]) -> Result<Vec<u8>>;
|
||||
/// Reconstruct an approximation (exact only for `Fp32`).
|
||||
fn decode(&self, blob: &[u8]) -> Result<Vec<f32>>;
|
||||
/// Symmetric blob×blob distance under this codec's metric.
|
||||
fn distance(&self, a: &[u8], b: &[u8]) -> Result<f32>;
|
||||
/// Prepare a query for repeated asymmetric scoring against blobs.
|
||||
fn make_query(&self, q: &[f32]) -> Result<Box<dyn EncodedQuery>>;
|
||||
}
|
||||
|
||||
/// A prepared query bound to one codec + metric.
|
||||
pub trait EncodedQuery: Send + Sync {
|
||||
/// Distance from the query to an encoded vector. Conventions match the
|
||||
/// index plane: Euclidean root, `1 − cos` (clamped ≥ 0), `−dot`
|
||||
/// (clamped ≥ 0), Manhattan sum.
|
||||
fn distance_to(&self, blob: &[u8]) -> f32;
|
||||
}
|
||||
|
||||
/// Construct a codec by kind. The single entry point bindings and storage
|
||||
/// use, so adding a codec never touches consumers.
|
||||
pub fn codec_for(
|
||||
kind: CodecKind,
|
||||
dim: usize,
|
||||
metric: DistanceMetric,
|
||||
rotation_seed: u64,
|
||||
) -> Result<Box<dyn VectorCodec>> {
|
||||
match kind {
|
||||
CodecKind::Fp32 => Ok(Box::new(codecs::Fp32Codec::new(dim, metric))),
|
||||
CodecKind::Fp16 => Ok(Box::new(codecs::Fp16Codec::new(dim, metric))),
|
||||
CodecKind::Int8 => Ok(Box::new(codecs::Int8Codec::new(dim, metric))),
|
||||
CodecKind::Turbo4 => Ok(Box::new(codecs::Turbo4PlaneCodec::new(
|
||||
dim,
|
||||
metric,
|
||||
rotation_seed,
|
||||
)?)),
|
||||
CodecKind::Pq | CodecKind::RaBitQ1 => Err(RuvectorError::InvalidParameter(format!(
|
||||
"{kind:?} codec is provided by its own crate and not yet wired \
|
||||
into the core plane (ADR-297 phase C)"
|
||||
))),
|
||||
}
|
||||
}
|
||||
|
||||
/// Provenance every stored vector carries (ADR-297 §7). Snapshots must be
|
||||
/// reproducible from (source, provenance) alone — which is why rotation
|
||||
/// seeds are recorded and codecs are versioned.
|
||||
#[derive(Debug, Clone, PartialEq, Serialize, Deserialize)]
|
||||
pub struct VectorProvenance {
|
||||
/// Embedding model identity (e.g. "text-embedding-3-large").
|
||||
pub model_id: Option<String>,
|
||||
/// Representation of the stored blob.
|
||||
pub codec: CodecKind,
|
||||
/// Version of that codec's layout/tables. Bump on any change that alters
|
||||
/// encoded bytes for identical input.
|
||||
pub codec_version: u16,
|
||||
/// Rotation seed for rotated codecs (Turbo4, RaBitQ).
|
||||
pub rotation_seed: Option<u64>,
|
||||
/// Dimensionality.
|
||||
pub dim: usize,
|
||||
/// Distance metric the codes were prepared for.
|
||||
pub metric: DistanceMetric,
|
||||
/// Hex SHA-256 of the source object, when known.
|
||||
pub source_hash: Option<String>,
|
||||
/// Migration lineage, oldest first (e.g. "fp32→turbo4@2026-08-06").
|
||||
pub lineage: Vec<String>,
|
||||
}
|
||||
|
||||
/// Current layout version of the in-core codecs.
|
||||
pub const CODEC_VERSION: u16 = 1;
|
||||
|
||||
impl VectorProvenance {
|
||||
/// Provenance for a freshly encoded vector under `codec`.
|
||||
pub fn new(codec: &dyn VectorCodec, metric: DistanceMetric) -> Self {
|
||||
Self {
|
||||
model_id: None,
|
||||
codec: codec.kind(),
|
||||
codec_version: CODEC_VERSION,
|
||||
rotation_seed: None,
|
||||
dim: codec.dim(),
|
||||
metric,
|
||||
source_hash: None,
|
||||
lineage: Vec::new(),
|
||||
}
|
||||
}
|
||||
|
||||
/// Record a migration step (old codec → new codec).
|
||||
pub fn push_migration(&mut self, entry: impl Into<String>) {
|
||||
self.lineage.push(entry.into());
|
||||
}
|
||||
}
|
||||
|
||||
/// Shared scalar distance with the same conventions as the HNSW plane
|
||||
/// (`index::hnsw::DistanceFn`): callers across codecs must agree on score
|
||||
/// semantics or adaptive escalation would compare incomparable numbers.
|
||||
pub(crate) fn metric_distance(metric: DistanceMetric, a: &[f32], b: &[f32]) -> f32 {
|
||||
use crate::simd_intrinsics as si;
|
||||
match metric {
|
||||
DistanceMetric::Euclidean => si::euclidean_distance_simd(a, b),
|
||||
DistanceMetric::Cosine => (1.0_f32 - si::cosine_similarity_simd(a, b)).max(0.0),
|
||||
DistanceMetric::DotProduct => (-si::dot_product_simd(a, b)).max(0.0),
|
||||
DistanceMetric::Manhattan => si::manhattan_distance_simd(a, b),
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn encoded_len_matches_kind_table() {
|
||||
for (kind, dim, expect) in [
|
||||
(CodecKind::Fp32, 128, 512),
|
||||
(CodecKind::Fp16, 128, 256),
|
||||
(CodecKind::Int8, 128, 136),
|
||||
(CodecKind::Turbo4, 128, 72),
|
||||
] {
|
||||
assert_eq!(kind.encoded_len(dim), Some(expect));
|
||||
let codec = codec_for(kind, dim, DistanceMetric::Euclidean, 42).unwrap();
|
||||
assert_eq!(codec.encoded_len(), expect);
|
||||
assert_eq!(codec.kind(), kind);
|
||||
}
|
||||
assert_eq!(CodecKind::Pq.encoded_len(128), None);
|
||||
assert!(codec_for(CodecKind::Pq, 128, DistanceMetric::Euclidean, 42).is_err());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn provenance_roundtrips_serde() {
|
||||
let codec = codec_for(CodecKind::Turbo4, 64, DistanceMetric::Cosine, 42).unwrap();
|
||||
let mut p = VectorProvenance::new(codec.as_ref(), DistanceMetric::Cosine);
|
||||
p.rotation_seed = Some(42);
|
||||
p.model_id = Some("test-model".into());
|
||||
p.push_migration("fp32→turbo4@2026-08-06");
|
||||
let json = serde_json::to_string(&p).unwrap();
|
||||
let back: VectorProvenance = serde_json::from_str(&json).unwrap();
|
||||
assert_eq!(p, back);
|
||||
assert_eq!(back.lineage.len(), 1);
|
||||
}
|
||||
}
|
||||
500
crates/ruvector-core/src/encoding/codecs.rs
Normal file
500
crates/ruvector-core/src/encoding/codecs.rs
Normal file
|
|
@ -0,0 +1,500 @@
|
|||
//! In-core implementations of the [`VectorCodec`] plane: Fp32, Fp16, Int8,
|
||||
//! and the Turbo4 adapter over `ruvector-turboquant` (ADR-297 §1).
|
||||
//!
|
||||
//! Fp32/Fp16/Int8 score by decoding — they are the verification / hot tiers,
|
||||
//! where fidelity matters more than per-candidate cost. Turbo4 scores
|
||||
//! directly on packed codes (never reconstructs), as required by ADR-296.
|
||||
|
||||
use super::{metric_distance, CodecKind, EncodedQuery, VectorCodec};
|
||||
use crate::error::{Result, RuvectorError};
|
||||
use crate::types::DistanceMetric;
|
||||
use ruvector_turboquant as tq;
|
||||
|
||||
fn check_dim(expected: usize, actual: usize) -> Result<()> {
|
||||
if expected != actual {
|
||||
return Err(RuvectorError::DimensionMismatch { expected, actual });
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
fn check_blob(expected: usize, actual: usize) -> Result<()> {
|
||||
if expected != actual {
|
||||
return Err(RuvectorError::InvalidInput(format!(
|
||||
"encoded blob length {actual} != expected {expected}"
|
||||
)));
|
||||
}
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Query that scores by decoding the candidate blob to f32 first.
|
||||
struct DecodingQuery {
|
||||
q: Vec<f32>,
|
||||
metric: DistanceMetric,
|
||||
decode: fn(&[u8]) -> Vec<f32>,
|
||||
}
|
||||
|
||||
impl EncodedQuery for DecodingQuery {
|
||||
fn distance_to(&self, blob: &[u8]) -> f32 {
|
||||
let v = (self.decode)(blob);
|
||||
metric_distance(self.metric, &self.q, &v)
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Fp32
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// Identity codec: 4 bytes/dim, exact. The verification anchor of the plane.
|
||||
pub struct Fp32Codec {
|
||||
dim: usize,
|
||||
metric: DistanceMetric,
|
||||
}
|
||||
|
||||
impl Fp32Codec {
|
||||
pub fn new(dim: usize, metric: DistanceMetric) -> Self {
|
||||
Self { dim, metric }
|
||||
}
|
||||
}
|
||||
|
||||
fn fp32_decode(blob: &[u8]) -> Vec<f32> {
|
||||
blob.chunks_exact(4)
|
||||
.map(|c| f32::from_le_bytes(c.try_into().unwrap()))
|
||||
.collect()
|
||||
}
|
||||
|
||||
impl VectorCodec for Fp32Codec {
|
||||
fn kind(&self) -> CodecKind {
|
||||
CodecKind::Fp32
|
||||
}
|
||||
fn dim(&self) -> usize {
|
||||
self.dim
|
||||
}
|
||||
fn encoded_len(&self) -> usize {
|
||||
self.dim * 4
|
||||
}
|
||||
fn encode(&self, v: &[f32]) -> Result<Vec<u8>> {
|
||||
check_dim(self.dim, v.len())?;
|
||||
Ok(v.iter().flat_map(|x| x.to_le_bytes()).collect())
|
||||
}
|
||||
fn decode(&self, blob: &[u8]) -> Result<Vec<f32>> {
|
||||
check_blob(self.encoded_len(), blob.len())?;
|
||||
Ok(fp32_decode(blob))
|
||||
}
|
||||
fn distance(&self, a: &[u8], b: &[u8]) -> Result<f32> {
|
||||
check_blob(self.encoded_len(), a.len())?;
|
||||
check_blob(self.encoded_len(), b.len())?;
|
||||
Ok(metric_distance(
|
||||
self.metric,
|
||||
&fp32_decode(a),
|
||||
&fp32_decode(b),
|
||||
))
|
||||
}
|
||||
fn make_query(&self, q: &[f32]) -> Result<Box<dyn EncodedQuery>> {
|
||||
check_dim(self.dim, q.len())?;
|
||||
Ok(Box::new(DecodingQuery {
|
||||
q: q.to_vec(),
|
||||
metric: self.metric,
|
||||
decode: fp32_decode,
|
||||
}))
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Fp16
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// IEEE 754 binary16: 2 bytes/dim. Hot-tier codec (~1e-3 relative error).
|
||||
/// Conversion is implemented in-crate (no `half` dependency) with
|
||||
/// round-to-nearest-even, matching hardware semantics.
|
||||
pub struct Fp16Codec {
|
||||
dim: usize,
|
||||
metric: DistanceMetric,
|
||||
}
|
||||
|
||||
impl Fp16Codec {
|
||||
pub fn new(dim: usize, metric: DistanceMetric) -> Self {
|
||||
Self { dim, metric }
|
||||
}
|
||||
}
|
||||
|
||||
/// f32 → binary16 bits, round-to-nearest-even; overflow → ±inf.
|
||||
pub fn f32_to_f16_bits(x: f32) -> u16 {
|
||||
let bits = x.to_bits();
|
||||
let sign = ((bits >> 16) & 0x8000) as u16;
|
||||
let exp = ((bits >> 23) & 0xFF) as i32;
|
||||
let mant = bits & 0x007F_FFFF;
|
||||
if exp == 255 {
|
||||
// Inf / NaN (preserve NaN-ness with a set mantissa bit).
|
||||
return sign | 0x7C00 | if mant != 0 { 0x0200 } else { 0 };
|
||||
}
|
||||
let unbiased = exp - 127;
|
||||
if unbiased > 15 {
|
||||
return sign | 0x7C00; // overflow → inf
|
||||
}
|
||||
if unbiased >= -14 {
|
||||
// Normal half.
|
||||
let m16 = (mant >> 13) as u16;
|
||||
let h = sign | (((unbiased + 15) as u16) << 10) | m16;
|
||||
let round = mant & 0x1FFF;
|
||||
// RNE; a carry out of the mantissa correctly increments the exponent.
|
||||
if round > 0x1000 || (round == 0x1000 && (m16 & 1) == 1) {
|
||||
h + 1
|
||||
} else {
|
||||
h
|
||||
}
|
||||
} else if unbiased >= -24 {
|
||||
// Subnormal half.
|
||||
let shift = (13 + (-14 - unbiased)) as u32;
|
||||
let full = mant | 0x0080_0000;
|
||||
let m16 = (full >> shift) as u16;
|
||||
let rem = full & ((1u32 << shift) - 1);
|
||||
let half = 1u32 << (shift - 1);
|
||||
let h = sign | m16;
|
||||
if rem > half || (rem == half && (m16 & 1) == 1) {
|
||||
h + 1
|
||||
} else {
|
||||
h
|
||||
}
|
||||
} else {
|
||||
sign // underflow → ±0
|
||||
}
|
||||
}
|
||||
|
||||
/// binary16 bits → f32 (exact).
|
||||
pub fn f16_bits_to_f32(h: u16) -> f32 {
|
||||
let sign = ((h & 0x8000) as u32) << 16;
|
||||
let exp = ((h >> 10) & 0x1F) as u32;
|
||||
let mant = (h & 0x03FF) as u32;
|
||||
let bits = if exp == 0 {
|
||||
if mant == 0 {
|
||||
sign
|
||||
} else {
|
||||
// Subnormal: renormalize.
|
||||
let mut e = 113u32; // 127 - 15 + 1
|
||||
let mut m = mant;
|
||||
while m & 0x0400 == 0 {
|
||||
m <<= 1;
|
||||
e -= 1;
|
||||
}
|
||||
sign | (e << 23) | ((m & 0x03FF) << 13)
|
||||
}
|
||||
} else if exp == 31 {
|
||||
sign | 0x7F80_0000 | (mant << 13)
|
||||
} else {
|
||||
sign | ((exp + 112) << 23) | (mant << 13)
|
||||
};
|
||||
f32::from_bits(bits)
|
||||
}
|
||||
|
||||
fn fp16_decode(blob: &[u8]) -> Vec<f32> {
|
||||
blob.chunks_exact(2)
|
||||
.map(|c| f16_bits_to_f32(u16::from_le_bytes(c.try_into().unwrap())))
|
||||
.collect()
|
||||
}
|
||||
|
||||
impl VectorCodec for Fp16Codec {
|
||||
fn kind(&self) -> CodecKind {
|
||||
CodecKind::Fp16
|
||||
}
|
||||
fn dim(&self) -> usize {
|
||||
self.dim
|
||||
}
|
||||
fn encoded_len(&self) -> usize {
|
||||
self.dim * 2
|
||||
}
|
||||
fn encode(&self, v: &[f32]) -> Result<Vec<u8>> {
|
||||
check_dim(self.dim, v.len())?;
|
||||
Ok(v.iter()
|
||||
.flat_map(|x| f32_to_f16_bits(*x).to_le_bytes())
|
||||
.collect())
|
||||
}
|
||||
fn decode(&self, blob: &[u8]) -> Result<Vec<f32>> {
|
||||
check_blob(self.encoded_len(), blob.len())?;
|
||||
Ok(fp16_decode(blob))
|
||||
}
|
||||
fn distance(&self, a: &[u8], b: &[u8]) -> Result<f32> {
|
||||
check_blob(self.encoded_len(), a.len())?;
|
||||
check_blob(self.encoded_len(), b.len())?;
|
||||
Ok(metric_distance(
|
||||
self.metric,
|
||||
&fp16_decode(a),
|
||||
&fp16_decode(b),
|
||||
))
|
||||
}
|
||||
fn make_query(&self, q: &[f32]) -> Result<Box<dyn EncodedQuery>> {
|
||||
check_dim(self.dim, q.len())?;
|
||||
Ok(Box::new(DecodingQuery {
|
||||
q: q.to_vec(),
|
||||
metric: self.metric,
|
||||
decode: fp16_decode,
|
||||
}))
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Int8
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
/// Per-vector min/scale uniform int8: `[min f32 | scale f32 | D bytes]`.
|
||||
/// Same scheme as the legacy `ScalarQuantized`, but blob-oriented so the
|
||||
/// plane can move it around like any other representation.
|
||||
pub struct Int8Codec {
|
||||
dim: usize,
|
||||
metric: DistanceMetric,
|
||||
}
|
||||
|
||||
impl Int8Codec {
|
||||
pub fn new(dim: usize, metric: DistanceMetric) -> Self {
|
||||
Self { dim, metric }
|
||||
}
|
||||
}
|
||||
|
||||
fn int8_decode(blob: &[u8]) -> Vec<f32> {
|
||||
let min = f32::from_le_bytes(blob[0..4].try_into().unwrap());
|
||||
let scale = f32::from_le_bytes(blob[4..8].try_into().unwrap());
|
||||
blob[8..].iter().map(|&u| min + u as f32 * scale).collect()
|
||||
}
|
||||
|
||||
impl VectorCodec for Int8Codec {
|
||||
fn kind(&self) -> CodecKind {
|
||||
CodecKind::Int8
|
||||
}
|
||||
fn dim(&self) -> usize {
|
||||
self.dim
|
||||
}
|
||||
fn encoded_len(&self) -> usize {
|
||||
self.dim + 8
|
||||
}
|
||||
fn encode(&self, v: &[f32]) -> Result<Vec<u8>> {
|
||||
check_dim(self.dim, v.len())?;
|
||||
let min = v.iter().copied().fold(f32::INFINITY, f32::min);
|
||||
let max = v.iter().copied().fold(f32::NEG_INFINITY, f32::max);
|
||||
let scale = if (max - min).abs() < f32::EPSILON {
|
||||
1.0
|
||||
} else {
|
||||
(max - min) / 255.0
|
||||
};
|
||||
let mut blob = Vec::with_capacity(self.encoded_len());
|
||||
blob.extend_from_slice(&min.to_le_bytes());
|
||||
blob.extend_from_slice(&scale.to_le_bytes());
|
||||
blob.extend(
|
||||
v.iter()
|
||||
.map(|&x| ((x - min) / scale).round().clamp(0.0, 255.0) as u8),
|
||||
);
|
||||
Ok(blob)
|
||||
}
|
||||
fn decode(&self, blob: &[u8]) -> Result<Vec<f32>> {
|
||||
check_blob(self.encoded_len(), blob.len())?;
|
||||
Ok(int8_decode(blob))
|
||||
}
|
||||
fn distance(&self, a: &[u8], b: &[u8]) -> Result<f32> {
|
||||
check_blob(self.encoded_len(), a.len())?;
|
||||
check_blob(self.encoded_len(), b.len())?;
|
||||
Ok(metric_distance(
|
||||
self.metric,
|
||||
&int8_decode(a),
|
||||
&int8_decode(b),
|
||||
))
|
||||
}
|
||||
fn make_query(&self, q: &[f32]) -> Result<Box<dyn EncodedQuery>> {
|
||||
check_dim(self.dim, q.len())?;
|
||||
Ok(Box::new(DecodingQuery {
|
||||
q: q.to_vec(),
|
||||
metric: self.metric,
|
||||
decode: int8_decode,
|
||||
}))
|
||||
}
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Turbo4 adapter
|
||||
// ---------------------------------------------------------------------------
|
||||
|
||||
fn to_turbo_metric(metric: DistanceMetric) -> Result<tq::Metric> {
|
||||
match metric {
|
||||
DistanceMetric::Euclidean => Ok(tq::Metric::Euclidean),
|
||||
DistanceMetric::Cosine => Ok(tq::Metric::Cosine),
|
||||
DistanceMetric::DotProduct => Ok(tq::Metric::DotProduct),
|
||||
DistanceMetric::Manhattan => Err(RuvectorError::InvalidParameter(
|
||||
"Turbo4 does not support the Manhattan metric".into(),
|
||||
)),
|
||||
}
|
||||
}
|
||||
|
||||
/// [`VectorCodec`] adapter over `ruvector_turboquant::Turbo4Codec` — scores
|
||||
/// on packed codes, never reconstructs (ADR-296).
|
||||
pub struct Turbo4PlaneCodec {
|
||||
inner: tq::Turbo4Codec,
|
||||
metric: tq::Metric,
|
||||
}
|
||||
|
||||
impl Turbo4PlaneCodec {
|
||||
pub fn new(dim: usize, metric: DistanceMetric, rotation_seed: u64) -> Result<Self> {
|
||||
let metric = to_turbo_metric(metric)?;
|
||||
let inner = tq::Turbo4Codec::new(dim, rotation_seed)
|
||||
.map_err(|e| RuvectorError::InvalidParameter(e.to_string()))?;
|
||||
Ok(Self { inner, metric })
|
||||
}
|
||||
}
|
||||
|
||||
struct Turbo4PlaneQuery {
|
||||
query: tq::Turbo4Query,
|
||||
metric: tq::Metric,
|
||||
dim: usize,
|
||||
}
|
||||
|
||||
impl EncodedQuery for Turbo4PlaneQuery {
|
||||
fn distance_to(&self, blob: &[u8]) -> f32 {
|
||||
tq::rescore(self.metric, &self.query, blob, self.dim)
|
||||
}
|
||||
}
|
||||
|
||||
impl VectorCodec for Turbo4PlaneCodec {
|
||||
fn kind(&self) -> CodecKind {
|
||||
CodecKind::Turbo4
|
||||
}
|
||||
fn dim(&self) -> usize {
|
||||
self.inner.dim()
|
||||
}
|
||||
fn encoded_len(&self) -> usize {
|
||||
self.inner.code_len()
|
||||
}
|
||||
fn encode(&self, v: &[f32]) -> Result<Vec<u8>> {
|
||||
self.inner
|
||||
.encode(v)
|
||||
.map_err(|e| RuvectorError::InvalidInput(e.to_string()))
|
||||
}
|
||||
fn decode(&self, blob: &[u8]) -> Result<Vec<f32>> {
|
||||
check_blob(self.encoded_len(), blob.len())?;
|
||||
Ok(self.inner.decode(blob))
|
||||
}
|
||||
fn distance(&self, a: &[u8], b: &[u8]) -> Result<f32> {
|
||||
check_blob(self.encoded_len(), a.len())?;
|
||||
check_blob(self.encoded_len(), b.len())?;
|
||||
Ok(tq::symmetric_distance(self.metric, a, b, self.inner.dim()))
|
||||
}
|
||||
fn make_query(&self, q: &[f32]) -> Result<Box<dyn EncodedQuery>> {
|
||||
let query = self
|
||||
.inner
|
||||
.encode_query(q)
|
||||
.map_err(|e| RuvectorError::InvalidInput(e.to_string()))?;
|
||||
Ok(Box::new(Turbo4PlaneQuery {
|
||||
query,
|
||||
metric: self.metric,
|
||||
dim: self.inner.dim(),
|
||||
}))
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::super::{codec_for, CodecKind};
|
||||
use super::*;
|
||||
|
||||
/// Decorrelated deterministic vectors (SplitMix64 per (seed, i)).
|
||||
fn test_vec(dim: usize, seed: u64) -> Vec<f32> {
|
||||
(0..dim as u64)
|
||||
.map(|i| {
|
||||
let mut z = (seed << 32 | i).wrapping_add(0x9E3779B97F4A7C15);
|
||||
z = (z ^ (z >> 30)).wrapping_mul(0xBF58476D1CE4E5B9);
|
||||
z = (z ^ (z >> 27)).wrapping_mul(0x94D049BB133111EB);
|
||||
((z >> 40) as f32 / (1u64 << 24) as f32) * 2.0 - 1.0
|
||||
})
|
||||
.collect()
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn f16_conversion_known_values() {
|
||||
for (x, bits) in [
|
||||
(0.0f32, 0x0000u16),
|
||||
(1.0, 0x3C00),
|
||||
(-2.0, 0xC000),
|
||||
(65504.0, 0x7BFF), // max finite half
|
||||
(0.5, 0x3800),
|
||||
] {
|
||||
assert_eq!(f32_to_f16_bits(x), bits, "encode {x}");
|
||||
assert_eq!(f16_bits_to_f32(bits), x, "decode {x}");
|
||||
}
|
||||
// Overflow → inf; NaN stays NaN; subnormals roundtrip.
|
||||
assert_eq!(f32_to_f16_bits(1e30), 0x7C00);
|
||||
assert!(f16_bits_to_f32(f32_to_f16_bits(f32::NAN)).is_nan());
|
||||
let sub = 3.0e-6f32;
|
||||
let rt = f16_bits_to_f32(f32_to_f16_bits(sub));
|
||||
assert!((rt - sub).abs() / sub < 0.05, "subnormal {sub} -> {rt}");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn f16_roundtrip_relative_error() {
|
||||
for &x in &[1.0f32, -1.5, 0.123, 999.25, -0.0004883, 3.1415926] {
|
||||
let rt = f16_bits_to_f32(f32_to_f16_bits(x));
|
||||
assert!(
|
||||
(rt - x).abs() / x.abs() < 1e-3,
|
||||
"roundtrip {x} -> {rt} exceeds half precision"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn all_codecs_roundtrip_and_score() {
|
||||
let dim = 64;
|
||||
let a = test_vec(dim, 1);
|
||||
let b = test_vec(dim, 7);
|
||||
let exact = metric_distance(DistanceMetric::Euclidean, &a, &b);
|
||||
|
||||
for kind in [
|
||||
CodecKind::Fp32,
|
||||
CodecKind::Fp16,
|
||||
CodecKind::Int8,
|
||||
CodecKind::Turbo4,
|
||||
] {
|
||||
let codec = codec_for(kind, dim, DistanceMetric::Euclidean, 42).unwrap();
|
||||
let ca = codec.encode(&a).unwrap();
|
||||
let cb = codec.encode(&b).unwrap();
|
||||
assert_eq!(ca.len(), codec.encoded_len(), "{kind:?} blob length");
|
||||
|
||||
// decode error bounded per codec class
|
||||
let da = codec.decode(&ca).unwrap();
|
||||
let err: f32 = a.iter().zip(&da).map(|(x, y)| (x - y).abs()).sum::<f32>() / dim as f32;
|
||||
let tol = match kind {
|
||||
CodecKind::Fp32 => 1e-9,
|
||||
CodecKind::Fp16 => 1e-3,
|
||||
CodecKind::Int8 => 0.01,
|
||||
_ => 0.2,
|
||||
};
|
||||
assert!(err <= tol, "{kind:?} mean decode error {err} > {tol}");
|
||||
|
||||
// symmetric + asymmetric agree with exact within codec tolerance
|
||||
let sym = codec.distance(&ca, &cb).unwrap();
|
||||
let asym = codec.make_query(&a).unwrap().distance_to(&cb);
|
||||
let dtol = match kind {
|
||||
CodecKind::Fp32 => 1e-4,
|
||||
CodecKind::Fp16 => 1e-2,
|
||||
CodecKind::Int8 => 0.05,
|
||||
_ => 0.30,
|
||||
};
|
||||
assert!(
|
||||
(sym - exact).abs() <= dtol * exact.max(1.0),
|
||||
"{kind:?} symmetric {sym} vs exact {exact}"
|
||||
);
|
||||
assert!(
|
||||
(asym - exact).abs() <= dtol * exact.max(1.0),
|
||||
"{kind:?} asymmetric {asym} vs exact {exact}"
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn turbo4_plane_matches_direct_codec() {
|
||||
let dim = 128;
|
||||
let v = test_vec(dim, 3);
|
||||
let plane = codec_for(CodecKind::Turbo4, dim, DistanceMetric::Cosine, 42).unwrap();
|
||||
let direct = tq::Turbo4Codec::new(dim, 42).unwrap();
|
||||
assert_eq!(plane.encode(&v).unwrap(), direct.encode(&v).unwrap());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn manhattan_rejected_only_for_turbo4() {
|
||||
assert!(codec_for(CodecKind::Turbo4, 64, DistanceMetric::Manhattan, 42).is_err());
|
||||
assert!(codec_for(CodecKind::Fp16, 64, DistanceMetric::Manhattan, 42).is_ok());
|
||||
}
|
||||
}
|
||||
|
|
@ -14,13 +14,45 @@
|
|||
|
||||
use crate::error::{Result, RuvectorError};
|
||||
use crate::index::VectorIndex;
|
||||
use crate::types::{DistanceMetric, HnswConfig, SearchResult, VectorId};
|
||||
use crate::types::{DistanceMetric, HnswConfig, SearchPolicy, SearchResult, VectorId};
|
||||
use dashmap::DashMap;
|
||||
use hnsw_rs::prelude::*;
|
||||
use parking_lot::RwLock;
|
||||
use ruvector_turboquant::{score, Metric, Turbo4Codec};
|
||||
use ruvector_turboquant::{score, Metric, Turbo4Codec, Turbo4Query};
|
||||
use std::sync::atomic::{AtomicU64, Ordering};
|
||||
use std::sync::Arc;
|
||||
|
||||
/// Escalation parameters derived from a [`SearchPolicy`] (ADR-297 §3/§9):
|
||||
/// relative score-margin below which a query is "uncertain", the ef widening
|
||||
/// factor per escalation round, and how many rounds are allowed.
|
||||
struct EscalationParams {
|
||||
margin_threshold: f32,
|
||||
ef_mult: usize,
|
||||
max_rounds: usize,
|
||||
}
|
||||
|
||||
impl EscalationParams {
|
||||
fn for_policy(policy: SearchPolicy) -> Self {
|
||||
match policy {
|
||||
SearchPolicy::Quality => Self {
|
||||
margin_threshold: 0.05,
|
||||
ef_mult: 3,
|
||||
max_rounds: 2,
|
||||
},
|
||||
SearchPolicy::Balanced => Self {
|
||||
margin_threshold: 0.02,
|
||||
ef_mult: 2,
|
||||
max_rounds: 1,
|
||||
},
|
||||
SearchPolicy::MaxCompression => Self {
|
||||
margin_threshold: 0.0,
|
||||
ef_mult: 1,
|
||||
max_rounds: 0,
|
||||
},
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
/// Distance functor over Turbo4 blobs. Chooses the kernel from blob lengths:
|
||||
/// hnsw_rs passes the search query straight through to `eval`, which is what
|
||||
/// makes true asymmetric traversal possible without forking hnsw_rs.
|
||||
|
|
@ -84,6 +116,10 @@ pub struct Turbo4HnswIndex {
|
|||
metric: Metric,
|
||||
dimensions: usize,
|
||||
rescore_multiplier: usize,
|
||||
escalation: EscalationParams,
|
||||
/// Adaptive-plane telemetry: total queries / queries that escalated.
|
||||
queries: AtomicU64,
|
||||
escalated: AtomicU64,
|
||||
}
|
||||
|
||||
impl Turbo4HnswIndex {
|
||||
|
|
@ -94,6 +130,7 @@ impl Turbo4HnswIndex {
|
|||
config: HnswConfig,
|
||||
rotation_seed: u64,
|
||||
rescore_multiplier: usize,
|
||||
policy: SearchPolicy,
|
||||
) -> Result<Self> {
|
||||
let metric = to_turbo_metric(metric)?;
|
||||
let codec = Turbo4Codec::new(dimensions, rotation_seed).map_err(|e| {
|
||||
|
|
@ -125,6 +162,9 @@ impl Turbo4HnswIndex {
|
|||
metric,
|
||||
dimensions,
|
||||
rescore_multiplier: rescore_multiplier.max(1),
|
||||
escalation: EscalationParams::for_policy(policy),
|
||||
queries: AtomicU64::new(0),
|
||||
escalated: AtomicU64::new(0),
|
||||
})
|
||||
}
|
||||
|
||||
|
|
@ -133,6 +173,59 @@ impl Turbo4HnswIndex {
|
|||
self.codec.code_len()
|
||||
}
|
||||
|
||||
/// Adaptive-plane telemetry: `(total_queries, escalated_queries)`.
|
||||
/// A healthy workload escalates only a small fraction (~5–15 %).
|
||||
pub fn adaptive_stats(&self) -> (u64, u64) {
|
||||
(
|
||||
self.queries.load(Ordering::Relaxed),
|
||||
self.escalated.load(Ordering::Relaxed),
|
||||
)
|
||||
}
|
||||
|
||||
/// One traversal + exact-rescore pass. Returns the full rescored
|
||||
/// candidate list, ascending — the margin between `list[k-1]` and
|
||||
/// `list[k]` is the stability signal for adaptive escalation.
|
||||
fn traverse_and_rescore(
|
||||
&self,
|
||||
inner: &Turbo4Inner,
|
||||
prepared: &Turbo4Query,
|
||||
fetch: usize,
|
||||
ef: usize,
|
||||
) -> Vec<SearchResult> {
|
||||
let neighbors = inner.hnsw.search(&prepared.blob, fetch, ef);
|
||||
let mut results: Vec<SearchResult> = neighbors
|
||||
.into_iter()
|
||||
.filter_map(|n| {
|
||||
let id = inner.idx_to_id.get(&n.d_id)?.clone();
|
||||
let code = inner.codes.get(&id)?;
|
||||
let dist = score::rescore(self.metric, prepared, code.value(), self.dimensions);
|
||||
Some(SearchResult {
|
||||
id,
|
||||
score: dist,
|
||||
vector: None,
|
||||
metadata: None,
|
||||
})
|
||||
})
|
||||
.collect();
|
||||
results.sort_unstable_by(|a, b| {
|
||||
a.score
|
||||
.partial_cmp(&b.score)
|
||||
.unwrap_or(std::cmp::Ordering::Equal)
|
||||
});
|
||||
results
|
||||
}
|
||||
|
||||
/// Relative margin between the last kept and first dropped candidate.
|
||||
/// `f32::INFINITY` when nothing is dropped (result already exhaustive).
|
||||
fn boundary_margin(results: &[SearchResult], k: usize) -> f32 {
|
||||
if results.len() <= k || k == 0 {
|
||||
return f32::INFINITY;
|
||||
}
|
||||
let kept = results[k - 1].score;
|
||||
let dropped = results[k].score;
|
||||
(dropped - kept) / kept.abs().max(1e-6)
|
||||
}
|
||||
|
||||
/// Search with an explicit `ef_search`.
|
||||
pub fn search_with_ef(
|
||||
&self,
|
||||
|
|
@ -159,33 +252,43 @@ impl Turbo4HnswIndex {
|
|||
if inner.codes.is_empty() {
|
||||
return Ok(vec![]);
|
||||
}
|
||||
self.queries.fetch_add(1, Ordering::Relaxed);
|
||||
|
||||
// Over-fetch for the exact rescoring pass.
|
||||
let fetch = k.saturating_mul(self.rescore_multiplier);
|
||||
let effective_ef = ef_search.max(fetch);
|
||||
let neighbors = inner.hnsw.search(&prepared.blob, fetch, effective_ef);
|
||||
// Base pass: over-fetch for the exact rescoring step.
|
||||
let mut fetch = k.saturating_mul(self.rescore_multiplier);
|
||||
let mut ef = ef_search.max(fetch);
|
||||
let mut results = self.traverse_and_rescore(&inner, &prepared, fetch, ef);
|
||||
|
||||
// Re-rank candidates with the exact f32 query against stored codes.
|
||||
let mut results: Vec<SearchResult> = neighbors
|
||||
.into_iter()
|
||||
.filter_map(|n| {
|
||||
let id = inner.idx_to_id.get(&n.d_id)?.clone();
|
||||
let code = inner.codes.get(&id)?;
|
||||
let dist = score::rescore(self.metric, &prepared, code.value(), self.dimensions);
|
||||
Some(SearchResult {
|
||||
id,
|
||||
score: dist,
|
||||
vector: None,
|
||||
metadata: None,
|
||||
})
|
||||
})
|
||||
.collect();
|
||||
// Adaptive escalation (ADR-297 §3): if the kept/dropped boundary sits
|
||||
// inside the quantization noise band, widen the search. Stop as soon
|
||||
// as the top-k membership is stable across rounds — that's the
|
||||
// candidate-stability signal, independent of absolute scores.
|
||||
let mut round = 0;
|
||||
let mut did_escalate = false;
|
||||
while round < self.escalation.max_rounds
|
||||
&& Self::boundary_margin(&results, k) < self.escalation.margin_threshold
|
||||
{
|
||||
did_escalate = true;
|
||||
fetch = fetch.saturating_mul(2);
|
||||
ef = ef.saturating_mul(self.escalation.ef_mult);
|
||||
let wider = self.traverse_and_rescore(&inner, &prepared, fetch, ef);
|
||||
|
||||
let stable = wider.len() >= k
|
||||
&& results.len() >= k
|
||||
&& wider[..k]
|
||||
.iter()
|
||||
.zip(&results[..k])
|
||||
.all(|(a, b)| a.id == b.id);
|
||||
results = wider;
|
||||
round += 1;
|
||||
if stable {
|
||||
break;
|
||||
}
|
||||
}
|
||||
if did_escalate {
|
||||
self.escalated.fetch_add(1, Ordering::Relaxed);
|
||||
}
|
||||
|
||||
results.sort_unstable_by(|a, b| {
|
||||
a.score
|
||||
.partial_cmp(&b.score)
|
||||
.unwrap_or(std::cmp::Ordering::Equal)
|
||||
});
|
||||
results.truncate(k);
|
||||
Ok(results)
|
||||
}
|
||||
|
|
@ -326,7 +429,14 @@ mod tests {
|
|||
|
||||
#[test]
|
||||
fn manhattan_is_rejected() {
|
||||
let err = Turbo4HnswIndex::new(64, DistanceMetric::Manhattan, HnswConfig::default(), 42, 4);
|
||||
let err = Turbo4HnswIndex::new(
|
||||
64,
|
||||
DistanceMetric::Manhattan,
|
||||
HnswConfig::default(),
|
||||
42,
|
||||
4,
|
||||
SearchPolicy::Balanced,
|
||||
);
|
||||
assert!(err.is_err());
|
||||
}
|
||||
|
||||
|
|
@ -380,7 +490,14 @@ mod tests {
|
|||
.map(|(i, v)| (format!("v{i}"), v.clone()))
|
||||
.collect();
|
||||
|
||||
let mut t4 = Turbo4HnswIndex::new(dim, DistanceMetric::Euclidean, config.clone(), 42, 8)?;
|
||||
let mut t4 = Turbo4HnswIndex::new(
|
||||
dim,
|
||||
DistanceMetric::Euclidean,
|
||||
config.clone(),
|
||||
42,
|
||||
8,
|
||||
SearchPolicy::Balanced,
|
||||
)?;
|
||||
t4.add_batch(entries.clone())?;
|
||||
assert_eq!(t4.len(), n);
|
||||
let mut f32_ix = HnswIndex::new(dim, DistanceMetric::Euclidean, config)?;
|
||||
|
|
@ -453,7 +570,14 @@ mod tests {
|
|||
.enumerate()
|
||||
.map(|(i, v)| (format!("v{i}"), v.clone()))
|
||||
.collect();
|
||||
let mut t4 = Turbo4HnswIndex::new(dim, DistanceMetric::Euclidean, config, 42, 8)?;
|
||||
let mut t4 = Turbo4HnswIndex::new(
|
||||
dim,
|
||||
DistanceMetric::Euclidean,
|
||||
config,
|
||||
42,
|
||||
8,
|
||||
SearchPolicy::Balanced,
|
||||
)?;
|
||||
t4.add_batch(entries)?;
|
||||
|
||||
let queries = gauss_vecs(20, dim, 12345);
|
||||
|
|
@ -486,7 +610,14 @@ mod tests {
|
|||
fn self_query_returns_self_first() -> Result<()> {
|
||||
let dim = 64;
|
||||
let config = HnswConfig::default();
|
||||
let mut index = Turbo4HnswIndex::new(dim, DistanceMetric::Cosine, config, 42, 4)?;
|
||||
let mut index = Turbo4HnswIndex::new(
|
||||
dim,
|
||||
DistanceMetric::Cosine,
|
||||
config,
|
||||
42,
|
||||
4,
|
||||
SearchPolicy::Balanced,
|
||||
)?;
|
||||
let vectors = gauss_vecs(50, dim, 3);
|
||||
for (i, v) in vectors.iter().enumerate() {
|
||||
index.add(format!("v{i}"), v.clone())?;
|
||||
|
|
@ -496,10 +627,132 @@ mod tests {
|
|||
Ok(())
|
||||
}
|
||||
|
||||
/// Adaptive escalation (ADR-297 §3): MaxCompression must never escalate;
|
||||
/// Quality must escalate on distance-concentrated (uncertain) queries and
|
||||
/// must not lose recall relative to the non-escalating policy.
|
||||
#[test]
|
||||
fn adaptive_escalation_follows_policy() -> Result<()> {
|
||||
let dim = 128;
|
||||
let n = 400;
|
||||
let config = HnswConfig {
|
||||
m: 16,
|
||||
ef_construction: 200,
|
||||
ef_search: 60,
|
||||
max_elements: 1000,
|
||||
};
|
||||
// i.i.d. Gaussian ⇒ tight boundary margins ⇒ escalation should fire.
|
||||
let vectors = gauss_vecs(n, dim, 21);
|
||||
let entries: Vec<_> = vectors
|
||||
.iter()
|
||||
.enumerate()
|
||||
.map(|(i, v)| (format!("v{i}"), v.clone()))
|
||||
.collect();
|
||||
|
||||
let mut fixed = Turbo4HnswIndex::new(
|
||||
dim,
|
||||
DistanceMetric::Euclidean,
|
||||
config.clone(),
|
||||
42,
|
||||
2,
|
||||
SearchPolicy::MaxCompression,
|
||||
)?;
|
||||
fixed.add_batch(entries.clone())?;
|
||||
let mut adaptive = Turbo4HnswIndex::new(
|
||||
dim,
|
||||
DistanceMetric::Euclidean,
|
||||
config,
|
||||
42,
|
||||
2,
|
||||
SearchPolicy::Quality,
|
||||
)?;
|
||||
adaptive.add_batch(entries)?;
|
||||
|
||||
let queries = gauss_vecs(25, dim, 4242);
|
||||
let (mut fixed_hits, mut adaptive_hits, mut total) = (0usize, 0usize, 0usize);
|
||||
for q in &queries {
|
||||
let mut truth: Vec<(usize, f32)> = vectors
|
||||
.iter()
|
||||
.enumerate()
|
||||
.map(|(i, v)| (i, l2(q, v)))
|
||||
.collect();
|
||||
truth.sort_by(|a, b| a.1.total_cmp(&b.1));
|
||||
let top10: std::collections::HashSet<String> =
|
||||
truth[..10].iter().map(|(i, _)| format!("v{i}")).collect();
|
||||
fixed_hits += fixed
|
||||
.search(q, 10)?
|
||||
.iter()
|
||||
.filter(|r| top10.contains(&r.id))
|
||||
.count();
|
||||
adaptive_hits += adaptive
|
||||
.search(q, 10)?
|
||||
.iter()
|
||||
.filter(|r| top10.contains(&r.id))
|
||||
.count();
|
||||
total += 10;
|
||||
}
|
||||
|
||||
let (fq, fe) = fixed.adaptive_stats();
|
||||
assert_eq!(fq, 25);
|
||||
assert_eq!(fe, 0, "MaxCompression must never escalate");
|
||||
|
||||
let (aq, ae) = adaptive.adaptive_stats();
|
||||
assert_eq!(aq, 25);
|
||||
assert!(
|
||||
ae > 0,
|
||||
"Quality policy should escalate on concentrated Gaussian queries"
|
||||
);
|
||||
assert!(
|
||||
adaptive_hits >= fixed_hits,
|
||||
"escalation must not lose recall: adaptive {adaptive_hits} vs fixed {fixed_hits} of {total}"
|
||||
);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
/// Easy well-separated queries must not trigger escalation under
|
||||
/// Balanced — the ~5–15 % escalation budget depends on margins staying
|
||||
/// wide when results are unambiguous.
|
||||
#[test]
|
||||
fn easy_queries_do_not_escalate() -> Result<()> {
|
||||
let dim = 64;
|
||||
let config = HnswConfig::default();
|
||||
let mut index = Turbo4HnswIndex::new(
|
||||
dim,
|
||||
DistanceMetric::Euclidean,
|
||||
config,
|
||||
42,
|
||||
4,
|
||||
SearchPolicy::Balanced,
|
||||
)?;
|
||||
// Well-separated clusters of exactly k members, so the kept/dropped
|
||||
// boundary falls BETWEEN clusters (wide margin). A boundary inside a
|
||||
// cluster is genuinely ambiguous and escalating on it is correct.
|
||||
let vectors = clustered_vecs(200, dim, 5, 0.05, 33);
|
||||
for (i, v) in vectors.iter().enumerate() {
|
||||
index.add(format!("v{i}"), v.iter().map(|x| x * 10.0).collect())?;
|
||||
}
|
||||
for j in 0..10 {
|
||||
let q: Vec<f32> = vectors[j * 20].iter().map(|x| x * 10.0).collect();
|
||||
index.search(&q, 5)?;
|
||||
}
|
||||
let (queries, escalated) = index.adaptive_stats();
|
||||
assert_eq!(queries, 10);
|
||||
assert!(
|
||||
escalated <= 2,
|
||||
"well-separated queries escalated {escalated}/10 times"
|
||||
);
|
||||
Ok(())
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn empty_index_is_safe() -> Result<()> {
|
||||
let index =
|
||||
Turbo4HnswIndex::new(64, DistanceMetric::Euclidean, HnswConfig::default(), 42, 4)?;
|
||||
let index = Turbo4HnswIndex::new(
|
||||
64,
|
||||
DistanceMetric::Euclidean,
|
||||
HnswConfig::default(),
|
||||
42,
|
||||
4,
|
||||
SearchPolicy::Balanced,
|
||||
)?;
|
||||
assert!(index.search(&vec![0.5; 64], 5)?.is_empty());
|
||||
Ok(())
|
||||
}
|
||||
|
|
|
|||
|
|
@ -37,6 +37,7 @@ pub mod agenticdb;
|
|||
|
||||
pub mod distance;
|
||||
pub mod embeddings;
|
||||
pub mod encoding;
|
||||
pub mod error;
|
||||
pub mod index;
|
||||
pub mod pdx;
|
||||
|
|
|
|||
|
|
@ -139,9 +139,26 @@ pub enum QuantizationConfig {
|
|||
/// exact rescoring pass. The recall/latency dial; default 4.
|
||||
#[serde(default = "default_turbo4_rescore_multiplier")]
|
||||
rescore_multiplier: usize,
|
||||
/// Outcome-level policy governing adaptive escalation (ADR-297 §3/§9).
|
||||
#[serde(default)]
|
||||
policy: SearchPolicy,
|
||||
},
|
||||
}
|
||||
|
||||
/// Outcome-level search policy (ADR-297 §9). Users pick a goal; the engine
|
||||
/// maps it to escalation thresholds, rescore pools, and verification tiers —
|
||||
/// no quantization knowledge required.
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq, Serialize, Deserialize, Default)]
|
||||
pub enum SearchPolicy {
|
||||
/// Maximum retrieval quality: aggressive escalation on uncertain queries.
|
||||
Quality,
|
||||
/// Default trade-off: escalate only when the result margin is unstable.
|
||||
#[default]
|
||||
Balanced,
|
||||
/// Minimum memory/latency: never escalate beyond the base quantized pass.
|
||||
MaxCompression,
|
||||
}
|
||||
|
||||
/// Default rotation seed for [`QuantizationConfig::Turbo4`].
|
||||
pub fn default_turbo4_rotation_seed() -> u64 {
|
||||
42
|
||||
|
|
|
|||
|
|
@ -83,7 +83,8 @@ impl VectorDB {
|
|||
Some(crate::types::QuantizationConfig::Turbo4 {
|
||||
rotation_seed,
|
||||
rescore_multiplier,
|
||||
}) => Some((*rotation_seed, *rescore_multiplier)),
|
||||
policy,
|
||||
}) => Some((*rotation_seed, *rescore_multiplier, *policy)),
|
||||
_ => None,
|
||||
};
|
||||
|
||||
|
|
@ -91,9 +92,9 @@ impl VectorDB {
|
|||
let mut index: Box<dyn VectorIndex> = if let Some(hnsw_config) = &options.hnsw_config {
|
||||
#[cfg(feature = "hnsw")]
|
||||
{
|
||||
if let Some((rotation_seed, rescore_multiplier)) = turbo4_params {
|
||||
if let Some((rotation_seed, rescore_multiplier, policy)) = turbo4_params {
|
||||
tracing::info!(
|
||||
"Turbo4 quantization active: {} bytes/vector instead of {}",
|
||||
"Turbo4 quantization active ({policy:?} policy): {} bytes/vector instead of {}",
|
||||
options.dimensions / 2 + 8,
|
||||
options.dimensions * 4
|
||||
);
|
||||
|
|
@ -103,6 +104,7 @@ impl VectorDB {
|
|||
hnsw_config.clone(),
|
||||
rotation_seed,
|
||||
rescore_multiplier,
|
||||
policy,
|
||||
)?) as Box<dyn VectorIndex>
|
||||
} else {
|
||||
Box::new(HnswIndex::new(
|
||||
|
|
@ -384,6 +386,7 @@ mod tests {
|
|||
options.quantization = Some(QuantizationConfig::Turbo4 {
|
||||
rotation_seed: 42,
|
||||
rescore_multiplier: 4,
|
||||
policy: SearchPolicy::Balanced,
|
||||
});
|
||||
options
|
||||
};
|
||||
|
|
|
|||
|
|
@ -74,6 +74,7 @@ impl From<JsQuantizationConfig> for QuantizationConfig {
|
|||
"turbo4" => QuantizationConfig::Turbo4 {
|
||||
rotation_seed: ruvector_core::types::default_turbo4_rotation_seed(),
|
||||
rescore_multiplier: ruvector_core::types::default_turbo4_rescore_multiplier(),
|
||||
policy: ruvector_core::types::SearchPolicy::default(),
|
||||
},
|
||||
_ => QuantizationConfig::Scalar,
|
||||
}
|
||||
|
|
|
|||
|
|
@ -212,6 +212,7 @@ mod tests {
|
|||
}
|
||||
|
||||
#[test]
|
||||
#[allow(clippy::identity_op, clippy::neg_multiply)] // literal per-dim products mirror the layout
|
||||
fn known_small_case() {
|
||||
// dim 4: dims 0,1 in low nibbles of bytes 0,1; dims 2,3 in high.
|
||||
// code: dim0=15 (level 127), dim1=0 (level -127), dim2=8 (6), dim3=7 (-6)
|
||||
|
|
|
|||
200
docs/adr/ADR-297-adaptive-compression-retrieval-plane.md
Normal file
200
docs/adr/ADR-297-adaptive-compression-retrieval-plane.md
Normal file
|
|
@ -0,0 +1,200 @@
|
|||
# ADR-297: Adaptive Compression & Retrieval Plane (ACRP)
|
||||
|
||||
- **Status**: Accepted
|
||||
- **Date**: 2026-08-06
|
||||
- **Extends**: ADR-296 (Turbo4 quantized vector datatype), ADR-254 (turbovec), ADR-026 (tiered routing)
|
||||
- **Related crates**: `ruvector-turboquant`, `ruvector-core`, `ruvector-rabitq`, `ruvector-pq-search`, `ruvector-attn-mincut`, `ruvector-agent-memory`, `rvf`
|
||||
|
||||
## Context
|
||||
|
||||
ADR-296 gives RuVector a faithful Turbo4 datatype: uniform 4-bit storage with
|
||||
direct packed scoring inside HNSW. That matches Qdrant's headline capability —
|
||||
but stopping there means *copying* Turbo4. The defensible advantage is an
|
||||
**adaptive compression and retrieval plane**: one system that chooses vector
|
||||
representations by workload, because retrieval errors are not equally costly
|
||||
across vectors, queries, or time.
|
||||
|
||||
Today's obstacles:
|
||||
|
||||
- Compression code is scattered across disconnected crates
|
||||
(`ruvector-core::quantization`, `ruvector-turboquant`, `ruvector-rabitq`,
|
||||
`ruvector-pq-search`, `ruvector-turbovec`, `ruvllm`), each with its own
|
||||
types; storage, HNSW, snapshots, WASM, and bindings cannot consume them
|
||||
interchangeably.
|
||||
- Precision is a single global choice; queries that need more bits pay the
|
||||
same as queries that need fewer, and vice versa.
|
||||
- No provenance: stored codes don't record which embedding model, codec
|
||||
version, or rotation seed produced them, so migrations and verification are
|
||||
not deterministic.
|
||||
|
||||
**Product claim to earn**: *RuVector automatically places each vector at the
|
||||
cheapest precision that preserves its retrieval value.*
|
||||
|
||||
## Decision
|
||||
|
||||
### 1. One `EncodedVector` interface for every representation
|
||||
|
||||
`ruvector-core::encoding` defines the unified codec plane:
|
||||
|
||||
```rust
|
||||
pub enum CodecKind { Fp32, Fp16, Int8, Turbo4, Pq, RaBitQ1 }
|
||||
|
||||
pub trait VectorCodec: Send + Sync {
|
||||
fn kind(&self) -> CodecKind;
|
||||
fn dim(&self) -> usize;
|
||||
fn encoded_len(&self) -> usize; // bytes per vector
|
||||
fn encode(&self, v: &[f32]) -> Result<Vec<u8>>;
|
||||
fn decode(&self, blob: &[u8]) -> Result<Vec<f32>>; // approximation
|
||||
fn distance(&self, a: &[u8], b: &[u8]) -> Result<f32>; // symmetric
|
||||
fn make_query(&self, q: &[f32]) -> Result<Box<dyn EncodedQuery>>;
|
||||
}
|
||||
|
||||
pub trait EncodedQuery: Send + Sync {
|
||||
fn distance_to(&self, blob: &[u8]) -> f32; // asymmetric
|
||||
}
|
||||
```
|
||||
|
||||
Storage, HNSW, DiskANN-style indexes, snapshots, WASM, and bindings consume
|
||||
`&dyn VectorCodec` / blobs — never concrete codec types. Core ships `Fp32`,
|
||||
`Fp16`, `Int8`, and `Turbo4` implementations; `Pq` and `RaBitQ1` implement the
|
||||
same traits from their own crates (core reserves the `CodecKind` variants so
|
||||
blob headers and provenance are stable). `ruvector-turboquant` becomes a
|
||||
non-optional core dependency: it is dependency-free and WASM-safe, so the
|
||||
codec plane exists on every build; only the HNSW index remains
|
||||
feature-gated.
|
||||
|
||||
### 2. Storage precision ≠ search precision
|
||||
|
||||
The active search plane composes per role (~5 bits/dim effective):
|
||||
|
||||
| Role | Representation |
|
||||
|------|----------------|
|
||||
| Source storage | Turbo4 (`D/2 + 8` B) |
|
||||
| Candidate index | RaBitQ 1-bit (`D/8` B) |
|
||||
| Rescoring | Turbo4 exact-LUT |
|
||||
| Critical verification | optional FP16 / FP32 |
|
||||
|
||||
Candidate generation traverses the graph on 1-bit codes (cheapest memory
|
||||
bandwidth), rescoring uses the Turbo4 codes, and a configurable verification
|
||||
tier re-checks results for critical policies.
|
||||
|
||||
### 3. Automatic precision selection (per query)
|
||||
|
||||
Measure query difficulty and spend bits only where needed:
|
||||
|
||||
- **Score margin** — relative gap between the last kept (k-th) and first
|
||||
dropped (k+1-th) rescored distances. Large margin ⇒ the Turbo4 result is
|
||||
already stable; return it.
|
||||
- **Candidate stability** — if an escalated pass (higher `efSearch`, larger
|
||||
rescore pool) changes the top-k membership, keep escalating; if not, stop.
|
||||
- **Critical policy** — callers can require exact (FP32 source or decode)
|
||||
verification.
|
||||
|
||||
Target: only ~5–15 % of queries take an escalation. The three-tier ladder is:
|
||||
Turbo4 answer → widened traversal (2–3× ef, 2× rescore pool) → high-precision
|
||||
verification.
|
||||
|
||||
### 4. Adaptive memory tiers (per vector, over time)
|
||||
|
||||
Promote by access frequency, demote by coldness — integrated with
|
||||
`ruvector-agent-memory` and the temporal-coherence machinery:
|
||||
|
||||
| Tier | Representation |
|
||||
|------|----------------|
|
||||
| Hot | FP16 |
|
||||
| Warm | Turbo4 |
|
||||
| Cold | PQ |
|
||||
| Archive | RaBitQ + source object reference |
|
||||
|
||||
### 5. Topology-aware precision allocation (per vector, by position)
|
||||
|
||||
Graph centrality and MinCut boundaries (via `ruvector-attn-mincut` /
|
||||
`ruvector-graph`) drive bit allocation: bridge vectors, rare concepts, and
|
||||
high-influence nodes get more bits; redundant vectors inside dense
|
||||
communities get fewer. Retrieval errors on hubs and bridges poison many
|
||||
queries; errors on redundant leaves poison almost none.
|
||||
|
||||
### 6. Streaming drift detection
|
||||
|
||||
Turbo4 is data-oblivious but its retrieval quality varies with model and
|
||||
dimension; PQ codebooks go stale outright. Track per collection: recall
|
||||
proxies (overlap between quantized and verified top-k on sampled queries),
|
||||
score distortion, embedding-norm distribution, and embedding model identity.
|
||||
Threshold failures trigger background re-encoding/migration.
|
||||
|
||||
### 7. Provenance (mandatory metadata)
|
||||
|
||||
Every stored vector carries a `VectorProvenance`:
|
||||
|
||||
```rust
|
||||
pub struct VectorProvenance {
|
||||
pub model_id: Option<String>, // embedding model identity
|
||||
pub codec: CodecKind,
|
||||
pub codec_version: u16,
|
||||
pub rotation_seed: Option<u64>,
|
||||
pub dim: usize,
|
||||
pub metric: DistanceMetric,
|
||||
pub source_hash: Option<String>, // hex sha-256 of source object
|
||||
pub lineage: Vec<String>, // migration history entries
|
||||
}
|
||||
```
|
||||
|
||||
Snapshots remain deterministic and independently verifiable through RVF: the
|
||||
same (source, provenance) always reproduces byte-identical codes (this is why
|
||||
ADR-296 forbids `rand`-derived rotations).
|
||||
|
||||
### 8. Honest benchmarking
|
||||
|
||||
Corpora: SIFT1M, GIST1M, Deep1M, DBpedia embeddings, RuFlo memory vectors,
|
||||
RuView spatial vectors. Metrics: Recall@1, Recall@10, NDCG, P50/P95/P99
|
||||
latency, build time, ingest rate, disk, RSS, energy per million queries.
|
||||
Harness lives in `ruvector-sota-bench`; results land in `bench_results/`.
|
||||
|
||||
### 9. Product surface: three policies
|
||||
|
||||
```rust
|
||||
pub enum SearchPolicy { Quality, Balanced, MaxCompression }
|
||||
```
|
||||
|
||||
Users pick an outcome, not an algorithm. Policies map to (rescore multiplier,
|
||||
escalation threshold, escalation rounds, verification tier); `Balanced` is
|
||||
the default. All lower-level knobs remain reachable for experts.
|
||||
|
||||
## Phases
|
||||
|
||||
| Phase | Scope | Status |
|
||||
|-------|-------|--------|
|
||||
| A | ADR-296 phases 1–2 (Turbo4 codec + applied HNSW integration) | done (PR #802) |
|
||||
| B | `EncodedVector`/`VectorCodec` plane in core (Fp32/Fp16/Int8/Turbo4); `VectorProvenance` schema; `SearchPolicy` + margin-based adaptive escalation v1 in the Turbo4 index | this ADR, first slice |
|
||||
| C | RaBitQ1 candidate cascade (traverse 1-bit, rescore Turbo4); PQ/RaBitQ codecs implementing the plane traits from their crates | next |
|
||||
| D | Provenance persisted in storage + RVF snapshot verification; drift monitors (recall proxy, distortion, norms, model id) | next |
|
||||
| E | Memory tiers (hot/warm/cold/archive) with promotion/demotion driven by access stats | next |
|
||||
| F | Topology-aware allocation (centrality/MinCut bit budgets) | next |
|
||||
| G | `ruvector-sota-bench` acceptance runs + **ablation** (below); NEON/AVX-512/WASM kernels from ADR-296 phase 3 | next |
|
||||
|
||||
## Acceptance test (ablation-gated)
|
||||
|
||||
Across at least three real workloads (one public benchmark, RuFlo memory, one
|
||||
embedding corpus):
|
||||
|
||||
- Adaptive mode must reduce total memory by **≥ 30 %** beyond uniform Turbo4,
|
||||
- keep Recall@10 within **0.5 pp**, and
|
||||
- keep P95 latency within **10 %** of the best fixed configuration.
|
||||
|
||||
The stated uncertainty — whether adaptive-precision complexity pays for
|
||||
itself over uniform Turbo4 — is resolved by this ablation, not argued. If the
|
||||
ablation fails, phases E/F stay experimental and the product surface ships
|
||||
uniform Turbo4 with per-query escalation only (phase B), which is already
|
||||
strictly better than a fixed pipeline.
|
||||
|
||||
## Consequences
|
||||
|
||||
- One interface ends the "N disconnected compression crates" problem; codecs
|
||||
become pluggable data, not architecture.
|
||||
- Per-query adaptivity means headline P50 latency reflects 4-bit scoring
|
||||
while tail quality is protected by escalation — the recall/latency curve
|
||||
dominates any fixed-precision point.
|
||||
- Complexity is contained: each phase lands behind the same trait plane, and
|
||||
the ablation gate prevents shipping adaptive machinery that doesn't pay.
|
||||
- Provenance-first storage makes migrations (drift, tier moves, codec
|
||||
upgrades) auditable and reversible.
|
||||
Loading…
Add table
Add a link
Reference in a new issue