SIMD / HNSW / AVX State-of-the-Art Audit
September 2, 2026 · View on GitHub
Date: 2026-02-24
Scope: velesdb-core SIMD kernels and HNSW implementation.
Executive conclusion
Short answer: partially yes.
- The codebase contains several modern optimizations (runtime SIMD dispatch, AVX2/AVX-512 kernels with FMA, masked tails, prefetch, fused cosine paths, and HNSW diversification hooks).
- But it is not yet fully state of the art versus current ANN leaders (Faiss/ScaNN/DiskANN/SPANN style stacks), mainly due to missing algorithmic layers and some implementation bottlenecks.
What is already strong
-
SIMD architecture-aware dispatch
- Runtime feature detection with cached dispatch (
OnceLock) and ISA-specific routing (AVX-512, AVX2+FMA, NEON, scalar). - Hot paths are inlined and specialized by vector length thresholds.
- Runtime feature detection with cached dispatch (
-
x86 intrinsics quality
- AVX2 and AVX-512 kernels implement multi-accumulator patterns for ILP.
- Masked-tail AVX-512 handling avoids scalar remainder loops.
- Fused cosine variants are present for large vectors.
-
HNSW search-side optimizations
- Standard top-down greedy descent + bottom-layer ef search.
- Multi-entry probing is implemented for harder queries.
- Software prefetch is used during neighbor expansion for larger dimensions.
-
Graph diversification support
- Neighbor selection includes an alpha-based diversification criterion (VAMANA-style idea).
Gaps vs state-of-the-art (priority order)
-
Not a full modern ANN stack yet
- No IVF/IMI layer, no graph-on-disk (DiskANN-style), no compressed-routing tiers, and no multi-stage candidate generation pipeline.
- Current design is primarily HNSW(+quantization), which is strong but not the full SOTA envelope for very large-scale / memory-constrained serving.
-
Native module maturity ambiguity
index/hnsw/native/mod.rsstill labels itself as PROTOTYPE, which conflicts with “production-grade” positioning and makes maturity status unclear.
-
Insertion path has avoidable vector cloning
get_vector()clones vectors (Vec<f32>) and is called repeatedly in insert/neighbor logic.- This creates allocation/copy overhead and reduces competitiveness during high-throughput builds.
-
Lock granularity / contention risk
- Core data structures (
vectors,layers) rely heavily onRwLockwith repeated lock/unlock phases. - At high concurrency, this is likely behind specialized lock-free or sharded-graph implementations.
- Core data structures (
-
SIMD coverage can be expanded
- AVX-512 detection currently hinges on
avx512f; no explicit exploitation of optional extensions (e.g., AVX-512 VNNI/BW/VL variants) for additional kernels. - No FP16/BF16 path for reduced bandwidth compute on capable hardware.
- AVX-512 detection currently hinges on
Recommendation roadmap
P0 (high impact, low/medium risk)
- Replace clone-heavy
get_vector()usage in native HNSW hot paths with borrowed/snapshot-based access. - Profile and rebalance lock scope in insert/search (especially neighbor update paths).
- Clarify production status in docs/comments for native HNSW module.
P1 (high impact)
- Add optional two-stage ANN mode (coarse quantizer/partition + HNSW rerank).
- Add stronger adaptive query-time policies (
ef_search, probes, rerank depth) based on latency SLO and query difficulty.
P2 (hardware frontier)
- Add additional ISA kernels where supported (e.g., FP16/BF16/vectorized quantized-dot paths).
- Consider dimension-specialized codegen (common dims: 384/768/1024/1536) for peak throughput.
Bottom line
- For a general-purpose Rust vector DB, the current SIMD and HNSW code is solid and modern.
- Relative to absolute SOTA in 2025+ ANN systems, it is one step short: excellent kernels and graph baseline, but missing some large-scale system-level algorithmic layers and a few hot-path engineering refinements.
Progress update (after follow-up implementation rounds)
Implemented from roadmap
-
P0 completed (core items):
- Clone-heavy vector retrieval paths in native HNSW were removed/refactored to snapshot/borrow patterns.
- Lock handling and lock-rank instrumentation were centralized with helper-based read snapshots.
- In-place neighbor mutation helpers were added to reduce read-clone-write churn and duplicate-edge risk.
- Native module status language was clarified away from prototype ambiguity.
-
P1 partially completed:
- Two-stage ANN behavior (candidate generation + exact SIMD rerank) has been added behind adaptive policy in
HnswIndex. - Query-time adaptation now considers quality profile,
ef_search,k, and dataset scale.
- Two-stage ANN behavior (candidate generation + exact SIMD rerank) has been added behind adaptive policy in
-
P2 partially completed:
- Runtime visibility for AVX-512 optional extensions (
VL,BW,VNNI) has been added.
- Runtime visibility for AVX-512 optional extensions (
Still open
- Full IVF/partitioned coarse stage (true IVF/IMI) is still not present.
- Disk-backed graph/search path (DiskANN-style) is still not present.
- FP16/BF16 compute paths and dimension-specialized kernel generation are still future work.
- A latency-aware rerank controller is in place (
rerank_latency_target_us+adapt_rerank_k_to_latency); a full benchmark-driven, recall-aware SLO controller is still future work.
Last updated: 2026-07-25 · Applies to: velesdb-core 6.0.0