Epic: Vector Datasets

#20 · open · 0 comments

View on GitHub ↗

redknightlois

## Goal Large-scale vector dataset providers for benchmarking RavenDB's vector search (HNSW) at realistic scale — from hundreds of thousands to hundreds of millions of vectors. ## Global invariants - This epic inherits the systemic architectural invariants from #13. - **Critical for this epic: SI-07 (fail fast on empty keyspace), SI-14 (vector search requires Corax), SI-12 (atomic key generation)** ## Related epics - #18 Dataset Import Pipeline — vector providers extend the dataset provider framework - #16 Workloads and Key Distribution — vector workloads consume metadata (query vectors, dimensions, ground truth) produced by these providers - #14 Load Generation Engine — the runner imports datasets before benchmarking begins ## Architectural invariants - Vector dataset providers produce `VectorWorkloadMetadata` containing query vectors, field name, dimensions, base vector count, and optional ground truth for recall computation. - Embeddings are always pre-computed in the dataset. The benchmark measures RavenDB's import and HNSW indexing performance, not embedding computation. - The existing ClinicalWords provider loads all vectors into memory. This works at 100K–200K scale but does not extend to datasets with millions or billions of entries. - Large-scale vector datasets (10M+ entries) require streaming import — on-the-fly decompression from the compressed source into RavenDB bulk insert, with no intermediate decompressed files on disk. - Import supports resume after failure. At billion-entry scale, failures during import are expected, not exceptional. - Dataset providers support tiered profiles so the same dataset can be used at different scales (e.g., 100K for smoke tests, 10M for CI, 899M for full benchmarks). - Import time and indexing time are measured separately and reported in the benchmark output. - **Vector index names encode HNSW parameters (M, efConstruction) when specified**, enabling side-by-side comparison of different HNSW configurations on the same dataset. Format: `{Collection}/ByEmbedding[Quantization]-{engine}[-m{M}-ef{efConstruction}]`. This ensures distinct indexes coexist without collision, and benchmark results can be correlated to specific HNSW tuning. ## Spec issues - [ ] #21 SPEC: SPHERE dataset provider — streaming import at TB scale

Comments