SPEC: SPHERE dataset provider — streaming import at TB scale

#21 · closed · 1 comments

View on GitHub ↗

redknightlois

Part of epic: #20 (Epic: Vector Datasets) Depends on: #18 (Dataset Import Pipeline) ## Inheritance - Inherits systemic invariants from #13 - **Critical for this spec: SI-07 (fail fast on empty keyspace), SI-14 (vector search requires Corax)** ## Slice boundary Included: SPHERE dataset download from Weaviate-prepared files, streaming decompression, bulk import into RavenDB with pre-computed embeddings, tiered profiles, resume support, import and indexing time measurement, recall measurement against brute-force ground truth. Excluded: Server-side embedding generation, Weaviate comparison benchmarking. ## Dataset SPHERE (Meta, CC-BY-NC 4.0) contains 134 million web documents chunked into 899 million 100-word passages. Weaviate prepared and hosts `.jsonl.tar.gz` files at multiple scales with pre-computed DPR embeddings (`facebook-dpr-ctx_encoder-single-nq-base`). Access requires authorization from Weaviate (`[email protected]`). Each JSONL line: `{"raw": "text", "sha": "hash", "title": "title", "url": "url", "vector": [...]}` | Profile | Lines | Compressed | Decompressed | |---------|-------|-----------|-------------| | 10k | 10,000 | ~40 MB | ~104 MB | | 100k | 100,000 | 403 MB | 1.04 GB | | 1m | 1,000,000 | 3.9 GB | 10.4 GB | | 10m | 10,000,000 | 39.1 GB | 104 GB | | 100m | 100,000,000 | 391 GB | 1.04 TB | | full | 899,000,000 | 3.4 TB | 9.36 TB | ## Behavior - The provider implements the dataset provider interface with profile-based database naming (e.g., `Sphere-10K`, `Sphere-100K`, `Sphere-1M`, `Sphere-Full`). Each profile uses its own database. - Import checks if the target database already exists and contains the expected document count. If so, import is skipped entirely. - Import streams directly from the compressed file: file/HTTP stream → GZip decompression → tar entry reader → line-by-line JSONL parsing → RavenDB bulk insert. No intermediate decompressed files on disk — the 9.36 TB decompressed size makes disk-based decompression infeasible. - Pre-computed DPR embeddings from the dataset are imported alongside the text fields. No server-side embedding generation is needed. - Vectors are stored as binary attachments (768D × 4 bytes = 3072 bytes per vector) rather than inline JSON, for storage efficiency. - Import tracks progress (line count, documents/sec) and supports resume by skipping already-imported lines on restart. - The provider measures and reports **import time** (wall clock from first document to last document inserted) and **indexing time** (wall clock from import completion to all indexes becoming non-stale). Both are included in the benchmark output. - The provider generates `VectorWorkloadMetadata` by sampling query vectors from the imported embeddings. - After import, the provider computes **brute-force ground truth** for the sampled query vectors: for each query vector, an exact nearest-neighbor search over all imported embeddings produces the true top-K results. This ground truth is stored in `VectorWorkloadMetadata.GroundTruth` and used to measure **recall@K** — the fraction of approximate HNSW results that appear in the true top-K. - Ground truth computation runs once after import and is cached in a document in the target database. For large profiles (10M+), ground truth is computed over a representative subset to keep computation tractable. ## Index naming convention Vector index names encode HNSW parameters when specified, enabling side-by-side comparison of different configurations: **Format:** `{Collection}/ByEmbedding[Quantization]-{engine}[-m{M}-ef{efConstruction}]` Examples: - `Passages/ByEmbedding-corax` — default HNSW parameters - `Passages/ByEmbedding-corax-m32-ef200` — M=32, efConstruction=200 - `Passages/ByEmbeddingInt8-corax-m16-ef100` — Int8 quantization with custom HNSW This is centralized in `VectorIndexNaming.GetIndexName()` and used consistently across index creation (dataset providers), query generation (benchmark runner, recall measurement), and the standalone recall command. ## Acceptance - [ ] `--dataset sphere --dataset-profile 10k` imports 10K SPHERE documents with their pre-computed vectors via streaming decompression into database `Sphere-10K` - [ ] `--dataset sphere --dataset-profile 100k` imports 100K SPHERE documents into database `Sphere-100K` - [ ] `--dataset sphere --dataset-profile 1m` imports 1M documents into database `Sphere-1M` with no intermediate decompressed file on disk - [ ] Running the same profile twice skips import — the existing database with the expected document count is detected and reused - [ ] Import resumes after interruption without re-importing already-stored documents - [ ] Import time (documents ingested) and indexing time (indexes non-stale) are measured separately and reported in the benchmark output - [ ] Vector search workload runs against the imported SPHERE data using the pre-computed DPR embeddings - [ ] The provider reports import progress (documents/sec, percentage) during long imports - [ ] Recall@K is computed by comparing approximate HNSW results against brute-force exact nearest neighbors for the sampled query vectors - [ ] Ground truth is cached in the target database and reused across benchmark runs - [ ] `--vector-edges M --vector-candidates EF` produces a distinctly named index (e.g., `Passages/ByEmbedding-corax-m32-ef200`) that coexists with default-parameter indexes - [ ] Standalone `recall` command measures recall@K without running a throughput benchmark

Comments

redknightlois

## Spec update: vectors as attachments The current implementation stores vectors inline as JSON arrays in the document body. At SPHERE scale (768D × 899M docs), this bloats document size and hurts storage/query performance. **Required change**: Store pre-computed DPR vectors as binary attachments (name: `"vector"`), not inline JSON fields. The index map becomes: ``` from d in docs.SphereDocuments let attachment = LoadAttachment(d, "vector") select new { d.Title, Embedding = CreateVector(attachment.GetContentAsStream()) } ``` BulkInsert supports this via `bulkInsert.AttachmentsFor(docId).Store("vector", stream)`. This applies to both SPHERE and ClinicalWords providers.