PolyGraph is a graph-based index for approximate nearest-neighbor search (ANNS) over multi-vector data under flexible query weights.
Each object contains multiple vector fields, and each query specifies a weight vector over these fields. PolyGraph selects representative weights, builds one edge group for each representative weight, prunes intra- and inter-group redundant edges, and performs weight-aware greedy search (WAGS) by activating only query-relevant edge groups.
PolyGraph is the official implementation of:
PolyGraph: An Efficient Multi-Vector Index for Approximate Nearest-Neighbor Search on Multi-Vector Data
Mengtong Xu, James Pan, Guoliang Li
The paper evaluates PolyGraph on four real-world million-scale multi-vector datasets. The original data sources are listed below. For convenience, we also provide the processed files, including base vectors, query vectors, and path-info files. All base and query vectors are stored in fvecs format. Please refer here for details about fvecs.
Dataset name <DATASET> |
Paper dataset | Raw Data Source | Number of fields | Dimensions | Path-info file | Download (base and query) |
|---|---|---|---|---|---|---|
ImageText |
Image-Text | Conceptual Captions | 2 | (768, 768) |
dataset/path_info_CC1MNorm_2field_100W.txt |
ImageText_data.tar.gz |
QA2 |
Question-Answer | LMSYS-Chat-1M | 4 | (384, 512, 384, 512) |
dataset/path_info_LMSYSNorm_4field_100W.txt |
QA2_data.tar.gz |
Wiki |
Wikipedia | Wikipedia Structured Contents | 6 | (384, 384, 384, 384, 384, 384) |
dataset/path_info_EnwikiNorm_6field_100W.txt |
Wiki_data.tar.gz |
Protein |
Protein | UniProt | 8 | (400, 128, 128, 128, 128, 128, 128, 320) |
dataset/path_info_ProteinNorm_8field_100W.txt |
Protein_data.tar.gz |
After downloading the processed files, place each dataset folder under dataset/VectorFiles/, or update the corresponding path-info file with your local file paths. PolyGraph uses the path-info file to locate the base and query vector files of each field.
Each path-info file contains one block per vector field. In the
/path/to/field_1_base.fvecs
/path/to/field_1_query.fvecs
/path/to/field_2_base.fvecs
/path/to/field_2_query.fvecs
All ground-truth files are stored in ivecs format. Please refer here for details about ivecs.
In the paper, PolyGraph uses a default query workload that contains all 2^m - 1 non-empty field-participation patterns for a dataset with m fields. In this implementation, active fields are assigned equal positive weights, which is rank-equivalent to the normalized workload
dataset/Ground-truth/<DATASET>/<i>-output.ivecs
If the corresponding ground-truth file is not available, PolyGraph computes the exact results by brute force.
Download links for the ground-truth files of the default query workload are listed below:
Dataset name <DATASET> |
Paper dataset | Number of ground-truth files | Download (ground-truth) |
|---|---|---|---|
ImageText |
Image-Text | 3 | ImageText_GT.tar.gz |
QA2 |
Question-Answer | 15 | QA2_GT.tar.gz |
Wiki |
Wikipedia | 63 | Wiki_GT.tar.gz |
Protein |
Protein | 255 | Protein_GT.tar.gz |
| Parameter | In-paper notation | Default / typical value | Description |
|---|---|---|---|
-total_sim_thresh |
0.95 |
Representative-weight coverage threshold. | |
-rela_sim_thresh |
0.5 |
Related-group threshold used during construction. | |
-search_rela_sim_thresh |
0.5 |
Edge-group activation threshold used during search. | |
-R_refine |
dataset-dependent | Construction out-degree budget. | |
-L_refine |
dataset-dependent | Construction candidate-list budget. | |
k |
20 |
Number of nearest neighbors for recall evaluation. |
- CMake >= 3.10
- C++14-compatible compiler
- OpenMP
- Boost >= 1.55
- Python 3
Clone the repository, create a Python virtual environment, install the Python dependencies, and compile PolyGraph with:
# Clone the repository
git clone https://github.com/TsinghuaDatabaseGroup/PolyGraph.git
cd PolyGraph
# Create a Python virtual environment and install the Python dependencies
python3 -m venv pythonEnv_ForPG
source pythonEnv_ForPG/bin/activate
pip install -r include/python_file/requirements.txt
# Create output directories
mkdir -p build myIndex include/python_file/nohup_logs include/python_file/backup_clusterGroups
# Compile
cd build
cmake ..
make -jRun the following command from the build/ directory to construct a PolyGraph index with the default parameter values (
./main PolyGraph <DATASET> buildTo specify different values for tau and c, use:
./main PolyGraph <DATASET> build -total_sim_thresh <tau> -rela_sim_thresh <c>After building the index, run the following command from the build/ directory to evaluate PolyGraph with WAGS under the default query workload. By default, WAGS uses the activation threshold c = 0.5. Note that to use the default query workload, set the first line of the corresponding dataset/useWeight/useWeightEachQuery_[...].txt file to 0 if the file exists. The command evaluates all non-empty field-participation patterns and reports the recall-latency performance following the experimental setting in the paper. The output logs include query latency, Recall@20, average query path length, candidate-set statistics, and the average number of distance evaluations.
./main PolyGraph <DATASET> all_recall_search 20To specify a different WAGS activation threshold c, use:
./main PolyGraph <DATASET> all_recall_search 20 -search_rela_sim_thresh <c>The repository also includes several built-in baselines used in the paper.
# Build an index
./main <INDEX_NAME> <DATASET> build
# Evaluate Recall@k
./main <INDEX_NAME> <DATASET> all_recall_search <k>Command name <INDEX_NAME> |
Paper name |
|---|---|
hnsw_fusion |
HNSW_Fusion |
vamana_fusion |
Vamana_Fusion |
vamana_equNoTotal |
Vamana_Union |
vamana_allWeight |
Vamana_Poly |
vamana_oracle |
Oracle_Vamana |
DEG and HJG are external baselines and should be run using their official repositories.