This repository contains the C++ experimental evaluation framework for the paper:
OnPair is a compression algorithm specifically designed for workloads requiring fast random access to individual strings in large collections. This benchmark suite provides comprehensive performance evaluation tools to compare OnPair against established compression methods.
For the standalone OnPair algorithm implementation, see: onpair_cpp
git clone --recurse-submodules https://github.com/gargiulofrancesco/compression_benchmark_cpp.git
cd compression_benchmark_cppTo run the benchmark, you must install the required libraries using vcpkg:
git clone https://github.com/microsoft/vcpkg.git
cd vcpkg
./bootstrap-vcpkg.sh
./vcpkg install simdjson brotli zlib lz4 liblzma zstd snappy# Create build directory
mkdir build
cd build
# Configure with CMake (adjust vcpkg path as needed)
cmake -DCMAKE_TOOLCHAIN_FILE=/path/to/vcpkg/scripts/buildsystems/vcpkg.cmake -DCMAKE_BUILD_TYPE=Release ..
# Build with release optimizations
cmake --build . --config ReleaseTo reproduce the experiments presented in the paper, use the provided Python script to download and process the standard datasets:
# Create a virtual environment
python -m venv venv
source venv/bin/activate
# Install dependencies
pip install -r scripts/requirements.txt
# Download and process datasets
python scripts/process_datasets.pyThe script will download and process the datasets into the JSON format required by the benchmark suite. The datasets will be saved in the data/ directory.
Evaluate a specific algorithm on a dataset:
./benchmark_individual <dataset.json> <algorithm> <output.json> [core_id]Example:
# Run onpair algorithm on example dataset with CPU core pinning
./benchmark_individual data/example.json onpair results.json 0Run all algorithms on all datasets in a directory:
./benchmark_all <dataset_directory> [core_id]Examples:
# Run complete benchmark suite with CPU core pinning
./benchmark_all data/ 0This generates a comprehensive performance comparison across all algorithms and datasets.
| Algorithm | Description |
|---|---|
raw |
Uncompressed baseline |
brotli |
Google's Brotli compression |
deflate |
DEFLATE compression algorithm |
lz4 |
LZ4 fast compression |
snappy |
Google's Snappy compression |
xz |
XZ compression |
zstd |
Facebook's Zstandard compression |
fsst |
Fast Static Symbol Table compression |
onpair |
OnPair algorithm |
Datasets must be JSON arrays of strings:
[
"user_12345",
"admin_67890",
"guest_11111",
"user_54321"
]The benchmark suite evaluates algorithms across four key dimensions:
| Metric | Description | Units |
|---|---|---|
| Compression Ratio | original_size / compressed_size |
Ratio |
| Compression Speed | Throughput during compression | MiB/s |
| Decompression Speed | Throughput during full decompression | MiB/s |
| Random Access Time | Average time per individual string access | nanoseconds |
Output Format: Results are exported as structured JSON for easy analysis and visualization.
This project is licensed under the MIT License - see the LICENSE file for details.
- Francesco Gargiulo - [email protected]
- Rossano Venturini - [email protected]
University of Pisa, Italy