This repository contains the implementation of FSST-LIKE-Matching, a high-performance string matching engine published at DaMoN 2026. This work is built upon the FSST compression algorithm; the link to the official repository is https://github.com/cwida/fsst.
The algorithms in the repository provide efficient pattern matching (specifically LIKE patterns) over FSST-compressed data. By avoiding decompression, we achieve high throughput for analytical workloads.
The project is organized into several key directories:
include/: Header files defining the core automata, codegen interfaces, and string search utilities.src/: Implementation of the pattern matching engine, including LLVM and C++ code generation.benchmark/: Performance measurement suite for datasets including IMDB, StackOverflow, and TPC-H.test/: Unit and integration tests for various matching scenarios (start, middle, end, and full pattern matching).fa-drawing/: A visualization tool (web-based) to draw and inspect the finite automata used in the matching process.
- Automata-based Matching: Uses specialized finite automata for
LIKEpattern evaluation. - Codegen Backend: Supports both LLVM-based JIT compilation and C++ source generation for matching kernels.
- FSST Integration: Direct matching on compressed data without full decompression.
- Multithreaded Benchmarking: Tools to measure throughput across multiple CPU cores.
- CMake (3.5+)
- LLVM16 (for LLVM codegen backend)
- Vectorscan / Hyperscan (for hybrid search support)
- C++20 compatible compiler
mkdir build && cd build
cmake ..
make -j$(nproc)To use the browser client, run the following command:
cd fa-drawing
./run_server.shAfter starting the server, open up index.html in a browser. You must provide three values in order to generate the automaton correctly:
- Pattern: The
LIKEpattern to generate the automaton for. - Symbol Table Path: The path to the binary of the symbol tables relative to the main folder of the repository; example symbol tables are provided in
data/. - Automaton Type: One of four possible values:
- start: For prefixes (the
%wildcard is implicit at the end). - middle: For substrings (the
%wildcard is implicit at both ends). - end: For suffixes (the
%wildcard is implicit at the start). - full: May contain multiple
LIKEsubpatterns and multiple%wildcards.
- start: For prefixes (the
- Compress File Path: If you want to generate a symbol table for a custom dataset, enter the path relative to the main folder of the repository. The format of the file is: one string entry per line.
Once the datasets have been installed, you can run the benchmark scripts by running the following commands:
cd benchmark
./run_benchmarks.sh