SparseFacetedSearch is a faceted searcher based on SimpleFacetedSearch. It uses DocID lists instead of bitmaps. Efficient memory usage for high cardinality sparsely populated facets.
Suitable for high cardinality, sparsely populated facets. i.e. There are a large number of facet values and each facet value is hit in a small percentage of documents. SimpleFacetedSearch holds a bitmap for each value representing whether that value is a hit in each document (approx 122KB per 1M documents per facet value).
Given the following:
, the set of unique facet values
, the number of unique facet values
number of hits for value
(i.e. the no. of documents the value appears in)
- d = total number of documents
then we have memory requirement given by:
fn(Simple):
bytes - hence memory increases as the product of documents * values
fn(Sparse):
bytes - hence memory increases in relation to the number of hits only.
So, if every document has exactly one value (e.g. a product category):
- the sum of all the hit counts is just the total number of documents, and thus the memory requirement is given by fn(Sparse) = 4d bytes.
- the point at which the two methods require equal memory usage is hence when
- since fn(Simple)
, for
, SparseFS requires less memory
More generally, the point where memory usage is equal for both methods is when n = 32 * hit-ratio, where hit-ratio is the average number of hits per document (across all values). Hence, if you have an average of 4 values per document then the break even point is 128 values. Above this number, SparseFS requires less memory than SimpleFS.