AndyPook/SparseFacetedSearch

Memory efficient faceted search for Lucene.Net

★ 0Forks 0C#GitHub ↗Compare

README

SparseFacetedSearch

SparseFacetedSearch is a faceted searcher based on SimpleFacetedSearch. It uses DocID lists instead of bitmaps. Efficient memory usage for high cardinality sparsely populated facets.

Suitable for high cardinality, sparsely populated facets. i.e. There are a large number of facet values and each facet value is hit in a small percentage of documents. SimpleFacetedSearch holds a bitmap for each value representing whether that value is a hit in each document (approx 122KB per 1M documents per facet value).

Memory requirements

Given the following:

  • equation, the set of unique facet values
  • equation, the number of unique facet values
  • equation number of hits for value equation (i.e. the no. of documents the value appears in)
  • d = total number of documents

then we have memory requirement given by:

fn(Simple): equation bytes - hence memory increases as the product of documents * values

fn(Sparse): equation bytes - hence memory increases in relation to the number of hits only.

So, if every document has exactly one value (e.g. a product category):

  • the sum of all the hit counts is just the total number of documents, and thus the memory requirement is given by fn(Sparse) = 4d bytes.
  • the point at which the two methods require equal memory usage is hence when equation
  • since fn(Simple) equation, for equation, SparseFS requires less memory

More generally, the point where memory usage is equal for both methods is when n = 32 * hit-ratio, where hit-ratio is the average number of hits per document (across all values). Hence, if you have an average of 4 values per document then the break even point is 128 values. Above this number, SparseFS requires less memory than SimpleFS.

Contributors

AndyPookkntajustjrobinson

Issues