Implementation of the PANDA join - note that this library implements the DDR join evaluation and excludes the CQ join evaluation included at the end of the paper.
My notes on the paper (covering the background and my intuition required for implementing the algorithm) are linked below
- PANDA Part 1: Preliminaries and Degree
- PANDA Part 2: The Core Shannon Inequality
- PANDA Part 3: The Core Algorithm
- create an application that takes a dependency on the
//src:panda_liblibrary - construct a PANDA subproblem (the spec file section below illustrates how this is done for the example provided in the paper) - include from here
- call
generate_ddr_feasible_outputto generate a feasible output for the passed PANDA subproblem - include from here
The PANDA subproblem may be directly parsed from a YAML specification file as illustrated in the example taken from the paper here - the specification file format used by the example defines the arguments to a PANDA subproblem. The example specification file is provided here. The file format requires the following fields:
GlobalSchema- the list of human-readable attribute/column names covering all input tables - note that the order of the names must correspond exactly with the type parameterization of the PANDA subproblemTables- a triple for each table in the input schema containing- the name of the table CSV file from which the table will be loaded (T)
- the degree constraint on the table (n)
- the corresponding w parameter in the Shannon inequality
OutputAttributes- a list of column name lists - each list corresponds to a relation in the output schema (Z)M- a list of monotonicity witnesses - each monotonicity contains- a list of column names - the conditioned attributes Y in Y|X
- a list of column names - the condition attributes X in Y|X
- a positive integer multiplicity of the monotonicity
S- a list of submodularity witnesses - each submodularity contains- a list of column names - the conditioned attributes Y in Y;Z|X
- a list of column names - the conditioned attributes Z in Y;Z|X
- a list of column names - the condition attributes X in Y;Z|X