randyzwitch/dataframe_mojo

★ 0Forks 0MojoGitHub ↗Compare

README

dataframe_mojo

An experimental, native CPU dataframe library for Mojo 1.2. Import it as dataframe. Storage and computation use Mojo and its standard library; there is no Python, pandas, Polars, or Arrow runtime dependency.

This is a working first implementation, not a production engine or a portable backend facade. The API is provisional.

Install

Add it to another Pixi workspace straight from GitHub:

[workspace]
preview = ["pixi-build"]

[dependencies]
mojo = ">=1.2,<1.3"
dataframe_mojo = { git = "https://github.com/randyzwitch/dataframe_mojo.git", tag = "v0.2.0" }
from dataframe import DataFrame, Series, Column, col

A compiled Mojo package only loads in the Mojo version that produced it, and pixi builds this package with the newest Mojo that [package.host-dependencies] in pixi.toml allows, so each tag targets one Mojo version. v0.2.0 targets 1.2; v0.1.3 is the last tag for 1.1.

Run

With Pixi installed, from this directory:

pixi run test          # runs every tests/test_*.mojo, in parallel
pixi run example
pixi run build
pixi run bench          # CPU benchmark suite; see the [benchmark guide](https://github.com/randyzwitch/dataframe_mojo/wiki/benchmarks)
pixi run bench-csv
./build/sales

Mojo 1.2 is required from 0.2.0. It is a nightly series today, so the nightly channel is one of the workspace channels and pixi install resolves it; a compiled package loads only in the Mojo version that built it, so a consumer on 1.1 stays on 0.1.3.

Supported platforms are Linux x86-64 and macOS arm64 (both tested in CI) and Linux aarch64 (resolved in pixi.lock, not yet tested in CI). Windows waits on Mojo support. The library requires a little-endian target and checks this at compile time. The original sales example prints:

east 90.0
west 180.0

Expressions

Expressions describe work without executing it. select, with_columns, filter, and group_by(...).agg(...) bind them against the input schema before running any expression kernel. The eager evaluator operates on bounded batches.

from dataframe import Column, DataFrame, Series, col, lit


def main() raises:
    var sales = DataFrame([
        Series("region", Column[String](["east", "west", "east", "west", "north"])),
        Series("amount", Column[Float64](
            [100, 80, -20, 120, 999],
            [True, True, True, True, False],
        )),
    ])
    var result = sales.filter(
        col("amount") > 0
    ).with_columns(
        (col("amount") * 0.9).alias("net")
    ).group_by("region", maintain_order=True).agg([
        col("net").sum().alias("revenue"),
        col("net").count().alias("sales"),
    ])
    print(result)
shape: (2, 3)
┌────────┬─────────┬───────┐
│ region ┆ revenue ┆ sales │
│ ---    ┆ ---     ┆ ---   │
│ str    ┆ f64     ┆ i64   │
╞════════╪═════════╪═══════╡
│ east   ┆ 90.0    ┆ 1     │
│ west   ┆ 180.0   ┆ 2     │
└────────┴─────────┴───────┘

Run pixi run example to see the sales, expression, and CSV pipelines. build/expressions and build/read_csv are the compiled examples.

CSV ingestion

read_csv reads local UTF-8 CSV files without Python or another dataframe runtime. read_csv(path) infers a schema from a sample (identifiers such as 007 and integers beyond Int64 stay strings); an explicit schema pins every type:

from dataframe import CsvField, CsvSchema, read_csv

var frame = read_csv(
    "sales.csv",
    CsvSchema([
        CsvField.string("region", False),
        CsvField.float64("amount"),
    ]),
)

It accepts LF or CRLF records, an optional UTF-8 BOM, quoted commas/newlines, and doubled quotes. Empty unquoted fields are null; quoted empty strings are values. Headers must exactly match the schema. Whitespace is preserved and typed fields do not trim it. Boolean values are exactly true or false. Float64 accepts the standard parser's values, including nan, and only explicit inf, +inf, -inf, Infinity, +Infinity, and -Infinity may be infinite.

write_csv(frame, path) writes the inverse format: reading it back with CsvSchema.of(frame) reproduces the frame exactly.

Parquet reading and writing

read_parquet(path) reads a local Parquet file; columns=[...] selects fields in the order given and row_groups=[...] selects row groups. Integer and float widths, strings, booleans, dates and timestamps carry over with their nulls. Dictionary columns arrive as plain strings, float16 as float32, and a timestamp keeps its time zone. Lists and structs map recursively. Decimal128 columns map to DataType.decimal(precision, scale); binary columns map to DataType.BINARY.

from dataframe import read_parquet, scan_parquet, col, lit

var frame = read_parquet("sales.parquet", columns=["region", "amount"])
var recent = (
    scan_parquet("sales.parquet")
    .filter(col("amount") > lit(1000.0))
    .select(["region", "amount"])
    .collect()
)

scan_parquet defers the read to collect(): only the columns the plan uses are decoded, and a filter with constant bounds reads the footer's row-group statistics first and skips row groups that cannot match. parquet_row_group_statistics(path) returns those bounds as a frame.

write_parquet(frame, "sales.parquet", compression="zstd", row_group_size=1_000_000) writes a local file, replacing an existing file at that path. Compression can be zstd (the default), snappy or uncompressed. All supported column types, including nested fields and temporal units, round-trip through Arrow schema metadata. Rebuild libdfparquet to add the writer entry point. See the Parquet guide for behavior and validation.

List and struct columns

DataType.list(inner) holds a variable number of values per row and DataType.struct(names, dtypes) holds one value of each named field per row, in the Arrow large_list and struct layouts. They come out of read_parquet and Arrow import, from str.split, and from pack_struct.

from dataframe import DataFrame, col

var tags = frame.with_columns(col("tags").str().split(",").alias("tag"))
var one_per_tag = tags.explode("tag")           # other columns repeat
var counts = tags.select_exprs([col("tag").list().len().alias("n")])
var packed = frame.pack_struct("point", ["x", "y"])
var xs = packed.select_exprs([col("point").field("x")])
var flat = packed.unnest("point")

The .list() namespace has len, get(i) (negative from the end, null when out of range), first, last, contains(value), join(separator), sum, min, max and mean; field(name) reads one struct field, null where the struct is null. col(x).implode() gathers a group's values into one list inside agg (or a whole column into one row), and as_struct([...]) packs expressions into a struct column. Struct columns can be group_by, unique and inner/left/semi/anti join keys; they compare field by field, with a null struct distinct from a struct of nulls. explode and unnest also exist on LazyFrame. Nested columns take part in select, filter, take, slice, concat, head, equals, display, and Arrow export and import; they cannot yet be sort keys, list columns cannot be keys at all, and neither can be cast, reduced (other than implode), used in when/then, or written to CSV. Each of those raises a clear error.

The reader is Arrow C++'s, built with only its Parquet parts and bundled codecs into libdfparquet, a 15 MB shared library with no dependencies beyond libc. It is loaded at run time, so nothing else in the package needs it. pixi run -e native build-dfparquet builds it into build/dfparquet/ (a C++20 compiler, cmake and ninja are needed); DATAFRAME_PARQUET_LIBRARY points read_parquet at a library elsewhere. See native/dfparquet/ and dataframe/parquet.mojo for the boundary a Mojo-native reader would replace.

The reader consumes bounded file buffers and retains tokenizer state across them. Its scalar structural scanner is the correctness reference for later SIMD scanning and record-boundary-aware parallel decoding. See the complete CSV contract.

Use typed literals: there is no implicit Int64/Float64 promotion. Expressions support + - * / // % **, comparisons < <= > >= plus .eq()/.ne(), unary -, abs, sqrt, exp, log, floor, ceil, round, clip, .alias(), Kleene & | ^ ~, is_null, is_nan, fill_null, fill_nan, coalesce, selectors (all, col([...]), exclude, by_dtype, nth), window operations (cum_sum, shift, diff, rank, rolling_*, forward_fill, interpolate, interpolate_by, cut, qcut, over), is_in, is_between, cast, when(...).then(...).otherwise(...), a .str() namespace with concat_str, and the reductions sum(min_count=0), count, null_count, len, min, max, mean, first, last, n_unique, var, std, median, quantile, any, all, arg_min, arg_max, mode, value_counts, skew and kurtosis, plus the two-column corr and cov. Equality is .eq() rather than Python-style ==. See the operator table in expressions.

select(expr) selects one expression; select_exprs([expr, ...]) selects several. select(["name", ...]) remains the name-only projection API, with no ambiguous empty-list overload. with_columns and agg accept an expression or a list. Sibling expressions all see the original input, not each other's aliases.

Expressions distinguish scalar, row-valued, and aggregate results. For example:

var totals = sales.select(col("amount").sum())  # one row
var repeated = sales.with_columns(col("amount").sum().alias("total"))
var deviations = sales.select(col("amount") - col("amount").sum())

Expression sums return zero for empty/all-null input, or null when the valid count is below min_count, following Narwhals' default. Integer sums accumulate exactly in 128 bits and check the final Int64 result, allowing parallel partial-state merging. Floating-point expression sums permit reassociation; bitwise reproducibility is not guaranteed. See the expression contract.

Implemented

  • Typed Column[T] with bit-packed validity and checked element access; Booleans and strings use the Arrow bit-packed and UTF-8 layouts.
  • Heterogeneous Series and runtime-schema DataFrame: Int8–Int64, UInt8–UInt64, Float32, Float64, Bool, String, Date, Datetime, Duration, and Time. The runtime tag is per column, not per cell.
  • Schema inspection, projection, indexed gathering, nullable Boolean filtering, and adding/replacing columns.
  • Vertical, diagonal, and horizontal concat, plus vstack/hstack.
  • unique, n_unique, is_duplicated, drop_nulls, and frame fill_null.
  • pivot and unpivot reshaping.
  • A Series API (operators, reductions, value_counts, unique, sort) backed by the same expression kernels.
  • Lazy queries (frame.lazy(), scan_csv) with predicate, projection, and slice pushdown and explain(); see lazy queries.
  • Bounded table display for DataFrame and Series (print(frame), to_string(max_rows=..., ...), glimpse()).
  • head/tail/slice/reverse, drop/rename/with_row_index, row and cell access through the tagged AnyValue, null counts, and structural equals.
  • Hash grouping by any number of columns or key expressions of any dtype.
  • Stable multi-column sorting with per-column direction and null placement, plus arg_sort, top_k, and bottom_k.
  • Hash joins on any number of keys of any dtype: inner, left, right, full, semi, anti, and cross, with left_on/right_on.
  • Expression IR, schema binding, scalar broadcasting, and grouped aggregates.
  • Explicit SIMD Float64 arithmetic/comparison kernels and mergeable reduction states.
  • Contract tests for validation, nulls, overflow, ordering, joins, empty shapes, and scalar/SIMD agreement across batch boundaries.

Useful calls:

var selected = frame.select(["region", "amount"])
var first = frame.head(3)
var cleaned = frame.drop("notes").rename({"amount": "revenue"})
var cell = frame.item(0, "region").string()
var sampled = frame.take([3, 0, 3])
var sorted = frame.sort(["region", "amount"], descending=[False, True], nulls_last=[True, True])
var joined = frame.join(regions, on="region", how="left")

Operations return new dataframes. column() and typed extraction methods (int64(), float64(), bool(), and string(), which returns a StringColumn) return immutable columns that share buffers with the frame; no copy is made. An empty dataframe can retain its row count: DataFrame([], height=10).

Deliberate limits

Execution is eager. Reductions, row-wise expressions, and filters run on worker threads for large inputs (DATAFRAME_THREADS sets the limit; 1 disables); hash grouping, joins, and sorting are single-threaded for now. Float64 expression arithmetic and comparisons use SIMD; checked integer arithmetic and reductions are currently scalar. Projections, slices, and column extraction share immutable buffers (zero-copy); computed results allocate new buffers, and unfused intermediates are still materialized per batch. String columns use the Arrow large_utf8 layout (one UTF-8 buffer plus Int64 offsets). Frames export to and import from the Arrow C Data Interface (see Arrow interchange); most dtypes export zero-copy. There are no performance claims yet. The underscore-prefixed fields are internal and must not be mutated by callers.

There is no GPU execution, index alignment, implicit dtype coercion, or dataframe backend adapter. Build input columns in memory for this first version.

Using the package

pixi run package precompiles dist/dataframe.mojoc; put its directory on the import path (mojo run -I dist app.mojo). Tagged releases attach the package and a generated API reference (pixi run docs). See the testing guide, the changelog, the stability policy, and the pandas/Polars migration guide.

See semantics for the behavior the tests promise and architecture and next steps for the intended progression. Historical benchmarks and one-off experiments are in the research wiki.

Contributors

randyzwitch

Issues