An experimental, native CPU dataframe library for Mojo 1.2. Import it as
dataframe. Storage and computation use Mojo and its standard library; there
is no Python, pandas, Polars, or Arrow runtime dependency.
This is a working first implementation, not a production engine or a portable backend facade. The API is provisional.
Add it to another Pixi workspace straight from GitHub:
[workspace]
preview = ["pixi-build"]
[dependencies]
mojo = ">=1.2,<1.3"
dataframe_mojo = { git = "https://github.com/randyzwitch/dataframe_mojo.git", tag = "v0.2.0" }from dataframe import DataFrame, Series, Column, colA compiled Mojo package only loads in the Mojo version that produced it, and
pixi builds this package with the newest Mojo that [package.host-dependencies]
in pixi.toml allows, so each tag targets one Mojo version. v0.2.0 targets
1.2; v0.1.3 is the last tag for 1.1.
With Pixi installed, from this directory:
pixi run test # runs every tests/test_*.mojo, in parallel
pixi run example
pixi run build
pixi run bench # CPU benchmark suite; see the [benchmark guide](https://github.com/randyzwitch/dataframe_mojo/wiki/benchmarks)
pixi run bench-csv
./build/sales
Mojo 1.2 is required from 0.2.0. It is a nightly series today, so the
nightly channel is one of the workspace channels and pixi install
resolves it; a compiled package loads only in the Mojo version that built
it, so a consumer on 1.1 stays on 0.1.3.
Supported platforms are Linux x86-64 and macOS arm64 (both tested in CI) and
Linux aarch64 (resolved in pixi.lock, not yet tested in CI). Windows waits on
Mojo support. The library requires a little-endian target and checks this at
compile time.
The original sales example prints:
east 90.0
west 180.0
Expressions describe work without executing it. select, with_columns,
filter, and group_by(...).agg(...) bind them against the input schema before
running any expression kernel. The eager evaluator operates on bounded batches.
from dataframe import Column, DataFrame, Series, col, lit
def main() raises:
var sales = DataFrame([
Series("region", Column[String](["east", "west", "east", "west", "north"])),
Series("amount", Column[Float64](
[100, 80, -20, 120, 999],
[True, True, True, True, False],
)),
])
var result = sales.filter(
col("amount") > 0
).with_columns(
(col("amount") * 0.9).alias("net")
).group_by("region", maintain_order=True).agg([
col("net").sum().alias("revenue"),
col("net").count().alias("sales"),
])
print(result)shape: (2, 3)
┌────────┬─────────┬───────┐
│ region ┆ revenue ┆ sales │
│ --- ┆ --- ┆ --- │
│ str ┆ f64 ┆ i64 │
╞════════╪═════════╪═══════╡
│ east ┆ 90.0 ┆ 1 │
│ west ┆ 180.0 ┆ 2 │
└────────┴─────────┴───────┘
Run pixi run example to see the sales, expression, and CSV pipelines. build/expressions and build/read_csv are the
compiled examples.
read_csv reads local UTF-8 CSV files without Python or another dataframe
runtime. read_csv(path) infers a schema from a sample (identifiers such as
007 and integers beyond Int64 stay strings); an explicit schema pins every
type:
from dataframe import CsvField, CsvSchema, read_csv
var frame = read_csv(
"sales.csv",
CsvSchema([
CsvField.string("region", False),
CsvField.float64("amount"),
]),
)It accepts LF or CRLF records, an optional UTF-8 BOM, quoted commas/newlines,
and doubled quotes. Empty unquoted fields are null; quoted empty strings are
values. Headers must exactly match the schema. Whitespace is preserved and
typed fields do not trim it. Boolean values are exactly true or false.
Float64 accepts the standard parser's values, including nan, and only explicit
inf, +inf, -inf, Infinity, +Infinity, and -Infinity may be infinite.
write_csv(frame, path) writes the inverse format: reading it back with
CsvSchema.of(frame) reproduces the frame exactly.
read_parquet(path) reads a local Parquet file; columns=[...] selects
fields in the order given and row_groups=[...] selects row groups.
Integer and float widths, strings, booleans, dates and timestamps carry
over with their nulls. Dictionary columns arrive as plain strings, float16
as float32, and a timestamp keeps its time zone. Lists and structs map recursively. Decimal128 columns map to DataType.decimal(precision, scale); binary columns map to DataType.BINARY.
from dataframe import read_parquet, scan_parquet, col, lit
var frame = read_parquet("sales.parquet", columns=["region", "amount"])
var recent = (
scan_parquet("sales.parquet")
.filter(col("amount") > lit(1000.0))
.select(["region", "amount"])
.collect()
)scan_parquet defers the read to collect(): only the columns the plan
uses are decoded, and a filter with constant bounds reads the footer's
row-group statistics first and skips row groups that cannot match.
parquet_row_group_statistics(path) returns those bounds as a frame.
write_parquet(frame, "sales.parquet", compression="zstd", row_group_size=1_000_000)
writes a local file, replacing an existing file at that path. Compression can
be zstd (the default), snappy or uncompressed. All supported column types,
including nested fields and temporal units, round-trip through Arrow schema
metadata. Rebuild libdfparquet to add the writer entry point. See the
Parquet guide for behavior and validation.
DataType.list(inner) holds a variable number of values per row and
DataType.struct(names, dtypes) holds one value of each named field per row,
in the Arrow large_list and struct layouts. They come out of
read_parquet and Arrow import, from str.split, and from pack_struct.
from dataframe import DataFrame, col
var tags = frame.with_columns(col("tags").str().split(",").alias("tag"))
var one_per_tag = tags.explode("tag") # other columns repeat
var counts = tags.select_exprs([col("tag").list().len().alias("n")])
var packed = frame.pack_struct("point", ["x", "y"])
var xs = packed.select_exprs([col("point").field("x")])
var flat = packed.unnest("point")The .list() namespace has len, get(i) (negative from the end, null when
out of range), first, last, contains(value), join(separator), sum,
min, max and mean; field(name) reads one struct field, null where the
struct is null. col(x).implode() gathers a group's values into one list
inside agg (or a whole column into one row), and as_struct([...]) packs
expressions into a struct column. Struct columns can be group_by, unique
and inner/left/semi/anti join keys; they compare field by field, with a null
struct distinct from a struct of nulls. explode and unnest also exist on
LazyFrame. Nested columns take part in select, filter, take, slice,
concat, head, equals, display, and Arrow export and import; they cannot
yet be sort keys, list columns cannot be keys at all, and neither can be
cast, reduced (other than implode), used in when/then, or written to
CSV. Each of those raises a clear error.
The reader is Arrow C++'s, built with only its Parquet parts and bundled
codecs into libdfparquet, a 15 MB shared library with no dependencies
beyond libc. It is loaded at run time, so nothing else in the package needs
it. pixi run -e native build-dfparquet builds it into build/dfparquet/
(a C++20 compiler, cmake and ninja are needed); DATAFRAME_PARQUET_LIBRARY
points read_parquet at a library elsewhere. See native/dfparquet/ and
dataframe/parquet.mojo for the boundary a Mojo-native reader would replace.
The reader consumes bounded file buffers and retains tokenizer state across them. Its scalar structural scanner is the correctness reference for later SIMD scanning and record-boundary-aware parallel decoding. See the complete CSV contract.
Use typed literals: there is no implicit Int64/Float64 promotion. Expressions
support + - * / // % **, comparisons < <= > >= plus .eq()/.ne(), unary
-, abs, sqrt, exp, log, floor, ceil, round, clip, .alias(),
Kleene & | ^ ~, is_null, is_nan, fill_null, fill_nan, coalesce,
selectors (all, col([...]), exclude, by_dtype, nth), window
operations (cum_sum, shift, diff, rank, rolling_*, forward_fill,
interpolate, interpolate_by, cut, qcut, over), is_in,
is_between, cast, when(...).then(...).otherwise(...), a .str()
namespace with concat_str, and the reductions sum(min_count=0), count,
null_count, len, min, max, mean, first, last, n_unique, var,
std, median, quantile, any, all, arg_min, arg_max, mode,
value_counts, skew and kurtosis, plus the two-column corr and cov. Equality is .eq() rather than
Python-style ==. See the operator table in expressions.
select(expr) selects one expression; select_exprs([expr, ...]) selects several.
select(["name", ...]) remains the name-only projection API, with no ambiguous
empty-list overload. with_columns and agg accept an expression or a list.
Sibling expressions all see the original input, not each other's aliases.
Expressions distinguish scalar, row-valued, and aggregate results. For example:
var totals = sales.select(col("amount").sum()) # one row
var repeated = sales.with_columns(col("amount").sum().alias("total"))
var deviations = sales.select(col("amount") - col("amount").sum())Expression sums return zero for empty/all-null input, or null when the valid
count is below min_count, following Narwhals' default. Integer sums accumulate exactly in
128 bits and check the final Int64 result, allowing parallel partial-state merging.
Floating-point expression sums permit reassociation; bitwise reproducibility is
not guaranteed. See the expression contract.
- Typed
Column[T]with bit-packed validity and checked element access; Booleans and strings use the Arrow bit-packed and UTF-8 layouts. - Heterogeneous
Seriesand runtime-schemaDataFrame: Int8–Int64, UInt8–UInt64, Float32, Float64, Bool, String, Date, Datetime, Duration, and Time. The runtime tag is per column, not per cell. - Schema inspection, projection, indexed gathering, nullable Boolean filtering, and adding/replacing columns.
- Vertical, diagonal, and horizontal
concat, plusvstack/hstack. unique,n_unique,is_duplicated,drop_nulls, and framefill_null.pivotandunpivotreshaping.- A
SeriesAPI (operators, reductions,value_counts,unique,sort) backed by the same expression kernels. - Lazy queries (
frame.lazy(),scan_csv) with predicate, projection, and slice pushdown andexplain(); see lazy queries. - Bounded table display for
DataFrameandSeries(print(frame),to_string(max_rows=..., ...),glimpse()). head/tail/slice/reverse,drop/rename/with_row_index, row and cell access through the taggedAnyValue, null counts, and structuralequals.- Hash grouping by any number of columns or key expressions of any dtype.
- Stable multi-column sorting with per-column direction and null placement,
plus
arg_sort,top_k, andbottom_k. - Hash joins on any number of keys of any dtype: inner, left, right, full,
semi, anti, and cross, with
left_on/right_on. - Expression IR, schema binding, scalar broadcasting, and grouped aggregates.
- Explicit SIMD Float64 arithmetic/comparison kernels and mergeable reduction states.
- Contract tests for validation, nulls, overflow, ordering, joins, empty shapes, and scalar/SIMD agreement across batch boundaries.
Useful calls:
var selected = frame.select(["region", "amount"])
var first = frame.head(3)
var cleaned = frame.drop("notes").rename({"amount": "revenue"})
var cell = frame.item(0, "region").string()
var sampled = frame.take([3, 0, 3])
var sorted = frame.sort(["region", "amount"], descending=[False, True], nulls_last=[True, True])
var joined = frame.join(regions, on="region", how="left")Operations return new dataframes. column() and typed extraction methods
(int64(), float64(), bool(), and string(), which returns a
StringColumn) return immutable columns that
share buffers with the frame; no copy is made.
An empty dataframe can retain its row count: DataFrame([], height=10).
Execution is eager. Reductions, row-wise expressions, and filters run on
worker threads for large inputs (DATAFRAME_THREADS sets the limit; 1
disables); hash grouping, joins, and sorting are single-threaded for now. Float64 expression arithmetic
and comparisons use SIMD; checked integer arithmetic and reductions are currently
scalar. Projections, slices, and column extraction share immutable buffers
(zero-copy); computed results allocate new buffers, and unfused intermediates
are still materialized per batch. String columns use the Arrow large_utf8
layout (one UTF-8 buffer plus Int64 offsets). Frames export to and import from
the Arrow C Data Interface (see Arrow interchange); most
dtypes export zero-copy. There are no performance claims yet.
The underscore-prefixed fields are internal and must not be mutated by callers.
There is no GPU execution, index alignment, implicit dtype coercion, or dataframe backend adapter. Build input columns in memory for this first version.
pixi run package precompiles dist/dataframe.mojoc; put its directory on the
import path (mojo run -I dist app.mojo). Tagged releases attach the package
and a generated API reference (pixi run docs). See the
testing guide, the changelog, the stability policy, and the
pandas/Polars migration guide.
See semantics for the behavior the tests promise and architecture and next steps for the intended progression. Historical benchmarks and one-off experiments are in the research wiki.