isVoid/EigenPrim

Numba CUDA bindings for Eigen linear algebra

★ 0Forks 0CudaGitHub ↗Compare

README

Eigenprim

Numba CUDA bindings for Eigen's fixed-size vector and matrix types, powered by numbast.

Quick Start

from eigenprim import Vector3f, Matrix3f, dot, norm, inverse, links
from numba import cuda
import numpy as np

@cuda.jit(link=links())
def kernel(out):
    a = Vector3f(1.0, 2.0, 3.0)
    b = Vector3f(4.0, 5.0, 6.0)
    out[0] = dot(a, b)     # type-dispatched
    out[1] = norm(a + b)   # operators + generic functions

out = np.zeros(2, dtype=np.float32)
kernel[1, 1](out)

Import types and functions, pass links() to @cuda.jit, and use them directly in the kernel.

Run all examples with:

pixi run run-examples

Individual examples can also run after installing EigenPrim:

python examples/01_vector_basics.py          # Vector3f dot, norm, add
python examples/02_all_types.py              # All 12 types with operators
python examples/03_point_cloud_transform.py  # Batch rigid-body transform
python examples/04_batch_linear_solve.py     # Batch Ax=b solve via inverse
python examples/05_templates.py              # Generic templated functions (experimental)
python examples/06_triangle_normals.py       # Surface normals: cross, normalized
python examples/07_covariance_matrix.py      # Outer products: outer, diagonal, trace
python examples/08_aabb_reduction.py         # Bounding box: cwise_min, cwise_max
python examples/09_double_precision.py       # Double-precision N-body: Vector3d
python examples/10_methods.py               # Method invocation: a.dot(b), M.inverse()

Examples 03-09 are realistic parallel patterns verified against numpy. See examples/README.md for details.

Setup

pip

Clone with submodules, then create an environment with Python 3.10+ and a CUDA-compatible host C++ compiler (g++/gcc or clang++/clang). EigenPrim uses the pinned Eigen checkout in thirdparty/eigen by default.

git clone --recurse-submodules <EigenPrim repo>
cd EigenPrim
python -m venv .venv
source .venv/bin/activate
python -m pip install -U pip
python -m pip install -e .
pytest tests/ -v

If the repository was cloned without submodules, initialize Eigen before building:

git submodule update --init --recursive thirdparty/eigen

The build uses scikit-build-core and native CMake CUDA targets to compile EigenPrim's packaged fatbins. With normal build isolation, pip installs the CUDA compiler build dependencies declared in pyproject.toml; if you disable build isolation, your environment must already provide nvcc on PATH.

pixi

pixi install
pixi run install-all  # installs EigenPrim and PyPI dependencies
pixi run test          # runs all examples + pytest
pixi run run-examples  # examples only

Pixi installs Eigen, CUDA Toolkit 13, a host compiler, and the Python package dependencies. EigenPrim depends on released numbast>=0.8.0 and ast_canopy>=0.8.0 packages; no local Numbast checkout or custom ast_canopy build is required.

Dependencies

  • numbast >= 0.8.0 — C++ to Numba binding generator
  • ast_canopy >= 0.8.0 — C++ declaration parser
  • Eigen headers, provided by default as a pinned thirdparty/eigen submodule
  • CUDA Toolkit 13 or the CUDA compiler Python build dependencies
  • A CUDA-compatible host C++ compiler (g++/gcc or clang++/clang) available in the environment; nvcc uses it when compiling EigenPrim's fatbin

Eigen Detection

During the wheel build, the cuda-activate build requirement puts the nvcc found by cuda.pathfinder on PATH before scikit-build-core configures CMake. CMake then uses normal CUDA language discovery, not a project-local nvcc command. Eigen detection prefers explicit overrides, then the pinned thirdparty/eigen submodule, then environment include paths such as CONDA_PREFIX/include/eigen3, PREFIX/include/eigen3, and common system include paths such as /usr/include/eigen3. Override with either EIGEN_INCLUDE_DIR or -Ccmake.define.EIGENPRIM_EIGEN_INCLUDE_DIR=/path/to/eigen3. Override EIGENPRIM_CUDA_ARCHITECTURES to choose which GPU architectures are embedded in the packaged fatbins.

Available Types

Vectors

Python name Eigen type Constructor
Vector2f Matrix<float,2,1> Vector2f(x, y)
Vector3f Matrix<float,3,1> Vector3f(x, y, z)
Vector4f Matrix<float,4,1> Vector4f(x, y, z, w)
Vector2h Matrix<half,2,1> Vector2h(float16(x), float16(y))
Vector3h Matrix<half,3,1> Vector3h(float16(x), float16(y), float16(z))
Vector4h Matrix<half,4,1> Vector4h(float16(x), ..., float16(w))
Vector2bf Matrix<bfloat16,2,1> Vector2bf(bf16(x), bf16(y))
Vector3bf Matrix<bfloat16,3,1> Vector3bf(bf16(x), bf16(y), bf16(z))
Vector4bf Matrix<bfloat16,4,1> Vector4bf(bf16(x), ..., bf16(w))
Vector2d Matrix<double,2,1> Vector2d(x, y)
Vector3d Matrix<double,3,1> Vector3d(x, y, z)
Vector4d Matrix<double,4,1> Vector4d(x, y, z, w)

Matrices

Python name Eigen type Constructor
Matrix2f Matrix<float,2,2> Matrix2f(c0r0, c0r1, c1r0, c1r1)
Matrix3f Matrix<float,3,3> Matrix3f(c0r0, c0r1, ..., c2r2) — 9 args
Matrix4f Matrix<float,4,4> Matrix4f(c0r0, c0r1, ..., c3r3) — 16 args
Matrix2h Matrix<half,2,2> Matrix2h(float16(c0r0), ..., float16(c1r1)) — 4 args
Matrix3h Matrix<half,3,3> Matrix3h(float16(c0r0), ..., float16(c2r2)) — 9 args
Matrix4h Matrix<half,4,4> Matrix4h(float16(c0r0), ..., float16(c3r3)) — 16 args
Matrix2bf Matrix<bfloat16,2,2> Matrix2bf(...) — 4 args
Matrix3bf Matrix<bfloat16,3,3> Matrix3bf(...) — 9 args
Matrix4bf Matrix<bfloat16,4,4> Matrix4bf(...) — 16 args
Matrix2d Matrix<double,2,2> Matrix2d(c0r0, c0r1, c1r0, c1r1)
Matrix3d Matrix<double,3,3> Matrix3d(c0r0, c0r1, ..., c2r2) — 9 args
Matrix4d Matrix<double,4,4> Matrix4d(c0r0, c0r1, ..., c3r3) — 16 args

Column-Major Storage

Matrix constructors take elements in column-major order (Eigen's native layout). For a 3x3 matrix:

Matrix3f(
    col0_row0, col0_row1, col0_row2,   # first column
    col1_row0, col1_row1, col1_row2,   # second column
    col2_row0, col2_row1, col2_row2,   # third column
)

Example — a diagonal matrix diag(2, 3, 4):

D = Matrix3f(2.0, 0.0, 0.0,   # col 0
             0.0, 3.0, 0.0,   # col 1
             0.0, 0.0, 4.0)   # col 2

Example — the matrix [[1,2,3],[4,5,6],[7,8,9]] (rows are 1-2-3, 4-5-6, 7-8-9):

M = Matrix3f(1.0, 4.0, 7.0,   # col 0: rows 0,1,2
             2.0, 5.0, 8.0,   # col 1: rows 0,1,2
             3.0, 6.0, 9.0)   # col 2: rows 0,1,2

Available Operations

Operations can be called four ways:

  • Methods: a.dot(b), M.inverse(), v.norm() — Eigen-style chaining
  • Operators: a + b, M @ v, v * 2.0
  • Generic functions: eigenprim.dot(a, b), eigenprim.inverse(M) — type-dispatched
  • Explicit functions: eigen_vec3f_dot(a, b) — when you need full control

All functions follow the naming pattern eigen_{type}_{op}, where type is any of the 24 type suffixes: vec2f..vec4d, vec2h..vec4h, vec2bf..vec4bf, mat2f..mat4d, mat2h..mat4h, mat2bf..mat4bf.

Vector Operations

Available for all 12 vector types (float, double, half, bfloat16 x 2D/3D/4D):

Operation Example Returns
add(a, b) eigen_vec3f_add(a, b) vector
sub(a, b) eigen_vec3f_sub(a, b) vector
dot(a, b) eigen_vec3f_dot(a, b) scalar
norm(v) eigen_vec3f_norm(v) scalar
squared_norm(v) eigen_vec3f_squared_norm(v) scalar
normalized(v) eigen_vec3f_normalized(v) vector
scale(v, s) eigen_vec3f_scale(v, 2.0) vector
cross(a, b) eigen_vec3f_cross(a, b) vector
cwise_product(a, b) eigen_vec3f_cwise_product(a, b) vector
cwise_abs(v) eigen_vec3f_cwise_abs(v) vector
cwise_min(a, b) eigen_vec3f_cwise_min(a, b) vector
cwise_max(a, b) eigen_vec3f_cwise_max(a, b) vector
sum(v) eigen_vec3f_sum(v) scalar
min_coeff(v) eigen_vec3f_min_coeff(v) scalar
max_coeff(v) eigen_vec3f_max_coeff(v) scalar
outer(a, b) eigen_vec3f_outer(a, b) matrix

cross is only available for 3D vectors (vec3f, vec3d).

Scalar return type matches the vector's element type: float for *f types, double for *d types.

Matrix Operations

Available for all 12 matrix types (float, double, half, bfloat16 x 2x2/3x3/4x4):

Operation Example Returns
add(a, b) eigen_mat3f_add(a, b) matrix
sub(a, b) eigen_mat3f_sub(a, b) matrix
mul(a, b) eigen_mat3f_mul(a, b) matrix
{mat}_{vec}_mul(m, v) eigen_mat3f_vec3f_mul(m, v) vector
determinant(m) eigen_mat3f_determinant(m) scalar
inverse(m) eigen_mat3f_inverse(m) matrix
transpose(m) eigen_mat3f_transpose(m) matrix
trace(m) eigen_mat3f_trace(m) scalar
cwise_product(a, b) eigen_mat3f_cwise_product(a, b) matrix
scale(m, s) eigen_mat3f_scale(m, 2.0) matrix
norm(m) eigen_mat3f_norm(m) scalar (Frobenius)
squared_norm(m) eigen_mat3f_squared_norm(m) scalar
diagonal(m) eigen_mat3f_diagonal(m) vector

Matrix-vector multiply uses the pattern eigen_{mattype}_{vectype}_mul — for example, eigen_mat4f_vec4f_mul(m, v).

Method Syntax

All operations are also available as methods directly on instances, so you can use Eigen-style chaining inside kernels:

Method Applies to Returns
v.dot(other) vectors scalar
v.cross(other) 3D vectors only vector
v.norm() vectors scalar
v.squared_norm() vectors scalar
v.normalized() vectors unit vector
v.scale(s) vectors, matrices same type
v.sum() vectors scalar
v.min_coeff() vectors scalar
v.max_coeff() vectors scalar
v.cwise_product(other) vectors, matrices same type
v.cwise_abs() vectors vector
v.cwise_min(other) vectors vector
v.cwise_max(other) vectors vector
v.outer(other) vectors matrix
M.determinant() matrices scalar
M.inverse() matrices matrix
M.transpose() matrices matrix
M.trace() matrices scalar
M.norm() matrices scalar (Frobenius)
M.squared_norm() matrices scalar
M.diagonal() matrices vector
M.vec_mul(v) matrices vector
@cuda.jit(link=links())
def kernel(out):
    a = Vector3f(1.0, 2.0, 3.0)
    b = Vector3f(4.0, 5.0, 6.0)

    out[0] = a.dot(b)                    # 32.0
    out[1] = a.norm()                    # 3.7417
    c = a.cross(b)
    u = c.normalized()                   # unit vector, chain methods

    M = Matrix3f(2.0, 0.0, 0.0,
                 0.0, 2.0, 0.0,
                 0.0, 0.0, 2.0)
    out[2] = M.determinant()             # 8.0
    out[3] = M.inverse().trace()         # 1.5
    out[4] = M.vec_mul(a).dot(a)         # 28.0

Methods and generic dispatch functions (dot(a, b), norm(v), ...) are exact aliases — they call the same underlying CUDA function.

Operator Syntax

Standard Python operators are overloaded for all 24 Eigen types (including half and bfloat16), so you can write natural expressions in kernels:

Syntax Vectors Matrices
a + b vector add matrix add
a - b vector sub matrix sub
v * 2.0 or 2.0 * v scalar multiply —
M @ N — matrix multiply
M @ v — matrix-vector multiply
@cuda.jit(link=links())
def kernel(out):
    a = Vector3f(1.0, 2.0, 3.0)
    b = Vector3f(4.0, 5.0, 6.0)
    c = a + b           # eigen_vec3f_add
    d = a * 2.0         # eigen_vec3f_scale
    e = 3.0 * a         # eigen_vec3f_scale (reversed)

    M = Matrix3f(...)
    v = M @ a            # eigen_mat3f_vec3f_mul
    R = M @ M            # eigen_mat3f_mul

Generic Functions

All operations are available as type-dispatched generic functions, accessible as eigenprim.{op} or via direct import. They dispatch to the correct eigen_{type}_{op} at JIT compile time based on argument types. Works across all 24 types (float, double, half, bfloat16).

from eigenprim import dot, norm, inverse, determinant, outer, diagonal

@cuda.jit(link=links())
def kernel(out):
    a = Vector3f(1.0, 2.0, 3.0)
    b = Vector3f(4.0, 5.0, 6.0)
    out[0] = dot(a, b)              # -> eigen_vec3f_dot
    out[1] = norm(a)                # -> eigen_vec3f_norm

    M = Matrix3f(...)
    out[2] = determinant(M)         # -> eigen_mat3f_determinant
    inv_M = inverse(M)              # -> eigen_mat3f_inverse
    d = diagonal(M)                 # -> eigen_mat3f_diagonal (returns Vector3f)
    P = outer(a, b)                 # -> eigen_vec3f_outer (returns Matrix3f)

Full list: add, sub, dot, cross, norm, squared_norm, normalized, scale, cwise_product, cwise_abs, cwise_min, cwise_max, sum, min_coeff, max_coeff, outer, mul, determinant, inverse, transpose, trace, diagonal, vec_mul.

Template Functions

Generic functions where Numba deduces the scalar type from arguments:

from eigenprim import templated_dot3, links
from numba import types

@cuda.jit(link=links())
def kernel(out):
    out[0] = templated_dot3(
        types.float32(1.0), types.float32(2.0), types.float32(3.0),
        types.float32(4.0), types.float32(5.0), types.float32(6.0),
    )

Execution Model (Thread-Level Primitives)

Every eigenprim operation is a per-thread __device__ primitive.

Each CUDA thread constructs and owns its own Eigen objects. There is no communication between threads, no shared memory access, and no synchronization involved. The mental model is:

one kernel thread  ->  one Eigen object  ->  one result

eigenprim has no warp- or block-level primitives (no reductions across threads, no shuffles, no barriers). For those, use CUDA intrinsics or CUB (cub::WarpReduce, cub::BlockReduce, etc.). The two layers compose cleanly:

@cuda.jit(link=links())
def kernel(vals, out):
    i = cuda.grid(1)
    if i >= vals.shape[0]: return
    # eigenprim: per-thread linear algebra
    a = Vector3f(vals[i, 0], vals[i, 1], vals[i, 2])
    scalar = dot(a, a)       # stays in this thread's registers
    # hand scalar off to CUB / shared memory for cross-thread work
    out[i] = scalar

Custom Bindings

For binding your own Eigen-wrapping CUDA headers, use the low-level API. The same flow supports both free functions and methods. The key rule is that NVRTC must see only Eigen-free declarations, while nvcc compiles the real Eigen implementation into a fatbin.

Step 1: Write Three Files

Implementation header (my_ops.cuh) — uses real Eigen, compiled by nvcc:

#pragma once
#include <Eigen/Dense>

struct MyVec {
  Eigen::Matrix<float, 3, 1> v;

  __host__ __device__ MyVec() {}

  __host__ __device__ MyVec(float x, float y, float z) : v(x, y, z) {}

  __device__ float dot(MyVec other) const {
    return v.dot(other.v);
  }
};

__device__ float my_dot(MyVec a, MyVec b) {
  return a.dot(b);
}

Declaration header (my_ops_decl.cuh) — NVRTC-safe, no Eigen. Methods are declared on the stub type and defined as small trampolines to standalone functions:

#pragma once

struct MyVec {
  float data[3];

  __host__ __device__ MyVec() {}

  __host__ __device__ MyVec(float x, float y, float z) {
    data[0] = x;
    data[1] = y;
    data[2] = z;
  }

  __device__ float dot(MyVec other) const;
};

__device__ float my_dot(MyVec a, MyVec b);

__device__ inline float MyVec::dot(MyVec other) const {
  return my_dot(*this, other);
}

Compilation unit (my_ops.cu):

#include "my_ops.cuh"

Step 2: Bind

from eigenprim import bind_eigen_header

bindings = bind_eigen_header(
    header="my_ops.cuh",
    decl_header="my_ops_decl.cuh",
    fatbin="my_ops.fatbin",
)

MyVec = bindings.types["MyVec"]
my_dot = bindings.functions["my_dot"]

Numbast binds the method because it appears on the NVRTC-safe MyVec declaration. At runtime, a.dot(b) enters the NVRTC-safe trampoline, which calls the standalone my_dot symbol linked from the nvcc-compiled fatbin.

If your public signatures expose Eigen::Matrix<...> directly instead of a wrapper type, pass stub_header and type_map so EigenPrim can register the matching NVRTC-safe stub type. Custom method names still need to appear on that stub type; for new method APIs, a wrapper type like MyVec is usually cleaner than trying to extend Eigen::Matrix itself.

Step 3: Use

from numba import cuda

@cuda.jit(link=bindings.links())
def kernel(out):
    a = MyVec(1.0, 2.0, 3.0)
    b = MyVec(4.0, 5.0, 6.0)
    out[0] = my_dot(a, b)  # free function
    out[1] = a.dot(b)      # method trampoline

The key requirement: NVRTC (used by numba-cuda) cannot compile Eigen headers. So you must separate declarations (NVRTC-safe stubs and trampolines) from implementations compiled by CUDA C++ into a fatbin. EigenPrim's built-in matrix and generic fatbins are compiled by native CMake CUDA targets during the wheel build; for custom bindings, compile your own .cu implementation and pass its fatbin path to bind_eigen_header. links() yields both for @cuda.jit(link=...) to link together.

License

Apache License 2.0. See LICENSE.

Contributors

isVoid

Issues