Numba CUDA bindings for Eigen's fixed-size vector and matrix types, powered by numbast.
from eigenprim import Vector3f, Matrix3f, dot, norm, inverse, links
from numba import cuda
import numpy as np
@cuda.jit(link=links())
def kernel(out):
a = Vector3f(1.0, 2.0, 3.0)
b = Vector3f(4.0, 5.0, 6.0)
out[0] = dot(a, b) # type-dispatched
out[1] = norm(a + b) # operators + generic functions
out = np.zeros(2, dtype=np.float32)
kernel[1, 1](out)Import types and functions, pass links() to @cuda.jit, and use them directly in the kernel.
Run all examples with:
pixi run run-examplesIndividual examples can also run after installing EigenPrim:
python examples/01_vector_basics.py # Vector3f dot, norm, add
python examples/02_all_types.py # All 12 types with operators
python examples/03_point_cloud_transform.py # Batch rigid-body transform
python examples/04_batch_linear_solve.py # Batch Ax=b solve via inverse
python examples/05_templates.py # Generic templated functions (experimental)
python examples/06_triangle_normals.py # Surface normals: cross, normalized
python examples/07_covariance_matrix.py # Outer products: outer, diagonal, trace
python examples/08_aabb_reduction.py # Bounding box: cwise_min, cwise_max
python examples/09_double_precision.py # Double-precision N-body: Vector3d
python examples/10_methods.py # Method invocation: a.dot(b), M.inverse()Examples 03-09 are realistic parallel patterns verified against numpy. See examples/README.md for details.
Clone with submodules, then create an environment with Python 3.10+ and a
CUDA-compatible host C++ compiler (g++/gcc or clang++/clang). EigenPrim
uses the pinned Eigen checkout in thirdparty/eigen by default.
git clone --recurse-submodules <EigenPrim repo>
cd EigenPrim
python -m venv .venv
source .venv/bin/activate
python -m pip install -U pip
python -m pip install -e .
pytest tests/ -vIf the repository was cloned without submodules, initialize Eigen before building:
git submodule update --init --recursive thirdparty/eigenThe build uses scikit-build-core and native CMake CUDA targets to compile
EigenPrim's packaged fatbins. With normal build isolation, pip installs the CUDA
compiler build dependencies declared in pyproject.toml; if you disable build
isolation, your environment must already provide nvcc on PATH.
pixi install
pixi run install-all # installs EigenPrim and PyPI dependencies
pixi run test # runs all examples + pytest
pixi run run-examples # examples onlyPixi installs Eigen, CUDA Toolkit 13, a host compiler, and the Python package
dependencies. EigenPrim depends on released numbast>=0.8.0 and
ast_canopy>=0.8.0 packages; no local Numbast checkout or custom ast_canopy
build is required.
- numbast >= 0.8.0 — C++ to Numba binding generator
- ast_canopy >= 0.8.0 — C++ declaration parser
- Eigen headers, provided by default as a pinned
thirdparty/eigensubmodule - CUDA Toolkit 13 or the CUDA compiler Python build dependencies
- A CUDA-compatible host C++ compiler (
g++/gccorclang++/clang) available in the environment;nvccuses it when compiling EigenPrim's fatbin
During the wheel build, the cuda-activate build requirement puts the nvcc
found by cuda.pathfinder on PATH before scikit-build-core configures CMake.
CMake then uses normal CUDA language discovery, not a project-local nvcc
command. Eigen detection prefers explicit overrides, then the pinned
thirdparty/eigen submodule, then environment include paths such as
CONDA_PREFIX/include/eigen3, PREFIX/include/eigen3, and common system include
paths such as /usr/include/eigen3. Override with either EIGEN_INCLUDE_DIR or
-Ccmake.define.EIGENPRIM_EIGEN_INCLUDE_DIR=/path/to/eigen3. Override
EIGENPRIM_CUDA_ARCHITECTURES to choose which GPU architectures are embedded in
the packaged fatbins.
| Python name | Eigen type | Constructor |
|---|---|---|
Vector2f |
Matrix<float,2,1> |
Vector2f(x, y) |
Vector3f |
Matrix<float,3,1> |
Vector3f(x, y, z) |
Vector4f |
Matrix<float,4,1> |
Vector4f(x, y, z, w) |
Vector2h |
Matrix<half,2,1> |
Vector2h(float16(x), float16(y)) |
Vector3h |
Matrix<half,3,1> |
Vector3h(float16(x), float16(y), float16(z)) |
Vector4h |
Matrix<half,4,1> |
Vector4h(float16(x), ..., float16(w)) |
Vector2bf |
Matrix<bfloat16,2,1> |
Vector2bf(bf16(x), bf16(y)) |
Vector3bf |
Matrix<bfloat16,3,1> |
Vector3bf(bf16(x), bf16(y), bf16(z)) |
Vector4bf |
Matrix<bfloat16,4,1> |
Vector4bf(bf16(x), ..., bf16(w)) |
Vector2d |
Matrix<double,2,1> |
Vector2d(x, y) |
Vector3d |
Matrix<double,3,1> |
Vector3d(x, y, z) |
Vector4d |
Matrix<double,4,1> |
Vector4d(x, y, z, w) |
| Python name | Eigen type | Constructor |
|---|---|---|
Matrix2f |
Matrix<float,2,2> |
Matrix2f(c0r0, c0r1, c1r0, c1r1) |
Matrix3f |
Matrix<float,3,3> |
Matrix3f(c0r0, c0r1, ..., c2r2) — 9 args |
Matrix4f |
Matrix<float,4,4> |
Matrix4f(c0r0, c0r1, ..., c3r3) — 16 args |
Matrix2h |
Matrix<half,2,2> |
Matrix2h(float16(c0r0), ..., float16(c1r1)) — 4 args |
Matrix3h |
Matrix<half,3,3> |
Matrix3h(float16(c0r0), ..., float16(c2r2)) — 9 args |
Matrix4h |
Matrix<half,4,4> |
Matrix4h(float16(c0r0), ..., float16(c3r3)) — 16 args |
Matrix2bf |
Matrix<bfloat16,2,2> |
Matrix2bf(...) — 4 args |
Matrix3bf |
Matrix<bfloat16,3,3> |
Matrix3bf(...) — 9 args |
Matrix4bf |
Matrix<bfloat16,4,4> |
Matrix4bf(...) — 16 args |
Matrix2d |
Matrix<double,2,2> |
Matrix2d(c0r0, c0r1, c1r0, c1r1) |
Matrix3d |
Matrix<double,3,3> |
Matrix3d(c0r0, c0r1, ..., c2r2) — 9 args |
Matrix4d |
Matrix<double,4,4> |
Matrix4d(c0r0, c0r1, ..., c3r3) — 16 args |
Matrix constructors take elements in column-major order (Eigen's native layout). For a 3x3 matrix:
Matrix3f(
col0_row0, col0_row1, col0_row2, # first column
col1_row0, col1_row1, col1_row2, # second column
col2_row0, col2_row1, col2_row2, # third column
)
Example — a diagonal matrix diag(2, 3, 4):
D = Matrix3f(2.0, 0.0, 0.0, # col 0
0.0, 3.0, 0.0, # col 1
0.0, 0.0, 4.0) # col 2Example — the matrix [[1,2,3],[4,5,6],[7,8,9]] (rows are 1-2-3, 4-5-6, 7-8-9):
M = Matrix3f(1.0, 4.0, 7.0, # col 0: rows 0,1,2
2.0, 5.0, 8.0, # col 1: rows 0,1,2
3.0, 6.0, 9.0) # col 2: rows 0,1,2Operations can be called four ways:
- Methods:
a.dot(b),M.inverse(),v.norm()— Eigen-style chaining - Operators:
a + b,M @ v,v * 2.0 - Generic functions:
eigenprim.dot(a, b),eigenprim.inverse(M)— type-dispatched - Explicit functions:
eigen_vec3f_dot(a, b)— when you need full control
All functions follow the naming pattern eigen_{type}_{op}, where type is any of the 24 type suffixes: vec2f..vec4d, vec2h..vec4h, vec2bf..vec4bf, mat2f..mat4d, mat2h..mat4h, mat2bf..mat4bf.
Available for all 12 vector types (float, double, half, bfloat16 x 2D/3D/4D):
| Operation | Example | Returns |
|---|---|---|
add(a, b) |
eigen_vec3f_add(a, b) |
vector |
sub(a, b) |
eigen_vec3f_sub(a, b) |
vector |
dot(a, b) |
eigen_vec3f_dot(a, b) |
scalar |
norm(v) |
eigen_vec3f_norm(v) |
scalar |
squared_norm(v) |
eigen_vec3f_squared_norm(v) |
scalar |
normalized(v) |
eigen_vec3f_normalized(v) |
vector |
scale(v, s) |
eigen_vec3f_scale(v, 2.0) |
vector |
cross(a, b) |
eigen_vec3f_cross(a, b) |
vector |
cwise_product(a, b) |
eigen_vec3f_cwise_product(a, b) |
vector |
cwise_abs(v) |
eigen_vec3f_cwise_abs(v) |
vector |
cwise_min(a, b) |
eigen_vec3f_cwise_min(a, b) |
vector |
cwise_max(a, b) |
eigen_vec3f_cwise_max(a, b) |
vector |
sum(v) |
eigen_vec3f_sum(v) |
scalar |
min_coeff(v) |
eigen_vec3f_min_coeff(v) |
scalar |
max_coeff(v) |
eigen_vec3f_max_coeff(v) |
scalar |
outer(a, b) |
eigen_vec3f_outer(a, b) |
matrix |
cross is only available for 3D vectors (vec3f, vec3d).
Scalar return type matches the vector's element type: float for *f types, double for *d types.
Available for all 12 matrix types (float, double, half, bfloat16 x 2x2/3x3/4x4):
| Operation | Example | Returns |
|---|---|---|
add(a, b) |
eigen_mat3f_add(a, b) |
matrix |
sub(a, b) |
eigen_mat3f_sub(a, b) |
matrix |
mul(a, b) |
eigen_mat3f_mul(a, b) |
matrix |
{mat}_{vec}_mul(m, v) |
eigen_mat3f_vec3f_mul(m, v) |
vector |
determinant(m) |
eigen_mat3f_determinant(m) |
scalar |
inverse(m) |
eigen_mat3f_inverse(m) |
matrix |
transpose(m) |
eigen_mat3f_transpose(m) |
matrix |
trace(m) |
eigen_mat3f_trace(m) |
scalar |
cwise_product(a, b) |
eigen_mat3f_cwise_product(a, b) |
matrix |
scale(m, s) |
eigen_mat3f_scale(m, 2.0) |
matrix |
norm(m) |
eigen_mat3f_norm(m) |
scalar (Frobenius) |
squared_norm(m) |
eigen_mat3f_squared_norm(m) |
scalar |
diagonal(m) |
eigen_mat3f_diagonal(m) |
vector |
Matrix-vector multiply uses the pattern eigen_{mattype}_{vectype}_mul — for example, eigen_mat4f_vec4f_mul(m, v).
All operations are also available as methods directly on instances, so you can use Eigen-style chaining inside kernels:
| Method | Applies to | Returns |
|---|---|---|
v.dot(other) |
vectors | scalar |
v.cross(other) |
3D vectors only | vector |
v.norm() |
vectors | scalar |
v.squared_norm() |
vectors | scalar |
v.normalized() |
vectors | unit vector |
v.scale(s) |
vectors, matrices | same type |
v.sum() |
vectors | scalar |
v.min_coeff() |
vectors | scalar |
v.max_coeff() |
vectors | scalar |
v.cwise_product(other) |
vectors, matrices | same type |
v.cwise_abs() |
vectors | vector |
v.cwise_min(other) |
vectors | vector |
v.cwise_max(other) |
vectors | vector |
v.outer(other) |
vectors | matrix |
M.determinant() |
matrices | scalar |
M.inverse() |
matrices | matrix |
M.transpose() |
matrices | matrix |
M.trace() |
matrices | scalar |
M.norm() |
matrices | scalar (Frobenius) |
M.squared_norm() |
matrices | scalar |
M.diagonal() |
matrices | vector |
M.vec_mul(v) |
matrices | vector |
@cuda.jit(link=links())
def kernel(out):
a = Vector3f(1.0, 2.0, 3.0)
b = Vector3f(4.0, 5.0, 6.0)
out[0] = a.dot(b) # 32.0
out[1] = a.norm() # 3.7417
c = a.cross(b)
u = c.normalized() # unit vector, chain methods
M = Matrix3f(2.0, 0.0, 0.0,
0.0, 2.0, 0.0,
0.0, 0.0, 2.0)
out[2] = M.determinant() # 8.0
out[3] = M.inverse().trace() # 1.5
out[4] = M.vec_mul(a).dot(a) # 28.0Methods and generic dispatch functions (dot(a, b), norm(v), ...) are exact aliases — they call the same underlying CUDA function.
Standard Python operators are overloaded for all 24 Eigen types (including half and bfloat16), so you can write natural expressions in kernels:
| Syntax | Vectors | Matrices |
|---|---|---|
a + b |
vector add | matrix add |
a - b |
vector sub | matrix sub |
v * 2.0 or 2.0 * v |
scalar multiply | — |
M @ N |
— | matrix multiply |
M @ v |
— | matrix-vector multiply |
@cuda.jit(link=links())
def kernel(out):
a = Vector3f(1.0, 2.0, 3.0)
b = Vector3f(4.0, 5.0, 6.0)
c = a + b # eigen_vec3f_add
d = a * 2.0 # eigen_vec3f_scale
e = 3.0 * a # eigen_vec3f_scale (reversed)
M = Matrix3f(...)
v = M @ a # eigen_mat3f_vec3f_mul
R = M @ M # eigen_mat3f_mulAll operations are available as type-dispatched generic functions, accessible as eigenprim.{op} or via direct import. They dispatch to the correct eigen_{type}_{op} at JIT compile time based on argument types. Works across all 24 types (float, double, half, bfloat16).
from eigenprim import dot, norm, inverse, determinant, outer, diagonal
@cuda.jit(link=links())
def kernel(out):
a = Vector3f(1.0, 2.0, 3.0)
b = Vector3f(4.0, 5.0, 6.0)
out[0] = dot(a, b) # -> eigen_vec3f_dot
out[1] = norm(a) # -> eigen_vec3f_norm
M = Matrix3f(...)
out[2] = determinant(M) # -> eigen_mat3f_determinant
inv_M = inverse(M) # -> eigen_mat3f_inverse
d = diagonal(M) # -> eigen_mat3f_diagonal (returns Vector3f)
P = outer(a, b) # -> eigen_vec3f_outer (returns Matrix3f)Full list: add, sub, dot, cross, norm, squared_norm, normalized, scale, cwise_product, cwise_abs, cwise_min, cwise_max, sum, min_coeff, max_coeff, outer, mul, determinant, inverse, transpose, trace, diagonal, vec_mul.
Generic functions where Numba deduces the scalar type from arguments:
from eigenprim import templated_dot3, links
from numba import types
@cuda.jit(link=links())
def kernel(out):
out[0] = templated_dot3(
types.float32(1.0), types.float32(2.0), types.float32(3.0),
types.float32(4.0), types.float32(5.0), types.float32(6.0),
)Every eigenprim operation is a per-thread __device__ primitive.
Each CUDA thread constructs and owns its own Eigen objects. There is no communication between threads, no shared memory access, and no synchronization involved. The mental model is:
one kernel thread -> one Eigen object -> one result
eigenprim has no warp- or block-level primitives (no reductions across threads, no shuffles, no barriers). For those, use CUDA intrinsics or CUB (cub::WarpReduce, cub::BlockReduce, etc.). The two layers compose cleanly:
@cuda.jit(link=links())
def kernel(vals, out):
i = cuda.grid(1)
if i >= vals.shape[0]: return
# eigenprim: per-thread linear algebra
a = Vector3f(vals[i, 0], vals[i, 1], vals[i, 2])
scalar = dot(a, a) # stays in this thread's registers
# hand scalar off to CUB / shared memory for cross-thread work
out[i] = scalarFor binding your own Eigen-wrapping CUDA headers, use the low-level API. The same flow supports both free functions and methods. The key rule is that NVRTC must see only Eigen-free declarations, while nvcc compiles the real Eigen implementation into a fatbin.
Implementation header (my_ops.cuh) — uses real Eigen, compiled by nvcc:
#pragma once
#include <Eigen/Dense>
struct MyVec {
Eigen::Matrix<float, 3, 1> v;
__host__ __device__ MyVec() {}
__host__ __device__ MyVec(float x, float y, float z) : v(x, y, z) {}
__device__ float dot(MyVec other) const {
return v.dot(other.v);
}
};
__device__ float my_dot(MyVec a, MyVec b) {
return a.dot(b);
}Declaration header (my_ops_decl.cuh) — NVRTC-safe, no Eigen. Methods are
declared on the stub type and defined as small trampolines to standalone
functions:
#pragma once
struct MyVec {
float data[3];
__host__ __device__ MyVec() {}
__host__ __device__ MyVec(float x, float y, float z) {
data[0] = x;
data[1] = y;
data[2] = z;
}
__device__ float dot(MyVec other) const;
};
__device__ float my_dot(MyVec a, MyVec b);
__device__ inline float MyVec::dot(MyVec other) const {
return my_dot(*this, other);
}Compilation unit (my_ops.cu):
#include "my_ops.cuh"from eigenprim import bind_eigen_header
bindings = bind_eigen_header(
header="my_ops.cuh",
decl_header="my_ops_decl.cuh",
fatbin="my_ops.fatbin",
)
MyVec = bindings.types["MyVec"]
my_dot = bindings.functions["my_dot"]Numbast binds the method because it appears on the NVRTC-safe MyVec
declaration. At runtime, a.dot(b) enters the NVRTC-safe trampoline, which
calls the standalone my_dot symbol linked from the nvcc-compiled fatbin.
If your public signatures expose Eigen::Matrix<...> directly instead of a
wrapper type, pass stub_header and type_map so EigenPrim can register the
matching NVRTC-safe stub type. Custom method names still need to appear on that
stub type; for new method APIs, a wrapper type like MyVec is usually cleaner
than trying to extend Eigen::Matrix itself.
from numba import cuda
@cuda.jit(link=bindings.links())
def kernel(out):
a = MyVec(1.0, 2.0, 3.0)
b = MyVec(4.0, 5.0, 6.0)
out[0] = my_dot(a, b) # free function
out[1] = a.dot(b) # method trampolineThe key requirement: NVRTC (used by numba-cuda) cannot compile Eigen headers.
So you must separate declarations (NVRTC-safe stubs and trampolines) from
implementations compiled by CUDA C++ into a fatbin. EigenPrim's built-in matrix
and generic fatbins are compiled by native CMake CUDA targets during the wheel
build; for custom bindings, compile your own .cu implementation and pass its
fatbin path to bind_eigen_header. links() yields both for
@cuda.jit(link=...) to link together.
Apache License 2.0. See LICENSE.