ilan-theodoro/affiners

Fast 3D affine transformations with trilinear interpolation using AVX2/AVX512 SIMD

★ 1Forks 0RustGitHub ↗Compare

README

affiners

Fast 3D affine transformations with trilinear interpolation using AVX2/AVX512 SIMD.

Performance

Data Type Throughput Memory/Voxel Speedup
f32 1.5 Gvoxels/s 8 bytes 1.0x
f16 1.2 Gvoxels/s 4 bytes -
u8 3.3 Gvoxels/s 2 bytes 2.2x

Use u8 for image data (microscopy, CT, MRI) to get 2.2x faster processing with 4x less memory!

Performance vs scipy

Benchmark on AMD Ryzen 9 9950X (32 threads):

Volume scipy affiners f32 affiners u8 Speedup (f32) Speedup (u8)
512³ 890 ms 89 ms 40 ms 10x 22x
1024³ 7.1 s 710 ms 320 ms 10x 22x

Installation

pip install affine-rs

Or build from source:

pip install .

Usage

Python

import numpy as np
import affiners

# Define 4x4 homogeneous transformation matrix
# Format: [[m00, m01, m02, tz],
#          [m10, m11, m12, ty],
#          [m20, m21, m22, tx],
#          [0,   0,   0,   1 ]]
matrix = np.array([
    [1.0, 0.25, 0.01, -10.0],
    [0.0, 1.0, 0.0, -5.0],
    [0.0, -0.02, 1.0, 8.0],
    [0.0, 0.0, 0.0, 1.0],
])  # Any numeric dtype works - auto-converted to float64

# affine_transform auto-dispatches based on input dtype
# Float32 data
input_f32 = np.random.rand(512, 512, 512).astype(np.float32)
output_f32 = affiners.affine_transform(input_f32, matrix)  # returns float32

# uint8 data (2.2x faster!)
input_u8 = np.random.randint(0, 256, (512, 512, 512), dtype=np.uint8)
output_u8 = affiners.affine_transform(input_u8, matrix)  # returns uint8

# float16 data (2x less memory)
input_f16 = input_f32.astype(np.float16)
output_f16 = affiners.affine_transform(input_f16, matrix)  # returns float16

# With custom output shape
output = affiners.affine_transform(input_f32, matrix, output_shape=(256, 256, 256))

Check Build Info

import affiners

print(affiners.__version__)  # '0.2.0'
print(affiners.build_info())
# {'version': '0.2.0', 'simd': {'avx2': True, 'avx512f': True, ...}, 
#  'backend_f32': 'avx512', 'backend_u8': 'avx2', 'num_threads': 32, ...}

Rust

use ndarray::{Array2, Array3};
use affiners::{affine_transform_3d_f32, affine_transform_3d_u8};

// Create 4x4 homogeneous transformation matrix (identity with translation)
let matrix = Array2::from_shape_vec((4, 4), vec![
    1.0, 0.0, 0.0, 10.0,  // tz = 10
    0.0, 1.0, 0.0, 20.0,  // ty = 20
    0.0, 0.0, 1.0, 30.0,  // tx = 30
    0.0, 0.0, 0.0, 1.0,
]).unwrap();

// Float32 (output same shape as input)
let input_f32 = Array3::<f32>::zeros((100, 100, 100));
let output_f32 = affine_transform_3d_f32(&input_f32.view(), &matrix.view(), None, 0.0);

// With custom output shape
let output_f32 = affine_transform_3d_f32(&input_f32.view(), &matrix.view(), Some((50, 50, 50)), 0.0);

// uint8 (2.2x faster!)
let input_u8 = Array3::<u8>::zeros((100, 100, 100));
let output_u8 = affine_transform_3d_u8(&input_u8.view(), &matrix.view(), None, 0);

API Reference

Python

Main function (auto-dispatches based on input dtype):

Function Input Type Description
affine_transform(input, matrix, output_shape, cval) float32, float16, uint8 Auto-dispatches, preserves dtype

Type-specific functions (for explicit control):

Function Input Type Description
affine_transform_f32(input, matrix, output_shape, cval) float32 Standard floating point
affine_transform_f16(input, matrix, output_shape, cval) uint16 Half precision (pass as .view(np.uint16))
affine_transform_u8(input, matrix, output_shape, cval) uint8 2.2x faster, 4x less memory
build_info() - Get version, SIMD features, and backend info

Parameters

  • input: 3D numpy array (C-contiguous)
  • matrix: 4x4 homogeneous transformation matrix (any numeric dtype, auto-converted to float64)
    [[m00, m01, m02, tz],
     [m10, m11, m12, ty],
     [m20, m21, m22, tx],
     [0,   0,   0,   1 ]]
    
  • output_shape: Optional tuple (z, y, x) for output dimensions (default: same as input)
  • cval: Constant value for out-of-bounds (default: 0)

When to Use Each Type

Use Case Recommended Type
Image data (CT, MRI, microscopy) u8 - 2.2x faster
Scientific floating-point data f32
Reduced memory footprint f16 or u8
Maximum precision f32

Memory Requirements

Volume Size f32 u8
512³ 1.1 GB 0.3 GB
1024³ 8.6 GB 2.1 GB
2048³ 68.7 GB 17.2 GB

Features

  • AVX2/AVX512 SIMD: Processes 8-16 values per iteration
  • Multi-threaded: Uses rayon for parallel z-slice processing
  • Memory efficient: u8 uses 4x less memory than f32
  • Python bindings: Via PyO3 and maturin
  • Zero-copy: Works directly with numpy arrays
  • Native compilation: Optimized for host CPU features (target-cpu=native)

License

BSD-3-Clause

Contributors

ilan-theodoro

Issues