Fast 3D affine transformations with trilinear interpolation using AVX2/AVX512 SIMD.
| Data Type | Throughput | Memory/Voxel | Speedup |
|---|---|---|---|
| f32 | 1.5 Gvoxels/s | 8 bytes | 1.0x |
| f16 | 1.2 Gvoxels/s | 4 bytes | - |
| u8 | 3.3 Gvoxels/s | 2 bytes | 2.2x |
Use u8 for image data (microscopy, CT, MRI) to get 2.2x faster processing with 4x less memory!
Benchmark on AMD Ryzen 9 9950X (32 threads):
| Volume | scipy | affiners f32 | affiners u8 | Speedup (f32) | Speedup (u8) |
|---|---|---|---|---|---|
| 512³ | 890 ms | 89 ms | 40 ms | 10x | 22x |
| 1024³ | 7.1 s | 710 ms | 320 ms | 10x | 22x |
pip install affine-rsOr build from source:
pip install .import numpy as np
import affiners
# Define 4x4 homogeneous transformation matrix
# Format: [[m00, m01, m02, tz],
# [m10, m11, m12, ty],
# [m20, m21, m22, tx],
# [0, 0, 0, 1 ]]
matrix = np.array([
[1.0, 0.25, 0.01, -10.0],
[0.0, 1.0, 0.0, -5.0],
[0.0, -0.02, 1.0, 8.0],
[0.0, 0.0, 0.0, 1.0],
]) # Any numeric dtype works - auto-converted to float64
# affine_transform auto-dispatches based on input dtype
# Float32 data
input_f32 = np.random.rand(512, 512, 512).astype(np.float32)
output_f32 = affiners.affine_transform(input_f32, matrix) # returns float32
# uint8 data (2.2x faster!)
input_u8 = np.random.randint(0, 256, (512, 512, 512), dtype=np.uint8)
output_u8 = affiners.affine_transform(input_u8, matrix) # returns uint8
# float16 data (2x less memory)
input_f16 = input_f32.astype(np.float16)
output_f16 = affiners.affine_transform(input_f16, matrix) # returns float16
# With custom output shape
output = affiners.affine_transform(input_f32, matrix, output_shape=(256, 256, 256))import affiners
print(affiners.__version__) # '0.2.0'
print(affiners.build_info())
# {'version': '0.2.0', 'simd': {'avx2': True, 'avx512f': True, ...},
# 'backend_f32': 'avx512', 'backend_u8': 'avx2', 'num_threads': 32, ...}use ndarray::{Array2, Array3};
use affiners::{affine_transform_3d_f32, affine_transform_3d_u8};
// Create 4x4 homogeneous transformation matrix (identity with translation)
let matrix = Array2::from_shape_vec((4, 4), vec![
1.0, 0.0, 0.0, 10.0, // tz = 10
0.0, 1.0, 0.0, 20.0, // ty = 20
0.0, 0.0, 1.0, 30.0, // tx = 30
0.0, 0.0, 0.0, 1.0,
]).unwrap();
// Float32 (output same shape as input)
let input_f32 = Array3::<f32>::zeros((100, 100, 100));
let output_f32 = affine_transform_3d_f32(&input_f32.view(), &matrix.view(), None, 0.0);
// With custom output shape
let output_f32 = affine_transform_3d_f32(&input_f32.view(), &matrix.view(), Some((50, 50, 50)), 0.0);
// uint8 (2.2x faster!)
let input_u8 = Array3::<u8>::zeros((100, 100, 100));
let output_u8 = affine_transform_3d_u8(&input_u8.view(), &matrix.view(), None, 0);Main function (auto-dispatches based on input dtype):
| Function | Input Type | Description |
|---|---|---|
affine_transform(input, matrix, output_shape, cval) |
float32, float16, uint8 | Auto-dispatches, preserves dtype |
Type-specific functions (for explicit control):
| Function | Input Type | Description |
|---|---|---|
affine_transform_f32(input, matrix, output_shape, cval) |
float32 | Standard floating point |
affine_transform_f16(input, matrix, output_shape, cval) |
uint16 | Half precision (pass as .view(np.uint16)) |
affine_transform_u8(input, matrix, output_shape, cval) |
uint8 | 2.2x faster, 4x less memory |
build_info() |
- | Get version, SIMD features, and backend info |
input: 3D numpy array (C-contiguous)matrix: 4x4 homogeneous transformation matrix (any numeric dtype, auto-converted to float64)[[m00, m01, m02, tz], [m10, m11, m12, ty], [m20, m21, m22, tx], [0, 0, 0, 1 ]]output_shape: Optional tuple (z, y, x) for output dimensions (default: same as input)cval: Constant value for out-of-bounds (default: 0)
| Use Case | Recommended Type |
|---|---|
| Image data (CT, MRI, microscopy) | u8 - 2.2x faster |
| Scientific floating-point data | f32 |
| Reduced memory footprint | f16 or u8 |
| Maximum precision | f32 |
| Volume Size | f32 | u8 |
|---|---|---|
| 512³ | 1.1 GB | 0.3 GB |
| 1024³ | 8.6 GB | 2.1 GB |
| 2048³ | 68.7 GB | 17.2 GB |
- AVX2/AVX512 SIMD: Processes 8-16 values per iteration
- Multi-threaded: Uses rayon for parallel z-slice processing
- Memory efficient: u8 uses 4x less memory than f32
- Python bindings: Via PyO3 and maturin
- Zero-copy: Works directly with numpy arrays
- Native compilation: Optimized for host CPU features (
target-cpu=native)
BSD-3-Clause