arcusbuilds/jax
Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
Graphics Processing Units Molecular Dynamics
High-performance multi-platform compiler for physics simulation
a language for fast, portable data-parallel computation
Performance-portable library for particle-based simulations
Metal programming in Julia
FlagGems is an operator library for large language models implemented in the Triton Language.
A Datacenter Scale Distributed Inference Serving Framework
Scalable, Portable and Distributed Gradient Boosting (GBDT, GBRT or GBM) Library, for Python, R, Java, Scala, C and more. Runs on single machine, Hadoop, Spark, Dask, Flink and DataFlow
Personal site for open-source work and career progress
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.
Ongoing research training transformer models at scale
[ICML2025] SpargeAttention: A training-free sparse attention that accelerates any model inference.
A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications.
Efficient Triton Kernels for LLM Training
Trixi.jl: Adaptive high-order numerical simulations of conservation laws in Julia
Fast kernel library for Diffusion inference with multiple compute backends.
Tokamax: A GPU and TPU kernel library.
Podman: A tool for managing OCI containers and pods.
🚀 Efficient implementations for emerging model architectures
Train transformer language models with reinforcement learning.
You like pytorch? You like micrograd? You love tinygrad! ❤️
eBPF-based Networking, Security, and Observability
Image-processing software for cryo-electron microscopy
Circuit IR Compilers and Tools
cuTile Rust provides a safe, tile-based kernel programming DSL for the Rust programming language. It features a safe host-side API for passing tensors to asynchronously executed kernel functions.
A cross-platform, safe, pure-Rust graphics API.
CUDA Core Compute Libraries
ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.