trevor-m/dynamo
A Datacenter Scale Distributed Inference Serving Framework
Sglang team at NVIDIA (Deep learning frameworks).
A Datacenter Scale Distributed Inference Serving Framework
DeepEP: an efficient expert-parallel communication library
DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
A REYES-style micropolygon renderer written in C++ which implements a subset of the RenderMan specification.
Open Source Continuous Inference Benchmarking Qwen3.5, DeepSeek, GPTOSS - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 vs H100 & soon™ TPUv6e/v7/Trainium2/3
NVIDIA Inference Benchmarks provide recipes in ready-to-use templates for evaluating platform speed. Validate your platform across specific AI use cases across hardware and software combinations.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
Cookbook of SGLang - Recipe
Implementation of "Fast Global Illumination Approximations on Deep G-Buffers" (Mara et. al, 2016) using C++, OpenGL, and GLSL
Procedurally generates a terrain rendered with OpenGL.
tf.image.resize_images has aliasing when downsampling and does not have gradients for bicubic mode. This implementation fixes those problems.
Tensorflow implementation of "Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network" (Ledig et al. 2017)
FlashInfer: Kernel Library for LLM Serving
A high-throughput and memory-efficient inference and serving engine for LLMs
SGLang is a fast serving framework for large language models and vision language models.
Using deep learning to distinguish between Tor and nonTor traffic
Open deep learning compiler stack for cpu, gpu and specialized accelerators
JAX-Toolbox
🤗 Transformers: State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX.
Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
An Open Source Machine Learning Framework for Everyone
A retargetable MLIR-based machine learning compiler and runtime toolkit.
PJRT plugin for interfacing the OpenXLA compiler to Jax, PyTorch/XLA and TensorFlow
A real-time photorealistic path tracer using OpenGL GPU Compute Shaders
A multithreaded Whitted ray tracer (C++) which supports reflection, refraction, shadows, interpolated textures and normals, color and intersection shaders, as well as Monte Carlo anti-aliasing, depth-of-field, and BSSSRDFs
A machine learning compiler for GPUs, CPUs, and ML accelerators
A MUD Server (Text based online multiplayer game) that uses Lua for scripting of skills and abilities.
The LLVM Project is a collection of modular and reusable compiler and toolchain technologies. Note: the repository does not accept github pull requests at this moment. Please submit your patches at http://reviews.llvm.org.
Optimized primitives for collective multi-GPU communication
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit floating point (FP8) precision on Hopper GPUs, to provide better performance with lower memory utilization in both training and inference.