JINO-ROHIT/blackwell_kernels
a collection of blackwell kernels
AI Researcher | Top 1% Kaggle Competitions Expert
a collection of blackwell kernels
a minimal paged attention implementation
SGLang Omni: High-Performance Multi-Stage Pipeline Framework for Omni Models
a simple c++ inference engine for gpt based architecture
writing really fast kernels
a personal collection of my notes for ml sys
learning different model design patterns for faster inference
fastcv is a CUDA rewrite of the opencv filters with python bindings
a LLM inference engine to run on consumer hardware
a series of benchmarking using dynamo for qwen3 30b a3b
Fast and memory-efficient exact attention
a repo to understand llama.cpp
a collection of kernels and optimizations
SGLang is a high-performance serving framework for large language models and multimodal models.
📚LeetCUDA: Modern CUDA Learn Notes with PyTorch for Beginners🐑, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.🎉
A Datacenter Scale Distributed Inference Serving Framework
A curriculum for learning about gpu performance engineering, from scratch to what the frontier AI labs do
serving a torch model using Celery, Redis and RabbitMQ to serve users asynchronously
a guide to learn and implement inference techniques from scratch
Tensors and Dynamic neural networks in Python with strong GPU acceleration
a playground autograd framework
c++ algorithms and problems
CUDA rewrite of a simple NN in cuda c++