HydraQYH/hp_rms_norm
High performance RMSNorm Implement by using SM Core Storage(Registers and Shared Memory)
NV/AMD GPU Kernel Performance Analysis & Optimization. Now in Bytedance(Data => AML => Seed). Used to be in AIACC Team@Alibaba Cloud.
High performance RMSNorm Implement by using SM Core Storage(Registers and Shared Memory)
FlyDSL is the Python front‑end of the project: Flexible LaYout DSL.
CUDA Templates for Linear Algebra Subroutines
Using a swizzled hierarchical layout for GEMM
SGLang is a fast serving framework for large language models and vision language models.
Reduce kernel based on CUTLASS CuTe and TMA.
High-level tracing language for Linux eBPF
Open deep learning compiler stack for cpu, gpu and specialized accelerators
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
FlashInfer: Kernel Library for LLM Serving
A high-throughput and memory-efficient inference and serving engine for LLMs
DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
Fast and memory-efficient exact attention
Benchmark sync primitives
Instruction-level benchmarks for NVGPUs
Prefetch experiments codes.
Expert Specialization MoE Solution based on CUTLASS
Some code snippet for CuTeDSL
Simple Code Snippet for TMA Multicast.
Cache operator demo code for peer access
This is a repository with some CUDA examples.
This is a repository with some OpenMP examples.
A repository with some MPI programming examples.
Puzzles for learning Triton, play it with minimal environment configuration!
Bpftrace scripts for measuring nivcsw overhead.
Practice codes for Understanding Software Dynamics.
Codes for YuhangOS
Homework for mlc.ai.