Mellonta/flash-attention
Fast and memory-efficient exact attention
Μέλλοντα
Fast and memory-efficient exact attention
Educational PyTorch implementations of linear attention, covering RetNet, Mamba-2, DeltaNet, Gated DeltaNet (GDN), Gated Delta Product (GDP), Gated Linear Attention (GLA), and Kimi Delta Attention (KDA).
Causal depthwise conv1d in CUDA, with a PyTorch interface
Mamba SSM architecture
🚀 Efficient implementations for emerging model architectures
NVIDIA Resiliency Extension is a python package for framework developers and users to implement fault-tolerant features. It improves the effective training time by minimizing the downtime due to failures and interruptions.
A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.
Ongoing research training transformer models at scale
SGLang is a fast serving framework for large language models and vision language models.
A userscript that adds danmaku to Plex Web
Optimized primitives for collective multi-GPU communication
A front-end library for displaying wechat messages with faces.