Romaosir/flash-attention
Fast and memory-efficient exact attention
Fast and memory-efficient exact attention
skills, prompts
DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
SGLang is a high-performance serving framework for large language models and multimodal models.