Master student at Zhejiang University, interested in data management, machine learning system
Repositories
KevinZeng08/flexible-flash-attention
Fast and memory-efficient exact attention
KevinZeng08/agent-gpu-skills
Personal Extensions for Agentic GPU Programming Skills
KevinZeng08/cuLA
CUDA kernels for linear attention variants, written in CuTe DSL and CUTLASS C++.
KevinZeng08/microbench-blackwell
KevinZeng08/MagiAttention
A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Training
KevinZeng08/DeepEP
DeepEP: an efficient expert-parallel communication library
KevinZeng08/nsys-parquet-perfetto-skill
Codex skill and Rust DataFusion converter for exporting Nsight Systems reports to Perfetto JSON
KevinZeng08/comm-microbench
Micro Benchmarks for GPU Communication
KevinZeng08/slime
slime is an LLM post-training framework for RL Scaling.
KevinZeng08/nccl
Optimized primitives for collective multi-GPU communication
KevinZeng08/vllm-omni
A framework for efficient model inference with omni-modality models
KevinZeng08/nvshmem
NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmers to perform one-sided communication from within CUDA kernels and on CUDA streams.
KevinZeng08/DeepGEMM
DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling
KevinZeng08/cutlass
CUDA Templates and Python DSLs for High-Performance Linear Algebra
KevinZeng08/flash-linear-attention
🚀 Efficient implementations of state-of-the-art linear attention models
KevinZeng08/verl
verl: Volcano Engine Reinforcement Learning for LLMs
KevinZeng08/lingbot-va
Causal video-action world model for generalist robot control
KevinZeng08/AReaL
Lightning-Fast RL for LLM Reasoning and Agents. Made Simple & Flexible.
KevinZeng08/FlashMLA
FlashMLA: Efficient Multi-head Latent Attention Kernels
KevinZeng08/Awesome-ML-SYS-Tutorial
My learning notes for ML SYS.
KevinZeng08/slime-agentic
A project implementing various agentic RL based on the Slime post-training framework
KevinZeng08/ROLL
An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
KevinZeng08/NVIDIA-Hopper-Benchmark
KevinZeng08/LightRFT
LightRFT (Light Reinforcement Fine-Tuning) is an advanced reinforcement learning fine-tuning framework designed for Large Language Models (LLMs) and Vision-Language Models (VLMs).
KevinZeng08/flashinfer
FlashInfer: Kernel Library for LLM Serving
KevinZeng08/zju-ai-course
KevinZeng08/LongCA-bench
KevinZeng08/tilelang
Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
KevinZeng08/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs