Wenxuan Li

@HalberdOfPineapple · User

GitHub profile ↗ · Compare

Research Intern @ MSRA | MPhil ACS @ University of Cambridge

University of Cambridge59 followers74 repositories

Repositories

HalberdOfPineapple/h100_gemm

A series of high-performance GEMM (General Matrix Multiply) implementations Iteratively optimised for H100 GPUs in Pure CUDA.

★ 0CudaForks 0

HalberdOfPineapple/FOCUS

[ICML 2026] Official implementation of "FOCUS: DLLMs Know How to Tame Their Compute Bound".

★ 0PythonForks 0

HalberdOfPineapple/sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

★ 0PythonForks 0

HalberdOfPineapple/MInference

[NeurIPS'24 Spotlight] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.

★ 0PythonForks 0

HalberdOfPineapple/tilelang

Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels

★ 0Forks 0

HalberdOfPineapple/LA-MCTS

High dimensional black-box optimizer using Latent Action Monte Carlo Tree Search algorithm

★ 0PythonForks 0