Qi Yuhang

@HydraQYH · User

GitHub profile ↗ · Compare

NV/AMD GPU Kernel Performance Analysis & Optimization. Now in Bytedance(Data => AML => Seed). Used to be in AIACC Team@Alibaba Cloud.

BytedanceHangzhou, China183 followers33 repositories

Repositories

HydraQYH/hp_rms_norm

High performance RMSNorm Implement by using SM Core Storage(Registers and Shared Memory)

★ 31CudaForks 1

HydraQYH/FlyDSL

FlyDSL is the Python front‑end of the project: Flexible LaYout DSL.

★ 0PythonForks 0

HydraQYH/sglang

SGLang is a fast serving framework for large language models and vision language models.

★ 2PythonForks 0

HydraQYH/tilelang_tvm

Open deep learning compiler stack for cpu, gpu and specialized accelerators

★ 0PythonForks 0

HydraQYH/tilelang

Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels

★ 0PythonForks 0

HydraQYH/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0PythonForks 0

HydraQYH/DeepGEMM

DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling

★ 0CudaForks 0

HydraQYH/mlc

Homework for mlc.ai.

★ 0Jupyter NotebookForks 0