CHEN Xi

@RubiaCx · User

GitHub profile ↗ · Compare

Hi~ o(* ̄▽ ̄*)ブ

@BABABeijing, China65 followers44 repositories

Repositories

RubiaCx/Megatron-LM

Ongoing research training transformer language models at scale, including: BERT & GPT-2

★ 0Forks 0

RubiaCx/MagiAttention

A Distributed Attention Towards Linear Scalability for Ultra-Long Context, Heterogeneous Data Training

★ 0Forks 0

RubiaCx/NotionNext

一个使用 NextJS + Notion API 实现的,部署在 Vercel 上的静态博客系统。

★ 0JavaScriptForks 0

RubiaCx/sglang

SGLang is a fast serving framework for large language models and vision language models.

★ 0PythonForks 0

RubiaCx/claude-code

An independent Python feature port of Claude Code, entirely rewritten from scratch. Educational Purpose only.

★ 0Forks 0

RubiaCx/FastVideo

A unified inference and post-training framework for accelerated video generation.

★ 0Forks 0

RubiaCx/LightX2V

Light Image Video Generation Inference Framework

★ 0PythonForks 0

RubiaCx/TransformerEngine

A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.

★ 0Forks 0

RubiaCx/infra-skills

A collection of specialized agent skills for AI infrastructure development, enabling Claude Code to write, optimize, and debug high-performance systems.

★ 0Forks 0

RubiaCx/x-attention

[ICML 2025] XAttention: Block Sparse Attention with Antidiagonal Scoring

★ 0Forks 0

RubiaCx/tilelang

Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels

★ 0C++Forks 0

RubiaCx/slime

slime is an LLM post-training framework for RL Scaling.

★ 0PythonForks 0

RubiaCx/quack

A Quirky Assortment of CuTe Kernels

★ 0PythonForks 0

RubiaCx/CogVideo

text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)

★ 0PythonForks 0

RubiaCx/FlagGems

FlagGems is an operator library for large language models implemented in Triton Language.

★ 0PythonForks 0

RubiaCx/tritonbench

Tritonbench is a collection of PyTorch custom operators with example inputs to measure their performance.

★ 0PythonForks 0

RubiaCx/SageAttention

Quantized Attention that achieves speedups of 2.1-3.1x and 2.7-5.1x compared to FlashAttention2 and xformers, respectively, without lossing end-to-end metrics across various models.

★ 0CudaForks 0

RubiaCx/CUDA-Learn-Notes

📚Tensor/CUDA Cores, 📖150+ CUDA Kernels, ⚡️⚡️toy-hgemm library with WMMA, MMA and CuTe (98%~100% TFLOPS of cuBLAS 🎉🎉).

★ 0Forks 0