Jinyan Chen

@liz-badada · User

GitHub profile ↗ · Compare

Hahahahahahahaha

NVIDIAShanghai, China25 followers39 repositories

Repositories

liz-badada/aisimulate

AISimulate predicts LLM serving behavior and searches for strong deployment configurations offline, without bringing up a GPU serving cluster

★ 0PythonForks 0

liz-badada/sglang

SGLang is a fast serving framework for large language models and vision language models.

★ 0PythonForks 0

liz-badada/dynamo

A Datacenter Scale Distributed Inference Serving Framework

★ 0RustForks 0

liz-badada/beat-ai

<Beat AI> 又名 <零生万物> , 是一本专属于软件开发工程师的 AI 入门圣经,手把手带你上手写 AI。从神经网络到大模型,从高层设计到微观原理,从工程实现到算法,学完后,你会发现 AI 也并不是想象中那么高不可攀、无法战胜,Just beat it !

★ 0Forks 0

liz-badada/DeepGEMM

DeepGEMM: clean and efficient FP8 GEMM kernels with fine-grained scaling

★ 0CudaForks 0

liz-badada/verl

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

★ 0PythonForks 0

liz-badada/ptx-isa-markdown

PTX ISA 9.1 documentation converted to searchable markdown. Includes Claude Code skill for CUDA development.

★ 0Forks 0

liz-badada/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0PythonForks 0

liz-badada/TileGym

Helpful kernel tutorials and examples for tile-based GPU programming

★ 0Forks 0

liz-badada/cutlass

CUDA Templates and Python DSLs for High-Performance Linear Algebra

★ 0C++Forks 0

liz-badada/AISystem

AISystem 主要是指AI系统,包括AI芯片、AI编译器、AI推理和训练框架等AI全栈底层技术

★ 0Forks 0

liz-badada/TensorRT-LLM

TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and support state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in performant way.

★ 0Forks 0

liz-badada/Mooncake

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

★ 0Forks 0

liz-badada/3FS

A high-performance distributed file system designed to address the challenges of AI training and inference workloads.

★ 0Forks 0

liz-badada/DeepEP

DeepEP: an efficient expert-parallel communication library

★ 0CudaForks 0