ziying

@z1ying ยท User

GitHub profile โ†— ยท Compare

Full-stack engineer exploring AI infra & open source ๐ŸŒฑ

San Francisco15 followers15 repositories

Repositories

z1ying/aibrix

Cost-efficient and pluggable Infrastructure components for GenAI inference

โ˜… 0Forks 0

z1ying/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

โ˜… 0PythonForks 0

z1ying/slime

slime is an LLM post-training framework for RL Scaling.

โ˜… 0Forks 0

z1ying/triton

Development repository for the Triton language and compiler

โ˜… 0MLIRForks 0

z1ying/RL-Kernel

Modern RL Post-training Infrastructure: Optimized for NVIDIA/AMD GPUs with a focus on vLLM and DeepSpeed integration, CUDA/ROCm/Triton kernels, and transparent hardware-aware scaling.

โ˜… 0Forks 0

z1ying/verl

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

โ˜… 0Forks 0

z1ying/unsloth

Unsloth Studio is a web UI for training and running open models like Gemma 4, Qwen3.6, DeepSeek, gpt-oss locally.

โ˜… 0Forks 0

z1ying/TensorRT-LLM

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

โ˜… 0Forks 0

z1ying/MiniCPM

MiniCPM4 & MiniCPM4.1: Ultra-Efficient LLMs on End Devices, achieving 3+ generation speedup on reasoning tasks

โ˜… 0Forks 0

z1ying/LMCache

Supercharge Your LLM with the Fastest KV Cache Layer

โ˜… 0PythonForks 0

z1ying/sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

โ˜… 0Forks 0

z1ying/vllm-omni

A framework for efficient model inference with omni-modality models

โ˜… 0PythonForks 0

z1ying/autoRAG

Automatic development for retrieval augmented generation system

โ˜… 0Forks 0