KaisennHu/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
M.Eng in CS @ HUST | Software Engineer @ Huawei | Research: Distributed Systems · Serverless · AI Systems
A high-throughput and memory-efficient inference and serving engine for LLMs
A high-performance and light-weight router for vLLM large scale deployment
A Datacenter Scale Distributed Inference Serving Framework
vLLM based inference framework for agentic workload.
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
RL training framework for diffusion and omni-modality models
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
SGLang is a high-performance serving framework for large language models and multimodal models.
An asynchronous streaming data management module for efficient post-training.
This is the design and implement of multi RL task pool-scheduler
PyTorch Single Controller