PhD student UCSD CSE, BS ZJU CS
Repositories
Viol2000/LR-private
Viol2000/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Viol2000/TensorRT-LLM
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and build TensorRT engines that contain state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that execute those TensorRT engines.
Viol2000/flash-attention-lookahead
Viol2000/sglang-rebase
SGLang is a fast serving framework for large language models and vision language models.
Viol2000/FastChat-llama3
Viol2000/lm-sys.github.io
Viol2000/fairseq
Facebook AI Research Sequence-to-Sequence Toolkit written in Python.
Viol2000/PatrickStar
PatrickStar enables Larger, Faster, Greener Pretrained Models for NLP and democratizes AI for everyone.