sychen52

@sychen52 · User

GitHub profile ↗ · Compare

7 followers25 repositories

Repositories

sychen52/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0PythonForks 0

sychen52/Model-Optimizer

A unified library of state-of-the-art model optimization techniques like quantization, pruning, distillation, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM or TensorRT to optimize inference speed.

★ 0Forks 0

sychen52/TensorRT-LLM

TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and support state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in performant way.

★ 0Forks 0

sychen52/gF-python-traceback

A vim plugin that allow you to jump to file with line number based on python traceback messages

★ 3Vim scriptForks 0

sychen52/ray

An open source framework that provides a simple, universal API for building distributed applications. Ray is packaged with RLlib, a scalable reinforcement learning library, and Tune, a scalable hyperparameter tuning library.

★ 0Forks 0