Thunderbeee/ZSCL
Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language Models
Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language Models
A high-throughput and memory-efficient inference and serving engine for LLMs
NVIDIA Inference Benchmarks provide recipes in ready-to-use templates for evaluating platform speed. Validate your platform across specific AI use cases across hardware and software combinations.
Research prototype of PRISM — a cost-efficient multi-LLM serving system with flexible time- and space-based GPU sharing.
Benchmark SGLang on SLURM
TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and support state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in performant way.
kvcached: Elastic KV cache for dynamic GPU sharing and efficient multi-LLM inference.
My learning notes/codes for ML SYS.
Code for MIT 6.S986 project
CUDA by practice
Reasoning with Language Model is Planning with World Model
LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models
The simplest, fastest repository for training/finetuning medium-sized GPTs.
Human-AI-Collaboration (cooking setting)