Mingyuan MA

@Thunderbeee · User

GitHub profile ↗ · Compare

UC BerkeleySan Francisco30 followers37 repositories

Repositories

Thunderbeee/ZSCL

Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language Models

★ 111PythonForks 9

Thunderbeee/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0PythonForks 0

Thunderbeee/nv-srt-slurm

NVIDIA Inference Benchmarks provide recipes in ready-to-use templates for evaluating platform speed. Validate your platform across specific AI use cases across hardware and software combinations.

★ 0Forks 0

Thunderbeee/multi-llm-serving

Research prototype of PRISM — a cost-efficient multi-LLM serving system with flexible time- and space-based GPU sharing.

★ 0Forks 0

Thunderbeee/TensorRT-LLM

TensorRT-LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and support state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT-LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in performant way.

★ 0C++Forks 0

Thunderbeee/kvcached

kvcached: Elastic KV cache for dynamic GPU sharing and efficient multi-LLM inference.

★ 0PythonForks 0

Thunderbeee/RAP

Reasoning with Language Model is Planning with World Model

★ 0PDDLForks 0

Thunderbeee/LLM-Adapters

LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models

★ 0PythonForks 0

Thunderbeee/nanoGPT

The simplest, fastest repository for training/finetuning medium-sized GPTs.

★ 0Forks 0