leo-pony/vllm-ascend
Community maintained hardware plugin for vLLM on Ascend
Community maintained hardware plugin for vLLM on Ascend
A high-throughput and memory-efficient inference and serving engine for LLMs
PR-level bisect tool for vllm-ascend regression detection
verl: Volcano Engine Reinforcement Learning for LLMs
githubGramaTest
Docker Model Runner
This repo hosts code for vLLM CI & Performance Benchmark infrastructure.
An example pipeline that runs a simple Bash script with artifacts and inline output.
Get up and running with Llama 3, Mistral, Gemma 2, and other large language models.
LLM inference in C/C++
The repository provides docs & api & tutorials & FAQ and all things like that.
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
The goal of RamaLama is to make working with AI boring.
LLaMA Factory Document
Shortnames project is collecting registry alias names for shortnames to fully specified container image names.
ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
https://cosdt.github.io
KV and LBA SSD userspace NVMe driver