nehaprakriya/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
A high-throughput and memory-efficient inference and serving engine for LLMs
AI Tensor Engine for ROCm
SGLang is a high-performance serving framework for large language models and multimodal models.
Open Source Continuous Inference Benchmarking Kimi K2.6, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 vs H100 & soon™ TPUv6e/v7/Trainium2/3
xDiT: A Scalable Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
MiniMax-M2.5 a8w8 GEMM tuning CSV for ROCm/aiter
Minimalistic large language model 3D-parallelism training
This repository contains IPs, Vitis kernels and software APIs that can be leveraged by Vitis users to build scale-out solutions on multiple Alveo cards.
[FPGA 2021, Best Paper Award] An automated floorplanning and pipelining tool for Vivado HLS.