19 followers16 repositories
Repositories
Add the support for AMD GPU platform. Scripts for fine-tuning Meta Llama3 with composable FSDP & PEFT methods to cover single/multi-node GPUs. Supports default & custom datasets for applications such as summarization and Q&A. Supporting a number of candid inference solutions such as HF TGI, VLLM for local or cloud deployment.
★ 0Forks 0
A framework for efficient model inference with omni-modality models
★ 0Forks 0
AI Tensor Engine for ROCm
★ 0Forks 0
Common recipes to run vLLM on AMD GPU
★ 0JavaScriptForks 0
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
★ 0PythonForks 0
Open Source Continuous Inference Benchmarking Qwen3.5, DeepSeek, GPTOSS - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 vs H100 & soon™ TPUv6e/v7/Trainium2/3
★ 0PythonForks 0
SGLang is a fast serving framework for large language models and vision language models.
★ 0Forks 0
A high-throughput and memory-efficient inference and serving engine for LLMs
★ 0PythonForks 0
Cookbook of SGLang - Recipe
★ 0JavaScriptForks 0
MiniMax M2.1, a SOTA model for real-world dev & agents.
★ 0Forks 0
MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model.
★ 0PythonForks 0
★ 0Forks 0
Kimi K2 is the large language model series developed by Moonshot AI team
★ 0Forks 0
GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
★ 0PythonForks 0
GLM-4.5: An open-source large language model designed for intelligent agents by Z.ai
★ 0PythonForks 0
★ 0PythonForks 0