QiuMike/VMsvga2ForQEMU
This is the VMware svga graphic card OSX driver for QEMU
This is the VMware svga graphic card OSX driver for QEMU
MacOS inside a Docker container.
SGLang is a fast serving framework for large language models and vision language models.
The source of LMSYS website and blogs
Tile-Based Runtime for Ultra-Low-Latency LLM Inference
A framework for efficient model inference with omni-modality models
LLM API 管理 & 分发系统,支持 OpenAI、Azure、Anthropic Claude、Google Gemini、DeepSeek、字节豆包、ChatGLM、文心一言、讯飞星火、通义千问、360 智脑、腾讯混元等主流模型,统一 API 适配,可用于 key 管理与二次分发。单可执行文件,提供 Docker 镜像,一键部署,开箱即用。LLM API management & key redistribution system, unifying multiple providers under a single API. Single binary, Docker-ready, with an English UI.
支持 AnyRouter、AgentRouter 的多平台多账号签到,理论兼容所有基于 NewAPI、OneAPI 的平台。
An independent Python feature port of Claude Code, entirely rewritting from scratch using oh-my-codex. Educational Purpose only.
A high-throughput and memory-efficient inference and serving engine for LLMs
TradingAgents: Multi-Agents LLM Financial Trading Framework
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
Community maintained hardware plugin for vLLM on Ascend
[NeurIPS'23] H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models.
A fast GPU memory copy library based on NVIDIA GPUDirect RDMA technology
Super-Efficient RLHF Training of LLMs with Parameter Reallocation
Unified KV Cache Compression Methods for Auto-Regressive Models
Documentation of NVIDIA chip/hardware interfaces
OpenR: An Open Source Framework for Advanced Reasoning with Large Language Models
Unified Efficient Fine-Tuning of 100+ LLMs (ACL 2024)
Artifacts for our NSDI'23 paper TGS
CUDA checkpoint and restore utility
Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity
KvikIO - High Performance File IO
PyTorch DataLoaders implemented with DALI for accelerating image preprocessing
A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications.