slin1237/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
As imagine the future
A high-throughput and memory-efficient inference and serving engine for LLMs
Shepherd Model Gateway
Lightweight coding agent that runs in your terminal
/home/tpounds
OME is a Kubernetes operator for enterprise-grade management and serving of Large Language Models (LLMs)
Run LLMs with MLX
A retargetable MLIR-based machine learning compiler and runtime toolkit.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
SGLang is a fast serving framework for large language models and vision language models.
Templating rendering and generation parsing for Cohere models.
HF transformer in rust
A High Performance Metadata System for Kubernetes
🙃 A delightful community-driven (with 1,300+ contributors) framework for managing your zsh configuration. Includes 200+ optional plugins (rails, git, OSX, hub, capistrano, brew, ant, php, python, etc), over 140 themes to spice up your morning, and an auto-update tool so that makes it easy to keep up with the latest updates from the community.
JobSet: a k8s native API for distributed ML training and HPC workloads
Production-Grade Container Scheduling and Management
repo for PRC hackathon - kickstart frontend
my own practices of leetcode
My own leetcode solutions by python