njhill/flashinfer
FlashInfer: Kernel Library for LLM Serving
FlashInfer: Kernel Library for LLM Serving
A high-throughput and memory-efficient inference and serving engine for LLMs
This repo hosts code for vLLM CI & Performance Benchmark infrastructure.
Abstracted helper classes providing consistent key-value store functionality, with zookeeper and etcd3 implementations
Alternative etcd3 java client
NVIDIA Inference Benchmarks provide recipes in ready-to-use templates for evaluating platform speed. Validate your platform across specific AI use cases across hardware and software combinations.
High-performance safetensors model loader
vLLM adapter for a TGIS-compatible gRPC server.
Redis for LLMs
Large Language Model Text Generation Inference
Python library for Synthetic Data Generation
Fast and memory-efficient exact attention
🏎️ Accelerate training and inference of 🤗 Transformers with easy to use hardware optimization tools
🤗 Transformers: State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX.
Fast Inference Solutions for BLOOM
Create, manipulate, apply and delete Kubernetes resource manifests at runtime
An inference server for your machine learning models, including support for multiple frameworks, multi-model serving and more
Distributed Model Serving Framework
Operator to install IBM Common Services
Unified runtime-adapter image of the sidecar containers which run in the modelmesh pods
KServe V2 Protocol Rest API Implementation Proxy
Serverless Inferencing on Kubernetes
High-performance netty and thrift-based microservice RPC library for Java
Distributed reliable key-value store for the most critical data of a distributed system
A ConcurrentLinkedHashMap for Java
A flexible, high-performance serving system for machine learning models
Argo Workflows: Get stuff done with Kubernetes.
Lightweight Kubernetes controllers as a service
The C based gRPC (C++, Python, Ruby, Objective-C, PHP, C#)