yuezhu1/spyre-inference
vLLM plugin for Spyre based on torch-spyre
vLLM plugin for Spyre based on torch-spyre
Documents describing interfaces between compiler/runtime and device
A high-throughput and memory-efficient inference and serving engine for LLMs
Intelligent Router for Mixture-of-Models
Systematic and comprehensive benchmarks for LLM systems.
Supercharge Your LLM with the Fastest KV Cache Layer
A RDMA-enabled Distributed Persistent Memory File System
User space POSIX-like file system in main memory
Models and examples built with TensorFlow
Computation using data flow graphs for scalable machine learning