First-year PHD student at UChicago. Working on @LMCache. [email protected]
Repositories
YaoJiayi/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
YaoJiayi/rag
This NVIDIA RAG blueprint serves as a reference solution for a foundational Retrieval Augmented Generation (RAG) pipeline.
YaoJiayi/utils
YaoJiayi/lmcache-vllm
The driver for LMCache core to run in vLLM
YaoJiayi/production-stack
Scale from single vLLM instance to distributed vLLM deployment without changing any application code.
YaoJiayi/LMCache
Prefill LLMs only once, re-use KV across instances
YaoJiayi/YaoJiayi.github.io
YaoJiayi/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
YaoJiayi/transformers_fuse
YaoJiayi/lmcache-server
YaoJiayi/text-generation-inference
Large Language Model Text Generation Inference
YaoJiayi/llvm-lang
YaoJiayi/oneflow
OneFlow is a deep learning framework designed to be user-friendly, scalable and efficient.
YaoJiayi/loghub
A large collection of system log datasets for AI-powered log analytics