FFY0/DefensiveKV
Official Implementation for [ICLR26] DefensiveKV: Taming the Fragility of KV Cache Eviction in LLM Inference
PhD Candidate @ USTC
Official Implementation for [ICLR26] DefensiveKV: Taming the Fragility of KV Cache Eviction in LLM Inference
The Official Implementation of Ada-KV [NeurIPS 2025]
SGLang is a high-performance serving framework for large language models and multimodal models.
Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as well as generating speech in real time.
The agent that grows with you
The Official Implementation of LegoSketch [ICML 2025]
SourceCode for MetaSketch
Open-source infrastructure for Computer-Use Agents. Sandboxes, SDKs, and benchmarks to train and evaluate AI agents that can control full desktops (macOS, Linux, Windows).
Github Pages template for academic personal websites, forked from mmistakes/minimal-mistakes
The fastest repo in history to surpass 50K stars ⭐, reaching the milestone in just 2 hours after publication. Better Harness Tools that make real things done. Now writing in Rust using oh-my-codex.
A custom implementation of [AdaKV](https://arxiv.org/abs/2407.11550) under the open-source project NVIDIA/kvpress. Original AdaKV code can be found in [repo](https://github.com/FFY0/AdaKV). Additionally, [NVIDIA/kvpress](https://github.com/NVIDIA/kvpress) also offers a simpler method to simulate the performance of AdaKV.
A high-throughput and memory-efficient inference and serving engine for LLMs
KV cache compression for high-throughput LLM inference