Yuan Feng

@FFY0 · User

GitHub profile ↗ · Compare

PhD Candidate @ USTC

30 followers14 repositories

Repositories

FFY0/DefensiveKV

Official Implementation for [ICLR26] DefensiveKV: Taming the Fragility of KV Cache Eviction in LLM Inference

★ 59PythonForks 4

FFY0/AdaKV

The Official Implementation of Ada-KV [NeurIPS 2025]

★ 140PythonForks 8

FFY0/sglang

SGLang is a high-performance serving framework for large language models and multimodal models.

★ 0Forks 0

FFY0/Qwen3-Omni

Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as well as generating speech in real time.

★ 0Forks 0

FFY0/cua

Open-source infrastructure for Computer-Use Agents. Sandboxes, SDKs, and benchmarks to train and evaluate AI agents that can control full desktops (macOS, Linux, Windows).

★ 0Forks 0

FFY0/yuanfeng.github.io

Github Pages template for academic personal websites, forked from mmistakes/minimal-mistakes

★ 0JavaScriptForks 0

FFY0/claw-code

The fastest repo in history to surpass 50K stars ⭐, reaching the milestone in just 2 hours after publication. Better Harness Tools that make real things done. Now writing in Rust using oh-my-codex.

★ 0Forks 0

FFY0/AdaKV-in-NVIDIA-kvpress

A custom implementation of [AdaKV](https://arxiv.org/abs/2407.11550) under the open-source project NVIDIA/kvpress. Original AdaKV code can be found in [repo](https://github.com/FFY0/AdaKV). Additionally, [NVIDIA/kvpress](https://github.com/NVIDIA/kvpress) also offers a simpler method to simulate the performance of AdaKV.

★ 3PythonForks 0

FFY0/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0Forks 0