Zhewen Li

@zhewenl · User

GitHub profile ↗ · Compare

@Inferact | @vllm-project

InferactSF Bay Areea63 followers18 repositories

Repositories

zhewenl/Mooncake

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

★ 0C++Forks 0

zhewenl/nixl

NVIDIA Inference Xfer Library (NIXL)

★ 0Forks 0

zhewenl/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0PythonForks 0

zhewenl/router

A high-performance and light-weight router for vLLM large scale deployment

★ 0RustForks 0

zhewenl/deep-swe

Measuring frontier coding agents on original, long-horizon engineering tasks

★ 0Forks 1

zhewenl/claude-code

Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.

★ 0Forks 0

zhewenl/everything-claude-code

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

★ 0Forks 0

zhewenl/superpowers

An agentic skills framework & software development methodology that works.

★ 0Forks 0

zhewenl/Model-Optimizer

A unified library of SOTA model optimization techniques like quantization, pruning, distillation, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

★ 0Forks 0

zhewenl/clawdbot

Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

★ 0Forks 0

zhewenl/sglang

SGLang is a fast serving framework for large language models and vision language models.

★ 0Forks 0

zhewenl/ci-infra

This repo hosts code for vLLM CI & Performance Benchmark infrastructure.

★ 0HCLForks 0

zhewenl/test-infra

This repository hosts code that supports the testing infrastructure for the PyTorch organization. For example, this repo hosts the logic to track disabled tests and slow tests, as well as our continuation integration jobs HUD/dashboard.

★ 0TypeScriptForks 0

zhewenl/pytorch

Tensors and Dynamic neural networks in Python with strong GPU acceleration

★ 0Forks 0

zhewenl/dynamo

A Datacenter Scale Distributed Inference Serving Framework

★ 0Forks 0