Misha Goin

@mgoin · User

GitHub profile ↗ · Compare

Open source inference optimization @vllm-project | @redhatofficial | @neuralmagic

@vllm-project @redhatofficialBoston535 followers86 repositories

Repositories

mgoin/blog

Public repo for HF blog posts

★ 0Forks 0

mgoin/tml-fa4

FA4-based Relative Attention Kernel developed by TML and Colfax

★ 0Forks 0

mgoin/speculators

A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM

★ 0PythonForks 0

mgoin/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

★ 0PythonForks 0

mgoin/Mooncake

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

★ 0C++Forks 0

mgoin/ci-infra

This repo hosts code for vLLM CI & Performance Benchmark infrastructure.

★ 0Forks 0

mgoin/advos

RISC-V OS in Rust with hardware support for SiFive's HiFive1 board

★ 1RustForks 1

mgoin/torch_bitmask

Implementations of bitmask compression for weight sparsity in PyTorch

★ 5PythonForks 2

mgoin/learned_indexes

Experiments on ideas proposed in Tim Kraska's "The Case for Learned Index Structures"

★ 10PythonForks 0

mgoin/meTile

python-based eDSL for efficient Metal Shading Language code generation

★ 1Forks 0

mgoin/SpecForge

Train speculative decoding models effortlessly and port them smoothly to SGLang serving.

★ 0Forks 0

mgoin/llm-d

llm-d is a Kubernetes-native high-performance distributed LLM inference framework

★ 0Forks 0