Open source inference optimization @vllm-project | @redhatofficial | @neuralmagic
Repositories
mgoin/hf_model_stats
mgoin/humming
mgoin/mgoin.github.io
mgoin/blog
Public repo for HF blog posts
mgoin/tml-fa4
FA4-based Relative Attention Kernel developed by TML and Colfax
mgoin/speculators
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
mgoin/vllm-dashboard
mgoin/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
mgoin/Mooncake
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
mgoin/ci-infra
This repo hosts code for vLLM CI & Performance Benchmark infrastructure.
mgoin/pingboard
mgoin/advos
RISC-V OS in Rust with hardware support for SiFive's HiFive1 board
mgoin/torch_bitmask
Implementations of bitmask compression for weight sparsity in PyTorch
mgoin/vllm-gguf-plugin
vLLM Quantization plugin for GGUF
mgoin/learned_indexes
Experiments on ideas proposed in Tim Kraska's "The Case for Learned Index Structures"
mgoin/InstantTensor
An ultra-fast, distributed Safetensors loader
mgoin/benchmark-arena
mgoin/meTile
python-based eDSL for efficient Metal Shading Language code generation
mgoin/shepherd
mgoin/binfer
inference in bun
mgoin/dvc.org
🔗 DVC website and documentation
mgoin/prime-rl
Async RL Training at Scale
mgoin/SpecForge
Train speculative decoding models effortlessly and port them smoothly to SGLang serving.
mgoin/loc-blocks-vllm
mgoin/vllm-ci-status
mgoin/flashinfer
FlashInfer: Kernel Library for LLM Serving
mgoin/llm-d
llm-d is a Kubernetes-native high-performance distributed LLM inference framework