aJupyter/tiny-llm

(๐Ÿšง WIP) a course of LLM inference serving on Apple Silicon for systems engineers.

โ˜… 1Forks 0GitHub โ†—Compare

Project website โ†—

README

tiny-llm - LLM Serving in a Week

CI (main)

Still WIP and in very early stage. A tutorial on LLM serving using MLX for system engineers. The codebase is solely (almost!) based on MLX array/matrix APIs without any high-level neural network APIs, so that we can build the model serving infrastructure from scratch and dig into the optimizations.

The goal is to learn the techniques behind efficiently serving a large language model (i.e., Qwen2 models).

Why MLX: nowadays it's easier to get a macOS-based local development environment than setting up an NVIDIA GPU.

Why Qwen2: this was the first LLM I've interacted with -- it's the go-to example in the vllm documentation. I spent some time looking at the vllm source code and built some knowledge around it.

Book

The tiny-llm book is available at https://skyzh.github.io/tiny-llm/. You can follow the guide and start building.

Community

You may join skyzh's Discord server and study with the tiny-llm community.

Join skyzh's Discord Server

Roadmap

Week + Chapter Topic Code Test Doc
1.1 Attention โœ… โœ… โœ…
1.2 RoPE โœ… โœ… โœ…
1.3 Grouped Query Attention โœ… โœ… โœ…
1.4 RMSNorm and MLP โœ… ๐Ÿšง ๐Ÿšง
1.5 Transformer Block โœ… ๐Ÿšง ๐Ÿšง
1.6 Load the Model โœ… ๐Ÿšง ๐Ÿšง
1.7 Generate Responses (aka Decoding) โœ… โœ… ๐Ÿšง
2.1 Key-Value Cache โœ… ๐Ÿšง ๐Ÿšง
2.2 Quantized Matmul and Linear - CPU โœ… ๐Ÿšง ๐Ÿšง
2.3 Quantized Matmul and Linear - GPU โœ… ๐Ÿšง ๐Ÿšง
2.4 Flash Attention 2 - CPU โœ… ๐Ÿšง ๐Ÿšง
2.5 Flash Attention 2 - GPU โœ… ๐Ÿšง ๐Ÿšง
2.6 Continuous Batching โœ… ๐Ÿšง ๐Ÿšง
2.7 Chunked Prefill โœ… ๐Ÿšง ๐Ÿšง
3.1 Paged Attention - Part 1 ๐Ÿšง ๐Ÿšง ๐Ÿšง
3.2 Paged Attention - Part 2 ๐Ÿšง ๐Ÿšง ๐Ÿšง
3.3 MoE (Mixture of Experts) ๐Ÿšง ๐Ÿšง ๐Ÿšง
3.4 Speculative Decoding ๐Ÿšง ๐Ÿšง ๐Ÿšง
3.5 Prefill-Decode Separation (requires two Macintosh devices) ๐Ÿšง ๐Ÿšง ๐Ÿšง
3.6 Parallelism ๐Ÿšง ๐Ÿšง ๐Ÿšง
3.7 AI Agent / Tool Calling ๐Ÿšง ๐Ÿšง ๐Ÿšง

Other topics not covered: quantized/compressed kv cache, prefix/prompt cache; sampling, fine tuning; smaller kernels (softmax, silu, etc)

Contributors

skyzhConnor1996shenxiangzhuangshivangsharma1

Issues