nil0x9/flash-muon
Flash-Muon: An Efficient Implementation of Muon Optimizer
LLM Pretrain
Flash-Muon: An Efficient Implementation of Muon Optimizer
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
A tutorial on modern GPU programming for machine learning systems
Learning TileLang with 10 puzzles!
A hackable markdown, Typst, latex, html(inline) & Asciidoc previewer for Neovim
high-performance linear attention kernel library built on TileLang
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Efficient Triton Kernels for LLM Training
OpenMMLab Foundational Library for Training Deep Learning Models
A Next-Generation Training Engine Built for Ultra-Large MoE Models
Minimal pretraining script for language modeling in PyTorch. Supporting torch compilation and DDP. It includes a model implementation and a data preprocessing.
Muon: An optimizer for hidden layers in neural networks
🤗 Transformers: State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX.
The simplest, fastest repository for training/finetuning medium-sized GPTs.
Personal Transformer models training library