Hyperion-shuo/picotron
Minimalistic 4D-parallelism distributed training framework for education purpose
Reinforcement learning, Game AI, LLM
Minimalistic 4D-parallelism distributed training framework for education purpose
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
A Next-Generation Training Engine Built for Ultra-Large MoE Models
A lightweight framework for building LLM-based agents
Skills for threat modeling, scanning, triage, patching, plus an autonomous scanning harness you can /customize
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
The official code repo of paper "Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training"
Student version of Assignment 1 for Stanford CS336 - Language Modeling From Scratch
Student version of Assignment 2 for Stanford CS336 - Language Modeling From Scratch
One leetcode everyday(python solution)
Implementation of advantage-weighted regression.
An Easy-to-use, Scalable and High-performance RLHF Framework based on Ray (PPO & GRPO & REINFORCE++ & TIS & vLLM & Ray & Dynamic Sampling & Async Agentic RL)
slime is an LLM post-training framework for RL Scaling.
Implement a reasoning LLM in PyTorch from scratch, step by step
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
The best ChatGPT that $100 can buy.
Quick illustration of how one can easily read books together with LLMs. It's great and I highly recommend it.
The simplest, fastest repository for training/finetuning medium-sized GPTs.
off-policy RL on long sequences
😼 优雅地使用基于 clash/mihomo 的代理环境
Build an email assistant with human-in-the-loop and memory
Domain Randomization via Entropy Maximization
Minimal, clean code for the Byte Pair Encoding (BPE) algorithm commonly used in LLM tokenization.
Example code for Fluent Python, 2nd edition (O'Reilly 2022)