none0663/slime
slime is a LLM post-training framework for RL Scaling.
slime is a LLM post-training framework for RL Scaling.
Your own personal AI assistant. Any OS. Any Platform. The lobster way. ๐ฆ
Config files for my GitHub profile.
veRL: Volcano Engine Reinforcement Learning for LLM
Bridge Megatron-Core to Hugging Face/Reinforcement Learning
Ongoing research training transformer models at scale
An Easy-to-use, Scalable and High-performance RLHF Framework (70B+ PPO Full Tuning & Iterative DPO & LoRA & RingAttention & RFT)
A high-performance distributed training framework for Reinforcement Learning