Wei Xiong

@WeiXiongUST · User

GitHub profile ↗ · Compare

Ph.D. Student in computer science at UIUC; machine learning theory and RLHF.

206 followers45 repositories

Repositories

WeiXiongUST/OpenRLHF

An Easy-to-use, Scalable and High-performance RLHF Framework (70B+ PPO Full Tuning & Iterative DPO & LoRA & RingAttention & RFT)

★ 0PythonForks 0

WeiXiongUST/RAFT

This is an official implementation of the Reward rAnked Fine-Tuning Algorithm (RAFT), also known as iterative best-of-n fine-tuning or rejection sampling fine-tuning.

★ 0Forks 0

WeiXiongUST/ToRA

ToRA is a series of Tool-integrated Reasoning LLM Agents designed to solve challenging mathematical reasoning problems by interacting with tools [ICLR'24].

★ 1PythonForks 1