WeiXiongUST/Building-Math-Agents-with-Multi-Turn-Iterative-Preference-Learning
This is an official implementation of the paper ``Building Math Agents with Multi-Turn Iterative Preference Learning'' with multi-turn DPO and KTO.
Ph.D. Student in computer science at UIUC; machine learning theory and RLHF.
This is an official implementation of the paper ``Building Math Agents with Multi-Turn Iterative Preference Learning'' with multi-turn DPO and KTO.
This is the code used for the paper "PMGT-VR: A decentralized proximal-gradient algorithmic framework with variance reduction", prepint.
A pipeline to improve skills of large language models
An adaptive sampling framework for Reinforce-style LLM post training.
Codebase for Iterative DPO Using Rule-based Rewards
Recipes to train the self-rewarding reasoning LLMs.
Library of contextual bandits algorithms
An Easy-to-use, Scalable and High-performance RLHF Framework (70B+ PPO Full Tuning & Iterative DPO & LoRA & RingAttention & RFT)
Recipes to train reward model for RLHF.
Go ahead and axolotl questions
This is an official implementation of the Reward rAnked Fine-Tuning Algorithm (RAFT), also known as iterative best-of-n fine-tuning or rejection sampling fine-tuning.
Robust recipes to align language models with human and AI preferences
A curated list of resources for using LLMs to develop more competitive grant applications.
ToRA is a series of Tool-integrated Reasoning LLM Agents designed to solve challenging mathematical reasoning problems by interacting with tools [ICLR'24].