YangRui2015/RiC
Code for the ICML 2024 paper "Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment"
Do less and do better.
Code for the ICML 2024 paper "Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference Adjustment"
Code for NeurIPS 2024 paper "Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs"
Code for ICLR 2022 paper Rethinking Goal-Conditioned Supervised Learning and Its Connection to Offline RL.
Code for NeurIPS 2022 paper "Robust offline Reinforcement Learning via Conservative Smoothing"
[ICLR 2024 Spotlight] Code for ICLR 2024 paper "Towards Robust Offline Reinforcement Learning under Diverse Data Corruption"
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
Implement many Sparse Reward algorithms in Gym Fetch environment
EmbodiedBench will be used for CVPR 2026 Workshop on Foundation Models Meet Embodied Agents (FMEA) Challenges
How to create a challenge on EvalAI?
ScaleCUA is the open-sourced computer use agents that can operate on cross-platform environments (Windows, macOS, Ubuntu, Android).
RLinf: Reinforcement Learning Infrastructure for Embodied and Agentic AI
2048 environment for Reinforcement Learning and DQN algorithm
A scalable asynchronous reinforcement learning implementation with in-flight weight updates.
Modular-HER is revised from OpenAI baselines and supports many improvements for Hindsight Experience Replay as modules.
A python tool to make a beautiful timeline by writing a txt file like diary.
Code for the ICML 2023 paper "What is Essential for Unseen Goal Generalization of Offline Goal-conditioned RL?".
Building a comprehensive and handy list of papers for GUI agents
Model-based Hindsight Experience Replay
GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
Course Material for the UG Course COMP4901Y
Hi there
RewardBench: the first evaluation tool for reward models.
add test function and my trained models to the original codes of random-network-distillation of openai