Tim-Siu/reft-exp
A research repo for experiments about Reinforcement Finetuning
@ NUS
A research repo for experiments about Reinforcement Finetuning
Code repo for "Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning"
Official repository for the paper "Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation"
Harbor is a framework for running agent evaluations and creating and using RL environments.
Public asset host for OPSD comprehensive report figures (rq_12_full_report)
Auto-refreshing token proxy for Feishu MCP Server
The assignment from a frontier LLM lab. Let's do something to nanochat!
verl: Volcano Engine Reinforcement Learning for LLMs
The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench verified!
slime is an LLM post-training framework for RL Scaling.
Democratizing Reinforcement Learning for LLMs
ICLR 25 SLLM: The Surprising Effectiveness of Randomness in LLM Pruning
Lighteval is your all-in-one toolkit for evaluating LLMs across multiple backends
Train transformer language models with reinforcement learning.
Fully open reproduction of DeepSeek-R1
MarkBind is a tool for generating content-heavy websites from source files in Markdown format
NUS CP4101 B.Comp. Dissertation (Final Year Project Report) Format
Prune transformer layers
Model Zoo For OpenCV DNN and Benchmarks.
A practice project on fine-tuning a segmentation network.
This is the project website for the TEAMMATES feedback management tool for education