yzlu0917/tf

★ 0Forks 0PythonGitHub ↗Compare

README

Training-Free GRPO (arXiv:2510.08191) — small-scale reproduction + a practical improvement

最新进展 / 最新结果:请直接看 PROGRESS.md(包含 SVAMP / GSM8K / ASDiv / MultiArith 的 multi-seed 结果、cost/quality Pareto、uncertainty gate、trigger filter、blacklist 等图表与结论)。

目标:基于 Training-Free Group Relative Policy Optimization(Training-Free GRPO, arXiv:2510.08191, 2025-10-09)做一个可在小模型/本地可跑的复现骨架,并针对论文里“experience library / token prior”这条路线提出一个更稳、更可落地的改进点,然后用本地 Qwen3-0.6B 做一轮实验验证。

这份仓库当前聚焦 数学推理(GSM8K),后续可以无缝迁移到 API 大模型做更强的 out-of-domain / tool-use 实验。


选题与 insight

Training-Free GRPO 的核心是:不更新模型参数,通过“组内相对优势”的信号,把 可复用的经验(natural language experiences) 迭代沉淀到一个经验库 E,在下一轮推理时把 E 注入 prompt,达到“像 GRPO 一样迁移输出分布”的效果。

我们观察到的关键问题(本地小模型复现时非常明显):

  1. 经验库是 free-form text,很容易出现“污染样本”(例如只输出一个数字、带 $ 的金额、半句没说完、重复/矛盾等)。
  2. 这些污染一旦被注入 system prompt,往往会显著伤害整体表现(比不加经验还差)。

改进方向(本项目落地实现):把“经验库治理 + 适用性路由”作为核心瓶颈,而不是一味加强检索。

  • Schema-first memory(WHEN/NOT/DO):把 free-form experience 升级为可路由 schema(when/when_not/tags/then),靠条件表达减少误用。
  • Training-free 的 curator 压缩/去毒:用远程 LLM(config.yaml)把 noisy library 压缩成小而干净的经验集(避免把质量判断 hardcode 在本地规则里)。
  • Uncertainty / confidence gate:先跑 baseline,再在低置信样本触发检索+注入(最大化 help−hurt,闭合 oracle gap)。
  • (可选 baseline)EQG:仍保留 rule-based 过滤脚本,用于做对照/消融;但主线不依赖 hardcode 规则。

代码结构

  • papers/2510.08191.pdf: 论文 PDF(已下载在本地便于查阅)
  • research_tfgrpo/: 最小可复现的模型封装/数据处理/experience & soft-prior 工具
  • scripts/train_tfgrpo.py: 生成 rollouts、做 winner/loser 对比、抽取经验、保存 experiences.json & logit_bias.npy
  • scripts/filter_experiences.py: Experience Quality Gate(过滤 + 去重)
  • scripts/curate_experiences_remote.py: 远程 curator(config.yaml)对经验库做 schema 化压缩/去毒(training-free)
  • scripts/refine_schema_memory_remote.py: 用 help/hurt 证据闭环治理 schema memory(training-free,支持 add/modify/delete/merge ops)
  • scripts/eval_gsm8k.py: GSM8K 评测(可选注入 experiences、可选 logit bias)

快速复现实验(使用你给的本地模型)

模型路径(用户提供):

/cephfs/shared/hf_cache/hub/models--Qwen--Qwen3-0.6B/snapshots/c1899de289a04d12100db370d81485cdf75e47ca

1) Baseline:不加经验

source /root/miniconda3/etc/profile.d/conda.sh
conda activate lean

python scripts/eval_gsm8k.py \
  --model_path /cephfs/shared/hf_cache/hub/models--Qwen--Qwen3-0.6B/snapshots/c1899de289a04d12100db370d81485cdf75e47ca \
  --split test --eval_size 50 --seed 0 \
  --out_json artifacts/baseline_eval_50.json

2) 训练一轮 Training-Free GRPO 风格的经验库

(小规模跑通为主;后续可加大 train_size / epochs / group_size)

python scripts/train_tfgrpo.py \
  --model_path /cephfs/shared/hf_cache/hub/models--Qwen--Qwen3-0.6B/snapshots/c1899de289a04d12100db370d81485cdf75e47ca \
  --train_size 40 --epochs 1 --group_size 4 --max_new_tokens 512 \
  --out_dir artifacts/tfgrpo_run2

会输出:

  • artifacts/tfgrpo_run2/experiences.json
  • artifacts/tfgrpo_run2/logit_bias.npy(当前实验表明这个简单 prior 不稳定,先作为探索保留)

3) Experience Quality Gate(过滤经验)

python scripts/filter_experiences.py \
  --in_json artifacts/tfgrpo_run2/experiences.json \
  --out_json artifacts/tfgrpo_run2/experiences_filtered.json

4) 评测:注入过滤后的经验

python scripts/eval_gsm8k.py \
  --model_path /cephfs/shared/hf_cache/hub/models--Qwen--Qwen3-0.6B/snapshots/c1899de289a04d12100db370d81485cdf75e47ca \
  --split test --eval_size 50 --seed 0 \
  --experiences_json artifacts/tfgrpo_run2/experiences_filtered.json \
  --out_json artifacts/exp_run2_filtered_eval_50.json

当前实验结果(同一评测切片:split=test, eval_size=50, seed=0)

Setting Accuracy
Baseline(无经验) 0.56 (artifacts/baseline_eval_50.json)
经验注入(run1, 未做质量过滤) 0.58 (artifacts/exp_only_eval_50.json)
经验注入(run2 + EQG 过滤) 0.60 (artifacts/exp_run2_filtered_eval_50.json)
仅 logit bias(当前 naive prior) 0.54 (artifacts/bias_only_eval_50.json)

说明:

  • 这是一个“小规模先跑通 pipeline”的验证,不代表最终结论;接下来要扩大 eval_size(甚至全量 test)并做多 seed。
  • run2 的原始经验里出现了类似 $62 这种“污染经验”,直接注入会显著变差;EQG 把它滤掉后,性能恢复并超过 baseline。

下一步把它变成完整 research project(建议 roadmap)

  1. 更强的经验表示:把 experience 从 free-form 文本改成结构化 DSL(Do/Don’t / Trigger → Action),提升可控性与可过滤性。
  2. 检索式经验注入:经验库变大后,按 query embedding 只取 top-k 相关经验,避免无关经验造成负迁移。
  3. 自动冲突检测/合并:用 embedding clustering + contradiction check 清理经验库(可用同一个 LLM 或轻量规则)。
  4. 更可信的“semantic advantage”:当没有 ground truth 时,用 pairwise judge(或自一致 + verifier)生成相对优势,贴近论文 setting。
  5. 更严谨评测:GSM8K 全量 + out-of-domain(SVAMP / AIME 子集等),报告均值/方差、置信区间,并做消融(无过滤/不同过滤策略/不同 top-k)。

如果你希望我下一步就做:我可以把 “检索式经验注入(top-k)” 和 “结构化经验 DSL” 也加进来,并在更大 eval_size 上跑一轮对比。

Contributors

yzlu0917

Issues