最新进展 / 最新结果:请直接看 PROGRESS.md(包含 SVAMP / GSM8K / ASDiv / MultiArith 的 multi-seed 结果、cost/quality Pareto、uncertainty gate、trigger filter、blacklist 等图表与结论)。
目标:基于 Training-Free Group Relative Policy Optimization(Training-Free GRPO, arXiv:2510.08191, 2025-10-09)做一个可在小模型/本地可跑的复现骨架,并针对论文里“experience library / token prior”这条路线提出一个更稳、更可落地的改进点,然后用本地 Qwen3-0.6B 做一轮实验验证。
这份仓库当前聚焦 数学推理(GSM8K),后续可以无缝迁移到 API 大模型做更强的 out-of-domain / tool-use 实验。
Training-Free GRPO 的核心是:不更新模型参数,通过“组内相对优势”的信号,把 可复用的经验(natural language experiences) 迭代沉淀到一个经验库 E,在下一轮推理时把 E 注入 prompt,达到“像 GRPO 一样迁移输出分布”的效果。
我们观察到的关键问题(本地小模型复现时非常明显):
- 经验库是 free-form text,很容易出现“污染样本”(例如只输出一个数字、带
$的金额、半句没说完、重复/矛盾等)。 - 这些污染一旦被注入 system prompt,往往会显著伤害整体表现(比不加经验还差)。
改进方向(本项目落地实现):把“经验库治理 + 适用性路由”作为核心瓶颈,而不是一味加强检索。
- Schema-first memory(WHEN/NOT/DO):把 free-form experience 升级为可路由 schema(
when/when_not/tags/then),靠条件表达减少误用。 - Training-free 的 curator 压缩/去毒:用远程 LLM(
config.yaml)把 noisy library 压缩成小而干净的经验集(避免把质量判断 hardcode 在本地规则里)。 - Uncertainty / confidence gate:先跑 baseline,再在低置信样本触发检索+注入(最大化 help−hurt,闭合 oracle gap)。
- (可选 baseline)EQG:仍保留 rule-based 过滤脚本,用于做对照/消融;但主线不依赖 hardcode 规则。
papers/2510.08191.pdf: 论文 PDF(已下载在本地便于查阅)research_tfgrpo/: 最小可复现的模型封装/数据处理/experience & soft-prior 工具scripts/train_tfgrpo.py: 生成 rollouts、做 winner/loser 对比、抽取经验、保存experiences.json&logit_bias.npyscripts/filter_experiences.py: Experience Quality Gate(过滤 + 去重)scripts/curate_experiences_remote.py: 远程 curator(config.yaml)对经验库做 schema 化压缩/去毒(training-free)scripts/refine_schema_memory_remote.py: 用 help/hurt 证据闭环治理 schema memory(training-free,支持 add/modify/delete/merge ops)scripts/eval_gsm8k.py: GSM8K 评测(可选注入 experiences、可选 logit bias)
模型路径(用户提供):
/cephfs/shared/hf_cache/hub/models--Qwen--Qwen3-0.6B/snapshots/c1899de289a04d12100db370d81485cdf75e47ca
source /root/miniconda3/etc/profile.d/conda.sh
conda activate lean
python scripts/eval_gsm8k.py \
--model_path /cephfs/shared/hf_cache/hub/models--Qwen--Qwen3-0.6B/snapshots/c1899de289a04d12100db370d81485cdf75e47ca \
--split test --eval_size 50 --seed 0 \
--out_json artifacts/baseline_eval_50.json(小规模跑通为主;后续可加大 train_size / epochs / group_size)
python scripts/train_tfgrpo.py \
--model_path /cephfs/shared/hf_cache/hub/models--Qwen--Qwen3-0.6B/snapshots/c1899de289a04d12100db370d81485cdf75e47ca \
--train_size 40 --epochs 1 --group_size 4 --max_new_tokens 512 \
--out_dir artifacts/tfgrpo_run2会输出:
artifacts/tfgrpo_run2/experiences.jsonartifacts/tfgrpo_run2/logit_bias.npy(当前实验表明这个简单 prior 不稳定,先作为探索保留)
python scripts/filter_experiences.py \
--in_json artifacts/tfgrpo_run2/experiences.json \
--out_json artifacts/tfgrpo_run2/experiences_filtered.jsonpython scripts/eval_gsm8k.py \
--model_path /cephfs/shared/hf_cache/hub/models--Qwen--Qwen3-0.6B/snapshots/c1899de289a04d12100db370d81485cdf75e47ca \
--split test --eval_size 50 --seed 0 \
--experiences_json artifacts/tfgrpo_run2/experiences_filtered.json \
--out_json artifacts/exp_run2_filtered_eval_50.json| Setting | Accuracy |
|---|---|
| Baseline(无经验) | 0.56 (artifacts/baseline_eval_50.json) |
| 经验注入(run1, 未做质量过滤) | 0.58 (artifacts/exp_only_eval_50.json) |
| 经验注入(run2 + EQG 过滤) | 0.60 (artifacts/exp_run2_filtered_eval_50.json) |
| 仅 logit bias(当前 naive prior) | 0.54 (artifacts/bias_only_eval_50.json) |
说明:
- 这是一个“小规模先跑通 pipeline”的验证,不代表最终结论;接下来要扩大
eval_size(甚至全量 test)并做多 seed。 run2的原始经验里出现了类似$62这种“污染经验”,直接注入会显著变差;EQG 把它滤掉后,性能恢复并超过 baseline。
- 更强的经验表示:把 experience 从 free-form 文本改成结构化 DSL(Do/Don’t / Trigger → Action),提升可控性与可过滤性。
- 检索式经验注入:经验库变大后,按 query embedding 只取 top-k 相关经验,避免无关经验造成负迁移。
- 自动冲突检测/合并:用 embedding clustering + contradiction check 清理经验库(可用同一个 LLM 或轻量规则)。
- 更可信的“semantic advantage”:当没有 ground truth 时,用 pairwise judge(或自一致 + verifier)生成相对优势,贴近论文 setting。
- 更严谨评测:GSM8K 全量 + out-of-domain(SVAMP / AIME 子集等),报告均值/方差、置信区间,并做消融(无过滤/不同过滤策略/不同 top-k)。
如果你希望我下一步就做:我可以把 “检索式经验注入(top-k)” 和 “结构化经验 DSL” 也加进来,并在更大 eval_size 上跑一轮对比。