tracked author

Haobo Wang (王皓波)

CriPO corresponding author and Zhejiang University assistant professor and PhD supervisor; his homepage records research in machine learning, data intelligence, and large language models. Explicit paper links are required because the name is common.

1 archived notes X: not-found Homepage

Related Notes

按论文归档时间排序,展示该作者在本站已经出现的材料。

归档

Enhancing Rubric based RL via Self Distillation

把评分量规聚合后的学习信号丢失拆成当前采样未覆盖和已满足但整体优势非正两类,用评分项条件自教师注入缺失行为,并以反事实自教师定位 token 后局部改写优势;两种 Qwen3 小模型在五项裁判评测中较 GRPO 平均提高 3.2 和 1.4 分,证据缺少多随机种子与完整硬件条件。

待审阅 2607.18082-cripo-rubric-rl-self-distillation Credit AssignmentOn-Policy DistillationRL Algorithm