tracked author

Yongchao Zhou

Coauthor of On-Policy Distillation of Language Models. The paper lists Google DeepMind and University of Toronto. Public OpenReview, Scholar, LinkedIn snippets, and X profile connect him to University of Toronto, Vector Institute, Google DeepMind, xAI, and micro1.

1 archived notes X: high HomepageX

Related Notes

按论文归档时间排序,展示该作者在本站已经出现的材料。

Jul 03, 2026

On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

GKD 把 LLM 蒸馏从固定参考答案上的 teacher forcing 推进到学生自生成轨迹上的教师分布匹配,用一个 $\lambda$ 混合离线数据和 on policy 数据,并允许 forward KL、reverse KL、JSD 等不同散度;它在 T5 系列的摘要、翻译、算术推理和 instruction tuning 上展示了稳定收益,是后续 OPD 系方法的重要早期基线。

2306.13649-on-policy-distillation-language-models On-Policy DistillationRL Algorithm