tracked author

Hailin Zhang (张海林)

MOPD coauthor and Xiaomi MiMo RL infrastructure researcher. Homepage identifies him as Hailin Zhang (张海林), currently building MiMo at Xiaomi, specializing in AI infrastructure and efficient, scalable, stable RL infrastructures for MiMo series models; it links GitHub HugoZHL and Google Scholar and states he earned a Peking University PhD in 2025 advised by Bin Cui. No high-confidence personal X account found in this pass.

1 archived notes X: not-found HomepageGitHub

Related Notes

按论文归档时间排序,展示该作者在本站已经出现的材料。

Jul 02, 2026

MOPD: Multi Teacher On Policy Distillation for Capability Integration in LLM Post Training

MOPD 把“多个领域 RL teacher 的能力整合”改写成一个 on policy token level distillation 问题:student 先用自己的当前策略生成轨迹,再让对应领域 teacher 在同一轨迹前缀上做 teacher forced prefill,最后用 reverse KL 或等价 policy gradient advantage 把 teacher 的局部偏好注入 student;关键成立条...

2606.30406-mopd-multi-teacher-on-policy-distillation On-Policy DistillationRL AlgorithmReasoning RL