tracked author

Zhihui Xie (谢知晖)

OPSD second author and HKU PhD student advised by Lingpeng Kong and Qi Liu. His homepage records a 2025 Meta Superintelligence Labs research internship and prior bachelor's and master's study at Shanghai Jiao Tong University.

1 archived notes X: high HomepageGitHubHugging FaceX

Representative Papers

来自作者已核验个人主页的重点论文;本站单篇归档见下方 Related Notes。

  1. 01 Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models ICML 2026 · 2026
  2. 02 Dream-Coder 7B: An Open Diffusion Language Model for Code arXiv · 2025
  3. 03 Dream 7B: Diffusion Large Language Models arXiv · 2025
  4. 04 Teaching Language Models to Critique via Reinforcement Learning ICML 2025 · 2025
  5. 05 Learning Versatile Skills with Curriculum Masking NeurIPS 2024 · 2024
  6. 06 VLRewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models CVPR 2025 · 2025
  7. 07 VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment EMNLP 2024 · 2024
  8. 08 Calibrating Reasoning in Language Models with Internal Consistency NeurIPS 2024 · 2024
  9. 09 Future-conditioned Unsupervised Pretraining for Decision Transformer ICML 2023 · 2023
  10. 10 Discovering Low-rank Subspaces for Language-agnostic Multilingual Representations EMNLP 2022 · 2022
  11. 11 Comparison-based Conversational Recommender System with Relative Bandit Feedback SIGIR 2021 · 2021

Related Notes

按论文归档时间排序,展示该作者在本站已经出现的材料。

归档

Self Distilled Reasoner: On Policy Self Distillation for Large Language Models

OPSD 让学生先在只看题目时采样自身轨迹,再用固定初始模型在参考解答上下文中对相同前缀给出的完整词表分布和逐词表项裁剪前向 KL 更新学生,在 Qwen3 1.7B/4B/8B 的三项数学平均分上高于基础模型、SFT 与 GRPO,但总计算、统计方差和领域外泛化仍未对齐。

待审阅 2601.18734-self-distilled-reasoner-opsd On-Policy DistillationReasoning RLTraining Stability