tracked author

Siyan Zhao

OPSD first and corresponding author and UCLA computer-science PhD student advised by Aditya Grover. The paper was completed at UCLA and during a part-time Meta internship; her homepage records research spanning LLM post-training, alignment and efficiency.

1 archived notes X: high HomepageGitHubHugging FaceX

Representative Papers

来自作者已核验个人主页的重点论文;本站单篇归档见下方 Related Notes。

  1. 01 Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models ICML 2026 · 2026
  2. 02 Inpainting-Guided Policy Optimization for Diffusion Large Language Models ICLR 2026 · 2026
  3. 03 d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement Learning NeurIPS 2025 · 2025
  4. 04 Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs ICLR 2025 · 2025
  5. 05 Probing the Decision Boundaries of In-context Learning in Large Language Models NeurIPS 2024 · 2024
  6. 06 Prepacking: A Simple Method for Fast Prefilling and Increased Throughput in Large Language Models AISTATS 2025 · 2025
  7. 07 Group Preference Optimization: Few-Shot Alignment of Large Language Models ICLR 2024 · 2024
  8. 08 Decision Stacks: Flexible Reinforcement Learning via Modular Generative Models NeurIPS 2023 · 2023

Related Notes

按论文归档时间排序,展示该作者在本站已经出现的材料。

归档

Self Distilled Reasoner: On Policy Self Distillation for Large Language Models

OPSD 让学生先在只看题目时采样自身轨迹,再用固定初始模型在参考解答上下文中对相同前缀给出的完整词表分布和逐词表项裁剪前向 KL 更新学生,在 Qwen3 1.7B/4B/8B 的三项数学平均分上高于基础模型、SFT 与 GRPO,但总计算、统计方差和领域外泛化仍未对齐。

待审阅 2601.18734-self-distilled-reasoner-opsd On-Policy DistillationReasoning RLTraining Stability