recurring author
Yuxin Zuo
2 archived notes
Related Notes
按论文归档时间排序,展示该作者在本站已经出现的材料。
归档 更新
Self Improving Agents in the Era of Experience: A Survey of Self to Meta Evolution
把 agent 自改进抽象为 trace to capability 流水线,覆盖 skills、memory、environment、model 与 meta layer。
归档 更新
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
把 RLVR 训练能力上限关联到 policy entropy 消耗,并分析 advantage update 的熵动力学。
待审阅 2505.22617-entropy-mechanism-rl-reasoning-language-models Training StabilityReasoning RLRL Theory