recurring author
Zhiyuan Liu
2 archived notes
Related Notes
按论文归档时间排序,展示该作者在本站已经出现的材料。
归档 更新
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
把 RLVR 训练能力上限关联到 policy entropy 消耗,并分析 advantage update 的熵动力学。
待审阅 2505.22617-entropy-mechanism-rl-reasoning-language-models Training StabilityReasoning RLRL Theory
归档 更新
From $f(x)$ and $g(x)$ to $f(g(x))$: LLMs Learn New Skills in RL by Composing Old Ones
在受控任务中证明 RL 可组合 base model 已掌握的 atomic skills,形成未见组合能力。