tracked author

Jiacheng Chen

Entropy Mechanism equal-contribution author. Homepage identifies him as a CUHK CSE Ph.D. student advised by Yu Cheng and Weiyang Liu, with MiniMax and Shanghai AI Lab internship/project work; page lists Entropy Mechanism and P1/MaxProof related reasoning RL work.

1 archived notes X: not-found HomepageGitHub

Related Notes

按论文归档时间排序,展示该作者在本站已经出现的材料。

Jun 21, 2026

The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

这篇论文把 reasoning LLM 的 RLVR 训练瓶颈重新表述为 policy entropy 的消耗过程:在没有 entropy / KL 干预时,reward 提升和 entropy 下降之间可以被经验式 $R= a\exp(\mathcal H)+b$ 拟合;进一步用 softmax policy 的 entropy dynamics 说明,高概率且高 advantage 的 token update 会持续降低 ent...

2505.22617-entropy-mechanism-rl-reasoning-language-models Training StabilityReasoning RLRL Theory