tracked author

Yueer Zhou

Coauthor of Maximum Likelihood Reinforcement Learning. Homepage identifies her as a Zhejiang University undergraduate, working closely with Andrea Zanette and Fahim Tajwar at CMU and with USC Viterbi GVL Lab; Andrea Zanette's homepage lists her as former intern to Stanford MSCS.

1 archived notes X: high HomepageGitHubX

Related Notes

按论文归档时间排序,展示该作者在本站已经出现的材料。

Jul 03, 2026

Maximum Likelihood Reinforcement Learning

MaxRL 把 binary outcome RLVR 改写为对成功 rollout 隐式 likelihood 的近似最大化:标准 RL 只优化 $pass@1$ 的一阶项,MaxRL 用 $N$ 条 rollout 中的成功样本数 $K$ 做归一化,得到对 $T=N$ 截断 maximum likelihood objective 的无偏 policy gradient estimator;实验显示它在 ImageNet toy ...

2602.02710-maximum-likelihood-reinforcement-learning RL AlgorithmRL TheoryReasoning RL