tracked author

Andrea Zanette

Corresponding author of Maximum Likelihood Reinforcement Learning. Homepage identifies him as an Assistant Professor at Carnegie Mellon University, joining ECE in Fall 2024 with a courtesy appointment in MLD; previously postdoctoral scholar at UC Berkeley and PhD at Stanford University.

1 archived notes X: not-found Homepage

Related Notes

按论文归档时间排序,展示该作者在本站已经出现的材料。

Jul 03, 2026

Maximum Likelihood Reinforcement Learning

MaxRL 把 binary outcome RLVR 改写为对成功 rollout 隐式 likelihood 的近似最大化:标准 RL 只优化 $pass@1$ 的一阶项,MaxRL 用 $N$ 条 rollout 中的成功样本数 $K$ 做归一化,得到对 $T=N$ 截断 maximum likelihood objective 的无偏 policy gradient estimator;实验显示它在 ImageNet toy ...

2602.02710-maximum-likelihood-reinforcement-learning RL AlgorithmRL TheoryReasoning RL