tracked author

John Schulman

Approximating KL Divergence blog author and PPO/TRPO/RLHF key researcher. Homepage identifies him as Thinking Machines cofounder and chief scientist, former Anthropic Alignment Science researcher, OpenAI cofounder, and former OpenAI post-training co-lead; X profile search result links the same joschu.net website. Also coauthor of Let's Verify Step by Step and Training Verifiers to Solve Math Word Problems.

2 archived notes X: high HomepageX

Related Notes

按论文归档时间排序,展示该作者在本站已经出现的材料。

Jun 23, 2026

Let's Verify Step by Step

这篇是 OpenAI reasoning verification 线的历史节点:它用 80 万 step level human labels 训练 PRM,在 MATH 500 题 held out subset 上用 Best of 1860 选择达到 78.2%,高于强 ORM 和 majority vote;核心价值是把“最终答案正确”拆成“每一步是否仍在正确推理轨道上”,并通过 PRM800K 让过程监督成为后续 reas...

2305.20050-lets-verify-step-by-step-process-supervision Process SupervisionVerifierReward Modeling
Jun 21, 2026

Approximating KL Divergence

这篇博客给出了 RL 实践中常用 KL 估计器 k1/k2/k3 的最小数学解释:在只能从 $q$ 采样并能计算 $p(x),q(x)$ 的场景下,$k 1= \log r$ 是无偏但高方差的 $\mathrm{KL}[q,p]$ 估计器;$k 2=\frac12(\log r)^2$ 是低方差但有偏的二阶近似;$k 3=(r 1) \log r$ 通过控制变量 $r 1$ 保持无偏、非负和低方差,其中 $r=p(x)/q(x)$。对...

2020-03-07-schulman-kl-divergence-approximations RL TheoryTraining Stability