tracked author

Rahul K. Arora

Beneficial RL coauthor. Homepage identifies him as an OpenAI researcher working on health and RL, a University of Calgary faculty member, and an early member of OpenAI's Health AI team; public work includes HealthBench and health model performance improvements.

1 archived notes X: not-found Homepage

Related Notes

按论文归档时间排序,展示该作者在本站已经出现的材料。

Jun 21, 2026

Reinforcement Learning Towards Broadly and Persistently Beneficial Models

这篇论文给 OpenAI 的 alignment post training 提供了一个正向版本的 emergent misalignment 实验:如果窄域有害训练能诱导跨域失配,那么用 5% 真实场景 beneficial trait data 加 RL reward 强化诚实、纠错、风险意识、公平和人类福祉等特质,也可能诱导跨域对齐收益。论文的证据强在 44/53 个 OOD 评测提升、health only 训练迁移到非健康安...

2026-06-18-openai-beneficial-rl AI SafetyReward ModelingAgent RL