tracked author

Akshay V. Jagadeesh

Corresponding author of OpenAI's Beneficial RL paper. Homepage identifies him as an OpenAI Research Scientist working on deeply and persistently beneficial AI with a focus on health and medicine; earlier research line connects computational neuroscience, Stanford and Harvard.

1 archived notes X: not-found HomepageGitHub

Related Notes

按论文归档时间排序,展示该作者在本站已经出现的材料。

Jun 21, 2026

Reinforcement Learning Towards Broadly and Persistently Beneficial Models

这篇论文给 OpenAI 的 alignment post training 提供了一个正向版本的 emergent misalignment 实验:如果窄域有害训练能诱导跨域失配,那么用 5% 真实场景 beneficial trait data 加 RL reward 强化诚实、纠错、风险意识、公平和人类福祉等特质,也可能诱导跨域对齐收益。论文的证据强在 44/53 个 OOD 评测提升、health only 训练迁移到非健康安...

2026-06-18-openai-beneficial-rl AI SafetyReward ModelingAgent RL