tracked author
Xuandong Zhao
VIMPO coauthor and recurring Berkeley reasoning RL/RLIF author. Homepage identifies him as a UC Berkeley postdoctoral researcher in BAIR/RDI working with Dawn Song, focused on machine learning, NLP, AI safety, scalable reinforcement learning and self-improvement; GitHub links the same homepage and X handle.