tracked author

Ju Huang

Alibaba Group author on ROLL, ROLL Flash, RollPacker and RollArt. Repeated official author lists support the project linkage.

2 archived notes X: not-found

Related Notes

按论文归档时间排序,展示该作者在本站已经出现的材料。

归档

Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning

在同一训练框架下、覆盖范围并不完全重合的 Qwen3 4B/8B Base 与对齐后数学强化学习实验中分别检查优势归一化、概率比裁剪、损失聚合和超长过滤,并将组内均值—批级标准差归一化与 token 级损失组成不训练价值模型的 Lite PPO;OpenReview 补充的一个 Qwen3 8B Base 三随机种子设置仍优于 GRPO 与 DAPO,但奖励范围异常、计算量未对齐和有限模型任务覆盖限制普适结论。

待审阅 2508.08221-tricks-or-traps-lite-ppo RL AlgorithmReasoning RLTraining Stability