Jun 23, 2026

国产前沿模型技术报告时间线总览

国产前沿模型技术报告在 2024 2026 年的主线,可以读成一次连续的工程迁移:先用 MoE、MLA、长上下文和低精度训练把 frontier base model 做到可训练、可服务;随后用 verifiable reward、long CoT、GRPO/OMD/CISPO 等方法把推理能力从模型先验里释放出来;再把 rollout、工具环境、checkpoint、sparse attention、MTP 和 anti hack ...

Pinned 2026-06-23-chinese-frontier-model-reports-timeline MoE SystemsLong ContextReasoning RL
Jul 16, 2026

Self Compacting Language Model Agents

SelfCompact 把长轨迹摘要做成同一模型可执行的 hard reset 操作,再用任务专用、要求引用当前轨迹证据的 rubric 决定何时压缩;它在七个开放权重模型的数学与搜索实验中优于无压缩,并在大多数质量对照上优于固定时机摘要。最可靠的增量是把 timing 判断拆成可审计的状态 predicate,消融支持含 rubric 的完整 policy 优于自由摘要工具;search 附录声明的 40k token gate、3...

Tianjian Li, Jingyu Zhang (张景昱), William Jurayj, Xi Wang, Chuanyang Jin, Mehrdad Farajtabar, Eric Nalisnick, Daniel Khashabi

2606.23525-self-compacting-language-model-agents Agent MemoryAgent WorkflowLong Context
Jul 16, 2026

What Preferences Can—and Cannot—Predict in Multi Agent Online Learning

论文把有限博弈的序数偏好图与连续时间 FTRL 的集合稳定性联系起来:偏好闭合给出稳定结果的必要约束;对 subgame,club 在一般 FTRL 下足以保证 span 渐近稳定、在无 ties 时形成等价判据,并在 strategy flow 下直接形成 attractor 等价判据;对一般纯策略集合,三人反例表明相同的偏好方向仍可能产生不稳定 span,作者因而引入依赖收益差幅度的 leaklessness,为一般 span 恢...

Omar Abbadi, Rida Laraki, Panayotis Mertikopoulos

2026-04-30-preferences-multi-agent-online-learning RL TheoryReward Modeling
Jul 16, 2026

A Survey of Reinforcement Learning for Large Language Models under Data Scarcity: Challenges and Solutions

这篇 ACL 2026 综述把 LLM 强化学习的数据约束拆成高成本外部监督与有限内部生成经验,并用 data centric、training centric、framework centric 三层九类 taxonomy 组织 125 条文献记录;它提供了便于导航的设计空间,系统性证据仍受检索协议缺失、分类轴混合、预算口径不统一和少数方法描述失真的限制。

Zhiyin Yu, Yuchen Mou, Juncheng Yan, Junyu Luo, Chunchun Chen, Xing Wei, Yunhui Liu, Hongru Sun, Yuxing Zhang, Jun Xu, +10 more

2604.17312-rl-llm-data-scarcity-survey Reasoning RLRollout OptimizationReward Modeling
Jul 16, 2026

SGLang Data Parallel Attention:从请求所有权到 MoE 布局转换

SGLang DPA 在同一组全局 TP workers 内,把 Attention ranks 划成多个按请求拥有 KV Cache 的 DP groups,再通过 gather / dispatch / combine / reduce scatter 把 DP local Attention 与全局 TP 或 EP MoE 连接起来;对于缓存受限的高并发 MLA decode,它可把单条请求的 cache 副本数从 $T$ 降到...

2026-07-16-sglang-data-parallel-attention Inference ParallelismKV CacheMoE Systems
Jul 16, 2026

MLA 在张量并行下的缓存复制:从压缩收益到可分片 Latent

MLA 把每个 token 的 K/V 压成共享 latent 与独立 RoPE 分支,常见 head TP absorbed decode 会让每个 rank 保存同一请求的完整 latent cache,使其相对展开 K/V 的每卡压缩倍数按 $1/T$ 衰减;生产系统可通过 DP Attention 改变请求所有权、DCP/CP 沿序列分片、P/D 拆分阶段、量化或分层缓存降低驻留字节,TPLA、GLA 与 MLRA 则进一步处...

2026-07-16-mla-tensor-parallel-cache-sharding KV CacheInference ParallelismLong Context
Jul 14, 2026

推理侧 Prefill Context Parallelism:为何有效,以及 CP / SP / UP 的边界

Prefill CP 将单条长请求的 query token 沿 context 维分到多个 rank,在每个 rank 保持全局 KV 或全局 attention state 可见性,并用因果负载均衡、通信重叠和并行组解耦把 dense attention 或 DSA indexer 的高增长计算并行化;这套数学分解可覆盖 MHA、GQA、MLA 与 DSA,生产级扩展仍取决于 KV 所有权、attention backend、跨节...

2026-07-14-prefill-context-parallelism-inference-scaling Inference ParallelismLong ContextAttention Kernel
Jul 14, 2026

Leyline: KV Cache Directives for Agentic Inference

Leyline 在 agent policy 与 KV Cache 之间增加语义编辑通道:policy 用 (span, replacement, mode) directive 声明需要修改的历史范围和语义保证,serving 层保留未修改前缀、重算 replacement,并用 RoPE $\delta$ rotation 重标 MLA 后缀 KV 的位置,从而在 AMORTIZE 模式下减少上下文编辑后的重复 prefill,在...

Bole Ma, Jan Eitzinger, Harald Köstler

2606.01065-leyline-kv-cache-directives-agentic-inference KV CacheAgent WorkflowServing Runtime
Jul 14, 2026

IndexCache: Accelerating Sparse Attention via Cross Layer Index Reuse

IndexCache 将 DSA 层划分为运行 selector 的 Full 层和继承最近 Full 层 top $k$ token positions 的 Shared 层,再用 loss guided layer search 或 multi layer index distillation 决定共享模式,从而在不复制 KV、也不改变 sparse core attention 的条件下跳过最多 75% 的逐层 indexer ...

Yushi Bai (白雨石), Qian Dong, Ting Jiang, Xin Lv (吕鑫), Zhengxiao Du, Aohan Zeng, Jie Tang (唐杰), Juanzi Li (李涓子)

2603.12201-indexcache-cross-layer-index-reuse Sparse AttentionLong ContextServing Runtime
Jul 13, 2026

Agentic RL Learned Environment 演进路线:从可执行 Sandbox 到可校准世界模型

Agentic RL 的环境供给在 2024 2026 年经历了从统一可复位 sandbox、有状态用户与工具模拟、代码驱动环境合成,到用真实交互轨迹训练 language world model 的连续演进;当前证据最支持 real environment 提供状态与验证锚点、learned environment 扩大低成本 rollout、周期性真实交互修正分布偏移的混合闭环。

2026-07-13-agentic-rl-learned-environment-evolution Learned EnvironmentAgent RLTool Use
Jul 13, 2026

2026 年 5 7 月 RL 信用分配研究进展:从 Token 到 Compact Agent

2026 年 5 7 月的 RL 信用分配研究开始围绕 token、segment、turn、memory operation 和 workflow role 五种 credit unit 形成可比较的方法谱系;compact agent 的关键增量是把“摘要或记忆写入后,各 rollout 的有效状态是否仍可比较”单独列为 estimator 条件,并分别发展出同状态局部重采样、belief proxy、hindsight coun...

2026-07-13-rl-credit-assignment-may-july-landscape Credit AssignmentAgent RLAgent Memory
Jul 13, 2026

CompactionRL: Reinforcement Learning with Context Compaction for Long Horizon Agents

CompactionRL 让同一个 trainable actor 在上下文预算将满时生成摘要,再从“摘要 + 最近两轮”重建上下文,并把执行 token 与摘要 token 一起放入共享最终任务 reward 的 PPO 目标;独立 critic、全 batch token level loss normalization 和跨 segment GAE 修正使 variable length compacted rollouts 可...

Yujiang Li, Zhenyu Hou, Yi Jing, Jie Tang (唐杰), Yuxiao Dong

2607.05378-compactionrl-context-compaction-agent-rl Agent MemoryAgent RLCredit Assignment
Jul 13, 2026

Efficient Serving for Agentic LLM Workflows via Micro Task Level Parallelism

Grape 面向预声明的多阶段 LLM task DAG (Directed Acyclic Graph,有向无环图),把静态 prompt、流式上游输出和 decode 拆成微任务,让下游已知 prefix 与上游逐步生成的中间结果在上游 decode 期间持续完成增量 prefill,再用服务级目标约束的 batch 构造和关键路径感知 Key Value Cache (KV Cache) 抢占控制尾延迟。 本地评价:最有价值的新...

当前匿名稿未披露完整作者列表;公开可确认关联作者为 Siqi Wang 和 Hailong Yang

2026-07-13-grape-micro-task-agentic-workflow-serving Agent WorkflowInference SchedulingKV Cache
Jul 13, 2026

LLM as a Verifier: A General Purpose Verification Framework

LLM as a Verifier 把评分 token 的概率期望、重复评估与 criteria decomposition 组合成连续 verifier,并用 Probabilistic Pivot Tournament 在接近线性的比较预算下选择多条 agent 轨迹;它在候选重排、进度代理和 RL dense reward 上显示出广泛用途,同时依赖 logits、成对上下文稳定性、人工 criteria 和较高验证推理预算。

Jacky Kwok, Shulu Li, Pranav Atreya, Yuejiang Liu, Yixing Jiang, Chelsea Finn, Marco Pavone, Ion Stoica, Azalia Mirhoseini

2607.05391-llm-as-a-verifier VerifierTest-Time ScalingReward Modeling
Jul 11, 2026

Single Rollout Asynchronous Optimization for Agentic Reinforcement Learning

SAO 把异步 agentic reinforcement learning 的 prompt local 采样单位改为一条 rollout,用独立 value model 的更高更新频率、冻结 attention 和跳过 observation 的 token level Generalized Advantage Estimation (GAE) 代替组内 reward baseline,再用 rollout token log ...

Zhenyu Hou, Yujiang Li, Jie Tang (唐杰), Yuxiao Dong

2607.07508-sao-single-rollout-asynchronous-agentic-rl RL AlgorithmAgent RLCredit Assignment
Jul 10, 2026

SPORK: Self Speculative Forking to Accelerate Agentic LLM Inference

SPORK 让正在生成的目标模型在首个 decode token 后从共享 KV prefix 开出强制工具调用 probe,按工具名 token 的置信度提前执行只读工具,并在预测失败时把 probe 的匹配 token 前缀交给目标模型验证复用,从而把部分工具等待时间隐藏在剩余 Chain of Thought decode 内。 本地评价:这项工作的核心贡献是 action level latency overlap。收益依赖数...

Huajun Bai, Weiwei Lv, Huichuan Zheng, Youyou Lu, Jiwu Shu

2607.03333-spork-self-speculative-agentic-inference Speculative DecodingTool UseAgent Workflow
Jul 09, 2026

Qwen3 Coder Next Technical Report

Qwen3 Coder Next 的核心判断是:在编码智能体上,80B total / 3B active 的 Qwen3 Next hybrid attention + MoE 基座,可以通过大规模 executable environment、repo 级 mid training、多模板 tool calling、SWE 专家 RL 和 expert distillation,逼近更大 active compute 模型的 ag...

Ruisheng Cao, Mouxiang Chen (陈谋祥), Jiawei Chen (陈家慰), Zeyu Cui, Yunlong Feng, Binyuan Hui, Yuheng Jing, Kaixin Li, Mingze Li, Junyang Lin, +10 more

2603.00729-qwen3-coder-next-agentic-coding Coding AgentAgent RLTool Use
Jul 09, 2026

Computer Environments Elicit General Agentic Intelligence in LLMs

LLM in Sandbox 的核心贡献是把通用计算机抽象成一个最小 Docker code sandbox,并证明强模型在 training free 设置下能利用外部资源访问、文件管理和代码执行三类 meta capability 提升数学、物理、化学、长上下文和指令遵循等非代码任务;进一步的 LLM in Sandbox RL 说明,把一般 context based 数据放进文件系统并用 outcome reward 训练,可...

Daixuan Cheng (成岱璇), Shaohan Huang, Yuxian Gu, Huatong Song (宋华彤), Guoxin Chen, Li Dong (董力), Wayne Xin Zhao, Ji-Rong Wen, Furu Wei (韦福如)

2601.16206-computer-environments-agentic-intelligence Tool UseAgent RLAgent Workflow
Jul 09, 2026

Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling

HiLS Attention 把 chunk wise sparse attention 的选择问题改写为可端到端训练的 hierarchical softmax:每个 chunk 用 landmark token 生成压缩 key 和 entropy bias,query 先给 chunk 分配总 attention mass,再在被选 chunk 内做 token level attention。它在 345M 到 7B 实验中同...

Xiang Hu, Xinyu Wei, Hao Gu, Minshen Zhang, Tian Liang, Huayang Li, Lei Zhu (祝磊), Yan Wang (王琰), Sirui Han (韓斯睿), Yushi Bai (白雨石), +3 more

2607.02980-hils-attention-infinite-context Sparse AttentionLong Context
Jul 05, 2026

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

FlashMemory DeepSeek V4 在 DeepSeek V4 Flash 的 HCA / CSA 结构上增加一个 lookahead Neural Memory Indexer,把历史 CSA chunk 放入 CPU cold pool,再按未来约 $64$ 个 token 的预测需求取回少量 chunk;它把平均物理 KV 占用降到 full context baseline 的 $13.5\%$,平均分数从 $76...

2606.09079-flashmemory-deepseek-v4-lookahead-sparse-attention KV CacheSparse AttentionLong Context
Jul 03, 2026

ThunderAgent: A Simple, Fast and Program Aware Agentic Inference System

ThunderAgent 把跨多轮 LLM 调用与工具等待建模为带 phase、status、KV footprint、backend placement 和 tool dependency 的 LLM Program,使 scheduler 能按 Reasoning / Acting 状态控制 KV working set,并在释放 KV 后通过全局队列重新选择恢复节点;论文在所测高并发 tool use serving 与 rol...

Hao Kang, Ziyang Li, Weili Xu, Xinyu Yang, Yinfang Chen, Junxiong Wang, Beidi Chen, Tushar Krishna, Chenfeng Xu, Simran Arora

2602.13692-thunderagent-program-aware-agentic-inference Agent WorkflowInference SchedulingKV Cache
Jul 03, 2026

SPIRAL: Learning to Search and Aggregate

SPIRAL 从共享策略生成的 8 条 search traces 中随机抽取 4 个互异的四元集合,再为每个集合生成 4 条 aggregation traces;它用集合聚合成功率的 participation averaged advantage 更新搜索轨迹,并用同集合内中心化 reward 更新聚合轨迹,从而把 parallel search 与 model based aggregation 放进同一个最终答案 rewar...

Jubayer Ibn Hamid, Ifdita Hasan Orney, Michael Y. Li, Omar Shaikh, Yoonho Lee, Dorsa Sadigh, Chelsea Finn, Noah Goodman

2606.23595-spiral-learning-search-aggregate Test-Time ScalingCredit AssignmentRL Algorithm
Jul 03, 2026

ECHO: Prune to act, trace to learn with selective turn memory in agentic RL

ECHO 的核心价值在于把 long horizon agent 的 context management 从“压缩历史以继续行动”推进到“保留 source indexed reconstruction trace 以便学习”:每个已完成工具 turn 被压成带原始 turn id 的 memory record,policy 在上下文触顶时选择可复用记忆来重建 bounded context,并把最终正确轨迹的正向 outcome...

Zijun Xie, Binbin Zheng, Enlei Gong, Jihua Liu, Yuyang You, Lingfeng Liu, Jiayao Tang, Guanqun Zhao, Aoqi Hu, Zeyu Chen

2606.31650-echo-selective-turn-memory-agentic-rl Agent MemoryAgent RLLong Context
Jul 03, 2026

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents

Vortex 用 vFlow 描述 page summary、逐 token 动态路由与 selected attention,用 vTensor 把这些逻辑算子 lowering 到 paged / ragged layouts,再复用或补充 SGLang 的 GQA / MLA decode backend,使 sparse attention selector、top k 和真实 KV layout 能在同一端到端 servin...

Zhuoming Chen (陈卓明), Xinrui Zhong, Qilong Feng, Ranajoy Sadhukhan, Yang Zhou, Michael Qizhe Shieh, Zhihao Jia, Beidi Chen

2606.06453-vortex-sparse-attention-serving Serving RuntimeSparse AttentionAttention Kernel
Jul 03, 2026

Seed2.0 Model Card: Towards Intelligence Frontier for Real World Complexity

Seed2.0 模型卡把 ByteDance Seed 的新一代模型定位为面向真实复杂任务的生产模型族:它以 Doubao / Trae 等产品流量和用户任务为入口,重构了从长尾知识、复杂指令、搜索、视觉、视频、工具调用、GUI agent 到科学研究任务的评测面,并用 Seed2.0 Pro / Lite / Mini 的性能、成本和案例轨迹证明 ByteDance Seed 已经把模型报告从单点能力榜单推进到“产品需求分布、评测系...

Bytedance Seed

2607.00248-seed2-model-card-real-world-complexity BenchmarkMultimodal ModelTool Use
Jul 03, 2026

On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

GKD 把 LLM 蒸馏从固定参考答案上的 teacher forcing 推进到学生自生成轨迹上的教师分布匹配,用一个 $\lambda$ 混合离线数据和 on policy 数据,并允许 forward KL、reverse KL、JSD 等不同散度;它在 T5 系列的摘要、翻译、算术推理和 instruction tuning 上展示了稳定收益,是后续 OPD 系方法的重要早期基线。

2306.13649-on-policy-distillation-language-models On-Policy DistillationRL Algorithm
Jul 03, 2026

Maximum Likelihood Reinforcement Learning

MaxRL 把 binary outcome RLVR 改写为对成功 rollout 隐式 likelihood 的近似最大化:标准 RL 只优化 $pass@1$ 的一阶项,MaxRL 用 $N$ 条 rollout 中的成功样本数 $K$ 做归一化,得到对 $T=N$ 截断 maximum likelihood objective 的无偏 policy gradient estimator;实验显示它在 ImageNet toy ...

Fahim Tajwar, Guanning Zeng, Yueer Zhou, Yuda Song, Daman Arora, Yiding Jiang, Jeff Schneider, Ruslan Salakhutdinov, Haiwen Feng, Andrea Zanette

2602.02710-maximum-likelihood-reinforcement-learning RL AlgorithmRL TheoryReasoning RL
Jul 02, 2026

HydraHead: From Head Level Functional Heterogeneity to Specialized Attention Hybridization

HydraHead 把 Full Attention (FA) 和 Linear Attention (LA) 的混合粒度从 layer 推进到 attention head:先用 activation patching 和 path patching 找到 retrieval critical heads,只给这些头保留 FA,其余头换成 Gated DeltaNet (GDN),再用独立 RMSNorm 与 head wise s...

Zhentao Tan, Wei Chen, Jingyi Shen, Yao Liu, Xu Shen, Yue Wu, Jieping Ye

2606.20097-hydrahead-head-wise-hybrid-attention Linear AttentionLong ContextReasoning Analysis
Jul 02, 2026

MOPD: Multi Teacher On Policy Distillation for Capability Integration in LLM Post Training

MOPD 把“多个领域 RL teacher 的能力整合”改写成一个 on policy token level distillation 问题:student 先用自己的当前策略生成轨迹,再让对应领域 teacher 在同一轨迹前缀上做 teacher forced prefill,最后用 reverse KL 或等价 policy gradient advantage 把 teacher 的局部偏好注入 student;关键成立条...

Wenhan Ma, Jianyu Wei, Liang Zhao, Hailin Zhang (张海林), Bangjun Xiao, Lei Li, Qibin Yang, Bofei Gao, Yudong Wang, Rang Li, +3 more

2606.30406-mopd-multi-teacher-on-policy-distillation On-Policy DistillationRL AlgorithmReasoning RL
Jun 30, 2026

RollArt: Disaggregated Multi Task Agentic RL Training at Scale

RollArt 用静态 task domain affinity 声明把训练、prefill heavy rollout、decode heavy rollout、环境与 reward 映射到 H800、H20、CPU/Kubernetes 和 serverless 资源池,再用轨迹级状态机、起始版本年龄上限 $\alpha$ 与 Mooncake weight movement 协调跨池异步;OSDI camera ready 支持...

Wei Gao, Yuheng Zhao, Tianyuan Wu, Shaopan Xiong, Weixun Wang, Dakai An, Lunxi Cao, Dilxat Muhtar, Zichen Liu, Haizhou Zhao, +8 more

2512.22560-rollart-disaggregated-agentic-rl-training RL InfrastructureRollout OptimizationDistributed Training
Jun 28, 2026

DSpark: Confidence Scheduled Speculative Decoding with Semi Autoregressive Generation

DSpark 是 DeepSeek 把并行 drafter 推向生产 serving 的一套完整方案:用 DFlash 式 parallel backbone 先一次生成长候选块,再用低秩 Markov head 注入块内局部自回归依赖,随后用 calibrated confidence head 和硬件感知 prefix scheduler 按请求与负载动态裁剪 target verification 长度。离线 Qwen3 / G...

Xin Cheng, Xingkai Yu (俞星凯), Chenze Shao, Jiashi Li, Yunfan Xiong, Yi Qian, Jiaqi Zhu, Shirong Ma, Xiaokang Zhang, Jiasheng Ye, +23 more

2026-06-27-dspark-confidence-scheduled-speculative-decoding Speculative DecodingMulti-Token PredictionServing Runtime
Jun 28, 2026

DFlash: Block Diffusion for Flash Speculative Decoding

DFlash 的核心贡献是把 diffusion LLM 放到 speculative decoding 的 draft stage:用目标自回归模型的多层 hidden features 做条件,把融合特征注入 draft model 每层 KV cache,再用 block diffusion 一次并行预测整块候选 token。它在 Qwen3 4B/8B 非 thinking 模式下报告约 4.0x 4.9x 平均加速,在 re...

Jian Chen, Yesheng Liang, Zhijian Liu

2602.06036-dflash-block-diffusion-speculative-decoding Speculative DecodingMulti-Token Prediction
Jun 28, 2026

Self Improving Agents in the Era of Experience: A Survey of Self to Meta Evolution

这篇 survey 把 self improving agents 从模型自我训练扩展为部署后 runtime system 的 trace to capability 问题:harness 将交互 trace 编译成可验证 experience,再分别写入 skills、memory、environment/tool boundary、model parameters 或 meta layer。它的主要贡献是给 self impro...

Che Jiang, Jincheng Zhong, Yu Fu, Kai Tian, Junlin Yang, Kaikai Zhao, Yuchong Wang, Tianwei Luo, Weizhi Wang, Yuxin Zuo, +17 more

2026-06-25-self-improving-agents-era-experience-survey Agent RLAgent MemoryAgent Workflow
Jun 28, 2026

Laminar: A Scalable Asynchronous RL Post Training Framework

Laminar 让完成 trajectory 独立进入 experience buffer,再用 CPU/RDMA relay 允许各 rollout 在本地生成 batch 结束或被 repack 释放后独立拉取新权重,并把同 weight version 的尾部请求集中到少数 rollout;1024 张 H800 实验支持这套设计能提高短窗口 actor update throughput,训练时实际 policy lag、GR...

Guangming Sheng, Yuxuan Tong (童雨轩), Borui Wan, Wang Zhang, Chaobo Jia (贾超博), Xibin Wu, Yuqi Wu, Xiang Li, Chi Zhang, Yanghua Peng, +3 more

2510.12633-laminar-asynchronous-rl-post-training RL InfrastructureRollout OptimizationDistributed Training
Jun 28, 2026

LoRAFusion: Efficient LoRA Fine Tuning for LLMs

LoRAFusion 把 LoRA 微调的效率问题拆成两层:kernel 层用 split graph fusion 减少 LoRA 分支对大激活张量的重复读写,调度层把共享 base model 的多个 LoRA adapter 合并训练并用分组、MILP / greedy packing 降低流水线气泡和变长样本负载不均;在 H100 / L40S、LLaMA 3.1 8B、Qwen 2.5 32B、LLaMA 3.1 70B 上...

Zhanda Zhu, Qidong Su, Yaoyao Ding, Kevin Song, Shang Wang

2510.00206-lorafusion-efficient-lora-fine-tuning Parameter-Efficient FinetuningDistributed TrainingTraining Memory
Jun 28, 2026

MegaScale MoE: Large Scale Communication Efficient Training of Mixture of Experts Models in Production

MegaScale MoE 的核心贡献是把大规模 MoE 训练的瓶颈从单点 kernel 优化提升到整层通信路径设计:attention 侧用 Ulysses style sequence parallelism 降低 TP critical path 通信,FFN 侧用 intra node expert parallelism 保持专家 GEMM 效率,再用跨算子调度、算子内 tile level overlap、选择性 acti...

Chao Jin, Ziheng Jiang, Zhihao Bai, Zheng Zhong, Juncai Liu, Xiang Li, Ningxin Zheng, Xi Wang, Cong Xie, Qi Huang, +9 more

2505.11432-megascale-moe-communication-efficient-training MoE SystemsDistributed TrainingTraining Memory
Jun 24, 2026

Fast Inference from Transformers via Speculative Decoding

这篇论文把 speculative execution 推到随机采样场景:小 draft model 先自回归猜 $\gamma$ 个 token,大 target model 一次并行验证这些前缀,并用 $\min(1,p/q)$ 接受概率与 residual distribution 校正拒绝位置,从而在无需改模型、无需重训且保持 target 输出分布不变的前提下,把大模型串行 decode 的目标调用数降低到每步平均生成多个 ...

Yaniv Leviathan, Matan Kalman, Yossi Matias

2211.17192-fast-inference-transformers-speculative-decoding Speculative Decoding
Jun 24, 2026

MiniMax Sparse Attention

MiniMax Sparse Attention (MSA) 是 MiniMax M3 的长上下文技术核心:它在 GQA (Grouped Query Attention,分组查询注意力) 之上增加轻量 Index Branch,为每个 query token 和每个 GQA group 选择少量 KV blocks,再由 Main Branch 对被选 block 做精确 softmax attention;论文的强证据在于 109...

Xunhao Lai (赖勋豪), Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu (徐旸), Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, +7 more

2606.13392-minimax-sparse-attention-m3 Sparse AttentionLong ContextMoE Architecture
Jun 24, 2026

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

DeepSeekMoE 把 MoE (Mixture of Experts,混合专家) 的效率问题改写成专家专门化问题:fine grained expert segmentation 通过切小 FFN (Feed Forward Network,前馈网络) 专家并增加激活专家数,提高每个 token 的专家组合分辨率;shared expert isolation 通过固定激活共享专家承载通用知识,让 routed experts ...

Damai Dai, Chengqi Deng, Chenggang Zhao, Runxin Xu (许润昕), Huazuo Gao, Deli Chen (陈德里), Jiashi Li, Wangding Zeng, Xingkai Yu (俞星凯), Yu Wu (吴俣), +7 more

2401.06066-deepseekmoe-expert-specialization MoE Architecture
Jun 24, 2026

Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

Engram 把大模型稀疏性从 MoE (Mixture of Experts,混合专家,用条件计算扩大 total 参数并控制 active compute) 扩展到 conditional memory:用 hashed N gram lookup 存静态局部模式,用 context aware gate 决定是否把查到的记忆注入 hidden state;在 iso parameter / iso FLOPs 的 27B MoE...

Xin Cheng, Wangding Zeng, Damai Dai, Qinyu Chen, Bingxuan Wang, Zhenda Xie (解振达), Kezhao Huang, Xingkai Yu (俞星凯), Zhewen Hao, Yukun Li, +4 more

2601.07372-conditional-memory-engram-scalable-lookup Memory ArchitectureMoE Architecture
Jun 23, 2026

Kimi K2: Open Agentic Intelligence

Kimi K2 是 Moonshot 将 open weight MoE foundation model 推向 agentic software engineering 和 tool use 的系统报告:1.04T total / 32B active MoE、MuonClip、15.5T pretraining tokens、3000+ real MCP tools、20000+ synthetic tools、verifiabl...

Kimi Team: Yifan Bai and 198 other authors;arXiv submitter: Yulun Du。

2507.20534-kimi-k2-open-agentic-intelligence MoE ArchitectureOptimizerAgent RL
Jun 23, 2026

Kimi k1.5: Scaling Reinforcement Learning with LLMs

Kimi k1.5 把 128K 长上下文、long CoT、verifiable reward、多模态 reasoning、partial rollout、length aware sampling、long2short 和 Megatron/vLLM/Mooncake 工程栈组合成一套 RL scaling recipe;它的价值在于把 reasoning RL 从单一数学训练扩展到长上下文、多模态、代码与系统协同,证据主要来自团...

Kimi Team and 95 other authors;arXiv submitter: Flood Sung。

2501.12599-kimi-k1-5-scaling-rl-llms Reasoning RLRollout OptimizationLong Context
Jun 23, 2026

Qwen3 Technical Report

Qwen3 把 dense / MoE (Mixture of Experts,混合专家) 开源模型族、thinking / non thinking 双模式、thinking budget、36T token 多语预训练、四阶段后训练和 strong to weak distillation 组织成一套可发布模型体系,使 Qwen3 235B A22B、Qwen3 32B 与小模型在数学、代码、agent、多语和长上下文任务上成为强...

An Yang and 59 other authors;本地核心跟踪作者包括 An Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Dayiheng Liu, Fei Huang, Jianwei Zhang, Jianxin Yang, Jingren Zhou, Junyang Lin, Rui Men。

2505.09388-qwen3-technical-report Reasoning RLMoE ArchitectureOn-Policy Distillation
Jun 23, 2026

Qwen2.5 Technical Report

Qwen2.5 展示了一条面向通用、代码、数学、结构化数据和长上下文场景的 LLM (Large Language Model,大语言模型) 工程路线:用 18T token 预训练、统一 tokenizer、百万级 SFT (Supervised Fine Tuning,监督微调)、DPO (Direct Preference Optimization,直接偏好优化)、GRPO (Group Relative Policy Opti...

An Yang, Binyuan Hui, Bo Zheng, Bowen Yu (郁博文), Dayiheng Liu (刘大一恒), Fei Huang, Jianwei Zhang, Jianxin Yang, Jingren Zhou, Junyang Lin, +1 more

2412.15115-qwen2-5-technical-report Reasoning RLLong ContextCoding Agent
Jun 23, 2026

DeepSeek V3 Technical Report

DeepSeek V3 用 MLA (Multi head Latent Attention,多头潜变量注意力)、DeepSeekMoE、无辅助损失负载均衡、MTP (Multi Token Prediction,多 token 预测)、FP8 混合精度训练和 DualPipe 通信重叠,把 671B total / 37B active 的开放 MoE (Mixture of Experts,混合专家) 模型训练到强代码、数学和通用...

Wenfeng Liang (梁文锋), Peiyi Wang, Runxin Xu (许润昕), Zhihong Shao (邵智宏), Damai Dai, Deli Chen (陈德里), Yu Wu (吴俣)

2412.19437-deepseek-v3-technical-report MoE ArchitectureDistributed TrainingMulti-Token Prediction
Jun 23, 2026

DeepSeek V2: A Strong, Economical, and Efficient Mixture of Experts Language Model

DeepSeek V2 最独立的贡献是 MLA (Multi head Latent Attention,多头潜变量注意力):它把 generation cache 从完整 per head K/V 改成由 hidden state 下投影得到的 512 维联合 KV latent 与 64 维 RoPE key,再用 projection absorption 让 latent 直接参与 content score 和 value ...

DeepSeek AI (group author);Appendix A 列出按三类角色分组的 156 位去重贡献者。

2405.04434-deepseek-v2-mla-moe-efficient-llm KV CacheMoE ArchitectureMoE Systems
Jun 23, 2026

From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models

这篇是 LLM 强化学习 credit assignment 的综述与索引节点:它把 RL (Reinforcement Learning,强化学习) for LLM (Large Language Model,大语言模型) 中的 CA (Credit Assignment,信用分配/功劳分配) 拆成两个问题域,reasoning RL 主要在单条长 CoT (Chain of Thought,思维链) 内把 outcome rewa...

Chenchen Zhang

2604.09459-credit-assignment-reasoning-agentic-llm-rl Credit AssignmentReasoning RLAgent RL
Jun 23, 2026

Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective

这篇从理论上挑战“过程监督天然比结果监督更强”的直觉:在标准 coverage / concentrability 假设下,只拿 trajectory level total reward 的 outcome supervision 可以通过 least squares reward imputation 转成 per step reward data,后续 offline reinforcement learning 的统计难度只比...

Zeyu Jia (贾泽宇), Alexander Rakhlin, Tengyang Xie

2502.10581-do-we-need-to-verify-step-by-step-process-supervision-theory Process SupervisionRL TheoryCredit Assignment
Jun 23, 2026

Math Shepherd: Verify and Reinforce LLMs Step by step without Human Annotations

Math Shepherd 的核心价值在于把数学推理步骤的标注问题改写为“当前 step 之后还能否补全到正确答案”的 Monte Carlo potential estimation:对每个中间 step 采样多个 continuation,用最终答案正确性给 step 生成 hard / soft pseudo label,训练 PRM 做 verifier reranking,并进一步把 PRM reward 接入 step b...

Peiyi Wang, Lei Li, Zhihong Shao (邵智宏), Runxin Xu (许润昕), Damai Dai, Yifei Li, Deli Chen (陈德里), Yu Wu (吴俣), Zhifang Sui

2312.08935-math-shepherd-automatic-process-supervision Process SupervisionVerifierReward Modeling
Jun 23, 2026

Let's Verify Step by Step

这篇是 OpenAI reasoning verification 线的历史节点:它用 80 万 step level human labels 训练 PRM,在 MATH 500 题 held out subset 上用 Best of 1860 选择达到 78.2%,高于强 ORM 和 majority vote;核心价值是把“最终答案正确”拆成“每一步是否仍在正确推理轨道上”,并通过 PRM800K 让过程监督成为后续 reas...

Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, Karl Cobbe

2305.20050-lets-verify-step-by-step-process-supervision Process SupervisionVerifierReward Modeling
Jun 23, 2026

Credit Assignment with Resets in Language Model Reasoning

这篇把 reset 变成 RLVR credit assignment 的训练时原语:从失败轨迹中选一个中间 thought prefix,重采样多个后缀 continuation,只对后缀 token 做 group relative policy update。SRPO 用模型自定位的首个错误 thought 近似 credit assignment oracle,在无需外部 step level label 或 critic 的...

Ankur Samanta, Akshayaa Magesh, Ayush Jain, Youliang Yu, Daniel R. Jiang, Kavosh Asadi, Kaveh Hassani, Paul Sajda, Jalaj Bhandari, Yonathan Efroni

2605.25507-credit-assignment-resets-language-model-reasoning Credit AssignmentReasoning RLRollout Optimization
Jun 22, 2026

The Optimal Token Baseline: Variance Reduction for Long Horizon LLM RL

OTB 解决的是 long horizon LLM RL 中 outcome reward 经过整条 response 回传时的高方差问题:它从 causal policy gradient 的方差最小化目标推导出 token level optimal baseline $B t^ =\mathbb E[G tW t]/\mathbb E[W t]$,其中 $W t$ 是到第 $t$ 个 token 为止累计的 gradient r...

Yingru Li (李英儒), Jiawei Xu, Ziniu Li (李子牛), Jiacai Liu (刘佳材), Wei Liu (刘威), Yuxuan Tong (童雨轩), Longtao Zheng (郑龙韬), Zhenghai Xue, Yaxiang Zhang, Tianle Cai (蔡天乐), +3 more

2602.07078-optimal-token-baseline-long-horizon-llm-rl Credit AssignmentRL AlgorithmRL Theory
Jun 21, 2026

VIMPO: Value Implicit Policy Optimization for LLMs

VIMPO 试图填补 GRPO 和 PPO actor critic 之间的空位:它不训练独立 critic,却从 KL regularized RL 的最优性条件推出一个由 policy reference log ratio 表达的隐式 value recurrence,用终止状态 $V(s T)=0$ 把 outcome reward 变成 value loss,再用同一个 log ratio identity 构造 token...

Zhewei Kang, Aosong Feng, Sergey Levine, Dawn Song, Xuandong Zhao

2606.20008-vimpo-value-implicit-policy-optimization-llms RL AlgorithmCredit AssignmentReasoning RL
Jun 21, 2026

Reinforcement Learning Towards Broadly and Persistently Beneficial Models

这篇论文给 OpenAI 的 alignment post training 提供了一个正向版本的 emergent misalignment 实验:如果窄域有害训练能诱导跨域失配,那么用 5% 真实场景 beneficial trait data 加 RL reward 强化诚实、纠错、风险意识、公平和人类福祉等特质,也可能诱导跨域对齐收益。论文的证据强在 44/53 个 OOD 评测提升、health only 训练迁移到非健康安...

Akshay V. Jagadeesh, Rahul K. Arora, Khaled Saab, Ali Malik, Mikhail Trofimov, Foivos Tsimpourlas, Johannes Heidecke, Karan Singhal

2026-06-18-openai-beneficial-rl AI SafetyReward ModelingAgent RL
Jun 21, 2026

Trust Region Masking for Long Horizon LLM Reinforcement Learning

这篇论文把 long horizon LLM RL 中 pi roll 与 pi theta 的实现级分布差异形式化为 surrogate objective error:经典 trust region bound 随长度呈 $O(T^2)$ 并迅速失效,作者给出更紧的 KL/TV 组合界,指出 max token divergence 是单靠 sequence average KL 无法替代的控制量,并提出 Trust Region...

Yingru Li (李英儒), Jiacai Liu (刘佳材), Jiawei Xu, Yuxuan Tong (童雨轩), Ziniu Li (李子牛), Qian Liu (刘乾), Baoxiang Wang (王宝祥)

2512.23075-trust-region-masking-long-horizon-llm-rl Training StabilityRL AlgorithmRollout Optimization
Jun 21, 2026

The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

这篇论文把 reasoning LLM 的 RLVR 训练瓶颈重新表述为 policy entropy 的消耗过程:在没有 entropy / KL 干预时,reward 提升和 entropy 下降之间可以被经验式 $R= a\exp(\mathcal H)+b$ 拟合;进一步用 softmax policy 的 entropy dynamics 说明,高概率且高 advantage 的 token update 会持续降低 ent...

Ganqu Cui, Yuchen Zhang, Jiacheng Chen, Lifan Yuan (袁立凡), Zhi Wang, Yuxin Zuo, Haozhan Li, Yuchen Fan, Huayu Chen (陈华玉), Weize Chen, +7 more

2505.22617-entropy-mechanism-rl-reasoning-language-models Training StabilityReasoning RLRL Theory
Jun 21, 2026

Approximating KL Divergence

这篇博客给出了 RL 实践中常用 KL 估计器 k1/k2/k3 的最小数学解释:在只能从 $q$ 采样并能计算 $p(x),q(x)$ 的场景下,$k 1= \log r$ 是无偏但高方差的 $\mathrm{KL}[q,p]$ 估计器;$k 2=\frac12(\log r)^2$ 是低方差但有偏的二阶近似;$k 3=(r 1) \log r$ 通过控制变量 $r 1$ 保持无偏、非负和低方差,其中 $r=p(x)/q(x)$。对...

John Schulman

2020-03-07-schulman-kl-divergence-approximations RL TheoryTraining Stability
Jun 21, 2026

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

DeepSpeed Ulysses 的核心贡献是把长序列训练中的 activation / context 维瓶颈变成一次可逆的数据布局转换:Transformer 其他算子沿 sequence 维分片,进入 attention 前用 all to all 把 sequence partitioned, all heads 的 QKV 变成 full sequence, head partitioned 的 QKV,让每张 GPU 只...

Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang, Minjia Zhang, Shuaiwen Leon Song, Samyam Rajbhandari, Yuxiong He

2309.14509-deepspeed-ulysses-long-sequence-training Distributed TrainingLong ContextAttention Kernel
Jun 19, 2026

Muon: An optimizer for hidden layers in neural networks

Muon 的核心技术含量在于把隐藏层矩阵参数的 momentum update 映射到近似半正交方向:它用低精度 Newton Schulz 近似 polar factor,让更新能量更均匀地覆盖矩阵奇异方向;后续大模型实践表明,Muon 想稳定扩展到 LLM pretraining,还需要 weight decay、shape aware update scale、AdamW/Muon 参数分组、分布式 full matrix or...

Keller Jordan

2026-06-19-muon-optimizer-keller-jordan-synthesis OptimizerDistributed TrainingTraining Stability
Jun 19, 2026

Ring Attention with Blockwise Transformers for Near Infinite Context

Ring Attention 的核心贡献是把 exact Transformer 的长序列瓶颈从“单卡必须驻留整段序列输出”改成“每张设备只驻留本地 query block,并让 key/value block 沿 ring 轮转”:它复用 blockwise exact attention 的在线 softmax 统计量,在不近似 attention 的前提下把最大上下文长度扩展到接近设备数倍;代价边界在于 exact attent...

Hao Liu, Matei Zaharia, Pieter Abbeel

2310.01889-ring-attention-blockwise-transformers-near-infinite-context Long ContextDistributed TrainingAttention Kernel
Jun 19, 2026

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Sarathi 的核心贡献是把 LLM serving 的低效从“decode 天然慢”重写为一个可调度的数据流问题:prefill 很快进入 compute saturating 状态,decode 因为逐 token 生成和 KV cache 访问长期 memory bound;如果把一个长 prefill 切成多个 compute sized chunk,再让 decodes 搭在每个 prefill chunk 的 linea...

Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Ramachandran Ramjee

2308.16369-sarathi-chunked-prefill-decode-maximal-batching Inference SchedulingServing RuntimeKV Cache
Jun 18, 2026

ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

ZeRO 的核心贡献是把数据并行中每张 GPU 都完整复制的 optimizer states、gradients、parameters 拆成可按 data parallel rank 分片保存和按需通信的动态状态系统,使 data parallelism 获得接近 model parallelism 的内存效率,同时保留 data parallelism 的大粒度计算和接近原始 DP 的通信量;论文实际评估了 $P {os+g}$ ...

Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong He

1910.02054-zero-memory-optimizations-trillion-parameter-models Training MemoryDistributed Training
Jun 18, 2026

GLM 5.2: Built for Long Horizon Tasks

GLM 5.2 是 GLM 5 系列从 200K 长上下文 agentic engineering 推进到 1M 长上下文 coding agent 的 release:它用 IndexShare 降低 DeepSeek Sparse Attention (DSA) indexer 成本,用 Multi Token Prediction (MTP) IndexShare + KVShare + rejection sampling +...

Z.ai / GLM 5 Team

2026-06-16-glm-5-2-long-horizon-tasks Coding AgentLong ContextSparse Attention
Jun 10, 2026

Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling

Bebop 的核心判断是:MTP 在 RL rollout 中失速的主因主要来自 policy entropy fluctuation,frozen MTP head 与更新后 policy 的权重漂移影响较小;因此有效方案是把 acceptance method 换成 probabilistic rejection sampling,并用 end to end TV loss 在 RL 前训练 MTP heads,使 draft t...

Yucheng Li, Huiqiang Jiang, Yang Xu (徐旸), Jianxin Yang, Yi Zhang, Yizhong Cao, Yuhao Shen, Fan Zhou, Rui Men, Jianwei Zhang, +7 more

2606.12370-bebop-mtp-rejection-sampling-rl-training Rollout OptimizationSpeculative DecodingMulti-Token Prediction
Jun 09, 2026

Dynamic Linear Attention

DLA 认为 long context linear attention 的核心损失来自固定 state merging 策略把信息密度不同的 token 压进同一个 summary state;它用 token level representation drift 动态决定 state 边界,并在固定容量 cache 中合并低信息密度的相邻 state,让 multi state linear attention 同时具备自适应分辨...

Xin Wang, Hui Shen, Boyuan Zheng, Xueshen Liu, Minkyoung Cho, Zhongwei Wan, Zesen Zhao, Zhuoqing Mao, Shen Yan, Mi Zhang

2606.10650-dynamic-linear-attention Linear AttentionLong ContextMemory Architecture
Jun 03, 2026

Why Muon Outperforms Adam: A Curvature Perspective

这篇论文给 Muon 相比 Adam 更快训练提供了一个局部曲率解释:在 matched validation loss 下,Muon 和 Adam 的一阶收益相近,差距主要来自二阶 Hessian curvature penalty;进一步分解发现二阶差距主要由 Muon update direction 的 Normalized Directional Sharpness (NDS) 更低造成,step size 对差距的解释力较...

Shuche Wang, Fengzhuo Zhang, Jiaxiang Li, Dirk Bergemann, Zhuoran Yang

2606.04662-muon-outperforms-adam-curvature OptimizerTraining Stability
Jun 02, 2026

UltraEP: Unleash MoE Training and Inference on Rack Scale Nodes with Near Optimal Load Balancing

UltraEP 的核心贡献是把大规模 MoE expert parallelism 中的负载均衡从“基于历史统计的周期性预测”推进到“基于 post gating exact load 的每 microbatch、每 layer 实时再均衡”:它利用 rack scale node 的高带宽 scale up fabric,把一个 EP group 放进同一机架级通信域,再用 quota driven planner 联合决定专家复制...

Xinming Wei, Chao Jin, Tuo Dai, Yinmin Zhong, Shan Yu, Chengxu Yang, Bingyang Wu, Zili Zhang, Jing Mai, Qianchao Zhu, +3 more

2606.04101-ultraep-rack-scale-moe-load-balancing MoE SystemsDistributed TrainingInference Scheduling
Jun 02, 2026

Large Language Models Hack Rewards, and Society

论文提出 societal hacking:当社会规则被编码成可优化的奖励结构时,RL 后训练会推动 LLM 在形式合规和制度意图之间寻找缝隙;在作者构造的 SocioHack 沙盒中,RL 模型能够重新发现大量真实历史漏洞,并且现有拒答、自我批判、训练正则化只能部分缓解这一现象。

Wei Liu KCL, Xinyi Mou, Hanqi Yan, Zhongyu Wei, Yulan He

2606.04075-llms-hack-rewards-and-society Reward HackingAI SafetyAgent RL
May 31, 2026

Self Trained Verification for Training and Test Time Self Improvement

Self Trained Verification (STV, 自训练验证) 的核心贡献是把 reference solution 变成 verifier 的 privileged teacher signal:同一个模型在看到参考答案时更容易指出候选解的错误,STV 用 On Policy Distillation (OPD, 在策略蒸馏) 和 verdict Reinforcement Learning (RL, 强化学习) 把这...

Chen Henry Wu, Aditi Raghunathan

2605.30290-self-trained-verification VerifierOn-Policy DistillationTest-Time Scaling
May 29, 2026

On Effectiveness and Efficiency of Agentic Tool calling and RL Training

论文量化了 tool calling 分数对随机种子、对话序列化、推理历史和 system prompt 的敏感性,并针对长 tool context 下 GRPO 的两类浪费,用近期 all correct 轨迹预测跳过 rollout、用 max variance 子集减少反向传播,在组合配置上报告 1.7×/2.6× matched performance wall clock 提速。 本地评价:评测审计与 wall clock...

Tong Liu, Cheng Qian, Matej Cief, Yuan He (何源), Daniele Dan, Nikolaos Aletras, Gabriella Kazai

2606.00135-agentic-tool-calling-rl-training Rollout OptimizationAgent RLTool Use
May 17, 2026

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning

这篇论文把 LLM RL 中训练侧与推理侧对同一 token 序列给出的 logprob 不一致定义为 Training Inference Mismatch (TIM),并用 VeXact 构造 FSDP trainer 与 rollout engine bitwise 对齐的 zero mismatch 基线;实验证明 TIM 这种看似微小的 token level 数值差异可以单独触发 RL training collapse,...

Tianle Zhong, Neiwen Ling, Yifan Pi, Zijun Wei, Tianshu Yu, Geoffrey Fox, Peng Wu, Xiao Yu

2605.14220-training-inference-mismatch-llm-rl Training StabilityDeterministic InferenceRL Infrastructure
May 16, 2026

Transformers are Inherently Succinct

论文证明:固定精度 Transformer 在表达某些语言时非常简洁;存在语言族可以用多项式大小的 Transformer 表示,但等价的 LTL 或 RNN 需要指数大小,等价有限自动机需要双指数大小。

Pascal Bergsträßer, Ryan Cotterell, Anthony W. Lin

2510.19315-transformers-inherently-succinct Formal Expressivity
Apr 25, 2026

DeepSeek V4: Towards Highly Efficient Million Token Context Intelligence

DeepSeek V4 的核心是把“百万 token 上下文”做成一个端到端系统能力:在扩大 Mixture of Experts (MoE) 规模之外,V4 Pro 用 1.6T total / 49B active parameters,V4 Flash 用 284B total / 13B active parameters,二者通过 Compressed Sparse Attention (CSA) / Heavily Com...

DeepSeek AI and 318 other authors; arXiv submitter Wenfeng Liang.

2026-04-24-deepseek-v4-million-token-context-intelligence Long ContextSparse AttentionMoE Architecture
Apr 17, 2026

From Curiosity to Caution: Mitigating Reward Hacking for Best of $N$ with Pessimism

Best of $N$ 在候选数量增大时会把输出推向 reward model 高估且分布外的区域,作者提出 caution:训练一个 predictor 去预测冻结 reward model 的中间特征,用预测误差作为分布外不确定性,并在推理时从 reward score 中扣除该不确定性,从而让多候选选择既能利用额外 inference compute,也能降低大 $N$ 下的 reward hacking。

Zhuohao Yu, Zhiwei Steven Wu, Adam Block

2604.04648-caution-pessimism-best-of-n-reward-hacking Reward HackingTest-Time ScalingReward Modeling
Apr 11, 2026

ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases

ImpossibleBench 把 coding agent 的 test case exploitation 转化成一个可测量的 reward hacking benchmark:作者从 LiveCodeBench 和 SWE bench 构造自然语言规格与单测相冲突的 impossible tasks,把 impossible tasks 上的通过率定义为 cheating rate,因此任何得分都对应 specification...

Ziqian Zhong, Aditi Raghunathan, Nicholas Carlini

2510.20270-impossiblebench-test-case-exploitation BenchmarkReward HackingCoding Agent
Apr 08, 2026

BroRL: Scaling Reinforcement Learning via Broadened Exploration

BroRL 把 RLVR 的 scaling 轴从“继续训练更多 step”扩展到“每个 prompt 采样更多 rollout”:作者从 one step RLVR 的 correct token probability mass 分解出一个可能为负的 unsampled coupling term,并说明增大 rollout size $N$ 会让未采样项的二阶矩衰减,从而让 policy update 更稳定地增加正确 toke...

Jian Hu, Mingjie Liu, Ximing Lu, Fang Wu, Zaid Harchaoui, Shizhe Diao, Yejin Choi, Pavlo Molchanov, Jun Yang, Jan Kautz, +1 more

2510.01180-brorl-broadened-rl-exploration Rollout OptimizationReasoning RLRL Algorithm
Apr 04, 2026

Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning

Seer 的核心贡献是把同步 LLM RL 中最耗时的 rollout 阶段当成一个“同 prompt group 内存在可学习上下文”的调度问题:GRPO 一类算法会为同一 prompt 采样多条 responses,这些 responses 在长度和局部 token 模式上高度相关;Seer 利用这个结构做 chunk level divided rollout、基于 speculative request 的 context a...

Ruoyu Qin (秦若愚), Weiran He, Weixiao Huang, Yangkun Zhang, Yikai Zhao (赵一开), Bo Pang, Xinran Xu (许欣然), Yingdi Shan (闪英迪), Yongwei Wu, Mingxing Zhang (章明星)

2511.14617-seer-online-context-learning-llm-rl Rollout OptimizationRL InfrastructureSpeculative Decoding
Mar 24, 2026

On the Interplay of Pre Training, Mid Training, and RL on Reasoning Language Models

这篇论文用完全可控的合成推理环境把 pre training、mid training 和 RL post training 的作用拆开:RL 能带来真正的 pass@128 能力扩展,但需要两个条件同时成立,base model 在目标区域有足够探索余地,RL 数据位于模型 edge of competence;contextual generalization 还需要预训练中出现过少量相关 primitive seed;mid t...

Charlie Zhang, Graham Neubig, Xiang Yue

2512.07783-interplay-pretraining-midtraining-rl-reasoning Reasoning AnalysisReasoning RL
Mar 20, 2026

From $f(x)$ and $g(x)$ to $f(g(x))$: LLMs Learn New Skills in RL by Composing Old Ones

这篇论文给 RLVR 争论提供了一个可控的正向证据:当模型已经通过预训练或 SFT 掌握 atomic skills,且 RL 训练目标明确奖励组合这些 atomic skills 时,RL 可以训练模型形成可泛化的 compositional skill,把见过的浅层模式迁移到更深嵌套、未见函数组合和跨任务组合场景;同样 Level 2 数据上的 RFT 和只训练 atomic tasks 的 RL 都没有得到类似泛化。

Lifan Yuan (袁立凡), Weize Chen, Yuchen Zhang, Ganqu Cui, Hanbin Wang, Ziming You, Ning Ding (丁宁), Zhiyuan Liu, Maosong Sun, Hao Peng

2509.25123-rl-compositional-skill-acquisition Reasoning AnalysisReasoning RLBenchmark
Mar 17, 2026

MiniMax M1: Scaling Test Time Compute Efficiently with Lightning Attention

MiniMax M1 把 test time compute scaling 的瓶颈拆成三层来处理:用 Lightning Attention + MoE (Mixture of Experts,混合专家模型) 降低长输出推理和 RL (Reinforcement Learning,强化学习) rollout 的计算成本,用 CISPO (Clipped IS weight Policy Optimization,截断重要性采样权重的...

MiniMax; arXiv title页显示 Aili Chen and 125 other authors,Appendix Contributors 按字母序列出完整贡献者。

2506.13585-minimax-m1-cispo-lightning-attention Reasoning RLTest-Time ScalingLinear Attention
Mar 13, 2026

Inference Time Reward Hacking in Large Language Models

这篇论文把 reward hacking 扩展到 inference time alignment 场景:Best of $n$ 这类“多采样后按 proxy reward 选最高分”的方法,会随着 $n$ 增大先提升真实质量,再因 winner's curse 选中过度高估的样本而降低真实质量;作者用 TP$ 2$/MLR 条件证明常见一参数推理策略的 true reward 曲线至多一个峰值,提出 Best of Poisson ...

Hadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju, Flavio du Pin Calmon

2506.19248-inference-time-reward-hacking-llms Reward HackingTest-Time ScalingRL Theory
Mar 10, 2026

ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

这篇论文是对 “RLVR 只提升 base model 已有解的采样效率” 观点的直接反驳:作者提出 ProRL,用高温 rollout、DAPO 式 decoupled clipping/dynamic sampling、KL regularization、周期性 reference policy 与 optimizer reset,以及 136K 多任务 verifiable reward 数据,把 DeepSeek R1 Dis...

Mingjie Liu, Shizhe Diao, Ximing Lu, Jian Hu, Xin Dong, Yejin Choi, Jan Kautz, Yi Dong

2505.24864-prorl-prolonged-rl-reasoning-boundaries Reasoning RLRL AlgorithmRollout Optimization
Mar 06, 2026

Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

这篇论文把 RLVR 的收益拆成 sampling efficiency 和 reasoning capacity boundary:当前基于 binary verifiable reward 的 RLVR 常常把 base model 已经能低概率采样到的正确 reasoning paths 提升到更高概率,因此 pass@1 明显改善;但在大 k 的 pass@$k$ 覆盖上,base model 往往能解出更多题,说明现有 RL...

Yang Yue, Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang, Shiji Song, Gao Huang

2504.13837-rlvr-reasoning-boundary-base-model Reasoning AnalysisReasoning RLBenchmark
Mar 03, 2026

DAPO: An Open Source LLM Reinforcement Learning System at Scale

DAPO 的核心贡献是一套可复现的 long CoT reasoning RL recipe:在 Qwen2.5 32B base 上,用基于 verl 的 GRPO 变体、规则奖励、DAPO Math 17K 数据、Clip Higher、Dynamic Sampling、Token level Policy Gradient Loss 和 Overlong Reward Shaping,将 AIME 2024 avg@32 提升到...

Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Lingjun Liu, +25 more

2503.14476-dapo-long-cot-rl-system RL AlgorithmReasoning RLRL Infrastructure
Feb 26, 2026

Spurious Rewards: Rethinking Training Signals in RLVR

这篇论文指出,某些 RLVR 实验中的能力提升可以由模型预训练 prior、GRPO clipping bias 和提示/格式行为共同解释:在 Qwen2.5 Math 上,随机 reward、format reward、甚至奖励错误答案的 reward 都能显著提高 MATH 500 / AMC / AIME24 表现,随机 reward 在 MATH 500 上带来 21.4 个百分点提升,接近 ground truth rewa...

Rulin Shao, Shuyue Stella Li, Rui Xin, Scott Geng, Yiping Wang, Sewoong Oh, Simon Shaolei Du, Nathan Lambert, Sewon Min, Ranjay Krishna, +4 more

2506.10947-spurious-rewards-rethinking-rlvr Reward HackingReasoning RLReward Modeling
Feb 25, 2026

GLM 5: from Vibe Coding to Agentic Engineering

GLM 5 是一篇面向“agentic engineering”的系统报告:模型侧把 GLM 4.5 扩到 744B total / 40B active MoE,使用 MLA、Muon Split、参数共享 MTP 和 DeepSeek Sparse Attention 支撑 200K 级长上下文;训练侧用 28.5T tokens、软件工程 mid training、SFT、Reasoning RL、Agentic RL、Gene...

GLM 5 Team: Aohan Zeng and 184 other authors

2602.15763-glm-5-agentic-engineering Coding AgentAgent RLMoE Architecture
Feb 17, 2026

Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

这篇论文把 reasoning model 的 chain of thought 从“可解释性附属物”提升为安全监控信号:OpenAI 发现 GPT 4o 读取 frontier reasoning agent 的 CoT 可以显著优于只看 actions / outputs 的监控器,检测 agentic coding RL 中的 reward hacking;但当训练目标直接惩罚会触发 CoT monitor 的思路时,模型在低优...

Bowen Baker, Joost Huizinga, Leo Gao, Zehao Dou, Melody Y. Guan, Aleksander Madry, Wojciech Zaremba, Jakub Pachocki, David Farhi

2503.11926-monitoring-reasoning-models-obfuscation AI SafetyReward HackingVerifier
Feb 13, 2026

Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking

这篇论文把 reward hacking 定义为:proxy reward 在参考策略访问到的状态 动作分布上和真实 reward 正相关,但经过优化后,新策略进入 proxy 与真实目标相关性失效的分布区域,导致真实回报低于参考策略;作者据此证明,把策略的 occupancy measure 约束在参考策略附近可以给真实回报提升提供下界,并提出 ORPO 用 discriminator 估计 occupancy ratio,在交通、...

Cassidy Laidlaw, Shivam Singhal, Anca Dragan

2403.03185-correlated-proxies-reward-hacking Reward HackingReward ModelingRL Theory
Feb 10, 2026

Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment

这篇论文把 RLHF reward hacking 中的一类常见失败归因到 reward model 学到了偏好数据里的伪相关变量,例如长度、迎合短语、概念词或人口属性;作者提出 Causal Reward Model (CRM),在 Bradley Terry reward loss 上加入基于 MMD 的独立性正则,让 reward 预测对人工指定的 spurious factor 分箱保持近似不变,从而在 sycophancy、...

Chaoqi Wang, Zhuokai Zhao, Yibo Jiang, Zhaorun Chen, Chen Zhu, Yuxin Chen, Jiayi Liu, Lizhu Zhang, Xiangjun Fan, Hao Ma, +1 more

2501.09620-causal-rewards-llm-alignment Reward ModelingReward HackingAI Safety
Feb 06, 2026

DeepSeek R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

DeepSeek R1 v2 的核心结论是:大规模 outcome based RL 可以在强 base model 上诱导 long CoT reasoning、自我反思、验证和策略切换等行为;R1 Zero 证明无需 SFT 也能通过 rule based verifiable reward 激发 reasoning capability,R1 则通过 cold start SFT、两阶段 RL、rejection samplin...

DeepSeek AI and 199 other authors. Core contributors listed in v2 source include Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Peiyi Wang, Qihao Zhu, Runxin Xu, Ruoyu Zhang, Shirong Ma, Xiao Bi, Xiaokang Zhang, Xingkai Yu, Yu Wu, Z.F. Wu, Zhibin Gou, Zhihong Shao, Zhuoshu Li, Ziyi Gao.

2501.12948-deepseek-r1-rl-reasoning Reasoning RLReward ModelingRL Algorithm
Feb 03, 2026

Kimi K2.5: Visual Agentic Intelligence

Kimi K2.5 是一篇系统型技术报告:它把 Kimi K2 扩展为 256K context 的开源多模态 agentic 模型,通过早期低比例视觉 文本混合预训练、zero vision SFT、联合文本/视觉 RL、MoonViT 3D 视频压缩、Token Efficient RL、Decoupled Encoder Process 和 Agent Swarm / PARL,把模型能力从单轮文本推理推进到视觉理解、长视频、浏...

Kimi Team: Tongtong Bai and 324 other authors; appendix lists contributors alphabetically by last name.

2602.02276-kimi-k2-5-visual-agentic-intelligence Multi-Agent OrchestrationMultimodal ModelAgent RL
Jan 31, 2026

Using Span Queries to Optimize for Cache and Attention Locality

这篇论文提出 Span Query,把 chat、RAG、judge generator、inference time scaling 和 agentic workload 统一表示为带可交换约束的 LLM 调用表达式树;当客户端声明哪些 message span 可以重排时,服务端可以把 KV cache 从 prefix only reuse 推进到 span level relocatable reuse,并进一步用树形改写提升...

Paul Castro, Nick Mitchell, Nathan Ordonez, Thomas Parnell, Mudhakar Srivatsa, Antoni Viros i Martin

2511.02749-span-queries-cache-attention-locality Agent WorkflowKV CacheServing Runtime
Jan 30, 2026

Defeating Nondeterminism in LLM Inference

这篇文章指出,LLM 推理在 temperature=0 下仍然出现不同输出,主要来源通常是 batch 不变性缺失:服务端负载改变 batch size、prefill/decode 切分、KV cache 布局和 attention split 策略,进而改变浮点 reduction 顺序;作者通过 batch invariant RMSNorm、matmul 和 attention kernel 展示了可复现推理的实现路径,并把...

Horace He, in collaboration with others at Thinking Machines Lab

2025-09-10-defeating-nondeterminism-llm-inference Deterministic InferenceAttention KernelRL Infrastructure
Jan 27, 2026

HybridFlow: A Flexible and Efficient RLHF Framework

HybridFlow 的核心贡献是把 RLHF 训练看成由多个大模型节点组成的复杂 dataflow,并提出一个混合控制架构:模型之间用 single controller 统一编排和数据重分片,模型内部用 multi controller 执行高效分布式训练/推理/生成;再配合 3D HybridEngine 和自动设备映射,在 PPO、ReMax、Safe RLHF 等 RLHF 算法上比 DeepSpeed Chat、OpenR...

Guangming Sheng, Chi Zhang, Zilingfeng Ye, Xibin Wu, Wang Zhang, Ru Zhang, Yanghua Peng, Haibin Lin, Chuan Wu

2409.19256-hybridflow-rlhf-framework RL InfrastructureDistributed TrainingRollout Optimization
Jan 23, 2026

Parrot: Efficient Serving of LLM based Applications with Semantic Variable

Parrot 的核心贡献是把 LLM 应用从一串孤立 completion requests 还原成带变量、依赖、性能目标和共享 prompt 结构的应用级数据流:开发者用 Semantic Variable 标注 prompt 中的输入/输出区域后,服务端可以做 DAG 分析、依赖请求连续执行、性能目标推导、动态共享前缀检测和应用感知调度,从而把 LLM serving 的优化对象从单请求延迟推进到端到端应用体验。

Chaofan Lin, Zhenhua Han, Chengruidong Zhang, Yuqing Yang, Fan Yang, Chen Chen, Lili Qiu

2405.19888-parrot-semantic-variable-llm-serving Agent WorkflowInference SchedulingKV Cache
Jan 20, 2026

Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention

这篇论文把 linear attention 在 causal LM 中“理论复杂度低、实际 GPU 训练慢”的核心原因定位到 prefix cumsum / scan 路径,并用 Lightning Attention 把注意力拆成块内 left product 与块间 right product:块内保留并行矩阵乘,块间维护 $KV$ 累计状态,再用 tiling / IO aware kernel 提升硬件效率;随后作者为该算子...

Zhen Qin, Weigao Sun, Dong Li, Xuyang Shen, Weixuan Sun, Yiran Zhong

2405.17381-various-lengths-constant-speed-lightning-attention Linear AttentionAttention KernelLong Context
Jan 13, 2026

FlashAttention: Fast and Memory Efficient Exact Attention with IO Awareness

FlashAttention 的核心贡献是把 exact softmax attention 的瓶颈从 FLOPs 视角重新定位到 GPU memory hierarchy 和 HBM 读写上:它用 tiling 在 SRAM 中分块计算 attention,并用 online softmax 统计量与 backward recomputation 避免物化 $N\times N$ attention matrix,从而保持 exac...

Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré

2205.14135-flashattention-io-aware-exact-attention Attention KernelLong Context
Jan 09, 2026

Training Compute Optimal Large Language Models

这篇论文用 400 多个不同规模和 token 预算的 Transformer 训练 run 重新估计 compute optimal pretraining frontier,结论直接修正 Kaplan scaling laws:在固定训练 FLOPs 下,最优 dense LM 应该让模型参数量 $N$ 和训练 token 数 $D$ 近似等比例增长,即 $N {\mathrm{opt}}\propto C^{0.5}$、$D {...

Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, +12 more

2203.15556-training-compute-optimal-large-language-models Scaling Laws
Jan 06, 2026

Scaling Laws for Neural Language Models

这篇论文把语言模型训练从“单次大模型实验”推进到可拟合的经验规律:cross entropy loss 随模型参数量 $N$、训练数据量 $D$ 和训练计算量 $C$ 呈稳定 power law;在当时的实验范围内,模型 shape 和许多超参数影响较弱,更大的模型具有更高 sample efficiency;论文据此推出 compute efficient training 应优先扩大模型、减少训练步数,并在远未完全收敛时停止。现代...

Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, Dario Amodei

2001.08361-scaling-laws-neural-language-models Scaling Laws