392 tracked profiles
9 recurring authors
104 archived notes

Tracked Profiles

这些作者已经有人类核验的主页、账号或机构信息。

6 notes Wenfeng Liang (梁文锋)

DeepSeek AI / High-Flyer

DeepSeekfoundation modelsefficient language modelsreasoning models
5 notes Damai Dai

DeepSeek AI / Peking University

DeepSeekMoEDeepSeek-V3mixture-of-expertslarge language models
5 notes Runxin Xu (许润昕)

DeepSeek AI / Peking University

DeepSeek-R1DeepSeek-V3DeepSeekMathDeepSeek Coder
5 notes Yu Wu (吴俣)

DeepSeek AI / Microsoft Research Asia

DeepSeek-R1GRPOLLM alignmentpost-training
4 notes Haibin Lin

ByteDance Seed

VERLHybridFlowLaminarasynchronous RL
4 notes Xingkai Yu (俞星凯)

DeepSeek AI / Nanjing University

DeepSeek-R1DeepSeekMoEEngramDSpark
4 notes Yuxuan Tong (童雨轩)

Tsinghua University / ByteDance Seed

verlDAPOLaminarasynchronous RL
3 notes Binyuan Hui

Qwen Team / Alibaba Group

QwenQwen-Codercode LLMsreasoning
3 notes Bo Zheng

Qwen Team / Alibaba Group

Qwenlarge language modelsfoundation modelspost-training
3 notes Chi Zhang

ByteDance / University of Southern California

VERLHybridFlowLaminarasynchronous RL
3 notes Deli Chen (陈德里)

DeepSeek AI / Peking University

DeepSeek seriesreasoning RLstep-by-step verificationLLM safety
3 notes Guangming Sheng

The University of Hong Kong / ByteDance Seed

VERLHybridFlowLaminarasynchronous RL
3 notes Jie Tang (唐杰)

Tsinghua University / Knowledge Engineering Group (KEG)

large language modelsagentic reinforcement learningcontext compactionIndexCache
3 notes Junyang Lin

Qwen Team / Alibaba Group

Qwenlarge language modelsmultimodal systemsagentic learning
3 notes Ning Ding (丁宁)

Tsinghua University / Shanghai AI Laboratory

reasoning intelligencescalable reinforcement learningPRIME-RLRLVR
3 notes Peiyi Wang

DeepSeek AI / Peking University

Math-ShepherdDeepSeekMathDeepSeek-R1reasoning RL
3 notes Wang Zhang

ByteDance

VERLHybridFlowLaminarasynchronous RL
3 notes Wangding Zeng

DeepSeek AI / Beijing University of Posts and Telecommunications

MLADeepSeek-V2DeepSeekMoEEngram
3 notes Xin Liu

ByteDance Seed / ByteDance

LaminarDAPOMegaScale-MoELLM systems
3 notes Yanghua Peng

ByteDance Seed / ByteDance

Seed InfraMegaScaleMegaScale-MoEHybridFlow
3 notes Zhifang Sui

Peking University

natural language processingcomputational linguisticsMath-ShepherdDeepSeekMoE
3 notes Zhihong Shao (邵智宏)

DeepSeek AI / Tsinghua University

DeepSeekMathDeepSeek-R1DeepSeek-ProverToRA
2 notes Aditi Raghunathan

Carnegie Mellon University / AI Reliability at CMU

AI safetyrobust machine learningverificationreward hacking
2 notes An Yang

Qwen Team / Alibaba Group

Qwenlarge language modelsfoundation modelspost-training
2 notes Baoxiang Wang (王宝祥)

The Chinese University of Hong Kong, Shenzhen / Vector Institute

reinforcement learninggame theorymachine learningtrust region masking
2 notes Bo Pang

Moonshot AI

Moonshot AIKimiKimi K2.5synchronous RL rollout
2 notes Bowen Baker

OpenAI / MIT

multi-agent reinforcement learningchain-of-thought monitoringreasoning model safetyprocess supervision
2 notes Bowen Yu (郁博文)

Qwen Team / Alibaba Group

Qwenpost-trainingautomated alignmentinstruction tuning
2 notes Bowen Zhou

Tsinghua University / Shanghai AI Laboratory

reasoning intelligencenatural language processingmultimodal AItrustworthy AI
2 notes Chao Jin

Peking University / Key Laboratory of High Confidence Software Technologies, Peking University

MegaScale-MoEUltraEPMoE systemsLLM training systems
2 notes Chelsea Finn

Stanford University / Physical Intelligence

SPIRALLLM-as-a-Verifierrobot learningmeta-learning
2 notes Chuan Wu

The University of Hong Kong / University of Toronto

LaminarHybridFlowdistributed systemscloud computing
2 notes Dayiheng Liu (刘大一恒)

Qwen Team / Alibaba Group

Qwenlarge language modelstext generationnatural language processing
2 notes Dongyan Zhao (赵东岩)

Peking University / Wangxuan Institute of Computer Technology

natural language processingsemantic data managementPeking University NLPEngram
2 notes Fan Zhou

Qwen Team / Alibaba Group

Qwen3-Coder-NextQwen3Qwen-Coderlarge language models
2 notes Fei Huang

Qwen Team / Alibaba Group

QwenTongyi Lablarge language modelsmultimodal models
2 notes Fuli Luo (罗福莉)

Xiaomi LLM Core / Xiaomi MiMo

MOPDMiMoDeepSeek-R1DeepSeekMoE
2 notes Furu Wei (韦福如)

Microsoft Research / Microsoft Research Asia

LLM-in-SandboxReOPDgeneral artificial intelligencelarge foundation models
2 notes Ganqu Cui

Shanghai AI Laboratory / Tsinghua University / THUNLP

RLVRLLM alignmentreinforcement learningPRIME-RL
2 notes Haitao Mi

Tencent / Tencent HY Team

FlashMemoryDeepSeek-V4Lookahead Sparse AttentionHiLS-Attention
2 notes Huazuo Gao

DeepSeek AI

DeepSeek-V2MLADeepSeekMoEefficient language models
2 notes Huishuai Zhang

Peking University / Microsoft Research Asia

large language modelsoptimizationdifferential privacyEngram
2 notes Jan Kautz

NVIDIA

NVIDIA Researchmachine learningcomputer visionProRL
2 notes Ji-Rong Wen

Gaoling School of Artificial Intelligence, Renmin University of China

LLM-in-Sandboxlong-horizon agentslarge language modelsweb search
2 notes Jiacai Liu (刘佳材)

Fudan University

reinforcement learningLLM post-trainingreasoning modelsRLVR
2 notes Jian Hu

NVIDIA / National Taiwan University

ProRLBroRLOpenRLHFRLHF systems
2 notes Jianwei Zhang

Qwen Team / Alibaba Group

Qwenlarge language modelsfoundation models
2 notes Jianxin Yang

Qwen Team / Alibaba Group

Qwenlarge language modelsMTPRL rollout acceleration
2 notes Jiashi Li

DeepSeek AI / Peking University

DSparkDeepSeekMoEspeculative decodingLLM systems
2 notes Jiawei Xu

The Chinese University of Hong Kong, Shenzhen

reinforcement learningLLM reinforcement learningtrust region maskingoptimal token baseline
2 notes Jingren Zhou

Alibaba Group / Alibaba Cloud

QwenAlibaba Cloudfoundation modelsAI infrastructure
2 notes John Schulman

Thinking Machines / Anthropic (former)

PPOTRPORLHFpost-training
2 notes Lei Bai (白磊)

Shanghai Artificial Intelligence Laboratory / Fudan University

AI for scienceautonomous scientific discoveryfoundation modelsspatiotemporal learning
2 notes Lei Li

The University of Hong Kong / Peking University

multimodal LLMsin-context learningLLM-as-a-judgeMath-Shepherd
2 notes Lei Zhu (祝磊)

Tencent HY Team / Tencent Hunyuan Frontier Lab

HiLS-Attentionnative sparse attentionlong-context modelingsparse attention
2 notes Li Dong (董力)

Microsoft Research / Microsoft Research Asia

LLM-in-SandboxReOPDlarge language modelsmachine intelligence
2 notes Lifan Yuan (袁立凡)

University of Illinois Urbana-Champaign / Tsinghua University / THUNLP

RLVRPRIME-RLprocess rewardscalable interaction
2 notes Mingjie Liu

NVIDIA / University of Texas at Austin

ProRLBroRLreinforcement learningreasoning models
2 notes Qian Liu (刘乾)

TikTok AI Innovation Center / xAI

natural language processingcode intelligenceLLM agentsreinforcement learning
2 notes Rui Men

Qwen Team / Alibaba Group

Qwenpretraininglarge language modelsnatural language processing
2 notes Ruoyu Qin (秦若愚)

Moonshot AI / Tsinghua University

LLM servingsynchronous RL rolloutKVCacheMooncake
2 notes Samyam Rajbhandari

Snowflake AI Research / Microsoft (DeepSpeed / ZeRO papers)

DeepSpeedZeRODeepSpeed Ulyssesdistributed training
2 notes Shizhe Diao

Thinking Machines Lab / NVIDIA Research

ProRLBroRLpost-trainingefficient language models
2 notes Tian Liang

Tencent / Tencent HY Team

FlashMemoryDeepSeek-V4Lookahead Sparse AttentionHiLS-Attention
2 notes Tri Dao

Princeton University / Together AI

FlashAttentionFlashAttention-2efficient attentionGPU kernels
2 notes Weiran He

Moonshot AI / MADSys Lab

LLM servingKVCacheMooncakeSeer
2 notes Weixiao Huang

Moonshot AI / Tsinghua University

cloud-native systemsAI infrastructureLLM servingSeer
2 notes Xiang Hu

Tencent / Tencent HY Team

FlashMemoryDeepSeek-V4Lookahead Sparse AttentionHiLS-Attention
2 notes Xiang Li

ByteDance Seed / ByteDance

LaminarMegaScaleMegaScale-MoELLM systems
2 notes Xibin Wu

ByteDance / ByteDance Seed

LaminarHybridFlowOpenRLHFverl
2 notes Ximing Lu

NVIDIA Research / University of Washington

ProRLBroRLreinforcement learningreasoning models
2 notes Xin Cheng

Peking University / DeepSeek AI

EngramDSparkspeculative decodingconditional memory
2 notes Xinran Xu (许欣然)

Moonshot AI

Moonshot AIKimiKimi K2.5LLM infrastructure
2 notes Yan Wang (王琰)

Independent Researchers / Tencent

FlashMemoryDeepSeek-V4Lookahead Sparse AttentionHiLS-Attention
2 notes Yang Xu (徐旸)

MiniMax / Zhejiang University

MiniMax Sparse AttentionMiniMax-M3model architectureefficient attention
2 notes Yangkun Zhang

Moonshot AI

Moonshot AIKimiKimi K2.5synchronous RL rollout
2 notes Yejin Choi

Stanford University / NVIDIA

language modelsreasoningcommonsense AIProRL
2 notes Yi Dong

NVIDIA

NVIDIA Researchreasoning modelsvirtual agentsProRL
2 notes Yikai Zhao (赵一开)

Moonshot AI / Tsinghua University

LLM servingdisaggregated systemsKimiSeer
2 notes Yingru Li (李英儒)

xAI / The Chinese University of Hong Kong

LLM reinforcement learningrollout correctiontrust region maskingoptimal token baseline
2 notes Yuchen Zhang

Peking University / Shanghai AI Laboratory

RLVRreasoning modelsagentsscalable reinforcement learning
2 notes Yujiang Li

Tsinghua University / Knowledge Engineering Group (KEG)

agentic reinforcement learningLLM post-trainingcoding agentsreasoning
2 notes Yushi Bai (白雨石)

Tencent HY Team / Tsinghua University

HiLS-AttentionIndexCacheIndexShareGLM-5
2 notes Yuxiao Dong

Tsinghua University / Knowledge Engineering Group (KEG)

large language modelsagentic reinforcement learningcontext compactionagents
2 notes Yuxiong He

Snowflake AI Research / Microsoft (DeepSpeed / ZeRO papers)

DeepSpeedZeRODeepSpeed UlyssesAI systems
2 notes Zhenda Xie (解振达)

DeepSeek AI / Tsinghua University

foundation modelspretrainingself-supervised learningDeepSeek
2 notes Zhenyu Hou

Tsinghua University / Knowledge Engineering Group (KEG)

agentic reinforcement learningLLM post-trainingreasoningagents
2 notes Ziniu Li (李子牛)

Tencent Hunyuan / The Chinese University of Hong Kong, Shenzhen

large language modelsreinforcement learningpost-trainingtrust region masking
1 notes Akshay V. Jagadeesh

OpenAI / Stanford University

AI alignmentbeneficial AIhealth AIcomputational neuroscience
1 notes Akshayaa Magesh

Meta AI / University of Illinois Urbana-Champaign

reinforcement learningrobust machine learningstatistical inferenceLLM reasoning
1 notes Alexander Rakhlin

Massachusetts Institute of Technology / University of Pennsylvania

reinforcement learning theoryonline learningstatistical learningoptimization
1 notes Amey Agrawal

Georgia Institute of Technology / Microsoft Research India (Sarathi internship)

LLM servingSystems for AISarathiSarathi-Serve
1 notes Andrea Zanette

Carnegie Mellon University / UC Berkeley

maximum likelihood reinforcement learningfoundation modelsreasoningalignment
1 notes Ankur Samanta

Columbia University / Meta AI

LLM reasoningRLVRself-improvementmultimodal foundation models
1 notes Aohan Zeng

Z.ai / Tsinghua University

IndexCacheGLM-5GLMlarge language models
1 notes Aoqi Hu

Baidu Inc.

ECHOagentic RLLLM agentsBaidu
1 notes Aosong Feng

Yale University / AWS Bedrock

agentic AIlong-horizon human-AI interactionLLM routingRLVR
1 notes Ashish Panwar

Microsoft Research India / Indian Institute of Science (PhD)

LLM servingsystemsmemory managementSarathi
1 notes Atri Rudra

University at Buffalo, SUNY

theoryalgorithmsIO lower boundsFlashAttention
1 notes Ayush Jain

Meta AI / Meta Applied Reinforcement Learning

reinforcement learningLLM reasoningagentsrobotics
1 notes Azalia Mirhoseini

Stanford University / Ricursive Intelligence

LLM-as-a-Verifierself-improving AItest-time scalingAI systems
1 notes Baohao Liao

Microsoft Research / University of Amsterdam

ReOPDLLM efficiencyreasoningagents
1 notes Baolin Peng

Microsoft Research

large language model agentscomplex reasoningplanningreinforcement learning
1 notes Beidi Chen

Carnegie Mellon University / Amazon Scholar

ThunderAgentagentic inferencelong-context efficiencyLLM serving
1 notes Bhargav S. Gulavani

Microsoft Research India / Microsoft

LLM servingsystemsSarathiSarathi-Serve
1 notes Bhavana Dalvi Mishra

Google / Allen Institute for AI

self-evolving agentsinteractive reasoningAI for scientific discoverynatural language processing
1 notes Binbin Zheng

Baidu Inc. / University of Science and Technology of China

ECHOagentic RLLLM agentslong-horizon tool use
1 notes Bo Zhang (张铂)

Shanghai Artificial Intelligence Laboratory

general AI agentsmultimodal reasoningautonomous scientific discoveryself-evolving LLMs
1 notes Borui Wan

The University of Hong Kong / ByteDance Seed

LaminarByteCheckpointmachine learning systemscheckpointing systems
1 notes Chang Chen

University of California, San Diego / PICASSO Lab

machine learning systemsdistributed trainingcommunication librariesLLM serving
1 notes Chaobo Jia (贾超博)

The Chinese University of Hong Kong / ByteDance Seed

Laminarasynchronous RLagent swarmlong-horizon agents
1 notes Che Jiang

Tsinghua University / Horizon Research

self-improving agentsagent harnessruntime adaptationagentic AI
1 notes Chen Henry Wu

Carnegie Mellon University / Tsinghua University (undergraduate)

language agentsverificationreward shapingAI safety
1 notes Chen Xu

Carnegie Mellon University / Renmin University of China

information retrievalfairness in recommender systemsAI and economicsLLM multi-agent systems
1 notes Chenchen Zhang

Independent Researcher

LLM reinforcement learningcredit assignmentagentic RLmulti-agent systems
1 notes Chenfeng Xu

Together AI / The University of Texas at Austin

ThunderAgentefficient generative AIML systemsrobotics
1 notes Cheng Qian

University of Illinois Urbana-Champaign / Tsinghua University

tool-calling RLlanguage agentsknowledge-intensive NLPreinforcement learning
1 notes Chenghao Zhang

Gaoling School of Artificial Intelligence, Renmin University of China

long-horizon agentsretrieval-augmented generationdeep research agentsmultimodal reasoning
1 notes Chenze Shao

DeepSeek AI / Tencent WeChat AI

DSparkspeculative decodingnon-autoregressive generationnatural language generation
1 notes Christof Monz

University of Amsterdam

ReOPDnatural language processingmachine translationinformation retrieval
1 notes Christopher Ré

Stanford University / Hazy Research

Hazy Researchdata-centric AILLM systemsFlashAttention
1 notes Chuanyang Jin

Johns Hopkins University / Social Cognitive AI Lab

social cognitive AITheory of Mindhuman interactionLLM agents
1 notes Chunyang Li

Tencent

FlashMemoryDeepSeek-V4Lookahead Sparse Attention
1 notes Daixuan Cheng (成岱璇)

Gaoling School of Artificial Intelligence, Renmin University of China / Microsoft Research

LLM-in-Sandboxagentic LLM trainingcomputer-use agentsreinforcement learning
1 notes Dakai An

HKUST

RL post-training systemsAI infrastructuredistributed systems
1 notes Daman Arora

Carnegie Mellon University / Microsoft Research India

maximum likelihood reinforcement learningreasoningcontinual learningsoftware agents
1 notes Daniel Khashabi

Johns Hopkins University / Center for Language and Speech Processing

natural language processinglanguage model reasoninglanguage model agentslong-context systems
1 notes Daniel R. Jiang

Meta AI / University of Pittsburgh

reinforcement learningRLHFlong-horizon agentsLLM post-training
1 notes Daniel Y. Fu

UC San Diego / Together AI

FlashAttentionefficient sequence modelingLLM systems
1 notes Daniele Dan

Amazon / University of Padua

tool-calling RLapplied machine learningnatural language processinglanguage agents
1 notes Dawn Song

UC Berkeley / BAIR

AI safetyAI securityagentic AIdeep learning
1 notes Dilxat Muhtar

Alibaba Group

agentic RL systemsRL post-trainingROLL
1 notes Dong Yu

Tencent / Tencent AI Lab

FlashMemoryDeepSeek-V4Lookahead Sparse Attentionspeech recognition
1 notes Dongyang Ma

Independent Researchers

FlashMemoryDeepSeek-V4Lookahead Sparse AttentionLLM agents
1 notes Dorsa Sadigh

Stanford University / Google DeepMind

SPIRALset reinforcement learningrobot learninghuman-robot interaction
1 notes Enlei Gong

Baidu Inc.

ECHOagentic RLLLM agentsBaidu
1 notes Eric Nalisnick

Johns Hopkins University

probabilistic machine learninguncertainty quantificationsafe and robust AIhuman-AI collaboration
1 notes Fahim Tajwar

Carnegie Mellon University / Stanford University

maximum likelihood reinforcement learningRLVRpost-trainingreasoning
1 notes Fangyun Wei

Microsoft Research Asia / Peking University

EAGLEspeculative decodingcomputer visionmultimodal learning
1 notes Flood Sung

XVI Robotics / Moonshot AI

Kimireinforcement learningagentic systemsrobotics
1 notes Gabriella Kazai

Amazon / Microsoft

tool-calling RLinformation retrievalhuman computationlanguage agents
1 notes Gang Li

Texas A&M University

reinforcement learningefficient reasoningreasoning language modelsoptimization
1 notes Gaotang Li

University of Illinois Urbana-Champaign / Google Cloud AI Research (RubricEM internship)

LLM post-trainingagentic reasoningreward modelingfoundation models
1 notes Ge Zhang

M-A-P / ByteDance Seed

natural language processingmultimodal intelligenceLLM reinforcement learning
1 notes Gennady Pekhimenko

University of Toronto / Vector Institute

LoRAFusionCentMLEcoSystemmachine learning systems
1 notes Guanning Zeng

Tsinghua University / Carnegie Mellon University

maximum likelihood reinforcement learningRLVRrepresentation learningreinforcement learning
1 notes Guanqun Zhao

Baidu Inc.

ECHOagentic RLLLM agentsBaidu
1 notes Guanting Dong (董冠霆)

Gaoling School of Artificial Intelligence, Renmin University of China / Beijing University of Posts and Telecommunications

long-horizon agentsagentic reinforcement learningdeep search agentsLLM alignment
1 notes Guoxin Chen

Gaoling School of Artificial Intelligence, Renmin University of China / Institute of Computing Technology, Chinese Academy of Sciences

LLM-in-Sandboxcoding agentsweb research agentslarge language models
1 notes Haichao Zhu

MiniMax

MiniMax Sparse AttentionMiniMax-M3MiniMax-M1MiniMax-M2
1 notes Hailin Zhang (张海林)

Xiaomi MiMo / Peking University

MiMoMOPDRL infrastructureMLSys
1 notes Hailong Yang

Beihang University / University of Michigan

Grapehigh-performance computingdeep learning systemscompilation
1 notes Haiwen Feng

UC Berkeley / Impossible, Inc.

maximum likelihood reinforcement learningmultimodal foundation modelsworld modelsinverse graphics
1 notes Haizhou Zhao

Alibaba Group

ROLLasynchronous RLagentic RL systemsRL post-training
1 notes Hanze Dong

Microsoft Research / Hong Kong University of Science and Technology

ReOPDfoundation model post-trainingalignmentgenerative modeling
1 notes Hao Cheng

Microsoft Research / University of Washington

large language model agentsconversational AItool-using agentsnatural language processing
1 notes Hao Gu

The Hong Kong University of Science and Technology

HiLS-Attentionnative sparse attentionlong-context modelinglarge language models
1 notes Hao Kang

Georgia Institute of Technology / MIT

ThunderAgentagentic inferenceML systemsefficient LLM systems
1 notes Hao Liu

Google DeepMind / Carnegie Mellon University

Ring AttentionBlockwise Parallel Transformerlong-context trainingreinforcement learning
1 notes Haobo Wang (王皓波)

Zhejiang University

machine learninglarge language modelsweakly supervised learningrubric-based reinforcement learning
1 notes Haohai Sun

MiniMax

MiniMax Sparse AttentionMiniMax-M3MiniMax-M1MiniMax-01
1 notes Harri Edwards

OpenAI (paper affiliation) / Google DeepMind

process supervisionreinforcement learningrandom network distillationgrokking
1 notes Huajun Bai

Tsinghua University / Tsinghua Storage Research Group

SPORKSystems for AIefficient LLM inferenceagentic inference
1 notes Huatong Song (宋华彤)

Gaoling School of Artificial Intelligence, Renmin University of China / IQuest Research

LLM-in-Sandboxagentic LLM trainingcomputer-use agentssoftware engineering agents
1 notes Huayang Li

Tencent HY Team / Tencent AI Lab

HiLS-Attentionnative sparse attentionlong-context modelingretrieval-augmented generation
1 notes Huayu Chen (陈华玉)

Tsinghua University / NVIDIA Deep Imagination Research group (internship)

reinforcement learningdeep generative modelsTianshouLLM/VLM RL
1 notes Huichuan Zheng

Tsinghua University / Tsinghua Storage Research Group

SPORKstorage systemsLLM trainingagentic inference
1 notes Hunter Lightman

OpenAI

process supervisionmathematical reasoningLLM reliabilityreasoning models
1 notes Idan Shenfeld

Massachusetts Institute of Technology / Improbable AI Lab

continual learningself-distillationreinforcement learninglanguage model post-training
1 notes Ifdita Hasan Orney

Stanford University

SPIRALset reinforcement learningreinforcement learninghuman-centered AI
1 notes Ilya Sutskever

Safe Superintelligence Inc. / OpenAI (cofounder, former chief scientist)

deep learningscalingOpenAIsafe superintelligence
1 notes Ion Stoica

University of California, Berkeley / Sky Computing Lab

LLM-as-a-VerifierAI systemsdistributed systemscloud computing
1 notes Jacky Kwok

Stanford University / University of California, Berkeley

LLM-as-a-Verifierverificationtest-time scalingrobot learning
1 notes Jalaj Bhandari

Meta AI / Columbia University

reinforcement learningoptimization theorypolicy gradienttemporal difference learning
1 notes Jan Leike

Anthropic / OpenAI (former)

AI alignmentRLHFscalable oversightsuperalignment
1 notes Jayashree Mohan

Microsoft Research India / University of Texas at Austin (PhD)

LLM servingsystemsstorage systemsSarathi
1 notes Jeff Schneider

Carnegie Mellon University / Robotics Institute

maximum likelihood reinforcement learningreinforcement learningdecision making with large language modelsself-driving cars
1 notes Jia Li

The Hong Kong University of Science and Technology (Guangzhou) / The Chinese University of Hong Kong

FlashMemoryDeepSeek-V4Lookahead Sparse Attentiondata science
1 notes Jiachen Yu

Tencent / Tsinghua University

FlashMemoryDeepSeek-V4Lookahead Sparse AttentionLLM reinforcement learning
1 notes Jiacheng Chen

The Chinese University of Hong Kong / Shanghai AI Laboratory

RLVRreasoning in NLPreinforcement learningMiniMax
1 notes Jiajie Jin (金佳杰)

Gaoling School of Artificial Intelligence, Renmin University of China

long-horizon agentsdeep research agentsknowledge utilizationverifiable agents
1 notes Jiajun Zhang

Qwen Team / Alibaba Group

Qwen3-Coder-NextQwen-Codercoding agentscompetitive coding
1 notes Jiamang Wang

Alibaba Group

machine learning infrastructuredistributed systemsROLLRL post-training
1 notes Jian Chen

UC San Diego / Z Lab

DFlashMagicDecspeculative decodingefficient LLM inference
1 notes Jianfeng Gao

Microsoft Research / Microsoft

agentic AItool useplanningmultimodal reasoning
1 notes Jiawei Chen (陈家慰)

Qwen Team / Alibaba Group

Qwen3-Coder-Nextcode LLMsagentic RLretrieval-augmented generation
1 notes Jiaxi Yang

Qwen Team / Alibaba Group

Qwen3-Coder-NextQwen3Qwen-Coderlarge language models
1 notes Jiayao Li

MiniMax

MiniMax Sparse AttentionMiniMax-M3MiniMax-M2inference kernels
1 notes Jiayao Tang

School of Mathematical Sciences, Peking University

ECHOagentic RLLLM agentslong-horizon tool use
1 notes Jieping Ye

Alibaba Group / Alibaba Cloud Data to Intelligence Lab

HydraHeadmachine learningdata miningartificial intelligence
1 notes Jihua Liu

Baidu Inc.

ECHOagentic RLLLM agentsBaidu
1 notes Jingkai Zhou

MiniMax / Hangzhou Dianzi University

MiniMax Sparse AttentionMiniMax-M3inference kernels
1 notes Jingyu Zhang (张景昱)

Johns Hopkins University / Apple AIML (research internship, 2026)

foundation model post-trainingAI alignmentLLM safetymulti-agent reinforcement learning
1 notes Jinkai Hu

MiniMax

MiniMax Sparse AttentionMiniMax-M3inference kernels
1 notes Jinxi Wei

Qwen Team / Alibaba Group

Qwen3-Coder-NextQwen-Codercoding agentstool calling
1 notes Jinyang Wu

Tsinghua University / Xi'an Jiaotong University

agent reinforcement learningon-policy skill distillationhindsight skillsmultimodal reasoning
1 notes Jiwu Shu

Tsinghua University / Tsinghua Storage Research Group

SPORKstorage systemsnonvolatile memorydistributed systems
1 notes Jonas Hübotter

ETH Zürich / Stanford University

self-distillationtest-time trainingreinforcement learningfoundation model specialization
1 notes Ju Huang

Alibaba Group

ROLLasynchronous RLrollout schedulingRL post-training
1 notes Juanzi Li (李涓子)

Tsinghua University / Knowledge Engineering Group (KEG)

IndexCachelarge language modelsknowledge graphsnatural language processing
1 notes Jubayer Ibn Hamid

Stanford University / Google

SPIRALset reinforcement learninginference computereinforcement learning
1 notes Junxiong Wang

Together AI / Cornell University

ThunderAgentTogether AIefficient RL rolloutsadaptive speculative decoding
1 notes Kaixin Li

Qwen Team / Alibaba Group

Qwen3-Coder-NextQwen-Codermultimodal agentsGUI agents
1 notes Kaiyan Zhang (张开颜)

Frontis.AI / Tsinghua University

self-improving agentsself-evolving agentsscalable reinforcement learningcollective intelligence
1 notes Karan Singhal

OpenAI / Google Research

health AIAGI benefitsAI safetymedical AI
1 notes Karl Cobbe

OpenAI / Stanford University

mathematical reasoningverifiersGSM8KPRM800K
1 notes Kashun Shum

Qwen Team / Alibaba Group

Qwen3-Coder-NextQwen-Codercoding agentslarge language models
1 notes Kaveh Hassani

Meta Superintelligence Labs / Meta

self-improving LLMsLLM reasoningmulti-agent reinforcement learningpost-training
1 notes Kavosh Asadi

Meta AI / Amazon

reinforcement learningvalue function optimizationRL theoryassistive AI agents
1 notes Kevin Song

University of Toronto / Vector Institute

LoRAFusioncomputer systemsmachine learning systemsGPU systems
1 notes Kewei Tu (屠可伟)

ShanghaiTech University

HiLS-Attentionhierarchical sparse attentionlong-context modelingnatural language processing
1 notes Khaled Saab

OpenAI / Google DeepMind

biomedical intelligencehealth AImedical AIAI alignment
1 notes Lei Feng (冯磊)

Southeast University / RIKEN

machine learningAI agentsAI safetyrobust learning
1 notes Lei Zhang

Qwen Team / Alibaba Group

Qwen3-Coder-NextQwen-Codersoftware engineering agentsexecutable environments
1 notes Leitian Tao

University of Wisconsin–Madison / Microsoft Research

large language model post-trainingAI agentsreinforcement learningreward modeling
1 notes Leo Liang

Tencent HY Team

HiLS-Attentionnative sparse attentionlong-context modeling
1 notes Liang Zhao

Xiaomi MiMo / Xiaomi LLM Core

MOPDMiMoLLM post-trainingMoE reinforcement learning
1 notes Lin Qu

Alibaba Group

distributed trainingmachine learning infrastructurerecommendation systemsROLL
1 notes Lingfeng Liu

School of Mathematical Sciences, Peking University

ECHOagentic RLLLM agentslong-horizon tool use
1 notes Longtao Zheng (郑龙韬)

ByteDance / Nanyang Technological University

LLM reinforcement learningcoding agentsmulti-agent reinforcement learningopen-ended agents
1 notes Lunbin Zeng

MiniMax / Huazhong University of Science and Technology

MiniMax Sparse AttentionMiniMax-M3MiniMax-M1MLLM
1 notes Lunxi Cao

HKUST

RL post-training systemsdistributed systemsagent sandbox runtime
1 notes Marco Pavone

Stanford University / NVIDIA Research

LLM-as-a-Verifierautonomous systemsroboticsautonomous driving
1 notes Matan Kalman

Google Research / Google

speculative decodingtransformer efficiencyimage editingLLM inference
1 notes Matei Zaharia

University of California, Berkeley / Sky Lab

distributed systemsdata systemsAI systemsRing Attention
1 notes Matej Cief

Amazon / Brno University of Technology

tool-calling RLnatural language processingmachine learninglanguage agents
1 notes Matthieu Geist

Google DeepMind / Google Brain France

reinforcement learningsequential decision makingon-policy distillationlanguage model alignment
1 notes Mehrdad Farajtabar

Apple

LLM reasoningplanningagentic LLMsinference efficiency
1 notes Mehul Damani

Massachusetts Institute of Technology / MIT-IBM Watson AI Lab

reinforcement learninglarge language modelsuncertainty calibrationtool use
1 notes Miao Peng

Tencent / The Hong Kong University of Science and Technology (Guangzhou)

FlashMemoryDeepSeek-V4Lookahead Sparse Attentiondata science
1 notes Michael Y. Li

Stanford University / Princeton University

SPIRALinference computelong-context memorystatistical machine learning
1 notes Mike Hang Wang

Microsoft Research

reinforcement learningAI agentsworld modelsmulti-agent reinforcement learning
1 notes Ming Lin

Oracle Cloud Infrastructure

reinforcement learningreasoning language modelsoptimizationgenerative AI
1 notes Mingxing Zhang (章明星)

Tsinghua University / MADSys Lab

memory systemsdistributed systemsKVCacheMooncake
1 notes Mingxuan Xia

Zhejiang University / ByteDance

large language modelsweak supervisionrubric-based reinforcement learningdata annotation
1 notes Mingze Li

Qwen Team / Alibaba Group

Qwen3-Coder-NextQwen3Qwen-Codersoftware engineering agents
1 notes Minshen Zhang

University of California, San Diego / ShanghaiTech University

HiLS-Attentionnative sparse attentionlong-context modelingmachine learning systems
1 notes Mouxiang Chen (陈谋祥)

Qwen Team / Alibaba Group

Qwen3-Coder-NextQwen-Coderlarge language modelscode generation
1 notes Nikolaos Aletras

University of Sheffield / Amazon

tool-calling RLnatural language processingcomputational social sciencelanguage model evaluation
1 notes Nino Vieillard

Google DeepMind / Google

on-policy distillationgeneralized knowledge distillationreinforcement learningRLHF
1 notes Nipun Kwatra

Microsoft Research India / Stanford University (PhD)

LLM servingdeep learning systemsAI infrastructureSarathi
1 notes Noah Goodman

Stanford University / MIT

SPIRALprobabilistic cognitionreinforcement learninglanguage model reasoning
1 notes Nuo Chen

Tencent / The Hong Kong University of Science and Technology (Guangzhou)

FlashMemoryDeepSeek-V4Lookahead Sparse AttentionLLM agents
1 notes Olivier Bachem

Google DeepMind / Google Brain

on-policy distillationGemmaopen-weight multimodal modelsreinforcement learning
1 notes Omar Shaikh

Stanford University / Georgia Institute of Technology

SPIRALhuman-AI groundingHCINLP
1 notes Paul Sajda

Columbia University / LIINC Lab

brain-computer interfacesneuroengineeringmachine learningdecision-making
1 notes Pengyu Zhao

MiniMax

MiniMax Sparse AttentionMiniMax-M3MiniMax-M2MiniMax-M1
1 notes Pieter Abbeel

University of California, Berkeley / Berkeley Robot Learning Lab

robot learningreinforcement learningAIRing Attention
1 notes Piotr Stanczyk

Google DeepMind / Google

on-policy distillationmachine learning infrastructurelanguage model distillation
1 notes Pranav Atreya

University of California, Berkeley / Berkeley Robot Learning Lab

LLM-as-a-Verifierrobot learninggeneralist policiesrobot evaluation
1 notes Pulkit Agrawal

Massachusetts Institute of Technology / MIT Computer Science and Artificial Intelligence Laboratory

reinforcement learningrobot learningcontinual learningself-supervised learning
1 notes Qian Dong

Tsinghua University / THUIR

IndexCacheGLM-5sparse attentionlong-context modeling
1 notes Qianxu Wang

University of California, San Diego / Peking University

machine learning systemshardware-software co-designAI accelerationLLM serving
1 notes Qiaorui Chen

NVIDIA

MiniMax Sparse AttentionGPU kernelsCUDAsparse attention
1 notes Qidong Su

University of Toronto / Vector Institute

LoRAFusionmachine learning systemscompilersdistributed systems
1 notes Qifan Zhang

Tencent / The Hong Kong University of Science and Technology (Guangzhou)

FlashMemoryDeepSeek-V4Lookahead Sparse AttentionLLM agents
1 notes Rahul K. Arora

OpenAI / University of Calgary

health AIreinforcement learningLLM evaluationclinical AI
1 notes Ramachandran Ramjee

Microsoft Research India / University of Massachusetts Amherst (PhD)

AI infrastructureLLM servingsystemsSarathi
1 notes Rishabh Agarwal

Google DeepMind / Mila

on-policy distillationgeneralized knowledge distillationreinforcement learningLLM distillation
1 notes Rui Gao

MiniMax / Nanjing University

MiniMax Sparse AttentionMiniMax-M3inference kernels
1 notes Ruisheng Cao

Qwen Team / Alibaba Group

Qwen3-Coder-Nextcoding agentssoftware engineering agentstext-to-SQL
1 notes Ruslan Salakhutdinov

Carnegie Mellon University / Machine Learning Department

maximum likelihood reinforcement learningdeep learningprobabilistic graphical modelslarge-scale optimization
1 notes Sabela Ramos

Google DeepMind / Google Research

on-policy distillationreinforcement learning infrastructurelanguage model distillationhigh-performance computing
1 notes Sam Ade Jacobs

Microsoft Research / Texas A&M University (Ph.D.)

DeepSpeedDeepSpeed UlyssesHPCAI systems
1 notes Sergey Levine

UC Berkeley EECS / RAIL Lab

reinforcement learningrobot learningdecision makingpolicy optimization
1 notes Shang Wang

University of Toronto / Vector Institute

LoRAFusionCentMLmachine learning systemsGPU systems
1 notes Shaohan Huang

Microsoft Research / Microsoft Research Asia

LLM-in-Sandboxlarge language modelsmultimodal large language modelsmodel architecture
1 notes Shaopan Xiong

Alibaba Group

RollArtRollPackerROLL Flashagentic RL systems
1 notes Sharon Li

University of Wisconsin–Madison

reliable AIagentic language-model systemsreinforcement learningreward modeling
1 notes Shen Yan

ByteDance Seed / ByteDance

Dynamic Linear Attentionmultimodal pretrainingvideo-text modelingvision-language models
1 notes Shulu Li

University of California, Berkeley / Fudan University

LLM-as-a-Verifierverificationrobotics systemstest-time scaling
1 notes Shuo He

Nanyang Technological University / University of Electronic Science and Technology of China

agentic reinforcement learningmulti-agent systemsAI safetymachine learning
1 notes Shuo Yang

Tsinghua University

agent reinforcement learningon-policy skill distillationhindsight skillslong-horizon agents
1 notes Simran Arora

Together AI / Stanford University

ThunderAgentAI systemsFrontier Performanceefficient AI
1 notes Siqi Wang

Beihang University

Grapeagentic LLM servingdistributed GNN trainingspeculative decoding
1 notes Siqing Wang

ByteDance

large language modelsreinforcement learningrubric-based reinforcement learningon-policy self-distillation
1 notes Siran Yang

Alibaba Group

AI infrastructuredistributed trainingRL post-trainingROLL
1 notes Sirui Han (韓斯睿)

The Hong Kong University of Science and Technology

HiLS-Attentionlarge language modelsAI for financemachine learning
1 notes Size Zheng

ByteDance Seed / Peking University

distributed AI systemscompute-communication overlapTriton-distributedTileLink
1 notes Songquan Zhu

MiniMax

MiniMax Sparse AttentionMiniMax-M3MiniMax-M1MiniMax-01
1 notes Stefano Ermon

Stanford University

machine learninggenerative modelsFlashAttention
1 notes Steven Swanson

University of California, San Diego / Non-Volatile Systems Lab

computer systemsmemory and storage systemscomputer architecturesystem reliability
1 notes Tao Ge (葛涛)

Microsoft Research / Microsoft

large language model post-trainingsynthetic dataefficient trainingagentic methods
1 notes Teddy Lee

OpenAI (paper affiliation)

human dataprocess supervisionmodel fine-tuningalignment data
1 notes Tengyang Xie

University of Wisconsin-Madison / Microsoft Research

reinforcement learningmachine learningartificial intelligencelong-horizon planning
1 notes Tianbao Yang

Texas A&M University

optimizationmachine learningreinforcement learningreasoning language models
1 notes Tianjian Li

Johns Hopkins University / Meta FAIR (research scientist intern, 2025 and 2026)

natural language processinglanguage model post-trainingadaptive data methodsreinforcement learning
1 notes Tianle Cai (蔡天乐)

Princeton University / Together AI

systems and architecture co-designlarge language modelsefficient trainingreinforcement learning
1 notes Tianyu Zhang

Huawei Noah’s Ark Lab

LLM inference accelerationspeculative decodingtest-time scalingMoE inference
1 notes Tianyuan Wu

HKUST

distributed trainingRL post-training systemsagent sandbox runtimeAI infrastructure
1 notes Ting Jiang

Z.ai

IndexCacheGLM-5sparse attentionlong-context modeling
1 notes Tong Liu

LMU Munich / Munich Center for Machine Learning

tool-calling RLagent evaluationreinforcement learningefficient post-training
1 notes Tushar Krishna

Georgia Institute of Technology / MIT

ThunderAgentcomputer architectureinterconnection networksdeep learning accelerators
1 notes Vineet Kosaraju

OpenAI / Stanford University

mathematical reasoningverifiersgeneral-purpose reasoningtrajectory forecasting
1 notes Vito Zhang

MiniMax / Peking University

MiniMax Sparse AttentionMiniMax-M3sparse attention
1 notes Wayne Xin Zhao

Gaoling School of Artificial Intelligence, Renmin University of China / Peking University

LLM-in-Sandboxlarge language modelsrecommender systemsnatural language processing
1 notes Wei Gao

HKUST / Alibaba Group

RollArtROLLRollPackeragentic RL systems
1 notes Wei Liu (刘威)

HKUST / DeepSeek

LLM reinforcement learningcoding agentsagentic systemsscalable methods
1 notes Wei Wang

HKUST / HKUST Big Data Institute

RollArtROLLRollPackerdistributed systems
1 notes Weili Xu

University of Illinois Urbana-Champaign / Zhejiang University

ThunderAgentmachine learning systemsLLM systemsagentic inference
1 notes Weiqi Xu

MiniMax

MiniMax Sparse AttentionMiniMax-M3MiniMax
1 notes Weixun Wang

Alibaba Group

RollArtRollPackerROLL Flashagentic RL systems
1 notes Wenbo Su

Alibaba Group

RL post-traininglarge language modelsdistributed trainingrecommendation systems
1 notes Wenhan Ma

Peking University / Xiaomi LLM Core

MOPDMiMoMoE reinforcement learningpost-training
1 notes Wenjie Wang (王文杰)

University of Science and Technology of China / National University of Singapore

recommender systemscausal inferencegenerative recommendationLLM agents
1 notes Wenlin Yao

Microsoft Research

foundation language modelsend-to-end reinforcement learningagentic frameworkstool use
1 notes Wenting Zhao

Qwen Team / Alibaba Group

Qwen3-Coder-NextQwen-Coderreasoningcoding agents
1 notes William Jurayj

Johns Hopkins University / Amazon AGI (applied scientist internship, 2026)

LLM uncertainty and calibrationtest-time scalingneuro-symbolic reasoningdeontic reasoning
1 notes Xi Wang

Johns Hopkins University

uncertainty quantificationefficient reasoningtest-time computeearly exiting
1 notes Xiao Luo

University of Wisconsin–Madison / University of California, Los Angeles (postdoctoral affiliation)

LLM post-trainingLLM agentsgraph machine learningdynamical systems
1 notes Xiaolong Li

MiniMax / Zhejiang University

MiniMax Sparse AttentionMiniMax-M3sparse attention
1 notes Xiaoshuai Song (宋晓帅)

Gaoling School of Artificial Intelligence, Renmin University of China / Beijing University of Posts and Telecommunications

long-horizon agentsweb information acquisitiontool useagent environments
1 notes Xin Jin

Peking University / Key Laboratory of High Confidence Software Technologies, Peking University

computer systemscomputer networkingLLM training systemsMegaScale-MoE
1 notes Xin Lv (吕鑫)

Zhipu AI / Tsinghua University

reasoning reinforcement learningreinforcement learning infrastructureagentic reinforcement learninglong-context modeling
1 notes Xinxing Xu

Microsoft Research / Microsoft Research Asia Singapore

ReOPDfoundation modelscomputer visionmachine learning
1 notes Xinyu Lin

National University of Singapore / Shandong University

recommender systemscausal inferencegenerative recommendationLLM agents
1 notes Xinyu Wei

ShanghaiTech University

HiLS-Attentionnative sparse attentionlong-context modeling
1 notes Xinyu Yang

Carnegie Mellon University

ThunderAgentInfiniAI Labmachine learning systemsfoundation model systems
1 notes Xu Shen

Alibaba Group / Tongyi Large Model Business Unit, Alibaba Token Hub

HydraHeadLLM architecturelong-context modelingmachine learning
1 notes Xuandong Zhao

UC Berkeley / BAIR

AI safetyreasoning LLMsscalable reinforcement learningself-improvement
1 notes Xunhao Lai (赖勋豪)

MiniMax / Peking University

MiniMax Sparse AttentionMiniMax-M3FlexPrefilllong-context inference
1 notes Xuwu Wang

Qwen Team / Alibaba Group

Qwen3-Coder-NextQwen-Codersoftware engineering agentsexecutable environments
1 notes Yan Chen

University of Virginia

reinforcement learningefficient reasoningreasoning language models
1 notes Yanbo Zhou

University of California, San Diego / Non-Volatile Systems Lab

cloud infrastructurememory and storage systemssystem reliabilitysoftware-hardware co-design
1 notes Yaniv Leviathan

Google Research / Google

speculative decodingtransformer efficiencygenerative UILLM inference
1 notes Yaoyao Ding

University of Toronto / Vector Institute

LoRAFusionGPU kernelsML systemsHidet
1 notes Yaxiang Zhang

National University of Singapore

LLM reinforcement learningoptimal token baselinemachine learning
1 notes Yesheng Liang

UC San Diego / Z Lab

DFlashParoQuantefficient machine learningLLM inference
1 notes Yi Jing

Tsinghua University / Xinya College

mechanistic interpretabilityLLM post-trainingagentic reinforcement learningcontext compaction
1 notes Yicheng Hu

University of Science and Technology of China

LLM agentsrecommender systemsAI for Sciencemulti-agent systems
1 notes Yiding Jiang

Carnegie Mellon University / Google DeepMind

maximum likelihood reinforcement learningpost-trainingGeminigeneralization
1 notes Yifan Bai

Kimi Team / Moonshot AI

Kimi K2Kimi K2.5Moonshot AIagentic intelligence
1 notes Yifei Li

The Ohio State University / Peking University

LLM agentsscientific coding agentsMath-ShepherdAutoSDT
1 notes Yinfang Chen

University of Illinois Urbana-Champaign

ThunderAgentagentic AIsystems and networkingautonomous clouds
1 notes Yingdi Shan (闪英迪)

Tsinghua University / MADSys Lab

distributed systemsstorage systemsLLM servingSeer
1 notes Yixing Jiang

Stanford University / Stanford Machine Learning Group

LLM-as-a-Verifiermedical agentsagent evaluationelectronic health records
1 notes Yonathan Efroni

Tel Aviv University / AAI Technologies

interactive learningreinforcement learningsequential decision makingagentic systems
1 notes Yongbin Li (李永彬)

Tongyi Lab / Alibaba Group

code LLMscoding agentsQoderLLM post-training
1 notes Yongchao Zhou

Google DeepMind / University of Toronto

on-policy distillationlanguage model distillationlarge language modelsreasoning
1 notes Yongpan Liu (刘勇攀)

Tsinghua University / Beijing National Research Center for Information Science and Technology

AI acceleratorshardware-software co-designenergy-efficient computingMoE inference acceleration
1 notes Yongwei Wu

Tsinghua University / MADSys Lab

parallel systemsdistributed systemsstorage systemsLLM serving
1 notes Yoonho Lee

Stanford University

SPIRALcontinual learningtext optimizationAI agents
1 notes Yossi Matias

Google Research / Google

Google Researchspeculative decodingalgorithmsLLM inference
1 notes Youliang Yu

Meta AI

reinforcement learningLLM post-trainingself-improvement
1 notes Youyou Lu

Tsinghua University / Tsinghua Storage Research Group

SPORKcomputer systemscomputer architecturememory systems
1 notes Yu Cheng

The Chinese University of Hong Kong / Shanghai AI Laboratory

efficient architecturesmodel compressionmultimodal learningRLVR
1 notes Yuan He (何源)

Amazon / University of Oxford

tool-calling RLagent post-trainingagent environmentsreinforcement learning infrastructure
1 notes Yuda Song

Carnegie Mellon University / FAIR Paris

maximum likelihood reinforcement learninginteractive decision-makingreinforcement learning theoryfoundation models
1 notes Yue Guan

University of California, San Diego / PICASSO Lab

efficient machine learningmodel compressionmachine learning systemsLLM serving
1 notes Yue Wu

Alibaba Group / Alibaba Cloud

HydraHeadlarge language modelsmachine learningimage and video processing
1 notes Yueer Zhou

Zhejiang University / Stanford University

maximum likelihood reinforcement learningself-improving foundation modelscontinual learninglow-rank adaptation
1 notes Yuejiang Liu

Stanford University / National University of Singapore

LLM-as-a-Verifierrobot learningworld modelsverification
1 notes Yufei Ding

University of California, San Diego / PICASSO Lab

machine learning systemsdomain-specific languagescompiler optimizationGPU systems
1 notes Yufeng Yang

MiniMax

MiniMax Sparse AttentionMiniMax-M3MiniMax-M1MiniMax-01
1 notes Yuhang Yang

Zhejiang University / ByteDance

large language modelstabular reasoningnatural language processingrubric-based reinforcement learning
1 notes Yuheng Jing

Qwen Team / Alibaba Group

Qwen3-Coder-NextQwen-Codermobile device agentscoding agents
1 notes Yuheng Zhao

HKUST / Alibaba Group

RollArtROLLRollPackeragentic RL systems
1 notes Yuhui Li

University of Waterloo / Peking University

EAGLEspeculative decodingdraft-model trainingLLM inference
1 notes Yujiu Yang

Tsinghua University / Tsinghua Shenzhen International Graduate School

FlashMemoryDeepSeek-V4Lookahead Sparse Attentionmultimedia
1 notes Yulun Du

Moonshot AI / Carnegie Mellon University

Kimi K2Kimi K2.5pretrainingMoonshot AI
1 notes Yunfan Xiong

Peking University / DeepSeek AI

DSparkspeculative decodingLLM serving
1 notes Yunlong Feng

Qwen Team / Alibaba Group

Qwen3-Coder-NextQwen-Codersoftware engineering agentsexecutable environments
1 notes Yuqi Wu

ByteDance Seed / Shanghai Jiao Tong University

Laminarasynchronous RLpost-training systems
1 notes Yura Burda

OpenAI / University of Toronto

mathematical reasoningreinforcement learningexplorationrandom network distillation
1 notes Yuxian Gu

Tsinghua University / Microsoft Research Asia

LLM-in-Sandboxefficient LLM traininglarge language modelsknowledge distillation
1 notes Yuyang Hu (扈煜阳)

Gaoling School of Artificial Intelligence, Renmin University of China / Beijing Academy of Artificial Intelligence

long-horizon agentsagent memoryself-evolving agentsAutoResearch agents
1 notes Yuyang You

School of Mathematical Sciences, Peking University

ECHOagentic RLLLM agentslong-horizon tool use
1 notes Zaifeng Pan (潘再峰)

University of California, San Diego / PICASSO Lab

machine learning systemsLLM servingagent servingKV cache
1 notes Zekun Li

MiniMax

MiniMax Sparse AttentionMiniMax-M3inference kernels
1 notes Zeyao Ma

Qwen Team / Alibaba Group

Qwen3-Coder-Nextcoding agentscode generationreasoning
1 notes Zeyu Chen

Baidu Inc.

ECHOagentic RLLLM agentsBaidu
1 notes Zeyu Cui

Qwen Team / Alibaba Group

Qwen3-Coder-NextQwen-Codercoding agentslarge language models
1 notes Zeyu Jia (贾泽宇)

Massachusetts Institute of Technology / Peking University

reinforcement learning theorymachine learning theoryoptimizationinformation theory
1 notes Zhanda Zhu

University of Toronto / Vector Institute

LoRAFusionLLM systemsGPU systemsLoRA fine-tuning
1 notes Zhengding Hu

University of California, San Diego / PICASSO Lab

high-performance computingmachine learning systemsLLM servingretrieval-augmented generation
1 notes Zhenghai Xue

Nanyang Technological University / Moonshot AI

reinforcement learningLLM agentstool-integrated reasoningdecision making
1 notes Zhengxiao Du

Z.ai / Tsinghua University

IndexCacheGLM-5large language modelssparse attention
1 notes Zhentao Tan

Alibaba Group / Alibaba Cloud

HydraHeadhead-wise hybrid attentioninterpretability-guided head selectionlong-context modeling
1 notes Zhewei Kang

UC Berkeley / University of Hong Kong

RLVRreasoning LLMsvalue-implicit policy optimizationself-certainty
1 notes Zhichao Wang

Tencent

FlashMemoryDeepSeek-V4Lookahead Sparse Attention
1 notes Zhicheng Dou (窦志成)

Gaoling School of Artificial Intelligence, Renmin University of China / Microsoft Research Asia

long-horizon agentsintelligent agentsdeep searchretrieval-augmented generation
1 notes Zhijian Liu

UC San Diego / Z Lab

DFlashefficient AILLM inferencemodel compression
1 notes Zhiyin Yu

Peking University / Shanghai Artificial Intelligence Laboratory

data-efficient reinforcement learningLLM post-trainingself-evolving LLMsAI for science
1 notes Zhonghai Wu (吴中海)

Peking University

big-data machine learningsoftware and systems securityprivacy computingembedded intelligent systems
1 notes Zibo Lin

Tencent

FlashMemoryDeepSeek-V4Lookahead Sparse Attention
1 notes Zichen Liu

Alibaba Group

ROLLasynchronous RLagentic RL systemsRL post-training
1 notes Ziheng Jiang

ByteDance Seed / ByteDance

MegaScaleMegaScale-MoEVERLLLM systems
1 notes Zijun Xie

School of Mathematical Sciences, Peking University / Baidu Inc.

ECHOagentic RLsource-indexed memorylong-horizon agents
1 notes Ziqian Zhong

Carnegie Mellon University / Transluce

AI safetyinterpretabilitycoding agentsreward hacking
1 notes Ziyang Li

Individual Researcher / Zhejiang University

ThunderAgentagentic inferenceLLM systems
1 notes Zongle Huang

Tsinghua University / Beijing National Research Center for Information Science and Technology

MoE inference accelerationLLM systemsAI accelerator architecturepipeline parallelism
1 notes Zongmeng Zhang

Qwen Team / Alibaba Group

Qwen3-Coder-Nextcoding agentslarge language modelsreinforcement learning

Recurring Authors

这些作者在当前归档中至少出现两次,页面由论文元数据自动汇总。