tracked author

Runheng Liu

EcoSpec co-first author at Beijing Institute of Technology and FlashBack first author. Explicit paper links are required to distinguish this NLP researcher from same-name authors in other fields.

1 archived notes X: not-found

Related Notes

按论文归档时间排序,展示该作者在本站已经出现的材料。

归档

Less Experts, Faster Decoding: Cost Aware Speculative Decoding for Mixture of Experts

EcoSpec 用轻量路由预测器估计草稿候选会新增的专家集合,并优先选择复用已有专家的验证路径;在八张 H200、Hugging Face Transformers、主要批量为一的研究原型中,三个大规模 MoE 的贪心解码平均加速从对应投机基线相对自回归的一点一零至一点二二倍提高到一点一五至一点三六倍,HBM 流量来自估算且批量为八时投机吞吐低于自回归。

待审阅 2607.12696-ecospec-cost-aware-moe-speculative-decoding Speculative DecodingMoE SystemsServing Runtime