tracked author

Yaniv Leviathan

Fast Inference from Transformers via Speculative Decoding equal-contribution and corresponding author. Public homepage identifies him as a Google Fellow and AI research lab lead, links the yanivle social profile, and lists speculative decoding as one of his major works.

1 archived notes X: high HomepageGitHubX

Related Notes

按论文归档时间排序,展示该作者在本站已经出现的材料。

Jun 24, 2026

Fast Inference from Transformers via Speculative Decoding

这篇论文把 speculative execution 推到随机采样场景:小 draft model 先自回归猜 $\gamma$ 个 token,大 target model 一次并行验证这些前缀,并用 $\min(1,p/q)$ 接受概率与 residual distribution 校正拒绝位置,从而在无需改模型、无需重训且保持 target 输出分布不变的前提下,把大模型串行 decode 的目标调用数降低到每步平均生成多个 ...

2211.17192-fast-inference-transformers-speculative-decoding Speculative Decoding