recurring author
Zhewen Hao
2 archived notes
Related Notes
按论文归档时间排序,展示该作者在本站已经出现的材料。
归档 更新
DSpark: Confidence Scheduled Speculative Decoding with Semi Autoregressive Generation
用 Markov head、置信度校准和硬件感知前缀调度,把并行 drafter 推进生产 serving。
待审阅 2026-06-27-dspark-confidence-scheduled-speculative-decoding Speculative DecodingMulti-Token PredictionServing Runtime
归档 更新
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
用 hashed N gram lookup 和 context aware gating 增加可离线扩展的 conditional memory。