tracked author

Nipun Kwatra

Sarathi coauthor; Microsoft Research profile identifies him as a Principal Researcher at MSR India working at the intersection of deep learning and systems. No high-confidence personal X account was found.

1 archived notes X: not-found Homepage

Related Notes

按论文归档时间排序,展示该作者在本站已经出现的材料。

Jun 19, 2026

SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Sarathi 的核心贡献是把 LLM serving 的低效从“decode 天然慢”重写为一个可调度的数据流问题:prefill 很快进入 compute saturating 状态,decode 因为逐 token 生成和 KV cache 访问长期 memory bound;如果把一个长 prefill 切成多个 compute sized chunk,再让 decodes 搭在每个 prefill chunk 的 linea...

2308.16369-sarathi-chunked-prefill-decode-maximal-batching Inference SchedulingServing RuntimeKV Cache