tracked author

Weixiao Huang

Seer author; GitHub profile states Tsinghua graduation and cloud-native / AI infra focus, and MADSys author page connects the same name to disaggregated LLM serving work. Kimi K2.5 appendix contributor list also includes this author. No high-confidence personal X account was found.

4 archived notes X: not-found Homepage

Related Notes

按论文归档时间排序,展示该作者在本站已经出现的材料。

归档

Attention Residuals

AttnRes 用输入相关的深度 softmax 聚合替代 PreNorm 的单位权重累积,并以块级表示和系统调度控制跨层状态成本;五个 194M–528M 激活参数 Kimi Linear MoE 设置均降低单次报告验证损失,48B/3B 的块级变体十五项下游数值十四升一平,但缺少随机种子、大规模 Full 对照和完整系统测量条件。

待审阅 2603.15031-attention-residuals Training StabilityTraining MemoryDistributed Training
归档

Kimi K3: Open Frontier Intelligence

Kimi K3 将三层 KDA 与一层全局注意力交错、跨块 Attention Residuals 和每个 token 激活 16 个路由专家的 Stable LatentMoE 组合为 2.78 万亿总参数、1042 亿激活参数、百万 token 上下文的原生多模态模型;报告的缩放律拟合把架构、数据与训练配方的合并收益估计为相对 Kimi K2 约 2.5 倍,尚未拆分单个组件贡献。

待审阅 2026-07-27-kimi-k3-open-frontier-intelligence MoE ArchitectureLinear AttentionLong Context