tracked author

Shen Yan

ByteDance Seed coauthor of Dynamic Linear Attention. LinkedIn identifies him as Research Scientist at ByteDance / Bytedance Seed responsible for multimodal pretraining; personal homepage records Research Scientist at Google DeepMind, Michigan State University PhD, and earlier work with ByteDance AML and Google Research. Homepage links the same X/Twitter handle.

2 archived notes X: high HomepageX

Related Notes

按论文归档时间排序,展示该作者在本站已经出现的材料。

归档

SMELT: Scaling Laws for Compute Matched MoE Looped Transformers

SMELT 在近似匹配每 token FLOPs、非嵌入参数量与 KV cache 的 MoE 对照中,将中间一半层重复执行两次并以缩窄隐藏维度和增加专家数补偿预算;四规模三稀疏度的独立缩放面拟合估计,在拟合覆盖的十的二十次方至十的二十一次方 FLOPs 区间达到相同验证损失可节省 6.8% 至 18.0% 训练计算,但报告使用专有数据与训练栈且未验证实际时延。

待审阅 2609.01343-smelt-compute-matched-moe-looped-transformers MoE ArchitectureScaling Laws