讲座标题(中文):
规模定律的普适性:涌现幂律拥有稳健的幂指数
讲座标题(英文):
Neural Scaling Universality: Emergent Power Laws Obey Robust Exponents
讲座摘要:
Abstract: Neural scaling laws are among the key observations driving the success of large language models. The mainstream view attributes power-law scaling in loss to power-law structures in data. In this talk, I present an alternative: power laws in loss can emerge from high-level architecture and data properties, yielding constant exponents robust to other details. Specifically, we identify a one-third time scaling due to Softmax and cross-entropy nonlinearities, an inverse width scaling due to representational superposition, and an inverse depth scaling due to ensemble averaging across Transformer layers. These fixed exponents, together with the mechanisms, define a universality class, that a broad range of architectures and datasets can fall into. With this new view, I will close by discussing implications on architecture and training choices in the current class, and possible pathways to other classes with better scaling.
个人简介:刘逸舟,现为麻省理工学院博士生,本科毕业于清华大学,主要研究复杂系统中的涌现行为,近期专注于大模型中的规模定律(neural scaling laws)。