讲座标题(中文):深入理解深度学习中的对抗样本现象:从模型表达能力与训练动力学视角
讲座标题(英文):Theoretical Understanding of Adversarial Examples in Deep Learning: Expressive Power and Training Dynamics
讲座摘要:In recent years, machine learning methods—especially deep learning—have shown exceptional performance in domains like computer vision, natural language processing, speech recognition, and game playing. However, deep neural networks still face fundamental limitations in robustness and reliability. A key issue is their vulnerability to adversarial examples—small perturbations that cause incorrect predictions while being imperceptible to humans. This poses significant concerns for deploying deep models in safety-critical applications, such as autonomous driving. In this talk, we aim to provide a theoretical account of adversarial examples in deep learning. Our analysis is grounded in two key perspectives: the expressive power of neural networks and the underlying principles of feature learning. By connecting these theoretical foundations, we seek to shed light on the mechanisms that give rise to adversarial vulnerability and offer insights into potential pathways for improving model robustness.
讲者信息:李柄辉,北京大学前沿交叉学科研究院国际机器学习研究中心2023级博士生,主要研究方向为深度学习的理论基础以及人工智能方法在数学问题中的应用。
相关工作:
Why Robust Generalization in Deep Learning is Difficult: Perspective of Expressive Power
https://arxiv.org/abs/2205.13863
Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural Networks
https://arxiv.org/abs/2410.10322
Adversarial Training Can Provably Improve Robustness: Theoretical Analysis of Feature Learning Process Under Structured Data
https://arxiv.org/abs/2410.08503