Publications
Journal Publications
On the Benefit of Width for Neural Networks: Disappearance of Bad Basins, with Tian Ding and Ruoyu Sun. SIAM Journal on Optimization, 2022, 32(3): 1728-1758.
Suboptimal Local Minima Exist for Wide Neural Networks with Smooth Activations, with Tian Ding and Ruoyu Sun. Mathematics of Operations Research, 2022, 47(4): 2784-2814.
The Global Landscape of Neural Networks: An Overview, with Ruoyu Sun, Shiyu Liang, Tian Ding and Rayadurgam Srikant. IEEE Signal Processing Magazine, 2020, 37(5): 95-108.
On a Faster R-Linear Convergence Rate of the Barzilai-Borwein Method, with Ruoyu Sun. To appear in Journal of Computational Mathematics.
Conference Proceedings
NTK-SAP: Improving Neural Network Pruning by Aligning Training Dynamics, with Yite Wang and Ruoyu Sun. In International Conference on Learning Representations, 2023.
RMSProp Converges with Proper Hyper-parameter, with Naichen Shi, Mingyi Hong and Ruoyu Sun. In International Conference on Learning Representations (Spotlight), 2021.
Preprints
Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training, with Zijian Zhang, Rizhen Hu, Athanasios Glentis, Chung-Yiu Yau, Hongzhou Lin and Mingyi Hong.
A Geometric Characterization of the Stationary Plateau for Two-Layer Neural Networks, with Tian Ding and Ruoyu Sun.
EMA-Nesterov: Stabilizing Nesterov’s Lookahead for Accelerated Deep Learning Optimization, with Chung-Yiu Yau, Athanasios Glentis, Valentyn Boreiko, Hoi-To Wai and Mingyi Hong.
Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates, with Athanasios Glentis, Chung-Yiu Yau and Mingyi Hong.
Working Papers
- Barzilai–Borwein Fails Superlinear Convergence on an Open Set of Quadratics for Every Dimension $n\geq 4$, with Xiaotian Jiang and Mingyi Hong.
