Publications

Journal Publications

  1. On the Benefit of Width for Neural Networks: Disappearance of Bad Basins, with Tian Ding and Ruoyu Sun. SIAM Journal on Optimization, 2022, 32(3): 1728-1758.

  2. Suboptimal Local Minima Exist for Wide Neural Networks with Smooth Activations, with Tian Ding and Ruoyu Sun. Mathematics of Operations Research, 2022, 47(4): 2784-2814.

  3. The Global Landscape of Neural Networks: An Overview, with Ruoyu Sun, Shiyu Liang, Tian Ding and Rayadurgam Srikant. IEEE Signal Processing Magazine, 2020, 37(5): 95-108.

  4. On a Faster R-Linear Convergence Rate of the Barzilai-Borwein Method, with Ruoyu Sun. To appear in Journal of Computational Mathematics.

Conference Proceedings

  1. NTK-SAP: Improving Neural Network Pruning by Aligning Training Dynamics, with Yite Wang and Ruoyu Sun. In International Conference on Learning Representations, 2023.

  2. RMSProp Converges with Proper Hyper-parameter, with Naichen Shi, Mingyi Hong and Ruoyu Sun. In International Conference on Learning Representations (Spotlight), 2021.

Preprints

  1. Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training, with Zijian Zhang, Rizhen Hu, Athanasios Glentis, Chung-Yiu Yau, Hongzhou Lin and Mingyi Hong.

  2. A Geometric Characterization of the Stationary Plateau for Two-Layer Neural Networks, with Tian Ding and Ruoyu Sun.

  3. EMA-Nesterov: Stabilizing Nesterov’s Lookahead for Accelerated Deep Learning Optimization, with Chung-Yiu Yau, Athanasios Glentis, Valentyn Boreiko, Hoi-To Wai and Mingyi Hong.

  4. Revisiting the Adam-SGD Gap in LLM Pre-Training: The Role of Large Effective Learning Rates, with Athanasios Glentis, Chung-Yiu Yau and Mingyi Hong.

Working Papers

  1. Barzilai–Borwein Fails Superlinear Convergence on an Open Set of Quadratics for Every Dimension $n\geq 4$, with Xiaotian Jiang and Mingyi Hong.