Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

A Mean-Field Analysis of Neural Stochastic Gradient Descent-Ascent for Functional Minimax Optimization

Дата публикации: 17-08-2026 20:26:00


This paper studies minimax optimization problems defined over infinite-dimensional function classes of over-parameterized two-layer neural networks. In particular, we consider the minimax optimization problem stemming from estimating linear functional equations defined by conditional expectations, where the objective functions are quadratic in the functional spaces. We address (i) the convergence of the stochastic gradient descent-ascent algorithm and (ii) the representation learning of the neural networks. We establish convergence in the mean-field regime by considering the continuous-time, infinite-width limit of the optimization dynamics.
Under this regime, stochastic gradient descent-ascent corresponds to a Wasserstein gradient flow over the space of probability measures defined over the space of neural network parameters. We prove that the Wasserstein gradient flow converges globally to a stationary point of the minimax objective at a $\mathcal{O}(T^{-1} + \alpha^{-1} ) $ sublinear rate, and additionally finds the solution to the functional equation when the regularizer of the minimax objective is strongly convex. Here $T$ denotes the time and $\alpha$ is a scaling parameter of the neural networks. In terms of representation learning, our results show that the feature representation induced by the neural networks may deviate from the initial representation by a factor of $\mathcal{O}(\alpha^{-1})$, measured by the Wasserstein distance. Finally, we apply our general results to concrete examples, including policy evaluation, nonparametric instrumental variable regression, and asset pricing.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1 A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks 010.9717-08-2026
2 Finite Neural Networks as Mixtures of Gaussian Processes: From Provable Error Bounds to Prior Selection 04.2317-08-2026
3 Statistical Learning Theory for Neural Operators 010.2117-08-2026
4 High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks 08.717-08-2026
5 Optimization and Generalization of Gradient Descent for Shallow ReLU Networks with Minimal Width 03.8417-08-2026
6 Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes 07.5217-08-2026
7 A Fully Parameter-Free Second-Order Algorithm for Convex-Concave Minimax Problems 013.1117-08-2026
8 Error Analysis for Deep ReLU Feedforward Density-Ratio Estimation with Bregman Divergence 08.7817-08-2026
9 Convergence of Decentralized Stochastic Subgradient-based Methods for Nonsmooth Nonconvex Optimization 08.7817-08-2026
10 Online Detection of Changes in Moment--Based Projections: When to Retrain Deep Learners or Update Portfolios? 07.6417-08-2026

Классификация: Пресс-релизы. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 9.82. Источник: jmlr.org.