Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Optimizing Attention with Mirror Descent: Generalized Max-Margin Token Selection

Дата публикации: 17-08-2026 20:26:00


Attention mechanisms have revolutionized several domains of artificial intelligence, such as natural language processing and computer vision, by enabling models to selectively focus on relevant parts of the input data. While recent work has characterized the optimization dynamics of gradient descent (GD) in attention-based models and the structural properties of its preferred solutions, less is known about more general optimization algorithms such as mirror descent (MD). In this paper, we investigate the convergence properties and implicit biases of a family of MD algorithms tailored for softmax attention mechanisms, with the potential function chosen as the $p$-th power of the $\ell_p$-norm. Specifically, we show that these algorithms converge in direction to a generalized hard-margin SVM with an $\ell_p$-norm objective when applied to a classification problem using a softmax attention model. Notably, our theoretical results reveal that the convergence rate is comparable to that of traditional GD in simpler models, despite the highly nonlinear and nonconvex nature of the present problem. Additionally, we delve into the joint optimization dynamics of the key-query matrix and the decoder, establishing conditions under which this complex joint optimization converges to their respective hard-margin SVM solutions. Lastly, our numerical experiments on real data demonstrate that MD algorithms improve generalization over standard GD and excel in optimal token selection.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1 Towards Convexity in Anomaly Detection: A New Formulation of SSLM with Unique Optimal Solutions 05.917-08-2026
2 Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes 07.5217-08-2026
3 A Mean-Field Analysis of Neural Stochastic Gradient Descent-Ascent for Functional Minimax Optimization 09.8217-08-2026
4 Graph-based Clustering Revisited: A Relaxation of Kernel k-Means Perspective 010.9417-08-2026
5 High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks 08.717-08-2026
6 The Sample Complexity of Parameter-Free Stochastic Convex Optimization 05.717-08-2026
7 Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks 0517-08-2026
8 Unsupervised Feature Selection via Nonnegative Orthogonal Constrained Regularized Minimization 05.3317-08-2026
9 Gradient Span Algorithms Make Predictable Progress in High Dimension 06.3817-08-2026
10 Extrapolated Markov Chain Oversampling Method for Imbalanced Text Classification 05.2317-08-2026

Классификация: Пресс-релизы. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 6.96. Источник: jmlr.org.