Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

A Two-Timescale Primal-Dual Framework for Reinforcement Learning via Online Dual Variable Guidance

Дата публикации: 17-08-2026 20:26:00


We study reinforcement learning by combining recent advances in regularized linear programming formulations with the classical theory of stochastic approximation.
Motivated by the challenge of designing algorithms that leverage off-policy data while maintaining on-policy exploration, we propose PGDA-RL, a novel primal-dual projected gradient descent-ascent algorithm for solving regularized Markov decision processes (MDPs). PGDA-RL integrates experience replay-based gradient estimation with a two-timescale decomposition of the underlying nested optimization problem.
The algorithm operates asynchronously, interacts with the environment through a single trajectory of correlated data, and updates its policy online in response to the dual variable associated with the occupancy measure of the underlying MDP. We prove that PGDA-RL converges almost surely to the optimal value function and policy of the regularized MDP. Our convergence analysis relies on tools from stochastic approximation theory and holds under weaker assumptions than those required by existing primal-dual RL approaches, notably removing the need for a simulator or a fixed behavioral policy.
Under a strengthened ergodicity assumption on the underlying Markov chain, we establish a last-iterate finite-time guarantee with $\widetilde{\mathcal O}(k^{-2/3})$ mean-square convergence, aligning with the best-known rates for two-timescale stochastic approximation methods under Markovian sampling and biased gradient estimates.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1 Almost Sure Convergence of Linear Temporal Difference Learning with Arbitrary Features 04.217-08-2026
2 Approximations and Learning for Continuous State and Action MDPs under Average Cost Criteria 04.117-08-2026
3 Graph-based Clustering Revisited: A Relaxation of Kernel k-Means Perspective 010.9417-08-2026
4 Convergence of Decentralized Stochastic Subgradient-based Methods for Nonsmooth Nonconvex Optimization 08.7817-08-2026
5 Finite-Time Decoupled Convergence in Nonlinear Two-Time-Scale Stochastic Approximation 09.6417-08-2026
6 Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes 07.5217-08-2026
7 A Single-Loop Stochastic Proximal Quasi-Newton Method for Large-Scale Nonsmooth Convex Optimization 0817-08-2026
8 Best Arm Identification with Minimal Regret 08.417-08-2026
9 High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks 08.717-08-2026
10 A Reinforcement Learning Approach in Multi-Phase Second-Price Auction Design 07.3417-08-2026

Классификация: Наука. Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 13.12. Источник: jmlr.org.