Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Learning to Play Two-Player Perfect-Information Games without Knowledge

Дата публикации: 17-08-2026 20:26:00


This paper introduces a set of techniques for learning game state evaluation functions through reinforcement learning. First, we generalize tree bootstrapping, i.e. learning the values of states encountered during search rather than restricting updates to states observed during matches, to the setting of reinforcement learning with non-linear function approximation. Second, we modifies Unbounded Best-First Minimax by extending best action sequences to terminal states. Third, we replace the traditional binary game outcome $+1/-1$ with richer reinforcement signals, including quick wins, delayed losses, and scoring. Fourth, we propose a completion mechanism that exploits state resolution.
Finally, we introduce a novel action-selection distribution, referred to as the ordinal distribution.
Experimental results show that each of these techniques contributes to substantial improvements in playing strength. We integrate them into a unified algorithm, Athénan, and compare it against ExIt, a leading self-play reinforcement learning approach without prior knowledge.
Our results demonstrate that Athénan consistently outperforms ExIt.
We further evaluate Athénan on the games Hex, Othello, and Arimaa, where it surpasses state-of-the-art performance without relying on domain-specific knowledge. In addition, we consider the single-player game Morpion Solitaire, in which Athénan again reaches state-of-the-art results under the same constraint.
Overall, these results show that reinforcement learning, when combined with the proposed techniques, can achieve state-of-the-art performance across a diverse range of games without the need for handcrafted heuristics or expert knowledge.

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1 A Reinforcement Learning Approach in Multi-Phase Second-Price Auction Design 07.3417-08-2026
2 Approximation-Free Differentiable Oblique Decision Trees 09.617-08-2026
3 End-to-End Deep Learning for Predicting Metric Space-Valued Outputs 010.6617-08-2026
4 The Sample Complexity of Parameter-Free Stochastic Convex Optimization 05.717-08-2026
5 High-Dimensional Analysis of Gradient Flow for Extensive-Width Quadratic Neural Networks 08.717-08-2026
6 Graph-based Clustering Revisited: A Relaxation of Kernel k-Means Perspective 010.9417-08-2026
7MO-Gymnasium - environments for reinforcement learning026.6719-07-2026
8 Near-optimal Delta-convex Estimation of Lipschitz Functions 09.7117-08-2026
9 Bridging Domain Invariance and Diversity: A Fine-Grained Risk Bound for Domain Generalization 0717-08-2026
10Equitable Domination in Turiyam Graphs with Network Applications [version 1; peer review: 3 approved]0701-06-2026

Классификация: . Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 9.01. Источник: jmlr.org.