STEEL: Singularity-aware Reinforcement Learning

Xiaohong Chen; Zhengling Qi; Runzhe Wan

STEEL: Singularity-aware Reinforcement Learning

Fiche du document

Auteurs

Date

30 janvier 2023

Discipline

Type de document

Textes imprimés

Périmètre

Publications

Identifiant

2301.13152

Source

arXiv - économie

Collection

arXiv

Organisation

Cornell University

Mots-clés Und

Statistics - Machine Learning Computer Science - Machine Learning Economics - Econometrics Statistics - Methodology

Sujets proches En

Supposition Assumption Continuum

Citer ce document

Xiaohong Chen et al., « STEEL: Singularity-aware Reinforcement Learning », arXiv - économie

Partage / Export

Résumé 0

Batch reinforcement learning (RL) aims at leveraging pre-collected data to find an optimal policy that maximizes the expected total rewards in a dynamic environment. Nearly all existing algorithms rely on the absolutely continuous assumption on the distribution induced by target policies with respect to the data distribution, so that the batch data can be used to calibrate target policies via the change of measure. However, the absolute continuity assumption could be violated in practice (e.g., no-overlap support), especially when the state-action space is large or continuous. In this paper, we propose a new batch RL algorithm without requiring absolute continuity in the setting of an infinite-horizon Markov decision process with continuous states and actions. We call our algorithm STEEL: SingulariTy-awarE rEinforcement Learning. Our algorithm is motivated by a new error analysis on off-policy evaluation, where we use maximum mean discrepancy, together with distributionally robust optimization, to characterize the error of off-policy evaluation caused by the possible singularity and to enable model extrapolation. By leveraging the idea of pessimism and under some mild conditions, we derive a finite-sample regret guarantee for our proposed algorithm without imposing absolute continuity. Compared with existing algorithms, by requiring only minimal data-coverage assumption, STEEL significantly improves the applicability and robustness of batch RL. Extensive simulation studies and one real experiment on personalized pricing demonstrate the superior performance of our method in dealing with possible singularity in batch RL.

STEEL: Singularity-aware Reinforcement Learning

Fiche du document

Mots-clés Und

Sujets proches En

Citer ce document

Partage / Export

Résumé 0

Par les mêmes auteurs

Sur les mêmes sujets

Sur les mêmes disciplines

Exporter en