ReportGem ReportGem

Academic paper

Sample Efficient Hierarchical Reinforcement Learning via Best Policy Identification

Authors: Anders Jonsson, Emilie Kaufmann, Gianmarco Tedeschi, Lorenzo SteccanellaPublished: 2026-07-31Paper ID: 2607.29294Category: cs.LGLicense: CC BY-SA 4.0

Abstract

We present HBPI-UCRL, a model-based algorithm for hierarchical reinforcement learning (HRL) that learns high-level and low-level policies in parallel. HBPI-UCRL exploits the fact that a high-level transition corresponds to a multi-step transition at the low level. We introduce two conditions on the low-level dynamics that are sufficient to make parallel HRL learnable. When these conditions hold, we prove that HBPI-UCRL has a polynomial sample complexity in the problem parameters. In the sparse-reward, goal-directed setting, our sample complexity upper bound for HBPI-UCRL is strictly lower than that of its non-hierarchical counterpart, providing theoretical justification for the empirical success of HRL.

This public page contains bibliographic metadata and the author abstract. Use the reader for licensed document access.

Open licensed paper reader