一项行之有效的强化学习抽样收集战略 (A Provably Efficient Sample Collection Strategy for Reinforcement Learning)

One of the challenges in online reinforcement learning (RL) is that the agent needs to trade off the exploration of the environment and the exploitation of the samples to optimize its behavior. Whether we optimize for regret, sample complexity, state-space coverage or model estimation, we need to strike a different exploration-exploitation trade-off. In this paper, we propose to tackle the exploration-exploitation problem following a decoupled approach composed of: 1) An "objective-specific" algorithm that (adaptively) prescribes how many samples to collect at which states, as if it has access to a generative model (i.e., a simulator of the environment); 2) An "objective-agnostic" sample collection exploration strategy responsible for generating the prescribed samples as fast as possible. Building on recent methods for exploration in the stochastic shortest path problem, we first provide an algorithm that, given as input the number of samples $b(s,a)$ needed in each state-action pair, requires $\tilde{O}(B D + D^{3/2} S^2 A)$ time steps to collect the $B=\sum_{s,a} b(s,a)$ desired samples, in any unknown communicating MDP with $S$ states, $A$ actions and diameter $D$. Then we show how this general-purpose exploration algorithm can be paired with "objective-specific" strategies that prescribe the sample requirements to tackle a variety of settings -- e.g., model estimation, sparse reward discovery, goal-free cost-free exploration in communicating MDPs -- for which we obtain improved or novel sample complexity guarantees.

翻译：在线强化学习(RL)的挑战之一是,代理商需要权衡环境勘探和样本开发,以优化其行为。无论我们为了遗憾、抽样复杂性、州-空间覆盖范围或模型估计而优化,我们都需要做出不同的勘探-开发权衡。在本文中,我们提议采用一种分解方法来解决勘探-开发问题,该方法包括:1)“目标特定”算法,该算法(可调整)规定在哪些样本中进行采集,以表明其是否具备基因模型(即环境模拟器);2“目标-敏感”样本采集勘探战略,负责尽可能快地生成规定的样本。在最新探索方法的基础上,我们首先提供一种算法,作为样本数量(美元,a)每一州行动模型需要的美元,要求以美元(美元)计算成本(D+D%3/2)成本(S%2 A),用于收集标定的样本的“目标-美元(美元)成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-成本-