Reinforcement Learning with Probabilistically Complete Exploration

Morere, Philippe; Francis, Gilad; Blau, Tom; Ramos, Fabio

Full-text links:

Download:

Current browse context:

cs.LG

< prev | next >

new | recent | 2001

Computer Science > Machine Learning

Title: Reinforcement Learning with Probabilistically Complete Exploration

Authors: Philippe Morere, Gilad Francis, Tom Blau, Fabio Ramos

(Submitted on 20 Jan 2020)

Abstract: Balancing exploration and exploitation remains a key challenge in reinforcement learning (RL). State-of-the-art RL algorithms suffer from high sample complexity, particularly in the sparse reward case, where they can do no better than to explore in all directions until the first positive rewards are found. To mitigate this, we propose Rapidly Randomly-exploring Reinforcement Learning (R3L). We formulate exploration as a search problem and leverage widely-used planning algorithms such as Rapidly-exploring Random Tree (RRT) to find initial solutions. These solutions are used as demonstrations to initialize a policy, then refined by a generic RL algorithm, leading to faster and more stable convergence. We provide theoretical guarantees of R3L exploration finding successful solutions, as well as bounds for its sampling complexity. We experimentally demonstrate the method outperforms classic and intrinsic exploration techniques, requiring only a fraction of exploration samples and achieving better asymptotic performance.

Subjects:	Machine Learning (cs.LG); Robotics (cs.RO); Machine Learning (stat.ML)
Cite as:	arXiv:2001.06940 [cs.LG]
	(or arXiv:2001.06940v1 [cs.LG] for this version)

Submission history

From: Tom Blau [view email]
[v1] Mon, 20 Jan 2020 02:11:24 GMT (2754kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2001.06940

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Machine Learning

Title: Reinforcement Learning with Probabilistically Complete Exploration

Submission history