Dyna-T: Dyna-Q and Upper Confidence Bounds Applied to Trees

Faycal, Tarek; Zito, Claudio

Full-text links:

Download:

Current browse context:

cs.LG

< prev | next >

new | recent | 2201

Computer Science > Machine Learning

Title: Dyna-T: Dyna-Q and Upper Confidence Bounds Applied to Trees

Authors: Tarek Faycal, Claudio Zito

(Submitted on 12 Jan 2022 (v1), last revised 19 Jan 2022 (this version, v2))

Abstract: In this work we present a preliminary investigation of a novel algorithm called Dyna-T. In reinforcement learning (RL) a planning agent has its own representation of the environment as a model. To discover an optimal policy to interact with the environment, the agent collects experience in a trial and error fashion. Experience can be used for learning a better model or improve directly the value function and policy. Typically separated, Dyna-Q is an hybrid approach which, at each iteration, exploits the real experience to update the model as well as the value function, while planning its action using simulated data from its model. However, the planning process is computationally expensive and strongly depends on the dimensionality of the state-action space. We propose to build a Upper Confidence Tree (UCT) on the simulated experience and search for the best action to be selected during the on-line learning process. We prove the effectiveness of our proposed method on a set of preliminary tests on three testbed environments from Open AI. In contrast to Dyna-Q, Dyna-T outperforms state-of-the-art RL agents in the stochastic environments by choosing a more robust action selection strategy.

Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2201.04502 [cs.LG]
	(or arXiv:2201.04502v2 [cs.LG] for this version)

Submission history

From: Claudio Zito [view email]
[v1] Wed, 12 Jan 2022 15:06:30 GMT (1559kb,D)
[v2] Wed, 19 Jan 2022 12:02:55 GMT (1559kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2201.04502

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Machine Learning

Title: Dyna-T: Dyna-Q and Upper Confidence Bounds Applied to Trees

Submission history