Latent Policies for Adversarial Imitation Learning

Wang, Tianyu; Karnwal, Nikhil; Atanasov, Nikolay

Full-text links:

Download:

Current browse context:

cs.LG

< prev | next >

new | recent | 2206

Computer Science > Machine Learning

Title: Latent Policies for Adversarial Imitation Learning

Authors: Tianyu Wang, Nikhil Karnwal, Nikolay Atanasov

(Submitted on 22 Jun 2022)

Abstract: This paper considers learning robot locomotion and manipulation tasks from expert demonstrations. Generative adversarial imitation learning (GAIL) trains a discriminator that distinguishes expert from agent transitions, and in turn use a reward defined by the discriminator output to optimize a policy generator for the agent. This generative adversarial training approach is very powerful but depends on a delicate balance between the discriminator and the generator training. In high-dimensional problems, the discriminator training may easily overfit or exploit associations with task-irrelevant features for transition classification. A key insight of this work is that performing imitation learning in a suitable latent task space makes the training process stable, even in challenging high-dimensional problems. We use an action encoder-decoder model to obtain a low-dimensional latent action space and train a LAtent Policy using Adversarial imitation Learning (LAPAL). The encoder-decoder model can be trained offline from state-action pairs to obtain a task-agnostic latent action representation or online, simultaneously with the discriminator and generator training, to obtain a task-aware latent action representation. We demonstrate that LAPAL training is stable, with near-monotonic performance improvement, and achieves expert performance in most locomotion and manipulation tasks, while a GAIL baseline converges slower and does not achieve expert performance in high-dimensional environments.

Comments:	8 pages, 5 figures
Subjects:	Machine Learning (cs.LG); Robotics (cs.RO)
Cite as:	arXiv:2206.11299 [cs.LG]
	(or arXiv:2206.11299v1 [cs.LG] for this version)

Submission history

From: Tianyu Wang [view email]
[v1] Wed, 22 Jun 2022 18:06:26 GMT (19436kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2206.11299

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Machine Learning

Title: Latent Policies for Adversarial Imitation Learning

Submission history