Task-Guided Inverse Reinforcement Learning Under Partial Information

Djeumou, Franck; Cubuktepe, Murat; Lennon, Craig; Topcu, Ufuk

Full-text links:

Download:

Current browse context:

cs.LG

< prev | next >

new | recent | 2105

Computer Science > Machine Learning

Title: Task-Guided Inverse Reinforcement Learning Under Partial Information

Authors: Franck Djeumou, Murat Cubuktepe, Craig Lennon, Ufuk Topcu

(Submitted on 28 May 2021 (v1), last revised 16 Dec 2021 (this version, v2))

Abstract: We study the problem of inverse reinforcement learning (IRL), where the learning agent recovers a reward function using expert demonstrations. Most of the existing IRL techniques make the often unrealistic assumption that the agent has access to full information about the environment. We remove this assumption by developing an algorithm for IRL in partially observable Markov decision processes (POMDPs). The algorithm addresses several limitations of existing techniques that do not take the information asymmetry between the expert and the learner into account. First, it adopts causal entropy as the measure of the likelihood of the expert demonstrations as opposed to entropy in most existing IRL techniques, and avoids a common source of algorithmic complexity. Second, it incorporates task specifications expressed in temporal logic into IRL. Such specifications may be interpreted as side information available to the learner a priori in addition to the demonstrations and may reduce the information asymmetry. Nevertheless, the resulting formulation is still nonconvex due to the intrinsic nonconvexity of the so-called forward problem, i.e., computing an optimal policy given a reward function, in POMDPs. We address this nonconvexity through sequential convex programming and introduce several extensions to solve the forward problem in a scalable manner. This scalability allows computing policies that incorporate memory at the expense of added computational cost yet also outperform memoryless policies. We demonstrate that, even with severely limited data, the algorithm learns reward functions and policies that satisfy the task and induce a similar behavior to the expert by leveraging the side information and incorporating memory into the policy.

Comments:	Initial submission to ICAPS 2022
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Optimization and Control (math.OC)
Cite as:	arXiv:2105.14073 [cs.LG]
	(or arXiv:2105.14073v2 [cs.LG] for this version)

Submission history

From: Franck Djeumou [view email]
[v1] Fri, 28 May 2021 19:36:54 GMT (172kb)
[v2] Thu, 16 Dec 2021 20:25:14 GMT (500kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2105.14073

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Machine Learning

Title: Task-Guided Inverse Reinforcement Learning Under Partial Information

Submission history