Dual Instrumental Method for Confounded Kernelized Bandits

Gong, Xueping; Zhang, Jiheng

Full-text links:

Download:

Current browse context:

cs.LG

< prev | next >

new | recent | 2209

Computer Science > Machine Learning

Title: Dual Instrumental Method for Confounded Kernelized Bandits

Authors: Xueping Gong, Jiheng Zhang

(Submitted on 7 Sep 2022)

Abstract: The contextual bandit problem is a theoretically justified framework with wide applications in various fields. While the previous study on this problem usually requires independence between noise and contexts, our work considers a more sensible setting where the noise becomes a latent confounder that affects both contexts and rewards. Such a confounded setting is more realistic and could expand to a broader range of applications. However, the unresolved confounder will cause a bias in reward function estimation and thus lead to a large regret. To deal with the challenges brought by the confounder, we apply the dual instrumental variable regression, which can correctly identify the true reward function. We prove the convergence rate of this method is near-optimal in two types of widely used reproducing kernel Hilbert spaces. Therefore, we can design computationally efficient and regret-optimal algorithms based on the theoretical guarantees for confounded bandit problems. The numerical results illustrate the efficacy of our proposed algorithms in the confounded bandit setting.

Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Machine Learning (stat.ML)
Cite as:	arXiv:2209.03224 [cs.LG]
	(or arXiv:2209.03224v1 [cs.LG] for this version)

Submission history

From: Xueping Gong [view email]
[v1] Wed, 7 Sep 2022 15:25:57 GMT (592kb)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2209.03224

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Machine Learning

Title: Dual Instrumental Method for Confounded Kernelized Bandits

Submission history