Fairness in Learning: Classic and Contextual Bandits

Joseph, Matthew; Kearns, Michael; Morgenstern, Jamie; Roth, Aaron

Full-text links:

Download:

Current browse context:

stat

< prev | next >

new | recent | 1605

Computer Science > Machine Learning

Title: Fairness in Learning: Classic and Contextual Bandits

Authors: Matthew Joseph, Michael Kearns, Jamie Morgenstern, Aaron Roth

(Submitted on 23 May 2016 (v1), last revised 7 Nov 2016 (this version, v2))

Abstract: We introduce the study of fairness in multi-armed bandit problems. Our fairness definition can be interpreted as demanding that given a pool of applicants (say, for college admission or mortgages), a worse applicant is never favored over a better one, despite a learning algorithm's uncertainty over the true payoffs. We prove results of two types.
First, in the important special case of the classic stochastic bandits problem (i.e., in which there are no contexts), we provide a provably fair algorithm based on "chained" confidence intervals, and provide a cumulative regret bound with a cubic dependence on the number of arms. We further show that any fair algorithm must have such a dependence. When combined with regret bounds for standard non-fair algorithms such as UCB, this proves a strong separation between fair and unfair learning, which extends to the general contextual case.
In the general contextual case, we prove a tight connection between fairness and the KWIK (Knows What It Knows) learning model: a KWIK algorithm for a class of functions can be transformed into a provably fair contextual bandit algorithm, and conversely any fair contextual bandit algorithm can be transformed into a KWIK learning algorithm. This tight connection allows us to provide a provably fair algorithm for the linear contextual bandit problem with a polynomial dependence on the dimension, and to show (for a different class of functions) a worst-case exponential gap in regret between fair and non-fair learning algorithms

Comments:	A condensed version of this work appears in the 30th Annual Conference on Neural Information Processing Systems (NIPS), 2016
Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:1605.07139 [cs.LG]
	(or arXiv:1605.07139v2 [cs.LG] for this version)

Submission history

From: Matthew Joseph [view email]
[v1] Mon, 23 May 2016 18:58:24 GMT (511kb,D)
[v2] Mon, 7 Nov 2016 15:49:05 GMT (150kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:1605.07139

Download:

Current browse context:

Change to browse by:

References & Citations

Bookmark

Computer Science > Machine Learning

Title: Fairness in Learning: Classic and Contextual Bandits

Submission history