We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:


Current browse context:


Change to browse by:

References & Citations


(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo ScienceWISE logo

Mathematics > Statistics Theory

Title: Inefficiency of Data Augmentation for Large Sample Imbalanced Data

Abstract: Many modern applications collect large sample size and highly imbalanced categorical data, with some categories being relatively rare. Bayesian hierarchical models are well motivated in such settings in providing an approach to borrow information to combat data sparsity, while quantifying uncertainty in estimation. However, a fundamental problem is scaling up posterior computation to massive sample sizes. In categorical data models, posterior computation commonly relies on data augmentation Gibbs sampling. In this article, we study computational efficiency of such algorithms in a large sample imbalanced regime, showing that mixing is extremely poor, with a spectral gap that converges to zero at a rate proportional to the square root of sample size or faster. This theoretical result is verified with empirical performance in simulations and an application to a computational advertising data set. In contrast, algorithms that bypass data augmentation show rapid mixing on the same dataset.
Subjects: Statistics Theory (math.ST); Computational Complexity (cs.CC); Computation (stat.CO)
MSC classes: 62
Cite as: arXiv:1605.05798 [math.ST]
  (or arXiv:1605.05798v1 [math.ST] for this version)

Submission history

From: James Johndrow [view email]
[v1] Thu, 19 May 2016 02:45:46 GMT (271kb,D)
[v2] Mon, 26 Jun 2017 15:06:27 GMT (513kb,D)

Link back to: arXiv, form interface, contact.