Distribution Compression in Near-linear Time

Shetty, Abhishek; Dwivedi, Raaz; Mackey, Lester

Full-text links:

Download:

Current browse context:

stat.ML

< prev | next >

new | recent | 2111

Statistics > Machine Learning

Title: Distribution Compression in Near-linear Time

Authors: Abhishek Shetty, Raaz Dwivedi, Lester Mackey

(Submitted on 15 Nov 2021 (v1), last revised 18 Oct 2022 (this version, v6))

Abstract: In distribution compression, one aims to accurately summarize a probability distribution $\mathbb{P}$ using a small number of representative points. Near-optimal thinning procedures achieve this goal by sampling $n$ points from a Markov chain and identifying $\sqrt{n}$ points with $\widetilde{\mathcal{O}}(1/\sqrt{n})$ discrepancy to $\mathbb{P}$. Unfortunately, these algorithms suffer from quadratic or super-quadratic runtime in the sample size $n$. To address this deficiency, we introduce Compress++, a simple meta-procedure for speeding up any thinning algorithm while suffering at most a factor of $4$ in error. When combined with the quadratic-time kernel halving and kernel thinning algorithms of Dwivedi and Mackey (2021), Compress++ delivers $\sqrt{n}$ points with $\mathcal{O}(\sqrt{\log n/n})$ integration error and better-than-Monte-Carlo maximum mean discrepancy in $\mathcal{O}(n \log^3 n)$ time and $\mathcal{O}( \sqrt{n} \log^2 n )$ space. Moreover, Compress++ enjoys the same near-linear runtime given any quadratic-time input and reduces the runtime of super-quadratic algorithms by a square-root factor. In our benchmarks with high-dimensional Monte Carlo samples and Markov chains targeting challenging differential equation posteriors, Compress++ matches or nearly matches the accuracy of its input algorithm in orders of magnitude less time.

Comments:	Accepted to ICLR 2022; An outdated proof of Theorem 2 was previously included in the appendix; this oversight is corrected in this version
Subjects:	Machine Learning (stat.ML); Data Structures and Algorithms (cs.DS); Machine Learning (cs.LG); Statistics Theory (math.ST); Methodology (stat.ME)
Cite as:	arXiv:2111.07941 [stat.ML]
	(or arXiv:2111.07941v6 [stat.ML] for this version)

Submission history

From: Lester Mackey [view email]
[v1] Mon, 15 Nov 2021 17:42:57 GMT (730kb,D)
[v2] Wed, 17 Nov 2021 01:49:21 GMT (734kb,D)
[v3] Thu, 24 Mar 2022 22:46:34 GMT (792kb,D)
[v4] Tue, 14 Jun 2022 12:36:23 GMT (789kb,D)
[v5] Tue, 13 Sep 2022 17:57:45 GMT (789kb,D)
[v6] Tue, 18 Oct 2022 01:29:37 GMT (789kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> stat > arXiv:2111.07941

Download:

Current browse context:

Change to browse by:

References & Citations

Bookmark

Statistics > Machine Learning

Title: Distribution Compression in Near-linear Time

Submission history