Distributed Bayesian Varying Coefficient Modeling Using a Gaussian Process Prior

Guhaniyogi, Rajarshi; Li, Cheng; Savitsky, Terrance D.; Srivastava, Sanvesh

Full-text links:

Download:

Current browse context:

stat.ME

< prev | next >

new | recent | 2006

Statistics > Methodology

Title: Distributed Bayesian Varying Coefficient Modeling Using a Gaussian Process Prior

Authors: Rajarshi Guhaniyogi, Cheng Li, Terrance D. Savitsky, Sanvesh Srivastava

(Submitted on 1 Jun 2020 (this version), latest version 26 Feb 2022 (v2))

Abstract: Varying coefficient models (VCMs) are widely used for estimating nonlinear regression functions in functional data models. Their Bayesian variants using Gaussian process (GP) priors on the functional coefficients, however, have received limited attention in massive data applications. This is primarily due to the prohibitively slow posterior computations using Markov chain Monte Carlo (MCMC) algorithms. We address this problem using a divide-and-conquer Bayesian approach that operates in three steps. The first step creates a large number of data subsets with much smaller sample sizes by sampling without replacement from the full data. The second step formulates VCM as a linear mixed-effects model and develops a data augmentation (DA)-type algorithm for obtaining MCMC draws of the parameters and predictions on all the subsets in parallel. The DA-type algorithm appropriately modifies the likelihood such that every subset posterior distribution is an accurate approximation of the corresponding true posterior distribution. The third step develops a combination algorithm for aggregating MCMC-based estimates of the subset posterior distributions into a single posterior distribution called the Aggregated Monte Carlo (AMC) posterior. Theoretically, we derive minimax optimal posterior convergence rates for the AMC posterior distributions of both the varying coefficients and the mean regression function. We provide quantification on the orders of subset sample sizes and the number of subsets according to the smoothness properties of the multivariate GP. The empirical results show that the combination schemes that satisfy our theoretical assumptions, including the one in the AMC algorithm, have better nominal coverage, shorter credible intervals, smaller mean square errors, and higher effective sample size than their main competitors across diverse simulations and in a real data analysis.

Subjects:	Methodology (stat.ME); Computation (stat.CO)
Cite as:	arXiv:2006.00783 [stat.ME]
	(or arXiv:2006.00783v1 [stat.ME] for this version)

Submission history

From: Cheng Li [view email]
[v1] Mon, 1 Jun 2020 08:16:45 GMT (2165kb,D)
[v2] Sat, 26 Feb 2022 01:43:59 GMT (2164kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> stat > arXiv:2006.00783v1

Download:

Current browse context:

Change to browse by:

References & Citations

Bookmark

Statistics > Methodology

Title: Distributed Bayesian Varying Coefficient Modeling Using a Gaussian Process Prior

Submission history