We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

stat.ME

Change to browse by:

References & Citations

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Statistics > Methodology

Title: Probabilistic Best Subset Selection by Gradient-Based Optimization

Abstract: In high-dimensional statistics, variable selection is an optimization problem aiming to recover the latent sparse pattern from all possible covariate combinations. In this paper, we transform the optimization problem from a discrete space to a continuous one via reparameterization. The new objective function is a reformulation of the exact $L_0$-regularized regression problem (a.k.a. best subset selection). In the framework of stochastic gradient descent, we propose a family of unbiased and efficient gradient estimators that are used to optimize the best subset selection objective and its variational lower bound. Under this family, we identify the estimator with non-vanishing signal-to-noise ratio and uniformly minimum variance. Theoretically we study the general conditions under which the method is guaranteed to converge to the ground truth in expectation. In a wide variety of synthetic and real data sets, the proposed method outperforms existing ones based on penalized regression or best subset selection, in both sparse pattern recovery and out-of-sample prediction. Our method can find the true regression model from thousands of covariates in a couple of seconds.
Subjects: Methodology (stat.ME)
Cite as: arXiv:2006.06448 [stat.ME]
  (or arXiv:2006.06448v1 [stat.ME] for this version)

Submission history

From: Mingzhang Yin [view email]
[v1] Thu, 11 Jun 2020 13:57:29 GMT (217kb,D)
[v2] Mon, 22 Jun 2020 18:28:46 GMT (217kb,D)
[v3] Fri, 7 Aug 2020 04:23:34 GMT (154kb,D)
[v4] Wed, 1 Jun 2022 01:59:07 GMT (215kb,D)

Link back to: arXiv, form interface, contact.