We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

stat

Change to browse by:

References & Citations

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Mathematics > Statistics Theory

Title: Estimation and Inference with Trees and Forests in High Dimensions

Abstract: We analyze the finite sample mean squared error (MSE) performance of regression trees and forests in the high dimensional regime with binary features, under a sparsity constraint. We prove that if only $r$ of the $d$ features are relevant for the mean outcome function, then shallow trees built greedily via the CART empirical MSE criterion achieve MSE rates that depend only logarithmically on the ambient dimension $d$. We prove upper bounds, whose exact dependence on the number relevant variables $r$ depends on the correlation among the features and on the degree of relevance. For strongly relevant features, we also show that fully grown honest forests achieve fast MSE rates and their predictions are also asymptotically normal, enabling asymptotically valid inference that adapts to the sparsity of the regression function.
Comments: Accepted for presentation at the Conference on Learning Theory (COLT) 2020
Subjects: Statistics Theory (math.ST); Machine Learning (cs.LG)
Cite as: arXiv:2007.03210 [math.ST]
  (or arXiv:2007.03210v2 [math.ST] for this version)

Submission history

From: Emmanouil Zampetakis [view email]
[v1] Tue, 7 Jul 2020 05:45:32 GMT (53kb)
[v2] Wed, 21 Oct 2020 18:44:09 GMT (58kb)

Link back to: arXiv, form interface, contact.