Current browse context:
stat.ML
Change to browse by:
References & Citations
Statistics > Machine Learning
Title: To Bag is to Prune
(Submitted on 17 Aug 2020 (this version), latest version 8 Jun 2021 (v4))
Abstract: It is notoriously hard to build a bad Random Forest (RF). Concurrently, RF is perhaps the only standard ML algorithm that blatantly overfits in-sample without any consequence out-of-sample. Standard arguments cannot rationalize this paradox. I propose a new explanation: bootstrap aggregation and model perturbation as implemented by RF automatically prune a (latent) true underlying tree. More generally, there is no need to tune the stopping point of a properly randomized ensemble of greedily optimized base learners. Thus, Boosting and MARS are eligible. I empirically demonstrate the property with simulations and real data by reporting that these new ensembles yield equivalent performance to their tuned counterparts.
Submission history
From: Philippe Goulet Coulombe [view email][v1] Mon, 17 Aug 2020 02:45:32 GMT (446kb,D)
[v2] Mon, 14 Sep 2020 04:10:02 GMT (460kb,D)
[v3] Fri, 5 Mar 2021 16:54:07 GMT (842kb,D)
[v4] Tue, 8 Jun 2021 21:54:35 GMT (1061kb,D)
Link back to: arXiv, form interface, contact.