Current browse context:
stat.ME
Change to browse by:
References & Citations
Statistics > Methodology
Title: Selective machine learning of doubly robust functionals
(Submitted on 5 Nov 2019 (v1), revised 29 Apr 2020 (this version, v2), latest version 3 Sep 2023 (v6))
Abstract: While model selection is a well-studied topic in parametric and nonparametric regression or density estimation, model selection of possibly high-dimensional nuisance parameters in semiparametric problems is far less developed. In this paper, we propose a selective machine learning framework for making inferences about a finite-dimensional functional defined on a semiparametric model, when the latter admits a doubly robust estimating function. We introduce two model selection criteria for bias reduction of functional of interest, each based on a novel definition of pseudo-risk for the functional that embodies this double robustness property and thus may be used to select the candidate model that is nearest to fulfilling this property even when all models are wrong. We establish an oracle property for a multi-fold cross-validation version of the new model selection criteria which states that our empirical criteria perform nearly as well as an oracle with a priori knowledge of the pseudo-risk for each candidate model. We also describe a smooth approximation to the selection criteria which allows for valid post-selection inference. Finally, we apply the approach to model selection of a semiparametric estimator of average treatment effect given an ensemble of candidate machine learners to account for confounding in an observational study.
Submission history
From: Yifan Cui [view email][v1] Tue, 5 Nov 2019 19:00:03 GMT (56kb,D)
[v2] Wed, 29 Apr 2020 03:35:30 GMT (50kb,D)
[v3] Wed, 29 Jul 2020 17:49:49 GMT (32kb,D)
[v4] Mon, 3 Aug 2020 18:43:09 GMT (27kb,D)
[v5] Mon, 12 Apr 2021 17:10:30 GMT (27kb)
[v6] Sun, 3 Sep 2023 07:36:36 GMT (30kb)
Link back to: arXiv, form interface, contact.