Repeated undersampling in PrInDT (RePrInDT): Variation in undersampling and prediction, and ranking of predictors in ensembles

Weihs, Claus; Buschfeld, Sarah

Full-text links:

Download:

Current browse context:

stat

< prev | next >

new | recent | 2108

Statistics > Applications

Title: Repeated undersampling in PrInDT (RePrInDT): Variation in undersampling and prediction, and ranking of predictors in ensembles

Authors: Claus Weihs, Sarah Buschfeld

(Submitted on 11 Aug 2021)

Abstract: In this paper, we extend our PrInDT method (Weihs & Buschfeld 2021a) towards undersampling with different percentages of the smaller and the larger classes (psmall and plarge), stratification of predictors, varying the prediction threshold, and measuring variable importance in ensembles. An application of these methods to a linguistic example suggests the following: 1. In undersampling, a careful selection of the percentages plarge and psmall is important for building models with high balanced accuracies; 2. Stratification of predictors does not majorly enhance balanced accuracies; 3. Lowering the prediction threshold for the smaller class turns out to be an alternative method to undersampling because it increases the likelihood of the smaller class being selected. Finally, we introduce a method for ranking predictor importance that allows for a straightforward interpretation of the results.

Subjects:	Applications (stat.AP)
Cite as:	arXiv:2108.05129 [stat.AP]
	(or arXiv:2108.05129v1 [stat.AP] for this version)

Submission history

From: Claus Weihs [view email]
[v1] Wed, 11 Aug 2021 10:15:06 GMT (17kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> stat > arXiv:2108.05129

Download:

Current browse context:

Change to browse by:

References & Citations

Bookmark

Statistics > Applications

Title: Repeated undersampling in PrInDT (RePrInDT): Variation in undersampling and prediction, and ranking of predictors in ensembles

Submission history