We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

stat.ME

Change to browse by:

References & Citations

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Statistics > Methodology

Title: Nearest Neighbor Imputation for Categorical Data by Weighting of Attributes

Abstract: Missing values are a common phenomenon in all areas of applied research. While various imputation methods are available for metrically scaled variables, methods for categorical data are scarce. An imputation method that has been shown to work well for high dimensional metrically scaled variables is the imputation by nearest neighbor methods. In this paper, we extend the weighted nearest neighbors approach to impute missing values in categorical variables. The proposed method, called $\mathtt{wNNSel_{cat}}$, explicitly uses the information on association among attributes. The performance of different imputation methods is compared in terms of the proportion of falsely imputed values. Simulation results show that the weighting of attributes yields smaller imputation errors than existing approaches. A variety of real data sets is used to support the results obtained by simulations.
Subjects: Methodology (stat.ME)
Cite as: arXiv:1710.01011 [stat.ME]
  (or arXiv:1710.01011v1 [stat.ME] for this version)

Submission history

From: Shahla Faisal [view email]
[v1] Tue, 3 Oct 2017 07:18:42 GMT (82kb,D)

Link back to: arXiv, form interface, contact.