We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

stat.ML

Change to browse by:

References & Citations

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Statistics > Machine Learning

Title: High-Dimensional Multi-Task Averaging and Application to Kernel Mean Embedding

Authors: Hannah Marienwald (TUB), Jean-Baptiste Fermanian (ENS Rennes), Gilles Blanchard (DATASHAPE, LMO, CNRS)
Abstract: We propose an improved estimator for the multi-task averaging problem, whose goal is the joint estimation of the means of multiple distributions using separate, independent data sets. The naive approach is to take the empirical mean of each data set individually, whereas the proposed method exploits similarities between tasks, without any related information being known in advance. First, for each data set, similar or neighboring means are determined from the data by multiple testing. Then each naive estimator is shrunk towards the local average of its neighbors. We prove theoretically that this approach provides a reduction in mean squared error. This improvement can be significant when the dimension of the input space is large, demonstrating a "blessing of dimensionality" phenomenon. An application of this approach is the estimation of multiple kernel mean embeddings, which plays an important role in many modern applications. The theoretical results are verified on artificial and real world data.
Subjects: Machine Learning (stat.ML); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Statistics Theory (math.ST)
Cite as: arXiv:2011.06794 [stat.ML]
  (or arXiv:2011.06794v1 [stat.ML] for this version)

Submission history

From: Gilles Blanchard [view email]
[v1] Fri, 13 Nov 2020 07:31:30 GMT (59kb)

Link back to: arXiv, form interface, contact.