References & Citations
Computer Science > Databases
Title: Approximate Selection with Guarantees using Proxies
(Submitted on 2 Apr 2020 (v1), last revised 3 Jan 2022 (this version, v4))
Abstract: Due to the falling costs of data acquisition and storage, researchers and industry analysts often want to find all instances of rare events in large datasets. For instance, scientists can cheaply capture thousands of hours of video, but are limited by the need to manually inspect long videos to identify relevant objects and events. To reduce this cost, recent work proposes to use cheap proxy models, such as image classifiers, to identify an approximate set of data points satisfying a data selection filter. Unfortunately, this recent work does not provide the statistical accuracy guarantees necessary in scientific and production settings.
In this work, we introduce novel algorithms for approximate selection queries with statistical accuracy guarantees. Namely, given a limited number of exact identifications from an oracle, often a human or an expensive machine learning model, our algorithms meet a minimum precision or recall target with high probability. In contrast, existing approaches can catastrophically fail in satisfying these recall and precision targets. We show that our algorithms can improve query result quality by up to 30x for both the precision and recall targets in both real and synthetic datasets.
Submission history
From: Daniel Kang [view email][v1] Thu, 2 Apr 2020 05:36:10 GMT (7628kb,D)
[v2] Sun, 12 Apr 2020 19:08:32 GMT (7628kb,D)
[v3] Thu, 23 Jul 2020 20:33:35 GMT (7604kb,D)
[v4] Mon, 3 Jan 2022 21:24:54 GMT (7604kb,D)
Link back to: arXiv, form interface, contact.