We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

cs.LG

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo ScienceWISE logo

Computer Science > Machine Learning

Title: Therapeutics Data Commons: Machine Learning Datasets and Tasks for Drug Discovery and Development

Abstract: Therapeutics machine learning is an emerging field with incredible opportunities for innovatiaon and impact. However, advancement in this field requires formulation of meaningful learning tasks and careful curation of datasets. Here, we introduce Therapeutics Data Commons (TDC), the first unifying platform to systematically access and evaluate machine learning across the entire range of therapeutics. To date, TDC includes 66 AI-ready datasets spread across 22 learning tasks and spanning the discovery and development of safe and effective medicines. TDC also provides an ecosystem of tools and community resources, including 33 data functions and types of meaningful data splits, 23 strategies for systematic model evaluation, 17 molecule generation oracles, and 29 public leaderboards. All resources are integrated and accessible via an open Python library. We carry out extensive experiments on selected datasets, demonstrating that even the strongest algorithms fall short of solving key therapeutics challenges, including real dataset distributional shifts, multi-scale modeling of heterogeneous data, and robust generalization to novel data points. We envision that TDC can facilitate algorithmic and scientific advances and considerably accelerate machine-learning model development, validation and transition into biomedical and clinical implementation. TDC is an open-science initiative available at this https URL
Comments: Published at NeurIPS 2021 Datasets and Benchmarks
Subjects: Machine Learning (cs.LG); Computers and Society (cs.CY); Biomolecules (q-bio.BM); Quantitative Methods (q-bio.QM)
Cite as: arXiv:2102.09548 [cs.LG]
  (or arXiv:2102.09548v2 [cs.LG] for this version)

Submission history

From: Kexin Huang [view email]
[v1] Thu, 18 Feb 2021 18:50:31 GMT (12366kb,D)
[v2] Sat, 28 Aug 2021 19:59:03 GMT (12921kb,D)

Link back to: arXiv, form interface, contact.