We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

cs.CL

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Computer Science > Computation and Language

Title: DiSCoMaT: Distantly Supervised Composition Extraction from Tables in Materials Science Articles

Abstract: A crucial component in the curation of KB for a scientific domain (e.g., materials science, foods & nutrition, fuels) is information extraction from tables in the domain's published research articles. To facilitate research in this direction, we define a novel NLP task of extracting compositions of materials (e.g., glasses) from tables in materials science papers. The task involves solving several challenges in concert, such as tables that mention compositions have highly varying structures; text in captions and full paper needs to be incorporated along with data in tables; and regular languages for numbers, chemical compounds and composition expressions must be integrated into the model. We release a training dataset comprising 4,408 distantly supervised tables, along with 1,475 manually annotated dev and test tables. We also present a strong baseline DISCOMAT, that combines multiple graph neural networks with several task-specific regular expressions, features, and constraints. We show that DISCOMAT outperforms recent table processing architectures by significant margins.
Comments: Accepted long paper at ACL 2023 (this https URL)
Subjects: Computation and Language (cs.CL); Materials Science (cond-mat.mtrl-sci); Information Retrieval (cs.IR)
Cite as: arXiv:2207.01079 [cs.CL]
  (or arXiv:2207.01079v4 [cs.CL] for this version)

Submission history

From: N M Anoop Krishnan [view email]
[v1] Sun, 3 Jul 2022 17:11:17 GMT (4026kb,D)
[v2] Sun, 10 Jul 2022 08:19:26 GMT (4014kb,D)
[v3] Sat, 24 Jun 2023 11:55:56 GMT (6076kb,D)
[v4] Sun, 28 Jan 2024 21:14:26 GMT (6213kb,D)

Link back to: arXiv, form interface, contact.