We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:


Current browse context:


Change to browse by:


References & Citations

DBLP - CS Bibliography


(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo ScienceWISE logo

Computer Science > Computation and Language

Title: Chemical Identification and Indexing in PubMed Articles via BERT and Text-to-Text Approaches

Abstract: The Biocreative VII Track-2 challenge consists of named entity recognition, entity-linking (or entity-normalization), and topic indexing tasks -- with entities and topics limited to chemicals for this challenge. Named entity recognition is a well-established problem and we achieve our best performance with BERT-based BioMegatron models. We extend our BERT-based approach to the entity linking task. After the second stage of pretraining BioBERT with a metric-learning loss strategy called self-alignment pretraining (SAP), we link entities based on the cosine similarity between their SAP-BioBERT word embeddings. Despite the success of our named entity recognition experiments, we find the chemical indexing task generally more challenging.
In addition to conventional NER methods, we attempt both named entity recognition and entity linking with a novel text-to-text or "prompt" based method that uses generative language models such as T5 and GPT. We achieve encouraging results with this new approach.
Comments: Submission to the BioCreative VII challenge - Track-2
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2111.15622 [cs.CL]
  (or arXiv:2111.15622v1 [cs.CL] for this version)

Submission history

From: Hoo Chang Shin [view email]
[v1] Tue, 30 Nov 2021 18:21:06 GMT (2781kb,D)

Link back to: arXiv, form interface, contact.