An Ensemble Approach to Acronym Extraction using Transformers

Sharma, Prashant; Saadany, Hadeel; Zilio, Leonardo; Kanojia, Diptesh; Orăsan, Constantin

Full-text links:

Download:

Current browse context:

cs.CL

< prev | next >

new | recent | 2201

Change to browse by:

Computer Science > Computation and Language

Title: An Ensemble Approach to Acronym Extraction using Transformers

Authors: Prashant Sharma, Hadeel Saadany, Leonardo Zilio, Diptesh Kanojia, Constantin Orăsan

(Submitted on 9 Jan 2022)

Abstract: Acronyms are abbreviated units of a phrase constructed by using initial components of the phrase in a text. Automatic extraction of acronyms from a text can help various Natural Language Processing tasks like machine translation, information retrieval, and text summarisation. This paper discusses an ensemble approach for the task of Acronym Extraction, which utilises two different methods to extract acronyms and their corresponding long forms. The first method utilises a multilingual contextual language model and fine-tunes the model to perform the task. The second method relies on a convolutional neural network architecture to extract acronyms and append them to the output of the previous method. We also augment the official training dataset with additional training samples extracted from several open-access journals to help improve the task performance. Our dataset analysis also highlights the noise within the current task dataset. Our approach achieves the following macro-F1 scores on test data released with the task: Danish (0.74), English-Legal (0.72), English-Scientific (0.73), French (0.63), Persian (0.57), Spanish (0.65), Vietnamese (0.65). We release our code and models publicly.

Comments:	Published at SDU@AAAI-22
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2201.03026 [cs.CL]
	(or arXiv:2201.03026v1 [cs.CL] for this version)

Submission history

From: Diptesh Kanojia [view email]
[v1] Sun, 9 Jan 2022 14:49:46 GMT (4620kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2201.03026

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computation and Language

Title: An Ensemble Approach to Acronym Extraction using Transformers

Submission history