We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

cs.CL

Change to browse by:

cs

References & Citations

DBLP - CS Bibliography

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Computer Science > Computation and Language

Title: Concept Extraction Using Pointer-Generator Networks

Abstract: Concept extraction is crucial for a number of downstream applications. However, surprisingly enough, straightforward single token/nominal chunk-concept alignment or dictionary lookup techniques such as DBpedia Spotlight still prevail. We propose a generic open-domain OOV-oriented extractive model that is based on distant supervision of a pointer-generator network leveraging bidirectional LSTMs and a copy mechanism. The model has been trained on a large annotated corpus compiled specifically for this task from 250K Wikipedia pages, and tested on regular pages, where the pointers to other pages are considered as ground truth concepts. The outcome of the experiments shows that our model significantly outperforms standard techniques and, when used on top of DBpedia Spotlight, further improves its performance. The experiments furthermore show that the model can be readily ported to other datasets on which it equally achieves a state-of-the-art performance.
Comments: Contribution to the Proceedings of the 22nd International Conference on Knowledge Engineering and Knowledge Management (EKAW 2020). A link to the final authenticated publication will be added once it is available online. Keywords: Open-domain discourse texts, Concept extraction, Pointer-generator neural network, Distant supervision
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2008.11295 [cs.CL]
  (or arXiv:2008.11295v1 [cs.CL] for this version)

Submission history

From: Alexander Shvets [view email]
[v1] Tue, 25 Aug 2020 22:28:14 GMT (1078kb,D)

Link back to: arXiv, form interface, contact.