We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:


Current browse context:


Change to browse by:


References & Citations

DBLP - CS Bibliography


(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo ScienceWISE logo

Computer Science > Computation and Language

Title: Predicting Themes within Complex Unstructured Texts: A Case Study on Safeguarding Reports

Abstract: The task of text and sentence classification is associated with the need for large amounts of labelled training data. The acquisition of high volumes of labelled datasets can be expensive or unfeasible, especially for highly-specialised domains for which documents are hard to obtain. Research on the application of supervised classification based on small amounts of training data is limited. In this paper, we address the combination of state-of-the-art deep learning and classification methods and provide an insight into what combination of methods fit the needs of small, domain-specific, and terminologically-rich corpora. We focus on a real-world scenario related to a collection of safeguarding reports comprising learning experiences and reflections on tackling serious incidents involving children and vulnerable adults. The relatively small volume of available reports and their use of highly domain-specific terminology makes the application of automated approaches difficult. We focus on the problem of automatically identifying the main themes in a safeguarding report using supervised classification approaches. Our results show the potential of deep learning models to simulate subject-expert behaviour even for complex tasks with limited labelled data.
Comments: 10 pages, 5 figures, workshop
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2010.14584 [cs.CL]
  (or arXiv:2010.14584v2 [cs.CL] for this version)

Submission history

From: Aleksandra Edwards Mrs [view email]
[v1] Tue, 27 Oct 2020 19:48:23 GMT (4394kb,D)
[v2] Thu, 29 Oct 2020 09:15:14 GMT (4393kb,D)

Link back to: arXiv, form interface, contact.