Explainability of Machine Learning Approaches in Forensic Linguistics: A Case Study in Geolinguistic Authorship Profiling

Roemling, Dana; Scherrer, Yves; Miletic, Aleksandra

Full-text links:

Download:

Current browse context:

cs.CL

< prev | next >

new | recent | 2404

Change to browse by:

Computer Science > Computation and Language

Title: Explainability of Machine Learning Approaches in Forensic Linguistics: A Case Study in Geolinguistic Authorship Profiling

Authors: Dana Roemling, Yves Scherrer, Aleksandra Miletic

(Submitted on 29 Apr 2024)

Abstract: Forensic authorship profiling uses linguistic markers to infer characteristics about an author of a text. This task is paralleled in dialect classification, where a prediction is made about the linguistic variety of a text based on the text itself. While there have been significant advances in the last years in variety classification (Jauhiainen et al., 2019) and state-of-the-art approaches reach accuracies of up to 100% depending on the similarity of varieties and the scope of prediction (e.g., Milne et al., 2012; Blodgett et al., 2017), forensic linguistics rarely relies on these approaches due to their lack of transparency (see Nini, 2023), amongst other reasons. In this paper we therefore explore explainability of machine learning approaches considering the forensic context. We focus on variety classification as a means of geolinguistic profiling of unknown texts. For this we work with an approach proposed by Xie et al. (2024) to extract the lexical items most relevant to the variety classifications. We find that the extracted lexical features are indeed representative of their respective varieties and note that the trained models also rely on place names for classifications.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2404.18510 [cs.CL]
	(or arXiv:2404.18510v1 [cs.CL] for this version)

Submission history

From: Dana Roemling [view email]
[v1] Mon, 29 Apr 2024 08:52:52 GMT (6892kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2404.18510

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computation and Language

Title: Explainability of Machine Learning Approaches in Forensic Linguistics: A Case Study in Geolinguistic Authorship Profiling

Submission history