References & Citations
Computer Science > Computation and Language
Title: Short Text Language Identification for Under Resourced Languages
(Submitted on 18 Nov 2019 (v1), last revised 22 Nov 2019 (this version, v2))
Abstract: The paper presents a hierarchical naive Bayesian and lexicon based classifier for short text language identification (LID) useful for under resourced languages. The algorithm is evaluated on short pieces of text for the 11 official South African languages some of which are similar languages. The algorithm is compared to recent approaches using test sets from previous works on South African languages as well as the Discriminating between Similar Languages (DSL) shared tasks' datasets. Remaining research opportunities and pressing concerns in evaluating and comparing LID approaches are also discussed.
Submission history
From: Bernardt Duvenhage [view email][v1] Mon, 18 Nov 2019 11:34:38 GMT (15kb)
[v2] Fri, 22 Nov 2019 04:53:48 GMT (15kb)
Link back to: arXiv, form interface, contact.