We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

cs.CL

Change to browse by:

cs

References & Citations

DBLP - CS Bibliography

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo ScienceWISE logo

Computer Science > Computation and Language

Title: MANorm: A Normalization Dictionary for Moroccan Arabic Dialect Written in Latin Script

Abstract: Social media user-generated text is actually the main resource for many NLP tasks. This text however, does not follow the standard rules of writing. Moreover, the use of dialect such as Moroccan Arabic in written communications increases further NLP tasks complexity. A dialect is a verbal language that does not have a standard orthography, which leads users to improvise spelling while writing. Thus, for the same word we can find multiple forms of transliterations. Subsequently, it is mandatory to normalize these different transliterations to one canonical word form. To reach this goal, we have exploited the powerfulness of word embedding models generated with a corpus of YouTube comments. Besides, using a Moroccan Arabic dialect dictionary that provides the canonical forms, we have built a normalization dictionary that we refer to as MANorm. We have conducted several experiments to demonstrate the efficiency of MANorm, which have shown its usefulness in dialect normalization.
Comments: The Fifth Arabic Natural Language Processing Workshop/COLING 2020
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2206.09167 [cs.CL]
  (or arXiv:2206.09167v1 [cs.CL] for this version)

Submission history

From: Randa Zarnoufi [view email]
[v1] Sat, 18 Jun 2022 10:17:46 GMT (613kb)

Link back to: arXiv, form interface, contact.