References & Citations
Computer Science > Computation and Language
Title: Open Korean Corpora: A Practical Report
(Submitted on 31 Dec 2020 (v1), last revised 16 May 2023 (this version, v2))
Abstract: Korean is often referred to as a low-resource language in the research community. While this claim is partially true, it is also because the availability of resources is inadequately advertised and curated. This work curates and reviews a list of Korean corpora, first describing institution-level resource development, then further iterate through a list of current open datasets for different types of tasks. We then propose a direction on how open-source dataset construction and releases should be done for less-resourced languages to promote research.
Submission history
From: Won Ik Cho [view email][v1] Thu, 31 Dec 2020 14:23:55 GMT (43kb,D)
[v2] Tue, 16 May 2023 17:08:24 GMT (61kb,D)
Link back to: arXiv, form interface, contact.