We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

cs.CL

Change to browse by:

cs

References & Citations

DBLP - CS Bibliography

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Computer Science > Computation and Language

Title: Normalizador Neural de Datas e Endereços

Abstract: Documents of any kind present a wide variety of date and address formats, in some cases dates can be written entirely in full or even have different types of separators. The pattern disorder in addresses is even greater due to the greater possibility of interchanging between streets, neighborhoods, cities and states. In the context of natural language processing, problems of this nature are handled by rigid tools such as ReGex or DateParser, which are efficient as long as the expected input is pre-configured. When these algorithms are given an unexpected format, errors and unwanted outputs happen. To circumvent this challenge, we present a solution with deep neural networks state of art T5 that treats non-preconfigured formats of dates and addresses with accuracy above 90% in some cases. With this model, our proposal brings generalization to the task of normalizing dates and addresses. We also deal with this problem with noisy data that simulates possible errors in the text.
Comments: 7 pages, in Portuguese, 5 tables
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2007.04300 [cs.CL]
  (or arXiv:2007.04300v2 [cs.CL] for this version)

Submission history

From: Paulo Finardi [view email]
[v1] Sat, 27 Jun 2020 20:24:35 GMT (9kb)
[v2] Thu, 9 Jul 2020 01:03:47 GMT (9kb)

Link back to: arXiv, form interface, contact.