A LayoutLMv3-Based Model for Enhanced Relation Extraction in Visually-Rich Documents

Adnan, Wiam; Tang, Joel; Zouggari, Yassine Bel Khayat; Laatiri, Seif Edinne; Lam, Laurent; Caspani, Fabien

Full-text links:

Download:

Current browse context:

cs.CL

< prev | next >

new | recent | 2404

Computer Science > Computation and Language

Title: A LayoutLMv3-Based Model for Enhanced Relation Extraction in Visually-Rich Documents

Authors: Wiam Adnan, Joel Tang, Yassine Bel Khayat Zouggari, Seif Edinne Laatiri, Laurent Lam, Fabien Caspani

(Submitted on 16 Apr 2024)

Abstract: Document Understanding is an evolving field in Natural Language Processing (NLP). In particular, visual and spatial features are essential in addition to the raw text itself and hence, several multimodal models were developed in the field of Visual Document Understanding (VDU). However, while research is mainly focused on Key Information Extraction (KIE), Relation Extraction (RE) between identified entities is still under-studied. For instance, RE is crucial to regroup entities or obtain a comprehensive hierarchy of data in a document. In this paper, we present a model that, initialized from LayoutLMv3, can match or outperform the current state-of-the-art results in RE applied to Visually-Rich Documents (VRD) on FUNSD and CORD datasets, without any specific pre-training and with fewer parameters. We also report an extensive ablation study performed on FUNSD, highlighting the great impact of certain features and modelization choices on the performances.

Comments:	Accepted at the International Conference on Document Analysis and Recognition (ICDAR 2024)
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2404.10848 [cs.CL]
	(or arXiv:2404.10848v1 [cs.CL] for this version)

Submission history

From: Joël Tang [view email]
[v1] Tue, 16 Apr 2024 18:50:57 GMT (522kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2404.10848

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computation and Language

Title: A LayoutLMv3-Based Model for Enhanced Relation Extraction in Visually-Rich Documents

Submission history