Current browse context:
cs.LG
Change to browse by:
References & Citations
Computer Science > Machine Learning
Title: DocParser: Hierarchical Structure Parsing of Document Renderings
(Submitted on 5 Nov 2019 (v1), last revised 25 Jan 2021 (this version, v2))
Abstract: Translating renderings (e. g. PDFs, scans) into hierarchical document structures is extensively demanded in the daily routines of many real-world applications. However, a holistic, principled approach to inferring the complete hierarchical structure of documents is missing. As a remedy, we developed "DocParser": an end-to-end system for parsing the complete document structure - including all text elements, nested figures, tables, and table cell structures. Our second contribution is to provide a dataset for evaluating hierarchical document structure parsing. Our third contribution is to propose a scalable learning framework for settings where domain-specific data are scarce, which we address by a novel approach to weak supervision that significantly improves the document structure parsing performance. Our experiments confirm the effectiveness of our proposed weak supervision: Compared to the baseline without weak supervision, it improves the mean average precision for detecting document entities by 39.1 % and improves the F1 score of classifying hierarchical relations by 35.8 %.
Submission history
From: Johannes Rausch [view email][v1] Tue, 5 Nov 2019 10:42:08 GMT (717kb,D)
[v2] Mon, 25 Jan 2021 11:54:38 GMT (9556kb,D)
Link back to: arXiv, form interface, contact.