LDRNet: Enabling Real-time Document Localization on Mobile Devices

Wu, Han; Qian, Holland; Wu, Huaming; van Moorsel, Aad

Full-text links:

Download:

Current browse context:

cs.CV

< prev | next >

new | recent | 2206

Computer Science > Computer Vision and Pattern Recognition

Title: LDRNet: Enabling Real-time Document Localization on Mobile Devices

Authors: Han Wu, Holland Qian, Huaming Wu, Aad van Moorsel

(Submitted on 5 Jun 2022 (v1), last revised 12 Oct 2023 (this version, v3))

Abstract: While Identity Document Verification (IDV) technology on mobile devices becomes ubiquitous in modern business operations, the risk of identity theft and fraud is increasing. The identity document holder is normally required to participate in an online video interview to circumvent impostors. However, the current IDV process depends on an additional human workforce to support online step-by-step guidance which is inefficient and expensive. The performance of existing AI-based approaches cannot meet the real-time and lightweight demands of mobile devices. In this paper, we address those challenges by designing an edge intelligence-assisted approach for real-time IDV. Aiming at improving the responsiveness of the IDV process, we propose a new document localization model for mobile devices, LDRNet, to Localize the identity Document in Real-time. On the basis of a lightweight backbone network, we build three prediction branches for LDRNet, the corner points prediction, the line borders prediction and the document classification. We design novel supplementary targets, the equal-division points, and use a new loss function named Line Loss, to improve the speed and accuracy of our approach. In addition to the IDV process, LDRNet is an efficient and reliable document localization alternative for all kinds of mobile applications. As a matter of proof, we compare the performance of LDRNet with other popular approaches on localizing general document datasets. The experimental results show that LDRNet runs at a speed up to 790 FPS which is 47x faster, while still achieving comparable Jaccard Index(JI) in single-model and single-scale tests.

Comments:	ECML-PKDD 2022 this https URL
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Performance (cs.PF)
Cite as:	arXiv:2206.02136 [cs.CV]
	(or arXiv:2206.02136v3 [cs.CV] for this version)

Submission history

From: Han Wu [view email]
[v1] Sun, 5 Jun 2022 09:39:12 GMT (62620kb,D)
[v2] Thu, 13 Oct 2022 10:10:51 GMT (36238kb,D)
[v3] Thu, 12 Oct 2023 13:55:06 GMT (18120kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2206.02136

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computer Vision and Pattern Recognition

Title: LDRNet: Enabling Real-time Document Localization on Mobile Devices

Submission history