BERT-CNN: a Hierarchical Patent Classifier Based on a Pre-Trained Language Model

Lu, Xiaolei; Ni, Bin

Full-text links:

Download:

PDF only

Current browse context:

cs.CL

< prev | next >

new | recent | 1911

Computer Science > Computation and Language

Title: BERT-CNN: a Hierarchical Patent Classifier Based on a Pre-Trained Language Model

Authors: Xiaolei Lu, Bin Ni

(Submitted on 3 Nov 2019)

Abstract: The automatic classification is a process of automatically assigning text documents to predefined categories. An accurate automatic patent classifier is crucial to patent inventors and patent examiners in terms of intellectual property protection, patent management, and patent information retrieval. We present BERT-CNN, a hierarchical patent classifier based on pre-trained language model by training the national patent application documents collected from the State Information Center, China. The experimental results show that BERT-CNN achieves 84.3% accuracy, which is far better than the two compared baseline methods, Convolutional Neural Networks and Recurrent Neural Networks. We didn't apply our model to the third and fourth hierarchical level of the International Patent Classification - "subclass" and "group".The visualization of the Attention Mechanism shows that BERT-CNN obtains new state-of-the-art results in representing vocabularies and semantics. This article demonstrates the practicality and effectiveness of BERT-CNN in the field of automatic patent classification.

Comments:	in Chinese
Subjects:	Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:1911.06241 [cs.CL]
	(or arXiv:1911.06241v1 [cs.CL] for this version)

Submission history

From: Xiaolei Lu [view email]
[v1] Sun, 3 Nov 2019 07:21:41 GMT (1305kb)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:1911.06241

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computation and Language

Title: BERT-CNN: a Hierarchical Patent Classifier Based on a Pre-Trained Language Model

Submission history