Transformer over Pre-trained Transformer for Neural Text Segmentation with Enhanced Topic Coherence

Lo, Kelvin; Jin, Yuan; Tan, Weicong; Liu, Ming; Du, Lan; Buntine, Wray

Full-text links:

Download:

Current browse context:

cs.CL

< prev | next >

new | recent | 2110

Change to browse by:

Computer Science > Computation and Language

Title: Transformer over Pre-trained Transformer for Neural Text Segmentation with Enhanced Topic Coherence

Authors: Kelvin Lo, Yuan Jin, Weicong Tan, Ming Liu, Lan Du, Wray Buntine

(Submitted on 14 Oct 2021)

Abstract: This paper proposes a transformer over transformer framework, called Transformer$^2$, to perform neural text segmentation. It consists of two components: bottom-level sentence encoders using pre-trained transformers, and an upper-level transformer-based segmentation model based on the sentence embeddings. The bottom-level component transfers the pre-trained knowledge learned from large external corpora under both single and pair-wise supervised NLP tasks to model the sentence embeddings for the documents. Given the sentence embeddings, the upper-level transformer is trained to recover the segmentation boundaries as well as the topic labels of each sentence. Equipped with a multi-task loss and the pre-trained knowledge, Transformer$^2$ can better capture the semantic coherence within the same segments. Our experiments show that (1) Transformer$^2$ manages to surpass state-of-the-art text segmentation models in terms of a commonly-used semantic coherence measure; (2) in most cases, both single and pair-wise pre-trained knowledge contribute to the model performance; (3) bottom-level sentence encoders pre-trained on specific languages yield better performance than those pre-trained on specific domains.

Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2110.07160 [cs.CL]
	(or arXiv:2110.07160v1 [cs.CL] for this version)

Submission history

From: Lan Du [view email]
[v1] Thu, 14 Oct 2021 05:26:39 GMT (1230kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2110.07160

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computation and Language

Title: Transformer over Pre-trained Transformer for Neural Text Segmentation with Enhanced Topic Coherence

Submission history