Towards unsupervised phone and word segmentation using self-supervised vector-quantized neural networks

Kamper, Herman; van Niekerk, Benjamin

Full-text links:

Download:

Current browse context:

cs.CL

< prev | next >

new | recent | 2012

Computer Science > Computation and Language

Title: Towards unsupervised phone and word segmentation using self-supervised vector-quantized neural networks

Authors: Herman Kamper, Benjamin van Niekerk

(Submitted on 14 Dec 2020 (v1), last revised 11 Jun 2021 (this version, v2))

Abstract: We investigate segmenting and clustering speech into low-bitrate phone-like sequences without supervision. We specifically constrain pretrained self-supervised vector-quantized (VQ) neural networks so that blocks of contiguous feature vectors are assigned to the same code, thereby giving a variable-rate segmentation of the speech into discrete units. Two segmentation methods are considered. In the first, features are greedily merged until a prespecified number of segments are reached. The second uses dynamic programming to optimize a squared error with a penalty term to encourage fewer but longer segments. We show that these VQ segmentation methods can be used without alteration across a wide range of tasks: unsupervised phone segmentation, ABX phone discrimination, same-different word discrimination, and as inputs to a symbolic word segmentation algorithm. The penalized dynamic programming method generally performs best. While performance on individual tasks is only comparable to the state-of-the-art in some cases, in all tasks a reasonable competing approach is outperformed at a substantially lower bitrate.

Comments:	Accepted to Interspeech 2021
Subjects:	Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2012.07551 [cs.CL]
	(or arXiv:2012.07551v2 [cs.CL] for this version)

Submission history

From: Herman Kamper [view email]
[v1] Mon, 14 Dec 2020 14:17:33 GMT (301kb,D)
[v2] Fri, 11 Jun 2021 12:12:43 GMT (301kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2012.07551

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computation and Language

Title: Towards unsupervised phone and word segmentation using self-supervised vector-quantized neural networks

Submission history