Once-for-All Sequence Compression for Self-Supervised Speech Models

Chen, Hsuan-Jui; Meng, Yen; Lee, Hung-yi

Full-text links:

Download:

Current browse context:

cs.CL

< prev | next >

new | recent | 2211

Computer Science > Computation and Language

Title: Once-for-All Sequence Compression for Self-Supervised Speech Models

Authors: Hsuan-Jui Chen, Yen Meng, Hung-yi Lee

(Submitted on 4 Nov 2022 (v1), revised 3 Jan 2023 (this version, v2), latest version 9 May 2023 (v4))

Abstract: The sequence length along the time axis is often the dominant factor of the computational cost of self-supervised speech models. Works have been proposed to reduce the sequence length for lowering the computational cost. However, different downstream tasks have different tolerance of sequence compressing, so a model that produces a fixed compressing rate may not fit all tasks. In this work, we introduce a once-for-all (OFA) sequence compression framework for self-supervised speech models that supports a continuous range of compressing rates. The framework is evaluated on various tasks, showing marginal degradation compared to the fixed compressing rate variants with a smooth performance-efficiency trade-off. We further explore adaptive compressing rate learning, demonstrating the ability to select task-specific preferred frame periods without needing a grid search.

Comments:	Submitted to ICASSP 2023
Subjects:	Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2211.02332 [cs.CL]
	(or arXiv:2211.02332v2 [cs.CL] for this version)

Submission history

From: Hsuan-Jui Chen [view email]
[v1] Fri, 4 Nov 2022 09:19:13 GMT (2196kb,D)
[v2] Tue, 3 Jan 2023 07:55:33 GMT (2749kb,D)
[v3] Wed, 15 Mar 2023 07:07:24 GMT (2237kb,D)
[v4] Tue, 9 May 2023 11:14:52 GMT (2233kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2211.02332v2

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computation and Language

Title: Once-for-All Sequence Compression for Self-Supervised Speech Models

Submission history