We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

eess.AS

Change to browse by:

References & Citations

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Low-Latency Online Speaker Diarization with Graph-Based Label Generation

Abstract: This paper introduces an online speaker diarization system that can handle long-time audio with low latency. We enable Agglomerative Hierarchy Clustering (AHC) to work in an online fashion by introducing a label matching algorithm. This algorithm solves the inconsistency between output labels and hidden labels that are generated each turn. To ensure the low latency in the online setting, we introduce a variant of AHC, namely chkpt-AHC, to cluster the speakers. In addition, we propose a speaker embedding graph to exploit a graph-based re-clustering method, further improving the performance. In the experiment, we evaluate our systems on both DIHARD3 and VoxConverse datasets. The experimental results show that our proposed online systems have better performance than our baseline online system and have comparable performance to our offline systems. We find out that the framework combining the chkpt-AHC method and the label matching algorithm works well in the online setting. Moreover, the chkpt-AHC method greatly reduces the time cost, while the graph-based re-clustering method helps improve the performance.
Comments: accepted by Odyssey 2022
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
Cite as: arXiv:2111.13803 [eess.AS]
  (or arXiv:2111.13803v4 [eess.AS] for this version)

Submission history

From: Yucong Zhang [view email]
[v1] Sat, 27 Nov 2021 03:34:34 GMT (3404kb,D)
[v2] Sun, 27 Feb 2022 07:17:26 GMT (4610kb,D)
[v3] Fri, 4 Mar 2022 14:25:48 GMT (4592kb,D)
[v4] Fri, 24 Jun 2022 06:21:28 GMT (4694kb,D)

Link back to: arXiv, form interface, contact.