Canonical Cortical Graph Neural Networks and its Application for Speech Enhancement in Future Audio-Visual Hearing Aids

Passos, Leandro A.; Papa, João Paulo; Hussain, Amir; Adeel, Ahsan

Full-text links:

Download:

Current browse context:

cs.SD

< prev | next >

new | recent | 2206

Computer Science > Sound

Title: Canonical Cortical Graph Neural Networks and its Application for Speech Enhancement in Future Audio-Visual Hearing Aids

Authors: Leandro A. Passos, João Paulo Papa, Amir Hussain, Ahsan Adeel

(Submitted on 6 Jun 2022 (v1), revised 22 Nov 2022 (this version, v2), latest version 31 Jan 2023 (v3))

Abstract: Despite the recent success of machine learning algorithms, most models still face several drawbacks when considering more complex tasks requiring interaction between different sources, such as multimodal input data and logical time sequence. On the other hand, the biological brain is highly sharpened in this sense, empowered to automatically manage and integrate such a stream of information through millions of years of evolution. In this context, this paper finds inspiration on recent discoveries on cortical circuits on the brain to propose a more biologically plausible self-supervised machine learning approach that combines multimodal information using intra-layer modulations together with Canonical Correlation Analysis, and a memory mechanism to keep track of temporal data, the so-called Canonical Cortical Graph Neural networks. The approach outperformed recent state-of-the-art results considering clean audio reconstruction and energy efficiency, described by a reduced and smother neuron firing rate distribution, suggesting the model as a suitable approach for speech enhancement in audio-visual hearing aid devices.

Subjects:	Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2206.02671 [cs.SD]
	(or arXiv:2206.02671v2 [cs.SD] for this version)

Submission history

From: Leandro Passos [view email]
[v1] Mon, 6 Jun 2022 15:20:07 GMT (1293kb)
[v2] Tue, 22 Nov 2022 11:21:19 GMT (2002kb)
[v3] Tue, 31 Jan 2023 14:14:49 GMT (2002kb)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2206.02671v2

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Sound

Title: Canonical Cortical Graph Neural Networks and its Application for Speech Enhancement in Future Audio-Visual Hearing Aids

Submission history