Sound

Authors and titles for cs.SD in Oct 2022, skipping first 350

[ total of 363 entries: 1-25 | ... | 276-300 | 301-325 | 326-350 | 351-363 ]
[ showing 25 entries per page: fewer | more | all ]

[351] arXiv:2210.16743 (cross-list from eess.AS) [pdf, other]: Title: WeKws: A production first small-footprint end-to-end Keyword Spotting Toolkit

Authors: Jie Wang, Menglong Xu, Jingyong Hou, Binbin Zhang, Xiao-Lei Zhang, Lei Xie, Fuping Pan

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[352] arXiv:2210.16871 (cross-list from eess.AS) [pdf, other]: Title: Improved acoustic-to-articulatory inversion using representations from pretrained self-supervised learning models

Authors: Sathvik Udupa, Siddarth C, Prasanta Kumar Ghosh

Comments: submitted to ICASSP 2023

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[353] arXiv:2210.16881 (cross-list from eess.AS) [pdf, other]: Title: Real-Time MRI Video synthesis from time aligned phonemes with sequence-to-sequence networks

Authors: Sathvik Udupa, Prasanta Kumar Ghosh

Comments: submitted to ICASSP 2023

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[354] arXiv:2210.17153 (cross-list from eess.AS) [pdf, other]: Title: The Importance of Accurate Alignments in End-to-End Speech Synthesis

Authors: Anusha Prakash, Hema A Murthy

Comments: Version 1 uploaded

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[355] arXiv:2210.17154 (cross-list from eess.AS) [pdf, other]: Title: Minimum Processing Near-end Listening Enhancement

Authors: Andreas Jonas Fuglsig, Jesper Jensen, Zheng-Hua Tan, Lars Søndergaard Bertelsen, Jens Christian Lindof, Jan Østergaard

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD); Signal Processing (eess.SP)
[356] arXiv:2210.17189 (cross-list from eess.AS) [src]: Title: DiaCorrect: End-to-end error correction for speaker diarization

Authors: Jiangyu Han, Yuhang Cao, Heng Lu, Yanhua Long

Comments: This paper has been superseded by arXiv:2309.08377 (merged from arXiv:2210.17189)

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[357] arXiv:2210.17287 (cross-list from eess.AS) [pdf, ps, other]: Title: Diffiner: A Versatile Diffusion-based Generative Refiner for Speech Enhancement

Authors: Ryosuke Sawata, Naoki Murata, Yuhta Takida, Toshimitsu Uesaka, Takashi Shibuya, Shusuke Takahashi, Yuki Mitsufuji

Comments: Accepted by Interspeech 2023

Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[358] arXiv:2210.17310 (cross-list from eess.AS) [pdf, other]: Title: Convolution-Based Channel-Frequency Attention for Text-Independent Speaker Verification

Authors: Jingyu Li, Yusheng Tian, Tan Lee

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[359] arXiv:2210.17316 (cross-list from eess.AS) [pdf, other]: Title: There is more than one kind of robustness: Fooling Whisper with adversarial examples

Authors: Raphael Olivier, Bhiksha Raj

Comments: Accepted at InterSpeech 2023

Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)
[360] arXiv:2210.17326 (cross-list from eess.AS) [pdf, other]: Title: Model Compression for DNN-based Speaker Verification Using Weight Quantization

Authors: Jingyu Li, Wei Liu, Zhaoyang Zhang, Jiong Wang, Tan Lee

Comments: Accepted by INTERSPEECH2023

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[361] arXiv:2210.17327 (cross-list from eess.AS) [pdf, other]: Title: Diffusion-based Generative Speech Source Separation

Authors: Robin Scheibler, Youna Ji, Soo-Whan Chung, Jaeuk Byun, Soyeon Choe, Min-Seok Choi

Comments: 5 pages, 3 figures, 2 tables. Submitted to ICASSP 2023

Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[362] arXiv:2210.17338 (cross-list from eess.AS) [pdf, ps, other]: Title: VoicePrivacy 2022 System Description: Speaker Anonymization with Feature-matched F0 Trajectories

Authors: Ünal Ege Gaznepoglu, Anna Leschanowsky, Nils Peters

Comments: 4 pages, 4 figures, 2 tables, submitted to VoicePrivacy Challenge 2022

Subjects: Audio and Speech Processing (eess.AS); Cryptography and Security (cs.CR); Sound (cs.SD)
[363] arXiv:2210.17456 (cross-list from eess.AS) [pdf, other]: Title: Audio-Visual Speech Enhancement and Separation by Utilizing Multi-Modal Self-Supervised Embeddings

Authors: I-Chun Chern, Kuo-Hsuan Hung, Yi-Ting Chen, Tassadaq Hussain, Mandar Gogate, Amir Hussain, Yu Tsao, Jen-Cheng Hou

Comments: ICASSP AMHAT 2023

Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)

[ total of 363 entries: 1-25 | ... | 276-300 | 301-325 | 326-350 | 351-363 ]
[ showing 25 entries per page: fewer | more | all ]

Disable MathJax (What is MathJax?)

Links to: arXiv, form interface, find, cs, 2404, contact, help (Access key information)

> cs > cs.SD

Sound

Authors and titles for cs.SD in Oct 2022, skipping first 350