We gratefully acknowledge support from
the Simons Foundation and member institutions.

Sound

Authors and titles for recent submissions

[ total of 83 entries: 1-33 | 34-66 | 67-83 ]
[ showing 33 entries per page: fewer | more | all ]

Fri, 7 Jun 2024

[1]  arXiv:2406.04140 [pdf, other]
Title: STraDa: A Singer Traits Dataset
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[2]  arXiv:2406.03822 [pdf, other]
Title: SilentCipher: Deep Audio Watermarking
Subjects: Sound (cs.SD); Cryptography and Security (cs.CR); Audio and Speech Processing (eess.AS)
[3]  arXiv:2406.03714 [pdf, other]
Title: Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
Comments: Accepted by Interspeech 2024
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[4]  arXiv:2406.03706 [pdf, other]
Title: Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
Comments: Accepted by Interspeech 2024
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[5]  arXiv:2406.03512 [pdf, other]
Title: Harder or Different? Understanding Generalization of Audio Deepfake Detection
Journal-ref: Interspeech 2024
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[6]  arXiv:2406.03510 [pdf, other]
Title: Speech-based Clinical Depression Screening: An Empirical Study
Comments: 5 pages, 3 figures
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[7]  arXiv:2406.04321 (cross-list from cs.CV) [pdf, other]
Title: VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling
Comments: The code and datasets will be available at this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD)
[8]  arXiv:2406.04269 (cross-list from eess.AS) [pdf, other]
Title: Beyond Performance Plateaus: A Comprehensive Study on Scalability in Speech Enhancement
Comments: 5 pages, 3 figures, 4 tables, Accepted by Interspeech 2024
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[9]  arXiv:2406.04212 (cross-list from eess.AS) [pdf, ps, other]
Title: Sound Event Bounding Boxes
Comments: Accepted for publication at Interspeech 2024
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[10]  arXiv:2406.03882 (cross-list from cs.CL) [pdf, other]
Title: Spontaneous Speech-Based Suicide Risk Detection Using Whisper and Large Language Models
Comments: Accepted by Interspeech 2024
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[11]  arXiv:2406.03872 (cross-list from cs.CL) [pdf, other]
Title: BLSP-Emo: Towards Empathetic Large Speech-Language Models
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[12]  arXiv:2406.03814 (cross-list from cs.CL) [pdf, other]
Title: Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[13]  arXiv:2406.03657 (cross-list from eess.AS) [pdf, other]
Title: UrBAN: Urban Beehive Acoustics and PheNotyping Dataset
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[14]  arXiv:2406.03637 (cross-list from eess.AS) [pdf, other]
Title: Style Mixture of Experts for Expressive Text-To-Speech Synthesis
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)
[15]  arXiv:2405.19334 (cross-list from cs.AI) [pdf, other]
Title: LLMs Meet Multimodal Generation and Editing: A Survey
Comments: 51 Pages with 16 Figures, 12 Tables, and 534 References. GitHub Repository at: this https URL
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD)

Thu, 6 Jun 2024 (showing first 18 of 24 entries)

[16]  arXiv:2406.03344 [pdf, other]
Title: Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
Comments: Code is available at this https URL
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[17]  arXiv:2406.03251 [pdf, other]
Title: ASoBO: Attentive Beamformer Selection for Distant Speaker Diarization in Meetings
Comments: 5 pages, 2 figures, 2 tables, accepted at Interspeech 2024
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[18]  arXiv:2406.03247 [pdf, other]
Title: Genuine-Focused Learning using Mask AutoEncoder for Generalized Fake Audio Detection
Comments: Accepted by INTERSPEECH 2024
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[19]  arXiv:2406.03240 [pdf, other]
Title: Generalized Source Tracing: Detecting Novel Audio Deepfake Algorithm with Real Emphasis and Fake Dispersion strategy
Comments: Accepted by INTERSPEECH 2024
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[20]  arXiv:2406.03237 [pdf, other]
Title: Generalized Fake Audio Detection via Deep Stable Learning
Comments: accepted by INTERSPEECH2024
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[21]  arXiv:2406.03138 [pdf, other]
Title: A Frame-based Attention Interpretation Method for Relevant Acoustic Feature Extraction in Long Speech Depression Detection
Comments: 5 pages, 3 figures. arXiv admin note: substantial text overlap with arXiv:2309.13476
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[22]  arXiv:2406.02963 [pdf, other]
Title: Dataset-Distillation Generative Model for Speech Emotion Recognition
Comments: Accepted at Interspeech 2024
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[23]  arXiv:2406.02940 [pdf, other]
Title: Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[24]  arXiv:2406.02897 [pdf, other]
Title: LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[25]  arXiv:2406.02565 [pdf, other]
Title: Sequence-to-sequence models in peer-to-peer learning: A practical application
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Multiagent Systems (cs.MA); Audio and Speech Processing (eess.AS)
[26]  arXiv:2406.03460 (cross-list from eess.AS) [pdf, other]
Title: The PESQetarian: On the Relevance of Goodhart's Law for Speech Enhancement
Comments: Accepted at Interspeech 2024
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[27]  arXiv:2406.03407 (cross-list from cs.LG) [pdf, other]
Title: Physics and geometry informed neural operator network with application to acoustic scattering
Comments: 20 pages of main text, 9 figures
Subjects: Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS); Computational Physics (physics.comp-ph)
[28]  arXiv:2406.03274 (cross-list from eess.AS) [pdf, other]
Title: Enhancing CTC-based speech recognition with diverse modeling units
Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Sound (cs.SD)
[29]  arXiv:2406.03120 (cross-list from eess.AS) [pdf, other]
Title: RevRIR: Joint Reverberant Speech and Room Impulse Response Embedding using Contrastive Learning with Application to Room Shape Classification
Comments: Accepted to Interspeech 2024
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[30]  arXiv:2406.03049 (cross-list from cs.CL) [pdf, other]
Title: StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
Comments: Accepted to ACL 2024 main conference, Project Page: this https URL
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[31]  arXiv:2406.02951 (cross-list from cs.CV) [pdf, other]
Title: AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection
Comments: Accepted to CVPR 2024
Subjects: Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[32]  arXiv:2406.02950 (cross-list from eess.AS) [pdf, other]
Title: 4D ASR: Joint Beam Search Integrating CTC, Attention, Transducer, and Mask Predict Decoders
Comments: submitted to IEEE/ACM Transactions on Audio Speech and Language Processing
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[33]  arXiv:2406.02925 (cross-list from eess.AS) [pdf, other]
Title: SYN2REAL: Leveraging Task Arithmetic for Mitigating Synthetic-Real Discrepancies in ASR Domain Adaptation
Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)
[ total of 83 entries: 1-33 | 34-66 | 67-83 ]
[ showing 33 entries per page: fewer | more | all ]

Disable MathJax (What is MathJax?)

Links to: arXiv, form interface, find, cs, new, 2406, contact, help  (Access key information)