Sound
Authors and titles for cs.SD in Jun 2022, skipping first 25
[ total of 221 entries: 1-50 | 26-75 | 76-125 | 126-175 | 176-221 ][ showing 50 entries per page: fewer | more | all ]
- [26] arXiv:2206.05408 [pdf, other]
-
Title: Multi-instrument Music Synthesis with Spectrogram DiffusionAuthors: Curtis Hawthorne, Ian Simon, Adam Roberts, Neil Zeghidour, Josh Gardner, Ethan Manilow, Jesse EngelSubjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
- [27] arXiv:2206.05876 [pdf, other]
-
Title: Description and Discussion on DCASE 2022 Challenge Task 2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring Applying Domain Generalization TechniquesAuthors: Kota Dohi, Keisuke Imoto, Noboru Harada, Daisuke Niizumi, Yuma Koizumi, Tomoya Nishida, Harsh Purohit, Takashi Endo, Masaaki Yamamoto, Yohei KawaguchiComments: arXiv admin note: substantial text overlap with arXiv:2106.04492Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS); Machine Learning (stat.ML)
- [28] arXiv:2206.05929 [pdf, other]
-
Title: Improvement of Serial Approach to Anomalous Sound Detection by Incorporating Two Binary Cross-Entropies for Outlier ExposureComments: 5 pages, 3 figures, 3 tables, EUSIPCO 2022Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [29] arXiv:2206.06057 [pdf, ps, other]
-
Title: Low-complexity deep learning frameworks for acoustic scene classificationSubjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
- [30] arXiv:2206.06117 [pdf]
-
Title: Optimizing musical chord inversions using the cartesian coordinate systemAuthors: Steve Mathew D AComments: 9 pages, 5 tablesSubjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [31] arXiv:2206.06126 [pdf, other]
-
Title: Robust Time Series Denoising with Learnable Wavelet Packet TransformComments: 15 pages, 13 figures, 8 tablesSubjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
- [32] arXiv:2206.06573 [pdf, ps, other]
-
Title: Speech intelligibility of simulated hearing loss sounds and its prediction using the Gammachirp Envelope Similarity Index (GESI)Comments: This paper was submitted to Interspeech 2022Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [33] arXiv:2206.06604 [pdf, other]
-
Title: WHIS: Hearing impairment simulator based on the gammachirp auditory filterbankAuthors: Toshio IrinoComments: This paper was submitted to Trends in Hearing on Jun 5, 2022Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [34] arXiv:2206.06680 [pdf, other]
-
Title: Exploring speaker enrolment for few-shot personalisation in emotional vocalisation predictionComments: Proceedings of the ICML Expressive Vocalizations Workshop and Competition held in conjunction with the $\mathit{39}^{th}$ International Conference on Machine Learning, Copyright 2022 by the author(s)Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
- [35] arXiv:2206.06908 [pdf, other]
-
Title: LPCSE: Neural Speech Enhancement through Linear Predictive CodingSubjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [36] arXiv:2206.07176 [pdf, other]
-
Title: Frequency-centroid features for word recognition of non-native English speakersComments: Published in IEEE Irish Signals & Systems Conference (ISSC), 2022Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
- [37] arXiv:2206.07229 [pdf, other]
-
Title: Accurate Emotion Strength Assessment for Seen and Unseen Speech Based on Data-Driven Deep LearningComments: To appear in INTERSPEECH 2022. 5 pages, 4 figures. Substantial text overlap with arXiv:2110.03156Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
- [38] arXiv:2206.07288 [pdf, other]
-
Title: Streaming non-autoregressive model for any-to-many voice conversionSubjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [39] arXiv:2206.07289 [pdf, other]
-
Title: Text-Aware End-to-end Mispronunciation Detection and DiagnosisComments: Rejected by Interspeech2022Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
- [40] arXiv:2206.07293 [pdf, other]
-
Title: FRCRN: Boosting Feature Representation using Frequency Recurrence for Monaural Speech EnhancementComments: The paper has been accepted by ICASSP 2022. 5 pages, 2 figures, 5 tablesSubjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [41] arXiv:2206.07340 [pdf, other]
- [42] arXiv:2206.07347 [pdf, other]
-
Title: On the Use of Deep Mask Estimation Module for Neural Source Separation SystemsComments: Accepted by Interspeech 2022Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [43] arXiv:2206.07511 [pdf, other]
-
Title: Investigating Multi-Feature Selection and Ensembling for Audio ClassificationSubjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
- [44] arXiv:2206.07860 [pdf, other]
-
Title: EPG2S: Speech Generation and Speech Enhancement based on Electropalatography and Audio Signals using Multimodal LearningComments: Accepted By IEEE Signal Processing LetterSubjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
- [45] arXiv:2206.07956 [pdf, other]
-
Title: Automatic Prosody Annotation with Pre-Trained Text-Speech ModelComments: accepted by INTERSPEECH2022Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
- [46] arXiv:2206.08007 [pdf, ps, other]
-
Title: DCASE 2022: Comparative Analysis Of CNNs For Acoustic Scene Classification Under Low-Complexity ConsiderationsSubjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
- [47] arXiv:2206.08039 [pdf, ps, other]
-
Title: Acoustic Modeling for End-to-End Empathetic Dialogue Speech Synthesis Using Linguistic and Prosodic Contexts of Dialogue HistoryComments: 5 pages, 3 figures, Accepted for INTERSPEECH2022Subjects: Sound (cs.SD); Computation and Language (cs.CL); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
- [48] arXiv:2206.08170 [pdf, other]
-
Title: Adversarial Privacy Protection on Speech EnhancementComments: 5 pages, 6 figuresSubjects: Sound (cs.SD); Cryptography and Security (cs.CR); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
- [49] arXiv:2206.08189 [pdf, other]
-
Title: Censer: Curriculum Semi-supervised Learning for Speech Recognition Based on Self-supervised Pre-trainingSubjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [50] arXiv:2206.08233 [pdf, other]
-
Title: Event-related data conditioning for acoustic event classificationComments: Accepted by INTERSPEECH 2022Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [51] arXiv:2206.08297 [pdf, other]
-
Title: GoodBye WaveNet -- A Language Model for Raw Audio with Context of 1/2 Million SamplesAuthors: Prateek VermaComments: 12 pages, 1 figure. Technical Report at Stanford University. Ongoing WorkSubjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
- [52] arXiv:2206.08312 [pdf, other]
-
Title: SoundSpaces 2.0: A Simulation Platform for Visual-Acoustic LearningAuthors: Changan Chen, Carl Schissler, Sanchit Garg, Philip Kobernik, Alexander Clegg, Paul Calamia, Dhruv Batra, Philip W Robinson, Kristen GraumanSubjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
- [53] arXiv:2206.08317 [pdf, other]
-
Title: Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech RecognitionComments: 5 pages, 3 figures, accepted by INTERSPEECH 2022Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
- [54] arXiv:2206.09131 [pdf, other]
-
Title: Tackling Spoofing-Aware Speaker Verification with Multi-Model FusionAuthors: Haibin Wu, Jiawen Kang, Lingwei Meng, Yang Zhang, Xixin Wu, Zhiyong Wu, Hung-yi Lee, Helen MengComments: Accepted by Odyssey 2022Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
- [55] arXiv:2206.09142 [pdf, other]
-
Title: Redundancy Reduction Twins Network: A Training framework for Multi-output Emotion RegressionComments: 5 pages, accepted by ICML Exvo workshopSubjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [56] arXiv:2206.09298 [pdf, ps, other]
-
Title: GMM based multi-stage Wiener filtering for low SNR speech enhancementComments: 5 pages, 3 figures, submitted to a conferenceSubjects: Sound (cs.SD); Robotics (cs.RO); Audio and Speech Processing (eess.AS)
- [57] arXiv:2206.09920 [pdf, other]
- [58] arXiv:2206.10175 [pdf, other]
-
Title: A Multi-grained based Attention Network for Semi-supervised Sound Event DetectionJournal-ref: INTERSPEECH 2022Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [59] arXiv:2206.10256 [pdf, other]
-
Title: Human-in-the-loop Speaker Adaptation for DNN-based Multi-speaker TTSComments: 5 pages, 3 figures, Accepted for INTERSPEECH2022Subjects: Sound (cs.SD); Human-Computer Interaction (cs.HC); Neural and Evolutionary Computing (cs.NE); Audio and Speech Processing (eess.AS)
- [60] arXiv:2206.10349 [pdf, ps, other]
-
Title: Joint Analysis of Acoustic Scenes and Sound Events Based on Multitask Learning with Dynamic Weight AdaptationComments: Submitted to Acoustical Science and TechnologySubjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [61] arXiv:2206.10421 [pdf, other]
-
Title: Rethinking Audio-visual Synchronization for Active Speaker DetectionComments: Accepted by IEEE International Workshop on Machine Learning for Signal Processing (MLSP 2022)Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
- [62] arXiv:2206.10695 [pdf, other]
-
Title: Exploring the Effectiveness of Self-supervised Learning and Classifier Chains in Emotion Recognition of Nonverbal VocalizationsComments: Accepted by the ICML Expressive Vocalizations Workshop and Competition 2022Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [63] arXiv:2206.10805 [pdf, other]
-
Title: Jointist: Joint Learning for Multi-instrument Transcription and Its ApplicationsAuthors: Kin Wai Cheuk, Keunwoo Choi, Qiuqiang Kong, Bochen Li, Minz Won, Amy Hung, Ju-Chiang Wang, Dorien HerremansComments: Submitted to ISMIRSubjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
- [64] arXiv:2206.11049 [pdf, other]
-
Title: Dynamic Restrained Uncertainty Weighting Loss for Multitask Learning of Vocal ExpressionAuthors: Meishu Song, Zijiang Yang, Andreas Triantafyllopoulos, Xin Jing, Vincent Karas, Xie Jiangjian, Zixing Zhang, Yamamoto Yoshiharu, Bjoern W. SchullerComments: 5 pagesSubjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
- [65] arXiv:2206.11066 [pdf, other]
-
Title: Radio2Speech: High Quality Speech Recovery from Radio Frequency SignalsComments: Accepted to INTERSPEECH 2022Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [66] arXiv:2206.11260 [pdf, other]
-
Title: Few-shot Long-Tailed Bird Audio RecognitionComments: LifeCLEF2022 (best paper award)Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
- [67] arXiv:2206.11567 [pdf]
-
Title: Restoring speech intelligibility for hearing aid users with deep learningAuthors: Peter Udo Diehl, Yosef Singer, Hannes Zilly, Uwe Schönfeld, Paul Meyer-Rachner, Mark Berry, Henning Sprekeler, Elias Sprengel, Annett Pudszuhn, Veit M. HofmannSubjects: Sound (cs.SD); Audio and Speech Processing (eess.AS); Neurons and Cognition (q-bio.NC)
- [68] arXiv:2206.11632 [pdf, other]
-
Title: Formant Estimation and Tracking using Probabilistic Heat-MapsComments: interspeech 2022Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [69] arXiv:2206.11643 [pdf, ps, other]
-
Title: Towards Green ASR: Lossless 4-bit Quantization of a Hybrid TDNN System on the 300-hr Switchboard CorpusComments: Interspeech 2022 Accepted. arXiv admin note: text overlap with arXiv:2111.14479Subjects: Sound (cs.SD); Human-Computer Interaction (cs.HC); Audio and Speech Processing (eess.AS)
- [70] arXiv:2206.11699 [pdf, ps, other]
-
Title: The SJTU X-LANCE Lab System for CNSRC 2022Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [71] arXiv:2206.11968 [pdf, other]
-
Title: Comparing supervised and self-supervised embedding for ExVo Multi-Task learning trackJournal-ref: Proceedings of the ICML 2022 Expressive Vocalizations Workshop and Competition: Recognizing, Generating, and Personalizing Vocal BurstsSubjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
- [72] arXiv:2206.12038 [pdf, other]
-
Title: BYOL-S: Learning Self-supervised Speech Representations by BootstrappingAuthors: Gasser Elbanna, Neil Scheidwasser-Clow, Mikolaj Kegler, Pierre Beckmann, Karl El Hajal, Milos CernakComments: Submitted to HEAR-PMLR 2021Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
- [73] arXiv:2206.12229 [pdf, other]
-
Title: Exact Prosody Cloning in Zero-Shot Multispeaker Text-to-SpeechComments: Accepted to IEEE SLT 2022Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
- [74] arXiv:2206.12230 [pdf, other]
-
Title: Deformable CNN and Imbalance-Aware Feature Learning for Singing Technique ClassificationComments: Accepted to INTERSPEECH2022Subjects: Sound (cs.SD); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
- [75] arXiv:2206.12320 [pdf, other]
-
Title: PoCaP Corpus: A Multimodal Dataset for Smart Operating Room Speech Assistant using Interventional Radiology Workflow AnalysisAuthors: Kubilay Can Demir, Matthias May, Axel Schmid, Michael Uder, Katharina Breininger, Tobias Weise, Andreas Maier, Seung Hee YangComments: 8 pages, 4 figures, Text, Speech and Dialogue 2022 ConferenceSubjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[ showing 50 entries per page: fewer | more | all ]
Disable MathJax (What is MathJax?)
Links to: arXiv, form interface, find, cs, 2302, contact, help (Access key information)