We gratefully acknowledge support from
the Simons Foundation and member institutions.

Audio and Speech Processing

Authors and titles for recent submissions

[ total of 70 entries: 1-50 | 51-70 ]
[ showing 50 entries per page: fewer | more | all ]

Fri, 7 Aug 2020

[1]  arXiv:2008.02689 [pdf, ps, other]
Title: Aalto's End-to-End DNN systems for the INTERSPEECH 2020 Computational Paralinguistics Challenge
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[2]  arXiv:2008.02686 [pdf, ps, other]
Title: Attentive Fusion Enhanced Audio-Visual Encoding for Transformer Based Robust Speech Recognition
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD)
[3]  arXiv:2008.02651 [pdf, other]
Title: Improving on-device speaker verification using federated learning with privacy
Comments: To appear in proceedings of INTERSPEECH 2020
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD); Machine Learning (stat.ML)
[4]  arXiv:2008.02603 [pdf, other]
Title: Data balancing for boosting performance of low-frequency classes in Spoken Language Understanding
Comments: accepted at InterSpeech 2020
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[5]  arXiv:2008.02519 [pdf]
Title: Spectral-change enhancement with prior SNR for the hearing impaired
Comments: Accepted by 23rd International Congress on Acoustics (ICA 2019), see this http URL
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[6]  arXiv:2008.02516 [pdf, other]
Title: FastLR: Non-Autoregressive Lipreading Model with Integrate-and-Fire
Comments: Accepted by ACM MM 2020
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Sound (cs.SD)
[7]  arXiv:2008.02493 [pdf, other]
Title: HooliGAN: Robust, High Quality Neural Vocoding
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[8]  arXiv:2008.02490 [pdf]
Title: PPSpeech: Phrase based Parallel End-to-End TTS System
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[9]  arXiv:2008.02487 [pdf, other]
Title: Shouted Speech Compensation for Speaker Verification Robust to Vocal Effort Conditions
Subjects: Audio and Speech Processing (eess.AS); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG); Sound (cs.SD)
[10]  arXiv:2008.02480 [pdf, other]
Title: Mixing-Specific Data Augmentation Techniques for Improved Blind Violin/Piano Source Separation
Comments: Accepted to IEEE 22nd International Workshop on Multimedia Signal Processing (MMSP 2020)
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[11]  arXiv:2008.02470 [pdf, other]
Title: Quantification of Transducer Misalignment in Ultrasound Tongue Imaging
Comments: 5 pages, accepted for publication at Interspeech 2020
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[12]  arXiv:2008.02439 [pdf, ps, other]
Title: Simultaneous measurement of time-invariant linear and nonlinear, and random and extra responses using frequency domain variant of velvet noise
Comments: 10 pages, 15 figures, APSIPA ASC 2020
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[13]  arXiv:2008.02371 [pdf, other]
Title: Recognition-Synthesis Based Non-Parallel Voice Conversion with Adversarial Learning
Comments: Accepted to INTERSPEECH 2020
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[14]  arXiv:2008.02323 [pdf, other]
Title: Hybrid Transformer/CTC Networks for Hardware Efficient Voice Triggering
Comments: INTERSPEECH, 2020
Subjects: Audio and Speech Processing (eess.AS); Human-Computer Interaction (cs.HC); Machine Learning (cs.LG); Sound (cs.SD)
[15]  arXiv:2008.02791 (cross-list from cs.SD) [pdf, other]
Title: Few-Shot Drum Transcription in Polyphonic Music
Comments: ISMIR 2020 camera-ready
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[16]  arXiv:2008.02734 (cross-list from cs.SD) [pdf, other]
Title: Exact, Parallelizable Dynamic Time Warping Alignment with Linear Memory
Comments: 12 Pages, 6 Figures, 1 Table, ISMIR 2020
Subjects: Sound (cs.SD); Information Retrieval (cs.IR); Machine Learning (cs.LG); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[17]  arXiv:1901.08025 (cross-list from cs.MM) [pdf, ps, other]
Title: Generalization of Spoofing Countermeasures: a Case Study with ASVspoof 2015 and BTAS 2016 Corpora
Journal-ref: Published in 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2017), New Orleans, LA, USA
Subjects: Multimedia (cs.MM); Audio and Speech Processing (eess.AS)

Thu, 6 Aug 2020

[18]  arXiv:2008.02098 [pdf, other]
Title: Speaker dependent acoustic-to-articulatory inversion using real-time MRI of the vocal tract
Comments: 5 pages, accepted for publication at Interspeech 2020. arXiv admin note: substantial text overlap with arXiv:2008.00889
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[19]  arXiv:2008.02070 [pdf, other]
Title: Content based singing voice source separation via strong conditioning using aligned phonemes
Comments: 21st International Society for Music Information Retrieval Conference 11-15 October 2020, Montreal, Canada
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[20]  arXiv:2008.02027 [pdf, other]
Title: Learning to Denoise Historical Music
Comments: ISMIR 2020
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG)
[21]  arXiv:2008.01832 [pdf, other]
Title: Future Vector Enhanced LSTM Language Model for LVCSR
Comments: Accepted by ASRU-2017
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[22]  arXiv:2008.02194 (cross-list from cs.SD) [pdf, other]
Title: On the Characterization of Expressive Performance in Classical Music: First Results of the Con Espressione Game
Comments: 8 pages, 2 figures, accepted for the 21st International Society for Music Information Retrieval Conference (ISMIR 2020)
Subjects: Sound (cs.SD); Information Retrieval (cs.IR); Audio and Speech Processing (eess.AS)
[23]  arXiv:2008.02069 (cross-list from cs.LG) [pdf, other]
Title: Data Cleansing with Contrastive Learning for Vocal Note Event Annotations
Comments: 21st International Society for Music Information Retrieval Conference 11-15 October 2020, Montreal, Canada
Subjects: Machine Learning (cs.LG); Information Retrieval (cs.IR); Sound (cs.SD); Audio and Speech Processing (eess.AS); Machine Learning (stat.ML)
[24]  arXiv:2008.02063 (cross-list from cs.CV) [pdf, other]
Title: Compact Graph Architecture for Speech Emotion Recognition
Authors: A. Shirian, T. Guha
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[25]  arXiv:2008.02011 (cross-list from cs.SD) [pdf, other]
Title: Neural Loop Combiner: Neural Network Models for Assessing the Compatibility of Loops
Comments: Accepted to the 21st International Society for Music Information Retrieval Conference (ISMIR 2020)
Subjects: Sound (cs.SD); Information Retrieval (cs.IR); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[26]  arXiv:2008.01951 (cross-list from cs.SD) [pdf, other]
Title: MusPy: A Toolkit for Symbolic Music Generation
Comments: Accepted by International Society for Music Information Retrieval Conference (ISMIR), 2020
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS); Machine Learning (stat.ML)

Wed, 5 Aug 2020

[27]  arXiv:2008.01698 [pdf, other]
Title: MIRNet: Learning multiple identities representations in overlapped speech
Comments: Accepted in Interspeech 2020
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[28]  arXiv:2008.01504 [pdf, other]
Title: "This is Houston. Say again, please". The Behavox system for the Apollo-11 Fearless Steps Challenge (phase II)
Comments: Accepted to Interspeech 2020
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[29]  arXiv:2008.01348 [pdf, other]
Title: Intra-class variation reduction of speaker representation in disentanglement framework
Comments: Accepted for INTERSPEECH 2020
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[30]  arXiv:2008.01300 [pdf, other]
Title: Weakly Supervised Construction of ASR Systems with Massive Video Data
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)
[31]  arXiv:2008.01160 [pdf, other]
Title: A Spectral Energy Distance for Parallel Speech Synthesis
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD); Machine Learning (stat.ML)
[32]  arXiv:2008.01077 [pdf, other]
Title: Self-attention encoding and pooling for speaker recognition
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[33]  arXiv:2008.01543 (cross-list from cs.CL) [pdf, other]
Title: Text-based classification of interviews for mental health -- juxtaposing the state of the art
Comments: 33 pages, 7 figures, belabBERT is available on this http URL
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD); Audio and Speech Processing (eess.AS); Machine Learning (stat.ML)
[34]  arXiv:2008.01532 (cross-list from cs.CL) [pdf, other]
Title: A Study on Effects of Implicit and Explicit Language Model Information for DBLSTM-CTC Based Handwriting Recognition
Comments: Accepted by ICDAR-2015
Subjects: Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
[35]  arXiv:2008.01490 (cross-list from cs.SD) [pdf, other]
Title: Expressive TTS Training with Frame and Style Reconstruction Loss
Comments: Submitted to IEEE/ACM Transactions on Audio, Speech and Language Processing
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[36]  arXiv:2008.01431 (cross-list from cs.SD) [pdf, other]
Title: Automatic Composition of Guitar Tabs by Transformers and Groove Modeling
Comments: Accepted at Proc. Int. Society for Music Information Retrieval Conf. 2020
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[37]  arXiv:2008.01393 (cross-list from cs.SD) [pdf, other]
Title: Neural Granular Sound Synthesis
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[38]  arXiv:2008.01370 (cross-list from cs.SD) [pdf]
Title: Timbre latent space: exploration and creative aspects
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[39]  arXiv:2008.01307 (cross-list from cs.SD) [pdf, other]
Title: The Jazz Transformer on the Front Line: Exploring the Shortcomings of AI-composed Music through Quantitative Measures
Comments: Accepted to the 21st International Society for Music Information Retrieval Conference (ISMIR 2020)
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[40]  arXiv:2008.01291 (cross-list from cs.LG) [pdf, other]
Title: Music SketchNet: Controllable Music Generation via Factorized Representations of Pitch and Rhythm
Comments: 8 pages, 8 figures, Proceedings of the 21st International Society for Music Information Retrieval Conference, ISMIR 2020
Subjects: Machine Learning (cs.LG); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS); Machine Learning (stat.ML)

Tue, 4 Aug 2020 (showing first 10 of 23 entries)

[41]  arXiv:2008.00953 [pdf, other]
Title: Modular End-to-end Automatic Speech Recognition Framework for Acoustic-to-word Model
Comments: Accepted by IEEE TASLP
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[42]  arXiv:2008.00889 [pdf, other]
Title: Speaker dependent articulatory-to-acoustic mapping using real-time MRI of the vocal tract
Comments: 5 pages, accepted for publication at Interspeech 2020
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD); Image and Video Processing (eess.IV)
[43]  arXiv:2008.00816 [pdf, other]
Title: Evolving Multi-Resolution Pooling CNN for Monaural Singing Voice Separation
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[44]  arXiv:2008.00781 [pdf, other]
Title: MusiCoder: A Universal Music-Acoustic Encoder Based on Transformers
Comments: 12 pages, submitted to MMM2021
Subjects: Audio and Speech Processing (eess.AS); Multimedia (cs.MM)
[45]  arXiv:2008.00768 [pdf, other]
Title: One Model, Many Languages: Meta-learning for Multilingual Text-to-Speech
Comments: Accepted to INTERSPEECH 2020; for the source files, see this https URL
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG)
[46]  arXiv:2008.00756 [pdf, other]
Title: Structure and Automatic Segmentation of Dhrupad Vocal Bandish Audio
Comments: Part of this work published in ISMIR 2020
Subjects: Audio and Speech Processing (eess.AS); Information Retrieval (cs.IR); Machine Learning (cs.LG)
[47]  arXiv:2008.00731 [pdf]
Title: Unsupervised Discovery of Recurring Speech Patterns Using Probabilistic Adaptive Metrics
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)
[48]  arXiv:2008.00702 [pdf, other]
Title: Multimodal Semi-supervised Learning Framework for Punctuation Prediction in Conversational Speech
Comments: Accepted for Interspeech 2020
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL)
[49]  arXiv:2008.00671 [pdf, other]
Title: TutorNet: Towards Flexible Knowledge Distillation for End-to-End Speech Recognition
Comments: 10 pages, 6 figures, 12 tables. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
Subjects: Audio and Speech Processing (eess.AS)
[50]  arXiv:2008.00667 [pdf, other]
Title: Learning Intonation Pattern Embeddings for Arabic Dialect Identification
Comments: Accepted for INTERSPEECH 2020
Subjects: Audio and Speech Processing (eess.AS)
[ total of 70 entries: 1-50 | 51-70 ]
[ showing 50 entries per page: fewer | more | all ]

Disable MathJax (What is MathJax?)

Links to: arXiv, form interface, find, eess, new, 2008, contact, help  (Access key information)