We gratefully acknowledge support from
the Simons Foundation and member institutions.

Audio and Speech Processing

Authors and titles for eess.AS in Apr 2019, skipping first 25

[ total of 167 entries: 1-25 | 26-50 | 51-75 | 76-100 | 101-125 | ... | 151-167 ]
[ showing 25 entries per page: fewer | more | all ]
[26]  arXiv:1904.05441 [pdf, other]
Title: ASVspoof 2019: Future Horizons in Spoofed and Fake Audio Detection
Journal-ref: Proc. Interspeech 2019
Subjects: Audio and Speech Processing (eess.AS); Cryptography and Security (cs.CR); Sound (cs.SD)
[27]  arXiv:1904.06086 [pdf, other]
Title: Unsupervised Speech Domain Adaptation Based on Disentangled Representation Learning for Robust Speech Recognition
Comments: Submitted to Interspeech 2019
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[28]  arXiv:1904.06157 [pdf, other]
Title: Examining the Mapping Functions of Denoising Autoencoders in Singing Voice Separation
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[29]  arXiv:1904.06478 [pdf, other]
Title: Low-Latency Speaker-Independent Continuous Speech Separation
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[30]  arXiv:1904.06588 [pdf, ps, other]
Title: Audio Compression Using Graph-based Transform
Comments: 2018 9th International Symposium on Telecommunications (IST)
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD); Signal Processing (eess.SP)
[31]  arXiv:1904.06648 [pdf, ps, other]
Title: A robust DOA estimation method for a linear microphone array under reverberant and noisy environments
Authors: Hao Wang, Jing Lu
Comments: 7 pages, 4 figures, 3 tables, 33 references
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[32]  arXiv:1904.06868 [pdf, ps, other]
Title: Singing voice synthesis based on convolutional neural networks
Comments: Singing voice samples (Japanese, English, Chinese): this https URL
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[33]  arXiv:1904.07294 [pdf, ps, other]
Title: RHR-Net: A Residual Hourglass Recurrent Neural Network for Speech Enhancement
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD); Signal Processing (eess.SP)
[34]  arXiv:1904.07386 [pdf, other]
[35]  arXiv:1904.07453 [pdf, other]
Title: Spoof detection using time-delay shallow neural network and feature switching
Journal-ref: 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 1011--1017
Subjects: Audio and Speech Processing (eess.AS); Cryptography and Security (cs.CR); Machine Learning (cs.LG); Sound (cs.SD)
[36]  arXiv:1904.07704 [pdf, other]
Title: SpeechYOLO: Detection and Localization of Speech Objects
Journal-ref: Interspeech 2019, pp. 4210-4214
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD); Machine Learning (stat.ML)
[37]  arXiv:1904.08104 [pdf, ps, other]
Title: RawNet: Advanced end-to-end deep neural network using raw waveforms for text-independent speaker verification
Comments: Accepted for oral presentation at Interspeech 2019, code available at this http URL
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[38]  arXiv:1904.08248 [pdf, ps, other]
Title: An Analysis of Speech Enhancement and Recognition Losses in Limited Resources Multi-talker Single Channel Audio-Visual ASR
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD); Machine Learning (stat.ML)
[39]  arXiv:1904.08775 [pdf, other]
Title: Few Shot Speaker Recognition using Deep Neural Networks
Subjects: Audio and Speech Processing (eess.AS); Machine Learning (cs.LG); Sound (cs.SD)
[40]  arXiv:1904.08779 [pdf, other]
Title: SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
Comments: 5 pages, 3 figures, 6 tables; v3: references added
Journal-ref: Proc. Interspeech 2019, 2613-2617
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD); Machine Learning (stat.ML)
[41]  arXiv:1904.09038 [pdf, other]
Title: Leveraging native language information for improved accented speech recognition
Comments: Accepted at Interspeech 2018
Subjects: Audio and Speech Processing (eess.AS)
[42]  arXiv:1904.09049 [pdf, other]
Title: An Investigation of End-to-End Multichannel Speech Recognition for Reverberant and Mismatch Conditions
Subjects: Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
[43]  arXiv:1904.10045 [pdf, other]
Title: Automatic Spelling Correction with Transformer for CTC-based End-to-End Speech Recognition
Comments: 6pages, 5 figures
Subjects: Audio and Speech Processing (eess.AS); Neural and Evolutionary Computing (cs.NE); Sound (cs.SD)
[44]  arXiv:1904.10134 [pdf, other]
Title: Replay attack detection with complementary high-resolution information using end-to-end DNN for the ASVspoof 2019 Challenge
Comments: Accepted for oral presentation at Interspeech 2019, code available at this https URL
Subjects: Audio and Speech Processing (eess.AS); Cryptography and Security (cs.CR); Sound (cs.SD)
[45]  arXiv:1904.10135 [pdf, other]
Title: Acoustic scene classification using teacher-student learning with soft-labels
Comments: Accepted for presentation at Interspeech 2019
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[46]  arXiv:1904.10408 [pdf, other]
Title: Towards joint sound scene and polyphonic sound event recognition
Comments: Accepted to Interspeech 2019
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[47]  arXiv:1904.10763 [pdf, other]
Title: The Analogue Computer as a Voltage-Controlled Synthesiser
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD)
[48]  arXiv:1904.10788 [pdf, other]
Title: Speech Emotion Recognition Using Multi-hop Attention Mechanism
Comments: 5 pages, Accepted as a conference paper at ICASSP 2019 (oral presentation)
Subjects: Audio and Speech Processing (eess.AS); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)
[49]  arXiv:1904.11130 [pdf, other]
Title: Latent Class Model with Application to Speaker Diarization
Subjects: Audio and Speech Processing (eess.AS)
[50]  arXiv:1904.12069 [pdf, ps, other]
Title: Improving Deep Speech Denoising by Noisy2Noisy Signal Mapping
Subjects: Audio and Speech Processing (eess.AS); Sound (cs.SD); Signal Processing (eess.SP)
[ total of 167 entries: 1-25 | 26-50 | 51-75 | 76-100 | 101-125 | ... | 151-167 ]
[ showing 25 entries per page: fewer | more | all ]

Disable MathJax (What is MathJax?)

Links to: arXiv, form interface, find, eess, 2404, contact, help  (Access key information)