We gratefully acknowledge support from
the Simons Foundation and member institutions.

Sound

Authors and titles for cs.SD in Feb 2023

[ total of 179 entries: 1-25 | 26-50 | 51-75 | 76-100 | ... | 176-179 ]
[ showing 25 entries per page: fewer | more | all ]
[1]  arXiv:2302.00286 [pdf, other]
Title: Jointist: Simultaneous Improvement of Multi-instrument Transcription and Music Source Separation via Joint Training
Comments: arXiv admin note: text overlap with arXiv:2206.10805
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[2]  arXiv:2302.00646 [pdf, other]
Title: Epic-Sounds: A Large-scale Dataset of Actions That Sound
Comments: 6 pages, 4 figures
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[3]  arXiv:2302.00868 [pdf, other]
Title: Speech Enhancement for Virtual Meetings on Cellular Networks
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[4]  arXiv:2302.01090 [pdf, other]
Title: Goniometers are a Powerful Acoustic Feature for Music Information Retrieval Tasks
Authors: Tim Ziemer
Subjects: Sound (cs.SD); Information Retrieval (cs.IR); Audio and Speech Processing (eess.AS)
[5]  arXiv:2302.02257 [pdf, other]
Title: Multi-Source Diffusion Models for Simultaneous Music Generation and Separation
Comments: Demo page: this https URL
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[6]  arXiv:2302.02845 [pdf, other]
Title: Audio Representation Learning by Distilling Video as Privileged Information
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[7]  arXiv:2302.02945 [pdf]
Title: Improved Vehicle Sub-type Classification for Acoustic Traffic Monitoring
Comments: Accepted at Twenty-Ninth National Conference on Communications(NCC) 23 - 26 February, Indian Institute of Technology Guwahati
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[8]  arXiv:2302.03540 [pdf, other]
Title: Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[9]  arXiv:2302.03917 [pdf, other]
Title: Noise2Music: Text-conditioned Music Generation with Diffusion Models
Comments: 15 pages
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[10]  arXiv:2302.04456 [pdf, other]
Title: ERNIE-Music: Text-to-Waveform Music Generation with Diffusion Models
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[11]  arXiv:2302.04469 [pdf, other]
Title: Joint Acoustic Echo Cancellation and Speech Dereverberation Using Kalman filters
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[12]  arXiv:2302.04577 [pdf, other]
Title: Incorporating Total Variation Regularization in the design of an intelligent Query by Humming system
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[13]  arXiv:2302.05393 [pdf, other]
Title: GTR-CTRL: Instrument and Genre Conditioning for Guitar-Focused Music Generation with Transformers
Comments: This preprint is licensed under a Creative Commons Attribution 4.0 International License (CC BY 4.0). The Version of Record of this contribution is published in Proceedings of EvoMUSART: International Conference on Computational Intelligence in Music, Sound, Art and Design (Part of EvoStar) 2023
Journal-ref: EvoMUSART: International Conference on Computational Intelligence in Music, Sound, Art and Design (Part of EvoStar) 2023
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Multimedia (cs.MM); Audio and Speech Processing (eess.AS)
[14]  arXiv:2302.05690 [pdf]
Title: Attention does not guarantee best performance in speech enhancement
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[15]  arXiv:2302.05693 [pdf]
Title: Local spectral attention for full-band speech enhancement
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[16]  arXiv:2302.05725 [pdf, other]
Title: Parameterizable Acoustical Modeling and Auralization of Cultural Heritage Sites based on Photogrammetry
Authors: Dominik Ukolov
Comments: 6 pages, 3 figures, 27th Conference on Cultural Heritage and New Technologies (Vienna, 2022)
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[17]  arXiv:2302.05940 [pdf, other]
Title: SemanticAC: Semantics-Assisted Framework for Audio Classification
Subjects: Sound (cs.SD); Artificial Intelligence (cs.AI); Audio and Speech Processing (eess.AS)
[18]  arXiv:2302.07640 [pdf, other]
Title: Detecting human and non-human vocal productions in large scale audio recordings
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS); Applications (stat.AP)
[19]  arXiv:2302.08095 [pdf, other]
Title: PAAPLoss: A Phonetic-Aligned Acoustic Parameter Loss for Speech Enhancement
Comments: Accepted at ICASSP 2023
Subjects: Sound (cs.SD); Computation and Language (cs.CL); Audio and Speech Processing (eess.AS)
[20]  arXiv:2302.08130 [pdf, other]
Title: Personalized Audio Quality Preference Prediction
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[21]  arXiv:2302.08136 [pdf, ps, other]
Title: An Attention-based Approach to Hierarchical Multi-label Music Instrument Classification
Comments: To appear at ICASSP 2023
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[22]  arXiv:2302.08137 [pdf, other]
Title: ACE-VC: Adaptive and Controllable Voice Conversion using Explicitly Disentangled Self-supervised Speech Representations
Comments: Published as a conference paper at ICASSP 2023
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[23]  arXiv:2302.08296 [pdf, other]
Title: QuickVC: Any-to-many Voice Conversion Using Inverse Short-time Fourier Transform for Faster Conversion
Subjects: Sound (cs.SD); Audio and Speech Processing (eess.AS)
[24]  arXiv:2302.08632 [pdf, other]
Title: jazznet: A Dataset of Fundamental Piano Patterns for Music Audio Machine Learning Research
Authors: Tosiron Adegbija
Comments: To Appear at IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2023
Subjects: Sound (cs.SD); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[25]  arXiv:2302.08650 [pdf, other]
Title: Gaussian-smoothed Imbalance Data Improves Speech Emotion Recognition
Comments: 5 pages
Subjects: Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
[ total of 179 entries: 1-25 | 26-50 | 51-75 | 76-100 | ... | 176-179 ]
[ showing 25 entries per page: fewer | more | all ]

Disable MathJax (What is MathJax?)

Links to: arXiv, form interface, find, cs, 2305, contact, help  (Access key information)