Ke Tan,DeLiang Wang
Ke Tan
The use of deep neural networks (DNNs) has dramatically elevated the performance of speech enhancement over the last decade. However, to achieve strong enhancement performance typically requires a large DNN, which is both memory and computa...
Ashutosh Pandey,DeLiang Wang
Ashutosh Pandey
Speech enhancement in the time domain is becoming increasingly popular in recent years, due to its capability to jointly enhance both the magnitude and the phase of speech. In this work, we propose a dense convolutional network (DCN) with s...
Meta-learning with Latent Space Clustering in Generative Adversarial Network for Speaker Diarization [0.03%]
基于生成对抗网络隐空间聚类的说话人聚类元学习方法
Monisankha Pal,Manoj Kumar,Raghuveer Peri et al.
Monisankha Pal et al.
The performance of most speaker diarization systems with x-vector embeddings is both vulnerable to noisy environments and lacks domain robustness. Earlier work on speaker diarization using generative adversarial network (GAN) with an encode...
Proportionate Adaptive Filtering Algorithms Derived Using an Iterative Reweighting Framework [0.03%]
基于迭代重加权框架的自适应滤波算法的比例化研究
Ching-Hua Lee,Bhaskar D Rao,Harinath Garudadri
Ching-Hua Lee
In this paper, based on sparsity-promoting regularization techniques from the sparse signal recovery (SSR) area, least mean square (LMS)-type sparse adaptive filtering algorithms are derived. The approach mimics the iterative reweighted ℓ ...
Speech Intelligibility Prediction using Spectro-Temporal Modulation Analysis [0.03%]
基于谱时调制分析的可懂度预测方法研究
Amin Edraki,Wai-Yip Chan,Jesper Jensen et al.
Amin Edraki et al.
Spectro-temporal modulations are believed to mediate the analysis of speech sounds in the human primary auditory cortex. Inspired by humans' robustness in comprehending speech in challenging acoustic environments, we propose an intrusive sp...
Robust Estimation of Hypernasality in Dysarthria with Acoustic Model Likelihood Features [0.03%]
基于声学模型似然特征的失语症患者鼻音异常程度的鲁棒估计方法研究
Michael Saxon,Ayush Tripathi,Yishan Jiao et al.
Michael Saxon et al.
Hypernasality is a common characteristic symptom across many motor-speech disorders. For voiced sounds, hypernasality introduces an additional resonance in the lower frequencies and, for unvoiced sounds, there is reduced articulatory precis...
On Cross-Corpus Generalization of Deep Learning Based Speech Enhancement [0.03%]
基于深度学习的语音增强在跨语料数据上的泛化性能分析
Ashutosh Pandey,DeLiang Wang
Ashutosh Pandey
In recent years, supervised approaches using deep neural networks (DNNs) have become the mainstream for speech enhancement. It has been established that DNNs generalize well to untrained noises and speakers if trained using a large number o...
Complex Spectral Mapping for Single- and Multi-Channel Speech Enhancement and Robust ASR [0.03%]
用于单通道和多通道语音增强及鲁棒ASR的复数谱映射方法
Zhong-Qiu Wang,Peidong Wang,DeLiang Wang
Zhong-Qiu Wang
This study proposes a complex spectral mapping approach for single- and multi-channel speech enhancement, where deep neural networks (DNNs) are used to predict the real and imaginary (RI) components of the direct-path signal from noisy and ...
Monaural Speech Dereverberation Using Temporal Convolutional Networks with Self Attention [0.03%]
基于自注意力的时序卷积网络单通道语音去混响方法
Yan Zhao,DeLiang Wang,Buye Xu et al.
Yan Zhao et al.
In daily listening environments, human speech is often degraded by room reverberation, especially under highly reverberant conditions. Such degradation poses a challenge for many speech processing systems, where the performance becomes much...
Zhong-Qiu Wang,DeLiang Wang
Zhong-Qiu Wang
This study investigates deep learning based single- and multi-channel speech dereverberation. For single-channel processing, we extend magnitude-domain masking and mapping based dereverberation to complex-domain mapping, where deep neural n...