Consonant-Vowel Transition Models Based on Deep Learning for Objective Evaluation of Articulation [0.03%]
基于深度学习的辅元音转换模型在构音客观评价中的应用研究
Vikram C Mathad,Julie M Liss,Kathy Chapman et al.
Vikram C Mathad et al.
Spectro-temporal dynamics of consonant-vowel (CV) transition regions are considered to provide robust cues related to articulation. In this work, we propose an objective measure of precise articulation, dubbed the objective articulation mea...
Self-attending RNN for Speech Enhancement to Improve Cross-corpus Generalization [0.03%]
自注意力循环神经网络在语音增强中的跨语料库泛化性能研究
Ashutosh Pandey,DeLiang Wang
Ashutosh Pandey
Deep neural networks (DNNs) represent the mainstream methodology for supervised speech enhancement, primarily due to their capability to model complex functions using hierarchical representations. However, a recent study revealed that DNNs ...
Neural Cascade Architecture with Triple-domain Loss for Speech Enhancement [0.03%]
基于三领域损失的神经级联架构在语音增强中的应用
Heming Wang,DeLiang Wang
Heming Wang
This paper proposes a neural cascade architecture to address the monaural speech enhancement problem. The cascade architecture is composed of three modules which optimize in turn enhanced speech with respect to the magnitude spectrogram, th...
Ingo R Titze,Anil Palaparthi
Ingo R Titze
A systematic variation of length and cross-sectional area of specific segments of the vocal tract (trachea to lips) was conducted computationally to quantify the effects of source-filter interaction. A one-dimensional Navier-Stokes (transmi...
Analysis and Calibration of Lombard Effect and Whisper for Speaker Recognition [0.03%]
说话人识别中的朗巴德效应与耳语分析及校准
Finnian Kelly,John H L Hansen
Finnian Kelly
Variations in vocal effort can create challenges for speaker recognition systems that are optimized for use with neutral speech. The Lombard effect and whisper are two commonly-occurring forms of vocal effort variation that result in non-ne...
Heming Wang,DeLiang Wang
Heming Wang
Speech super-resolution (SR) aims to increase the sampling rate of a given speech signal by generating high-frequency components. This paper proposes a convolutional neural network (CNN) based SR model that takes advantage of information fr...
Evaluation of Glottal Inverse Filtering Algorithms Using a Physiologically Based Articulatory Speech Synthesizer [0.03%]
基于生理学的发音合成器在声门逆滤波算法评估中的应用研究
Yu-Ren Chien,Daryush D Mehta,Jón Guðnason et al.
Yu-Ren Chien et al.
Glottal inverse filtering aims to estimate the glottal airflow signal from a speech signal for applications such as speaker recognition and clinical voice assessment. Nonetheless, evaluation of inverse filtering algorithms has been challeng...
Ashutosh Pandey,DeLiang Wang
Ashutosh Pandey
This paper proposes a new learning mechanism for a fully convolutional neural network (CNN) to address speech enhancement in the time domain. The CNN takes as input the time frames of noisy utterance and outputs the time frames of the enhan...
Multi-microphone Complex Spectral Mapping for Utterance-wise and Continuous Speech Separation [0.03%]
基于多通道复数谱映射的语音分离算法
Zhong-Qiu Wang,Peidong Wang,DeLiang Wang
Zhong-Qiu Wang
We propose multi-microphone complex spectral mapping, a simple way of applying deep learning for time-varying non-linear beamforming, for speaker separation in reverberant conditions. We aim at both speaker separation and dereverberation. O...
Deep Learning Based Real-time Speech Enhancement for Dual-microphone Mobile Phones [0.03%]
基于深度学习的双麦克风手机实时语音增强技术
Ke Tan,Xueliang Zhang,DeLiang Wang
Ke Tan
In mobile speech communication, speech signals can be severely corrupted by background noise when the far-end talker is in a noisy acoustic environment. To suppress background noise, speech enhancement systems are typically integrated into ...