Learning Complex Spectral Mapping with Gated Convolutional Recurrent Networks for Monaural Speech Enhancement [0.03%]
基于门控卷积循环神经网络的单通道语音增强复杂谱映射学习方法
Ke Tan,DeLiang Wang
Ke Tan
Phase is important for perceptual quality of speech. However, it seems intractable to directly estimate phase spectra through supervised learning due to their lack of spectrotemporal structure in it. Complex spectral mapping aims to estimat...
Divide and Conquer: A Deep CASA Approach to Talker-independent Monaural Speaker Separation [0.03%]
分而治之:一种端到端的说话人无关单通道语音分离方法
Yuzhou Liu,DeLiang Wang
Yuzhou Liu
We address talker-independent monaural speaker separation from the perspectives of deep learning and computational auditory scene analysis (CASA). Specifically, we decompose the multi-speaker separation task into the stages of simultaneous ...
Deep Learning for Talker-dependent Reverberant Speaker Separation: An Empirical Study [0.03%]
基于深度学习的说话人依赖型混响声源分离方法实验研究
Masood Delfarah,DeLiang Wang
Masood Delfarah
Speaker separation refers to the problem of separating speech signals from a mixture of simultaneous speakers. Previous studies are limited to addressing the speaker separation problem in anechoic conditions. This paper addresses the proble...
Modal and non-modal voice quality classification using acoustic and electroglottographic features [0.03%]
基于声学特征和电门换图的模态与非模态音色分类
Michal Borsky,Daryush D Mehta,Jarrad H Van Stan et al.
Michal Borsky et al.
The goal of this study was to investigate the performance of different feature types for voice quality classification using multiple classifiers. The study compared the COVAREP feature set; which included glottal source features, frequency ...
Ashwin Bellur,Mounya Elhilali
Ashwin Bellur
One of the unique characteristics of human hearing is its ability to recognize acoustic objects even in presence of severe noise and distortions. In this work, we explore two mechanisms underlying this ability: 1) redundant mapping of acous...
The Temporal Limits Encoder as a Sound Coding Strategy for Bilateral Cochlear Implants [0.03%]
双侧人工耳蜗植入的声码策略的时域极限编码器
Alan Kan,Qinglin Meng
Alan Kan
The difference in binaural benefit between bilateral cochlear implant (CI) users and normal hearing (NH) listeners has typically been attributed to CI sound coding strategies not encoding the acoustic fine structure (FS) interaural time dif...
Causal Deep CASA for Monaural Talker-Independent Speaker Separation [0.03%]
因果深度CASAC在独立说话人的单通道语音分离中的应用
Yuzhou Liu,DeLiang Wang
Yuzhou Liu
Talker-independent monaural speaker separation aims to separate concurrent speakers from a single-microphone recording. Inspired by human auditory scene analysis (ASA) mechanisms, a two-stage deep CASA approach has been proposed recently to...
Structured Sparse Spectral Transforms and Structural Measures for Voice Conversion [0.03%]
基于结构稀疏谱变换和结构测度的音色转换方法
Yunxin Zhao,Mili Kuruvilla-Dugdale,Minguang Song
Yunxin Zhao
We investigate a structured sparse spectral transform method for voice conversion (VC) to perform frequency warping and spectral shaping simultaneously on high-dimensional (D) STRAIGHT spectra. Learning a large transform matrix for high-D d...
Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation [0.03%]
超越理想时频幅度掩蔽的端到端说话人分离模型_conv-tasnet
Yi Luo,Nima Mesgarani
Yi Luo
Single-channel, speaker-independent speech separation methods have recently seen great progress. However, the accuracy, latency, and computational cost of such methods remain insufficient. The majority of the previous methods have formulate...
Gated Residual Networks with Dilated Convolutions for Monaural Speech Enhancement [0.03%]
基于扩张卷积的残差网路在单通道语音增强中的应用
Ke Tan,Jitong Chen,DeLiang Wang
Ke Tan
For supervised speech enhancement, contextual information is important for accurate mask estimation or spectral mapping. However, commonly used deep neural networks (DNNs) are limited in capturing temporal contexts. To leverage long-term co...