← 返回资源分享
Multi-Decoder DPRNN: High Accuracy Source Counting and Separation
论文
论文
发布时间2020-11-24
发表arXiv:2011.12022
作者:Mark Hasegawa-Johnson,Junzhe Zhu,Raymond Yeh
详细介绍
We propose an end-to-end trainable approach to single-channel speech separation with unknown number of speakers. Our approach extends the MulCat source separation backbone with additional output heads: a count-head to infer the number of speakers, and decoder-heads for reconstructing the original signals. Beyond the model, we also propose a metric on how to evaluate source separation with variable number of speakers. Specifically, we cleared up the issue on how to evaluate the quality when the ground-truth hasmore or less speakers than the ones predicted by the model. We evaluate our approach on the WSJ0-mix datasets, with mixtures up to five speakers. We demonstrate that our approach outperforms state-of-the-art in counting the number of speakers and remains competitive in quality of reconstructed signals.
代码仓库 (3)
asteroid-team/asteroid/tree/master/egs/wsj0-mix-var/Multi-Decoder-DPRNN官方PyTorch
JunzheJosephZhu/MultiDecoder-DPRNNPyTorch
JunzheJosephZhu/Multi-Decoder-DPRNNPyTorch
