
Module-Based End-to-End Distant Speech Processing: A Case Study of Far-Field Automatic Speech Recognition
背景动机
概 览





端到端系统
端到端模块化系统
a.现有方法。与模块化系统类似,该系统通常由前端模块、特征提取模块和语音识别后端模块组成,但为了解决前文所提到的模块化系统在模块间失配问题上的不足,其训练过程则采用了端到端系统的设计,即仅利用最终的语音识别训练准则对所有模块进行联合优化,同时保持每个模块的可解释性与功能。为了实现这一点,现有研究通常采用基于混合方法设计的前端模块,如基于神经网络的波束形成器、基于神经网络的加权预测误差(Weighted Prediction Error,WPE)去混响算法等[4][20],借助其强误差容忍度和泛化能力,确保在利用语音识别准则进行联合优化的过程中保持原有的信号增强功能。此外,在优化过程中还可能面临训练稳定性问题,因为涉及多个模块的同时优化,在特定模块中的梯度更新可能由于不稳定的数值运算(如复数矩阵求导)而出现极端梯度,导致训练不稳定甚至无法收敛。现有研究通常可采用对角加载项、掩码底值、优化运算算子等多项技术[21]来缓解上述问题。这类系统目前已在单说话人[20]和多说话人场景[22]中均得到了验证,并进一步拓展至将基于预训练模型的特征提取模块引入联合优化过程中[23],取得了显著性能提升。
实验案例

总 结
参考文献
[1] E. C. Cherry, “Some experiments on the recognition of speech, with one and with two ears,” The Journal of the Acoustical Society of America, vol. 25, no. 5, pp. 975–979, 1953.
[2] R. Haeb-Umbach, J. Heymann, L. Drude, S. Watanabe, M. Delcroix, and T. Nakatani, “Far-field automatic speech recognition,” Proceedings of the IEEE, vol. 109, no. 2, pp. 124–148, 2021.
[3] R. Haeb-Umbach, S. Watanabe, T. Nakatani, M. Bacchiani, B. Hoffmeister, M. L. Seltzer, H. Zen, and M. Souden, “Speech processing for digital home assistants: Combining signal processing with deep-learning techniques,” IEEE Signal Processing Magazine, vol. 36, no. 6, pp. 111–124, 2019.
[4] M. L. Seltzer, B. Raj, and R. M. Stern, “Likelihood-maximizing beamforming for robust hands-free speech recognition,” IEEE Transactions on Speech and Audio Processing, vol. 12, no. 5, pp. 489–498, 2004.
[5] S. Boll, “Suppression of acoustic noise in speech using spectral subtraction,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 27, no. 2, pp. 113–120, 1979.
[6] B. D. Van Veen, and K. M. Buckley, “Beamforming: A versatile approach to spatial filtering,” IEEE ASSP Magazine, vol. 5, no. 2, pp. 4–24, 1988.
[7] S. Choi, A. Cichocki, H.-M. Park, and S.-Y. Lee, “Blind source separation and independent component analysis: A review,” Neural Information Processing-Letters and Reviews, vol. 6, no. 1, pp. 1–57, 2005.
[8] D. Wang, and G. J. Brown, “Computational auditory scene analysis: Principles, algorithms, and applications,” Wiley-IEEE Press, 2006.
[9] W. Jiang, and K. Yu, “Speech enhancement with integration of neural homomorphic synthesis and spectral masking,” IEEE/ACM Trans. ASLP., vol. 31, pp. 1758–1770, 2023.
[10] S. Watanabe, M. Mandel, J. Barker, E. Vincent, et al., “CHiME-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings,” in Proc. 6th International Workshop on Speech Processing in Everyday Environments (CHiME 2020), 2020, pp. 1–7.
[11] H. Hermansky, D. P. Ellis, and S. Sharma, “Tandem connectionist feature extraction for conventional HMM systems,” in Proc. IEEE ICASSP, 2000, pp. 1635–1638.
[12] A. Mohamed, H.-y. Lee, L. Borgholt, J. D. Havtorn, J. Edin, C. Igel, K. Kirchhoff, S.-W. Li, K. Livescu, L. Maaløe, T. N. Sainath, and S. Watanabe, “Self-supervised speech representation learning: A review,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1179–1210, 2022.
[13] G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Processing Magazine, vol. 29, no. 6, pp. 82–97, 2012.
[14] J. Li, “Recent advances in end-to-end automatic speech recognition,” APSIPA Transactions on Signal and Information Processing, vol. 11, no. 1, 2022.
[15] Y. Shinohara, “Adversarial multi-task learning of deep neural networks for robust speech recognition,” in Proc. ISCA Interspeech, 2016, pp. 2369–2372.
[16] N. Kanda, Y. Fujita, S. Horiguchi, R. Ikeshita, K. Nagamatsu, and S. Watanabe, “Acoustic modeling for distant multi-talker speech recognition with single-and multi-channel branches,” in Proc. IEEE ICASSP, 2019, pp. 6630–6634.
[17] F.-J. Chang, M. Radfar, A. Mouchtaris, B. King, and S. Kunzmann, “End-to-end multi-channel transformer for speech recognition,” in Proc. IEEE ICASSP, 2021, pp. 5884–5888.
[18] M. Kolbaek, D. Yu, Z. H. Tan, and J. Jensen, “Multitalker speech separation with utterance-level permutation invariant training of deep recurrent neural networks,” IEEE/ACM Trans. ASLP., vol. 25, no. 10, pp. 1901–1913, 2017.
[19] N. Kanda, Y. Gaur, X. Wang, Z. Meng, and T. Yoshioka, “Serialized output training for end-to-end overlapped speech recognition,” Proc. ISCA Interspeech, pp. 2797–2801, 2020.
[20] T. Ochiai, S. Watanabe, T. Hori, and J. R. Hershey, “Multichannel end-to-end speech recognition,” in Proc. ICML, 2017, pp. 2632–2641.
[21] W. Zhang, X. Chang, C. Boeddeker, T. Nakatani, S. Watanabe, and Y. Qian, “End-to-end dereverberation, beamforming, and speech recognition in a cocktail party,” IEEE/ACM Trans. ASLP., vol. 30, pp. 3173–3188, 2022.
[22] X. Chang, W. Zhang, Y. Qian, J. Le Roux, and S. Watanabe, “MIMO-Speech: End-to-end multi-channel multi-speaker speech recognition,” in Proc. IEEE ASRU, 2019, pp. 237–244.
[23] Y. Masuyama, X. Chang, W. Zhang, S. Cornell, Z.-Q. Wang, N. Ono, Y. Qian, and S. Watanabe, “Exploring the integration of speech separation and recognition with self-supervised learning representation,” in Proc. IEEE WASPAA, 2023, pp. 1–5.
