
背景动机
以Whisper为代表的预训练的多语言语音识别模型已经达到了很好的效果。但在将这些模型迁移到新的特定语上需要消耗大量算力并且有灾难性遗忘的问题。为解决这两个问题,我们探究了保持原有语种的性能的同时,进行新语种微调的策略。特别地,为了减少训练所需算力,我们的研究首先对比了多种基于LoRA(Low Rank Adaptation)的PEFT(Parameter Efficenet Finetuning)方法的效果,以及它们各自受的灾难性遗忘现象影响的程度。为了解决灾难性遗忘问题,我们利用原始模型的LoRA参数近似原始模型的梯度空间,来对新的样本进行正交梯度下降优化。同时,我们还引入了一个可学习的秩系数来提升训练效率。我们的实验在一个用中文微调的Whisper模型上对维吾尔语和藏语进行迁移,以更小的参数量获得了更好的性能。
我们的方案有以下几大优点。i)无需复习(Rehersal-free):由于隐私原因,我们无法获得原始模型训练所使用的数据,因此我们需要根据原始模型的参数进行继续学习。ii) 参数量小:我们需要继续学习过程中使用的参数量远小于原始模型。iii) 无需任务id:我们需要继续学习的模型无需在任务id来帮助它将原始领域的测试样例和新的领域的测试样例进行区分。
问题定义

方案

图1 整体结构图
正交梯度下降

动态参数分配

实验
表1 使用的数据集

实验配置:我们使用huggingface的PEFT进行所有的实验。Whisper的base和large模型参数量分别为74M和1550M。对于PEFT模型,我们对{
实验结果
与其他方案的对比


探索传统微调的局限性
我们首先研究了在 Whisper 中添加一种新语言(维吾尔语)的可能性,采用了各种不同规模的基线方法:全参数微调、LoRA、AdaLoRA 和冻结编码器在 LLM 适应性中看到的方法。我们用维吾尔语组成的数据集对不同大小的 Whisper 模型进行了微调。本实验关注的语言是汉语和维吾尔语。我们用 Thuyg-20 测试维吾尔语的性能增益。使用 Thuyg-20 测试维吾尔语的性能提升,同时使用 Aishell-1和一个内部中文测试集测试汉语的性能下降。结果详见表 2。较小的 Whisper 模型对灾难性遗忘的容忍度较低,基础模型识别的是Aishell-1 数据集识别为维吾尔语。一个更具挑战性的数据集对灾难性遗忘更为敏感、因为遭受灾难性遗忘的 Whisper 模型仍能仍能将 Aishell-1 转译成中文,但内部中文数据集却被转译成维吾尔语。冻结编码器策略的好处是用较少的可训练参数实现了类似的 WER。虽然冻结编码器策略的好处是以较少的可训练参数实现了类似的 WER,但它对添加新语言时出现的灾难性遗忘现象并没有显著改善。
正交子空间学习的有效性

我们利用这个经过中文微调的的 Whisper-large模型对灾难性遗忘现象进行了进一步研究。在对中文数据进行了 20,000 步的初始训练后,我们又使用各种方法对维吾尔语和藏语数据集各进行了 5,000 步的额外训练。图 2 显示了 WER/CER 的变化。

总结
本文从继续学习视角,以正交梯度下降方案将Whisper预训练模型应用于维吾尔语和藏语语音识别。我们首先探究了多种微调方案以及它们对之前任务产生的遗忘问题。之后,我们验证了模型的Lora矩阵可以用来指导后续的继续学习,同事使用一个可学习的秩矩阵来动态分配训练参数来提高总体性能。这让完全无法获得预训练数据的基座模型进行避免本来任务的遗忘的微调成为可能,提高了模型的可重复使用性。
参考文献
[1] S. Vander Eeckt and H. Van Hamme, “Using adapters to overcome catastrophic forgetting in end-to-end automatic speech recognition,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, pp. 1–5.
[2] M. Yang, I. R. Lane, and S. Watanabe, “Online continual learning of end-to-end speech recognition models,” in Interspeech 2022, 23rd Annual Conference of the International Speech Communication Association, Incheon, Korea, 18-22 September 2022, H. Ko and J. H. L. Hansen, Eds. ISCA, 2022, pp. 2668–2672.
[3] X. Wang, T. Chen, Q. Ge, H. Xia, R. Bao, R. Zheng, Q. Zhang, T. Gui, and X. Huang, “Orthogonal subspace learning for language model continual learning,” in Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, H. Bouamor, J. Pino, and K. Bali, Eds. Association for Computational Linguistics, 2023, pp. 10 658–10 671.
[4] B. Zhang, H. Lv, P. Guo, Q. Shao, C. Yang, L. Xie, X. Xu, H. Bu, X. Chen, C. Zeng et al., “Wenetspeech: A 10000+ hours multidomain mandarin corpus for speech recognition,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, pp. 6182–6186.
[5] H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “Aishell-1: An opensource mandarin speech corpus and a speech recognition baseline,” in 2017 20th conference of the oriental chapter of the international coordinating committee on speech databases and speech I/O systems and assessment (O-COCOSDA). IEEE, 2017, pp. 1–5.
[6] R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” in Proceedings of The 12th Language Resources and Evaluation Conference, LREC 2020, Marseille, France, May 11-16, 2020, pp. 4218–4222.
[7] A. Roze, S. Yin, Z. Zhang, D. Wang, and A. Hamdulla, “Thugy20: A free uyghur speech database,” in NCMMSC’15, 2015.
[8] Y. Zhao, X. Xu, J. Yue, W. Song, X. Li, L. Wu, and Q. Ji, “An open speech resource for tibetan multi-dialect and multitask recognition,” International Journal of Computational Science and Engineering, vol. 22, no. 2-3, pp. 297–304, 2020.
[9] J. N. Senyan Li, Guanyu Li, “XBMU-AMDO31:An open source of Amdo Tibetan speech database and speech recognition baseline system,” in National Conference on Man-Machine Speech Communication,NCMMSC2022, 2022.
[10] U. Hermjakob, J. May, and K. Knight, “Out-of-the-box universal Romanization tool uroman,” in Proceedings of ACL 2018, System Demonstrations, F. Liu and T. Solorio, Eds. Melbourne, Australia: Association for Computational Linguistics, Jul. 2018, pp. 13–18.
[11] R. Aljundi, F. Babiloni, M. Elhoseiny, M. Rohrbach, and T. Tuytelaars, “Memory aware synapses: Learning what (not) to forget,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 139–154.
[12] S. Vander Eeckt and H. Van Hamme, “Using adapters to overcome catastrophic forgetting in end-to-end automatic speech recognition,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, pp. 1–5.
[13] W. Liu, Y. Qin, Z. Peng, and T. Lee, “Sparsely shared lora on Whisper for child speech recognition,” arXiv preprint arXiv:2309.11756, 2023.
[14] M. Yang, I. R. Lane, and S. Watanabe, “Online continual learning of end-to-end speech recognition models,” in Interspeech 2022, 23rd Annual Conference of the International Speech Communication Association, Incheon, Korea, 18-22 September 2022, H. Ko and J. H. L. Hansen, Eds. ISCA, 2022, pp. 2668–2672.
