采样点的降低,会让听感没有原来16KHz的细腻,这同样会影响着模型的识别性能,因为能建模的信息量直接少了一半;而且跨信道比对的分数,相比同信道(8KHz vs 8KHz)的会出现一定的分布差异,这会影响着阈值的设定和校准。除了16KHz vs 8KHz跨信道识别,还有这些常见的cross-domain场合:
transfer learning,这里有很多细方向。对于神经网络来说,有少量标签可以做fine-tuning和multi-task learning,无标签的可以做semi-supervised learning自我学习等等;对于backend来说,可以对PLDA做adaptation,可以换mean.vec,甚至是对raw embeddings做CORAL等等。
Voiceai Systems to NIST Sre19 Evaluation: Robust Speaker Recognition on Conversational Telephone Speech目前有从raw xvector上源头进行改造,有从负责打分的PLDA上改造,有从分数结果上改,图来源:Rongjin Li, Dongpeng Chen, and Weibin Zhang, “Voiceai systems to nist sre19 evaluation: Robust speaker recognition on conversational telephone speech,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, pp. 6459–6463.
Daniel Garcia-Romero, Alan McCree, Stephen Shum, Niko Brummer, and Carlos Vaquero, “Unsupervised domain adaptation for i-vector speaker recognition,” in Proceedings of Odyssey: The Speaker and Language Recognition Workshop, 2014, vol. 8.
Kong Aik Lee, Qiongqiong Wang, and Takafumi Koshinaka, “The coral+ algorithm for unsupervised domain adaptation of plda,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2019, pp. 5821–5825.
对于raw xvector,CORAL和fDA的工作比较出名:
Baochen Sun, Jiashi Feng, and Kate Saenko, “Return of frustratingly easy domain adaptation,” in Proceedings of the AAAI Conference on Artificial Intelligence, 2016, vol. 30.Md Jahangir Alam, Gautam Bhattacharya, and Patrick Kenny, “Speaker verification in mismatched conditions with frustratingly easy domain adaptation.,” in Odyssey, 2018, vol. 2018, pp. 176–180.Pierre-Michel Bousquet and Mickael Rouvier, “On robustness of unsupervised domain adaptation for speaker recognition,” in InterSpeech, 2019.