本文经arXiv每日学术速递授权转载
【1】 The Effect of Perceptual Metrics on Music Representation Learning for Genre Classification
标题: 知觉预设对流派分类音乐表示学习的影响
作者:Tashi Namgyal,Alexander Hepburn,Raul Santos-Rodriguez,Valero Laparra,Jesus Malo
备注:arXiv admin note: text overlap with arXiv:2312.03455
链接:点击下载PDF文件
【2】 Using LLM for Real-Time Transcription and Summarization of Doctor-Patient Interactions into ePuskesmas in Indonesia
标题: 使用LLM实时转录和总结印度尼西亚ePuskesmas中的医患互动
作者:Azmul Asmar Irfan,Nur Ahmad Khatim,Mansur M. Arief
链接:点击下载PDF文件
【3】 Weighted Cross-entropy for Low-Resource Languages in Multilingual Speech Recognition
标题: 多语言语音识别中低资源语言的加权交叉信息
作者:Andrés Piñeiro-Martín,Carmen García-Mateo,Laura Docío-Fernández,María del Carmen López-Pérez,Georg Rehm
Journal-ref:Proceedings of Interspeech 2024
链接:点击下载PDF文件
【4】 Spelling Correction through Rewriting of Non-Autoregressive ASR Lattices
标题: 通过重写非自回归ASC格进行拼写纠正
作者:Leonid Velikovich,Christopher Li,Diamantino Caseiro,Shankar Kumar,Pat Rondon,Kandarp Joshi,Xavier Velez
备注:8 pages, 7 figures
链接:点击下载PDF文件
【5】 FastTalker: Jointly Generating Speech and Conversational Gestures from Text
标题: FastTalker:从文本联合生成语音和对话手势
作者:Zixin Guo,Jian Zhang
备注:European Conference on Computer Vision Workshop
链接:点击下载PDF文件
【6】 Revisiting Acoustic Features for Robust ASR
标题: 重新审视稳健的ASB的声学特征
作者:Muhammad A. Shah,Bhiksha Raj
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
【7】 MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events
标题: MT 2 KD:面向语音、扬声器和音频事件的通用编码器
作者:Xiaoyu Yang,Qiujia Li,Chao Zhang,Phil Woodland
链接:点击下载PDF文件
【8】 Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling
标题: 具有多视图伪标签的语音半监督认知状态分类
作者:Yuanchao Li,Zixing Zhang,Jing Han,Peter Bell,Catherine Lai
链接:点击下载PDF文件
【9】 Cross-lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
标题: 跨语言语音情感识别:人类与自我监督模型
作者:Zhichen Han,Tianqi Geng,Hui Feng,Jiahong Yuan,Korin Richmond,Yuanchao Li
链接:点击下载PDF文件
【10】 Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings
标题: 多通道多方会议的模块化扬声器拨号中的融合空间线索
作者:Ruoyu Wang,Shutong Niu,Gaobin Yang,Jun Du,Shuangqing Qian,Tian Gao,Jia Pan
备注:5 pages, Submitted to ICASSP 2025
链接:点击下载PDF文件
【11】 Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions
标题: 基于语言模型的文本转语音中的情感维度控制:跨越人类情感的广泛范围
作者:Kun Zhou,You Zhang,Shengkui Zhao,Hao Wang,Zexu Pan,Dianwen Ng,Chong Zhang,Chongjia Ni,Yukun Ma,Trung Hieu Nguyen,Jia Qi Yip,Bin Ma
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
【12】 Speech Recognition Rescoring with Large Speech-Text Foundation Models
标题: 使用大语音文本基础模型的语音识别重新评分
作者:Prashanth Gurunath Shivakumar,Jari Kolehmainen,Aditya Gourav,Yi Gu,Ankur Gandhe,Ariya Rastrow,Ivan Bulyko
链接:点击下载PDF文件
【13】 Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation
标题: 启用听觉大语言模型进行自动语音质量评估
作者:Siyin Wang,Wenyi Yu,Yudong Yang,Changli Tang,Yixuan Li,Jimin Zhuang,Xianzhao Chen,Xiaohai Tian,Jun Zhang,Guangzhi Sun,Lu Lu,Chao Zhang
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
【14】 Towards Within-Class Variation in Alzheimer's Disease Detection from Spontaneous Speech
标题: 自发言语检测阿尔茨海默病的班内变异
作者:Jiawen Kang,Dongrui Han,Lingwei Meng,Jingyan Zhou,Jinchao Li,Xixin Wu,Helen Meng
链接:点击下载PDF文件
【15】 A Literature Review of Keyword Spotting Technologies for Urdu
标题: 乌尔都语关键词识别技术文献综述
作者:Syed Muhammad Aqdas Rizvi
链接:点击下载PDF文件
【16】 How Redundant Is the Transformer Stack in Speech Representation Models?
标题: 语音表示模型中的Transformer Stack有多冗余?
作者:Teresa Dorszewski,Albert Kjøller Jacobsen,Lenka Tětková,Lars Kai Hansen
链接:点击下载PDF文件
【17】 Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
标题: 在计算预算下有效训练自我监督语音基础模型
作者:Andy T. Liu,Yi-Cheng Lin,Haibin Wu,Stefan Winkler,Hung-yi Lee
备注:To appear in SLT 2024
链接:点击下载PDF文件
【18】 Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
标题: 无监督单词发现:使用集群进行边界检测与动态编程
作者:Simon Malan,Benjamin van Niekerk,Herman Kamper
备注:3 figures, 3 tables
链接:点击下载PDF文件
标题: MT 2 KD:面向语音、扬声器和音频事件的通用编码器
作者:Xiaoyu Yang,Qiujia Li,Chao Zhang,Phil Woodland
链接:点击下载PDF文件
【2】 Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling
标题: 具有多视图伪标签的语音半监督认知状态分类
作者:Yuanchao Li,Zixing Zhang,Jing Han,Peter Bell,Catherine Lai
链接:点击下载PDF文件
【3】 Cross-lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
标题: 跨语言语音情感识别:人类与自我监督模型
作者:Zhichen Han,Tianqi Geng,Hui Feng,Jiahong Yuan,Korin Richmond,Yuanchao Li
链接:点击下载PDF文件
【4】 Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings
标题: 多通道多方会议的模块化扬声器拨号中的融合空间线索
作者:Ruoyu Wang,Shutong Niu,Gaobin Yang,Jun Du,Shuangqing Qian,Tian Gao,Jia Pan
备注:5 pages, Submitted to ICASSP 2025
链接:点击下载PDF文件
【5】 Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions
标题: 基于语言模型的文本转语音中的情感维度控制:跨越人类情感的广泛范围
作者:Kun Zhou,You Zhang,Shengkui Zhao,Hao Wang,Zexu Pan,Dianwen Ng,Chong Zhang,Chongjia Ni,Yukun Ma,Trung Hieu Nguyen,Jia Qi Yip,Bin Ma
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
【6】 Speech Recognition Rescoring with Large Speech-Text Foundation Models
标题: 使用大语音文本基础模型的语音识别重新评分
作者:Prashanth Gurunath Shivakumar,Jari Kolehmainen,Aditya Gourav,Yi Gu,Ankur Gandhe,Ariya Rastrow,Ivan Bulyko
链接:点击下载PDF文件
【7】 Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation
标题: 启用听觉大语言模型进行自动语音质量评估
作者:Siyin Wang,Wenyi Yu,Yudong Yang,Changli Tang,Yixuan Li,Jimin Zhuang,Xianzhao Chen,Xiaohai Tian,Jun Zhang,Guangzhi Sun,Lu Lu,Chao Zhang
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
【8】 Towards Within-Class Variation in Alzheimer's Disease Detection from Spontaneous Speech
标题: 自发言语检测阿尔茨海默病的班内变异
作者:Jiawen Kang,Dongrui Han,Lingwei Meng,Jingyan Zhou,Jinchao Li,Xixin Wu,Helen Meng
链接:点击下载PDF文件
【9】 A Literature Review of Keyword Spotting Technologies for Urdu
标题: 乌尔都语关键词识别技术文献综述
作者:Syed Muhammad Aqdas Rizvi
链接:点击下载PDF文件
【10】 How Redundant Is the Transformer Stack in Speech Representation Models?
标题: 语音表示模型中的Transformer Stack有多冗余?
作者:Teresa Dorszewski,Albert Kjøller Jacobsen,Lenka Tětková,Lars Kai Hansen
链接:点击下载PDF文件
【11】 Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
标题: 在计算预算下有效训练自我监督语音基础模型
作者:Andy T. Liu,Yi-Cheng Lin,Haibin Wu,Stefan Winkler,Hung-yi Lee
备注:To appear in SLT 2024
链接:点击下载PDF文件
【12】 The Effect of Perceptual Metrics on Music Representation Learning for Genre Classification
标题: 知觉预设对流派分类音乐表示学习的影响
作者:Tashi Namgyal,Alexander Hepburn,Raul Santos-Rodriguez,Valero Laparra,Jesus Malo
备注:arXiv admin note: text overlap with arXiv:2312.03455
链接:点击下载PDF文件
【13】 Using LLM for Real-Time Transcription and Summarization of Doctor-Patient Interactions into ePuskesmas in Indonesia
标题: 使用LLM实时转录和总结印度尼西亚ePuskesmas中的医患互动
作者:Azmul Asmar Irfan,Nur Ahmad Khatim,Mansur M. Arief
链接:点击下载PDF文件
【14】 Weighted Cross-entropy for Low-Resource Languages in Multilingual Speech Recognition
标题: 多语言语音识别中低资源语言的加权交叉信息
作者:Andrés Piñeiro-Martín,Carmen García-Mateo,Laura Docío-Fernández,María del Carmen López-Pérez,Georg Rehm
Journal-ref:Proceedings of Interspeech 2024
链接:点击下载PDF文件
【15】 Spelling Correction through Rewriting of Non-Autoregressive ASR Lattices
标题: 通过重写非自回归ASC格进行拼写纠正
作者:Leonid Velikovich,Christopher Li,Diamantino Caseiro,Shankar Kumar,Pat Rondon,Kandarp Joshi,Xavier Velez
备注:8 pages, 7 figures
链接:点击下载PDF文件
【16】 FastTalker: Jointly Generating Speech and Conversational Gestures from Text
标题: FastTalker:从文本联合生成语音和对话手势
作者:Zixin Guo,Jian Zhang
备注:European Conference on Computer Vision Workshop
链接:点击下载PDF文件
【17】 Revisiting Acoustic Features for Robust ASR
标题: 重新审视稳健的ASB的声学特征
作者:Muhammad A. Shah,Bhiksha Raj
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
标题: 知觉预设对流派分类音乐表示学习的影响
作者:Tashi Namgyal,Alexander Hepburn,Raul Santos-Rodriguez,Valero Laparra,Jesus Malo
备注:arXiv admin note: text overlap with arXiv:2312.03455
链接:点击下载PDF文件
摘要:自然信号的主观质量可以用客观感知度量来近似。感知度量旨在近似人类观察者的感知行为,通常反映自然信号和神经通路中的结构。用感知度量作为损失函数训练的模型可以从这些度量中的结构中捕获感知上有意义的特征。我们证明,使用从感知损失训练的自动编码器中提取的特征可以提高音乐理解任务(即流派分类)的性能,而不是在学习分类器时直接使用这些指标作为距离。这一结果表明,改进的泛化到新的信号时,使用知觉指标作为损失函数的表示学习。摘要:The subjective quality of natural signals can be approximated with objective perceptual metrics. Designed to approximate the perceptual behaviour of human observers, perceptual metrics often reflect structures found in natural signals and neurological pathways. Models trained with perceptual metrics as loss functions can capture perceptually meaningful features from the structures held within these metrics. We demonstrate that using features extracted from autoencoders trained with perceptual losses can improve performance on music understanding tasks, i.e. genre classification, over using these metrics directly as distances when learning a classifier. This result suggests improved generalisation to novel signals when using perceptual metrics as loss functions for representation learning.
【2】 Using LLM for Real-Time Transcription and Summarization of Doctor-Patient Interactions into ePuskesmas in Indonesia
标题: 使用LLM实时转录和总结印度尼西亚ePuskesmas中的医患互动
作者:Azmul Asmar Irfan,Nur Ahmad Khatim,Mansur M. Arief
链接:点击下载PDF文件
摘要:导致Puskesmas效率低下的关键问题之一是医患互动的耗时性。医生需要进行全面的咨询,其中包括诊断病人的病情,提供治疗建议,并将详细的笔记转录成医疗记录。在语言背景不同的地区,医生往往不得不提出澄清问题,进一步延长了这个过程。虽然诊断是必不可少的,但转录和总结通常可以使用人工智能自动化,以提高时间效率,帮助医生提高护理质量,并实现早期诊断和干预。本文提出了一种使用本地化的大语言模型(LLM)来转录,翻译和总结医患对话的解决方案。我们利用Whisper模型进行转录和GPT-3,将它们汇总为ePuskemas医疗记录格式。该系统是作为现有Web浏览器扩展的附加组件实现的,允许医生在交谈时填写患者表格。通过利用该解决方案进行实时转录、翻译和总结,医生可以缩短患者护理的周转时间,同时提高记录的质量,使其在未来的访问中变得更加详细和有见地。这项创新解决了印度尼西亚设施过度拥挤和医疗保健提供者行政负担等挑战。我们相信,这一解决方案将帮助医生节省时间,提供更好的护理,并生成更准确的医疗记录,这是实现医疗保健现代化的重要一步,也是确保患者即使在资源有限的情况下也能获得及时、高质量护理的重要一步。摘要:One of the key issues contributing to inefficiency in Puskesmas is the time-consuming nature of doctor-patient interactions. Doctors need to conduct thorough consultations, which include diagnosing the patient's condition, providing treatment advice, and transcribing detailed notes into medical records. In regions with diverse linguistic backgrounds, doctors often have to ask clarifying questions, further prolonging the process. While diagnosing is essential, transcription and summarization can often be automated using AI to improve time efficiency and help doctors enhance care quality and enable early diagnosis and intervention. This paper proposes a solution using a localized large language model (LLM) to transcribe, translate, and summarize doctor-patient conversations. We utilize the Whisper model for transcription and GPT-3 to summarize them into the ePuskemas medical records format. This system is implemented as an add-on to an existing web browser extension, allowing doctors to fill out patient forms while talking. By leveraging this solution for real-time transcription, translation, and summarization, doctors can improve the turnaround time for patient care while enhancing the quality of records, which become more detailed and insightful for future visits. This innovation addresses challenges like overcrowded facilities and the administrative burden on healthcare providers in Indonesia. We believe this solution will help doctors save time, provide better care, and produce more accurate medical records, representing a significant step toward modernizing healthcare and ensuring patients receive timely, high-quality care, even in resource-constrained settings.
【3】 Weighted Cross-entropy for Low-Resource Languages in Multilingual Speech Recognition
标题: 多语言语音识别中低资源语言的加权交叉信息
作者:Andrés Piñeiro-Martín,Carmen García-Mateo,Laura Docío-Fernández,María del Carmen López-Pérez,Georg Rehm
Journal-ref:Proceedings of Interspeech 2024
链接:点击下载PDF文件
摘要:本文讨论了将低资源语言集成到多语言自动语音识别(ASR)系统中的挑战。我们引入了加权交叉熵的一种新应用,通常用于不平衡的数据集,以促进在持续多语言学习的背景下将低资源语言集成到预训练的多语言ASR模型中。我们在五种高资源语言和一种低资源语言上对Whisper多语言ASR模型进行了微调,采用了语言加权动态交叉熵和数据增强。结果显示,与没有应用我们的方法的微调模型相比,低资源语言的单词错误率(WER)降低了6.69%,与原始Whisper模型相比,WER降低了48.86%。此外,我们的方法在六种语言中平均降低了3.29%的WER,表明高资源语言没有退化。摘要:This paper addresses the challenge of integrating low-resource languages into multilingual automatic speech recognition (ASR) systems. We introduce a novel application of weighted cross-entropy, typically used for unbalanced datasets, to facilitate the integration of low-resource languages into pre-trained multilingual ASR models within the context of continual multilingual learning. We fine-tune the Whisper multilingual ASR model on five high-resource languages and one low-resource language, employing language-weighted dynamic cross-entropy and data augmentation. The results show a remarkable 6.69% word error rate (WER) reduction for the low-resource language compared to the fine-tuned model without applying our approach, and a 48.86% WER reduction compared to the original Whisper model. In addition, our approach yields an average WER reduction of 3.29% across the six languages, showing no degradation for the high-resource languages.
【4】 Spelling Correction through Rewriting of Non-Autoregressive ASR Lattices
标题: 通过重写非自回归ASC格进行拼写纠正
作者:Leonid Velikovich,Christopher Li,Diamantino Caseiro,Shankar Kumar,Pat Rondon,Kandarp Joshi,Xavier Velez
备注:8 pages, 7 figures
链接:点击下载PDF文件
摘要:对于端到端自动语音识别(ASR)模型,识别个人或罕见短语可能很困难。提高准确性的一种有希望的方法是通过ASR网格的拼写校正(或重写),其中潜在的错误识别短语被替换为声学上相似和上下文相关的替代品。然而,由于非自回归、上下文无关的波束搜索产生的噪声假设,用联结主义时间分类(CTC)训练的ASR模型的重写是具有挑战性的。 我们提出了一种有限状态转换器(FST)技术重写基于转换器的CTC模型生成的字段格。我们的算法进行直接从词片到音素的字素到音素(G2P)转换,避免明确的单词表示和利用CTC晶格的丰富性。我们的方法不需要重新训练或修改ASR模型。我们实现了高达15.2%的相对减少句子错误率(SER)的测试集与上下文相关的实体。摘要:For end-to-end Automatic Speech Recognition (ASR) models, recognizing personal or rare phrases can be hard. A promising way to improve accuracy is through spelling correction (or rewriting) of the ASR lattice, where potentially misrecognized phrases are replaced with acoustically similar and contextually relevant alternatives. However, rewriting is challenging for ASR models trained with connectionist temporal classification (CTC) due to noisy hypotheses produced by a non-autoregressive, context-independent beam search. We present a finite-state transducer (FST) technique for rewriting wordpiece lattices generated by Transformer-based CTC models. Our algorithm performs grapheme-to-phoneme (G2P) conversion directly from wordpieces into phonemes, avoiding explicit word representations and exploiting the richness of the CTC lattice. Our approach requires no retraining or modification of the ASR model. We achieved up to a 15.2% relative reduction in sentence error rate (SER) on a test set with contextually relevant entities.
【5】 FastTalker: Jointly Generating Speech and Conversational Gestures from Text
标题: FastTalker:从文本联合生成语音和对话手势
作者:Zixin Guo,Jian Zhang
备注:European Conference on Computer Vision Workshop
链接:点击下载PDF文件
摘要:从文本脚本生成3D人类手势和语音对于创建逼真的谈话化身至关重要。一种解决方案是利用文本到语音(TTS)和语音到手势(STG)的单独管道,但是这种方法遭受语音和手势的不良对齐以及缓慢的推理时间。在本文中,我们介绍FastTalker,一个高效和有效的框架,同时产生高品质的语音音频和3D人类手势在高推理速度。我们的关键见解是重用的中间功能,从语音合成手势生成,因为这些功能包含更精确的节奏信息比功能重新提取生成的语音。具体来说,1)我们提出了一个端到端的框架,同时生成语音波形和全身手势,使用中间语音特征,如音高,起始,能量和持续时间直接用于手势解码; 2)我们重新设计了因果网络架构,以消除对未来输入的依赖性,以用于实际应用; 3)我们采用基于强化学习的神经结构搜索(NAS),通过优化我们的网络结构来提高性能和推理速度。在BEAT 2数据集上的实验结果表明,FastTalker在语音合成和手势生成方面都达到了最先进的性能,在NVIDIA 3090上每秒处理语音和手势的速度为0.17秒。摘要:Generating 3D human gestures and speech from a text script is critical for creating realistic talking avatars. One solution is to leverage separate pipelines for text-to-speech (TTS) and speech-to-gesture (STG), but this approach suffers from poor alignment of speech and gestures and slow inference times. In this paper, we introduce FastTalker, an efficient and effective framework that simultaneously generates high-quality speech audio and 3D human gestures at high inference speeds. Our key insight is reusing the intermediate features from speech synthesis for gesture generation, as these features contain more precise rhythmic information than features re-extracted from generated speech. Specifically, 1) we propose an end-to-end framework that concurrently generates speech waveforms and full-body gestures, using intermediate speech features such as pitch, onset, energy, and duration directly for gesture decoding; 2) we redesign the causal network architecture to eliminate dependencies on future inputs for real applications; 3) we employ Reinforcement Learning-based Neural Architecture Search (NAS) to enhance both performance and inference speed by optimizing our network architecture. Experimental results on the BEAT2 dataset demonstrate that FastTalker achieves state-of-the-art performance in both speech synthesis and gesture generation, processing speech and gestures in 0.17 seconds per second on an NVIDIA 3090.
【6】 Revisiting Acoustic Features for Robust ASR
标题: 重新审视稳健的ASB的声学特征
作者:Muhammad A. Shah,Bhiksha Raj
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:自动语音识别(ASR)系统必须对现实世界环境中存在的各种类型的噪声具有鲁棒性,包括环境噪声,房间脉冲响应,特殊效果以及恶意行为者的攻击(对抗性攻击)。最近的工作试图通过开发新的深度神经网络(DNN)并为其管理各种训练数据集,同时使用相对简单的声学特征来提高准确性和鲁棒性。虽然这种方法提高了对训练数据中存在的噪声类型的鲁棒性,但它对不可见噪声的鲁棒性有限,对对抗性攻击的鲁棒性可以忽略不计。在本文中,我们重新审视了早期作品的方法,这些作品开发了受生物听觉感知启发的声学特征,可用于执行准确和强大的ASR。相比之下,具体来说,我们评估ASR的准确性和鲁棒性的几个生物启发的声学特征。除了一些功能,从以前的作品,如伽玛滤波器组功能(GammSpec),我们还提出了两个新的声学功能,称为频率掩蔽频谱图(FreqMask)和伽玛频谱图的差异(DoGSpec)来模拟神经心理现象的频率掩蔽和横向抑制。在不同模型和数据集上的实验表明:(1)DoGSpec比非常流行的对数梅尔频谱图(LogMelSpec)具有更好的鲁棒性,并且精度下降最小;(2)GammSpec在语音鲁棒性基准测试中对非对抗性噪声具有更好的精度和鲁棒性,但在对抗性攻击方面优于DoGSpec。摘要:Automatic Speech Recognition (ASR) systems must be robust to the myriad types of noises present in real-world environments including environmental noise, room impulse response, special effects as well as attacks by malicious actors (adversarial attacks). Recent works seek to improve accuracy and robustness by developing novel Deep Neural Networks (DNNs) and curating diverse training datasets for them, while using relatively simple acoustic features. While this approach improves robustness to the types of noise present in the training data, it confers limited robustness against unseen noises and negligible robustness to adversarial attacks. In this paper, we revisit the approach of earlier works that developed acoustic features inspired by biological auditory perception that could be used to perform accurate and robust ASR. In contrast, Specifically, we evaluate the ASR accuracy and robustness of several biologically inspired acoustic features. In addition to several features from prior works, such as gammatone filterbank features (GammSpec), we also propose two new acoustic features called frequency masked spectrogram (FreqMask) and difference of gammatones spectrogram (DoGSpec) to simulate the neuro-psychological phenomena of frequency masking and lateral suppression. Experiments on diverse models and datasets show that (1) DoGSpec achieves significantly better robustness than the highly popular log mel spectrogram (LogMelSpec) with minimal accuracy degradation, and (2) GammSpec achieves better accuracy and robustness to non-adversarial noises from the Speech Robust Bench benchmark, but it is outperformed by DoGSpec against adversarial attacks.
【7】 MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events
标题: MT 2 KD:面向语音、扬声器和音频事件的通用编码器
作者:Xiaoyu Yang,Qiujia Li,Chao Zhang,Phil Woodland
链接:点击下载PDF文件
摘要:随着深度学习的进步,用于语音和音频处理的端到端(E2E)单任务模型的性能一直在不断提高。 然而,构建一个在多个任务上具有高性能的通用模型仍然具有挑战性,因为不同的语音和音频处理任务通常需要不同的训练数据、输入特征或模型架构来实现最佳性能。 在这项工作中,MT2KD,一种新的两阶段多任务学习框架,提出了建立一个通用的语音和音频编码器,共同执行三个基本任务:自动语音识别(ASR),音频标记(AT)和说话人验证(SV)。在第一阶段中,多教师知识蒸馏(KD)被应用于使用相同的未标记数据将三个单任务高性能教师编码器的特征空间对齐到单个学生编码器中。在第二阶段,通过初始化第一阶段的模型并在每个单个任务的单独标记数据上进行训练来进行多任务监督微调。 实验表明,提出的多任务训练管道的性能显着优于从头开始进行多任务学习训练的基线模型。最终的系统在ASR、AT和SV上实现了良好的性能:与性能最好的单任务编码器相比,ASR上的相对字错误率增加不到4%,AT上的平均平均精度仅低1.9,SV上的绝对等错误率高0.23%,仅使用66M的总模型参数。摘要:With the advances in deep learning, the performance of end-to-end (E2E) single-task models for speech and audio processing has been constantly improving. However, it is still challenging to build a general-purpose model with high performance on multiple tasks, since different speech and audio processing tasks usually require different training data, input features, or model architectures to achieve optimal performance. In this work, MT2KD, a novel two-stage multi-task learning framework is proposed to build a general-purpose speech and audio encoder that jointly performs three fundamental tasks: automatic speech recognition (ASR), audio tagging (AT) and speaker verification (SV). In the first stage, multi-teacher knowledge distillation (KD) is applied to align the feature spaces of three single-task high-performance teacher encoders into a single student encoder using the same unlabelled data. In the second stage, multi-task supervised fine-tuning is carried out by initialising the model from the first stage and training on the separate labelled data of each single task. Experiments demonstrate that the proposed multi-task training pipeline significantly outperforms a baseline model trained with multi-task learning from scratch. The final system achieves good performance on ASR, AT and SV: with less than 4% relative word-error-rate increase on ASR, only 1.9 lower mean averaged precision on AT and 0.23% absolute higher equal error rate on SV compared to the best-performing single-task encoders, using only a 66M total model parameters.
【8】 Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling
标题: 具有多视图伪标签的语音半监督认知状态分类
作者:Yuanchao Li,Zixing Zhang,Jing Han,Peter Bell,Catherine Lai
链接:点击下载PDF文件
摘要:缺乏标记数据是语音分类任务中的常见挑战,特别是那些需要大量主观评估的任务,例如认知状态分类。在这项工作中,我们提出了一个半监督学习(SSL)框架,引入了一种新的多视图伪标记方法,该方法利用声学和语言特征来选择最有信心的数据来训练分类模型。在声学上,未标记的数据与使用Frechet音频距离的标记数据进行比较,该距离是从多个音频编码器生成的嵌入计算的。在语言学上,大型语言模型会根据我们提出的特定任务知识来修改自动语音识别transmittance和预测标签。当来自两个源的伪标签对齐时,识别高置信度数据,而不匹配则被视为低置信度数据。然后训练双峰分类器以迭代地标记低置信度数据,直到满足预定义的标准。我们评估我们的SSL框架的情绪识别和痴呆症检测任务。实验结果表明,与仅使用30%标记数据的完全监督学习相比,我们的方法具有竞争力的性能,并且显著优于两个选定的基线。摘要:The lack of labeled data is a common challenge in speech classification tasks, particularly those requiring extensive subjective assessment, such as cognitive state classification. In this work, we propose a Semi-Supervised Learning (SSL) framework, introducing a novel multi-view pseudo-labeling method that leverages both acoustic and linguistic characteristics to select the most confident data for training the classification model. Acoustically, unlabeled data are compared to labeled data using the Frechet audio distance, calculated from embeddings generated by multiple audio encoders. Linguistically, large language models are prompted to revise automatic speech recognition transcriptions and predict labels based on our proposed task-specific knowledge. High-confidence data are identified when pseudo-labels from both sources align, while mismatches are treated as low-confidence data. A bimodal classifier is then trained to iteratively label the low-confidence data until a predefined criterion is met. We evaluate our SSL framework on emotion recognition and dementia detection tasks. Experimental results demonstrate that our method achieves competitive performance compared to fully supervised learning using only 30% of the labeled data and significantly outperforms two selected baselines.
【9】 Cross-lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
标题: 跨语言语音情感识别:人类与自我监督模型
作者:Zhichen Han,Tianqi Geng,Hui Feng,Jiahong Yuan,Korin Richmond,Yuanchao Li
链接:点击下载PDF文件
摘要:利用自监督学习(SSL)模型进行语音情感识别(SER)已被证明是有效的,但有限的研究探索了跨语言的情况。本研究提出了人类表现和SSL模型之间的比较分析,从逐层分析和探索单语,跨语言和迁移学习环境中的参数有效的微调策略开始。我们进一步比较SER能力的模型和人类在话语和段水平。此外,我们还通过人的评价研究了方言对跨语言SER的影响。我们的研究结果表明,模型,适当的知识转移,可以适应目标语言,并实现与本族语者相当的性能。我们还证明了显着的影响,方言对SER的个人没有事先的语言和语言学背景。此外,人类和模型在不同的情绪下表现出不同的行为。这些结果为SSL模型的跨语言SER能力提供了新的见解,强调了它们与人类情感感知的相似性和差异性。摘要:Utilizing Self-Supervised Learning (SSL) models for Speech Emotion Recognition (SER) has proven effective, yet limited research has explored cross-lingual scenarios. This study presents a comparative analysis between human performance and SSL models, beginning with a layer-wise analysis and an exploration of parameter-efficient fine-tuning strategies in monolingual, cross-lingual, and transfer learning contexts. We further compare the SER ability of models and humans at both utterance- and segment-levels. Additionally, we investigate the impact of dialect on cross-lingual SER through human evaluation. Our findings reveal that models, with appropriate knowledge transfer, can adapt to the target language and achieve performance comparable to native speakers. We also demonstrate the significant effect of dialect on SER for individuals without prior linguistic and paralinguistic background. Moreover, both humans and models exhibit distinct behaviors across different emotions. These results offer new insights into the cross-lingual SER capabilities of SSL models, underscoring both their similarities to and differences from human emotion perception.
【10】 Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings
标题: 多通道多方会议的模块化扬声器拨号中的融合空间线索
作者:Ruoyu Wang,Shutong Niu,Gaobin Yang,Jun Du,Shuangqing Qian,Tian Gao,Jia Pan
备注:5 pages, Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:None摘要:Although fully end-to-end speaker diarization systems have made significant progress in recent years, modular systems often achieve superior results in real-world scenarios due to their greater adaptability and robustness. Historically, modular speaker diarization methods have seldom discussed how to leverage spatial cues from multi-channel speech. This paper proposes a three-stage modular system to enhance single-channel neural speaker diarization systems and recognition performance by utilizing spatial cues from multi-channel speech to provide more accurate initialization for each stage of neural speaker diarization (NSD) decoding: (1) Overlap detection and continuous speech separation (CSS) on multi-channel speech are used to obtain cleaner single speaker speech segments for clustering, followed by the first NSD decoding pass. (2) The results from the first pass initialize a complex Angular Central Gaussian Mixture Model (cACGMM) to estimate speaker-wise masks on multi-channel speech, and through Overlap-add and Mask-to-VAD, achieve initialization with lower speaker error (SpkErr), followed by the second NSD decoding pass. (3) The second decoding results are used for guided source separation (GSS), recognizing and filtering short segments containing less one word to obtain cleaner speech segments, followed by re-clustering and the final NSD decoding pass. We presented the progressively explored evaluation results from the CHiME-8 NOTSOFAR-1 (Natural Office Talkers in Settings Of Far-field Audio Recordings) challenge, demonstrating the effectiveness of our system and its contribution to improving recognition performance. Our final system achieved the first place in the challenge.
【11】 Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions
标题: 基于语言模型的文本转语音中的情感维度控制:跨越人类情感的广泛范围
作者:Kun Zhou,You Zhang,Shengkui Zhao,Hao Wang,Zexu Pan,Dianwen Ng,Chong Zhang,Chongjia Ni,Yukun Ma,Trung Hieu Nguyen,Jia Qi Yip,Bin Ma
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:由于情感的固有复杂性以及情感语音数据集和模型的局限性,当前的情感文本到语音(TTS)系统在模仿广泛的人类情感方面面临挑战。本文提出了一个TTS框架,有利于控制的乐趣,唤醒,和优势,并可以合成的情感风格的多样性,而不需要任何情感语音数据在TTS训练。我们训练的情感属性预测器只使用分类标签的语音数据,与心理学研究,并结合锚定降维的自我监督学习(SSL)功能。TTS框架通过自回归语言模型将文本输入转换为语音标记,并使用伪情感维度来指导细粒度声学细节的并行预测。在LibriTTS数据集上进行的实验表明,我们的框架可以通过有效控制情感维度来合成具有增强的自然度和各种情感风格的语音,即使在TTS训练期间不包含任何情感语音。摘要:Current emotional text-to-speech (TTS) systems face challenges in mimicking a broad spectrum of human emotions due to the inherent complexity of emotions and limitations in emotional speech datasets and models. This paper proposes a TTS framework that facilitates control over pleasure, arousal, and dominance, and can synthesize a diversity of emotional styles without requiring any emotional speech data during TTS training. We train an emotional attribute predictor using only categorical labels from speech data, aligning with psychological research and incorporating anchored dimensionality reduction on self-supervised learning (SSL) features. The TTS framework converts text inputs into phonetic tokens via an autoregressive language model and uses pseudo-emotional dimensions to guide the parallel prediction of fine-grained acoustic details. Experiments conducted on the LibriTTS dataset demonstrate that our framework can synthesize speech with enhanced naturalness and a variety of emotional styles by effectively controlling emotional dimensions, even without the inclusion of any emotional speech during TTS training.
【12】 Speech Recognition Rescoring with Large Speech-Text Foundation Models
标题: 使用大语音文本基础模型的语音识别重新评分
作者:Prashanth Gurunath Shivakumar,Jari Kolehmainen,Aditya Gourav,Yi Gu,Ankur Gandhe,Ariya Rastrow,Ivan Bulyko
链接:点击下载PDF文件
摘要:大型语言模型(LLM)已经证明了通过利用大量文本数据来理解人类语言的能力。自动语音识别(ASR)系统通常受到可用的转录语音数据的限制,并受益于使用LLM的第二遍重新评分。近年来,多模态大型语言模型,特别是语音和文本基础模型,已经表现出强大的口语理解能力。语音-文本基础模型利用语音和文本模态中的大量未标记和标记数据来建模人类语言。在这项工作中,我们提出了新的技术,使用多模态LLM的ASR重新评分。我们还探索了区分训练,以进一步提高基础模型的再评分性能。我们证明了语音-文本LLM中的跨模态知识转移可以有益于重新评分。我们的实验证明了高达20%的相对改善耳语大ASR和高达15%的相对改善,在纯文本LLM。摘要:Large language models (LLM) have demonstrated the ability to understand human language by leveraging large amount of text data. Automatic speech recognition (ASR) systems are often limited by available transcribed speech data and benefit from a second pass rescoring using LLM. Recently multi-modal large language models, particularly speech and text foundational models have demonstrated strong spoken language understanding. Speech-Text foundational models leverage large amounts of unlabelled and labelled data both in speech and text modalities to model human language. In this work, we propose novel techniques to use multi-modal LLM for ASR rescoring. We also explore discriminative training to further improve the foundational model rescoring performance. We demonstrate cross-modal knowledge transfer in speech-text LLM can benefit rescoring. Our experiments demonstrate up-to 20% relative improvements over Whisper large ASR and up-to 15% relative improvements over text-only LLM.
【13】 Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation
标题: 启用听觉大语言模型进行自动语音质量评估
作者:Siyin Wang,Wenyi Yu,Yudong Yang,Changli Tang,Yixuan Li,Jimin Zhuang,Xianzhao Chen,Xiaohai Tian,Jun Zhang,Guangzhi Sun,Lu Lu,Chao Zhang
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:语音质量评估通常需要从多个方面评估音频,例如平均意见得分(MOS)和说话人相似度(SIM)等,这对于使用为单个任务设计的一个小模型来覆盖可能是具有挑战性的。在本文中,我们建议利用最近推出的听觉大语言模型(LLM)的自动语音质量评估。通过使用特定任务的提示,听觉LLM被微调来预测MOS,SIM和A B测试结果,这些测试结果通常用于评估文本到语音系统。此外,微调的听觉LLM能够生成自然语言描述,评估噪音,失真,不连续性和整体质量等方面,提供更多可解释的输出。在NISQA、BVCC、SOMOS和VoxSim语音质量数据集上进行了大量的实验,使用了开源的听觉LLM,如SALMONN、Qwen-Audio和Qwen 2-Audio。对于自然语言描述任务,还评估了商业模型Google Gemini 1.5 Pro。结果表明,与最先进的特定于任务的小模型相比,听觉LLM在预测MOS和SIM方面具有竞争力的性能,同时在A B测试和自然语言描述中也提供了有希望的结果。我们的数据处理脚本和微调模型检查点将在验收后发布。摘要:Speech quality assessment typically requires evaluating audio from multiple aspects, such as mean opinion score (MOS) and speaker similarity (SIM) etc., which can be challenging to cover using one small model designed for a single task. In this paper, we propose leveraging recently introduced auditory large language models (LLMs) for automatic speech quality assessment. By employing task-specific prompts, auditory LLMs are finetuned to predict MOS, SIM and A B testing results, which are commonly used for evaluating text-to-speech systems. Additionally, the finetuned auditory LLM is able to generate natural language descriptions assessing aspects like noisiness, distortion, discontinuity, and overall quality, providing more interpretable outputs. Extensive experiments have been performed on the NISQA, BVCC, SOMOS and VoxSim speech quality datasets, using open-source auditory LLMs such as SALMONN, Qwen-Audio, and Qwen2-Audio. For the natural language descriptions task, a commercial model Google Gemini 1.5 Pro is also evaluated. The results demonstrate that auditory LLMs achieve competitive performance compared to state-of-the-art task-specific small models in predicting MOS and SIM, while also delivering promising results in A B testing and natural language descriptions. Our data processing scripts and finetuned model checkpoints will be released upon acceptance.
【14】 Towards Within-Class Variation in Alzheimer's Disease Detection from Spontaneous Speech
标题: 自发言语检测阿尔茨海默病的班内变异
作者:Jiawen Kang,Dongrui Han,Lingwei Meng,Jingyan Zhou,Jinchao Li,Xixin Wu,Helen Meng
链接:点击下载PDF文件
摘要:阿尔茨海默病(AD)检测已经成为一个很有前途的研究领域,其采用机器学习分类模型来区分患有AD的个体和没有AD的个体。与传统的分类任务不同,我们确定类内变异是AD检测中的一个关键挑战:AD患者表现出一系列认知障碍。鉴于许多AD检测任务缺乏细粒度的标签,简单的二进制分类可能会忽略两个关键方面:类内差异和实例级不平衡。前者迫使模型将具有不同程度损伤的AD样本映射到单个诊断标签,而忽略认知功能的某些变化。而后者使模型偏向于过度代表的严重程度。这项工作提出了应对这些挑战的早期努力。我们提出了两种新的方法:软目标蒸馏(SoTD)和实例级重新平衡(InRe),分别针对两个问题。在ADReSS和ADReSSo数据集上的实验表明,该方法显著提高了检测精度.进一步的分析表明,SoTD有效地利用了多个组件模型的优势,而InRe大大消除了模型过拟合。这些发现为开发更强大和可靠的AD检测模型提供了见解。摘要:Alzheimer's Disease (AD) detection has emerged as a promising research area that employs machine learning classification models to distinguish between individuals with AD and those without. Unlike conventional classification tasks, we identify within-class variation as a critical challenge in AD detection: individuals with AD exhibit a spectrum of cognitive impairments. Given that many AD detection tasks lack fine-grained labels, simplistic binary classification may overlook two crucial aspects: within-class differences and instance-level imbalance. The former compels the model to map AD samples with varying degrees of impairment to a single diagnostic label, disregarding certain changes in cognitive function. While the latter biases the model towards overrepresented severity levels. This work presents early efforts to address these challenges. We propose two novel methods: Soft Target Distillation (SoTD) and Instance-level Re-balancing (InRe), targeting two problems respectively. Experiments on the ADReSS and ADReSSo datasets demonstrate that the proposed methods significantly improve detection accuracy. Further analysis reveals that SoTD effectively harnesses the strengths of multiple component models, while InRe substantially alleviates model over-fitting. These findings provide insights for developing more robust and reliable AD detection models.
【15】 A Literature Review of Keyword Spotting Technologies for Urdu
标题: 乌尔都语关键词识别技术文献综述
作者:Syed Muhammad Aqdas Rizvi
链接:点击下载PDF文件
摘要:这篇文献综述调查了关键字定位(KWS)技术的进步,特别是专注于乌尔都语,巴基斯坦的低资源语言(LRL),其中有复杂的语音。尽管语音技术在全球范围内取得了长足的进步,但乌尔都语仍面临着独特的挑战,需要更定制的解决方案。该综述追溯了从基础高斯混合模型到深度神经网络和Transformers等复杂神经架构的演变,突出了重要的里程碑,例如集成多任务学习和利用未标记数据的自我监督方法。它研究了新兴技术在多语言和资源受限环境中增强KWS系统性能的作用,强调了满足乌尔都语等语言的创新需求。因此,本次审查强调需要针对具体情况的研究,解决乌尔都语和类似的URL的固有复杂性,并通过这些语言的语音技术更具包容性的方法进行沟通的地区的手段。摘要:This literature review surveys the advancements of keyword spotting (KWS) technologies, specifically focusing on Urdu, Pakistan's low-resource language (LRL), which has complex phonetics. Despite the global strides in speech technology, Urdu presents unique challenges requiring more tailored solutions. The review traces the evolution from foundational Gaussian Mixture Models to sophisticated neural architectures like deep neural networks and transformers, highlighting significant milestones such as integrating multi-task learning and self-supervised approaches that leverage unlabeled data. It examines emerging technologies' role in enhancing KWS systems' performance within multilingual and resource-constrained settings, emphasizing the need for innovations that cater to languages like Urdu. Thus, this review underscores the need for context-specific research addressing the inherent complexities of Urdu and similar URLs and the means of regions communicating through such languages for a more inclusive approach to speech technology.
【16】 How Redundant Is the Transformer Stack in Speech Representation Models?
标题: 语音表示模型中的Transformer Stack有多冗余?
作者:Teresa Dorszewski,Albert Kjøller Jacobsen,Lenka Tětková,Lars Kai Hansen
链接:点击下载PDF文件
摘要:自监督语音表示模型,特别是那些利用Transformer架构的模型,已经在语音识别、说话人识别和情感检测等各种任务中表现出了卓越的性能。最近对Transformer模型的研究揭示了层之间的高冗余和显著修剪的潜力,我们将在这里研究基于transformer的语音表示模型。我们使用三个相似性度量:余弦相似性,中心内核对齐,和相互最近邻对齐的语音表示模型中的层相似性进行了详细的分析。我们的研究结果揭示了一个高相似性的块状结构,表明两个主要的处理步骤和显着的冗余层。我们证明了修剪基于transformer的语音表示模型的有效性,而不需要后训练,实现了高达40%的减少Transformer层,同时保持超过95%的模型的预测能力。此外,我们采用了知识蒸馏方法来取代整个Transformer堆栈与模仿层,减少网络大小95-98%和推理时间高达94%。计算负载的大幅降低并没有造成相当大的性能损失,这表明Transformer堆栈对于语音表示模型的下游应用程序几乎完全冗余。摘要:Self-supervised speech representation models, particularly those leveraging transformer architectures, have demonstrated remarkable performance across various tasks such as speech recognition, speaker identification, and emotion detection. Recent studies on transformer models revealed a high redundancy between layers and the potential for significant pruning, which we will investigate here for transformer-based speech representation models. We perform a detailed analysis of layer similarity in speech representation models using three similarity metrics: cosine similarity, centered kernel alignment, and mutual nearest-neighbor alignment. Our findings reveal a block-like structure of high similarity, suggesting two main processing steps and significant redundancy of layers. We demonstrate the effectiveness of pruning transformer-based speech representation models without the need for post-training, achieving up to 40% reduction in transformer layers while maintaining over 95% of the model's predictive capacity. Furthermore, we employ a knowledge distillation method to substitute the entire transformer stack with mimicking layers, reducing the network size 95-98% and the inference time by up to 94%. This substantial decrease in computational load occurs without considerable performance loss, suggesting that the transformer stack is almost completely redundant for downstream applications of speech representation models.
【17】 Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
标题: 在计算预算下有效训练自我监督语音基础模型
作者:Andy T. Liu,Yi-Cheng Lin,Haibin Wu,Stefan Winkler,Hung-yi Lee
备注:To appear in SLT 2024
链接:点击下载PDF文件
摘要:尽管它们取得了令人印象深刻的成功,但训练基础模型的计算成本仍然很高。本文研究如何在有限的计算资源下,利用自监督学习(SSL)有效地训练语音基础模型。我们研究了SSL中影响预算的关键因素,包括模型架构、模型大小和数据大小。我们的目标是为理解语音基础模型的训练动态做出分析步骤。我们基准SSL的目标在一个完全可比的设置,并发现其他因素更显着的SSL的成功。我们的研究结果表明,在相同的计算和参数预算下,更苗条的模型架构优于常见的小型架构。我们证明了预训练数据的大小仍然至关重要,即使在SSL训练期间进行数据增强,因为在有限的数据上迭代时性能会受到影响。最后,我们确定了模型大小和数据大小之间的权衡,突出了给定计算预算的最佳模型大小。摘要:Despite their impressive success, training foundation models remains computationally costly. This paper investigates how to efficiently train speech foundation models with self-supervised learning (SSL) under a limited compute budget. We examine critical factors in SSL that impact the budget, including model architecture, model size, and data size. Our goal is to make analytical steps toward understanding the training dynamics of speech foundation models. We benchmark SSL objectives in an entirely comparable setting and find that other factors contribute more significantly to the success of SSL. Our results show that slimmer model architectures outperform common small architectures under the same compute and parameter budget. We demonstrate that the size of the pre-training data remains crucial, even with data augmentation during SSL training, as performance suffers when iterating over limited data. Finally, we identify a trade-off between model size and data size, highlighting an optimal model size for a given compute budget.
【18】 Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
标题: 无监督单词发现:使用集群进行边界检测与动态编程
作者:Simon Malan,Benjamin van Niekerk,Herman Kamper
备注:3 figures, 3 tables
链接:点击下载PDF文件
摘要:我们研究了将未标记的语音分割成类似单词的片段并将其聚集到词典中的长期问题。之前的几种方法使用评分模型与动态规划相结合来找到最佳分割。在这里,我们提出了一个更简单的策略:我们预测词边界使用相邻的自监督功能之间的相异性,然后我们聚类预测段构建一个词典。为了进行公平的比较,我们更新了旧的ES-KMeans动态规划方法,具有更好的功能和边界约束。在五种语言的ZeroSpeech基准测试中,与新的ES-KMeans+方法相比,我们的简单方法给出了类似的最先进的结果,同时速度快了近五倍。摘要:We look at the long-standing problem of segmenting unlabeled speech into word-like segments and clustering these into a lexicon. Several previous methods use a scoring model coupled with dynamic programming to find an optimal segmentation. Here we propose a much simpler strategy: we predict word boundaries using the dissimilarity between adjacent self-supervised features, then we cluster the predicted segments to construct a lexicon. For a fair comparison, we update the older ES-KMeans dynamic programming method with better features and boundary constraints. On the five-language ZeroSpeech benchmarks, our simple approach gives similar state-of-the-art results compared to the new ES-KMeans+ method, while being almost five times faster.
eess.AS音频处理
【1】 MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events标题: MT 2 KD:面向语音、扬声器和音频事件的通用编码器
作者:Xiaoyu Yang,Qiujia Li,Chao Zhang,Phil Woodland
链接:点击下载PDF文件
摘要:随着深度学习的进步,用于语音和音频处理的端到端(E2E)单任务模型的性能一直在不断提高。 然而,构建一个在多个任务上具有高性能的通用模型仍然具有挑战性,因为不同的语音和音频处理任务通常需要不同的训练数据、输入特征或模型架构来实现最佳性能。 在这项工作中,MT2KD,一种新的两阶段多任务学习框架,提出了建立一个通用的语音和音频编码器,共同执行三个基本任务:自动语音识别(ASR),音频标记(AT)和说话人验证(SV)。在第一阶段中,多教师知识蒸馏(KD)被应用于使用相同的未标记数据将三个单任务高性能教师编码器的特征空间对齐到单个学生编码器中。在第二阶段,通过初始化第一阶段的模型并在每个单个任务的单独标记数据上进行训练来进行多任务监督微调。 实验表明,所提出的多任务训练管道显着优于从零开始多任务学习训练的基线模型。最终的系统在ASR、AT和SV上实现了良好的性能:与性能最好的单任务编码器相比,ASR上的相对字错误率增加不到4%,AT上的平均平均精度仅低1.9,SV上的绝对等错误率高0.23%,仅使用66M的总模型参数。摘要:With the advances in deep learning, the performance of end-to-end (E2E) single-task models for speech and audio processing has been constantly improving. However, it is still challenging to build a general-purpose model with high performance on multiple tasks, since different speech and audio processing tasks usually require different training data, input features, or model architectures to achieve optimal performance. In this work, MT2KD, a novel two-stage multi-task learning framework is proposed to build a general-purpose speech and audio encoder that jointly performs three fundamental tasks: automatic speech recognition (ASR), audio tagging (AT) and speaker verification (SV). In the first stage, multi-teacher knowledge distillation (KD) is applied to align the feature spaces of three single-task high-performance teacher encoders into a single student encoder using the same unlabelled data. In the second stage, multi-task supervised fine-tuning is carried out by initialising the model from the first stage and training on the separate labelled data of each single task. Experiments demonstrate that the proposed multi-task training pipeline significantly outperforms a baseline model trained with multi-task learning from scratch. The final system achieves good performance on ASR, AT and SV: with less than 4% relative word-error-rate increase on ASR, only 1.9 lower mean averaged precision on AT and 0.23% absolute higher equal error rate on SV compared to the best-performing single-task encoders, using only a 66M total model parameters.
【2】 Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling
标题: 具有多视图伪标签的语音半监督认知状态分类
作者:Yuanchao Li,Zixing Zhang,Jing Han,Peter Bell,Catherine Lai
链接:点击下载PDF文件
摘要:缺乏标记数据是语音分类任务中的常见挑战,特别是那些需要大量主观评估的任务,例如认知状态分类。在这项工作中,我们提出了一个半监督学习(SSL)框架,引入了一种新的多视图伪标记方法,该方法利用声学和语言特征来选择最有信心的数据来训练分类模型。在声学上,未标记的数据与使用Frechet音频距离的标记数据进行比较,该距离是从多个音频编码器生成的嵌入计算的。在语言学上,大型语言模型会根据我们提出的特定任务知识来修改自动语音识别transmittance和预测标签。当来自两个源的伪标签对齐时,识别高置信度数据,而不匹配则被视为低置信度数据。然后训练双峰分类器以迭代地标记低置信度数据,直到满足预定义的标准。我们评估我们的SSL框架的情绪识别和痴呆症检测任务。实验结果表明,与仅使用30%标记数据的完全监督学习相比,我们的方法具有竞争力的性能,并且显著优于两个选定的基线。摘要:The lack of labeled data is a common challenge in speech classification tasks, particularly those requiring extensive subjective assessment, such as cognitive state classification. In this work, we propose a Semi-Supervised Learning (SSL) framework, introducing a novel multi-view pseudo-labeling method that leverages both acoustic and linguistic characteristics to select the most confident data for training the classification model. Acoustically, unlabeled data are compared to labeled data using the Frechet audio distance, calculated from embeddings generated by multiple audio encoders. Linguistically, large language models are prompted to revise automatic speech recognition transcriptions and predict labels based on our proposed task-specific knowledge. High-confidence data are identified when pseudo-labels from both sources align, while mismatches are treated as low-confidence data. A bimodal classifier is then trained to iteratively label the low-confidence data until a predefined criterion is met. We evaluate our SSL framework on emotion recognition and dementia detection tasks. Experimental results demonstrate that our method achieves competitive performance compared to fully supervised learning using only 30% of the labeled data and significantly outperforms two selected baselines.
【3】 Cross-lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
标题: 跨语言语音情感识别:人类与自我监督模型
作者:Zhichen Han,Tianqi Geng,Hui Feng,Jiahong Yuan,Korin Richmond,Yuanchao Li
链接:点击下载PDF文件
摘要:利用自监督学习(SSL)模型进行语音情感识别(SER)已被证明是有效的,但有限的研究探索了跨语言的情况。本研究提出了人类表现和SSL模型之间的比较分析,从逐层分析和探索单语,跨语言和迁移学习环境中的参数有效的微调策略开始。我们进一步比较SER能力的模型和人类在话语和段水平。此外,我们还通过人的评价研究了方言对跨语言SER的影响。我们的研究结果表明,模型,适当的知识转移,可以适应目标语言,并实现与本族语者相当的性能。我们还证明了显着的影响,方言对SER的个人没有事先的语言和语言学背景。此外,人类和模型在不同的情绪下表现出不同的行为。这些结果为SSL模型的跨语言SER能力提供了新的见解,强调了它们与人类情感感知的相似性和差异性。摘要:Utilizing Self-Supervised Learning (SSL) models for Speech Emotion Recognition (SER) has proven effective, yet limited research has explored cross-lingual scenarios. This study presents a comparative analysis between human performance and SSL models, beginning with a layer-wise analysis and an exploration of parameter-efficient fine-tuning strategies in monolingual, cross-lingual, and transfer learning contexts. We further compare the SER ability of models and humans at both utterance- and segment-levels. Additionally, we investigate the impact of dialect on cross-lingual SER through human evaluation. Our findings reveal that models, with appropriate knowledge transfer, can adapt to the target language and achieve performance comparable to native speakers. We also demonstrate the significant effect of dialect on SER for individuals without prior linguistic and paralinguistic background. Moreover, both humans and models exhibit distinct behaviors across different emotions. These results offer new insights into the cross-lingual SER capabilities of SSL models, underscoring both their similarities to and differences from human emotion perception.
【4】 Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings
标题: 多通道多方会议的模块化扬声器拨号中的融合空间线索
作者:Ruoyu Wang,Shutong Niu,Gaobin Yang,Jun Du,Shuangqing Qian,Tian Gao,Jia Pan
备注:5 pages, Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:尽管近年来完全端到端的扬声器日志化系统已经取得了重大进展,但模块化系统由于其更大的适应性和鲁棒性,通常在现实世界的场景中取得更好的效果。从历史上看,模块化说话人日记方法很少讨论如何利用多通道语音的空间线索。本文提出了一种三阶段模块化系统,通过利用来自多通道语音的空间线索为神经说话人日记化(NSD)解码的每个阶段提供更准确的初始化,来增强单通道神经说话人日记化系统和识别性能:(1)对多通道语音进行重叠检测和连续语音分离(CSS),得到更清晰的单说话人语音段进行聚类,接着是第一NSD解码通道。(2)第一遍的结果初始化一个复杂的角中心高斯混合模型(cACGMM),以估计多通道语音上的扬声器掩码,并通过重叠相加和掩码到VAD,实现具有较低扬声器误差(SpkErr)的初始化,然后是第二NSD解码遍。(3)第二解码结果用于引导源分离(GSS),识别和过滤包含少于一个词的短段以获得更干净的语音段,随后进行重新聚类和最终的NSD解码通道。我们介绍了CHiME-8 NOTSOFAR-1(远场录音环境中的自然办公室谈话者)挑战赛的逐步探索评估结果,展示了我们系统的有效性及其对提高识别性能的贡献。我们最终的系统在挑战赛中获得了第一名。摘要:Although fully end-to-end speaker diarization systems have made significant progress in recent years, modular systems often achieve superior results in real-world scenarios due to their greater adaptability and robustness. Historically, modular speaker diarization methods have seldom discussed how to leverage spatial cues from multi-channel speech. This paper proposes a three-stage modular system to enhance single-channel neural speaker diarization systems and recognition performance by utilizing spatial cues from multi-channel speech to provide more accurate initialization for each stage of neural speaker diarization (NSD) decoding: (1) Overlap detection and continuous speech separation (CSS) on multi-channel speech are used to obtain cleaner single speaker speech segments for clustering, followed by the first NSD decoding pass. (2) The results from the first pass initialize a complex Angular Central Gaussian Mixture Model (cACGMM) to estimate speaker-wise masks on multi-channel speech, and through Overlap-add and Mask-to-VAD, achieve initialization with lower speaker error (SpkErr), followed by the second NSD decoding pass. (3) The second decoding results are used for guided source separation (GSS), recognizing and filtering short segments containing less one word to obtain cleaner speech segments, followed by re-clustering and the final NSD decoding pass. We presented the progressively explored evaluation results from the CHiME-8 NOTSOFAR-1 (Natural Office Talkers in Settings Of Far-field Audio Recordings) challenge, demonstrating the effectiveness of our system and its contribution to improving recognition performance. Our final system achieved the first place in the challenge.
【5】 Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions
标题: 基于语言模型的文本转语音中的情感维度控制:跨越人类情感的广泛范围
作者:Kun Zhou,You Zhang,Shengkui Zhao,Hao Wang,Zexu Pan,Dianwen Ng,Chong Zhang,Chongjia Ni,Yukun Ma,Trung Hieu Nguyen,Jia Qi Yip,Bin Ma
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:由于情感的固有复杂性以及情感语音数据集和模型的局限性,当前的情感文本到语音(TTS)系统在模仿广泛的人类情感方面面临挑战。本文提出了一种TTS框架,该框架有助于控制愉悦、唤醒和支配,并且可以在TTS训练期间合成多种情感风格,而无需任何情感语音数据。我们训练的情感属性预测器只使用分类标签的语音数据,与心理学研究,并结合锚定降维的自我监督学习(SSL)功能。TTS框架通过自回归语言模型将文本输入转换为语音标记,并使用伪情感维度来指导细粒度声学细节的并行预测。在LibriTTS数据集上进行的实验表明,我们的框架可以通过有效控制情感维度来合成具有增强的自然度和各种情感风格的语音,即使在TTS训练期间不包含任何情感语音。摘要:Current emotional text-to-speech (TTS) systems face challenges in mimicking a broad spectrum of human emotions due to the inherent complexity of emotions and limitations in emotional speech datasets and models. This paper proposes a TTS framework that facilitates control over pleasure, arousal, and dominance, and can synthesize a diversity of emotional styles without requiring any emotional speech data during TTS training. We train an emotional attribute predictor using only categorical labels from speech data, aligning with psychological research and incorporating anchored dimensionality reduction on self-supervised learning (SSL) features. The TTS framework converts text inputs into phonetic tokens via an autoregressive language model and uses pseudo-emotional dimensions to guide the parallel prediction of fine-grained acoustic details. Experiments conducted on the LibriTTS dataset demonstrate that our framework can synthesize speech with enhanced naturalness and a variety of emotional styles by effectively controlling emotional dimensions, even without the inclusion of any emotional speech during TTS training.
【6】 Speech Recognition Rescoring with Large Speech-Text Foundation Models
标题: 使用大语音文本基础模型的语音识别重新评分
作者:Prashanth Gurunath Shivakumar,Jari Kolehmainen,Aditya Gourav,Yi Gu,Ankur Gandhe,Ariya Rastrow,Ivan Bulyko
链接:点击下载PDF文件
摘要:大型语言模型(LLM)已经证明了通过利用大量文本数据来理解人类语言的能力。自动语音识别(ASR)系统通常受到可用的转录语音数据的限制,并受益于使用LLM的第二遍重新评分。近年来,多模态大型语言模型,特别是语音和文本基础模型,已经表现出强大的口语理解能力。语音-文本基础模型利用语音和文本模态中的大量未标记和标记数据来建模人类语言。在这项工作中,我们提出了新的技术,使用多模态LLM的ASR重新评分。我们还探索了区分训练,以进一步提高基础模型的再评分性能。我们证明了语音-文本LLM中的跨模态知识转移可以有益于重新评分。我们的实验证明了高达20%的相对改善耳语大ASR和高达15%的相对改善,在纯文本LLM。摘要:Large language models (LLM) have demonstrated the ability to understand human language by leveraging large amount of text data. Automatic speech recognition (ASR) systems are often limited by available transcribed speech data and benefit from a second pass rescoring using LLM. Recently multi-modal large language models, particularly speech and text foundational models have demonstrated strong spoken language understanding. Speech-Text foundational models leverage large amounts of unlabelled and labelled data both in speech and text modalities to model human language. In this work, we propose novel techniques to use multi-modal LLM for ASR rescoring. We also explore discriminative training to further improve the foundational model rescoring performance. We demonstrate cross-modal knowledge transfer in speech-text LLM can benefit rescoring. Our experiments demonstrate up-to 20% relative improvements over Whisper large ASR and up-to 15% relative improvements over text-only LLM.
【7】 Enabling Auditory Large Language Models for Automatic Speech Quality Evaluation
标题: 启用听觉大语言模型进行自动语音质量评估
作者:Siyin Wang,Wenyi Yu,Yudong Yang,Changli Tang,Yixuan Li,Jimin Zhuang,Xianzhao Chen,Xiaohai Tian,Jun Zhang,Guangzhi Sun,Lu Lu,Chao Zhang
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:语音质量评估通常需要从多个方面评估音频,例如平均意见得分(MOS)和说话人相似度(SIM)等,这对于使用为单个任务设计的一个小模型来覆盖可能是具有挑战性的。在本文中,我们建议利用最近推出的听觉大语言模型(LLM)的自动语音质量评估。通过使用特定于任务的提示,对听觉LLM进行微调以预测MOS、SIM和A B测试结果,这些测试结果通常用于评估文本到语音系统。此外,微调的听觉LLM能够生成自然语言描述,评估噪音、失真、不连续性和整体质量等方面,提供更多可解释的输出。在NISQA、BVCC、SOMOS和VoxSim语音质量数据集上进行了大量的实验,使用了开源的听觉LLM,如SALMONN、Qwen-Audio和Qwen 2-Audio。对于自然语言描述任务,还评估了商业模型Google Gemini 1.5 Pro。结果表明,与最先进的特定于任务的小模型相比,听觉LLM在预测MOS和SIM方面具有竞争力的性能,同时在A B测试和自然语言描述中也提供了有希望的结果。我们的数据处理脚本和微调模型检查点将在验收后发布。摘要:Speech quality assessment typically requires evaluating audio from multiple aspects, such as mean opinion score (MOS) and speaker similarity (SIM) etc., which can be challenging to cover using one small model designed for a single task. In this paper, we propose leveraging recently introduced auditory large language models (LLMs) for automatic speech quality assessment. By employing task-specific prompts, auditory LLMs are finetuned to predict MOS, SIM and A B testing results, which are commonly used for evaluating text-to-speech systems. Additionally, the finetuned auditory LLM is able to generate natural language descriptions assessing aspects like noisiness, distortion, discontinuity, and overall quality, providing more interpretable outputs. Extensive experiments have been performed on the NISQA, BVCC, SOMOS and VoxSim speech quality datasets, using open-source auditory LLMs such as SALMONN, Qwen-Audio, and Qwen2-Audio. For the natural language descriptions task, a commercial model Google Gemini 1.5 Pro is also evaluated. The results demonstrate that auditory LLMs achieve competitive performance compared to state-of-the-art task-specific small models in predicting MOS and SIM, while also delivering promising results in A B testing and natural language descriptions. Our data processing scripts and finetuned model checkpoints will be released upon acceptance.
【8】 Towards Within-Class Variation in Alzheimer's Disease Detection from Spontaneous Speech
标题: 自发言语检测阿尔茨海默病的班内变异
作者:Jiawen Kang,Dongrui Han,Lingwei Meng,Jingyan Zhou,Jinchao Li,Xixin Wu,Helen Meng
链接:点击下载PDF文件
摘要:阿尔茨海默病(AD)检测已经成为一个很有前途的研究领域,其采用机器学习分类模型来区分患有AD的个体和没有AD的个体。与传统的分类任务不同,我们确定类内变异是AD检测中的一个关键挑战:AD患者表现出一系列认知障碍。鉴于许多AD检测任务缺乏细粒度的标签,简单的二进制分类可能会忽略两个关键方面:类内差异和实例级不平衡。前者迫使模型将具有不同程度损伤的AD样本映射到单个诊断标签,而忽略认知功能的某些变化。而后者使模型偏向于过度代表的严重程度。这项工作提出了应对这些挑战的早期努力。我们提出了两种新的方法:软目标蒸馏(SoTD)和实例级重新平衡(InRe),分别针对两个问题。在ADReSS和ADReSSo数据集上的实验表明,该方法显著提高了检测精度.进一步的分析表明,SoTD有效地利用了多个组件模型的优势,而InRe大大消除了模型过拟合。这些发现为开发更强大和可靠的AD检测模型提供了见解。摘要:Alzheimer's Disease (AD) detection has emerged as a promising research area that employs machine learning classification models to distinguish between individuals with AD and those without. Unlike conventional classification tasks, we identify within-class variation as a critical challenge in AD detection: individuals with AD exhibit a spectrum of cognitive impairments. Given that many AD detection tasks lack fine-grained labels, simplistic binary classification may overlook two crucial aspects: within-class differences and instance-level imbalance. The former compels the model to map AD samples with varying degrees of impairment to a single diagnostic label, disregarding certain changes in cognitive function. While the latter biases the model towards overrepresented severity levels. This work presents early efforts to address these challenges. We propose two novel methods: Soft Target Distillation (SoTD) and Instance-level Re-balancing (InRe), targeting two problems respectively. Experiments on the ADReSS and ADReSSo datasets demonstrate that the proposed methods significantly improve detection accuracy. Further analysis reveals that SoTD effectively harnesses the strengths of multiple component models, while InRe substantially alleviates model over-fitting. These findings provide insights for developing more robust and reliable AD detection models.
【9】 A Literature Review of Keyword Spotting Technologies for Urdu
标题: 乌尔都语关键词识别技术文献综述
作者:Syed Muhammad Aqdas Rizvi
链接:点击下载PDF文件
摘要:这篇文献综述调查了关键字识别(KWS)技术的进步,特别关注巴基斯坦的低资源语言(LRL)乌尔都语,它具有复杂的语音。尽管语音技术在全球范围内取得了长足的进步,但乌尔都语仍面临着独特的挑战,需要更定制的解决方案。该综述追溯了从基础高斯混合模型到深度神经网络和Transformers等复杂神经架构的演变,突出了重要的里程碑,例如集成多任务学习和利用未标记数据的自我监督方法。它研究了新兴技术在多语言和资源受限环境中增强KWS系统性能的作用,强调了满足乌尔都语等语言的创新需求。因此,本次审查强调需要针对具体情况的研究,解决乌尔都语和类似的URL的固有复杂性,并通过这些语言的语音技术更具包容性的方法进行沟通的地区的手段。摘要:This literature review surveys the advancements of keyword spotting (KWS) technologies, specifically focusing on Urdu, Pakistan's low-resource language (LRL), which has complex phonetics. Despite the global strides in speech technology, Urdu presents unique challenges requiring more tailored solutions. The review traces the evolution from foundational Gaussian Mixture Models to sophisticated neural architectures like deep neural networks and transformers, highlighting significant milestones such as integrating multi-task learning and self-supervised approaches that leverage unlabeled data. It examines emerging technologies' role in enhancing KWS systems' performance within multilingual and resource-constrained settings, emphasizing the need for innovations that cater to languages like Urdu. Thus, this review underscores the need for context-specific research addressing the inherent complexities of Urdu and similar URLs and the means of regions communicating through such languages for a more inclusive approach to speech technology.
【10】 How Redundant Is the Transformer Stack in Speech Representation Models?
标题: 语音表示模型中的Transformer Stack有多冗余?
作者:Teresa Dorszewski,Albert Kjøller Jacobsen,Lenka Tětková,Lars Kai Hansen
链接:点击下载PDF文件
摘要:自监督语音表示模型,特别是那些利用Transformer架构的模型,已经在语音识别、说话人识别和情感检测等各种任务中表现出了卓越的性能。最近对Transformer模型的研究揭示了层之间的高冗余和显著修剪的潜力,我们将在这里研究基于transformer的语音表示模型。我们使用三个相似性度量:余弦相似性,中心内核对齐,和相互最近邻对齐的语音表示模型中的层相似性进行了详细的分析。我们的研究结果揭示了一个高相似性的块状结构,表明两个主要的处理步骤和显着的冗余层。我们证明了修剪基于transformer的语音表示模型的有效性,而不需要后训练,实现了高达40%的减少Transformer层,同时保持超过95%的模型的预测能力。此外,我们采用了知识蒸馏方法来取代整个Transformer堆栈与模仿层,减少网络大小95-98%和推理时间高达94%。这种计算负载的显著降低在没有相当大的性能损失的情况下发生,表明Transformer堆栈对于语音表示模型的下游应用几乎完全是冗余的。摘要:Self-supervised speech representation models, particularly those leveraging transformer architectures, have demonstrated remarkable performance across various tasks such as speech recognition, speaker identification, and emotion detection. Recent studies on transformer models revealed a high redundancy between layers and the potential for significant pruning, which we will investigate here for transformer-based speech representation models. We perform a detailed analysis of layer similarity in speech representation models using three similarity metrics: cosine similarity, centered kernel alignment, and mutual nearest-neighbor alignment. Our findings reveal a block-like structure of high similarity, suggesting two main processing steps and significant redundancy of layers. We demonstrate the effectiveness of pruning transformer-based speech representation models without the need for post-training, achieving up to 40% reduction in transformer layers while maintaining over 95% of the model's predictive capacity. Furthermore, we employ a knowledge distillation method to substitute the entire transformer stack with mimicking layers, reducing the network size 95-98% and the inference time by up to 94%. This substantial decrease in computational load occurs without considerable performance loss, suggesting that the transformer stack is almost completely redundant for downstream applications of speech representation models.
【11】 Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
标题: 在计算预算下有效训练自我监督语音基础模型
作者:Andy T. Liu,Yi-Cheng Lin,Haibin Wu,Stefan Winkler,Hung-yi Lee
备注:To appear in SLT 2024
链接:点击下载PDF文件
摘要:尽管它们取得了令人印象深刻的成功,但训练基础模型的计算成本仍然很高。本文研究如何在有限的计算资源下,利用自监督学习(SSL)有效地训练语音基础模型。我们研究了SSL中影响预算的关键因素,包括模型架构、模型大小和数据大小。我们的目标是为理解语音基础模型的训练动态做出分析步骤。我们基准SSL的目标在一个完全可比的设置,并发现其他因素更显着的SSL的成功。我们的研究结果表明,在相同的计算和参数预算下,更苗条的模型架构优于常见的小型架构。我们证明,即使在SSL训练期间进行数据增强,预训练数据的大小仍然至关重要,因为在有限的数据上迭代时性能会受到影响。最后,我们确定了模型大小和数据大小之间的权衡,突出了给定计算预算的最佳模型大小。摘要:Despite their impressive success, training foundation models remains computationally costly. This paper investigates how to efficiently train speech foundation models with self-supervised learning (SSL) under a limited compute budget. We examine critical factors in SSL that impact the budget, including model architecture, model size, and data size. Our goal is to make analytical steps toward understanding the training dynamics of speech foundation models. We benchmark SSL objectives in an entirely comparable setting and find that other factors contribute more significantly to the success of SSL. Our results show that slimmer model architectures outperform common small architectures under the same compute and parameter budget. We demonstrate that the size of the pre-training data remains crucial, even with data augmentation during SSL training, as performance suffers when iterating over limited data. Finally, we identify a trade-off between model size and data size, highlighting an optimal model size for a given compute budget.
【12】 The Effect of Perceptual Metrics on Music Representation Learning for Genre Classification
标题: 知觉预设对流派分类音乐表示学习的影响
作者:Tashi Namgyal,Alexander Hepburn,Raul Santos-Rodriguez,Valero Laparra,Jesus Malo
备注:arXiv admin note: text overlap with arXiv:2312.03455
链接:点击下载PDF文件
摘要:自然信号的主观质量可以用客观感知度量来近似。感知度量旨在近似人类观察者的感知行为,通常反映自然信号和神经通路中的结构。用感知度量作为损失函数训练的模型可以从这些度量中的结构中捕获感知上有意义的特征。我们证明,使用从感知损失训练的自动编码器中提取的特征可以提高音乐理解任务(即流派分类)的性能,而不是在学习分类器时直接使用这些指标作为距离。这一结果表明,改进的泛化到新的信号时,使用知觉指标作为损失函数的表示学习。摘要:The subjective quality of natural signals can be approximated with objective perceptual metrics. Designed to approximate the perceptual behaviour of human observers, perceptual metrics often reflect structures found in natural signals and neurological pathways. Models trained with perceptual metrics as loss functions can capture perceptually meaningful features from the structures held within these metrics. We demonstrate that using features extracted from autoencoders trained with perceptual losses can improve performance on music understanding tasks, i.e. genre classification, over using these metrics directly as distances when learning a classifier. This result suggests improved generalisation to novel signals when using perceptual metrics as loss functions for representation learning.
【13】 Using LLM for Real-Time Transcription and Summarization of Doctor-Patient Interactions into ePuskesmas in Indonesia
标题: 使用LLM实时转录和总结印度尼西亚ePuskesmas中的医患互动
作者:Azmul Asmar Irfan,Nur Ahmad Khatim,Mansur M. Arief
链接:点击下载PDF文件
摘要:导致Puskesmas效率低下的关键问题之一是医患互动的耗时性。医生需要进行全面的咨询,其中包括诊断病人的病情,提供治疗建议,并将详细的笔记转录成医疗记录。在语言背景不同的地区,医生往往不得不提出澄清问题,进一步延长了这个过程。虽然诊断是必不可少的,但转录和总结通常可以使用人工智能自动化,以提高时间效率,帮助医生提高护理质量,并实现早期诊断和干预。本文提出了一种使用本地化的大语言模型(LLM)来转录,翻译和总结医患对话的解决方案。我们利用Whisper模型进行转录和GPT-3,将它们汇总为ePuskemas医疗记录格式。该系统作为现有网络浏览器扩展的附加组件实现,允许医生在交谈时填写患者表格。通过利用该解决方案进行实时转录、翻译和总结,医生可以缩短患者护理的周转时间,同时提高记录的质量,使其在未来的访问中变得更加详细和有见地。这项创新解决了印度尼西亚设施过度拥挤和医疗保健提供者行政负担等挑战。我们相信,这一解决方案将帮助医生节省时间,提供更好的护理,并生成更准确的医疗记录,这是实现医疗保健现代化的重要一步,也是确保患者即使在资源有限的情况下也能获得及时、高质量护理的重要一步。摘要:One of the key issues contributing to inefficiency in Puskesmas is the time-consuming nature of doctor-patient interactions. Doctors need to conduct thorough consultations, which include diagnosing the patient's condition, providing treatment advice, and transcribing detailed notes into medical records. In regions with diverse linguistic backgrounds, doctors often have to ask clarifying questions, further prolonging the process. While diagnosing is essential, transcription and summarization can often be automated using AI to improve time efficiency and help doctors enhance care quality and enable early diagnosis and intervention. This paper proposes a solution using a localized large language model (LLM) to transcribe, translate, and summarize doctor-patient conversations. We utilize the Whisper model for transcription and GPT-3 to summarize them into the ePuskemas medical records format. This system is implemented as an add-on to an existing web browser extension, allowing doctors to fill out patient forms while talking. By leveraging this solution for real-time transcription, translation, and summarization, doctors can improve the turnaround time for patient care while enhancing the quality of records, which become more detailed and insightful for future visits. This innovation addresses challenges like overcrowded facilities and the administrative burden on healthcare providers in Indonesia. We believe this solution will help doctors save time, provide better care, and produce more accurate medical records, representing a significant step toward modernizing healthcare and ensuring patients receive timely, high-quality care, even in resource-constrained settings.
【14】 Weighted Cross-entropy for Low-Resource Languages in Multilingual Speech Recognition
标题: 多语言语音识别中低资源语言的加权交叉信息
作者:Andrés Piñeiro-Martín,Carmen García-Mateo,Laura Docío-Fernández,María del Carmen López-Pérez,Georg Rehm
Journal-ref:Proceedings of Interspeech 2024
链接:点击下载PDF文件
摘要:本文讨论了将低资源语言集成到多语言自动语音识别(ASR)系统中的挑战。我们引入了加权交叉熵的一种新应用,通常用于不平衡的数据集,以促进在持续多语言学习的背景下将低资源语言集成到预训练的多语言ASR模型中。我们在五种高资源语言和一种低资源语言上对Whisper多语言ASR模型进行了微调,采用了语言加权动态交叉熵和数据增强。结果显示,与没有应用我们的方法的微调模型相比,低资源语言的单词错误率(WER)降低了6.69%,与原始Whisper模型相比,WER降低了48.86%。此外,我们的方法使六种语言的WER平均降低了3.29%,表明高资源语言的WER没有下降。摘要:This paper addresses the challenge of integrating low-resource languages into multilingual automatic speech recognition (ASR) systems. We introduce a novel application of weighted cross-entropy, typically used for unbalanced datasets, to facilitate the integration of low-resource languages into pre-trained multilingual ASR models within the context of continual multilingual learning. We fine-tune the Whisper multilingual ASR model on five high-resource languages and one low-resource language, employing language-weighted dynamic cross-entropy and data augmentation. The results show a remarkable 6.69% word error rate (WER) reduction for the low-resource language compared to the fine-tuned model without applying our approach, and a 48.86% WER reduction compared to the original Whisper model. In addition, our approach yields an average WER reduction of 3.29% across the six languages, showing no degradation for the high-resource languages.
【15】 Spelling Correction through Rewriting of Non-Autoregressive ASR Lattices
标题: 通过重写非自回归ASC格进行拼写纠正
作者:Leonid Velikovich,Christopher Li,Diamantino Caseiro,Shankar Kumar,Pat Rondon,Kandarp Joshi,Xavier Velez
备注:8 pages, 7 figures
链接:点击下载PDF文件
摘要:对于端到端自动语音识别(ASR)模型,识别个人或罕见短语可能很困难。提高准确性的一种有希望的方法是通过ASR网格的拼写校正(或重写),其中潜在的错误识别短语被替换为声学上相似和上下文相关的替代品。然而,由于非自回归、上下文无关的波束搜索产生的噪声假设,用联结主义时间分类(CTC)训练的ASR模型的重写是具有挑战性的。 我们提出了一种有限状态转换器(FST)技术重写基于转换器的CTC模型生成的字段格。我们的算法进行直接从词片到音素的字素到音素(G2P)转换,避免明确的单词表示和利用CTC晶格的丰富性。我们的方法不需要重新训练或修改ASR模型。我们实现了高达15.2%的相对减少句子错误率(SER)的测试集与上下文相关的实体。摘要:For end-to-end Automatic Speech Recognition (ASR) models, recognizing personal or rare phrases can be hard. A promising way to improve accuracy is through spelling correction (or rewriting) of the ASR lattice, where potentially misrecognized phrases are replaced with acoustically similar and contextually relevant alternatives. However, rewriting is challenging for ASR models trained with connectionist temporal classification (CTC) due to noisy hypotheses produced by a non-autoregressive, context-independent beam search. We present a finite-state transducer (FST) technique for rewriting wordpiece lattices generated by Transformer-based CTC models. Our algorithm performs grapheme-to-phoneme (G2P) conversion directly from wordpieces into phonemes, avoiding explicit word representations and exploiting the richness of the CTC lattice. Our approach requires no retraining or modification of the ASR model. We achieved up to a 15.2% relative reduction in sentence error rate (SER) on a test set with contextually relevant entities.
【16】 FastTalker: Jointly Generating Speech and Conversational Gestures from Text
标题: FastTalker:从文本联合生成语音和对话手势
作者:Zixin Guo,Jian Zhang
备注:European Conference on Computer Vision Workshop
链接:点击下载PDF文件
摘要:从文本脚本生成3D人类手势和语音对于创建逼真的谈话化身至关重要。一种解决方案是利用文本到语音(TTS)和语音到手势(STG)的单独管道,但是这种方法遭受语音和手势的不良对齐以及缓慢的推理时间。在本文中,我们介绍FastTalker,一个高效和有效的框架,同时产生高品质的语音音频和3D人类手势在高推理速度。我们的关键见解是重用的中间功能,从语音合成手势生成,因为这些功能包含更精确的节奏信息比功能重新提取生成的语音。具体来说,1)我们提出了一个端到端的框架,同时生成语音波形和全身手势,使用中间语音特征,如音高,起始,能量和持续时间直接用于手势解码; 2)我们重新设计了因果网络架构,以消除对未来输入的依赖性,以用于实际应用; 3)我们采用基于强化学习的神经结构搜索(NAS),通过优化我们的网络结构来提高性能和推理速度。在BEAT 2数据集上的实验结果表明,FastTalker在语音合成和手势生成方面都达到了最先进的性能,在NVIDIA 3090上每秒处理语音和手势的速度为0.17秒。摘要:Generating 3D human gestures and speech from a text script is critical for creating realistic talking avatars. One solution is to leverage separate pipelines for text-to-speech (TTS) and speech-to-gesture (STG), but this approach suffers from poor alignment of speech and gestures and slow inference times. In this paper, we introduce FastTalker, an efficient and effective framework that simultaneously generates high-quality speech audio and 3D human gestures at high inference speeds. Our key insight is reusing the intermediate features from speech synthesis for gesture generation, as these features contain more precise rhythmic information than features re-extracted from generated speech. Specifically, 1) we propose an end-to-end framework that concurrently generates speech waveforms and full-body gestures, using intermediate speech features such as pitch, onset, energy, and duration directly for gesture decoding; 2) we redesign the causal network architecture to eliminate dependencies on future inputs for real applications; 3) we employ Reinforcement Learning-based Neural Architecture Search (NAS) to enhance both performance and inference speed by optimizing our network architecture. Experimental results on the BEAT2 dataset demonstrate that FastTalker achieves state-of-the-art performance in both speech synthesis and gesture generation, processing speech and gestures in 0.17 seconds per second on an NVIDIA 3090.
【17】 Revisiting Acoustic Features for Robust ASR
标题: 重新审视稳健的ASB的声学特征
作者:Muhammad A. Shah,Bhiksha Raj
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:自动语音识别(ASR)系统必须对现实世界环境中存在的各种类型的噪声具有鲁棒性,包括环境噪声,房间脉冲响应,特殊效果以及恶意行为者的攻击(对抗性攻击)。最近的工作试图通过开发新的深度神经网络(DNN)并为其管理各种训练数据集,同时使用相对简单的声学特征来提高准确性和鲁棒性。虽然这种方法提高了对训练数据中存在的噪声类型的鲁棒性,但它对不可见噪声的鲁棒性有限,对对抗性攻击的鲁棒性可以忽略不计。在本文中,我们重新审视了早期作品的方法,这些作品开发了受生物听觉感知启发的声学特征,可用于执行准确和强大的ASR。相比之下,具体来说,我们评估ASR的准确性和鲁棒性的几个生物启发的声学特征。除了一些功能,从以前的作品,如伽玛滤波器组功能(GammSpec),我们还提出了两个新的声学功能,称为频率掩蔽频谱图(FreqMask)和伽玛频谱图的差异(DoGSpec)来模拟神经心理现象的频率掩蔽和横向抑制。在不同模型和数据集上的实验表明:(1)DoGSpec比非常流行的对数梅尔频谱图(LogMelSpec)具有更好的鲁棒性,并且精度下降最小;(2)GammSpec在语音鲁棒性基准测试中对非对抗性噪声具有更好的精度和鲁棒性,但在对抗性攻击方面优于DoGSpec。摘要:Automatic Speech Recognition (ASR) systems must be robust to the myriad types of noises present in real-world environments including environmental noise, room impulse response, special effects as well as attacks by malicious actors (adversarial attacks). Recent works seek to improve accuracy and robustness by developing novel Deep Neural Networks (DNNs) and curating diverse training datasets for them, while using relatively simple acoustic features. While this approach improves robustness to the types of noise present in the training data, it confers limited robustness against unseen noises and negligible robustness to adversarial attacks. In this paper, we revisit the approach of earlier works that developed acoustic features inspired by biological auditory perception that could be used to perform accurate and robust ASR. In contrast, Specifically, we evaluate the ASR accuracy and robustness of several biologically inspired acoustic features. In addition to several features from prior works, such as gammatone filterbank features (GammSpec), we also propose two new acoustic features called frequency masked spectrogram (FreqMask) and difference of gammatones spectrogram (DoGSpec) to simulate the neuro-psychological phenomena of frequency masking and lateral suppression. Experiments on diverse models and datasets show that (1) DoGSpec achieves significantly better robustness than the highly popular log mel spectrogram (LogMelSpec) with minimal accuracy degradation, and (2) GammSpec achieves better accuracy and robustness to non-adversarial noises from the Speech Robust Bench benchmark, but it is outperformed by DoGSpec against adversarial attacks.
机器翻译,仅供参考
![]()
