今日论文合集:cs.SD语音13篇,eess.AS音频处理17篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 AudioLens: A Closer Look at Auditory Attribute Perception of Large  Audio-Language Models
标题: AudioLens:仔细观察大型音频语言模型的听觉属性感知
链接:https://arxiv.org/abs/2506.05140
作者: Chih-Kai Yang,  Neo Ho,  Yi-Jyun Lee,  Hung-yi Lee 
备注:8 pages, 5 figures, 3 tables
摘要:理解大型音频语言模型(LALM)的内部机制对于解释其行为和提高性能至关重要。这项工作提出了第一次深入分析LALM内部如何感知和识别听觉属性。通过在三个最先进的LALM上应用词汇投影,我们跟踪了属性信息如何在层和标记位置之间演变。我们发现,当识别失败时,属性信息通常会随着层深度的增加而减少,并且在较早的层中解决属性与更好的准确性相关。此外,LALM严重依赖于查询听觉输入来预测属性,而不是在属性提及位置处的隐藏状态中聚集必要的信息。基于我们的研究结果,我们展示了一种方法来提高LALM。我们的研究结果为听觉属性处理提供了见解,为未来的改进铺平了道路。
摘要:Understanding the internal mechanisms of large audio-language models (LALMs) is crucial for interpreting their behavior and improving performance. This work presents the first in-depth analysis of how LALMs internally perceive and recognize auditory attributes. By applying vocabulary projection on three state-of-the-art LALMs, we track how attribute information evolves across layers and token positions. We find that attribute information generally decreases with layer depth when recognition fails, and that resolving attributes at earlier layers correlates with better accuracy. Moreover, LALMs heavily rely on querying auditory inputs for predicting attributes instead of aggregating necessary information in hidden states at attribute-mentioning positions. Based on our findings, we demonstrate a method to enhance LALMs. Our results offer insights into auditory attribute processing, paving the way for future improvements.


【2】 The NTNU System at the S&I Challenge 2025 SLA Open Track

标题: NTNU系统参加2025年S & I挑战赛SLA公开赛
链接:https://arxiv.org/abs/2506.05121
作者: Hong-Yun Lin,  Tien-Hong Lo,  Yu-Hsuan Fang,  Jhen-Ke Lin,  Chung-Chun Wang,  Hao-Chien Lu,  Berlin Chen 
备注:submitted to the ISCA SLaTE-2025 Workshop
摘要:最近一系列关于口语评估(SLA)的研究采用神经模型(如BERT和wav 2 vec 2.0(W2 V))来评估跨语言和声学模态的口语水平。虽然这两种模式有效地捕捉相关的口语能力的功能,每个表现出特定的模态的局限性。基于BERT的方法依赖于ASR成绩单,这往往无法捕捉韵律和语音线索的二语习得。相比之下,基于W2 V的方法擅长建模声学特征,但缺乏语义可解释性。为了克服这些限制,我们提出了一个系统,通过分数融合策略将W2 V与Phi-4多模态大语言模型(MLLM)集成在一起。该系统在Speak & Improve Challenge 2025的官方测试集上实现了0.375的均方根误差(RMSE),在比赛中获得第二名。作为比较,排名第一、第三和官方基线系统的RMSE分别为0.364、0.384和0.444。
摘要:A recent line of research on spoken language assessment (SLA) employs neural models such as BERT and wav2vec 2.0 (W2V) to evaluate speaking proficiency across linguistic and acoustic modalities. Although both models effectively capture features relevant to oral competence, each exhibits modality-specific limitations. BERT-based methods rely on ASR transcripts, which often fail to capture prosodic and phonetic cues for SLA. In contrast, W2V-based methods excel at modeling acoustic features but lack semantic interpretability. To overcome these limitations, we propose a system that integrates W2V with Phi-4 multimodal large language model (MLLM) through a score fusion strategy. The proposed system achieves a root mean square error (RMSE) of 0.375 on the official test set of the Speak & Improve Challenge 2025, securing second place in the competition. For comparison, the RMSEs of the top-ranked, third-ranked, and official baseline systems are 0.364, 0.384, and 0.444, respectively.


【3】 Survey on the Evaluation of Generative Models in Music

标题: 音乐生成模型评价调查
链接:https://arxiv.org/abs/2506.05104
作者: Alexander Lerch,  Claire Arthur,  Nick Bryan-Kinns,  Corey Ford,  Qianyi Sun,  Ashvala Vinay 
备注:Submitted to ACM CSUR, 26-Jun-2024
摘要:近年来,对音乐生成系统的研究受到了相当大的关注和增长。已经进行了各种尝试来系统地评估这种系统。我们提供了一个跨学科的审查的共同评价目标,方法和指标的系统输出和模型的可用性,包括主观和客观的方法,定性和定量的方法,以及经验和计算方法的评估。我们讨论的优势和挑战,这种方法从音乐学,工程和人机交互的角度来看。
摘要:Research on generative systems in music has seen considerable attention and growth in recent years. A variety of attempts have been made to systematically evaluate such systems. We provide an interdisciplinary review of the common evaluation targets, methodologies, and metrics for the evaluation of both system output and model usability, covering subjective and objective approaches, qualitative and quantitative approaches, as well as empirical and computational methods. We discuss the advantages and challenges of such approaches from a musicological, an engineering, and an HCI perspective.


【4】 Better Semi-supervised Learning for Multi-domain ASR Through Incremental  Retraining and Data Filtering

标题: 通过增量再训练和数据过滤实现更好的多域ASB半监督学习
链接:https://arxiv.org/abs/2506.04981
作者: Andres Carofilis,  Pradeep Rangappa,  Srikanth Madikeri,  Shashi Kumar,  Sergio Burdisso,  Jeena Prakash,  Esau Villatoro-Tello,  Petr Motlicek,  Bidisha Sharma,  Kadri Hacioglu,  Shankar Venkatesan,  Saurabh Vyas,  Andreas Stolcke 
备注:Accepted at Interspeech 2025, Netherlands
摘要:当标记数据稀缺时,对特定领域的预训练ASR模型进行微调具有挑战性。但是来自相关领域的未标记的音频和标记的数据通常是可用的。我们提出了一种增量式半监督学习管道,首先将一个小的域内标记集和一个密切相关领域的辅助数据集集成在一起,与没有辅助数据相比,相对提高了4%。然后应用基于多模型共识或命名实体识别(NER)的过滤来选择和迭代地细化伪标签,与随机选择相比,表现出较慢的性能饱和。在多领域Wow呼叫中心和Fisher英语语料库上的测试结果表明,该方法优于单步微调。基于递归的过滤优于其他方法,与随机选择的单步微调相比,Wow和Fisher的相对改进分别高达22.3%和24.8%。NER是第二好的滤波器,以较低的计算成本提供有竞争力的性能。
摘要:Fine-tuning pretrained ASR models for specific domains is challenging when labeled data is scarce. But unlabeled audio and labeled data from related domains are often available. We propose an incremental semi-supervised learning pipeline that first integrates a small in-domain labeled set and an auxiliary dataset from a closely related domain, achieving a relative improvement of 4% over no auxiliary data. Filtering based on multi-model consensus or named entity recognition (NER) is then applied to select and iteratively refine pseudo-labels, showing slower performance saturation compared to random selection. Evaluated on the multi-domain Wow call center and Fisher English corpora, it outperforms single-step fine-tuning. Consensus-based filtering outperforms other methods, providing up to 22.3% relative improvement on Wow and 24.8% on Fisher over single-step fine-tuning with random selection. NER is the second-best filter, providing competitive performance at a lower computational cost.


【5】 Improving AI-generated music with user-guided training

标题: 通过用户指导训练改善人工智能生成的音乐
链接:https://arxiv.org/abs/2506.04852
作者: Vishwa Mohan Singh,  Sai Anirudh Aryasomayajula,  Ahan Chatterjee,  Beste Aydemir,  Rifat Mehreen Amin 
备注:Select for presentation in HHAI 2025
摘要:人工智能音乐生成技术发展迅速,扩散和自回归算法等模型实现了高保真输出。这些工具可以改变风格,混合乐器,或隔离它们。由于声音可以被可视化为频谱图,因此图像生成算法可以被应用于生成新颖的音乐。然而,这些算法通常是在固定的数据集上训练的,这使得它们很难准确地解释和响应用户输入。这是特别成问题的,因为音乐是高度主观的,并且需要图像生成所不提供的个性化水平。在这项工作中,我们提出了一个人的计算方法,逐步提高这些算法的性能的基础上,用户交互。人工计算元素涉及聚合和选择用户评级,以用作微调模型的损失函数。我们采用了一种遗传算法,结合用户反馈,以提高最初在固定数据集上训练的模型的基线性能。这种方法的有效性是通过每次迭代的用户评分的平均增加来衡量的。在试点测试中,第一次迭代显示,与基线相比,平均评级增加了0.2。第二次迭代进一步改进了这一点,实现了比第一次迭代额外增加0.39。
摘要:AI music generation has advanced rapidly, with models like diffusion and autoregressive algorithms enabling high-fidelity outputs. These tools can alter styles, mix instruments, or isolate them. Since sound can be visualized as spectrograms, image-generation algorithms can be applied to generate novel music. However, these algorithms are typically trained on fixed datasets, which makes it challenging for them to interpret and respond to user input accurately. This is especially problematic because music is highly subjective and requires a level of personalization that image generation does not provide. In this work, we propose a human-computation approach to gradually improve the performance of these algorithms based on user interactions. The human-computation element involves aggregating and selecting user ratings to use as the loss function for fine-tuning the model. We employ a genetic algorithm that incorporates user feedback to enhance the baseline performance of a model initially trained on a fixed dataset. The effectiveness of this approach is measured by the average increase in user ratings with each iteration. In the pilot test, the first iteration showed an average rating increase of 0.2 compared to the baseline. The second iteration further improved upon this, achieving an additional increase of 0.39 over the first iteration.


【6】 MMSU: A Massive Multi-task Spoken Language 

Understanding and Reasoning  Benchmark

标题: MMSU:大规模多任务口语理解和推理基准
链接:https://arxiv.org/abs/2506.04779
作者: Dingdong Wang,  Jincenzi Wu,  Junan Li,  Dongchao Yang,  Xueyuan Chen,  Tianhua Zhang,  Helen Meng 
备注:MMSU benchmark is available at this https URL Evaluation Code is available at this https URL
摘要:语音本身包含丰富的声学信息,远远超出了文本语言。在现实世界的口语理解中,有效的解释通常需要整合语义意义(例如,内容),语言特征(例如,情绪、速度、音高)和语音特征(例如,韵律,语调,节奏),这些都嵌入在语音中。虽然最近的多模态语音大语言模型(SpeechLLM)在处理音频信息方面表现出了显着的能力,但它们在自然语音中执行细粒度感知和复杂推理的能力在很大程度上尚未被探索。为了解决这一差距,我们引入了MMSU,这是一个专门为口语理解和推理而设计的综合基准。MMSU包括5,000个精心策划的音频问答三元组,涉及47个不同的任务。为了使我们的基准建立在语言学理论上,我们系统地结合了广泛的语言现象,包括语音学、韵律学、修辞学、句法学、语义学和语言学。通过对14个高级SpeechLLM的严格评估,我们发现现有模型有很大的改进空间,突出了未来优化的有意义的方向。MMSU为口语理解的综合评估建立了新的标准,为开发更复杂的人-AI语音交互系统提供了有价值的见解。MMSU基准可在https://huggingface.co/datasets/ddwang2000/MMSU上获得。评估代码可在https://github.com/dingdongwang/MMSU_Bench上获得。
摘要:Speech inherently contains rich acoustic information that extends far beyond the textual language. In real-world spoken language understanding, effective interpretation often requires integrating semantic meaning (e.g., content), paralinguistic features (e.g., emotions, speed, pitch) and phonological characteristics (e.g., prosody, intonation, rhythm), which are embedded in speech. While recent multimodal Speech Large Language Models (SpeechLLMs) have demonstrated remarkable capabilities in processing audio information, their ability to perform fine-grained perception and complex reasoning in natural speech remains largely unexplored. To address this gap, we introduce MMSU, a comprehensive benchmark designed specifically for understanding and reasoning in spoken language. MMSU comprises 5,000 meticulously curated audio-question-answer triplets across 47 distinct tasks. To ground our benchmark in linguistic theory, we systematically incorporate a wide range of linguistic phenomena, including phonetics, prosody, rhetoric, syntactics, semantics, and paralinguistics. Through a rigorous evaluation of 14 advanced SpeechLLMs, we identify substantial room for improvement in existing models, highlighting meaningful directions for future optimization. MMSU establishes a new standard for comprehensive assessment of spoken language understanding, providing valuable insights for developing more sophisticated human-AI speech interaction systems. MMSU benchmark is available at https://huggingface.co/datasets/ddwang2000/MMSU. Evaluation Code is available at https://github.com/dingdongwang/MMSU_Bench.


【7】 LLM-based phoneme-to-grapheme for phoneme-based speech recognition

标题: 基于LLM的音素到字形,用于基于音素的语音识别
链接:https://arxiv.org/abs/2506.04711
作者: Te Ma,  Min Bi,  Saierdaer Yusuyin,  Hao Huang,  Zhijian Ou 
备注:Interspeech 2025
摘要:在自动语音识别(ASR)中,基于音素的多语种预训练和跨语种微调是有吸引力的,其高数据效率和竞争力的结果相比,基于子词的模型。然而,基于加权有限状态转换器(WFST)的解码受限于其复杂的流水线和无法利用大型语言模型(LLM)。因此,我们提出了基于LLM的音素到字素(LLM-P2 G)解码的音素为基础的ASR,包括语音到音素(S2 P)和音素到字素(P2 G)。一个挑战是,在级联S2 P和P2 G中似乎存在信息丢失。为了解决这一挑战,我们提出了两种训练策略:带噪声音素的数据增强(DANP)和随机前K边缘化(TKM)训练和解码。我们的实验结果表明,LLM-P2 G优于WFST为基础的系统在跨语言ASR波兰语和德语,相对WER分别减少3.6%和6.9%。
摘要:In automatic speech recognition (ASR), phoneme-based multilingual pre-training and crosslingual fine-tuning is attractive for its high data efficiency and competitive results compared to subword-based models. However, Weighted Finite State Transducer (WFST) based decoding is limited by its complex pipeline and inability to leverage large language models (LLMs). Therefore, we propose LLM-based phoneme-to-grapheme (LLM-P2G) decoding for phoneme-based ASR, consisting of speech-to-phoneme (S2P) and phoneme-to-grapheme (P2G). A challenge is that there seems to have information loss in cascading S2P and P2G. To address this challenge, we propose two training strategies: data augmentation with noisy phonemes (DANP), and randomized top-$K$ marginalized (TKM) training and decoding. Our experimental results show that LLM-P2G outperforms WFST-based systems in crosslingual ASR for Polish and German, by relative WER reductions of 3.6% and 6.9% respectively.


【8】 LESS: Large Language Model Enhanced Semi-Supervised Learning for Speech  Foundational Models

标题: LESS:大语言模型语音基础模型的增强半监督学习
链接:https://arxiv.org/abs/2506.04586
作者: Wen Ding,  Fan Qian 
摘要:我们引入了LESS(大型语言模型增强半监督学习),这是一个多功能的框架,它利用大型语言模型(LLM)来纠正从野外数据中生成的伪标签。在LESS框架内,来自无监督数据的自动语音识别(ASR)或自动语音翻译(AST)的伪标记文本由LLM进行细化,并通过数据过滤策略进行增强,以优化LLM知识传输效率。在普通话ASR和西班牙语到英语AST任务上的实验表明,LESS在Wenet Speech测试集上实现了3.77%的绝对WER降低,在Callhome和Fisher测试集上的BLEU得分分别为34.0和64.7。这些结果验证了LESS在不同语言、任务和领域中的适应性。使用各种LLM和提示配置进行的消融研究为利用LLM衍生的知识用于语音处理应用提供了新的见解。
摘要:We introduce LESS (Large Language Model Enhanced Semi-supervised Learning), a versatile framework that leverages Large Language Models (LLMs) to correct pseudo labels generated from in-the-wild data. Within the LESS framework, pseudo-labeled text from Automatic Speech Recognition (ASR) or Automatic Speech Translation (AST) of the unsupervised data is refined by an LLM, and augmented by a data filtering strategy to optimize LLM knowledge transfer efficiency. Experiments on both Mandarin ASR and Spanish-to-English AST tasks show that LESS achieves a notable absolute WER reduction of 3.77% on the Wenet Speech test set, as well as BLEU scores of 34.0 and 64.7 on Callhome and Fisher test sets respectively. These results validate the adaptability of LESS across different languages, tasks, and domains. Ablation studies conducted with various LLMs and prompt configurations provide novel insights into leveraging LLM-derived knowledge for speech processing applications.


【9】 Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit  and Explicit Grapheme Conditioning

标题: 通过隐式和显式字形条件处理语音的字形连贯音素和韵律注释
链接:https://arxiv.org/abs/2506.04527
作者: Hien Ohnaka,  Yuma Shirahata,  Byeongseon Park,  Ryuichi Yamamoto 
备注:5 pages, 2 figures, and 4 tables, accepted to INTERSPEECH 2025
摘要:我们提出了一个模型,以获得语音的音素和韵律标签是连贯的字素。与之前简单地用标签微调预训练的ASR模型的方法不同,所提出的模型通过两种方法来在相应的字素上调节标签生成:1)通过使用预训练的BERT特征的提示编码器添加隐式字素调节。2)在推理过程中,对与字素不一致的标签假设进行显式剪枝。这些方法使得能够获得语音、标签和字素的并行数据,其适用于各种下游任务,诸如文本到语音和从文本的口音估计。实验表明,该方法显著提高了字素与预测标签的一致性。此外,口音估计任务的实验证实,所提出的方法创建的并行数据有效地提高了估计精度。
摘要:We propose a model to obtain phonemic and prosodic labels of speech that are coherent with graphemes. Unlike previous methods that simply fine-tune a pre-trained ASR model with the labels, the proposed model conditions the label generation on corresponding graphemes by two methods: 1) Add implicit grapheme conditioning through prompt encoder using pre-trained BERT features. 2) Explicitly prune the label hypotheses inconsistent with the grapheme during inference. These methods enable obtaining parallel data of speech, the labels, and graphemes, which is applicable to various downstream tasks such as text-to-speech and accent estimation from text. Experiments showed that the proposed method significantly improved the consistency between graphemes and the predicted labels. Further, experiments on accent estimation task confirmed that the created parallel data by the proposed method effectively improve the estimation accuracy.


【10】 Benchmarking Time-localized Explanations for Audio Classification Models

标题: 音频分类模型的时间本地化简化基准
链接:https://arxiv.org/abs/2506.04391
作者: Cecilia Bolaños,  Leonardo Pepino,  Martin Meza,  Luciana Ferrer 
摘要:大多数现代音频处理方法都是不透明的,因为它们不为自己的决定提供解释。为此,人们提出了各种方法来解释这些模型产生的输出。好的解释可以产生关于数据或模型的有趣见解,并增加对系统的信任。不幸的是,评估解释的质量远非微不足道,因为对于大多数任务来说,没有明确的地面真理解释可供参考。在这项工作中,我们提出了一个基准的音频分类模型,使用目标事件的时间注释作为地面实况解释的代理时间本地化的解释。我们使用这个基准系统地优化和比较各种方法的模型不可知的事后解释,在某些情况下,获得接近完美的解释。最后,我们说明了实用的解释,发现虚假的相关性。
摘要:Most modern approaches for audio processing are opaque, in the sense that they do not provide an explanation for their decisions. For this reason, various methods have been proposed to explain the outputs generated by these models. Good explanations can result in interesting insights about the data or the model, as well as increase trust in the system. Unfortunately, evaluating the quality of explanations is far from trivial since, for most tasks, there is no clear ground truth explanation to use as reference. In this work, we propose a benchmark for time-localized explanations for audio classification models that uses time annotations of target events as a proxy for ground truth explanations. We use this benchmark to systematically optimize and compare various approaches for model-agnostic post-hoc explanation, obtaining, in some cases, close to perfect explanations. Finally, we illustrate the utility of the explanations for uncovering spurious correlations.


【11】 Domain Adaptation Method and Modality Gap Impact in Audio-Text Models  for Prototypical Sound Classification

标题: 原型声音分类音频文本模型中的领域适应方法和情态差距影响
链接:https://arxiv.org/abs/2506.04376
作者: Emiliano Acevedo,  Martín Rocamora,  Magdalena Fuentes 
备注:Accepted at INTERSPEECH 2025
摘要:音频文本模型广泛用于zero-shot环境声音分类,因为它们减轻了对注释数据的需求。然而,我们表明,他们的表现严重下降,在存在背景声源。我们的分析表明,这种退化主要是由背景音景的SNR水平驱动的,并且与背景类型无关。为了解决这个问题,我们提出了一种新的方法,量化和集成的背景源的贡献到分类过程中,提高性能,而不需要模型重新训练。我们的域自适应技术提高了在各种背景和SNR条件下的准确性。此外,我们分析了音频和文本嵌入之间的模态差距,表明缩小这一差距可以提高分类性能。该方法有效地概括了最先进的原型方法,展示了其可扩展性和鲁棒性,为不同的环境。
摘要:Audio-text models are widely used in zero-shot environmental sound classification as they alleviate the need for annotated data. However, we show that their performance severely drops in the presence of background sound sources. Our analysis reveals that this degradation is primarily driven by SNR levels of background soundscapes, and independent of background type. To address this, we propose a novel method that quantifies and integrates the contribution of background sources into the classification process, improving performance without requiring model retraining. Our domain adaptation technique enhances accuracy across various backgrounds and SNR conditions. Moreover, we analyze the modality gap between audio and text embeddings, showing that narrowing this gap improves classification performance. The method generalizes effectively across state-of-the-art prototypical approaches, showcasing its scalability and robustness for diverse environments.


【12】 Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot  Accent Robustness in Low-Resource ASR

标题: 低资源ASB中说话者数量、持续时间和口音多样性对Zero-Shot口音稳健性的影响
链接:https://arxiv.org/abs/2506.04364
作者: Zheng-Xin Yong,  Vineel Pratap,  Michael Auli,  Jean Maillard 
备注:Accepted to INTERSPEECH 2025
摘要:为了构建一个可以为世界上每个人服务的自动语音识别(ASR)系统,ASR需要对各种口音(包括看不见的口音)具有鲁棒性。我们系统地研究了训练数据中的三个不同变量--说话者数量、每个说话者的音频持续时间以及口音的多样性--如何影响低资源训练机制中对看不见口音的ASR鲁棒性。我们观察到,对于固定的ASR培训小时数,增加演讲者的数量(这意味着每个演讲者贡献较少)比每个演讲者贡献的小时数更有益。我们还观察到,更多的扬声器可以通过扩展小时数来提高ASR性能。令人惊讶的是,我们观察到最小的好处,优先考虑不同口音的扬声器时,扬声器的数量受到控制。我们的工作表明,从业者应该优先考虑增加新语言的ASR训练数据组成中的说话者数量。
摘要:To build an automatic speech recognition (ASR) system that can serve everyone in the world, the ASR needs to be robust to a wide range of accents including unseen accents. We systematically study how three different variables in training data -- the number of speakers, the audio duration per each individual speaker, and the diversity of accents -- affect ASR robustness towards unseen accents in a low-resource training regime. We observe that for a fixed number of ASR training hours, it is more beneficial to increase the number of speakers (which means each speaker contributes less) than the number of hours contributed per speaker. We also observe that more speakers enables ASR performance gains from scaling number of hours. Surprisingly, we observe minimal benefits to prioritizing speakers with different accents when the number of speakers is controlled. Our work suggests that practitioners should prioritize increasing the speaker count in ASR training data composition for new languages.


【13】 Can we reconstruct a dysarthric voice with the large speech model Parler  TTS?

标题: 我们可以使用大型语音模型Parler TTC重建发音障碍的声音吗?
链接:https://arxiv.org/abs/2506.04397
作者: Ariadna Sanchez,  Simon King 
备注:Accepted at Interspeech 2025
摘要:语言障碍会使交流变得困难,甚至对那些发展它们的人来说是不可能的。个性化的文本到语音是一个有吸引力的选择,作为一个沟通的援助。我们尝试语音重建使用一个大的语音模型,我们产生一个近似的构音障碍扬声器的声音之前,他们的条件发作。特别是,我们调查是否一个国家的最先进的大型语音模型,Parler TTS,可以生成清晰的语音,同时保持扬声器的身份。我们管理一个数据集,并使用相关的说话者和可懂度信息对其进行注释,并使用它来微调模型。我们的研究结果表明,该模型确实可以学习从这种具有挑战性的数据的分布中生成,但难以控制可懂度并保持一致的说话人身份。我们提出了未来的方向,以提高这类模型的可控性,语音重建任务。
摘要:Speech disorders can make communication hard or even impossible for those who develop them. Personalised Text-to-Speech is an attractive option as a communication aid. We attempt voice reconstruction using a large speech model, with which we generate an approximation of a dysarthric speaker's voice prior to the onset of their condition. In particular, we investigate whether a state-of-the-art large speech model, Parler TTS, can generate intelligible speech while maintaining speaker identity. We curate a dataset and annotate it with relevant speaker and intelligibility information, and use this to fine-tune the model. Our results show that the model can indeed learn to generate from the distribution of this challenging data, but struggles to control intelligibility and to maintain consistent speaker identity. We propose future directions to improve controllability of this class of model, for the voice reconstruction task.


eess.AS音频处理


【1】 Multivariate Probabilistic Assessment of Speech Quality

标题: 言语质量的多元概率评估
链接:https://arxiv.org/abs/2506.04890
作者: Fredrik Cumlin,  Xinyu Liang,  Victor Ungureanu,  Chandan K. A. Reddy,  Christian Schüldt,  Saikat Chatterjee 
备注:Accepted at Interspeech 2025
摘要:平均意见分数(MOS)是评估语音质量的标准度量,但其单一的焦点无法识别特定的失真时,观察到低分数。NISQA数据集通过提供四个额外维度的评级来解决这一限制:噪音,着色,不连续性和响度,以及MOS。在本文中,我们扩展了探索单变量MOS估计的多变量框架建模这些尺寸联合使用多元高斯分布。我们的方法利用Cholesky分解来预测协方差,而不施加限制性的假设和概率仿射变换扩展到一个多变量的背景下。实验结果表明,我们的模型在点估计方面与最先进的方法不相上下,同时独特地提供了跨语音质量维度的不确定性和相关性估计。这可以更好地诊断语音质量差,并提供有针对性的改进。
摘要:The mean opinion score (MOS) is a standard metric for assessing speech quality, but its singular focus fails to identify specific distortions when low scores are observed. The NISQA dataset addresses this limitation by providing ratings across four additional dimensions: noisiness, coloration, discontinuity, and loudness, alongside MOS. In this paper, we extend the explored univariate MOS estimation to a multivariate framework by modeling these dimensions jointly using a multivariate Gaussian distribution. Our approach utilizes Cholesky decomposition to predict covariances without imposing restrictive assumptions and extends probabilistic affine transformations to a multivariate context. Experimental results show that our model performs on par with state-of-the-art methods in point estimation, while uniquely providing uncertainty and correlation estimates across speech quality dimensions. This enables better diagnosis of poor speech quality and informs targeted improvements.


【2】 EMO-Debias: Benchmarking Gender Debiasing Techniques in Multi-Label  Speech Emotion Recognition

标题: EMO-Debias:多标签语音情感识别中的性别去偏见技术基准
链接:https://arxiv.org/abs/2506.04652
作者: Yi-Cheng Lin,  Huang-Cheng Chou,  Yu-Hsuan Li Liang,  Hung-yi Lee 
备注:8 pages
摘要:语音情感识别(SER)系统往往表现出性别偏见。然而,在这种多标签的情况下,现有的去偏置方法的有效性和鲁棒性仍然没有得到充分的研究。为了解决这一差距,我们提出了EMO-Debias,一个大规模的比较13个去偏置方法应用于多标签SER。我们的研究包括预处理,正则化,对抗学习,有偏见的学习者和分布式鲁棒优化技术。在采取行动和自然主义的情感数据集上进行的实验,使用WavLM和XLSR表示,在性别不平衡的条件下评估每种方法。我们的分析量化了公平性和准确性之间的权衡,确定了哪些方法可以在不影响整体模型性能的情况下持续减少性别性能差距。这些发现为选择有效的去偏策略提供了可操作的见解,并突出了数据集分布的影响。
摘要:Speech emotion recognition (SER) systems often exhibit gender bias. However, the effectiveness and robustness of existing debiasing methods in such multi-label scenarios remain underexplored. To address this gap, we present EMO-Debias, a large-scale comparison of 13 debiasing methods applied to multi-label SER. Our study encompasses techniques from pre-processing, regularization, adversarial learning, biased learners, and distributionally robust optimization. Experiments conducted on acted and naturalistic emotion datasets, using WavLM and XLSR representations, evaluate each method under conditions of gender imbalance. Our analysis quantifies the trade-offs between fairness and accuracy, identifying which approaches consistently reduce gender performance gaps without compromising overall model performance. The findings provide actionable insights for selecting effective debiasing strategies and highlight the impact of dataset distributions.


【3】 Towards Efficient Speech-Text Jointly Decoding within One Speech  Language Model

标题: 实现一种语音语言模型内高效的语音-文本联合解码
链接:https://arxiv.org/abs/2506.04518
作者: Haibin Wu,  Yuxuan Hu,  Ruchao Fan,  Xiaofei Wang,  Kenichi Kumatani,  Bo Ren,  Jianwei Yu,  Heng Lu,  Lijuan Wang,  Yao Qian,  Jinyu Li 
摘要:语音语言模型(Speech LM)能够在单个模型中实现端到端的语音文本建模,为口语对话系统提供了一个有前途的方向。语音-文本联合解码模式的选择对解码性能、效率和对齐质量起着至关重要的作用。在这项工作中,我们系统地比较代表性的联合语音文本解码策略,包括交错,并行生成的范例下,使用相同的基础语言模型,语音标记和训练数据的控制实验设置。我们的研究结果表明,交错的方法实现了最佳的对齐。然而,由于长令牌序列长度,它遭受缓慢的推理。为了解决这个问题,我们提出了一种新的早期停止交织(ESI)模式,不仅显着加快解码,但也产生略好的性能。此外,我们还策划了高质量的问答(QA)数据集,以进一步提高语音QA性能。
摘要:Speech language models (Speech LMs) enable end-to-end speech-text modelling within a single model, offering a promising direction for spoken dialogue systems. The choice of speech-text jointly decoding paradigm plays a critical role in performance, efficiency, and alignment quality. In this work, we systematically compare representative joint speech-text decoding strategies-including the interleaved, and parallel generation paradigms-under a controlled experimental setup using the same base language model, speech tokenizer and training data. Our results show that the interleaved approach achieves the best alignment. However it suffers from slow inference due to long token sequence length. To address this, we propose a novel early-stop interleaved (ESI) pattern that not only significantly accelerates decoding but also yields slightly better performance. Additionally, we curate high-quality question answering (QA) datasets to further improve speech QA performance.


【4】 French Listening Tests for the Assessment of Intelligibility, Quality,  and Identity of Body-Conducted Speech Enhancement

标题: 评估身体引导语音增强的可理解性、质量和身份的法语听力测试
链接:https://arxiv.org/abs/2506.04495
作者: Thomas Joubaud,  Julien Hauret,  Véronique Zimpfer,  Éric Bavu 
备注:Submitted to Interspeech 2025 (accepted)
摘要:本研究评估极端带宽扩展网络(EBEN)模型的身体传导传感器通过听力测试。使用Vibravox数据集,我们使用法国改良押韵测试评估可懂度,使用MUSHRA(具有隐藏参考和锚点的多刺激)协议评估语音质量,并使用A/B识别任务评估说话人身份保留。这些实验涉及男性和女性说话者,他们用前额加速度计、刚性入耳式麦克风和喉部麦克风记录。结果证实EBEN提高了语音质量和可懂度。当应用于女性说话人的喉咙麦克风录音时,它会稍微降低说话人识别性能。研究结果还表明,短时客观可懂度(STOI)和身体传导语音的感知质量之间存在相关性,而使用ECAPA 2-TDNN的说话人验证与识别性能保持一致。没有测试指标可靠地预测EBEN的可懂度的影响。
摘要:This study evaluates the Extreme Bandwidth Extension Network (EBEN) model on body-conduction sensors through listening tests. Using the Vibravox dataset, we assess intelligibility with a French Modified Rhyme Test, speech quality with a MUSHRA (MUltiple Stimuli with Hidden Reference and Anchor) protocol and speaker identity preservation with an A/B identification task. The experiments involved male and female speakers recorded with a forehead accelerometer, rigid in-ear and throat microphones. The results confirm that EBEN enhances both speech quality and intelligibility. It slightly degrades speaker identification performance when applied to female speakers' throat microphone recordings. The findings also demonstrate a correlation between Short-Time Objective Intelligibility (STOI) and perceived quality in body-conducted speech, while speaker verification using ECAPA2-TDNN aligns well with identification performance. No tested metric reliably predicts EBEN's effect on intelligibility.


【5】 Bringing Interpretability to Neural Audio Codecs

标题: 为神经音频编解码器带来可解释性
链接:https://arxiv.org/abs/2506.04492
作者: Samir Sadok,  Julien Hauret,  Éric Bavu 
备注:Submitted to Interspeech 2025 (accepted)
摘要:神经音频编解码器的出现越来越受欢迎,因为它们有可能用Transformers有效地建模音频。这种先进的编解码器表示从高度连续波形到低采样离散单元的音频。与语义单元相比,声学单元可能缺乏可解释性,因为它们的训练目标主要集中在重建性能上。本文提出了一个两步的方法来探索编解码器令牌内的语音信息的编码。分析阶段的主要目标是更深入地了解语音属性(如内容、身份和音高)是如何编码的。然后,合成阶段训练AnCoGen网络用于编解码器的事后解释,以直接从相应的令牌中提取语音属性。
摘要:The advent of neural audio codecs has increased in popularity due to their potential for efficiently modeling audio with transformers. Such advanced codecs represent audio from a highly continuous waveform to low-sampled discrete units. In contrast to semantic units, acoustic units may lack interpretability because their training objectives primarily focus on reconstruction performance. This paper proposes a two-step approach to explore the encoding of speech information within the codec tokens. The primary goal of the analysis stage is to gain deeper insight into how speech attributes such as content, identity, and pitch are encoded. The synthesis stage then trains an AnCoGen network for post-hoc explanation of codecs to extract speech attributes from the respective tokens directly.


【6】 Can we reconstruct a dysarthric voice with the large speech model Parler  TTS?

标题: 我们可以使用大型语音模型Parler TTC重建发音障碍的声音吗?
链接:https://arxiv.org/abs/2506.04397
作者: Ariadna Sanchez,  Simon King 
备注:Accepted at Interspeech 2025
摘要:语言障碍会使交流变得困难,甚至对那些发展它们的人来说是不可能的。个性化的文本到语音是一个有吸引力的选择,作为一个沟通的援助。我们尝试语音重建使用一个大的语音模型,我们产生一个近似的构音障碍扬声器的声音之前,他们的条件发作。特别是,我们调查是否一个国家的最先进的大型语音模型,Parler TTS,可以生成清晰的语音,同时保持扬声器的身份。我们管理一个数据集,并使用相关的说话者和可懂度信息对其进行注释,并使用它来微调模型。我们的研究结果表明,该模型确实可以学习从这种具有挑战性的数据的分布中生成,但难以控制可懂度并保持一致的说话人身份。我们提出了未来的方向,以提高这类模型的可控性,语音重建任务。
摘要:Speech disorders can make communication hard or even impossible for those who develop them. Personalised Text-to-Speech is an attractive option as a communication aid. We attempt voice reconstruction using a large speech model, with which we generate an approximation of a dysarthric speaker's voice prior to the onset of their condition. In particular, we investigate whether a state-of-the-art large speech model, Parler TTS, can generate intelligible speech while maintaining speaker identity. We curate a dataset and annotate it with relevant speaker and intelligibility information, and use this to fine-tune the model. Our results show that the model can indeed learn to generate from the distribution of this challenging data, but struggles to control intelligibility and to maintain consistent speaker identity. We propose future directions to improve controllability of this class of model, for the voice reconstruction task.


【7】 Phi-Omni-ST: A multimodal language model for direct speech-to-speech  translation

标题: Phi-Omni-ST:用于直接语音到语音翻译的多模式语言模型
链接:https://arxiv.org/abs/2506.04392
作者: Yuxuan Hu,  Haibin Wu,  Ruchao Fan,  Xiaofei Wang,  Heng Lu,  Yao Qian,  Jinyu Li 
摘要:语音感知语言模型(LM)已经证明了在理解口语的同时生成基于文本的响应的能力。然而,使他们能够有效地产生语音输出仍然是一个挑战。在本文中,我们提出了Phi-Omni-ST,一个用于直接语音到语音翻译(ST)的多模态LM,建立在开源的Phi-4 MM模型上。Phi-Omni-ST通过使用音频Transformer头生成翻译语音来扩展其前身,该头预测音频令牌,相对于文本令牌具有延迟,然后是用于波形合成的流式声码器。我们在CVSS-C数据集上的实验结果证明了Phi-Omni-ST的卓越性能,显著超过了在相同数据集上训练的现有基线模型。此外,当我们扩大训练数据和模型大小时,Phi-Omni-ST达到了与当前SOTA模型相当的性能。
摘要:Speech-aware language models (LMs) have demonstrated capabilities in understanding spoken language while generating text-based responses. However, enabling them to produce speech output efficiently and effectively remains a challenge. In this paper, we present Phi-Omni-ST, a multimodal LM for direct speech-to-speech translation (ST), built on the open-source Phi-4 MM model. Phi-Omni-ST extends its predecessor by generating translated speech using an audio transformer head that predicts audio tokens with a delay relative to text tokens, followed by a streaming vocoder for waveform synthesis. Our experimental results on the CVSS-C dataset demonstrate Phi-Omni-ST's superior performance, significantly surpassing existing baseline models trained on the same dataset. Furthermore, when we scale up the training data and the model size, Phi-Omni-ST reaches on-par performance with the current SOTA model.


【8】 AudioLens: A Closer Look at Auditory Attribute Perception of Large  Audio-Language Models

标题: AudioLens:仔细观察大型音频语言模型的听觉属性感知
链接:https://arxiv.org/abs/2506.05140
作者: Chih-Kai Yang,  Neo Ho,  Yi-Jyun Lee,  Hung-yi Lee 
备注:8 pages, 5 figures, 3 tables
摘要:理解大型音频语言模型(LALM)的内部机制对于解释其行为和提高性能至关重要。这项工作提出了第一次深入分析LALM内部如何感知和识别听觉属性。通过在三个最先进的LALM上应用词汇投影,我们跟踪了属性信息如何在层和标记位置之间演变。我们发现,当识别失败时,属性信息通常会随着层深度的增加而减少,并且在较早的层中解决属性与更好的准确性相关。此外,LALM严重依赖于查询听觉输入来预测属性,而不是在属性提及位置处的隐藏状态中聚集必要的信息。基于我们的研究结果,我们展示了一种方法来提高LALM。我们的研究结果为听觉属性处理提供了见解,为未来的改进铺平了道路。
摘要:Understanding the internal mechanisms of large audio-language models (LALMs) is crucial for interpreting their behavior and improving performance. This work presents the first in-depth analysis of how LALMs internally perceive and recognize auditory attributes. By applying vocabulary projection on three state-of-the-art LALMs, we track how attribute information evolves across layers and token positions. We find that attribute information generally decreases with layer depth when recognition fails, and that resolving attributes at earlier layers correlates with better accuracy. Moreover, LALMs heavily rely on querying auditory inputs for predicting attributes instead of aggregating necessary information in hidden states at attribute-mentioning positions. Based on our findings, we demonstrate a method to enhance LALMs. Our results offer insights into auditory attribute processing, paving the way for future improvements.


【9】 The NTNU System at the S&I Challenge 2025 SLA Open Track

标题: NTNU系统参加2025年S & I挑战赛SLA公开赛
链接:https://arxiv.org/abs/2506.05121
作者: Hong-Yun Lin,  Tien-Hong Lo,  Yu-Hsuan Fang,  Jhen-Ke Lin,  Chung-Chun Wang,  Hao-Chien Lu,  Berlin Chen 
备注:submitted to the ISCA SLaTE-2025 Workshop
摘要:最近一系列关于口语评估(SLA)的研究采用神经模型(如BERT和wav 2 vec 2.0(W2 V))来评估跨语言和声学模态的口语水平。虽然这两种模式有效地捕捉相关的口语能力的功能,每个表现出特定的模态的局限性。基于BERT的方法依赖于ASR成绩单,这往往无法捕捉韵律和语音线索的二语习得。相比之下,基于W2 V的方法擅长建模声学特征,但缺乏语义可解释性。为了克服这些限制,我们提出了一个系统,通过分数融合策略将W2 V与Phi-4多模态大语言模型(MLLM)集成在一起。该系统在Speak & Improve Challenge 2025的官方测试集上实现了0.375的均方根误差(RMSE),在比赛中获得第二名。作为比较,排名第一、第三和官方基线系统的RMSE分别为0.364、0.384和0.444。
摘要:A recent line of research on spoken language assessment (SLA) employs neural models such as BERT and wav2vec 2.0 (W2V) to evaluate speaking proficiency across linguistic and acoustic modalities. Although both models effectively capture features relevant to oral competence, each exhibits modality-specific limitations. BERT-based methods rely on ASR transcripts, which often fail to capture prosodic and phonetic cues for SLA. In contrast, W2V-based methods excel at modeling acoustic features but lack semantic interpretability. To overcome these limitations, we propose a system that integrates W2V with Phi-4 multimodal large language model (MLLM) through a score fusion strategy. The proposed system achieves a root mean square error (RMSE) of 0.375 on the official test set of the Speak & Improve Challenge 2025, securing second place in the competition. For comparison, the RMSEs of the top-ranked, third-ranked, and official baseline systems are 0.364, 0.384, and 0.444, respectively.


【10】 Better Semi-supervised Learning for Multi-domain ASR Through Incremental  Retraining and Data Filtering

标题: 通过增量再训练和数据过滤实现更好的多域ASB半监督学习
链接:https://arxiv.org/abs/2506.04981
作者: Andres Carofilis,  Pradeep Rangappa,  Srikanth Madikeri,  Shashi Kumar,  Sergio Burdisso,  Jeena Prakash,  Esau Villatoro-Tello,  Petr Motlicek,  Bidisha Sharma,  Kadri Hacioglu,  Shankar Venkatesan,  Saurabh Vyas,  Andreas Stolcke 
备注:Accepted at Interspeech 2025, Netherlands
摘要:当标记数据稀缺时,对特定领域的预训练ASR模型进行微调具有挑战性。但是来自相关领域的未标记的音频和标记的数据通常是可用的。我们提出了一种增量式半监督学习管道,首先将一个小的域内标记集和一个密切相关领域的辅助数据集集成在一起,与没有辅助数据相比,相对提高了4%。然后应用基于多模型共识或命名实体识别(NER)的过滤来选择和迭代地细化伪标签,与随机选择相比,表现出较慢的性能饱和。在多领域Wow呼叫中心和Fisher英语语料库上的测试结果表明,该方法优于单步微调。基于递归的过滤优于其他方法,与随机选择的单步微调相比,Wow和Fisher的相对改进分别高达22.3%和24.8%。NER是第二好的滤波器,以较低的计算成本提供有竞争力的性能。
摘要:Fine-tuning pretrained ASR models for specific domains is challenging when labeled data is scarce. But unlabeled audio and labeled data from related domains are often available. We propose an incremental semi-supervised learning pipeline that first integrates a small in-domain labeled set and an auxiliary dataset from a closely related domain, achieving a relative improvement of 4% over no auxiliary data. Filtering based on multi-model consensus or named entity recognition (NER) is then applied to select and iteratively refine pseudo-labels, showing slower performance saturation compared to random selection. Evaluated on the multi-domain Wow call center and Fisher English corpora, it outperforms single-step fine-tuning. Consensus-based filtering outperforms other methods, providing up to 22.3% relative improvement on Wow and 24.8% on Fisher over single-step fine-tuning with random selection. NER is the second-best filter, providing competitive performance at a lower computational cost.


【11】 A Practitioner's Guide to Building ASR Models for Low-Resource  Languages: A Case Study on Scottish Gaelic

标题: 为低资源语言构建SVR模型的从业者指南:苏格兰盖尔语的案例研究
链接:https://arxiv.org/abs/2506.04915
作者: Ondřej Klejch,  William Lamb,  Peter Bell 
备注:Accepted to Interspeech 2025
摘要:为低资源语言开发ASR系统的一种有效方法是微调现有的多语言端到端模型。当原始模型已经在来自许多语言的大量数据上进行了训练时,微调可以在有限的训练数据下有效,即使所讨论的语言不存在于原始训练数据中。公共领域E2E模型的可用性鼓励了这种微调方法,并被广泛认为会带来最先进的结果。然而,本文对这一信念提出了挑战。我们证明了一种将混合障碍与自监督模型相结合的方法可以在有限的训练数据下产生更好的性能。这种组合允许通过持续的自我监督预训练和半监督训练更好地利用所有可用的语音和文本数据。我们以苏格兰盖尔语为基准,实现了WER相对于我们最好的微调Whisper模型减少32%。
摘要:An effective approach to the development of ASR systems for low-resource languages is to fine-tune an existing multilingual end-to-end model. When the original model has been trained on large quantities of data from many languages, fine-tuning can be effective with limited training data, even when the language in question was not present in the original training data. The fine-tuning approach has been encouraged by the availability of public-domain E2E models and is widely believed to lead to state-of-the-art results. This paper, however, challenges that belief. We show that an approach combining hybrid HMMs with self-supervised models can yield substantially better performance with limited training data. This combination allows better utilisation of all available speech and text data through continued self-supervised pre-training and semi-supervised training. We benchmark our approach on Scottish Gaelic, achieving WER reductions of 32% relative over our best fine-tuned Whisper model.


【12】 Improving AI-generated music with user-guided training

标题: 通过用户指导训练改善人工智能生成的音乐
链接:https://arxiv.org/abs/2506.04852
作者: Vishwa Mohan Singh,  Sai Anirudh Aryasomayajula,  Ahan Chatterjee,  Beste Aydemir,  Rifat Mehreen Amin 
备注:Select for presentation in HHAI 2025
摘要:人工智能音乐生成发展迅速,扩散和自回归算法等模型实现了高保真度输出。这些工具可以改变风格,混合乐器,或隔离它们。由于声音可以被可视化为频谱图,因此图像生成算法可以被应用于生成新颖的音乐。然而,这些算法通常是在固定的数据集上训练的,这使得它们很难准确地解释和响应用户输入。这是特别成问题的,因为音乐是高度主观的,并且需要图像生成所不提供的个性化水平。在这项工作中,我们提出了一个人的计算方法,逐步提高这些算法的性能的基础上,用户交互。人工计算元素涉及聚合和选择用户评级,以用作微调模型的损失函数。我们采用了一种遗传算法,结合用户反馈,以提高最初在固定数据集上训练的模型的基线性能。这种方法的有效性是通过每次迭代的用户评分的平均增加来衡量的。在试点测试中,第一次迭代显示,与基线相比,平均评级增加了0.2。第二次迭代进一步改进了这一点,实现了比第一次迭代额外增加0.39。
摘要:AI music generation has advanced rapidly, with models like diffusion and autoregressive algorithms enabling high-fidelity outputs. These tools can alter styles, mix instruments, or isolate them. Since sound can be visualized as spectrograms, image-generation algorithms can be applied to generate novel music. However, these algorithms are typically trained on fixed datasets, which makes it challenging for them to interpret and respond to user input accurately. This is especially problematic because music is highly subjective and requires a level of personalization that image generation does not provide. In this work, we propose a human-computation approach to gradually improve the performance of these algorithms based on user interactions. The human-computation element involves aggregating and selecting user ratings to use as the loss function for fine-tuning the model. We employ a genetic algorithm that incorporates user feedback to enhance the baseline performance of a model initially trained on a fixed dataset. The effectiveness of this approach is measured by the average increase in user ratings with each iteration. In the pilot test, the first iteration showed an average rating increase of 0.2 compared to the baseline. The second iteration further improved upon this, achieving an additional increase of 0.39 over the first iteration.


【13】 IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech  translation

标题: IWWE 2025低资源Bhojpuri到印地语语音翻译的IIIQUAL-BUT系统
链接:https://arxiv.org/abs/2506.04714
作者: Bhavana Akkiraju,  Aishwarya Pothula,  Santosh Kesiraju,  Anil Kumar Vuppala 
备注:Paper is accepted to IWSLT2025
摘要:本文介绍了III-BUT提交给IWIT 2025共享任务的语音翻译低资源Bhojpuri-Hindi语言对。我们探讨了超参数优化和数据增强技术对针对这一特定任务进行微调的无源M4 T模型性能的影响。我们系统地研究了一系列超参数,包括学习率计划、更新步骤数、预热步骤、标签平滑和批量大小;并报告了它们对翻译质量的影响。为了解决数据不足的问题,我们应用了速度扰动和SpecAugment,并研究了它们对翻译质量的影响。我们还研究了跨语言信号的使用,通过联合训练与马拉地语和Bhojpuri语音数据。我们的实验表明,仔细选择的超参数和简单而有效的增强技术的应用显着提高低资源设置的性能。我们还分析了翻译假设,以了解影响BLEU翻译质量的各种错误。
摘要:This paper presents the submission of IIITH-BUT to the IWSLT 2025 shared task on speech translation for the low-resource Bhojpuri-Hindi language pair. We explored the impact of hyperparameter optimisation and data augmentation techniques on the performance of the SeamlessM4T model fine-tuned for this specific task. We systematically investigated a range of hyperparameters including learning rate schedules, number of update steps, warm-up steps, label smoothing, and batch sizes; and report their effect on translation quality. To address data scarcity, we applied speed perturbation and SpecAugment and studied their effect on translation quality. We also examined the use of cross-lingual signal through joint training with Marathi and Bhojpuri speech data. Our experiments reveal that careful selection of hyperparameters and the application of simple yet effective augmentation techniques significantly improve performance in low-resource settings. We also analysed the translation hypotheses to understand various kinds of errors that impacted the translation quality in terms of BLEU.


【14】 LLM-based phoneme-to-grapheme for phoneme-based speech recognition

标题: 基于LLM的音素到字形,用于基于音素的语音识别
链接:https://arxiv.org/abs/2506.04711
作者: Te Ma,  Min Bi,  Saierdaer Yusuyin,  Hao Huang,  Zhijian Ou 
备注:Interspeech 2025
摘要:在自动语音识别(ASR)中,基于音素的多语种预训练和跨语种微调是有吸引力的,其高数据效率和竞争力的结果相比,基于子词的模型。然而,基于加权有限状态转换器(WFST)的解码受限于其复杂的流水线和无法利用大型语言模型(LLM)。因此,我们提出了基于LLM的音素到字素(LLM-P2 G)解码的音素为基础的ASR,包括语音到音素(S2 P)和音素到字素(P2 G)。一个挑战是,在级联S2 P和P2 G中似乎存在信息丢失。为了解决这一挑战,我们提出了两种训练策略:带噪声音素的数据增强(DANP)和随机前K边缘化(TKM)训练和解码。我们的实验结果表明,LLM-P2 G优于WFST为基础的系统在跨语言ASR波兰语和德语,相对WER分别减少3.6%和6.9%。
摘要:In automatic speech recognition (ASR), phoneme-based multilingual pre-training and crosslingual fine-tuning is attractive for its high data efficiency and competitive results compared to subword-based models. However, Weighted Finite State Transducer (WFST) based decoding is limited by its complex pipeline and inability to leverage large language models (LLMs). Therefore, we propose LLM-based phoneme-to-grapheme (LLM-P2G) decoding for phoneme-based ASR, consisting of speech-to-phoneme (S2P) and phoneme-to-grapheme (P2G). A challenge is that there seems to have information loss in cascading S2P and P2G. To address this challenge, we propose two training strategies: data augmentation with noisy phonemes (DANP), and randomized top-$K$ marginalized (TKM) training and decoding. Our experimental results show that LLM-P2G outperforms WFST-based systems in crosslingual ASR for Polish and German, by relative WER reductions of 3.6% and 6.9% respectively.


【15】 LESS: Large Language Model Enhanced Semi-Supervised Learning for Speech  Foundational Models

标题: LESS:大语言模型语音基础模型的增强半监督学习
链接:https://arxiv.org/abs/2506.04586
作者: Wen Ding,  Fan Qian 
摘要:我们引入了LESS(大型语言模型增强半监督学习),这是一个多功能的框架,它利用大型语言模型(LLM)来纠正从野外数据中生成的伪标签。在LESS框架内,来自无监督数据的自动语音识别(ASR)或自动语音翻译(AST)的伪标记文本由LLM进行细化,并通过数据过滤策略进行增强,以优化LLM知识传输效率。在普通话ASR和西班牙语到英语AST任务上的实验表明,LESS在Wenet Speech测试集上实现了3.77%的绝对WER降低,在Callhome和Fisher测试集上的BLEU得分分别为34.0和64.7。这些结果验证了LESS在不同语言、任务和领域中的适应性。使用各种LLM和即时配置进行的消融研究为利用LLM衍生知识进行语音处理应用提供了新的见解。
摘要:We introduce LESS (Large Language Model Enhanced Semi-supervised Learning), a versatile framework that leverages Large Language Models (LLMs) to correct pseudo labels generated from in-the-wild data. Within the LESS framework, pseudo-labeled text from Automatic Speech Recognition (ASR) or Automatic Speech Translation (AST) of the unsupervised data is refined by an LLM, and augmented by a data filtering strategy to optimize LLM knowledge transfer efficiency. Experiments on both Mandarin ASR and Spanish-to-English AST tasks show that LESS achieves a notable absolute WER reduction of 3.77% on the Wenet Speech test set, as well as BLEU scores of 34.0 and 64.7 on Callhome and Fisher test sets respectively. These results validate the adaptability of LESS across different languages, tasks, and domains. Ablation studies conducted with various LLMs and prompt configurations provide novel insights into leveraging LLM-derived knowledge for speech processing applications.


【16】 Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit  and Explicit Grapheme Conditioning

标题: 通过隐式和显式字形条件处理语音的字形连贯音素和韵律注释
链接:https://arxiv.org/abs/2506.04527
作者: Hien Ohnaka,  Yuma Shirahata,  Byeongseon Park,  Ryuichi Yamamoto 
备注:5 pages, 2 figures, and 4 tables, accepted to INTERSPEECH 2025
摘要:我们提出了一个模型,以获得语音的音素和韵律标签是连贯的字素。与之前简单地用标签微调预训练的ASR模型的方法不同,所提出的模型通过两种方法来在相应的字素上调节标签生成:1)通过使用预训练的BERT特征的提示编码器添加隐式字素调节。2)在推理过程中,对与字素不一致的标签假设进行显式剪枝。这些方法使得能够获得语音、标签和字素的并行数据,其适用于各种下游任务,诸如文本到语音和从文本的口音估计。实验表明,该方法显著提高了字素与预测标签的一致性。此外,口音估计任务的实验证实,所提出的方法创建的并行数据有效地提高了估计精度。
摘要:We propose a model to obtain phonemic and prosodic labels of speech that are coherent with graphemes. Unlike previous methods that simply fine-tune a pre-trained ASR model with the labels, the proposed model conditions the label generation on corresponding graphemes by two methods: 1) Add implicit grapheme conditioning through prompt encoder using pre-trained BERT features. 2) Explicitly prune the label hypotheses inconsistent with the grapheme during inference. These methods enable obtaining parallel data of speech, the labels, and graphemes, which is applicable to various downstream tasks such as text-to-speech and accent estimation from text. Experiments showed that the proposed method significantly improved the consistency between graphemes and the predicted labels. Further, experiments on accent estimation task confirmed that the created parallel data by the proposed method effectively improve the estimation accuracy.


【17】 Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot  Accent Robustness in Low-Resource ASR

标题: 低资源ASB中说话者数量、持续时间和口音多样性对Zero-Shot口音稳健性的影响
链接:https://arxiv.org/abs/2506.04364
作者: Zheng-Xin Yong,  Vineel Pratap,  Michael Auli,  Jean Maillard 
备注:Accepted to INTERSPEECH 2025
摘要:为了构建一个可以为世界上每个人服务的自动语音识别(ASR)系统,ASR需要对各种口音(包括看不见的口音)具有鲁棒性。我们系统地研究了训练数据中的三个不同变量--说话者数量、每个说话者的音频持续时间以及口音的多样性--如何影响低资源训练机制中对看不见口音的ASR鲁棒性。我们观察到,对于固定的ASR培训小时数,增加演讲者的数量(这意味着每个演讲者贡献较少)比每个演讲者贡献的小时数更有益。我们还观察到,更多的扬声器可以通过扩展小时数来提高ASR性能。令人惊讶的是,我们观察到最小的好处,优先考虑不同口音的扬声器时,扬声器的数量受到控制。我们的工作表明,从业者应该优先考虑增加新语言的ASR训练数据组成中的说话者数量。
摘要:To build an automatic speech recognition (ASR) system that can serve everyone in the world, the ASR needs to be robust to a wide range of accents including unseen accents. We systematically study how three different variables in training data -- the number of speakers, the audio duration per each individual speaker, and the diversity of accents -- affect ASR robustness towards unseen accents in a low-resource training regime. We observe that for a fixed number of ASR training hours, it is more beneficial to increase the number of speakers (which means each speaker contributes less) than the number of hours contributed per speaker. We also observe that more speakers enables ASR performance gains from scaling number of hours. Surprisingly, we observe minimal benefits to prioritizing speakers with different accents when the number of speakers is controlled. Our work suggests that practitioners should prioritize increasing the speaker count in ASR training data composition for new languages.


机器翻译由腾讯交互翻译提供,仅供参考