微信公众号:arXiv_Daily
cs.SD语音
标题: 迈向阿拉伯语发音评估的统一基准:以《古兰经》背诵为案例研究
链接:https://arxiv.org/abs/2506.07722
备注:Accepted Interspeech 2025 and ArabicNLP Shared Task 2025
摘要:我们提出了一个统一的基准,在现代标准阿拉伯语(MSA)使用古兰经背诵作为案例研究的发音错误检测。我们的方法为推进阿拉伯语发音评估奠定了基础,提供了一个全面的管道,涵盖了数据处理,针对MSA发音的细微差别开发了一个专门的音素集,并为此任务创建了第一个公开可用的测试集,我们称之为古兰经错误发音基准(QuranMB.v1)。此外,我们评估了几个基线模型,以提供初步的性能见解,从而突出的承诺和固有的挑战,在评估MSA发音。通过建立这个标准化的框架,我们的目标是促进进一步的研究和发展发音评估在阿拉伯语技术和相关应用。
摘要:We present a unified benchmark for mispronunciation detection in Modern Standard Arabic (MSA) using Qur'anic recitation as a case study. Our approach lays the groundwork for advancing Arabic pronunciation assessment by providing a comprehensive pipeline that spans data processing, the development of a specialized phoneme set tailored to the nuances of MSA pronunciation, and the creation of the first publicly available test set for this task, which we term as the Qur'anic Mispronunciation Benchmark (QuranMB.v1). Furthermore, we evaluate several baseline models to provide initial performance insights, thereby highlighting both the promise and the challenges inherent in assessing MSA pronunciation. By establishing this standardized framework, we aim to foster further research and development in pronunciation assessment in Arabic language technology and related applications.
【2】 Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
链接:https://arxiv.org/abs/2506.07646
备注:Accepted to INTERSPEECH 2025
摘要:在本文中,我们提出了一种方法来注释音素和韵律标签上一个给定的音频转录对,旨在构建日本的文本到语音(TTS)数据集。我们的方法涉及微调大规模预训练的自动语音识别(ASR)模型,以地面实况成绩单为条件,同时输出短语级字素和注释标签。为了进一步纠正错误的音素标记,我们采用解码策略,利用字典先验知识。客观的评价结果表明,我们提出的方法优于以往的方法,只依赖于文本或音频。主观评价结果表明,语音合成的TTS模型,使用我们的方法标注的标签训练的语音的自然度,是与人工标注训练的模型。
摘要:In this paper, we propose a method for annotating phonemic and prosodic labels on a given audio-transcript pair, aimed at constructing Japanese text-to-speech (TTS) datasets. Our approach involves fine-tuning a large-scale pre-trained automatic speech recognition (ASR) model, conditioned on ground truth transcripts, to simultaneously output phrase-level graphemes and annotation labels. To further correct errors in phonemic labeling, we employ a decoding strategy that utilizes dictionary prior knowledge. The objective evaluation results demonstrate that our proposed method outperforms previous approaches relying solely on text or audio. The subjective evaluation results indicate that the naturalness of speech synthesized by the TTS model, trained with labels annotated using our method, is comparable to that of a model trained with manual annotations.
【3】 Generative Voice Bursts during Phone Call
链接:https://arxiv.org/abs/2506.07526
备注:12 pages, 2 figures
摘要:在紧急情况下,传统的移动电话无法将紧急语音消息传送到已经参与另一呼叫的被叫方。标准呼叫等待警报不提供等待呼叫的紧急程度或内容。本文提出了一种新的方法,用于传输生成语音突发短,上下文感知的音频消息在正在进行的呼叫,从预先授权或动态优先级的呼叫者。通过利用生成AI技术,当呼叫者由于丧失能力或环境限制而无法说话时,系统自动从上下文输入例如位置,健康数据,图像,背景噪音生成语音消息。该解决方案集成了语音、文本和优先级推理机制,允许高优先级紧急消息绕过传统的呼叫等待障碍。该方法采用GPT Neo等模型来生成文本,这些文本被合成为音频,并以可配置的间隔G秒发送,并计数N次,从而确保最小的中断,同时保持紧迫性。该方法具有跨电信、移动终端制造和应急通信平台的重大影响的潜力。
摘要:In critical situations, conventional mobile telephony fails to convey emergency voice messages to a callee already engaged in another call. The standard call waiting alert does not provide the urgency or content of the waiting call. This paper proposes a novel method for transmitting Generative Voice Bursts short, context aware audio messages during ongoing calls, from either preauthorized or dynamically prioritized callers. By leveraging generative AI techniques, the system automatically generates spoken messages from contextual inputs example like location, health data, images, background noise when the caller is unable to speak due to incapacitation or environmental constraints. The solution incorporates voice, text, and priority inference mechanisms, allowing high priority emergency messages to bypass conventional call waiting barriers. The approach employs models such as GPT Neo for generative text, which is synthesized into audio and delivered in configurable intervals G seconds and counts N times, ensuring minimal disruption while preserving urgency. This method holds potential for significant impact across telecom, mobile device manufacturing, and emergency communication platforms.
【4】 LeVo: High-Quality Song Generation with Multi-Preference Alignment
链接:https://arxiv.org/abs/2506.07520
摘要:大型语言模型(LLM)和音频语言模型的最新进展显着改善了音乐生成,特别是在歌词到歌曲的生成。然而,现有的方法仍然与歌曲的复杂组成和高质量数据的稀缺作斗争,导致音质,音乐性,指令遵循和声乐乐器和谐方面的限制。为了应对这些挑战,我们引入了LeVo,一个基于LM的框架,由LeLM和音乐编解码器组成。LeLM能够对两种类型的令牌进行并行建模:混合令牌,代表人声和伴奏的组合音频,以实现人声-乐器和谐,以及双轨令牌,分别对人声和伴奏进行编码,以生成高质量的歌曲。它采用了两个解码器专用的Transformers和一个模块化的扩展训练策略,以防止不同令牌类型之间的干扰。为了进一步提高音乐性和指令遵循,我们引入了一个多偏好对齐方法的基础上直接偏好优化(DPO)。这种方法通过半自动数据构建过程和DPO后训练来处理不同的人类偏好。实验结果表明,LeVo一贯优于现有的方法在客观和主观指标。消融研究进一步证明了我们设计的有效性。音频示例可在https://levo-demo.github.io/上获取。
摘要:Recent advances in large language models (LLMs) and audio language models have significantly improved music generation, particularly in lyrics-to-song generation. However, existing approaches still struggle with the complex composition of songs and the scarcity of high-quality data, leading to limitations in sound quality, musicality, instruction following, and vocal-instrument harmony. To address these challenges, we introduce LeVo, an LM-based framework consisting of LeLM and a music codec. LeLM is capable of parallelly modeling two types of tokens: mixed tokens, which represent the combined audio of vocals and accompaniment to achieve vocal-instrument harmony, and dual-track tokens, which separately encode vocals and accompaniment for high-quality song generation. It employs two decoder-only transformers and a modular extension training strategy to prevent interference between different token types. To further enhance musicality and instruction following, we introduce a multi-preference alignment method based on Direct Preference Optimization (DPO). This method handles diverse human preferences through a semi-automatic data construction process and DPO post-training. Experimental results demonstrate that LeVo consistently outperforms existing methods on both objective and subjective metrics. Ablation studies further justify the effectiveness of our designs. Audio examples are available at https://levo-demo.github.io/.
【5】 Towards Energy-Efficient and Low-Latency Voice-Controlled Smart Homes: A Proposal for Offline Speech Recognition and IoT Integration
链接:https://arxiv.org/abs/2506.07494
摘要:智能家居系统基于人工智能语音识别和物联网技术,使人们能够通过口头命令控制设备,使人们的生活更加高效。然而,现有的人工智能语音识别服务主要部署在互联网上的云平台上。当用户发出命令时,像“Amazon Echo”这样的语音识别设备将通过许多网络节点发布录音,到达多个服务器,然后通过互联网接收响应。这种机制存在几个问题,包括不必要的能量消耗、通信延迟和单点故障的风险。在这份立场文件中,我们提出了一个基于离线语音识别和物联网技术的智能家居概念:1)将离线关键字识别(KWS)技术集成到具有有限资源硬件的家用电器中,使它们能够理解用户的语音命令; 2)设计一个具有分散架构的本地物联网网络来管理和连接各种设备,增强系统的鲁棒性和可扩展性。这一基于离线语音识别和物联网技术的智能家居提案将允许用户在家中任何地方使用低延迟语音控制,而无需依赖互联网,并提供更好的可扩展性和能源可持续性。
摘要:The smart home systems, based on AI speech recognition and IoT technology, enable people to control devices through verbal commands and make people's lives more efficient. However, existing AI speech recognition services are primarily deployed on cloud platforms on the Internet. When users issue a command, speech recognition devices like ``Amazon Echo'' will post a recording through numerous network nodes, reach multiple servers, and then receive responses through the Internet. This mechanism presents several issues, including unnecessary energy consumption, communication latency, and the risk of a single-point failure. In this position paper, we propose a smart home concept based on offline speech recognition and IoT technology: 1) integrating offline keyword spotting (KWS) technologies into household appliances with limited resource hardware to enable them to understand user voice commands; 2) designing a local IoT network with decentralized architecture to manage and connect various devices, enhancing the robustness and scalability of the system. This proposal of a smart home based on offline speech recognition and IoT technology will allow users to use low-latency voice control anywhere in the home without depending on the Internet and provide better scalability and energy sustainability.
【6】 An introduction to pitch strength in contemporary popular music analysis and production
链接:https://arxiv.org/abs/2506.07473
备注:In Music 2024, Innovation in Music Conference, 14-16 June, 2024, Kristiania University College, Oslo, Norway
摘要:音乐信息检索区分音乐的低级和高级描述。目前的生成式AI模型依赖于比录音室音乐家熟悉的控制更高级别的文本描述。音高强度是当代流行音乐的一个低级感知参数,可能是使这种人工智能模型更适合音乐制作的一个特征。信号和感知分析表明,音高强度(1)在歌曲之间和歌曲内部变化显著;(2)有助于小尺度和大尺度结构;(3)有助于处理复调不和谐;(4)可能是从感知丰富性的角度来看可听到的高次谐波的特征。
摘要:Music information retrieval distinguishes between low- and high-level descriptions of music. Current generative AI models rely on text descriptions that are higher level than the controls familiar to studio musicians. Pitch strength, a low-level perceptual parameter of contemporary popular music, may be one feature that could make such AI models more suited to music production. Signal and perceptual analyses suggest that pitch strength (1) varies significantly across and inside songs; (2) contributes to both small- and large-scale structure; (3) contributes to the handling of polyphonic dissonance; and (4) may be a feature of upper harmonics made audible in a perspective of perceptual richness.
【7】 Lightweight Joint Audio-Visual Deepfake Detection via Single-Stream Multi-Modal Learning Framework
链接:https://arxiv.org/abs/2506.07358
摘要:Deepfakes是人工智能合成的多媒体数据,可能会被滥用来传播错误信息。Deepfake生成涉及视觉和音频操作。为了检测视听深度伪造,以前的研究通常使用两个相对独立的子模型来分别学习音频和视觉特征,并随后将其融合用于深度伪造检测。然而,这可能未充分利用音频和视觉特征之间的固有相关性。此外,利用两个孤立的特征学习子模型可能会导致冗余的神经层,使得整个模型对于资源受限的环境来说效率低下且不切实际。 在这项工作中,我们通过单流多模态学习框架设计了一个用于音频视频深度伪造检测的轻量级网络。具体来说,我们引入了协作视听学习模块,以在学习视觉和音频功能的同时有效地集成多模式信息。通过迭代使用这个块,我们的单流网络实现了跨层多模态特征的连续融合。因此,我们的网络可以有效地捕获视觉和音频功能,而无需过多的块堆叠,从而实现轻量级的网络设计。此外,我们提出了一个多模态分类模块,可以提高依赖的视觉和音频分类器的模态内容。它还增强了视频分类器对音频和视觉模态之间的失配的整体抵抗力。我们在DF-TIMIT,FakeAVCeleb和DFDC基准数据集上进行实验。与最先进的视听联合检测方法相比,我们的方法非常轻量级,只有0.48M参数,但它在单模态和多模态deepfake以及看不见的deepfake类型中都具有优势。
摘要:Deepfakes are AI-synthesized multimedia data that may be abused for spreading misinformation. Deepfake generation involves both visual and audio manipulation. To detect audio-visual deepfakes, previous studies commonly employ two relatively independent sub-models to learn audio and visual features, respectively, and fuse them subsequently for deepfake detection. However, this may underutilize the inherent correlations between audio and visual features. Moreover, utilizing two isolated feature learning sub-models can result in redundant neural layers, making the overall model inefficient and impractical for resource-constrained environments. In this work, we design a lightweight network for audio-visual deepfake detection via a single-stream multi-modal learning framework. Specifically, we introduce a collaborative audio-visual learning block to efficiently integrate multi-modal information while learning the visual and audio features. By iteratively employing this block, our single-stream network achieves a continuous fusion of multi-modal features across its layers. Thus, our network efficiently captures visual and audio features without the need for excessive block stacking, resulting in a lightweight network design. Furthermore, we propose a multi-modal classification module that can boost the dependence of the visual and audio classifiers on modality content. It also enhances the whole resistance of the video classifier against the mismatches between audio and visual modalities. We conduct experiments on the DF-TIMIT, FakeAVCeleb, and DFDC benchmark datasets. Compared to state-of-the-art audio-visual joint detection methods, our method is significantly lightweight with only 0.48M parameters, yet it achieves superiority in both uni-modal and multi-modal deepfakes, as well as in unseen types of deepfakes.
【8】 Speech Recognition on TV Series with Video-guided Post-Correction
链接:https://arxiv.org/abs/2506.07323
摘要:自动语音识别(ASR)在深度学习方面取得了巨大成功,推动了会话人工智能、媒体转录和辅助技术的进步。然而,ASR系统仍然在复杂的环境中挣扎,例如电视连续剧,其中重叠的语音,特定领域的术语和长距离上下文依赖性对转录准确性构成了重大挑战。现有的多模态方法无法利用视频中丰富的时间和上下文信息来校正ASR输出。为了解决这个问题,我们提出了一种新的多模态后校正框架,利用从视频中提取的上下文线索来细化ASR transmittance。我们的框架包括两个阶段:ASR生成和基于视频的后期校正,其中第一阶段产生初始成绩单,第二阶段使用基于视频的上下文信息提取和上下文感知ASR校正来校正错误。我们采用视频大多模态模型(VLMM)提取关键的上下文信息,使用量身定制的提示,然后将其与大语言模型(LLM)集成,以完善ASR输出。我们评估我们的方法在多模态基准的电视剧ASR,并证明其有效性,提高ASR的性能,利用基于视频的上下文,以提高转录的准确性,在复杂的多媒体环境。
摘要:Automatic Speech Recognition (ASR) has achieved remarkable success with deep learning, driving advancements in conversational artificial intelligence, media transcription, and assistive technologies. However, ASR systems still struggle in complex environments such as TV series, where overlapping speech, domain-specific terminology, and long-range contextual dependencies pose significant challenges to transcription accuracy. Existing multimodal approaches fail to correct ASR outputs with the rich temporal and contextual information available in video. To address this limitation, we propose a novel multimodal post-correction framework that refines ASR transcriptions by leveraging contextual cues extracted from video. Our framework consists of two stages: ASR Generation and Video-based Post-Correction, where the first stage produces the initial transcript and the second stage corrects errors using Video-based Contextual Information Extraction and Context-aware ASR Correction. We employ the Video-Large Multimodal Model (VLMM) to extract key contextual information using tailored prompts, which is then integrated with a Large Language Model (LLM) to refine the ASR output. We evaluate our method on a multimodal benchmark for TV series ASR and demonstrate its effectiveness in improving ASR performance by leveraging video-based context to enhance transcription accuracy in complex multimedia environments.
【9】 Towards Generalized Source Tracing for Codec-Based Deepfake Speech
链接:https://arxiv.org/abs/2506.07294
备注:Submitted to IEEE ASRU 2025
摘要:最近尝试对基于编解码器的深度伪造语音(CodecFake)进行源跟踪,由基于神经音频编解码器的语音生成(CoSG)模型生成,表现出次优性能。然而,如何使用模拟的CoSG数据训练源跟踪模型,同时保持对真实CoSG生成的音频的强大性能仍然是一个开放的挑战。在本文中,我们表明,仅在编解码器重新合成的数据上训练的模型往往过拟合非语音区域,并且难以推广到看不见的内容。为了缓解这些挑战,我们引入了语义声源跟踪网络(SASTNet),它联合利用Whisper进行语义特征编码,并利用Wav2vec2与AudioMAE进行声学特征编码。我们提出的SASTNet在CodecFake+数据集的CoSG测试集上实现了最先进的性能,证明了其可靠源跟踪的有效性。
摘要:Recent attempts at source tracing for codec-based deepfake speech (CodecFake), generated by neural audio codec-based speech generation (CoSG) models, have exhibited suboptimal performance. However, how to train source tracing models using simulated CoSG data while maintaining strong performance on real CoSG-generated audio remains an open challenge. In this paper, we show that models trained solely on codec-resynthesized data tend to overfit to non-speech regions and struggle to generalize to unseen content. To mitigate these challenges, we introduce the Semantic-Acoustic Source Tracing Network (SASTNet), which jointly leverages Whisper for semantic feature encoding and Wav2vec2 with AudioMAE for acoustic feature encoding. Our proposed SASTNet achieves state-of-the-art performance on the CoSG test set of the CodecFake+ dataset, demonstrating its effectiveness for reliable source tracing.
【10】 Methods for pitch analysis in contemporary popular music: Vitalic's use of tones that do not operate on the principle of acoustic resonance
链接:https://arxiv.org/abs/2506.07207
摘要:Vitalic是一位电子音乐制作人,自2001年以来一直活跃。Vitalic的2005年曲目“No Fun”的主要合成器部分是由一系列单一的不和谐音调组成的,可以同时唤起两种旋律。这一部分作为一个起点,检查维塔利奇的使用音调,不运作的原则,声学共振。这项研究认为,引起两个或两个以上的同时音高的音调,并检查各种不和谐的部分布局。维塔利克音乐之外的例子也表明,类似的音调属性可以在当代流行音乐的其他地方找到。
摘要:Vitalic is an electronic music producer who has been active since 2001. Vitalic's 2005 track "No Fun" features a main synthesiser part built from a sequence of single inharmonic tones that evoke two simultaneous melodies. This part serves as a starting point for examining Vitalic's use of tones that do not operate on the principle of acoustic resonance. The study considers tones that evoke two or more simultaneous pitches and examines various inharmonic partial layouts. Examples outside Vitalic's music are also provided to suggest that similar tone properties can be found elsewhere in contemporary popular music.
【11】 Audio synthesizer inversion in symmetric parameter spaces with approximately equivariant flow matching
链接:https://arxiv.org/abs/2506.07199
备注:Accepted at ISMIR 2025
摘要:许多音频合成器可以在给定不同参数配置的情况下产生相同的信号,这意味着从声音到参数的反演是一个固有的不适定问题。我们表明,这在很大程度上是由于固有的对称性的合成器,并特别关注置换不变性。首先,我们证明了在一个合成任务,回归点估计置换对称性降低性能,即使使用置换不变的损失函数或破环算法。然后,查看等价的解决方案的概率分布的模式,我们表明,条件生成模型大大提高了性能。此外,承认隐式参数分布的不变性,我们发现,通过使用置换等变连续归一化流,性能得到进一步提高。为了适应复杂的对称性在真正的合成器,我们还提出了一个宽松的等方差策略,自适应地发现相关的对称性数据。将我们的方法应用于Surge XT,一种用于现实世界音频制作的全功能开源合成器,我们发现我们的方法在音频重建指标上优于回归和生成基线。
摘要:Many audio synthesizers can produce the same signal given different parameter configurations, meaning the inversion from sound to parameters is an inherently ill-posed problem. We show that this is largely due to intrinsic symmetries of the synthesizer, and focus in particular on permutation invariance. First, we demonstrate on a synthetic task that regressing point estimates under permutation symmetry degrades performance, even when using a permutation-invariant loss function or symmetry-breaking heuristics. Then, viewing equivalent solutions as modes of a probability distribution, we show that a conditional generative model substantially improves performance. Further, acknowledging the invariance of the implicit parameter distribution, we find that performance is further improved by using a permutation equivariant continuous normalizing flow. To accommodate intricate symmetries in real synthesizers, we also propose a relaxed equivariance strategy that adaptively discovers relevant symmetries from data. Applying our method to Surge XT, a full-featured open source synthesizer used in real world audio production, we find our method outperforms regression and generative baselines across audio reconstruction metrics.
【12】 Technical Report: A Practical Guide to Kaldi ASR Optimization
链接:https://arxiv.org/abs/2506.07149
摘要:本技术报告介绍了基于Kaldi的自动语音识别(ASR)系统的创新优化,重点关注声学模型增强,超参数调整和语言模型效率。我们开发了一个自定义的Conformer块,集成了多流TDNN-F结构,实现了卓越的特征提取和时间建模。我们的方法包括先进的数据增强技术和动态超参数优化,以提高性能并减少过拟合。此外,我们提出了强大的语言模型管理策略,采用贝叶斯优化和$n$-gram修剪,以确保相关性和计算效率。这些系统性改进显著提高了ASR的准确性和鲁棒性,优于现有方法,并为各种语音识别场景提供了可扩展的解决方案。本报告强调了战略优化在保持Kaldi在快速发展的技术环境中的适应性和竞争力方面的重要性。
摘要:This technical report introduces innovative optimizations for Kaldi-based Automatic Speech Recognition (ASR) systems, focusing on acoustic model enhancement, hyperparameter tuning, and language model efficiency. We developed a custom Conformer block integrated with a multistream TDNN-F structure, enabling superior feature extraction and temporal modeling. Our approach includes advanced data augmentation techniques and dynamic hyperparameter optimization to boost performance and reduce overfitting. Additionally, we propose robust strategies for language model management, employing Bayesian optimization and $n$-gram pruning to ensure relevance and computational efficiency. These systematic improvements significantly elevate ASR accuracy and robustness, outperforming existing methods and offering a scalable solution for diverse speech recognition scenarios. This report underscores the importance of strategic optimizations in maintaining Kaldi's adaptability and competitiveness in rapidly evolving technological landscapes.
【13】 RBA-FE: A Robust Brain-Inspired Audio Feature Extractor for Depression Diagnosis
链接:https://arxiv.org/abs/2506.07118
备注:14 pages
摘要:本文使用改进的分层网络架构,提出了一种用于抑郁症诊断的鲁棒脑启发音频特征提取器(RBA-FE)模型。大多数深度学习模型在基于图像的诊断任务中实现了最先进的性能,忽略了对应的音频特征。为了应对噪声挑战,RBA-FE利用从原始音频中提取的六个声学特征,捕获空间特征和时间依赖性。这种混合属性有助于缓解其他学习模型(如深度残差收缩网络)中音频特征提取的精度限制。为了处理噪声问题,我们的模型采用了一种改进的尖峰神经元模型,称为自适应速率平滑泄漏积分和发射(ARSLIF)。ARSLIF模型模拟了大脑注意系统中“细胞信号选择性的重新调谐”机制,增强了模型对音频数据中环境噪声的鲁棒性。实验结果表明,RBA-FE在MODMA数据集上达到了最先进的准确率,精确率,准确率,召回率和F1得分分别为0.8750,0.8974,0.8750和0.8750。在AVEC 2014和DAIC-WOZ数据集上进行的大量实验都显示出噪声鲁棒性的增强。通过比较进一步表明,ARSLIF神经元模型在对抑郁音频数据的特征提取中提出了异常放电模式,提供了大脑启发的可解释性。
摘要:This article proposes a robust brain-inspired audio feature extractor (RBA-FE) model for depression diagnosis, using an improved hierarchical network architecture. Most deep learning models achieve state-of-the-art performance for image-based diagnostic tasks, ignoring the counterpart audio features. In order to tailor the noise challenge, RBA-FE leverages six acoustic features extracted from the raw audio, capturing both spatial characteristics and temporal dependencies. This hybrid attribute helps alleviate the precision limitation in audio feature extraction within other learning models like deep residual shrinkage networks. To deal with the noise issues, our model incorporates an improved spiking neuron model, called adaptive rate smooth leaky integrate-and-fire (ARSLIF). The ARSLIF model emulates the mechanism of ``retuning of cellular signal selectivity" in the brain attention systems, which enhances the model robustness against environmental noises in audio data. Experimental results demonstrate that RBA-FE achieves state-of-the-art accuracy on the MODMA dataset, respectively with 0.8750, 0.8974, 0.8750 and 0.8750 in precision, accuracy, recall and F1 score. Extensive experiments on the AVEC2014 and DAIC-WOZ datasets both show enhancements in noise robustness. It is further indicated by comparison that the ARSLIF neuron model suggest the abnormal firing pattern within the feature extraction on depressive audio data, offering brain-inspired interpretability.
【14】 Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
链接:https://arxiv.org/abs/2506.07081
摘要:准确、低延迟的端点对于有效的口语对话系统至关重要。虽然传统的端点通常依赖于基于频谱的音频特征,但这项工作提出了使用流式低比特率神经音频编解码器(NAC)特征的多回合对话的实时语音端点,该特征建立在神经音频编解码器的最新进展基础上。为了进一步减少截止误差,我们引入了一种新的标签延迟训练方案。在160 ms的固定中值延迟下,我们的NAC和标签延迟组合方法实现了显着的相对截止误差降低:与基线方法相比,单流端点为42.7%,双流配置为37.5%。最后,我们展示了与基于编解码器的预训练语音大语言模型的有效集成,将其中值响应时间提高了1200 ms,并将其截止误差降低了35%。
摘要:Accurate, low-latency endpointing is crucial for effective spoken dialogue systems. While traditional endpointers often rely on spectrum-based audio features, this work proposes real-time speech endpointing for multi-turn dialogues using streaming, low-bitrate Neural Audio Codec (NAC) features, building upon recent advancements in neural audio codecs. To further reduce cutoff errors, we introduce a novel label delay training scheme. At a fixed median latency of 160 ms, our combined NAC and label delay approach achieves significant relative cutoff error reductions: 42.7% for a single-stream endpointer and 37.5% for a two-stream configuration, compared to baseline methods. Finally, we demonstrate efficient integration with a codec-based pretrained speech large language model, improving its median response time by 1200 ms and reducing its cutoff error by 35%.
【15】 E-BATS: Efficient Backpropagation-Free Test-Time Adaptation for Speech Foundation Models
链接:https://arxiv.org/abs/2506.07078
备注:Under Review
摘要:当部署在涉及声学域偏移(例如背景噪声和说话者口音)的真实场景中时,语音基础模型会遇到显著的性能下降。测试时自适应(TTA)最近出现了一个可行的策略,以解决这种域的变化,在推理时,而不需要访问源数据或标签。然而,现有的TTA方法,特别是那些依赖于反向传播的方法,是内存密集型的,限制了它们在语音任务和资源受限环境中的适用性。虽然反向传播免费的方法提供了提高的效率,现有的表现出较差的精度。这是因为它们主要是为视觉任务而开发的,而视觉任务与语音任务的制定、噪声特性和模型架构有着根本的不同,这带来了独特的可移植性挑战。在本文中,我们介绍了E-BATS,第一个高效的BAckpropagation-free TTA框架明确设计的语音基础模型。E-BATS通过三个关键组成部分实现了自适应有效性和记忆效率之间的平衡:(i)用于基于前向传递的特征对齐的轻量级即时自适应,(ii)用于捕获全局(话语级)和局部分布偏移(令牌级)的多尺度损失,以及(iii)用于跨话语稳定自适应的测试时间指数移动平均机制。在跨越16种声学条件的四个噪声语音数据集上进行的实验证明了一致的改进,与基于反向传播的方法相比,无反向传播基线的准确率提高了4.1%-13.5%,GPU内存节省了2.0-6.4倍。通过在声学变化下实现可扩展和鲁棒的自适应,这项工作为在现实环境中为实际语音处理系统开发更有效的自适应方法铺平了道路。
摘要:Speech Foundation Models encounter significant performance degradation when deployed in real-world scenarios involving acoustic domain shifts, such as background noise and speaker accents. Test-time adaptation (TTA) has recently emerged as a viable strategy to address such domain shifts at inference time without requiring access to source data or labels. However, existing TTA approaches, particularly those relying on backpropagation, are memory-intensive, limiting their applicability in speech tasks and resource-constrained settings. Although backpropagation-free methods offer improved efficiency, existing ones exhibit poor accuracy. This is because they are predominantly developed for vision tasks, which fundamentally differ from speech task formulations, noise characteristics, and model architecture, posing unique transferability challenges. In this paper, we introduce E-BATS, the first Efficient BAckpropagation-free TTA framework designed explicitly for speech foundation models. E-BATS achieves a balance between adaptation effectiveness and memory efficiency through three key components: (i) lightweight prompt adaptation for a forward-pass-based feature alignment, (ii) a multi-scale loss to capture both global (utterance-level) and local distribution shifts (token-level) and (iii) a test-time exponential moving average mechanism for stable adaptation across utterances. Experiments conducted on four noisy speech datasets spanning sixteen acoustic conditions demonstrate consistent improvements, with 4.1%-13.5% accuracy gains over backpropagation-free baselines and 2.0-6.4 times GPU memory savings compared to backpropagation-based methods. By enabling scalable and robust adaptation under acoustic variability, this work paves the way for developing more efficient adaptation approaches for practical speech processing systems in real-world environments.
【16】 Insights on Harmonic Tones from a Generative Music Experiment
链接:https://arxiv.org/abs/2506.07073
备注:15th International Workshop on Machine Learning and Music, September 9, 2024, Vilnius, Lithuania
摘要:生成式音乐AI的最终目的是音乐制作。工作室实验室是跨学科艺术科学分支中的一种社会形式,是一种利用人工智能音乐模型推进音乐制作的方式。在一个涉及研究人员、音乐制作人和一个用于音乐生成类似音乐的音频的人工智能模型的工作室实验中,人们观察到制作人使用模型的输出来用一个谐波复音来传达两个或更多个音高,这反过来又表明模型已经学会了使用谐波复音的单声道序列来生成结构化和连贯的同时旋律线。这些发现促使人们重新思考长期以来关于人类是否可以将谐波视为不同音高的争论,并强调生成人工智能不仅可以增强音乐创造力,还有助于更深入地理解音乐。
摘要:The ultimate purpose of generative music AI is music production. The studio-lab, a social form within the art-science branch of cross-disciplinarity, is a way to advance music production with AI music models. During a studio-lab experiment involving researchers, music producers, and an AI model for music generating bass-like audio, it was observed that the producers used the model's output to convey two or more pitches with a single harmonic complex tone, which in turn revealed that the model had learned to generate structured and coherent simultaneous melodic lines using monophonic sequences of harmonic complex tones. These findings prompt a reconsideration of the long-standing debate on whether humans can perceive harmonics as distinct pitches and highlight how generative AI can not only enhance musical creativity but also contribute to a deeper understanding of music.
【17】 "In This Environment, As That Speaker": A Text-Driven Framework for Multi-Attribute Speech Conversion
备注:Accepted by Interspeech2025
摘要:我们提出了TES-VC(文本驱动的环境和扬声器可控语音转换),一个文本驱动的语音转换框架与独立控制扬声器音色和环境声学。TES-VC处理目标语音和环境的同时文本输入,准确地生成与所描述的音色/环境匹配的语音,同时保留源内容。通过潜在扩散模型对具有解耦的声音/环境特征的合成数据进行训练,我们的方法消除了属性之间的干扰。基于检索的音色控制(RBTC)模块可以使用抽象描述进行精确操作,而无需配对数据。实验证明TES-VC能有效地生成在音质和环境上都符合语境的语音,具有较高的内容保持率和较好的可控性,具有广泛的应用前景。
摘要:We propose TES-VC (Text-driven Environment and Speaker controllable Voice Conversion), a text-driven voice conversion framework with independent control of speaker timbre and environmental acoustics. TES-VC processes simultaneous text inputs for target voice and environment, accurately generating speech matching described timbre/environment while preserving source content. Trained on synthetic data with decoupled vocal/environment features via latent diffusion modeling, our method eliminates interference between attributes. The Retrieval-Based Timbre Control (RBTC) module enables precise manipulation using abstract descriptions without paired data. Experiments confirm TES-VC effectively generates contextually appropriate speech in both timbre and environment with high content retention and superior controllability which demonstrates its potential for widespread applications.
【18】 Automatic Speech Recognition of African American English: Lexical and Contextual Effects
链接:https://arxiv.org/abs/2506.06888
备注:submitted to Interspeech 2025
摘要:自动语音识别(ASR)模型经常与非洲裔美国人英语(AAE)中的语音,音韵和形态句法特征相斗争。本研究的重点是两个关键的AAE变量:辅音集群减少(CCR)和ING减少。它检查CCR和ING减少的存在是否会增加ASR错误识别。随后,它调查是否没有外部语言模型(LM)的端到端的ASR系统更受词汇邻域效应和上下文的可预测性相比,LM系统。区域非裔美国人语言语料库(CORAAL)使用wav2vec 2.0(有和没有LM)进行转录。CCR和ING减少检测使用蒙特利尔强制对齐(MFA)发音扩展。分析表明,CCR和ING对词错误率(WER)的影响很小,但显着,并表明在没有LM的ASR系统中存在更强的词汇邻居效应。
摘要:Automatic Speech Recognition (ASR) models often struggle with the phonetic, phonological, and morphosyntactic features found in African American English (AAE). This study focuses on two key AAE variables: Consonant Cluster Reduction (CCR) and ING-reduction. It examines whether the presence of CCR and ING-reduction increases ASR misrecognition. Subsequently, it investigates whether end-to-end ASR systems without an external Language Model (LM) are more influenced by lexical neighborhood effect and less by contextual predictability compared to systems with an LM. The Corpus of Regional African American Language (CORAAL) was transcribed using wav2vec 2.0 with and without an LM. CCR and ING-reduction were detected using the Montreal Forced Aligner (MFA) with pronunciation expansion. The analysis reveals a small but significant effect of CCR and ING on Word Error Rate (WER) and indicates a stronger presence of lexical neighborhood effect in ASR systems without LMs.
【19】 Multimodal Spatial Language Maps for Robot Navigation and Manipulation
链接:https://arxiv.org/abs/2506.06862
备注:accepted to International Journal of Robotics Research (IJRR). 24 pages, 18 figures. The paper contains texts from VLMaps(arXiv:2210.05714) and AVLMaps(arXiv:2303.07522). The project page is this https URL
摘要:将语言接地到导航代理的观察可以利用预先训练的多模态基础模型来将感知与对象或事件描述相匹配。然而,以前的方法仍然断开与环境映射,缺乏几何地图的空间精度,或忽视视觉以外的附加模态信息。为了解决这个问题,我们提出了多模态空间语言地图作为空间地图表示,融合了预先训练的多模态特征与环境的3D重建。我们使用标准探索自主构建这些地图。我们提出了两个实例,我们的地图,这是视觉语言地图(VLMaps)和他们的扩展视听语言地图(AVLMaps)通过添加音频信息。当与大型语言模型(LLM)结合时,VLMaps可以(i)将自然语言命令翻译成开放词汇空间目标(例如,“在沙发和电视之间”)直接定位在地图中,以及(ii)跨不同的机器人实施例共享以按需生成定制的障碍物地图。基于上述功能,AVLMaps通过引入统一的3D空间表示来扩展VLMaps,该空间表示通过融合来自预训练的多模态基础模型的特征来集成音频,视觉和语言线索。这使得机器人能够基于多模态目标查询(例如,文本、图像或音频片段)到用于导航的空间位置。此外,不同的感官输入的结合显着提高目标消歧在模糊的环境。在模拟和现实世界中的实验表明,我们的多模态空间语言地图,使zero-shot空间和多模态目标导航和提高召回率50%,在模糊的情况下。这些功能扩展到移动机器人和桌面机械手,支持由视觉,音频和空间提示引导的导航和交互。
摘要:Grounding language to a navigating agent's observations can leverage pretrained multimodal foundation models to match perceptions to object or event descriptions. However, previous approaches remain disconnected from environment mapping, lack the spatial precision of geometric maps, or neglect additional modality information beyond vision. To address this, we propose multimodal spatial language maps as a spatial map representation that fuses pretrained multimodal features with a 3D reconstruction of the environment. We build these maps autonomously using standard exploration. We present two instances of our maps, which are visual-language maps (VLMaps) and their extension to audio-visual-language maps (AVLMaps) obtained by adding audio information. When combined with large language models (LLMs), VLMaps can (i) translate natural language commands into open-vocabulary spatial goals (e.g., "in between the sofa and TV") directly localized in the map, and (ii) be shared across different robot embodiments to generate tailored obstacle maps on demand. Building upon the capabilities above, AVLMaps extend VLMaps by introducing a unified 3D spatial representation integrating audio, visual, and language cues through the fusion of features from pretrained multimodal foundation models. This enables robots to ground multimodal goal queries (e.g., text, images, or audio snippets) to spatial locations for navigation. Additionally, the incorporation of diverse sensory inputs significantly enhances goal disambiguation in ambiguous environments. Experiments in simulation and real-world settings demonstrate that our multimodal spatial language maps enable zero-shot spatial and multimodal goal navigation and improve recall by 50% in ambiguous scenarios. These capabilities extend to mobile robots and tabletop manipulators, supporting navigation and interaction guided by visual, audio, and spatial cues.
【20】 Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs
链接:https://arxiv.org/abs/2506.06820
摘要:音频大语言模型(AudioLLM)在语音识别和翻译等语义任务中取得了很好的效果,但在建模情感等非语言线索方面仍然有限。现有的方法通常将情绪理解视为分类问题,对预测背后的基本原理几乎没有深入了解。在这项工作中,我们探索情感推理,这是一种利用AudioLLM的生成能力,通过产生语义对齐,以证据为基础的解释来增强情感识别的策略。为了在多任务AudioLLM中支持这一点,我们引入了一个统一的框架,该框架结合了推理增强的数据监督,双编码器架构和任务交替训练。这种方法使AudioLLM能够有效地学习不同的任务,同时结合情感推理。IEMOCAP和MELD上的实验表明,我们的方法不仅提高了情感预测的准确性,但也提高了生成的响应的一致性和证据基础。
摘要:Audio Large Language Models (AudioLLMs) have achieved strong results in semantic tasks like speech recognition and translation, but remain limited in modeling paralinguistic cues such as emotion. Existing approaches often treat emotion understanding as a classification problem, offering little insight into the underlying rationale behind predictions. In this work, we explore emotion reasoning, a strategy that leverages the generative capabilities of AudioLLMs to enhance emotion recognition by producing semantically aligned, evidence-grounded explanations. To support this in multitask AudioLLMs, we introduce a unified framework combining reasoning-augmented data supervision, dual-encoder architecture, and task-alternating training. This approach enables AudioLLMs to effectively learn different tasks while incorporating emotional reasoning. Experiments on IEMOCAP and MELD show that our approach not only improves emotion prediction accuracy but also enhances the coherence and evidential grounding of the generated responses.
【21】 SynHate: Detecting Hate Speech in Synthetic Deepfake Audio
链接:https://arxiv.org/abs/2506.06772
备注:Accepted in Interspeech 2025
摘要:Deepfake音频和仇恨言论的兴起,由先进的文本到语音技术提供支持,威胁到在线安全。SynHate是第一个用于检测合成音频中仇恨言论的多语言数据集,涵盖37种语言。SynHate使用了一种新颖的四类方案:真实正常,真实仇恨,假正常和假仇恨。它基于MuTox和ADIMA数据集构建,捕捉了全球和印度的各种仇恨言论模式。我们评估了五种领先的自监督模型(Whisper-small/medium,XLS-R,AST,mHuBERT),发现不同语言的性能差异显著,Whisper-small整体表现最好。跨数据集泛化仍然是一个挑战。通过发布SynHate和基线代码,我们的目标是推进针对合成仇恨言论的强大,文化敏感和多语言解决方案。该数据集可在https://www.iab-rubric.org/resources上获得。
摘要:The rise of deepfake audio and hate speech, powered by advanced text-to-speech, threatens online safety. We present SynHate, the first multilingual dataset for detecting hate speech in synthetic audio, spanning 37 languages. SynHate uses a novel four-class scheme: Real-normal, Real-hate, Fake-normal, and Fake-hate. Built from MuTox and ADIMA datasets, it captures diverse hate speech patterns globally and in India. We evaluate five leading self-supervised models (Whisper-small/medium, XLS-R, AST, mHuBERT), finding notable performance differences by language, with Whisper-small performing best overall. Cross-dataset generalization remains a challenge. By releasing SynHate and baseline code, we aim to advance robust, culturally sensitive, and multilingual solutions against synthetic hate speech. The dataset is available at https://www.iab-rubric.org/resources.
【22】 Can Quantized Audio Language Models Perform Zero-Shot Spoofing Detection?
链接:https://arxiv.org/abs/2506.06756
备注:Accepted in Interspeech 2025
摘要:量化对于在资源受限的环境中有效地部署大型音频语言模型(LALM)至关重要。然而,它对复杂任务的影响,如zero-shot音频欺骗检测,仍然没有得到充分的研究。这项研究评估了五种LALM的zero-shot能力,GAMA,LTU-AS,MERaLiON,Qwen-Audio和SALMONN,跨越三个不同的数据集:ASVspoof 2019,In-the-Wild和WaveFake,并研究了它们对量化的鲁棒性(FP 32,FP 16,INT 8)。尽管初始欺骗检测精度很高,但我们的分析表明,所有模型对欺骗分类都存在严重的预测偏差,使其实际性能等同于随机分类。有趣的是,与FP 32相比,量化到FP 16精度导致的性能下降可以忽略不计,有效地将内存和计算需求减半,而不会对精度产生实质性影响。然而,INT 8量化加剧了模型偏差,显著降低了平衡精度。这些发现强调了关键的架构限制,并强调FP 16量化是一种最佳权衡,为实际部署和未来模型改进提供了指导方针。
摘要:Quantization is essential for deploying large audio language models (LALMs) efficiently in resource-constrained environments. However, its impact on complex tasks, such as zero-shot audio spoofing detection, remains underexplored. This study evaluates the zero-shot capabilities of five LALMs, GAMA, LTU-AS, MERaLiON, Qwen-Audio, and SALMONN, across three distinct datasets: ASVspoof2019, In-the-Wild, and WaveFake, and investigates their robustness to quantization (FP32, FP16, INT8). Despite high initial spoof detection accuracy, our analysis demonstrates severe predictive biases toward spoof classification across all models, rendering their practical performance equivalent to random classification. Interestingly, quantization to FP16 precision resulted in negligible performance degradation compared to FP32, effectively halving memory and computational requirements without materially impacting accuracy. However, INT8 quantization intensified model biases, significantly degrading balanced accuracy. These findings highlight critical architectural limitations and emphasize FP16 quantization as an optimal trade-off, providing guidelines for practical deployment and future model refinement.
【23】 A Fast and Lightweight Model for Causal Audio-Visual Speech Separation
链接:https://arxiv.org/abs/2506.06689
备注:8 pages, 5 figures
摘要:视听语音分离(AVSS)的目的是利用听觉和视觉(嘴唇运动)线索从混合信号中提取目标语音信号。然而,大多数现有的AVSS方法表现出复杂的体系结构,并依赖于未来的上下文,离线操作,这使得它们不适合实时应用。受RTFSNet流水线的启发,我们提出了一种新的流式AVSS模型,命名为Swift-Net,它增强了实时应用所需的因果处理能力。Swift-Net采用轻量级的视觉特征提取模块和高效的融合模块进行视听融合。此外,Swift-Net采用分组SRU来整合不同特征空间的历史信息,从而提高历史信息的利用效率。我们进一步提出了一个因果转换模板,以促进非因果AVSS模型转换成因果对应。在三个标准基准数据集(LRS 2,LRS 3和VoxCeleb 2)上的实验表明,在因果条件下,我们提出的Swift-Net表现出出色的性能,突出了这种方法在复杂环境中处理语音的潜力。
摘要:Audio-visual speech separation (AVSS) aims to extract a target speech signal from a mixed signal by leveraging both auditory and visual (lip movement) cues. However, most existing AVSS methods exhibit complex architectures and rely on future context, operating offline, which renders them unsuitable for real-time applications. Inspired by the pipeline of RTFSNet, we propose a novel streaming AVSS model, named Swift-Net, which enhances the causal processing capabilities required for real-time applications. Swift-Net adopts a lightweight visual feature extraction module and an efficient fusion module for audio-visual integration. Additionally, Swift-Net employs Grouped SRUs to integrate historical information across different feature spaces, thereby improving the utilization efficiency of historical information. We further propose a causal transformation template to facilitate the conversion of non-causal AVSS models into causal counterparts. Experiments on three standard benchmark datasets (LRS2, LRS3, and VoxCeleb2) demonstrated that under causal conditions, our proposed Swift-Net exhibited outstanding performance, highlighting the potential of this method for processing speech in complex environments.
【24】 CAtCh: Cognitive Assessment through Cookie Thief
链接:https://arxiv.org/abs/2506.06603
摘要:已经开发了几种机器学习算法用于从自发语音预测阿尔茨海默病和相关痴呆症(ADRD)。然而,这些算法都没有被翻译为预测更广泛的认知障碍(CI),这在某些情况下是ADRD的前兆和风险因素。在本文中,我们评估了几种最初提出用于预测ADRD的基于语音的开源方法,以及用于从患者录音中预测CI的任务的多模态情感分析方法。结果表明,多模态的方法优于单峰的CI预测,和基于声学的方法比基于语言学的表现更好。具体而言,可解释的声学特征有关的影响和韵律被发现显着优于基于BERT的语言特征和可解释的语言特征,分别。为这项研究开发的所有代码都可以在https://github.com/JTColonel/catch上获得。
摘要:Several machine learning algorithms have been developed for the prediction of Alzheimer's disease and related dementia (ADRD) from spontaneous speech. However, none of these algorithms have been translated for the prediction of broader cognitive impairment (CI), which in some cases is a precursor and risk factor of ADRD. In this paper, we evaluated several speech-based open-source methods originally proposed for the prediction of ADRD, as well as methods from multimodal sentiment analysis for the task of predicting CI from patient audio recordings. Results demonstrated that multimodal methods outperformed unimodal ones for CI prediction, and that acoustics-based approaches performed better than linguistics-based ones. Specifically, interpretable acoustic features relating to affect and prosody were found to significantly outperform BERT-based linguistic features and interpretable linguistic features, respectively. All the code developed for this study is available at https://github.com/JTColonel/catch.
【25】 Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models
链接:https://arxiv.org/abs/2506.06537
备注:Accepted on INTERSPEECH2025
摘要:音视频分割(Audiovisual segmentation,AVS)的目的是识别与声源对应的视觉区域,在视频理解、监控和人机交互等方面发挥着重要作用。传统的AVS方法依赖于大规模的像素级注释,这是昂贵和耗时的获得。为了解决这个问题,我们提出了一种新的zero-shot AVS框架,通过利用多个预训练模型来消除特定于任务的训练。我们的方法集成了音频,视觉和文本表示来弥合模态差距,实现精确的声源分割,而无需AVS特定的注释。我们系统地探索了连接预训练模型的不同策略,并评估了它们在多个数据集上的有效性。实验结果表明,我们的框架实现了最先进的zero-shot AVS性能,凸显了多模式模型集成对于细粒度视听分割的有效性。
摘要:Audiovisual segmentation (AVS) aims to identify visual regions corresponding to sound sources, playing a vital role in video understanding, surveillance, and human-computer interaction. Traditional AVS methods depend on large-scale pixel-level annotations, which are costly and time-consuming to obtain. To address this, we propose a novel zero-shot AVS framework that eliminates task-specific training by leveraging multiple pretrained models. Our approach integrates audio, vision, and text representations to bridge modality gaps, enabling precise sound source segmentation without AVS-specific annotations. We systematically explore different strategies for connecting pretrained models and evaluate their efficacy across multiple datasets. Experimental results demonstrate that our framework achieves state-of-the-art zero-shot AVS performance, highlighting the effectiveness of multimodal model integration for finegrained audiovisual segmentation.
【26】 TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment
链接:https://arxiv.org/abs/2506.06343
摘要:语音语言模型的最新进展在构建智能语音助手方面显示出了可喜的成果。然而,大多数现有的方法依赖于大规模配对的语音文本数据和广泛的计算资源,这在可扩展性和可访问性方面带来了挑战。在本文中,我们提出了\textbf{TESU-LLM},这是一个新的框架,可以只使用文本数据来训练支持语音的语言模型。我们的关键见解是利用一个统一的编码器,映射语义等效的文本和语音输入到一个共享的潜在空间。通过将编码器输出与LLM的嵌入空间通过轻量级投影网络对齐,我们使模型能够从纯文本监督推广到基于语音的推理。尽管TESU-LLM只接受文本训练,但它在各种语音相关基准测试中表现出色,与使用大规模多模态数据集和大量计算资源训练的基线方法相当。这些结果突出了我们的方法的有效性和效率,为在没有语音数据的情况下构建语音LLM提供了一条可扩展的路径。
摘要:Recent advances in speech-enabled language models have shown promising results in building intelligent voice assistants. However, most existing approaches rely on large-scale paired speech-text data and extensive computational resources, which pose challenges in terms of scalability and accessibility. In this paper, we present \textbf{TESU-LLM}, a novel framework that enables training speech-capable language models using only text data. Our key insight is to leverage a unified encoder that maps semantically equivalent text and speech inputs to a shared latent space. By aligning the encoder output with the embedding space of a LLM via a lightweight projection network, we enable the model to generalize from text-only supervision to speech-based inference. Despite being trained exclusively on text, TESU-LLM achieves strong performance on various speech-related benchmarks, comparable to baseline methods trained with large-scale multimodal datasets and substantial computational resources. These results highlight the effectiveness and efficiency of our approach, offering a scalable path toward building speech LLMs without speech data.
【27】 Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
链接:https://arxiv.org/abs/2506.07515
备注:Accepted at INTERSPEECH 2025
摘要:本文提出了一种新的框架,多说话人自动语音识别,而无需辅助信息。序列化输出训练(SOT)是一种广泛使用的方法,由于说话人分配失败而导致识别错误。虽然结合辅助信息,如令牌级时间戳,可以提高识别准确性,从自然的会话语音中提取这样的信息仍然具有挑战性。为了解决这一限制,我们提出了说话者可区分CTC(SD-CTC),CTC的扩展,联合分配一个令牌及其相应的扬声器标签到每个帧。我们进一步将SD-CTC集成到SOT框架中,使SOT模型能够仅使用重叠语音和transmittance来学习说话人区分。实验比较表明,使用SD-CTC和SOT的多任务学习将SOT模型的错误率降低了26%,并实现了与依赖辅助信息的最先进方法相当的性能。
摘要:This paper presents a novel framework for multi-talker automatic speech recognition without the need for auxiliary information. Serialized Output Training (SOT), a widely used approach, suffers from recognition errors due to speaker assignment failures. Although incorporating auxiliary information, such as token-level timestamps, can improve recognition accuracy, extracting such information from natural conversational speech remains challenging. To address this limitation, we propose Speaker-Distinguishable CTC (SD-CTC), an extension of CTC that jointly assigns a token and its corresponding speaker label to each frame. We further integrate SD-CTC into the SOT framework, enabling the SOT model to learn speaker distinction using only overlapping speech and transcriptions. Experimental comparisons show that multi-task learning with SD-CTC and SOT reduces the error rate of the SOT model by 26% and achieves performance comparable to state-of-the-art methods relying on auxiliary information.
【28】 Accurate analysis of the pitch pulse-based magnitude/phase structure of natural vowels and assessment of three lightweight time/frequency voicing restoration methods
链接:https://arxiv.org/abs/2506.06675
备注:58 pages, 17 figures, 8 tables
摘要:耳语是在声带没有被使用时产生的,无论是故意的,还是由于暂时或永久的声音条件。自然语音和耳语音之间的本质区别在于,由于声带振动而存在于前者的某些区域(称为浊音区域)中的周期性信号分量在后者中缺失。从耳语音恢复自然语音需要精细的信号处理过程,如果它们可以在低资源便携式设备上实时和即时地实现,则这些信号处理过程特别有用,从而利用语音产生和相关模型的已建立的源滤波器范例。本文讨论了两个相互交织的挑战,这两个挑战是告知和实现这一设想的技术实现的关键。第一个挑战涉及表征和建模的谐波相位/幅度结构的一个序列中的一个单独的音高周期的自然语音的浊音区,包括持续或共同发音的元音的演变。本文提出了一种新的算法分割个别音高脉冲,然后使用它来获得说明性的结果突出持续和共同发音元音之间的重要差异,并建议实用的合成发声方法。第二个挑战涉及基于模型的合成语音。描述了三种实施方案,它们的信号重建方法不同:频域,组合的频域和时域,以及生理启发单独产生的声门激励脉冲的单独滤波。这三种替代品进行了比较,客观上使用说明性的例子,主观上使用听力测试的结果,涉及合成语音的持续和共同发音的元音在词的上下文中。
摘要:Whispered speech is produced when the vocal folds are not used, either intentionally, or due to a temporary or permanent voice condition. The essential difference between natural speech and whispered speech is that periodic signal components that exist in certain regions of the former, called voiced regions, as a consequence of the vibration of the vocal folds, are missing in the latter. The restoration of natural speech from whispered speech requires delicate signal processing procedures that are especially useful if they can be implemented on low-resourced portable devices, in real-time, and on-the-fly, taking advantage of the established source-filter paradigm of voice production and related models. This paper addresses two challenges that are intertwined and are key in informing and making viable this envisioned technological realization. The first challenge involves characterizing and modeling the evolution of the harmonic phase/magnitude structure of a sequence of individual pitch periods in a voiced region of natural speech comprising sustained or co-articulated vowels. This paper proposes a novel algorithm segmenting individual pitch pulses, which is then used to obtain illustrative results highlighting important differences between sustained and co-articulated vowels, and suggesting practical synthetic voicing approaches. The second challenge involves model-based synthetic voicing. Three implementation alternatives are described that differ in their signal reconstruction approaches: frequency-domain, combined frequency and time-domain, and physiologically-inspired separate filtering of glottal excitation pulses individually generated. The three alternatives are compared objectively using illustrative examples, and subjectively using the results of listening tests involving synthetic voicing of sustained and co-articulated vowels in word context.
【1】 Unified Semi-Supervised Pipeline for Automatic Speech Recognition
链接:https://arxiv.org/abs/2506.07659
备注:Accepted to Interspeech 2025
摘要:自动语音识别一直是一个长期的研究领域,由于标记数据集的稀缺性,人们一直致力于集成半监督学习。然而,大多数先前的工作都集中在使用现有数据集改进学习算法,而没有为跨新数据集或语言的大规模半监督训练提供完整的公共框架。在这项工作中,我们介绍了一个完全开源的半监督训练框架,包括整个管道:从未标记的数据收集到伪标记和模型训练。我们的方法使可扩展的数据集创建任何语言使用公共可用的语音数据下知识共享许可证。我们还提出了一种新的伪标签算法,TopIPL,并评估它在低资源(葡萄牙语,亚美尼亚语)和高资源(西班牙语)设置。值得注意的是,TopIPL实现了葡萄牙语18-40%,亚美尼亚语5-16%和西班牙语2-8%的相对WER改善。
摘要:Automatic Speech Recognition has been a longstanding research area, with substantial efforts dedicated to integrating semi-supervised learning due to the scarcity of labeled datasets. However, most prior work has focused on improving learning algorithms using existing datasets, without providing a complete public framework for large-scale semi-supervised training across new datasets or languages. In this work, we introduce a fully open-source semi-supervised training framework encompassing the entire pipeline: from unlabeled data collection to pseudo-labeling and model training. Our approach enables scalable dataset creation for any language using publicly available speech data under Creative Commons licenses. We also propose a novel pseudo-labeling algorithm, TopIPL, and evaluate it in both low-resource (Portuguese, Armenian) and high-resource (Spanish) settings. Notably, TopIPL achieves relative WER improvements of 18-40% for Portuguese, 5-16% for Armenian, and 2-8% for Spanish.
【2】 SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement
链接:https://arxiv.org/abs/2506.07634
备注:Submitted to NeurIPS2025
摘要:创作结构连贯、乐器和声乐元素和谐的音乐仍然是歌曲创作中的一个重大挑战。现有的语言模型和基于扩散的方法通常难以平衡全局一致性和局部保真度,导致输出缺乏音乐性或遭受不连贯的进展和不匹配的歌词。本文介绍了$\textbf{SongBloom}$,全长歌曲生成,利用自回归草图和扩散为基础的细化交错范例的一种新的框架。SongBloom采用自回归扩散模型,将扩散模型的高保真度与语言模型的可扩展性相结合。具体来说,它是将音乐小品由短到长逐步延伸,细节由粗到细逐步细化。交错生成范式有效地整合了先前的语义和声学上下文来指导生成过程。实验结果表明,SongBloom在主观和客观指标上都优于现有方法,并达到了与最先进的商业音乐生成平台相当的性能。音频样本可以在我们的演示页面上找到:https://cypress-yang.github.io/SongBloom\_demo。
摘要:Generating music with coherent structure, harmonious instrumental and vocal elements remains a significant challenge in song generation. Existing language models and diffusion-based methods often struggle to balance global coherence with local fidelity, resulting in outputs that lack musicality or suffer from incoherent progression and mismatched lyrics. This paper introduces $\textbf{SongBloom}$, a novel framework for full-length song generation that leverages an interleaved paradigm of autoregressive sketching and diffusion-based refinement. SongBloom employs an autoregressive diffusion model that combines the high fidelity of diffusion models with the scalability of language models. Specifically, it gradually extends a musical sketch from short to long and refines the details from coarse to fine-grained. The interleaved generation paradigm effectively integrates prior semantic and acoustic context to guide the generation process. Experimental results demonstrate that SongBloom outperforms existing methods across both subjective and objective metrics and achieves performance comparable to the state-of-the-art commercial music generation platforms. Audio samples are available on our demo page: https://cypress-yang.github.io/SongBloom\_demo.
【3】 Bayesian Learning for Domain-Invariant Speaker Verification and Anti-Spoofing
链接:https://arxiv.org/abs/2506.07536
备注:Accepted to Interspeech2025
摘要:在真实世界中的域失配条件下,自动说话人确认(ASV)和反欺骗的性能会严重下降。放松的实例频率归一化(RFN),它规范化的频率分量的基础上沿时间和通道轴的特征统计,是一种很有前途的方法,以减少域依赖的特征映射的说话人嵌入网络。我们主张不同的频率应该得到不同的权重,并且应该考虑到由于域移动而导致的权重的不确定性。为此,我们建议利用变分推理来模拟权重的后验分布,从而得到贝叶斯加权RFN(BWRFN)。这种方法克服了固定权重RFN的局限性,使其在域失配条件下更有效。在跨数据集ASV、跨TTS反欺骗和欺骗鲁棒ASV上的大量实验表明,BWRFN明显优于WRFN和RFN。
摘要:The performance of automatic speaker verification (ASV) and anti-spoofing drops seriously under real-world domain mismatch conditions. The relaxed instance frequency-wise normalization (RFN), which normalizes the frequency components based on the feature statistics along the time and channel axes, is a promising approach to reducing the domain dependence in the feature maps of a speaker embedding network. We advocate that the different frequencies should receive different weights and that the weights' uncertainty due to domain shift should be accounted for. To these ends, we propose leveraging variational inference to model the posterior distribution of the weights, which results in Bayesian weighted RFN (BWRFN). This approach overcomes the limitations of fixed-weight RFN, making it more effective under domain mismatch conditions. Extensive experiments on cross-dataset ASV, cross-TTS anti-spoofing, and spoofing-robust ASV show that BWRFN is significantly better than WRFN and RFN.
【4】 Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
链接:https://arxiv.org/abs/2506.07515
备注:Accepted at INTERSPEECH 2025
摘要:本文提出了一种新的框架,多说话人自动语音识别,而无需辅助信息。序列化输出训练(SOT)是一种广泛使用的方法,由于说话人分配失败而导致识别错误。虽然结合辅助信息,如令牌级时间戳,可以提高识别准确性,从自然的会话语音中提取这样的信息仍然具有挑战性。为了解决这一限制,我们提出了说话者可区分CTC(SD-CTC),CTC的扩展,联合分配一个令牌及其相应的扬声器标签到每个帧。我们进一步将SD-CTC集成到SOT框架中,使SOT模型能够仅使用重叠语音和transmittance来学习说话人区分。实验比较表明,使用SD-CTC和SOT的多任务学习将SOT模型的错误率降低了26%,并实现了与依赖辅助信息的最先进方法相当的性能。
摘要:This paper presents a novel framework for multi-talker automatic speech recognition without the need for auxiliary information. Serialized Output Training (SOT), a widely used approach, suffers from recognition errors due to speaker assignment failures. Although incorporating auxiliary information, such as token-level timestamps, can improve recognition accuracy, extracting such information from natural conversational speech remains challenging. To address this limitation, we propose Speaker-Distinguishable CTC (SD-CTC), an extension of CTC that jointly assigns a token and its corresponding speaker label to each frame. We further integrate SD-CTC into the SOT framework, enabling the SOT model to learn speaker distinction using only overlapping speech and transcriptions. Experimental comparisons show that multi-task learning with SD-CTC and SOT reduces the error rate of the SOT model by 26% and achieves performance comparable to state-of-the-art methods relying on auxiliary information.
【5】 Multi-Distillation from Speech and Music Representation Models
链接:https://arxiv.org/abs/2506.07237
备注:8 pages, 1 figures
摘要:现实世界的音频通常混合语音和音乐,但模型通常只处理一个域。本文介绍了一个多教师蒸馏框架,将语音和音乐模型统一为一个模型,同时显着减少模型大小。我们的方法利用了特定领域的教师模型的优势,例如用于语音的HuBERT和用于音乐的MERT,并探索了平衡这两个领域的各种策略。在不同任务的实验表明,我们的模型匹配特定领域的模型的性能,跨域蒸馏的有效性。此外,我们进行了Few-Shot学习实验,强调了在标记数据有限的真实场景中对通用模型的需求。我们的研究结果表明,我们的模型不仅表现与专业模型,但也优于他们在Few-Shot的情况下,证明了跨域的方法是必要的和有效的不同的任务,有限的数据。
摘要:Real-world audio often mixes speech and music, yet models typically handle only one domain. This paper introduces a multi-teacher distillation framework that unifies speech and music models into a single one while significantly reducing model size. Our approach leverages the strengths of domain-specific teacher models, such as HuBERT for speech and MERT for music, and explores various strategies to balance both domains. Experiments across diverse tasks demonstrate that our model matches the performance of domain-specific models, showing the effectiveness of cross-domain distillation. Additionally, we conduct few-shot learning experiments, highlighting the need for general models in real-world scenarios where labeled data is limited. Our results show that our model not only performs on par with specialized models but also outperforms them in few-shot scenarios, proving that a cross-domain approach is essential and effective for diverse tasks with limited data.
【6】 Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding
链接:https://arxiv.org/abs/2506.07233
摘要:大型音频语言模型(LALM)可以将音频和文本作为输入,并回答有关音频的问题。虽然之前的LALM在标准基准测试中表现出了强大的性能,但有令人震惊的证据表明LALM可以使音频中呈现的内容产生幻觉。为了减轻LALM的幻觉,我们引入了音频感知解码(AAD),这是一种轻量级的推理时间策略,它使用对比解码来比较有和没有音频上下文的令牌预测logits。通过对比解码,AAD提升当音频存在时概率增加的令牌。我们在具有三个LALM的对象幻觉数据集上进行了实验,结果表明AAD将F1分数提高了0.046到0.428。我们还表明,AAD可以将Clotho-AQA等一般音频QA数据集的准确性提高5.4%至10.3%。我们进行全面的消融研究,以了解AAD中每个组件的有效性。
摘要:Large Audio-Language Models (LALMs) can take audio and text as the inputs and answer questions about the audio. While prior LALMs have shown strong performance on standard benchmarks, there has been alarming evidence that LALMs can hallucinate what is presented in the audio. To mitigate the hallucination of LALMs, we introduce Audio-Aware Decoding (AAD), a lightweight inference-time strategy that uses contrastive decoding to compare the token prediction logits with and without the audio context. By contrastive decoding, AAD promotes the tokens whose probability increases when the audio is present. We conduct our experiment on object hallucination datasets with three LALMs and show that AAD improves the F1 score by 0.046 to 0.428. We also show that AAD can improve the accuracy on general audio QA datasets like Clotho-AQA by 5.4% to 10.3%. We conduct thorough ablation studies to understand the effectiveness of each component in AAD.
【7】 Rhythm Features for Speaker Identification
链接:https://arxiv.org/abs/2506.06834
摘要:虽然深度学习模型在说话人识别任务中表现出了强大的性能,但它们主要依赖于从频谱图或原始波形中经验学习的低级音频特征。然而,先前的研究表明,特殊的说话风格严重影响语言单位的时间结构的语音信号(节奏)。这使得节奏成为一个强大但在很大程度上被忽视的候选语音身份特征。在本文中,我们通过应用深度学习方法从节奏特征执行与文本无关的说话人识别来测试这一假设。我们的研究结果支持说话人识别任务的节奏信息的有用性,但也表明,高主题内的可变性,在特设语音会降低其有效性。
摘要:While deep learning models have demonstrated robust performance in speaker recognition tasks, they primarily rely on low-level audio features learned empirically from spectrograms or raw waveforms. However, prior work has indicated that idiosyncratic speaking styles heavily influence the temporal structure of linguistic units in speech signals (rhythm). This makes rhythm a strong yet largely overlooked candidate for a speech identity feature. In this paper, we test this hypothesis by applying deep learning methods to perform text-independent speaker identification from rhythm features. Our findings support the usefulness of rhythmic information for speaker recognition tasks but also suggest that high intra-subject variability in ad-hoc speech can degrade its effectiveness.
【8】 Neural Spectral Band Generation for Audio Coding
链接:https://arxiv.org/abs/2506.06732
备注:Accepted to Interspeech 2025
摘要:音频带宽扩展是重构带宽受限的音频信号的丢失的高频分量的任务,其中由于包括信道容量和数据约束的若干原因,带宽限制是音频信号的常见问题。虽然传统的频谱带复制是音频带宽扩展的成熟参数方法,但是SBR通常需要粗糙的特征提取和重构技术,这导致在处理各种类型的音频信号时的限制。同时,已经提出了许多基于深度神经网络的音频带宽扩展方法。这些基于DNN的方法通常被称为盲BWE,因为这些方法不依赖于从原始信号提取的先验信息,并且仅利用给定的低频带信号来估计丢失的高频分量。为了用DNN取代传统的SBR,简单地采用现有的基于DNN的方法会由于这些方法的盲目性而导致性能不佳。我提出的研究提出了一种新的参数化非盲带宽扩展方法,因为基于DNN的边信息提取和基于DNN的带宽扩展仅在音频编码管道的前端和末端执行。
摘要:Audio bandwidth extension is the task of reconstructing missing high frequency components of bandwidth-limited audio signals, where bandwidth limitation is a common issue for audio signals due to several reasons, including channel capacity and data constraints. While conventional spectral band replication is a well-established parametric approach to audio bandwidth extension, the SBR usually entails coarse feature extraction and reconstruction techniques, which leads to limitations when processing various types of audio signals. In parallel, numerous deep neural network-based audio bandwidth extension methods have been proposed. These DNN-based methods are usually referred to as blind BWE, as these methods do not rely on prior information extracted from original signals, and only utilize given low frequency band signals to estimate missing high frequency components. In order to replace conventional SBR with DNNs, simply adopting existing DNN-based methodologies results in suboptimal performance due to the blindness of these methods. My proposed research suggests a new approach to parametric non-blind bandwidth extension, as DNN-based side information extraction and DNN-based bandwidth extension are performed only at the front and end of the audio coding pipeline.
【9】 Exploring Length Generalization For Transformer-based Speech Enhancement
链接:https://arxiv.org/abs/2506.06697
备注:14 pages; Accepted by TASLP
摘要:事实证明,Transformer网络架构在语音增强方面是有效的。然而,作为其核心模块,自我注意具有二次复杂度,这使得它不适用于长语音话语的训练。在实际场景中,语音增强模型通常需要在运行时间上对噪声语音执行,该运行时间基本上长于训练话语。基于transformer的语音增强模型如何推广到长语音仍然是一个挑战。在本文中,进行了广泛的实证研究,探讨模型的长度泛化能力。特别是,我们对四个训练目标进行语音增强实验,并用五个指标进行评估。我们的研究表明,位置编码是一种有效的工具,以抑制话语长度对语音增强的影响。我们首先探讨了几种现有的位置编码方法,结果表明,相对位置编码方法表现出更好的长度泛化性能比绝对位置编码方法。此外,我们还探索了一种更简单、更有效的位置编码方案,即LearnLin,该方案仅为每个注意力头部使用一个可训练参数来缩放时间帧之间的真实相对位置,从而学习对短期或长期的不同偏好。这些头部的依赖性。结果表明,我们的建议具有良好的长度泛化能力与其他国家的最先进的位置编码策略相比,具有相当或更高的性能。
摘要:Transformer network architecture has proven effective in speech enhancement. However, as its core module, self-attention suffers from quadratic complexity, making it infeasible for training on long speech utterances. In practical scenarios, speech enhancement models are often required to perform on noisy speech at run-time that is substantially longer than the training utterances. It remains a challenge how a Transformer-based speech enhancement model can generalize to long speech utterances. In this paper, extensive empirical studies are conducted to explore the model's length generalization ability. In particular, we conduct speech enhancement experiments on four training objectives and evaluate with five metrics. Our studies establish that positional encoding is an effective instrument to dampen the effect of utterance length on speech enhancement. We first explore several existing positional encoding methods, and the results show that relative positional encoding methods exhibit a better length generalization property than absolute positional encoding methods. Additionally, we also explore a simpler and more effective positional encoding scheme, i.e. LearnLin, that uses only one trainable parameter for each attention head to scale the real relative position between time frames, which learns the different preferences on short- or long-term dependencies of these heads. The results demonstrate that our proposal exhibits excellent length generalization ability with comparable or superior performance than other state-of-the-art positional encoding strategies.
【10】 Accurate analysis of the pitch pulse-based magnitude/phase structure of natural vowels and assessment of three lightweight time/frequency voicing restoration methods
链接:https://arxiv.org/abs/2506.06675
备注:58 pages, 17 figures, 8 tables
摘要:耳语是在声带没有被使用时产生的,无论是故意的,还是由于暂时或永久的声音条件。自然语音和耳语音之间的本质区别在于,由于声带振动而存在于前者的某些区域(称为浊音区域)中的周期性信号分量在后者中缺失。从耳语音恢复自然语音需要精细的信号处理过程,如果它们可以在低资源便携式设备上实时和即时地实现,则这些信号处理过程特别有用,从而利用语音产生和相关模型的已建立的源滤波器范例。本文讨论了两个相互交织的挑战,这两个挑战是告知和实现这一设想的技术实现的关键。第一个挑战涉及表征和建模的谐波相位/幅度结构的一个序列中的一个单独的音高周期的自然语音的浊音区,包括持续或共同发音的元音的演变。本文提出了一种新的算法分割个别音高脉冲,然后使用它来获得说明性的结果突出持续和共同发音元音之间的重要差异,并建议实用的合成发声方法。第二个挑战涉及基于模型的合成语音。描述了三种实施方案,其信号重建方法不同:频域、频域和时域组合以及对单独生成的声门激励脉冲进行生理启发的单独滤波。这三种替代品进行了比较,客观上使用说明性的例子,主观上使用听力测试的结果,涉及合成语音的持续和共同发音的元音在词的上下文中。
摘要:Whispered speech is produced when the vocal folds are not used, either intentionally, or due to a temporary or permanent voice condition. The essential difference between natural speech and whispered speech is that periodic signal components that exist in certain regions of the former, called voiced regions, as a consequence of the vibration of the vocal folds, are missing in the latter. The restoration of natural speech from whispered speech requires delicate signal processing procedures that are especially useful if they can be implemented on low-resourced portable devices, in real-time, and on-the-fly, taking advantage of the established source-filter paradigm of voice production and related models. This paper addresses two challenges that are intertwined and are key in informing and making viable this envisioned technological realization. The first challenge involves characterizing and modeling the evolution of the harmonic phase/magnitude structure of a sequence of individual pitch periods in a voiced region of natural speech comprising sustained or co-articulated vowels. This paper proposes a novel algorithm segmenting individual pitch pulses, which is then used to obtain illustrative results highlighting important differences between sustained and co-articulated vowels, and suggesting practical synthetic voicing approaches. The second challenge involves model-based synthetic voicing. Three implementation alternatives are described that differ in their signal reconstruction approaches: frequency-domain, combined frequency and time-domain, and physiologically-inspired separate filtering of glottal excitation pulses individually generated. The three alternatives are compared objectively using illustrative examples, and subjectively using the results of listening tests involving synthetic voicing of sustained and co-articulated vowels in word context.
【11】 AS-ASR: A Lightweight Framework for Aphasia-Specific Automatic Speech Recognition
链接:https://arxiv.org/abs/2506.06566
备注:Under review
摘要:本文提出了AS-ASR,这是一种基于Whisper-tiny的轻量级失语症特定语音识别框架,专为边缘设备上的低资源部署而量身定制。我们的方法引入了一种混合训练策略,该策略以不同的比例系统地结合了标准和失语症语音,实现了强大的泛化,以及一种基于GPT-4的参考增强方法,该方法可以改进嘈杂的失语症成绩单,提高监督质量。我们在多个数据混合配置和评估设置中进行了广泛的实验。结果表明,我们的微调模型显着优于zero-shot基线,减少WER失语症的语音超过30%,同时保持对标准语音的性能。该框架为现实世界的无序语音识别提供了一个可扩展的,高效的解决方案。
摘要:This paper proposes AS-ASR, a lightweight aphasia-specific speech recognition framework based on Whisper-tiny, tailored for low-resource deployment on edge devices. Our approach introduces a hybrid training strategy that systematically combines standard and aphasic speech at varying ratios, enabling robust generalization, and a GPT-4-based reference enhancement method that refines noisy aphasic transcripts, improving supervision quality. We conduct extensive experiments across multiple data mixing configurations and evaluation settings. Results show that our fine-tuned model significantly outperforms the zero-shot baseline, reducing WER on aphasic speech by over 30% while preserving performance on standard speech. The proposed framework offers a scalable, efficient solution for real-world disordered speech recognition.
【12】 W4S4: WaLRUS Meets S4 for Long-Range Sequence Modeling
链接:https://arxiv.org/abs/2506.07920
备注:10 pages, 2 figures, 3 tables
摘要:状态空间模型(SSM)已经成为序列建模的强大组件,通过线性递归和卷积计算有效地处理长期依赖关系。然而,它们的有效性在很大程度上取决于状态矩阵的选择和初始化。在这项工作中,我们建立在SaFARI框架和现有的WaLRUS SSM引入一个新的变种,W4S4(WaLRUS S4),一类新的SSM构造的冗余小波框架。WaLRUS允许稳定的对角化,并支持快速的内核计算,而不需要低秩近似,使其在理论上和计算效率。我们表明,WaLRUS在长期范围内保留的信息明显优于基于HiPPO的SSM,无论是在隔离和集成到深层架构,如S4。我们的实验表明,在延迟重建任务,分类基准和长距离序列建模的一致改进,确认高质量,结构化的初始化,使基于小波的状态动态提供了显着的优势,现有的替代品。WaLRUS为下一代基于SSM的深度模型提供了可扩展的通用基础。
摘要:State Space Models (SSMs) have emerged as powerful components for sequence modeling, enabling efficient handling of long-range dependencies via linear recurrence and convolutional computation. However, their effectiveness depends heavily on the choice and initialization of the state matrix. In this work, we build on the SaFARi framework and existing WaLRUS SSMs to introduce a new variant, W4S4 (WaLRUS for S4), a new class of SSMs constructed from redundant wavelet frames. WaLRUS admits a stable diagonalization and supports fast kernel computation without requiring low-rank approximations, making it both theoretically grounded and computationally efficient. We show that WaLRUS retains information over long horizons significantly better than HiPPO-based SSMs, both in isolation and when integrated into deep architectures such as S4. Our experiments demonstrate consistent improvements across delay reconstruction tasks, classification benchmarks, and long-range sequence modeling, confirming that high-quality, structured initialization enabled by wavelet-based state dynamic offers substantial advantages over existing alternatives. WaLRUS provides a scalable and versatile foundation for the next generation of deep SSM-based models.
【13】 Towards a Unified Benchmark for Arabic Pronunciation Assessment: Quranic Recitation as Case Study
链接:https://arxiv.org/abs/2506.07722
备注:Accepted Interspeech 2025 and ArabicNLP Shared Task 2025
摘要:我们提出了一个统一的基准,在现代标准阿拉伯语(MSA)使用古兰经背诵作为案例研究的发音错误检测。我们的方法为推进阿拉伯语发音评估奠定了基础,提供了一个全面的管道,涵盖了数据处理,针对MSA发音的细微差别开发了一个专门的音素集,并为此任务创建了第一个公开可用的测试集,我们称之为古兰经错误发音基准(QuranMB.v1)。此外,我们评估了几个基线模型,以提供初步的性能见解,从而突出的承诺和固有的挑战,在评估MSA发音。通过建立这个标准化的框架,我们的目标是促进进一步的研究和发展发音评估在阿拉伯语技术和相关应用。
摘要:We present a unified benchmark for mispronunciation detection in Modern Standard Arabic (MSA) using Qur'anic recitation as a case study. Our approach lays the groundwork for advancing Arabic pronunciation assessment by providing a comprehensive pipeline that spans data processing, the development of a specialized phoneme set tailored to the nuances of MSA pronunciation, and the creation of the first publicly available test set for this task, which we term as the Qur'anic Mispronunciation Benchmark (QuranMB.v1). Furthermore, we evaluate several baseline models to provide initial performance insights, thereby highlighting both the promise and the challenges inherent in assessing MSA pronunciation. By establishing this standardized framework, we aim to foster further research and development in pronunciation assessment in Arabic language technology and related applications.
【14】 Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
链接:https://arxiv.org/abs/2506.07646
备注:Accepted to INTERSPEECH 2025
摘要:在本文中,我们提出了一种方法来注释音素和韵律标签上一个给定的音频转录对,旨在构建日本的文本到语音(TTS)数据集。我们的方法涉及微调大规模预训练的自动语音识别(ASR)模型,以地面实况成绩单为条件,同时输出短语级字素和注释标签。为了进一步纠正错误的音素标记,我们采用解码策略,利用字典先验知识。客观的评价结果表明,我们提出的方法优于以往的方法,只依赖于文本或音频。主观评价结果表明,语音合成的TTS模型,使用我们的方法标注的标签训练的语音的自然度,是与人工标注训练的模型。
摘要:In this paper, we propose a method for annotating phonemic and prosodic labels on a given audio-transcript pair, aimed at constructing Japanese text-to-speech (TTS) datasets. Our approach involves fine-tuning a large-scale pre-trained automatic speech recognition (ASR) model, conditioned on ground truth transcripts, to simultaneously output phrase-level graphemes and annotation labels. To further correct errors in phonemic labeling, we employ a decoding strategy that utilizes dictionary prior knowledge. The objective evaluation results demonstrate that our proposed method outperforms previous approaches relying solely on text or audio. The subjective evaluation results indicate that the naturalness of speech synthesized by the TTS model, trained with labels annotated using our method, is comparable to that of a model trained with manual annotations.
【15】 Generative Voice Bursts during Phone Call
链接:https://arxiv.org/abs/2506.07526
备注:12 pages, 2 figures
摘要:在紧急情况下,传统的移动电话无法将紧急语音消息传送到已经参与另一呼叫的被叫方。标准呼叫等待警报不提供等待呼叫的紧急程度或内容。本文提出了一种新的方法,用于传输生成语音突发短,上下文感知的音频消息在正在进行的呼叫,从预先授权或动态优先级的呼叫者。通过利用生成AI技术,当呼叫者由于丧失能力或环境限制而无法说话时,系统自动从上下文输入例如位置,健康数据,图像,背景噪音生成语音消息。该解决方案集成了语音、文本和优先级推理机制,允许高优先级紧急消息绕过传统的呼叫等待障碍。该方法采用GPT Neo等模型来生成文本,这些文本被合成为音频,并以可配置的间隔G秒发送,并计数N次,从而确保最小的中断,同时保持紧迫性。该方法具有跨电信、移动终端制造和应急通信平台的重大影响的潜力。
摘要:In critical situations, conventional mobile telephony fails to convey emergency voice messages to a callee already engaged in another call. The standard call waiting alert does not provide the urgency or content of the waiting call. This paper proposes a novel method for transmitting Generative Voice Bursts short, context aware audio messages during ongoing calls, from either preauthorized or dynamically prioritized callers. By leveraging generative AI techniques, the system automatically generates spoken messages from contextual inputs example like location, health data, images, background noise when the caller is unable to speak due to incapacitation or environmental constraints. The solution incorporates voice, text, and priority inference mechanisms, allowing high priority emergency messages to bypass conventional call waiting barriers. The approach employs models such as GPT Neo for generative text, which is synthesized into audio and delivered in configurable intervals G seconds and counts N times, ensuring minimal disruption while preserving urgency. This method holds potential for significant impact across telecom, mobile device manufacturing, and emergency communication platforms.
【16】 LeVo: High-Quality Song Generation with Multi-Preference Alignment
链接:https://arxiv.org/abs/2506.07520
摘要:大型语言模型(LLM)和音频语言模型的最新进展显着改善了音乐生成,特别是在歌词到歌曲的生成。然而,现有的方法仍然与歌曲的复杂组成和高质量数据的稀缺作斗争,导致音质,音乐性,指令遵循和声乐乐器和谐方面的限制。为了应对这些挑战,我们引入了LeVo,一个基于LM的框架,由LeLM和音乐编解码器组成。LeLM能够对两种类型的令牌进行并行建模:混合令牌,代表人声和伴奏的组合音频,以实现人声-乐器和谐,以及双轨令牌,分别对人声和伴奏进行编码,以生成高质量的歌曲。它采用了两个解码器专用的Transformers和一个模块化的扩展训练策略,以防止不同令牌类型之间的干扰。为了进一步提高音乐性和指令遵循,我们引入了一个多偏好对齐方法的基础上直接偏好优化(DPO)。这种方法通过半自动数据构建过程和DPO后训练来处理不同的人类偏好。实验结果表明,LeVo一贯优于现有的方法在客观和主观指标。消融研究进一步证明了我们设计的有效性。音频示例可在https://levo-demo.github.io/上获得。
摘要:Recent advances in large language models (LLMs) and audio language models have significantly improved music generation, particularly in lyrics-to-song generation. However, existing approaches still struggle with the complex composition of songs and the scarcity of high-quality data, leading to limitations in sound quality, musicality, instruction following, and vocal-instrument harmony. To address these challenges, we introduce LeVo, an LM-based framework consisting of LeLM and a music codec. LeLM is capable of parallelly modeling two types of tokens: mixed tokens, which represent the combined audio of vocals and accompaniment to achieve vocal-instrument harmony, and dual-track tokens, which separately encode vocals and accompaniment for high-quality song generation. It employs two decoder-only transformers and a modular extension training strategy to prevent interference between different token types. To further enhance musicality and instruction following, we introduce a multi-preference alignment method based on Direct Preference Optimization (DPO). This method handles diverse human preferences through a semi-automatic data construction process and DPO post-training. Experimental results demonstrate that LeVo consistently outperforms existing methods on both objective and subjective metrics. Ablation studies further justify the effectiveness of our designs. Audio examples are available at https://levo-demo.github.io/.
【17】 Towards Energy-Efficient and Low-Latency Voice-Controlled Smart Homes: A Proposal for Offline Speech Recognition and IoT Integration
链接:https://arxiv.org/abs/2506.07494
摘要:智能家居系统基于人工智能语音识别和物联网技术,使人们能够通过口头命令控制设备,使人们的生活更加高效。然而,现有的人工智能语音识别服务主要部署在互联网上的云平台上。当用户发出命令时,像“Amazon Echo”这样的语音识别设备将通过许多网络节点发布录音,到达多个服务器,然后通过互联网接收响应。这种机制存在几个问题,包括不必要的能源消耗、通信延迟和单点故障的风险。在这份立场文件中,我们提出了一个基于离线语音识别和物联网技术的智能家居概念:1)将离线关键字识别(KWS)技术集成到具有有限资源硬件的家用电器中,使它们能够理解用户的语音命令; 2)设计一个具有分散架构的本地物联网网络来管理和连接各种设备,增强系统的鲁棒性和可扩展性。这一基于离线语音识别和物联网技术的智能家居提案将允许用户在家中任何地方使用低延迟语音控制,而无需依赖互联网,并提供更好的可扩展性和能源可持续性。
摘要:The smart home systems, based on AI speech recognition and IoT technology, enable people to control devices through verbal commands and make people's lives more efficient. However, existing AI speech recognition services are primarily deployed on cloud platforms on the Internet. When users issue a command, speech recognition devices like ``Amazon Echo'' will post a recording through numerous network nodes, reach multiple servers, and then receive responses through the Internet. This mechanism presents several issues, including unnecessary energy consumption, communication latency, and the risk of a single-point failure. In this position paper, we propose a smart home concept based on offline speech recognition and IoT technology: 1) integrating offline keyword spotting (KWS) technologies into household appliances with limited resource hardware to enable them to understand user voice commands; 2) designing a local IoT network with decentralized architecture to manage and connect various devices, enhancing the robustness and scalability of the system. This proposal of a smart home based on offline speech recognition and IoT technology will allow users to use low-latency voice control anywhere in the home without depending on the Internet and provide better scalability and energy sustainability.
【18】 An introduction to pitch strength in contemporary popular music analysis and production
链接:https://arxiv.org/abs/2506.07473
备注:In Music 2024, Innovation in Music Conference, 14-16 June, 2024, Kristiania University College, Oslo, Norway
摘要:音乐信息检索区分音乐的低级和高级描述。目前的生成式AI模型依赖于比录音室音乐家熟悉的控制更高级别的文本描述。音高强度是当代流行音乐的一个低级感知参数,可能是使这种人工智能模型更适合音乐制作的一个特征。信号和感知分析表明,音高强度(1)在歌曲之间和歌曲内部变化显著;(2)有助于小尺度和大尺度结构;(3)有助于处理复调不和谐;(4)可能是从感知丰富性的角度来看可听到的高次谐波的特征。
摘要:Music information retrieval distinguishes between low- and high-level descriptions of music. Current generative AI models rely on text descriptions that are higher level than the controls familiar to studio musicians. Pitch strength, a low-level perceptual parameter of contemporary popular music, may be one feature that could make such AI models more suited to music production. Signal and perceptual analyses suggest that pitch strength (1) varies significantly across and inside songs; (2) contributes to both small- and large-scale structure; (3) contributes to the handling of polyphonic dissonance; and (4) may be a feature of upper harmonics made audible in a perspective of perceptual richness.
【19】 Lightweight Joint Audio-Visual Deepfake Detection via Single-Stream Multi-Modal Learning Framework
链接:https://arxiv.org/abs/2506.07358
摘要:Deepfakes是人工智能合成的多媒体数据,可能会被滥用来传播错误信息。Deepfake生成涉及视觉和音频操作。为了检测视听深度伪造,以前的研究通常使用两个相对独立的子模型来分别学习音频和视觉特征,并随后将其融合用于深度伪造检测。然而,这可能未充分利用音频和视觉特征之间的固有相关性。此外,利用两个孤立的特征学习子模型可能会导致冗余的神经层,使得整个模型对于资源受限的环境来说效率低下且不切实际。 在这项工作中,我们通过单流多模态学习框架设计了一个用于音频视频深度伪造检测的轻量级网络。具体来说,我们引入了协作视听学习模块,以在学习视觉和音频功能的同时有效地集成多模式信息。通过迭代使用这个块,我们的单流网络实现了跨层多模态特征的连续融合。因此,我们的网络可以有效地捕获视觉和音频功能,而无需过多的块堆叠,从而实现轻量级的网络设计。此外,我们提出了一个多模态分类模块,可以提高依赖的视觉和音频分类器的模态内容。它还增强了视频分类器对音频和视觉模态之间的失配的整体抵抗力。我们在DF-TIMIT,FakeAVCeleb和DFDC基准数据集上进行实验。与最先进的视听联合检测方法相比,我们的方法非常轻量级,只有0.48M参数,但它在单模态和多模态deepfake以及看不见的deepfake类型中都具有优势。
摘要:Deepfakes are AI-synthesized multimedia data that may be abused for spreading misinformation. Deepfake generation involves both visual and audio manipulation. To detect audio-visual deepfakes, previous studies commonly employ two relatively independent sub-models to learn audio and visual features, respectively, and fuse them subsequently for deepfake detection. However, this may underutilize the inherent correlations between audio and visual features. Moreover, utilizing two isolated feature learning sub-models can result in redundant neural layers, making the overall model inefficient and impractical for resource-constrained environments. In this work, we design a lightweight network for audio-visual deepfake detection via a single-stream multi-modal learning framework. Specifically, we introduce a collaborative audio-visual learning block to efficiently integrate multi-modal information while learning the visual and audio features. By iteratively employing this block, our single-stream network achieves a continuous fusion of multi-modal features across its layers. Thus, our network efficiently captures visual and audio features without the need for excessive block stacking, resulting in a lightweight network design. Furthermore, we propose a multi-modal classification module that can boost the dependence of the visual and audio classifiers on modality content. It also enhances the whole resistance of the video classifier against the mismatches between audio and visual modalities. We conduct experiments on the DF-TIMIT, FakeAVCeleb, and DFDC benchmark datasets. Compared to state-of-the-art audio-visual joint detection methods, our method is significantly lightweight with only 0.48M parameters, yet it achieves superiority in both uni-modal and multi-modal deepfakes, as well as in unseen types of deepfakes.
【20】 Speech Recognition on TV Series with Video-guided Post-Correction
链接:https://arxiv.org/abs/2506.07323
摘要:自动语音识别(ASR)在深度学习方面取得了巨大成功,推动了会话人工智能、媒体转录和辅助技术的进步。然而,ASR系统仍然在复杂的环境中挣扎,例如电视连续剧,其中重叠的语音,特定领域的术语和长距离上下文依赖性对转录准确性构成了重大挑战。现有的多模态方法无法利用视频中丰富的时间和上下文信息来校正ASR输出。为了解决这个问题,我们提出了一种新的多模态后校正框架,利用从视频中提取的上下文线索来细化ASR transmittance。我们的框架包括两个阶段:ASR生成和基于视频的后期校正,其中第一阶段产生初始成绩单,第二阶段使用基于视频的上下文信息提取和上下文感知ASR校正来校正错误。我们采用视频大多模态模型(VLMM)提取关键的上下文信息,使用量身定制的提示,然后将其与大语言模型(LLM)集成,以完善ASR输出。我们评估我们的方法在多模态基准的电视剧ASR,并证明其有效性,提高ASR的性能,利用基于视频的上下文,以提高转录的准确性,在复杂的多媒体环境。
摘要:Automatic Speech Recognition (ASR) has achieved remarkable success with deep learning, driving advancements in conversational artificial intelligence, media transcription, and assistive technologies. However, ASR systems still struggle in complex environments such as TV series, where overlapping speech, domain-specific terminology, and long-range contextual dependencies pose significant challenges to transcription accuracy. Existing multimodal approaches fail to correct ASR outputs with the rich temporal and contextual information available in video. To address this limitation, we propose a novel multimodal post-correction framework that refines ASR transcriptions by leveraging contextual cues extracted from video. Our framework consists of two stages: ASR Generation and Video-based Post-Correction, where the first stage produces the initial transcript and the second stage corrects errors using Video-based Contextual Information Extraction and Context-aware ASR Correction. We employ the Video-Large Multimodal Model (VLMM) to extract key contextual information using tailored prompts, which is then integrated with a Large Language Model (LLM) to refine the ASR output. We evaluate our method on a multimodal benchmark for TV series ASR and demonstrate its effectiveness in improving ASR performance by leveraging video-based context to enhance transcription accuracy in complex multimedia environments.
【21】 Towards Generalized Source Tracing for Codec-Based Deepfake Speech
链接:https://arxiv.org/abs/2506.07294
备注:Submitted to IEEE ASRU 2025
摘要:最近尝试对基于编解码器的深度伪造语音(CodecFake)进行源跟踪,由基于神经音频编解码器的语音生成(CoSG)模型生成,表现出次优性能。然而,如何使用模拟的CoSG数据训练源跟踪模型,同时保持对真实CoSG生成的音频的强大性能仍然是一个开放的挑战。在本文中,我们表明,仅在编解码器重新合成的数据上训练的模型往往过拟合非语音区域,并且难以推广到看不见的内容。为了缓解这些挑战,我们引入了语义声源跟踪网络(SASTNet),它联合利用Whisper进行语义特征编码,并利用Wav2vec2与AudioMAE进行声学特征编码。我们提出的SASTNet在CodecFake+数据集的CoSG测试集上实现了最先进的性能,证明了其可靠源跟踪的有效性。
摘要:Recent attempts at source tracing for codec-based deepfake speech (CodecFake), generated by neural audio codec-based speech generation (CoSG) models, have exhibited suboptimal performance. However, how to train source tracing models using simulated CoSG data while maintaining strong performance on real CoSG-generated audio remains an open challenge. In this paper, we show that models trained solely on codec-resynthesized data tend to overfit to non-speech regions and struggle to generalize to unseen content. To mitigate these challenges, we introduce the Semantic-Acoustic Source Tracing Network (SASTNet), which jointly leverages Whisper for semantic feature encoding and Wav2vec2 with AudioMAE for acoustic feature encoding. Our proposed SASTNet achieves state-of-the-art performance on the CoSG test set of the CodecFake+ dataset, demonstrating its effectiveness for reliable source tracing.
【22】 Methods for pitch analysis in contemporary popular music: Vitalic's use of tones that do not operate on the principle of acoustic resonance
链接:https://arxiv.org/abs/2506.07207
摘要:Vitalic是一位电子音乐制作人,自2001年以来一直活跃。Vitalic的2005年曲目“No Fun”的主要合成器部分是由一系列单一的不和谐音调组成的,可以同时唤起两种旋律。这一部分作为一个起点,检查维塔利奇的使用音调,不运作的原则,声学共振。这项研究认为,引起两个或两个以上的同时音高的音调,并检查各种不和谐的部分布局。维塔利克音乐之外的例子也表明,类似的音调属性可以在当代流行音乐的其他地方找到。
摘要:Vitalic is an electronic music producer who has been active since 2001. Vitalic's 2005 track "No Fun" features a main synthesiser part built from a sequence of single inharmonic tones that evoke two simultaneous melodies. This part serves as a starting point for examining Vitalic's use of tones that do not operate on the principle of acoustic resonance. The study considers tones that evoke two or more simultaneous pitches and examines various inharmonic partial layouts. Examples outside Vitalic's music are also provided to suggest that similar tone properties can be found elsewhere in contemporary popular music.
【23】 Audio synthesizer inversion in symmetric parameter spaces with approximately equivariant flow matching
链接:https://arxiv.org/abs/2506.07199
备注:Accepted at ISMIR 2025
摘要:许多音频合成器可以在给定不同参数配置的情况下产生相同的信号,这意味着从声音到参数的反演是一个固有的不适定问题。我们表明,这在很大程度上是由于固有的对称性的合成器,并特别关注置换不变性。首先,我们证明了在一个合成任务,回归点估计置换对称性降低性能,即使使用置换不变的损失函数或破环算法。然后,查看等价的解决方案的概率分布的模式,我们表明,条件生成模型大大提高了性能。此外,承认隐式参数分布的不变性,我们发现,通过使用置换等变连续归一化流,性能得到进一步提高。为了适应复杂的对称性在真正的合成器,我们还提出了一个宽松的等方差策略,自适应地发现相关的对称性数据。将我们的方法应用于Surge XT,一种用于现实世界音频制作的全功能开源合成器,我们发现我们的方法在音频重建指标上优于回归和生成基线。
摘要:Many audio synthesizers can produce the same signal given different parameter configurations, meaning the inversion from sound to parameters is an inherently ill-posed problem. We show that this is largely due to intrinsic symmetries of the synthesizer, and focus in particular on permutation invariance. First, we demonstrate on a synthetic task that regressing point estimates under permutation symmetry degrades performance, even when using a permutation-invariant loss function or symmetry-breaking heuristics. Then, viewing equivalent solutions as modes of a probability distribution, we show that a conditional generative model substantially improves performance. Further, acknowledging the invariance of the implicit parameter distribution, we find that performance is further improved by using a permutation equivariant continuous normalizing flow. To accommodate intricate symmetries in real synthesizers, we also propose a relaxed equivariance strategy that adaptively discovers relevant symmetries from data. Applying our method to Surge XT, a full-featured open source synthesizer used in real world audio production, we find our method outperforms regression and generative baselines across audio reconstruction metrics.
【24】 Technical Report: A Practical Guide to Kaldi ASR Optimization
链接:https://arxiv.org/abs/2506.07149
摘要:本技术报告介绍了基于Kaldi的自动语音识别(ASR)系统的创新优化,重点关注声学模型增强,超参数调整和语言模型效率。我们开发了一个自定义的Conformer块,集成了多流TDNN-F结构,实现了卓越的特征提取和时间建模。我们的方法包括先进的数据增强技术和动态超参数优化,以提高性能并减少过拟合。此外,我们提出了强大的语言模型管理策略,采用贝叶斯优化和$n$-gram修剪,以确保相关性和计算效率。这些系统性改进显著提高了ASR的准确性和鲁棒性,优于现有方法,并为各种语音识别场景提供了可扩展的解决方案。本报告强调了战略优化在保持Kaldi在快速发展的技术环境中的适应性和竞争力方面的重要性。
摘要:This technical report introduces innovative optimizations for Kaldi-based Automatic Speech Recognition (ASR) systems, focusing on acoustic model enhancement, hyperparameter tuning, and language model efficiency. We developed a custom Conformer block integrated with a multistream TDNN-F structure, enabling superior feature extraction and temporal modeling. Our approach includes advanced data augmentation techniques and dynamic hyperparameter optimization to boost performance and reduce overfitting. Additionally, we propose robust strategies for language model management, employing Bayesian optimization and $n$-gram pruning to ensure relevance and computational efficiency. These systematic improvements significantly elevate ASR accuracy and robustness, outperforming existing methods and offering a scalable solution for diverse speech recognition scenarios. This report underscores the importance of strategic optimizations in maintaining Kaldi's adaptability and competitiveness in rapidly evolving technological landscapes.
【25】 RBA-FE: A Robust Brain-Inspired Audio Feature Extractor for Depression Diagnosis
链接:https://arxiv.org/abs/2506.07118
备注:14 pages
摘要:本文提出了一个强大的脑启发音频特征提取器(RBA-FE)模型的抑郁症诊断,使用改进的分层网络架构。大多数深度学习模型在基于图像的诊断任务中实现了最先进的性能,忽略了对应的音频特征。为了应对噪声挑战,RBA-FE利用从原始音频中提取的六个声学特征,捕获空间特征和时间依赖性。这种混合属性有助于缓解其他学习模型(如深度残差收缩网络)中音频特征提取的精度限制。为了处理噪声问题,我们的模型采用了一种改进的尖峰神经元模型,称为自适应速率平滑泄漏积分和发射(ARSLIF)。ARSLIF模型模拟了大脑注意系统中“细胞信号选择性的重新调谐”机制,增强了模型对音频数据中环境噪声的鲁棒性。实验结果表明,RBA-FE在MODMA数据集上达到了最先进的准确率,精确率,准确率,召回率和F1得分分别为0.8750,0.8974,0.8750和0.8750。在AVEC 2014和DAIC-WOZ数据集上进行的大量实验都显示出噪声鲁棒性的增强。通过比较进一步表明,ARSLIF神经元模型在对抑郁音频数据的特征提取中提出了异常放电模式,提供了大脑启发的可解释性。
摘要:This article proposes a robust brain-inspired audio feature extractor (RBA-FE) model for depression diagnosis, using an improved hierarchical network architecture. Most deep learning models achieve state-of-the-art performance for image-based diagnostic tasks, ignoring the counterpart audio features. In order to tailor the noise challenge, RBA-FE leverages six acoustic features extracted from the raw audio, capturing both spatial characteristics and temporal dependencies. This hybrid attribute helps alleviate the precision limitation in audio feature extraction within other learning models like deep residual shrinkage networks. To deal with the noise issues, our model incorporates an improved spiking neuron model, called adaptive rate smooth leaky integrate-and-fire (ARSLIF). The ARSLIF model emulates the mechanism of ``retuning of cellular signal selectivity" in the brain attention systems, which enhances the model robustness against environmental noises in audio data. Experimental results demonstrate that RBA-FE achieves state-of-the-art accuracy on the MODMA dataset, respectively with 0.8750, 0.8974, 0.8750 and 0.8750 in precision, accuracy, recall and F1 score. Extensive experiments on the AVEC2014 and DAIC-WOZ datasets both show enhancements in noise robustness. It is further indicated by comparison that the ARSLIF neuron model suggest the abnormal firing pattern within the feature extraction on depressive audio data, offering brain-inspired interpretability.
【26】 Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
链接:https://arxiv.org/abs/2506.07081
摘要:准确、低延迟的端点对于有效的口语对话系统至关重要。虽然传统的端点通常依赖于基于频谱的音频特征,但这项工作提出了使用流式低比特率神经音频编解码器(NAC)特征的多回合对话的实时语音端点,该特征建立在神经音频编解码器的最新进展基础上。为了进一步减少截止误差,我们引入了一种新的标签延迟训练方案。在160 ms的固定中值延迟下,我们的NAC和标签延迟组合方法实现了显着的相对截止误差降低:与基线方法相比,单流端点为42.7%,双流配置为37.5%。最后,我们展示了与基于编解码器的预训练语音大语言模型的有效集成,将其中值响应时间提高了1200 ms,并将其截止误差降低了35%。
摘要:Accurate, low-latency endpointing is crucial for effective spoken dialogue systems. While traditional endpointers often rely on spectrum-based audio features, this work proposes real-time speech endpointing for multi-turn dialogues using streaming, low-bitrate Neural Audio Codec (NAC) features, building upon recent advancements in neural audio codecs. To further reduce cutoff errors, we introduce a novel label delay training scheme. At a fixed median latency of 160 ms, our combined NAC and label delay approach achieves significant relative cutoff error reductions: 42.7% for a single-stream endpointer and 37.5% for a two-stream configuration, compared to baseline methods. Finally, we demonstrate efficient integration with a codec-based pretrained speech large language model, improving its median response time by 1200 ms and reducing its cutoff error by 35%.
【27】 E-BATS: Efficient Backpropagation-Free Test-Time Adaptation for Speech Foundation Models
链接:https://arxiv.org/abs/2506.07078
备注:Under Review
摘要:当部署在涉及声学域偏移(例如背景噪声和说话者口音)的真实场景中时,语音基础模型会遇到显著的性能下降。测试时自适应(TTA)最近出现了一个可行的策略,以解决这种域的变化,在推理时,而不需要访问源数据或标签。然而,现有的TTA方法,特别是那些依赖于反向传播的方法,是内存密集型的,限制了它们在语音任务和资源受限环境中的适用性。虽然反向传播免费的方法提供了提高的效率,现有的表现出较差的精度。这是因为它们主要是为视觉任务而开发的,这与语音任务公式、噪声特征和模型架构有根本不同,带来了独特的可移植性挑战。在本文中,我们介绍了E-BATS,第一个高效的BAckpropagation-free TTA框架明确设计的语音基础模型。E-BATS通过三个关键组成部分实现了自适应有效性和记忆效率之间的平衡:(i)用于基于前向传递的特征对齐的轻量级即时自适应,(ii)用于捕获全局(话语级)和局部分布偏移(令牌级)的多尺度损失,以及(iii)用于跨话语稳定自适应的测试时间指数移动平均机制。在跨越16种声学条件的四个噪声语音数据集上进行的实验证明了一致的改进,与基于反向传播的方法相比,无反向传播基线的准确率提高了4.1%-13.5%,GPU内存节省了2.0-6.4倍。通过在声学变化下实现可扩展和鲁棒的自适应,这项工作为在现实环境中为实际语音处理系统开发更有效的自适应方法铺平了道路。
摘要:Speech Foundation Models encounter significant performance degradation when deployed in real-world scenarios involving acoustic domain shifts, such as background noise and speaker accents. Test-time adaptation (TTA) has recently emerged as a viable strategy to address such domain shifts at inference time without requiring access to source data or labels. However, existing TTA approaches, particularly those relying on backpropagation, are memory-intensive, limiting their applicability in speech tasks and resource-constrained settings. Although backpropagation-free methods offer improved efficiency, existing ones exhibit poor accuracy. This is because they are predominantly developed for vision tasks, which fundamentally differ from speech task formulations, noise characteristics, and model architecture, posing unique transferability challenges. In this paper, we introduce E-BATS, the first Efficient BAckpropagation-free TTA framework designed explicitly for speech foundation models. E-BATS achieves a balance between adaptation effectiveness and memory efficiency through three key components: (i) lightweight prompt adaptation for a forward-pass-based feature alignment, (ii) a multi-scale loss to capture both global (utterance-level) and local distribution shifts (token-level) and (iii) a test-time exponential moving average mechanism for stable adaptation across utterances. Experiments conducted on four noisy speech datasets spanning sixteen acoustic conditions demonstrate consistent improvements, with 4.1%-13.5% accuracy gains over backpropagation-free baselines and 2.0-6.4 times GPU memory savings compared to backpropagation-based methods. By enabling scalable and robust adaptation under acoustic variability, this work paves the way for developing more efficient adaptation approaches for practical speech processing systems in real-world environments.
【28】 Insights on Harmonic Tones from a Generative Music Experiment
链接:https://arxiv.org/abs/2506.07073
备注:15th International Workshop on Machine Learning and Music, September 9, 2024, Vilnius, Lithuania
摘要:生成式音乐AI的最终目的是音乐制作。工作室实验室是跨学科艺术科学分支中的一种社会形式,是一种利用人工智能音乐模型推进音乐制作的方式。在一个涉及研究人员、音乐制作人和一个用于音乐生成类似音乐的音频的人工智能模型的工作室实验中,人们观察到制作人使用模型的输出来用一个谐波复音来传达两个或更多个音高,这反过来又表明模型已经学会了使用谐波复音的单声道序列来生成结构化和连贯的同时旋律线。这些发现促使人们重新思考长期以来关于人类是否可以将谐波视为不同音高的争论,并强调生成人工智能不仅可以增强音乐创造力,还有助于更深入地理解音乐。
摘要:The ultimate purpose of generative music AI is music production. The studio-lab, a social form within the art-science branch of cross-disciplinarity, is a way to advance music production with AI music models. During a studio-lab experiment involving researchers, music producers, and an AI model for music generating bass-like audio, it was observed that the producers used the model's output to convey two or more pitches with a single harmonic complex tone, which in turn revealed that the model had learned to generate structured and coherent simultaneous melodic lines using monophonic sequences of harmonic complex tones. These findings prompt a reconsideration of the long-standing debate on whether humans can perceive harmonics as distinct pitches and highlight how generative AI can not only enhance musical creativity but also contribute to a deeper understanding of music.
【29】 "In This Environment, As That Speaker": A Text-Driven Framework for Multi-Attribute Speech Conversion
备注:Accepted by Interspeech2025
摘要:我们提出了TES-VC(文本驱动的环境和扬声器可控语音转换),一个文本驱动的语音转换框架与独立控制扬声器音色和环境声学。TES-VC处理目标语音和环境的同时文本输入,准确地生成与所描述的音色/环境匹配的语音,同时保留源内容。通过潜在扩散模型对具有解耦的声音/环境特征的合成数据进行训练,我们的方法消除了属性之间的干扰。基于检索的音色控制(RBTC)模块可以使用抽象描述进行精确操作,而无需配对数据。实验证明TES-VC能有效地生成在音质和环境上都符合语境的语音,具有较高的内容保持率和较好的可控性,具有广泛的应用前景。
摘要:We propose TES-VC (Text-driven Environment and Speaker controllable Voice Conversion), a text-driven voice conversion framework with independent control of speaker timbre and environmental acoustics. TES-VC processes simultaneous text inputs for target voice and environment, accurately generating speech matching described timbre/environment while preserving source content. Trained on synthetic data with decoupled vocal/environment features via latent diffusion modeling, our method eliminates interference between attributes. The Retrieval-Based Timbre Control (RBTC) module enables precise manipulation using abstract descriptions without paired data. Experiments confirm TES-VC effectively generates contextually appropriate speech in both timbre and environment with high content retention and superior controllability which demonstrates its potential for widespread applications.
【30】 Automatic Speech Recognition of African American English: Lexical and Contextual Effects
链接:https://arxiv.org/abs/2506.06888
备注:submitted to Interspeech 2025
摘要:自动语音识别(ASR)模型经常难以处理非裔美国英语(AAE)中的语音、音韵和形态句法特征。本研究的重点是两个关键的AAE变量:辅音集群减少(CCR)和ING减少。它检查CCR和ING减少的存在是否会增加ASR错误识别。随后,它调查是否没有外部语言模型(LM)的端到端的ASR系统更受词汇邻域效应和上下文的可预测性相比,LM系统。使用wav2vec 2.0在有和没有LM的情况下转录区域非裔美国人语言语料库(CORAAL)。CCR和ING减少检测使用蒙特利尔强制对齐(MFA)发音扩展。分析表明,CCR和ING对词错误率(WER)的影响很小,但显着,并表明在没有LM的ASR系统中存在更强的词汇邻居效应。
摘要:Automatic Speech Recognition (ASR) models often struggle with the phonetic, phonological, and morphosyntactic features found in African American English (AAE). This study focuses on two key AAE variables: Consonant Cluster Reduction (CCR) and ING-reduction. It examines whether the presence of CCR and ING-reduction increases ASR misrecognition. Subsequently, it investigates whether end-to-end ASR systems without an external Language Model (LM) are more influenced by lexical neighborhood effect and less by contextual predictability compared to systems with an LM. The Corpus of Regional African American Language (CORAAL) was transcribed using wav2vec 2.0 with and without an LM. CCR and ING-reduction were detected using the Montreal Forced Aligner (MFA) with pronunciation expansion. The analysis reveals a small but significant effect of CCR and ING on Word Error Rate (WER) and indicates a stronger presence of lexical neighborhood effect in ASR systems without LMs.
【31】 Multimodal Spatial Language Maps for Robot Navigation and Manipulation
链接:https://arxiv.org/abs/2506.06862
备注:accepted to International Journal of Robotics Research (IJRR). 24 pages, 18 figures. The paper contains texts from VLMaps(arXiv:2210.05714) and AVLMaps(arXiv:2303.07522). The project page is this https URL
摘要:将语言接地到导航代理的观察可以利用预先训练的多模态基础模型来将感知与对象或事件描述相匹配。然而,以前的方法仍然断开与环境映射,缺乏几何地图的空间精度,或忽视视觉以外的附加模态信息。为了解决这个问题,我们提出了多模态空间语言地图作为空间地图表示,融合了预先训练的多模态特征与环境的3D重建。我们使用标准探索自主构建这些地图。我们提出了两个实例,我们的地图,这是视觉语言地图(VLMaps)和他们的扩展视听语言地图(AVLMaps)通过添加音频信息。当与大型语言模型(LLM)结合时,VLMaps可以(i)将自然语言命令翻译成开放词汇空间目标(例如,“在沙发和电视之间”)直接定位在地图中,以及(ii)跨不同的机器人实施例共享以按需生成定制的障碍物地图。基于上述功能,AVLMaps通过引入统一的3D空间表示来扩展VLMaps,该空间表示通过融合来自预训练的多模态基础模型的特征来集成音频,视觉和语言线索。这使得机器人能够基于多模态目标查询(例如,文本、图像或音频片段)到用于导航的空间位置。此外,不同的感官输入的结合显着提高目标消歧在模糊的环境。在模拟和现实世界中的实验表明,我们的多模态空间语言地图,使zero-shot空间和多模态目标导航和提高召回率50%,在模糊的情况下。这些功能扩展到移动机器人和桌面机械手,支持由视觉,音频和空间提示引导的导航和交互。
摘要:Grounding language to a navigating agent's observations can leverage pretrained multimodal foundation models to match perceptions to object or event descriptions. However, previous approaches remain disconnected from environment mapping, lack the spatial precision of geometric maps, or neglect additional modality information beyond vision. To address this, we propose multimodal spatial language maps as a spatial map representation that fuses pretrained multimodal features with a 3D reconstruction of the environment. We build these maps autonomously using standard exploration. We present two instances of our maps, which are visual-language maps (VLMaps) and their extension to audio-visual-language maps (AVLMaps) obtained by adding audio information. When combined with large language models (LLMs), VLMaps can (i) translate natural language commands into open-vocabulary spatial goals (e.g., "in between the sofa and TV") directly localized in the map, and (ii) be shared across different robot embodiments to generate tailored obstacle maps on demand. Building upon the capabilities above, AVLMaps extend VLMaps by introducing a unified 3D spatial representation integrating audio, visual, and language cues through the fusion of features from pretrained multimodal foundation models. This enables robots to ground multimodal goal queries (e.g., text, images, or audio snippets) to spatial locations for navigation. Additionally, the incorporation of diverse sensory inputs significantly enhances goal disambiguation in ambiguous environments. Experiments in simulation and real-world settings demonstrate that our multimodal spatial language maps enable zero-shot spatial and multimodal goal navigation and improve recall by 50% in ambiguous scenarios. These capabilities extend to mobile robots and tabletop manipulators, supporting navigation and interaction guided by visual, audio, and spatial cues.
【32】 Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs
链接:https://arxiv.org/abs/2506.06820
摘要:音频大语言模型(AudioLLM)在语音识别和翻译等语义任务中取得了很好的效果,但在建模情感等非语言线索方面仍然有限。现有的方法通常将情绪理解视为分类问题,对预测背后的基本原理几乎没有深入了解。在这项工作中,我们探索情感推理,这是一种利用AudioLLM的生成能力,通过产生语义对齐,以证据为基础的解释来增强情感识别的策略。为了在多任务AudioLLM中支持这一点,我们引入了一个统一的框架,该框架结合了推理增强的数据监督,双编码器架构和任务交替训练。这种方法使AudioLLM能够有效地学习不同的任务,同时结合情感推理。IEMOCAP和MELD上的实验表明,我们的方法不仅提高了情感预测的准确性,但也提高了生成的响应的一致性和证据基础。
摘要:Audio Large Language Models (AudioLLMs) have achieved strong results in semantic tasks like speech recognition and translation, but remain limited in modeling paralinguistic cues such as emotion. Existing approaches often treat emotion understanding as a classification problem, offering little insight into the underlying rationale behind predictions. In this work, we explore emotion reasoning, a strategy that leverages the generative capabilities of AudioLLMs to enhance emotion recognition by producing semantically aligned, evidence-grounded explanations. To support this in multitask AudioLLMs, we introduce a unified framework combining reasoning-augmented data supervision, dual-encoder architecture, and task-alternating training. This approach enables AudioLLMs to effectively learn different tasks while incorporating emotional reasoning. Experiments on IEMOCAP and MELD show that our approach not only improves emotion prediction accuracy but also enhances the coherence and evidential grounding of the generated responses.
【33】 SynHate: Detecting Hate Speech in Synthetic Deepfake Audio
链接:https://arxiv.org/abs/2506.06772
备注:Accepted in Interspeech 2025
摘要:Deepfake音频和仇恨言论的兴起,由先进的文本到语音技术提供支持,威胁到在线安全。SynHate是第一个用于检测合成音频中仇恨言论的多语言数据集,涵盖37种语言。SynHate使用了一种新颖的四类方案:真实正常,真实仇恨,假正常和假仇恨。它基于MuTox和ADIMA数据集构建,捕捉了全球和印度的各种仇恨言论模式。我们评估了五种领先的自监督模型(Whisper-small/medium,XLS-R,AST,mHuBERT),发现不同语言的性能差异显著,Whisper-small整体表现最好。跨数据集泛化仍然是一个挑战。通过发布SynHate和基线代码,我们的目标是推进针对合成仇恨言论的强大,文化敏感和多语言解决方案。该数据集可在https://www.iab-rubric.org/resources上获得。
摘要:The rise of deepfake audio and hate speech, powered by advanced text-to-speech, threatens online safety. We present SynHate, the first multilingual dataset for detecting hate speech in synthetic audio, spanning 37 languages. SynHate uses a novel four-class scheme: Real-normal, Real-hate, Fake-normal, and Fake-hate. Built from MuTox and ADIMA datasets, it captures diverse hate speech patterns globally and in India. We evaluate five leading self-supervised models (Whisper-small/medium, XLS-R, AST, mHuBERT), finding notable performance differences by language, with Whisper-small performing best overall. Cross-dataset generalization remains a challenge. By releasing SynHate and baseline code, we aim to advance robust, culturally sensitive, and multilingual solutions against synthetic hate speech. The dataset is available at https://www.iab-rubric.org/resources.
【34】 Can Quantized Audio Language Models Perform Zero-Shot Spoofing Detection?
链接:https://arxiv.org/abs/2506.06756
备注:Accepted in Interspeech 2025
摘要:量化对于在资源受限的环境中有效地部署大型音频语言模型(LALM)至关重要。然而,它对复杂任务的影响,如zero-shot音频欺骗检测,仍然没有得到充分的研究。这项研究评估了五种LALM的zero-shot能力,GAMA,LTU-AS,MERaLiON,Qwen-Audio和SALMONN,跨越三个不同的数据集:ASVspoof 2019,In-the-Wild和WaveFake,并研究了它们对量化的鲁棒性(FP 32,FP 16,INT 8)。尽管初始欺骗检测精度很高,但我们的分析表明,所有模型对欺骗分类都存在严重的预测偏差,使其实际性能等同于随机分类。有趣的是,与FP 32相比,量化到FP 16精度导致的性能下降可以忽略不计,有效地将内存和计算需求减半,而不会对准确性产生重大影响。然而,INT 8量化加剧了模型偏差,显著降低了平衡精度。这些发现强调了关键的架构限制,并强调FP 16量化是一种最佳权衡,为实际部署和未来模型改进提供了指导方针。
摘要:Quantization is essential for deploying large audio language models (LALMs) efficiently in resource-constrained environments. However, its impact on complex tasks, such as zero-shot audio spoofing detection, remains underexplored. This study evaluates the zero-shot capabilities of five LALMs, GAMA, LTU-AS, MERaLiON, Qwen-Audio, and SALMONN, across three distinct datasets: ASVspoof2019, In-the-Wild, and WaveFake, and investigates their robustness to quantization (FP32, FP16, INT8). Despite high initial spoof detection accuracy, our analysis demonstrates severe predictive biases toward spoof classification across all models, rendering their practical performance equivalent to random classification. Interestingly, quantization to FP16 precision resulted in negligible performance degradation compared to FP32, effectively halving memory and computational requirements without materially impacting accuracy. However, INT8 quantization intensified model biases, significantly degrading balanced accuracy. These findings highlight critical architectural limitations and emphasize FP16 quantization as an optimal trade-off, providing guidelines for practical deployment and future model refinement.
【35】 A Fast and Lightweight Model for Causal Audio-Visual Speech Separation
链接:https://arxiv.org/abs/2506.06689
备注:8 pages, 5 figures
摘要:视听语音分离(AVSS)的目的是利用听觉和视觉(嘴唇运动)线索从混合信号中提取目标语音信号。然而,大多数现有的AVSS方法表现出复杂的体系结构,并依赖于未来的上下文,离线操作,这使得它们不适合实时应用。受RTFSNet流水线的启发,我们提出了一种新的流式AVSS模型,命名为Swift-Net,它增强了实时应用所需的因果处理能力。Swift-Net采用轻量级的视觉特征提取模块和高效的融合模块进行视听融合。此外,Swift-Net采用分组SRU来整合不同特征空间的历史信息,从而提高历史信息的利用效率。我们进一步提出了一个因果转换模板,以促进非因果AVSS模型转换成因果对应。在三个标准基准数据集(LRS 2,LRS 3和VoxCeleb 2)上的实验表明,在因果条件下,我们提出的Swift-Net表现出出色的性能,突出了这种方法在复杂环境中处理语音的潜力。
摘要:Audio-visual speech separation (AVSS) aims to extract a target speech signal from a mixed signal by leveraging both auditory and visual (lip movement) cues. However, most existing AVSS methods exhibit complex architectures and rely on future context, operating offline, which renders them unsuitable for real-time applications. Inspired by the pipeline of RTFSNet, we propose a novel streaming AVSS model, named Swift-Net, which enhances the causal processing capabilities required for real-time applications. Swift-Net adopts a lightweight visual feature extraction module and an efficient fusion module for audio-visual integration. Additionally, Swift-Net employs Grouped SRUs to integrate historical information across different feature spaces, thereby improving the utilization efficiency of historical information. We further propose a causal transformation template to facilitate the conversion of non-causal AVSS models into causal counterparts. Experiments on three standard benchmark datasets (LRS2, LRS3, and VoxCeleb2) demonstrated that under causal conditions, our proposed Swift-Net exhibited outstanding performance, highlighting the potential of this method for processing speech in complex environments.
【36】 CAtCh: Cognitive Assessment through Cookie Thief
链接:https://arxiv.org/abs/2506.06603
摘要:已经开发了几种机器学习算法,用于根据自发言语预测阿尔茨海默病和相关痴呆症(ADRD)。然而,这些算法都没有被翻译为预测更广泛的认知障碍(CI),这在某些情况下是ADRD的前兆和风险因素。在本文中,我们评估了几种最初提出用于预测ADRD的基于语音的开源方法,以及用于从患者录音中预测CI的任务的多模态情感分析方法。结果表明,多模态的方法优于单峰的CI预测,和基于声学的方法比基于语言学的表现更好。具体而言,可解释的声学特征有关的影响和韵律被发现显着优于基于BERT的语言特征和可解释的语言特征,分别。为这项研究开发的所有代码都可以在https://github.com/JTColonel/catch上获得。
摘要:Several machine learning algorithms have been developed for the prediction of Alzheimer's disease and related dementia (ADRD) from spontaneous speech. However, none of these algorithms have been translated for the prediction of broader cognitive impairment (CI), which in some cases is a precursor and risk factor of ADRD. In this paper, we evaluated several speech-based open-source methods originally proposed for the prediction of ADRD, as well as methods from multimodal sentiment analysis for the task of predicting CI from patient audio recordings. Results demonstrated that multimodal methods outperformed unimodal ones for CI prediction, and that acoustics-based approaches performed better than linguistics-based ones. Specifically, interpretable acoustic features relating to affect and prosody were found to significantly outperform BERT-based linguistic features and interpretable linguistic features, respectively. All the code developed for this study is available at https://github.com/JTColonel/catch.
【37】 Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models
链接:https://arxiv.org/abs/2506.06537
备注:Accepted on INTERSPEECH2025
摘要:音视频分割(Audiovisual segmentation,AVS)的目的是识别与声源对应的视觉区域,在视频理解、监控和人机交互等方面发挥着重要作用。传统的AVS方法依赖于大规模的像素级注释,这是昂贵和耗时的获得。为了解决这个问题,我们提出了一种新的zero-shot AVS框架,通过利用多个预训练模型来消除特定于任务的训练。我们的方法集成了音频,视觉和文本表示来弥合模态差距,实现精确的声源分割,而无需AVS特定的注释。我们系统地探索了连接预训练模型的不同策略,并评估了它们在多个数据集上的有效性。实验结果表明,我们的框架实现了最先进的zero-shot AVS性能,凸显了多模式模型集成对于细粒度视听分割的有效性。
摘要:Audiovisual segmentation (AVS) aims to identify visual regions corresponding to sound sources, playing a vital role in video understanding, surveillance, and human-computer interaction. Traditional AVS methods depend on large-scale pixel-level annotations, which are costly and time-consuming to obtain. To address this, we propose a novel zero-shot AVS framework that eliminates task-specific training by leveraging multiple pretrained models. Our approach integrates audio, vision, and text representations to bridge modality gaps, enabling precise sound source segmentation without AVS-specific annotations. We systematically explore different strategies for connecting pretrained models and evaluate their efficacy across multiple datasets. Experimental results demonstrate that our framework achieves state-of-the-art zero-shot AVS performance, highlighting the effectiveness of multimodal model integration for finegrained audiovisual segmentation.
【38】 TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment
链接:https://arxiv.org/abs/2506.06343
摘要:语音语言模型的最新进展在构建智能语音助手方面显示出了可喜的成果。然而,大多数现有的方法依赖于大规模配对的语音文本数据和广泛的计算资源,这在可扩展性和可访问性方面带来了挑战。在本文中,我们提出了\textbf{TESU-LLM},这是一个新的框架,可以只使用文本数据来训练支持语音的语言模型。我们的关键见解是利用一个统一的编码器,映射语义等效的文本和语音输入到一个共享的潜在空间。通过将编码器输出与LLM的嵌入空间通过轻量级投影网络对齐,我们使模型能够从纯文本监督推广到基于语音的推理。尽管TESU-LLM只接受文本训练,但它在各种语音相关基准测试中表现出色,与使用大规模多模态数据集和大量计算资源训练的基线方法相当。这些结果突出了我们的方法的有效性和效率,为在没有语音数据的情况下构建语音LLM提供了一条可扩展的路径。
摘要:Recent advances in speech-enabled language models have shown promising results in building intelligent voice assistants. However, most existing approaches rely on large-scale paired speech-text data and extensive computational resources, which pose challenges in terms of scalability and accessibility. In this paper, we present \textbf{TESU-LLM}, a novel framework that enables training speech-capable language models using only text data. Our key insight is to leverage a unified encoder that maps semantically equivalent text and speech inputs to a shared latent space. By aligning the encoder output with the embedding space of a LLM via a lightweight projection network, we enable the model to generalize from text-only supervision to speech-based inference. Despite being trained exclusively on text, TESU-LLM achieves strong performance on various speech-related benchmarks, comparable to baseline methods trained with large-scale multimodal datasets and substantial computational resources. These results highlight the effectiveness and efficiency of our approach, offering a scalable path toward building speech LLMs without speech data.
机器翻译由腾讯交互翻译提供,仅供参考
