微信公众号:arXiv_Daily
cs.SD语音
【1】VGGSounder: Audio-Visual Evaluations for Foundation Models
标题:VGG Sounder:基础模型的视听评估
链接:https://arxiv.org/abs/2508.08237
备注:Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) 2025
摘要:视听基础模型的出现强调了可靠评估其多模态理解的重要性。VGGSounder数据集通常用作评估视听分类的基准。然而,我们的分析确定了VGGSounder的几个局限性,包括不完整的标签,部分重叠的类,和错位的方式。这些导致对听觉和视觉能力的扭曲评估。为了解决这些限制,我们引入了VGGSounder,这是一个全面重新注释的多标签测试集,扩展了VGGSound,专门用于评估视听基础模型。VGGSounder具有详细的模态注释功能,可以精确分析特定模态的性能。此外,我们揭示了模型的局限性,通过分析性能下降时,添加另一个输入模态与我们的新的模态混淆度量。
摘要:The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding. The VGGSounder dataset is commonly used as a benchmark for evaluation audio-visual classification. However, our analysis identifies several limitations of VGGSounder, including incomplete labelling, partially overlapping classes, and misaligned modalities. These lead to distorted evaluations of auditory and visual capabilities. To address these limitations, we introduce VGGSounder, a comprehensively re-annotated, multi-label test set that extends VGGSound and is specifically designed to evaluate audio-visual foundation models. VGGSounder features detailed modality annotations, enabling precise analyses of modality-specific performance. Furthermore, we reveal model limitations by analysing performance degradation when adding another input modality with our new modality confusion metric.
【2】Pindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine Grained-Localization
标题:别碰它!用于鲁棒检测和细粒度定位的音频和视觉Deepfake对策
链接:https://arxiv.org/abs/2508.08141
摘要:视频和音频生成领域正在以最先进的新方法蓬勃发展。新技术的快速发展凸显了对检测视频中合成内容的强大解决方案的需求。特别是,当在视觉、音频或这两个领域中通过局部操纵进行细粒度更改时,这些细微的修改给检测算法带来了挑战。本文提出了deepfake视频分类和定位问题的解决方案。这些方法被提交给ACM 1M Deepfakes检测挑战赛,在时间定位任务中获得了最佳性能,并在评估数据集的TestA分割的分类任务中获得了前四名。
摘要:The field of visual and audio generation is burgeoning with new state-of-the-art methods. This rapid proliferation of new techniques underscores the need for robust solutions for detecting synthetic content in videos. In particular, when fine-grained alterations via localized manipulations are performed in visual, audio, or both domains, these subtle modifications add challenges to the detection algorithms. This paper presents solutions for the problems of deepfake video classification and localization. The methods were submitted to the ACM 1M Deepfakes Detection Challenge, achieving the best performance in the temporal localization task and a top four ranking in the classification task for the TestA split of the evaluation dataset.
【3】Audio-Thinker: Guiding Audio Language Model When and How to Think via Reinforcement Learning
标题:音频思想者:通过强化学习指导音频语言模型何时以及如何思考
链接:https://arxiv.org/abs/2508.08039
备注:preprint
摘要:大型语言模型、多模态大型语言模型和大型音频语言模型(LALM)的最新进展通过基于规则的奖励强化学习显著提高了它们的推理能力。然而,显式推理过程尚未显示出对音频问题回答的显着好处,有效地利用深度推理仍然是一个开放的挑战,LALM仍然达不到人类水平的语义语言推理。为了解决这些限制,我们提出了Audio-Thinker,这是一个强化学习框架,旨在增强LALM的推理能力,重点是提高适应性,一致性和有效性。我们的方法引入了自适应的思考准确性奖励,使模型能够根据任务复杂性动态调整其推理策略。此外,我们还引入了一个外部奖励模型来评估推理过程的整体一致性和质量,并辅以基于思维的奖励,帮助模型在训练过程中区分有效和有缺陷的推理路径。实验结果表明,我们的Audio-Thinker模型优于现有的面向推理的LALM在各种基准任务,表现出优越的推理和泛化能力。
摘要:Recent advancements in large language models, multimodal large language models, and large audio language models (LALMs) have significantly improved their reasoning capabilities through reinforcement learning with rule-based rewards. However, the explicit reasoning process has yet to show significant benefits for audio question answering, and effectively leveraging deep reasoning remains an open challenge, with LALMs still falling short of human-level auditory-language reasoning. To address these limitations, we propose Audio-Thinker, a reinforcement learning framework designed to enhance the reasoning capabilities of LALMs, with a focus on improving adaptability, consistency, and effectiveness. Our approach introduces an adaptive think accuracy reward, enabling the model to adjust its reasoning strategies based on task complexity dynamically. Furthermore, we incorporate an external reward model to evaluate the overall consistency and quality of the reasoning process, complemented by think-based rewards that help the model distinguish between valid and flawed reasoning paths during training. Experimental results demonstrate that our Audio-Thinker model outperforms existing reasoning-oriented LALMs across various benchmark tasks, exhibiting superior reasoning and generalization capabilities.
【4】Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
标题:为结构障碍性语音识别弥合ASC和LLM:自我监督和生成方法的基准
链接:https://arxiv.org/abs/2508.08027
摘要:语音识别(ASR)由于音素失真和高可变性。虽然像Wav 2 Vec、HuBERT和Whisper这样的自我监督ASR模型已经显示出了希望,但它们在构音障碍语音中的有效性仍不清楚。这项研究系统地基准这些模型与不同的解码策略,包括CTC,seq 2seq和LLM增强解码(BART,GPT-2,Vicuna)。我们的贡献包括(1)对构音障碍语音的ASR架构进行基准测试,(2)引入基于LLM的解码以提高可懂度,(3)分析跨数据集的泛化,以及(4)提供对跨严重级别的识别错误的见解。研究结果强调,LLM增强解码改善构音障碍的ASR利用音素恢复和语法纠正的语言限制。
摘要:Speech Recognition (ASR) due to phoneme distortions and high variability. While self-supervised ASR models like Wav2Vec, HuBERT, and Whisper have shown promise, their effectiveness in dysarthric speech remains unclear. This study systematically benchmarks these models with different decoding strategies, including CTC, seq2seq, and LLM-enhanced decoding (BART,GPT-2, Vicuna). Our contributions include (1) benchmarking ASR architectures for dysarthric speech, (2) introducing LLM-based decoding to improve intelligibility, (3) analyzing generalization across datasets, and (4) providing insights into recognition errors across severity levels. Findings highlight that LLM-enhanced decoding improves dysarthric ASR by leveraging linguistic constraints for phoneme restoration and grammatical correction.
【5】Exploring Procedural Data Generation for Automatic Acoustic Guitar Fingerpicking Transcription
标题:探索自动声学吉他指纹识别转录的程序数据生成
链接:https://arxiv.org/abs/2508.07987
备注:Accepted to the 6th Conference on AI Music Creativity (AIMC), 2025
摘要:原声吉他指弹演奏的自动转录仍然是一项具有挑战性的任务,由于缺乏标记的训练数据和与音乐录音相关的法律约束。这项工作研究了一个程序化的数据生成管道,作为训练转录模型的真实音频记录的替代方案。我们的方法通过四个阶段来合成训练数据:基于知识的指法谱组成,可调性能渲染,使用扩展的Karplus-Strong算法进行物理建模,以及包括混响和失真在内的音频增强。我们训练和评估基于CRNN的笔记跟踪模型的真实和合成数据集,证明程序数据可以用来实现合理的笔记跟踪结果。使用少量的真实数据进行微调进一步提高了转录的准确性,比只在真实记录上训练的模型有所改进。这些结果突出了程序生成的音频数据稀缺的音乐信息检索任务的潜力。
摘要:Automatic transcription of acoustic guitar fingerpicking performances remains a challenging task due to the scarcity of labeled training data and legal constraints connected with musical recordings. This work investigates a procedural data generation pipeline as an alternative to real audio recordings for training transcription models. Our approach synthesizes training data through four stages: knowledge-based fingerpicking tablature composition, MIDI performance rendering, physical modeling using an extended Karplus-Strong algorithm, and audio augmentation including reverb and distortion. We train and evaluate a CRNN-based note-tracking model on both real and synthetic datasets, demonstrating that procedural data can be used to achieve reasonable note-tracking results. Finetuning with a small amount of real data further enhances transcription accuracy, improving over models trained exclusively on real recordings. These results highlight the potential of procedurally generated audio for data-scarce music information retrieval tasks.
【6】Joint Transcription of Acoustic Guitar Strumming Directions and Chords
标题:原声吉他弹奏方向和和弦的联合抄写
链接:https://arxiv.org/abs/2508.07973
备注:Accepted to the 26th International Society for Music Information Retrieval Conference (ISMIR), 2025
摘要:吉他弹奏的自动转录是音乐信息检索(MIR)中一个代表性不足且具有挑战性的任务,特别是用于从音频信号中提取弹奏方向和和弦进行。虽然现有的方法显示出希望,但它们的有效性往往受到有限数据集的阻碍。在这项工作中,我们通过引入一个新的数据集和一个基于深度学习的转录模型,将多模式方法扩展到吉他弹奏转录。我们使用ESP32智能手表运动传感器和结构化录音协议收集了90分钟的真实吉他录音,并辅以4小时标记的弹奏音频的合成数据集。卷积递归神经网络(CRNN)模型被训练用于检测弹奏事件,对其方向进行分类,并仅使用麦克风音频识别相应的和弦。我们的评估表明,与基线发作检测算法相比,该算法有了显着的改进,结合合成和真实数据的混合方法在弹奏动作检测和和弦分类方面都达到了最高的准确度。这些结果突出了深度学习在强大的吉他弹奏转录方面的潜力,并为自动节奏吉他分析开辟了新的途径。
摘要:Automatic transcription of guitar strumming is an underrepresented and challenging task in Music Information Retrieval (MIR), particularly for extracting both strumming directions and chord progressions from audio signals. While existing methods show promise, their effectiveness is often hindered by limited datasets. In this work, we extend a multimodal approach to guitar strumming transcription by introducing a novel dataset and a deep learning-based transcription model. We collect 90 min of real-world guitar recordings using an ESP32 smartwatch motion sensor and a structured recording protocol, complemented by a synthetic dataset of 4h of labeled strumming audio. A Convolutional Recurrent Neural Network (CRNN) model is trained to detect strumming events, classify their direction, and identify the corresponding chords using only microphone audio. Our evaluation demonstrates significant improvements over baseline onset detection algorithms, with a hybrid method combining synthetic and real-world data achieving the highest accuracy for both strumming action detection and chord classification. These results highlight the potential of deep learning for robust guitar strumming transcription and open new avenues for automatic rhythm guitar analysis.
【7】SCDF: A Speaker Characteristics DeepFake Speech Dataset for Bias Analysis
标题:SCDF:用于偏见分析的说话者特征DeepFake语音数据集
链接:https://arxiv.org/abs/2508.07944
摘要:尽管人们越来越关注deepfake语音检测,但在语音领域,偏见和公平性方面仍然没有得到充分的研究。为了解决这一差距,我们引入了Speaker Characteristics Deepfake(SCDF)数据集:一种新颖的、注释丰富的资源,可以系统地评估Deepfake语音检测中的人口统计学偏见。SCDF包含超过237 000个话语,男女发言者的代表性均衡,跨越五种语言和广泛的年龄范围。我们评估了几个国家的最先进的检测器,并表明扬声器的特性显着影响检测性能,揭示了性别,语言,年龄和合成器类型的差异。这些发现强调了偏见意识发展的必要性,并为建立符合道德和监管标准的非歧视性deepfake检测系统提供了基础。
摘要:Despite growing attention to deepfake speech detection, the aspects of bias and fairness remain underexplored in the speech domain. To address this gap, we introduce the Speaker Characteristics Deepfake (SCDF) dataset: a novel, richly annotated resource enabling systematic evaluation of demographic biases in deepfake speech detection. SCDF contains over 237,000 utterances in a balanced representation of both male and female speakers spanning five languages and a wide age range. We evaluate several state-of-the-art detectors and show that speaker characteristics significantly influence detection performance, revealing disparities across sex, language, age, and synthesizer type. These findings highlight the need for bias-aware development and provide a foundation for building non-discriminatory deepfake detection systems aligned with ethical and regulatory standards.
【8】Filling MIDI Velocity using U-Net Image Colorizer
标题:使用U-Net图像着色器填充收件箱速度
链接:https://arxiv.org/abs/2508.07751
备注:12 pages, submitted to CMMR2025 conference
摘要:现代音乐制作者通常使用乐器数字接口(Musical Instrument Digital Interface)来存储他们的音乐作品。然而,使用数字软件创建的MIDI文件可能缺乏人类表演的表现力特征,基本上留下了速度参数-音符响度的控制-未定义,默认值为平坦值。填充音频速度的任务被称为音频速度预测,其使用回归模型通过仅调整该参数来增强音乐表现力。在本文中,我们介绍了U-Net,一个广泛采用的架构,在图像彩色化,这项任务。通过将MIDI数据概念化为图像,我们采用窗口注意力并开发了一个自定义损失函数来解决MIDI转换图像的稀疏性。目前的数据集可用性限制了我们的实验钢琴数据。在MAESTRO v3和SMD数据集上进行评估,我们提出的填充速度的方法在定量指标和定性听力测试方面都优于以前的方法。
摘要:Modern music producers commonly use MIDI (Musical Instrument Digital Interface) to store their musical compositions. However, MIDI files created with digital software may lack the expressive characteristics of human performances, essentially leaving the velocity parameter - a control for note loudness - undefined, which defaults to a flat value. The task of filling MIDI velocity is termed MIDI velocity prediction, which uses regression models to enhance music expressiveness by adjusting only this parameter. In this paper, we introduce the U-Net, a widely adopted architecture in image colorization, to this task. By conceptualizing MIDI data as images, we adopt window attention and develop a custom loss function to address the sparsity of MIDI-converted images. Current dataset availability restricts our experiments to piano data. Evaluated on the MAESTRO v3 and SMD datasets, our proposed method for filling MIDI velocity outperforms previous approaches in both quantitative metrics and qualitative listening tests.
【9】AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition
标题:AD-AVSR:用于鲁棒视听语音识别的非对称双流增强
链接:https://arxiv.org/abs/2508.07608
备注:Accepted by the ACM MM 2025 Workshop on SVC
摘要:视听语音识别(AVSR)结合视听模态来提高语音识别,特别是在嘈杂的环境中。然而,大多数现有的方法部署的单向增强或对称融合的方式,这限制了他们的能力,捕捉异构和互补的视听数据的相关性,特别是在非对称信息条件下。为了解决这些差距,我们引入了一个新的AVSR框架称为AD-AVSR的基础上双向模态增强。具体来说,我们首先引入音频双流编码策略,从多个角度丰富音频表示,并有意建立不对称性,以支持后续的跨模态交互。增强过程涉及两个关键组件,用于在音频引导下增强视觉表示的音频感知视觉细化模块,以及使用视觉线索细化音频表示的跨模态噪声抑制掩蔽模块,协同导致闭环和双向信息流。为了进一步增强相关鲁棒性,我们采用基于阈值的选择机制来过滤掉不相关或弱相关的视听对。在LRS 2和LRS 3数据集上的大量实验结果表明,我们的AD-AVSR在性能和噪声鲁棒性方面始终优于SOTA方法,突出了我们模型设计的有效性。
摘要:Audio-visual speech recognition (AVSR) combines audio-visual modalities to improve speech recognition, especially in noisy environments. However, most existing methods deploy the unidirectional enhancement or symmetric fusion manner, which limits their capability to capture heterogeneous and complementary correlations of audio-visual data-especially under asymmetric information conditions. To tackle these gaps, we introduce a new AVSR framework termed AD-AVSR based on bidirectional modality enhancement. Specifically, we first introduce the audio dual-stream encoding strategy to enrich audio representations from multiple perspectives and intentionally establish asymmetry to support subsequent cross-modal interactions. The enhancement process involves two key components, Audio-aware Visual Refinement Module for enhanced visual representations under audio guidance, and Cross-modal Noise Suppression Masking Module which refines audio representations using visual cues, collaboratively leading to the closed-loop and bidirectional information flow. To further enhance correlation robustness, we adopt a threshold-based selection mechanism to filter out irrelevant or weakly correlated audio-visual pairs. Extensive experimental results on the LRS2 and LRS3 datasets indicate that our AD-AVSR consistently surpasses SOTA methods in both performance and noise robustness, highlighting the effectiveness of our model design.
【10】Voice Pathology Detection Using Phonation
标题:使用发音的声音病理检测
链接:https://arxiv.org/abs/2508.07587
备注:17 Pages, 11 Figures
摘要:语音障碍严重影响沟通和生活质量,需要早期准确诊断。传统的方法,如喉镜检查是侵入性的,主观的,往往难以接近。这项研究提出了一种非侵入性的,基于机器学习的框架,用于使用发声数据检测语音病理。 从Saarbrucken语音数据库的发声数据进行了分析,使用声学特征,如梅尔频率倒谱系数(MFCC),色度特征,梅尔频谱图。循环神经网络(RNN),包括LSTM和注意力机制,将样本分类为正常和病理类别。数据增强技术,包括音调偏移和高斯噪声添加,增强了模型的泛化能力,而预处理则确保了信号质量。基于尺度的特征,如H“older和Hurst指数,进一步捕获信号的不规则性和长期依赖性。 拟议的框架提供了一种非侵入性的自动诊断工具,用于早期检测语音病理,支持AI驱动的医疗保健,并改善患者的治疗效果。
摘要:Voice disorders significantly affect communication and quality of life, requiring an early and accurate diagnosis. Traditional methods like laryngoscopy are invasive, subjective, and often inaccessible. This research proposes a noninvasive, machine learning-based framework for detecting voice pathologies using phonation data. Phonation data from the Saarbr\"ucken Voice Database are analyzed using acoustic features such as Mel Frequency Cepstral Coefficients (MFCCs), chroma features, and Mel spectrograms. Recurrent Neural Networks (RNNs), including LSTM and attention mechanisms, classify samples into normal and pathological categories. Data augmentation techniques, including pitch shifting and Gaussian noise addition, enhance model generalizability, while preprocessing ensures signal quality. Scale-based features, such as H\"older and Hurst exponents, further capture signal irregularities and long-term dependencies. The proposed framework offers a noninvasive, automated diagnostic tool for early detection of voice pathologies, supporting AI-driven healthcare, and improving patient outcomes.
【11】Exploring Efficient Directional and Distance Cues for Regional Speech Separation
标题:探索区域语音分离的有效方向和距离线索
链接:https://arxiv.org/abs/2508.07563
备注:This paper has been accepted by Interspeech 2025
摘要:在本文中,我们介绍了一种基于神经网络的方法,使用麦克风阵列的区域语音分离。该方法利用新颖的空间线索,不仅从指定的方向,而且在定义的距离内提取声源。具体来说,我们的方法采用了改进的延迟和求和技术来获得方向线索,大大增强了来自目标方向的信号。我们通过将直达混响比纳入输入特征来进一步增强分离,使模型能够更好地区分指定距离内和指定距离外的源。实验结果表明,我们提出的方法导致在多个客观指标的实质性收益。此外,我们的方法在CHiME-8 MMCSG数据集上实现了最先进的性能,该数据集记录在真实世界的对话场景中,强调了其在实际应用中语音分离的有效性。
摘要:In this paper, we introduce a neural network-based method for regional speech separation using a microphone array. This approach leverages novel spatial cues to extract the sound source not only from specified direction but also within defined distance. Specifically, our method employs an improved delay-and-sum technique to obtain directional cues, substantially enhancing the signal from the target direction. We further enhance separation by incorporating the direct-to-reverberant ratio into the input features, enabling the model to better discriminate sources within and beyond a specified distance. Experimental results demonstrate that our proposed method leads to substantial gains across multiple objective metrics. Furthermore, our method achieves state-of-the-art performance on the CHiME-8 MMCSG dataset, which was recorded in real-world conversational scenarios, underscoring its effectiveness for speech separation in practical applications.
【12】A Small-footprint Acoustic Echo Cancellation Solution for Mobile Full-Duplex Speech Interactions
标题:一种用于移动全速语音交互的小占地面积声学回声消除解决方案
链接:https://arxiv.org/abs/2508.07561
备注:This paper is accepted to ICASSP 2025
摘要:在全双工语音交互系统中,有效的回声消除(AEC)是恢复回声污染语音的关键。本文提出了一种基于神经网络的AEC解决方案,以解决移动场景中的挑战,不同的硬件,非线性失真和长延迟。我们首先采用不同的数据增强策略,以提高模型在各种环境中的鲁棒性。此外,渐进式学习,逐步提高AEC的有效性,从而在语音质量的显着改善。为了进一步优化AEC的下游应用,我们引入了一种新的后处理策略,采用专门为语音活动检测(VAD)和自动语音识别(ASR)等任务设计的定制参数,从而提高其整体功效。最后,我们的方法采用了一个小的足迹模型与流推理,使移动设备上的无缝部署。实验结果表明,该方法在回声回波损耗增强和语音质量的感知评估的有效性,以及在VAD和ASR的结果显着改善。
摘要:In full-duplex speech interaction systems, effective Acoustic Echo Cancellation (AEC) is crucial for recovering echo-contaminated speech. This paper presents a neural network-based AEC solution to address challenges in mobile scenarios with varying hardware, nonlinear distortions and long latency. We first incorporate diverse data augmentation strategies to enhance the model's robustness across various environments. Moreover, progressive learning is employed to incrementally improve AEC effectiveness, resulting in a considerable improvement in speech quality. To further optimize AEC's downstream applications, we introduce a novel post-processing strategy employing tailored parameters designed specifically for tasks such as Voice Activity Detection (VAD) and Automatic Speech Recognition (ASR), thus enhancing their overall efficacy. Finally, our method employs a small-footprint model with streaming inference, enabling seamless deployment on mobile devices. Empirical results demonstrate effectiveness of the proposed method in Echo Return Loss Enhancement and Perceptual Evaluation of Speech Quality, alongside significant improvements in both VAD and ASR results.
【13】Think Before You Talk: Enhancing Meaningful Dialogue Generation in Full-Duplex Speech Language Models with Planning-Inspired Text Guidance
标题:说话前先思考:利用受规划启发的文本指导增强全复式语音语言模型中有意义的对话生成
链接:https://arxiv.org/abs/2508.07375
备注:Work in progress
摘要:全双工语音语言模型(FD-SLMs)是专门的基础模型,旨在通过对中断、反向通道和重叠语音等复杂的对话动态进行建模来实现自然、实时的语音交互,而端到端(e2 e)FD-SLMs利用现实世界的双通道对话数据来捕捉细致入微的两个说话者对话模式,以实现类似人类的交互。然而,他们面临着一个关键的挑战-由于长时间的语音序列和有限的高质量口语对话数据,与纯文本对话相比,他们的会话能力往往会下降。虽然文本引导的语音生成可以缓解这些问题,但当将文本引导集成到双通道音频流中时,它会遇到时间和长度问题,从而破坏自然交互所必需的精确时间对齐。为了解决这些挑战,我们提出了TurnGuide,这是一种新颖的规划启发方法,它通过动态地将助理语音分割成对话回合并在语音输出之前生成回合级文本指导来模仿人类会话规划,从而有效地解决了插入时间和长度的挑战。大量的实验表明,我们的方法显着提高e2 e FD-SLM的会话能力,使他们能够生成语义有意义的和连贯的语音,同时保持自然的会话流。演示可在https://dreamtheater123.github.io/TurnGuide-Demo/上获得。代码将在https://github.com/dreamtheater123/TurnGuide上提供。
摘要:Full-Duplex Speech Language Models (FD-SLMs) are specialized foundation models designed to enable natural, real-time spoken interactions by modeling complex conversational dynamics such as interruptions, backchannels, and overlapping speech, and End-to-end (e2e) FD-SLMs leverage real-world double-channel conversational data to capture nuanced two-speaker dialogue patterns for human-like interactions. However, they face a critical challenge -- their conversational abilities often degrade compared to pure-text conversation due to prolonged speech sequences and limited high-quality spoken dialogue data. While text-guided speech generation could mitigate these issues, it suffers from timing and length issues when integrating textual guidance into double-channel audio streams, disrupting the precise time alignment essential for natural interactions. To address these challenges, we propose TurnGuide, a novel planning-inspired approach that mimics human conversational planning by dynamically segmenting assistant speech into dialogue turns and generating turn-level text guidance before speech output, which effectively resolves both insertion timing and length challenges. Extensive experiments demonstrate our approach significantly improves e2e FD-SLMs' conversational abilities, enabling them to generate semantically meaningful and coherent speech while maintaining natural conversational flow. Demos are available at https://dreamtheater123.github.io/TurnGuide-Demo/. Code will be available at https://github.com/dreamtheater123/TurnGuide.
【14】Keyword Mamba: Spoken Keyword Spotting with State Space Models
标题:关键词曼巴:使用状态空间模型进行口语关键词定位
链接:https://arxiv.org/abs/2508.07363
备注:Under peer review
摘要:关键词识别是语音处理中的一项重要任务。它广泛应用于语音助手和智能设备。CNN、RNN和Transformers等深度学习模型在KWS中表现良好。然而,他们往往难以处理长期模式,同时保持效率。在这项工作中,我们提出了关键字曼巴,一个新的架构KWS。它使用了一个叫做Mamba的神经状态空间模型(SSM)。我们沿着时间轴应用Mamba,并探索它如何取代Transformer模型中的自我注意部分。我们在Google Speech Commands数据集上测试了我们的模型。结果表明,Keyword Mamba算法具有较强的识别精度,且参数少,计算量小.据我们所知,这是第一次状态空间模型已被用于KWS。这些结果表明,曼巴在语音相关的任务有很强的潜力。
摘要:Keyword spotting (KWS) is an essential task in speech processing. It is widely used in voice assistants and smart devices. Deep learning models like CNNs, RNNs, and Transformers have performed well in KWS. However, they often struggle to handle long-term patterns and stay efficient at the same time. In this work, we present Keyword Mamba, a new architecture for KWS. It uses a neural state space model (SSM) called Mamba. We apply Mamba along the time axis and also explore how it can replace the self-attention part in Transformer models. We test our model on the Google Speech Commands datasets. The results show that Keyword Mamba reaches strong accuracy with fewer parameters and lower computational cost. To our knowledge, this is the first time a state space model has been used for KWS. These results suggest that Mamba has strong potential in speech-related tasks.
【15】How Does a Deep Neural Network Look at Lexical Stress?
标题:深度神经网络如何看待词汇压力?
链接:https://arxiv.org/abs/2508.07229
备注:10 pages, 4 figures, submitted to the Journal of the Acoustical Society of America (JASA)
摘要:尽管神经网络在语音处理方面取得了成功,但它们通常像黑匣子一样运行,这就引发了一个问题:是什么影响了它们的决策,以及我们如何解释它们?本文从词汇重音的角度来探讨这一问题。从阅读和自发语音中自动构建了英语双音节词数据集。训练了几种卷积神经网络(CNN)架构,以从缺乏最小重音对的双音节词(例如,初始应力Wallet,最终应力exTEND),在保持测试数据上达到高达92%的准确度。分层相关传播(LRP),一种用于CNN可解释性分析的技术,揭示了对保留最小对(PROtest与proTEST)的预测受到重读音节与非重读音节信息的最强烈影响,特别是重读元音的频谱特性。然而,分类器也关注整个单词的信息。提出了一个特定功能的相关性分析,其结果表明,我们表现最好的分类器的强烈影响,强调元音的第一和第二共振峰,有一些证据表明,其音高和第三共振峰也作出贡献。这些结果揭示了深度学习从自然发生的数据中获取分布式线索的能力,扩展了基于高度控制刺激的传统语音工作。
摘要:Despite their success in speech processing, neural networks often operate as black boxes, prompting the question: what informs their decisions, and how can we interpret them? This work examines this issue in the context of lexical stress. A dataset of English disyllabic words was automatically constructed from read and spontaneous speech. Several Convolutional Neural Network (CNN) architectures were trained to predict stress position from a spectrographic representation of disyllabic words lacking minimal stress pairs (e.g., initial stress WAllet, final stress exTEND), achieving up to 92% accuracy on held-out test data. Layerwise Relevance Propagation (LRP), a technique for CNN interpretability analysis, revealed that predictions for held-out minimal pairs (PROtest vs. proTEST ) were most strongly influenced by information in stressed versus unstressed syllables, particularly the spectral properties of stressed vowels. However, the classifiers also attended to information throughout the word. A feature-specific relevance analysis is proposed, and its results suggest that our best-performing classifier is strongly influenced by the stressed vowel's first and second formants, with some evidence that its pitch and third formant also contribute. These results reveal deep learning's ability to acquire distributed cues to stress from naturally occurring data, extending traditional phonetic work based around highly controlled stimuli.
【16】Noise-Robust Sound Event Detection and Counting via Language-Queried Sound Separation
标题:通过百分比查询声音分离进行噪音稳健的声音事件检测和计数
链接:https://arxiv.org/abs/2508.07176
摘要:大多数声音事件检测(SED)系统在干净的数据集上表现良好,但在嘈杂的环境中会显着下降。语音查询音频源分离(LASS)模型显示出通过分离目标事件来实现鲁棒SED的前景;现有方法需要精心的多阶段训练,并且缺乏对目标事件的明确指导。为了解决这些挑战,我们引入了事件外观检测(EAD),一种基于计数的方法,可以在剪辑和帧级别上计数事件的发生。基于EAD,我们提出了一个基于协同训练的多任务学习框架,EAD和SED,以提高SED在噪声环境中的性能。首先,SED努力学习与EAD相同的模式。然后,基于任务的约束设计,以提高预测的一致性之间的SED和EAD。该框架为LASS模型提供了更可靠的剪辑级预测,并增强了时间戳检测能力。在DESED和WildDESED数据集上的实验表明,与现有方法相比,性能更好,在更高的噪声水平下优势更加明显。
摘要:Most sound event detection (SED) systems perform well on clean datasets but degrade significantly in noisy environments. Language-queried audio source separation (LASS) models show promise for robust SED by separating target events; existing methods require elaborate multi-stage training and lack explicit guidance for target events. To address these challenges, we introduce event appearance detection (EAD), a counting-based approach that counts event occurrences at both the clip and frame levels. Based on EAD, we propose a co-training-based multi-task learning framework for EAD and SED to enhance SED's performance in noisy environments. First, SED struggles to learn the same patterns as EAD. Then, a task-based constraint is designed to improve prediction consistency between SED and EAD. This framework provides more reliable clip-level predictions for LASS models and strengthens timestamp detection capability. Experiments on DESED and WildDESED datasets demonstrate better performance compared to existing methods, with advantages becoming more pronounced at higher noise levels.
【17】Acoustic source depth estimation method based on a single hydrophone in Arctic underwater
标题:北极水下基于单个水下听音器的源深度估计方法
链接:https://arxiv.org/abs/2508.07157
摘要:基于简正波和射线理论,讨论了地表声源和表层接收的特点,探讨了基于简正波和射线的深度估计方法,提出了一种基于模态频率上限的深度估计方法。通过数据验证,讨论了不同方法的适用性和局限性。对于表面折射法模波导,通过翘曲变换可以实现模式分离。基于简正波振幅随频率和个数变化的特性,通过匹配振幅信息可以估计声源深度。基于特征函数随频率的空间变化特性,提出了一种与简正波截止频率相匹配的声源深度估计方法。对于北极深海,通过对深层反演声线轨迹的分析,得到接收端声线到达结构,通过匹配射线到达时间差,可以估计声源深度。实验数据验证了声场模式和声源深度估计方法的有效性。
摘要:Based on the normal mode and ray theory, this article discusses the characteristics of surface sound source and reception at the surface layer, and explores depth estimation methods based on normal modes and rays, and proposes a depth estimation method based on the upper limit of modal frequency. Data verification is conducted to discuss the applicability and limitations of different methods. For the surface refracted normal mode waveguide, modes can be separated through warping transformation. Based on the characteristics of normal mode amplitude variation with frequency and number, the sound source depth can be estimated by matching amplitude information. Based on the spatial variation characteristics of eigenfunctions with frequency, a sound source depth estimation method matching the cutoff frequency of normal modes is proposed. For the deep Arctic sea, the sound ray arrival structure at the receiving end is obtained through the analysis of deep inversion sound ray trajectories, and the sound source depth can be estimated by matching the time difference of ray arrivals. Experimental data is used to verify the sound field patterns and the effectiveness of the sound source depth estimation method.
【18】Inversion of Arctic dual-channel sound speed profile based on random airgun signal
标题:基于随机气枪信号的北极双通道音速剖面倒置
链接:https://arxiv.org/abs/2508.07152
摘要:针对加拿大海盆和北极楚科奇高原独特的双通道声速剖面,基于折射简正波在双通道声速剖面下的传播特性,提出了一种利用折射简正波反演双通道声速剖面的方法。该方法针对双通道声速剖面的特点,提出了一种双参数的双通道声速剖面表示方法。针对双通道声速剖面下折射简正波的色散结构特征,提出了一种色散结构提取方法。将声速剖面的参数表示方法与频散结构提取方法相结合,提出了一种双通道声速剖面的反演方法。针对长距离声传播中常见的声速剖面水平变化情况,提出了一种水平变化双通道声速剖面的反演方法。最后,本文利用北极低频远距离声传播实验验证了双通道声速剖面反演方法的有效性。与以往的声速剖面反演方法相比,本文提出的方法具有反演参数少、反演速度快的优点。该方法只需单个水听器被动接收随机气枪信号即可实现,并解决了声速剖面水平变化的反演问题。它具有成本低、部署方便、计算速度快等显著优点。
摘要:For the unique dual-channel sound speed profiles of the Canadian Basin and the Chukchi Plateau in the Arctic, based on the propagation characteristics of refracted normal modes under dual-channel sound speed profiles, an inversion method using refracted normal modes for dual-channel sound speed profiles is proposed. This method proposes a dual-parameter representation method for dual-channel sound speed profiles, tailored to the characteristics of dual-channel sound speed profiles. A dispersion structure extraction method is proposed for the dispersion structure characteristics of refracted normal modes under dual-channel sound speed profiles. Combining the parameter representation method of sound speed profiles and the dispersion structure extraction method, an inversion method for dual-channel sound speed profiles is proposed. For the common horizontal variation of sound speed profiles in long-distance acoustic propagation, a method for inverting horizontally varying dual-channel sound speed profiles is proposed. Finally, this article verifies the effectiveness of the dual-channel sound speed profile inversion method using the Arctic low-frequency long-range acoustic propagation experiment. Compared with previous sound speed profile inversion methods, the method proposed in this article has the advantages of fewer inversion parameters and faster inversion speed. It can be implemented using only a single hydrophone passively receiving random air gun signals, and it also solves the inversion problem of horizontal variation of sound speed profiles. It has significant advantages such as low cost, easy deployment, and fast computation speed.
【19】SEF-MK: Speaker-Embedding-Free Voice Anonymization through Multi-k-means Quantization
标题:SEF-MK:通过多k均值量化实现无扬声器嵌入语音模拟
链接:https://arxiv.org/abs/2508.07086
备注:8 pages, 3 figures, accepted by 2025 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
摘要:语音匿名通过隐藏身份同时保留语言和非语言内容来保护说话者隐私。自监督学习(SSL)表示编码语言特征,但保留说话人特征。我们提出了一种新的扬声器嵌入自由的框架称为SEF-MK。SEF-MK不是使用在整个数据集上训练的单个k-means模型,而是通过随机选择多个k-means模型中的一个来匿名化每个话语的SSL表示,每个k-means模型都在不同的说话者子集上训练。我们从攻击者和用户的角度探讨这种方法。大量的实验表明,相比于单个k-means模型,具有多个k-means模型的SEF-MK从用户的角度更好地保留了语言和情感内容。然而,从攻击者的角度来看,利用多个k-means模型提高了隐私攻击的有效性。这些见解可以帮助用户设计语音匿名系统,以减轻攻击者的威胁。
摘要:Voice anonymization protects speaker privacy by concealing identity while preserving linguistic and paralinguistic content. Self-supervised learning (SSL) representations encode linguistic features but preserve speaker traits. We propose a novel speaker-embedding-free framework called SEF-MK. Instead of using a single k-means model trained on the entire dataset, SEF-MK anonymizes SSL representations for each utterance by randomly selecting one of multiple k-means models, each trained on a different subset of speakers. We explore this approach from both attacker and user perspectives. Extensive experiments show that, compared to a single k-means model, SEF-MK with multiple k-means models better preserves linguistic and emotional content from the user's viewpoint. However, from the attacker's perspective, utilizing multiple k-means models boosts the effectiveness of privacy attacks. These insights can aid users in designing voice anonymization systems to mitigate attacker threats.
【20】Whisfusion: Parallel ASR Decoding via a Diffusion Transformer
标题:Whisfusion:通过扩散Transformer进行并行ASB解码
链接:https://arxiv.org/abs/2508.07048
备注:16 pages, 9 figures
摘要:快速自动语音识别(ASR)对于实时字幕和会议转录等延迟敏感型应用至关重要。然而,由于自回归(AR)解码器的顺序性质和非自回归(NAR)方法的上下文限制,真正的并行ASR解码仍然具有挑战性。虽然现代ASR编码器可以一次处理长达30秒的音频,但AR解码器仍然会顺序生成令牌,从而造成延迟瓶颈。我们提出了Whisfusion,这是第一个将预训练的Whisper编码器与文本扩散解码器融合的框架。这种NAR架构通过在每个解码步骤并行处理整个声学上下文来解决AR延迟瓶颈。通过参数高效微调(PEFT)训练的轻量级交叉注意力适配器将这两种模式连接起来。我们还引入了一个批并行,多步解码策略,通过增加候选人的数量来提高准确性,同时对速度的影响最小。Whisfusion仅在LibriSpeech(960小时)上进行微调,WER低于Whisper-tiny(8.3% vs. 9.7%),并在短音频上提供相当的延迟。对于较长的话语(> 20秒),它比AR基线快2.6倍,为长形式ASR建立了一个新的高效工作点。实现和培训脚本可在https://github.com/taeyoun811/Whisfusion上获得。
摘要:Fast Automatic Speech Recognition (ASR) is critical for latency-sensitive applications such as real-time captioning and meeting transcription. However, truly parallel ASR decoding remains challenging due to the sequential nature of autoregressive (AR) decoders and the context limitations of non-autoregressive (NAR) methods. While modern ASR encoders can process up to 30 seconds of audio at once, AR decoders still generate tokens sequentially, creating a latency bottleneck. We propose Whisfusion, the first framework to fuse a pre-trained Whisper encoder with a text diffusion decoder. This NAR architecture resolves the AR latency bottleneck by processing the entire acoustic context in parallel at every decoding step. A lightweight cross-attention adapter trained via parameter-efficient fine-tuning (PEFT) bridges the two modalities. We also introduce a batch-parallel, multi-step decoding strategy that improves accuracy by increasing the number of candidates with minimal impact on speed. Fine-tuned solely on LibriSpeech (960h), Whisfusion achieves a lower WER than Whisper-tiny (8.3% vs. 9.7%), and offers comparable latency on short audio. For longer utterances (>20s), it is up to 2.6x faster than the AR baseline, establishing a new, efficient operating point for long-form ASR. The implementation and training scripts are available at https://github.com/taeyoun811/Whisfusion.
【21】Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody
标题:Maestro-EVC:由参考文献和显式韵律引导的可控情感语音转换
链接:https://arxiv.org/abs/2508.06890
备注:Accepted at ASRU 2025
摘要:情感语音转换(EVC)旨在修改语音的情感风格,同时保留其语言内容。在实际的EVC中,可控性,即使用不同参考独立控制说话者身份和情感风格的能力,是至关重要的。然而,现有的方法往往很难完全解开这些属性,并缺乏建模细粒度的情感表达,如时间动态的能力。我们提出Maestro-EVC,一个可控的EVC框架,使独立的控制内容,扬声器的身份,和情感有效地解开每个属性从单独的参考。我们进一步引入了一个时间的情感表示和一个明确的韵律建模与韵律增强鲁棒捕获和转移的时间动态的目标情感,即使在韵律不匹配的条件下。实验结果表明,Maestro-EVC实现了高质量,可控,情感表达的语音合成。
摘要:Emotional voice conversion (EVC) aims to modify the emotional style of speech while preserving its linguistic content. In practical EVC, controllability, the ability to independently control speaker identity and emotional style using distinct references, is crucial. However, existing methods often struggle to fully disentangle these attributes and lack the ability to model fine-grained emotional expressions such as temporal dynamics. We propose Maestro-EVC, a controllable EVC framework that enables independent control of content, speaker identity, and emotion by effectively disentangling each attribute from separate references. We further introduce a temporal emotion representation and an explicit prosody modeling with prosody augmentation to robustly capture and transfer the temporal dynamics of the target emotion, even under prosody-mismatched conditions. Experimental results confirm that Maestro-EVC achieves high-quality, controllable, and emotionally expressive speech synthesis.
【22】Text to Speech System for Meitei Mayek Script
标题:Meitei Mayek脚本的文本到语音系统
链接:https://arxiv.org/abs/2508.06870
摘要:本文介绍了一个文本到语音(TTS)系统的开发Manipuri语言使用Meitei Mayek脚本。利用Tacotron 2和HiFi-GAN,我们引入了一个适合支持音调语音学和资源不足的语言环境的神经TTS架构。我们为Meitei Mayek开发了一个音素映射到ARPAbet,策划了一个单扬声器数据集,并通过主观和客观指标验证了可理解和自然的语音合成。该系统为Manipuri的语言保存和技术融入奠定了基础。
摘要:This paper presents the development of a Text-to-Speech (TTS) system for the Manipuri language using the Meitei Mayek script. Leveraging Tacotron 2 and HiFi-GAN, we introduce a neural TTS architecture adapted to support tonal phonology and under-resourced linguistic environments. We develop a phoneme mapping for Meitei Mayek to ARPAbet, curate a single-speaker dataset, and demonstrate intelligible and natural speech synthesis, validated through subjective and objective metrics. This system lays the groundwork for linguistic preservation and technological inclusion of Manipuri.
【23】MMFformer: Multimodal Fusion Transformer Network for Depression Detection
标题:MMFformer:用于抑郁检测的多峰融合Transformer网络
链接:https://arxiv.org/abs/2508.06701
备注:Accepted for the 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Vienna, Austria
摘要:抑郁症是一种严重的心理健康疾病,严重影响个人的健康和生活质量,因此早期发现对于适当的护理和治疗至关重要。检测抑郁症通常很困难,因为它主要基于临床访谈期间的主观评估。因此,抑郁症的早期诊断,由于社交网络的内容,已成为一个突出的研究领域。用户生成信息的广泛性和多样性构成了一个重大挑战,限制了准确提取相关时间信息和有效融合多种模式的数据。本文介绍了MMFformer,一个多模态抑郁症检测网络,旨在检索抑郁症的时空高层次模式,从多模态社交媒体信息。具有剩余连接的Transformer网络从视频中捕获空间特征,并且利用Transformer编码器来设计音频中的重要时间动态。此外,融合架构融合提取的特征,通过后期和中期融合策略,找出最相关的多式联运相关性。最后,在两个大规模抑郁症检测数据集上对所提出的网络进行了评估,结果清楚地表明,它超越了现有的最先进的方法,D-Vlog数据集的F1分数提高了13.92%,LMVD数据集提高了7.74%。该守则可在https://github.com/rezwanh001/Large-Scale-Multimodal-Depression-Detection上公开查阅。
摘要:Depression is a serious mental health illness that significantly affects an individual's well-being and quality of life, making early detection crucial for adequate care and treatment. Detecting depression is often difficult, as it is based primarily on subjective evaluations during clinical interviews. Hence, the early diagnosis of depression, thanks to the content of social networks, has become a prominent research area. The extensive and diverse nature of user-generated information poses a significant challenge, limiting the accurate extraction of relevant temporal information and the effective fusion of data across multiple modalities. This paper introduces MMFformer, a multimodal depression detection network designed to retrieve depressive spatio-temporal high-level patterns from multimodal social media information. The transformer network with residual connections captures spatial features from videos, and a transformer encoder is exploited to design important temporal dynamics in audio. Moreover, the fusion architecture fused the extracted features through late and intermediate fusion strategies to find out the most relevant intermodal correlations among them. Finally, the proposed network is assessed on two large-scale depression detection datasets, and the results clearly reveal that it surpasses existing state-of-the-art approaches, improving the F1-Score by 13.92% for D-Vlog dataset and 7.74% for LMVD dataset. The code is made available publicly at https://github.com/rezwanh001/Large-Scale-Multimodal-Depression-Detection.
【24】AutoMashup: Automatic Music Mashups Creation
标题:AutoMashup:自动音乐Mashup创建
链接:https://arxiv.org/abs/2508.06516
备注:None
摘要:我们将介绍AutoMashup,这是一个基于源分离、音乐分析和兼容性评估的自动mashup创建系统。我们建议使用COCOLA来评估分离的词干之间的兼容性,并调查通用的预训练音频模型(CLAP和MERT)是否可以支持音轨对兼容性的zero-shot估计。我们的研究结果表明,混搭兼容性是不对称的-它取决于分配给每个轨道(人声或伴奏)的角色-并且当前的嵌入无法重现COCOLA测量的感知一致性。这些发现强调了在mashup创建中用于兼容性估计的通用音频表示的局限性。
摘要:We introduce AutoMashup, a system for automatic mashup creation based on source separation, music analysis, and compatibility estimation. We propose using COCOLA to assess compatibility between separated stems and investigate whether general-purpose pretrained audio models (CLAP and MERT) can support zero-shot estimation of track pair compatibility. Our results show that mashup compatibility is asymmetric -- it depends on the role assigned to each track (vocals or accompaniment) -- and that current embeddings fail to reproduce the perceptual coherence measured by COCOLA. These findings underline the limitations of general-purpose audio representations for compatibility estimation in mashup creation.
【25】MSU-Bench: Towards Understanding the Conversational Multi-talker Scenarios
标题:MSU长凳:了解对话式多人交谈场景
链接:https://arxiv.org/abs/2508.08155
摘要:口语理解(SLU)已经从传统的单任务方法发展到大型音频语言模型(LALM)解决方案。然而,大多数现有的语音基准集中在单个扬声器或孤立的任务,忽视了现实世界中常见的多扬声器对话所带来的挑战。我们介绍MSU-Bench,这是一个全面的基准测试,用于评估以说话者为中心的设计的多说话者会话理解。我们的分层框架包括四个渐进的层次:单说话人静态属性理解,单说话人动态属性理解,多说话人背景理解,多说话人交互理解。这种结构确保所有任务都以说话者为中心,从基本感知到多个说话者的复杂推理。通过评估MSU-Bench上最先进的模型,我们证明了随着任务复杂性在基准测试的各个层次上的增加,所有模型都表现出显着的性能下降。我们还观察到开源模型和闭源商业模型之间存在持续的能力差距,特别是在多发言人交互推理方面。这些发现验证了MSU-Bench在现实多说话人环境中评估和推进会话理解的有效性。演示可以在补充材料中找到。
摘要:Spoken Language Understanding (SLU) has progressed from traditional single-task methods to large audio language model (LALM) solutions. Yet, most existing speech benchmarks focus on single-speaker or isolated tasks, overlooking the challenges posed by multi-speaker conversations that are common in real-world scenarios. We introduce MSU-Bench, a comprehensive benchmark for evaluating multi-speaker conversational understanding with a speaker-centric design. Our hierarchical framework covers four progressive tiers: single-speaker static attribute understanding, single-speaker dynamic attribute understanding, multi-speaker background understanding, and multi-speaker interaction understanding. This structure ensures all tasks are grounded in speaker-centric contexts, from basic perception to complex reasoning across multiple speakers. By evaluating state-of-the-art models on MSU-Bench, we demonstrate that as task complexity increases across the benchmark's tiers, all models exhibit a significant performance decline. We also observe a persistent capability gap between open-source models and closed-source commercial ones, particularly in multi-speaker interaction reasoning. These findings validate the effectiveness of MSU-Bench for assessing and advancing conversational understanding in realistic multi-speaker environments. Demos can be found in the supplementary material.
【26】Auditory Intelligence: Understanding the World Through Sound
标题:听觉智力:通过声音了解世界
链接:https://arxiv.org/abs/2508.07829
备注:Position paper without experimental/quantitative validation. Not submitted to any journal/conference
摘要:听觉智能的最新进展已经产生了用于声音事件检测(SED)、声学场景分类(ASC)、自动音频字幕(AAC)和音频问答(AQA)的高性能系统。然而,这些任务在很大程度上仍然局限于表面层面的解释--捕捉发生了什么,但不捕捉原因、它意味着什么,或者它在上下文中如何展开。我提出了一个概念性的听觉智能作为一个分层的,位于过程中,包括感知,推理和互动的重构。为了实例化这一观点,我介绍了四个认知启发的任务范式ASPIRE,SODA,AUX,和AUGMENT-这些结构的听觉理解跨时频模式字幕,分层事件/场景描述,因果解释,和目标驱动的解释,分别。总之,这些范例提供了一个更可概括、更可解释、更符合人类的听觉智能的路线图,并旨在促进更广泛的讨论,即机器理解声音意味着什么。
摘要:Recent progress in auditory intelligence has yielded high-performing systems for sound event detection (SED), acoustic scene classification (ASC), automated audio captioning (AAC), and audio question answering (AQA). Yet these tasks remain largely constrained to surface-level recognition-capturing what happened but not why, what it implies, or how it unfolds in context. I propose a conceptual reframing of auditory intelligence as a layered, situated process that encompasses perception, reasoning, and interaction. To instantiate this view, I introduce four cognitively inspired task paradigms-ASPIRE, SODA, AUX, and AUGMENT-those structure auditory understanding across time-frequency pattern captioning, hierarchical event/scene description, causal explanation, and goal-driven interpretation, respectively. Together, these paradigms provide a roadmap toward more generalizable, explainable, and human-aligned auditory intelligence, and are intended to catalyze a broader discussion of what it means for machines to understand sound.
【27】Score-Informed BiLSTM Correction for Refining MIDI Velocity in Automatic Piano Transcription
标题:基于分数的BiLSTM修正,以完善自动钢琴抄写中的乐谱速度
链接:https://arxiv.org/abs/2508.07757
备注:4 pages; rejected paper by WASPAA2025
摘要:录音带是一种存储音乐的现代标准,记录音符是如何演奏的。许多钢琴演奏都有相应的在线乐谱。其中一些是由原始表演者创作的,在电子钢琴上录制音频,而另一些则是通过手动转录。近年来,自动音乐转录(AMT)迅速发展,使机器能够从音频转录音乐。然而,这些过渡往往需要进一步的纠正。假设一个完美的定时校正,我们专注于响度校正方面的响度速度(一个参数,在响度控制)。这个任务可以通过分数通知的预测速度估计来完成,这已经经历了几个发展。虽然之前的方法引入了专门构建的模型来重新估计MIDI速度,从而取代了AMT估计,但我们提出了一个BiLSTM校正模块来细化AMT估计的速度。虽然我们没有达到最先进的性能,我们验证了我们的方法在著名的AMT系统,高分辨率钢琴转录(HPT),并取得了显着的改善。
摘要:MIDI is a modern standard for storing music, recording how musical notes are played. Many piano performances have corresponding MIDI scores available online. Some of these are created by the original performer, recording on an electric piano alongside the audio, while others are through manual transcription. In recent years, automatic music transcription (AMT) has rapidly advanced, enabling machines to transcribe MIDI from audio. However, these transcriptions often require further correction. Assuming a perfect timing correction, we focus on the loudness correction in terms of MIDI velocity (a parameter in MIDI for loudness control). This task can be approached through score-informed MIDI velocity estimation, which has undergone several developments. While previous approaches introduced specifically built models to re-estimate MIDI velocity, thereby replacing AMT estimates, we propose a BiLSTM correction module to refine AMT-estimated velocity. Although we did not reach state-of-the-art performance, we validated our method on the well-known AMT system, the high-resolution piano transcription (HPT), and achieved significant improvements.
【28】Real-time CARFAC Cochlea Model Acceleration on FPGA for Underwater Acoustic Sensing Systems
标题:实时CARFAC为水下声学传感系统提供基于DSP的模型加速
链接:https://arxiv.org/abs/2508.07523
备注:5 pages, 6 figures
摘要:本文提出了一种实时,节能的嵌入式系统实现了一个阵列的级联非对称谐振器与快速压缩(CARFAC)的耳蜗模型的水下声音分析。该系统基于AMD Kria KV260模块化系统(SoM)构建,在处理器上集成了基于Rust的软件框架,用于与多个水听器输入进行实时接口和同步,并在现场可编程门阵列(FPGA)上硬件加速实现CARFAC模型,用于实时声音预处理。与以前的工作相比,CARFAC加速器实现了改进的可扩展性和处理速度,同时通过优化的时分复用,流水线设计和消除昂贵的除法电路来减少资源使用。实验结果表明,13.5%的硬件利用率为一个单一的64通道CARFAC实例和3.11 W的全板功耗时,实时处理256 kHz的输入信号。
摘要:This paper presents a real-time, energy-efficient embedded system implementing an array of Cascade of Asymmetric Resonators with Fast-Acting Compression (CARFAC) cochlea models for underwater sound analysis. Built on the AMD Kria KV260 System-on-Module (SoM), the system integrates a Rust-based software framework on the processor for real-time interfacing and synchronization with multiple hydrophone inputs, and a hardware-accelerated implementation of the CARFAC models on a Field-Programmable Gate Array (FPGA) for real-time sound pre-processing. Compared to prior work, the CARFAC accelerator achieves improved scalability and processing speed while reducing resource usage through optimized time-multiplexing, pipelined design, and elimination of costly division circuits. Experimental results demonstrate 13.5% hardware utilization for a single 64-channel CARFAC instance and a whole board power consumption of 3.11 W when processing a 256 kHz input signal in real time.
【29】FlexCTC: GPU-powered CTC Beam Decoding with advanced Contextual Abilities
标题:Flexctc:具有高级上下文能力的基于图形处理器的CTC束解码
链接:https://arxiv.org/abs/2508.07315
备注:Accepted to Automatic Speech Recognition and Understanding Workshop (ASRU) 2025
摘要:虽然波束搜索比贪婪解码提高了语音识别质量,但标准实现速度很慢,通常是顺序的,并且受到CPU的限制。为了充分利用现代硬件功能,我们提出了一种新的开源FlexCTC工具包,用于完全基于GPU的波束解码,专为连接主义时间分类(CTC)模型设计。它完全用Python和PyTorch开发,为传统的C++、CUDA或基于WFST的解码器提供了一种快速、用户友好和可扩展的替代方案。该工具包具有高性能,完全批量GPU实现,消除了CPU-GPU同步,并通过CUDA图形最大限度地减少了内核启动开销。它还支持高级上下文化技术,包括GPU驱动的N-gram语言模型融合和短语级提升。这些功能可以实现准确和高效的解码,使其适合研究和生产使用。
摘要:While beam search improves speech recognition quality over greedy decoding, standard implementations are slow, often sequential, and CPU-bound. To fully leverage modern hardware capabilities, we present a novel open-source FlexCTC toolkit for fully GPU-based beam decoding, designed for Connectionist Temporal Classification (CTC) models. Developed entirely in Python and PyTorch, it offers a fast, user-friendly, and extensible alternative to traditional C++, CUDA, or WFST-based decoders. The toolkit features a high-performance, fully batched GPU implementation with eliminated CPU-GPU synchronization and minimized kernel launch overhead via CUDA Graphs. It also supports advanced contextualization techniques, including GPU-powered N-gram language model fusion and phrase-level boosting. These features enable accurate and efficient decoding, making them suitable for both research and production use.
【30】ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
标题:ParaNoise-SV:具有语音增强和噪音提取并行联合学习的噪音稳健说话人验证的集成方法
链接:https://arxiv.org/abs/2508.07219
备注:5 pages, 3 figures, accepted to Interspeech 2025
摘要:噪声鲁棒的说话人确认利用语音增强(SE)和说话人确认(SV)的联合学习来提高鲁棒性。然而,主流的方法依赖于隐式噪声抑制,其难以将噪声与说话者特征分离,因为它们在训练期间没有明确区分噪声与语音。虽然集成SE和SV有所帮助,但它在有效处理噪声方面仍然有限。与此同时,最近的SE研究表明,明确建模噪声,而不仅仅是抑制它,增强噪声弹性。考虑到这一点,我们提出了ParaNoise-SV,双U网结合了噪声提取(NE)网络和语音增强(SE)网络。NE U-Net明确地对噪声进行建模,而SE U-Net通过并行连接在NE的指导下细化语音,保留说话者相关特征。实验结果表明,ParaNoise-SV实现了一个相对较低的8.4%的等错误率(EER)比以前的联合SE-SV模型。
摘要:Noise-robust speaker verification leverages joint learning of speech enhancement (SE) and speaker verification (SV) to improve robustness. However, prevailing approaches rely on implicit noise suppression, which struggles to separate noise from speaker characteristics as they do not explicitly distinguish noise from speech during training. Although integrating SE and SV helps, it remains limited in handling noise effectively. Meanwhile, recent SE studies suggest that explicitly modeling noise, rather than merely suppressing it, enhances noise resilience. Reflecting this, we propose ParaNoise-SV, with dual U-Nets combining a noise extraction (NE) network and a speech enhancement (SE) network. The NE U-Net explicitly models noise, while the SE U-Net refines speech with guidance from NE through parallel connections, preserving speaker-relevant features. Experimental results show that ParaNoise-SV achieves a relatively 8.4% lower equal error rate (EER) than previous joint SE-SV models.
【31】TurboBias: Universal ASR Context-Biasing powered by GPU-accelerated Phrase-Boosting Tree
标题:TurboBias:由GOP加速的短语增强树提供支持的通用ASB上下文偏置
链接:https://arxiv.org/abs/2508.07014
备注:Accepted by ASRU 2025
摘要:识别特定的关键短语是语境化自动语音识别(ASR)的一项重要任务。然而,大多数现有的上下文偏置方法具有与附加模型训练的必要性相关联的限制,显著减慢解码过程,或者约束ASR系统类型的选择。本文提出了一个通用的ASR上下文偏置框架,支持所有主要类型:CTC,传感器和注意力编码器-解码器模型。该框架基于GPU加速的单词提升树,这使其能够在浅层融合模式下用于贪婪和波束搜索解码,而不会出现明显的速度下降,即使有大量的关键短语(多达20 K个项目)。实验结果表明,该方法具有很高的效率,在准确性和解码速度上超过了开源的上下文偏置方法。我们的上下文偏置框架作为NeMo工具包的一部分是开源的。
摘要:Recognizing specific key phrases is an essential task for contextualized Automatic Speech Recognition (ASR). However, most existing context-biasing approaches have limitations associated with the necessity of additional model training, significantly slow down the decoding process, or constrain the choice of the ASR system type. This paper proposes a universal ASR context-biasing framework that supports all major types: CTC, Transducers, and Attention Encoder-Decoder models. The framework is based on a GPU-accelerated word boosting tree, which enables it to be used in shallow fusion mode for greedy and beam search decoding without noticeable speed degradation, even with a vast number of key phrases (up to 20K items). The obtained results showed high efficiency of the proposed method, surpassing the considered open-source context-biasing approaches in accuracy and decoding speed. Our context-biasing framework is open-sourced as a part of the NeMo toolkit.
【1】MSU-Bench: Towards Understanding the Conversational Multi-talker Scenarios
标题:MSU长凳:了解对话式多人交谈场景
链接:https://arxiv.org/abs/2508.08155
摘要:Spoken Language Understanding (SLU) has progressed from traditional single-task methods to large audio language model (LALM) solutions. Yet, most existing speech benchmarks focus on single-speaker or isolated tasks, overlooking the challenges posed by multi-speaker conversations that are common in real-world scenarios. We introduce MSU-Bench, a comprehensive benchmark for evaluating multi-speaker conversational understanding with a speaker-centric design. Our hierarchical framework covers four progressive tiers: single-speaker static attribute understanding, single-speaker dynamic attribute understanding, multi-speaker background understanding, and multi-speaker interaction understanding. This structure ensures all tasks are grounded in speaker-centric contexts, from basic perception to complex reasoning across multiple speakers. By evaluating state-of-the-art models on MSU-Bench, we demonstrate that as task complexity increases across the benchmark's tiers, all models exhibit a significant performance decline. We also observe a persistent capability gap between open-source models and closed-source commercial ones, particularly in multi-speaker interaction reasoning. These findings validate the effectiveness of MSU-Bench for assessing and advancing conversational understanding in realistic multi-speaker environments. Demos can be found in the supplementary material.
【2】G-IFT: A Gated Linear Unit adapter with Iterative Fine-Tuning for Low-Resource Children's Speaker Verification
标题:G-TIP:具有迭代微调功能的门控线性单元适配器,用于低资源儿童扬声器验证
链接:https://arxiv.org/abs/2508.07836
摘要:由于声学失配,在成人语音上训练的说话人验证(SV)系统在儿童的SV上通常表现不佳,并且有限的儿童语音数据使得微调不是很有效。在本文中,我们提出了一个创新的框架,一个门控线性单元适配器与迭代微调(G-IFT),以提高高资源的成人语音域和低资源的儿童语音域之间的知识转移效率。在该框架中,首先在预训练的说话人嵌入模型和分类器之间插入门控线性单元适配器。然后,分类器,适配器,和预训练的说话人嵌入模型依次优化迭代的方式。该框架对于SV系统的底层架构的类型是不可知的。我们使用OGI和MyST数据集对ECAPA-TDNN,ResNet和X-vector架构进行的实验表明,与基线方法相比,G-IFT框架在等错误率方面产生了一致的降低。
摘要:Speaker Verification (SV) systems trained on adults speech often underperform on children's SV due to the acoustic mismatch, and limited children speech data makes fine-tuning not very effective. In this paper, we propose an innovative framework, a Gated Linear Unit adapter with Iterative Fine-Tuning (G-IFT), to enhance knowledge transfer efficiency between the high-resource adults speech domain and the low-resource children's speech domain. In this framework, a Gated Linear Unit adapter is first inserted between the pre-trained speaker embedding model and the classifier. Then the classifier, adapter, and pre-trained speaker embedding model are optimized sequentially in an iterative way. This framework is agnostic to the type of the underlying architecture of the SV system. Our experiments on ECAPA-TDNN, ResNet, and X-vector architectures using the OGI and MyST datasets demonstrate that the G-IFT framework yields consistent reductions in Equal Error Rates compared to baseline methods.
【3】Auditory Intelligence: Understanding the World Through Sound
标题:听觉智力:通过声音了解世界
链接:https://arxiv.org/abs/2508.07829
摘要:听觉智能的最新进展已经产生了用于声音事件检测(SED)、声学场景分类(ASC)、自动音频字幕(AAC)和音频问答(AQA)的高性能系统。然而,这些任务在很大程度上仍然局限于表面层面的解释--捕捉发生了什么,但不捕捉原因、它意味着什么,或者它在上下文中如何展开。我提出了一个概念性的听觉智能作为一个分层的,位于过程中,包括感知,推理和互动的重构。为了实例化这一观点,我介绍了四个认知启发的任务范式ASPIRE,SODA,AUX,和AUGMENT-这些结构的听觉理解跨时频模式字幕,分层事件/场景描述,因果解释,和目标驱动的解释,分别。总之,这些范例提供了一个更可概括、更可解释、更符合人类的听觉智能的路线图,并旨在促进更广泛的讨论,即机器理解声音意味着什么。
摘要:Recent progress in auditory intelligence has yielded high-performing systems for sound event detection (SED), acoustic scene classification (ASC), automated audio captioning (AAC), and audio question answering (AQA). Yet these tasks remain largely constrained to surface-level recognition-capturing what happened but not why, what it implies, or how it unfolds in context. I propose a conceptual reframing of auditory intelligence as a layered, situated process that encompasses perception, reasoning, and interaction. To instantiate this view, I introduce four cognitively inspired task paradigms-ASPIRE, SODA, AUX, and AUGMENT-those structure auditory understanding across time-frequency pattern captioning, hierarchical event/scene description, causal explanation, and goal-driven interpretation, respectively. Together, these paradigms provide a roadmap toward more generalizable, explainable, and human-aligned auditory intelligence, and are intended to catalyze a broader discussion of what it means for machines to understand sound.
【4】Score-Informed BiLSTM Correction for Refining MIDI Velocity in Automatic Piano Transcription
标题:基于分数的BiLSTM修正,以完善自动钢琴抄写中的乐谱速度
链接:https://arxiv.org/abs/2508.07757
摘要:录音带是一种存储音乐的现代标准,记录音符是如何演奏的。许多钢琴演奏都有相应的在线乐谱。其中一些是由原始表演者创作的,在电子钢琴上录制音频,而另一些则是通过手动转录。近年来,自动音乐转录(AMT)迅速发展,使机器能够从音频转录音乐。然而,这些过渡往往需要进一步的纠正。假设一个完美的定时校正,我们专注于响度校正方面的响度速度(一个参数,在响度控制)。这个任务可以通过分数通知的预测速度估计来完成,这已经经历了几个发展。虽然以前的方法引入了专门构建的模型来重新估计平均速度,从而取代AMT估计,但我们提出了一个BiLSTM校正模块来改进AMT估计的速度。虽然我们没有达到最先进的性能,我们验证了我们的方法在著名的AMT系统,高分辨率钢琴转录(HPT),并取得了显着的改善。
摘要:MIDI is a modern standard for storing music, recording how musical notes are played. Many piano performances have corresponding MIDI scores available online. Some of these are created by the original performer, recording on an electric piano alongside the audio, while others are through manual transcription. In recent years, automatic music transcription (AMT) has rapidly advanced, enabling machines to transcribe MIDI from audio. However, these transcriptions often require further correction. Assuming a perfect timing correction, we focus on the loudness correction in terms of MIDI velocity (a parameter in MIDI for loudness control). This task can be approached through score-informed MIDI velocity estimation, which has undergone several developments. While previous approaches introduced specifically built models to re-estimate MIDI velocity, thereby replacing AMT estimates, we propose a BiLSTM correction module to refine AMT-estimated velocity. Although we did not reach state-of-the-art performance, we validated our method on the well-known AMT system, the high-resolution piano transcription (HPT), and achieved significant improvements.
【5】Is GAN Necessary for Mel-Spectrogram-based Neural Vocoder?
标题:GAN对于基于Mel频谱图的神经声码器来说是否必要?
链接:https://arxiv.org/abs/2508.07711
摘要:最近,主流的基于mel频谱图的神经声码器依赖于生成对抗网络(GAN)用于高保真语音生成,例如,HiFi-GAN和BigVGAN。然而,GAN的使用限制了训练效率和模型复杂度。因此,本文提出了一种新的自由GAN声码器,旨在回答的问题,GAN是否是必要的梅尔频谱的神经声码器。FreeGAN采用幅度-相位串行预测框架,消除了对GAN训练的需要。它集成了幅度先验输入、SNAKE-ConvNeXt v2主干和频率加权反缠绕相位损耗,以补偿因缺少GAN而导致的性能损失。实验结果证实,FreeGAN的语音质量与先进的基于GAN的声码器相当,同时显着提高了训练效率和复杂度。利用我们提出的方法,其他基于显式相位预测的神经声码器也可以在没有GAN的情况下工作。
摘要:Recently, mainstream mel-spectrogram-based neural vocoders rely on generative adversarial network (GAN) for high-fidelity speech generation, e.g., HiFi-GAN and BigVGAN. However, the use of GAN restricts training efficiency and model complexity. Therefore, this paper proposes a novel FreeGAN vocoder, aiming to answer the question of whether GAN is necessary for mel-spectrogram-based neural vocoders. The FreeGAN employs an amplitude-phase serial prediction framework, eliminating the need for GAN training. It incorporates amplitude prior input, SNAKE-ConvNeXt v2 backbone and frequency-weighted anti-wrapping phase loss to compensate for the performance loss caused by the absence of GAN. Experimental results confirm that the speech quality of FreeGAN is comparable to that of advanced GAN-based vocoders, while significantly improving training efficiency and complexity. Other explicit-phase-prediction-based neural vocoders can also work without GAN, leveraging our proposed methods.
【6】UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling
标题:UniFlow:通过连续生成建模统一语音前端任务
链接:https://arxiv.org/abs/2508.07558
摘要:生成建模最近在图像、视频和音频领域取得了显著的成功,展示了统一表示学习的强大功能。然而,语音前端任务,如语音增强(SE),目标说话人提取(TSE),声学回声消除(AEC)和语言查询源分离(LASS)仍然主要是由不同的,特定于任务的解决方案。这种碎片化会导致冗余的工程工作、不一致的性能和有限的可扩展性。为了解决这一差距,我们引入了UniFlow,这是一个统一的框架,它采用连续生成建模来处理共享潜在空间中的各种语音前端任务。具体而言,UniFlow利用波形变分自动编码器(VAE)来学习原始音频的紧凑潜在表示,再加上预测潜在更新的扩散Transformer(DiT)。为了在训练过程中区分语音处理任务,采用由任务ID索引的可学习条件嵌入来实现最大参数共享,同时保持特定于任务的适应性。为了平衡模型性能和计算效率,我们研究并比较了三个生成目标:去噪扩散,流量匹配和潜在域内的平均流量。我们在多个公共基准测试中验证了UniFlow,证明了其在最先进的基线上的一致收益。UniFlow的统一潜在公式和条件设计使其易于扩展到新任务,为构建和扩展生成语音处理管道提供了集成基础。为了促进未来的研究,我们将开源我们的代码库。
摘要:Generative modeling has recently achieved remarkable success across image, video, and audio domains, demonstrating powerful capabilities for unified representation learning. Yet speech front-end tasks such as speech enhancement (SE), target speaker extraction (TSE), acoustic echo cancellation (AEC), and language-queried source separation (LASS) remain largely tackled by disparate, task-specific solutions. This fragmentation leads to redundant engineering effort, inconsistent performance, and limited extensibility. To address this gap, we introduce UniFlow, a unified framework that employs continuous generative modeling to tackle diverse speech front-end tasks in a shared latent space. Specifically, UniFlow utilizes a waveform variational autoencoder (VAE) to learn a compact latent representation of raw audio, coupled with a Diffusion Transformer (DiT) that predicts latent updates. To differentiate the speech processing task during the training, learnable condition embeddings indexed by a task ID are employed to enable maximal parameter sharing while preserving task-specific adaptability. To balance model performance and computational efficiency, we investigate and compare three generative objectives: denoising diffusion, flow matching, and mean flow within the latent domain. We validate UniFlow on multiple public benchmarks, demonstrating consistent gains over state-of-the-art baselines. UniFlow's unified latent formulation and conditional design make it readily extensible to new tasks, providing an integrated foundation for building and scaling generative speech processing pipelines. To foster future research, we will open-source our codebase.
【7】Real-time CARFAC Cochlea Model Acceleration on FPGA for Underwater Acoustic Sensing Systems
标题:实时CARFAC为水下声学传感系统提供基于DSP的模型加速
链接:https://arxiv.org/abs/2508.07523
摘要:本文提出了一种实时,节能的嵌入式系统实现了一个阵列的级联非对称谐振器与快速压缩(CARFAC)的耳蜗模型的水下声音分析。该系统基于AMD Kria KV260模块化系统(SoM)构建,在处理器上集成了基于Rust的软件框架,用于与多个水听器输入进行实时接口和同步,并在现场可编程门阵列(FPGA)上硬件加速实现CARFAC模型,用于实时声音预处理。与以前的工作相比,CARFAC加速器实现了改进的可扩展性和处理速度,同时通过优化的时分复用,流水线设计和消除昂贵的除法电路来减少资源使用。实验结果表明,13.5%的硬件利用率为一个单一的64通道CARFAC实例和3.11 W的全板功耗时,实时处理256 kHz的输入信号。
摘要:This paper presents a real-time, energy-efficient embedded system implementing an array of Cascade of Asymmetric Resonators with Fast-Acting Compression (CARFAC) cochlea models for underwater sound analysis. Built on the AMD Kria KV260 System-on-Module (SoM), the system integrates a Rust-based software framework on the processor for real-time interfacing and synchronization with multiple hydrophone inputs, and a hardware-accelerated implementation of the CARFAC models on a Field-Programmable Gate Array (FPGA) for real-time sound pre-processing. Compared to prior work, the CARFAC accelerator achieves improved scalability and processing speed while reducing resource usage through optimized time-multiplexing, pipelined design, and elimination of costly division circuits. Experimental results demonstrate 13.5% hardware utilization for a single 64-channel CARFAC instance and a whole board power consumption of 3.11 W when processing a 256 kHz input signal in real time.
【8】Scalable Controllable Accented TTS
标题:可扩展可控制可强调的TTC
链接:https://arxiv.org/abs/2508.07426
摘要:我们解决了扩展重音TTS系统的挑战,扩展了它们的功能,包括更大量的训练数据和更广泛的重音标签,即使是在传统的TTS数据集中表现不佳或未标记的重音。为了实现这一目标,我们采用两种策略:1。通过语音地理定位模型进行口音标签发现,该模型自动从原始语音数据推断口音标签,而不仅仅依赖于人类注释; 2.通过kNN语音转换增强音色,以增加数据多样性和模型鲁棒性。这些策略在CommonVoice上得到了验证,在那里,我们对XTTS-v2进行了微调,以适应带有口音的TTS,并使用地理位置发现或增强口音标签。我们证明,由此产生的重音TTS模型不仅优于XTTS-v2微调自我报告的口音标签在CommonVoice,但也现有的重音TTS基准。
摘要:We tackle the challenge of scaling accented TTS systems, expanding their capabilities to include much larger amounts of training data and a wider variety of accent labels, even for accents that are poorly represented or unlabeled in traditional TTS datasets. To achieve this, we employ two strategies: 1. Accent label discovery via a speech geolocation model, which automatically infers accent labels from raw speech data without relying solely on human annotation; 2. Timbre augmentation through kNN voice conversion to increase data diversity and model robustness. These strategies are validated on CommonVoice, where we fine-tune XTTS-v2 for accented TTS with accent labels discovered or enhanced using geolocation. We demonstrate that the resulting accented TTS model not only outperforms XTTS-v2 fine-tuned on self-reported accent labels in CommonVoice, but also existing accented TTS benchmarks.
【9】KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features
标题:KLASSify将验证:使用基于SSL的音频和手工制作的视觉特征进行视听Deepfake检测
链接:https://arxiv.org/abs/2508.07337
摘要:音频驱动的说话头部生成器和高级文本到语音(TTS)模型的快速发展导致了更复杂的时间深度伪造。这些进展强调了对能够检测和定位deepfake的强大方法的需求,即使在新的、看不见的攻击场景下也是如此。目前最先进的deepfake检测器虽然准确,但通常计算成本很高,并且很难推广到新的操作技术。为了应对这些挑战,我们提出了针对AV-Deepfake 1 M 2025挑战的多模式方法。对于视觉模态,我们利用手工制作的功能来提高可解释性和适应性。对于音频模态,我们采用自监督学习(SSL)骨干与图形注意力网络相结合,以捕获丰富的音频表示,提高检测鲁棒性。我们的方法在性能和实际部署之间取得了平衡,专注于弹性和潜在的可解释性。在AV-Deepfake 1 M ++数据集上,我们的多模态系统在deepfake分类任务中实现了92.78%的AUC,在仅使用音频模态的时间定位中实现了0.3536的IoU。
摘要:The rapid development of audio-driven talking head generators and advanced Text-To-Speech (TTS) models has led to more sophisticated temporal deepfakes. These advances highlight the need for robust methods capable of detecting and localizing deepfakes, even under novel, unseen attack scenarios. Current state-of-the-art deepfake detectors, while accurate, are often computationally expensive and struggle to generalize to novel manipulation techniques. To address these challenges, we propose multimodal approaches for the AV-Deepfake1M 2025 challenge. For the visual modality, we leverage handcrafted features to improve interpretability and adaptability. For the audio modality, we adapt a self-supervised learning (SSL) backbone coupled with graph attention networks to capture rich audio representations, improving detection robustness. Our approach strikes a balance between performance and real-world deployment, focusing on resilience and potential interpretability. On the AV-Deepfake1M++ dataset, our multimodal system achieves AUC of 92.78% for deepfake classification task and IoU of 0.3536 for temporal localization using only the audio modality.
【10】FlexCTC: GPU-powered CTC Beam Decoding with advanced Contextual Abilities
标题:Flexctc:具有高级上下文能力的基于图形处理器的CTC束解码
链接:https://arxiv.org/abs/2508.07315
摘要:虽然波束搜索比贪婪解码提高了语音识别质量,但标准实现速度很慢,通常是顺序的,并且受到CPU的限制。为了充分利用现代硬件功能,我们提出了一种新的开源FlexCTC工具包,用于完全基于GPU的波束解码,专为连接主义时间分类(CTC)模型设计。它完全用Python和PyTorch开发,为传统的C++、CUDA或基于WFST的解码器提供了一种快速、用户友好和可扩展的替代方案。该工具包具有高性能,完全批量GPU实现,消除了CPU-GPU同步,并通过CUDA图形最大限度地减少了内核启动开销。它还支持高级上下文化技术,包括GPU驱动的N-gram语言模型融合和短语级提升。这些功能可以实现准确和高效的解码,使其适合研究和生产使用。
摘要:While beam search improves speech recognition quality over greedy decoding, standard implementations are slow, often sequential, and CPU-bound. To fully leverage modern hardware capabilities, we present a novel open-source FlexCTC toolkit for fully GPU-based beam decoding, designed for Connectionist Temporal Classification (CTC) models. Developed entirely in Python and PyTorch, it offers a fast, user-friendly, and extensible alternative to traditional C++, CUDA, or WFST-based decoders. The toolkit features a high-performance, fully batched GPU implementation with eliminated CPU-GPU synchronization and minimized kernel launch overhead via CUDA Graphs. It also supports advanced contextualization techniques, including GPU-powered N-gram language model fusion and phrase-level boosting. These features enable accurate and efficient decoding, making them suitable for both research and production use.
【11】XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation
标题:XSYS RAG:使用检索增强生成的强度可控的跨语言情感转移
链接:https://arxiv.org/abs/2508.07302
摘要:Zero-shot emotion transfer in cross-lingual speech synthesis refers to generating speech in a target language, where the emotion is expressed based on reference speech from a different source language.However, this task remains challenging due to the scarcity of parallel multilingual emotional corpora, the presence of foreign accent artifacts, and the difficulty of separating emotion from language-specific prosodic features.In this paper, we propose XEmoRAG, a novel framework to enable zero-shot emotion transfer from Chinese to Thai using a large language model (LLM)-based model, without relying on parallel emotional data.XEmoRAG extracts language-agnostic emotional embeddings from Chinese speech and retrieves emotionally matched Thai utterances from a curated emotional database, enabling controllable emotion transfer without explicit emotion labels. Additionally, a flow-matching alignment module minimizes pitch and duration mismatches, ensuring natural prosody. It also blends Chinese timbre into the Thai synthesis, enhancing rhythmic accuracy and emotional expression, while preserving speaker characteristics and emotional consistency.Experimental results show that XEmoRAG synthesizes expressive and natural Thai speech using only Chinese reference audio, without requiring explicit emotion labels.These results highlight XEmoRAG's capability to achieve flexible and low-resource emotional transfer across languages.Our demo is available at https://tlzuo-lesley.github.io/Demo-page/.
【12】A Survey on Non-Intrusive ASR Refinement: From Output-Level Correction to Full-Model Distillation
标题:非侵入性ASB精制研究:从输出级修正到全模型蒸馏
链接:https://arxiv.org/abs/2508.07285
摘要:Automatic Speech Recognition (ASR) has become an integral component of modern technology, powering applications such as voice-activated assistants, transcription services, and accessibility tools. Yet ASR systems continue to struggle with the inherent variability of human speech, such as accents, dialects, and speaking styles, as well as environmental interference, including background noise. Moreover, domain-specific conversations often employ specialized terminology, which can exacerbate transcription errors. These shortcomings not only degrade raw ASR accuracy but also propagate mistakes through subsequent natural language processing pipelines. Because redesigning an ASR model is costly and time-consuming, non-intrusive refinement techniques that leave the model's architecture unchanged have become increasingly popular. In this survey, we systematically review current non-intrusive refinement approaches and group them into five classes: fusion, re-scoring, correction, distillation, and training adjustment. For each class, we outline the main methods, advantages, drawbacks, and ideal application scenarios. Beyond method classification, this work surveys adaptation techniques aimed at refining ASR in domain-specific contexts, reviews commonly used evaluation datasets along with their construction processes, and proposes a standardized set of metrics to facilitate fair comparisons. Finally, we identify open research gaps and suggest promising directions for future work. By providing this structured overview, we aim to equip researchers and practitioners with a clear foundation for developing more robust, accurate ASR refinement pipelines.
【13】Lessons Learnt: Revisit Key Training Strategies for Effective Speech Emotion Recognition in the Wild
标题:吸取的教训:重温野外有效语音情感识别的关键训练策略
链接:https://arxiv.org/abs/2508.07282
摘要:在这项研究中,我们重新审视了机器学习中经常被忽视的关键训练策略,以支持更深入的架构。具体来说,我们探索平衡策略,激活功能和微调技术,以提高语音情感识别(SER)的自然条件。我们的研究结果表明,简单的修改提高泛化与最小的架构变化。我们的多模态融合模型,整合了这些优化,达到了0.6953的效价CCC,这是任务2:情感属性回归中的最佳效价分数。值得注意的是,在单模态设置中分别微调RoBERTa和WavLM,然后在不训练主干提取器的情况下进行特征融合,产生最高的效价性能。此外,焦点丢失和激活功能在不增加复杂性的情况下显著提高了性能。这些结果表明,细化核心组件,而不是深化模型,导致更强大的SER在野外。
摘要:In this study, we revisit key training strategies in machine learning often overlooked in favor of deeper architectures. Specifically, we explore balancing strategies, activation functions, and fine-tuning techniques to enhance speech emotion recognition (SER) in naturalistic conditions. Our findings show that simple modifications improve generalization with minimal architectural changes. Our multi-modal fusion model, integrating these optimizations, achieves a valence CCC of 0.6953, the best valence score in Task 2: Emotional Attribute Regression. Notably, fine-tuning RoBERTa and WavLM separately in a single-modality setting, followed by feature fusion without training the backbone extractor, yields the highest valence performance. Additionally, focal loss and activation functions significantly enhance performance without increasing complexity. These results suggest that refining core components, rather than deepening models, leads to more robust SER in-the-wild.
【14】ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
标题:ParaNoise-SV:具有语音增强和噪音提取并行联合学习的噪音稳健说话人验证的集成方法
链接:https://arxiv.org/abs/2508.07219
摘要:噪声鲁棒的说话人确认利用语音增强(SE)和说话人确认(SV)的联合学习来提高鲁棒性。然而,主流的方法依赖于隐式噪声抑制,其难以将噪声与说话者特征分离,因为它们在训练期间没有明确区分噪声与语音。虽然集成SE和SV有所帮助,但它在有效处理噪声方面仍然有限。与此同时,最近的SE研究表明,明确建模噪声,而不仅仅是抑制它,增强噪声弹性。考虑到这一点,我们提出了ParaNoise-SV,双U网结合了噪声提取(NE)网络和语音增强(SE)网络。NE U-Net明确地对噪声进行建模,而SE U-Net通过并行连接在NE的指导下细化语音,保留说话者相关特征。实验结果表明,ParaNoise-SV实现了一个相对较低的8.4%的等错误率(EER)比以前的联合SE-SV模型。
摘要:Noise-robust speaker verification leverages joint learning of speech enhancement (SE) and speaker verification (SV) to improve robustness. However, prevailing approaches rely on implicit noise suppression, which struggles to separate noise from speaker characteristics as they do not explicitly distinguish noise from speech during training. Although integrating SE and SV helps, it remains limited in handling noise effectively. Meanwhile, recent SE studies suggest that explicitly modeling noise, rather than merely suppressing it, enhances noise resilience. Reflecting this, we propose ParaNoise-SV, with dual U-Nets combining a noise extraction (NE) network and a speech enhancement (SE) network. The NE U-Net explicitly models noise, while the SE U-Net refines speech with guidance from NE through parallel connections, preserving speaker-relevant features. Experimental results show that ParaNoise-SV achieves a relatively 8.4% lower equal error rate (EER) than previous joint SE-SV models.
【15】TurboBias: Universal ASR Context-Biasing powered by GPU-accelerated Phrase-Boosting Tree
标题:TurboBias:由GOP加速的短语增强树提供支持的通用ASB上下文偏置
链接:https://arxiv.org/abs/2508.07014
摘要:识别特定的关键短语是语境化自动语音识别(ASR)的一项重要任务。然而,大多数现有的上下文偏置方法具有与附加模型训练的必要性相关联的限制,显著减慢解码过程,或者约束ASR系统类型的选择。本文提出了一个通用的ASR上下文偏置框架,支持所有主要类型:CTC,传感器和注意力编码器-解码器模型。该框架基于GPU加速的单词提升树,这使其能够在浅层融合模式下用于贪婪和波束搜索解码,而不会出现明显的速度下降,即使有大量的关键短语(多达20 K个项目)。实验结果表明,该方法具有很高的效率,在准确性和解码速度上超过了开源的上下文偏置方法。我们的上下文偏置框架作为NeMo工具包的一部分是开源的。
摘要:Recognizing specific key phrases is an essential task for contextualized Automatic Speech Recognition (ASR). However, most existing context-biasing approaches have limitations associated with the necessity of additional model training, significantly slow down the decoding process, or constrain the choice of the ASR system type. This paper proposes a universal ASR context-biasing framework that supports all major types: CTC, Transducers, and Attention Encoder-Decoder models. The framework is based on a GPU-accelerated word boosting tree, which enables it to be used in shallow fusion mode for greedy and beam search decoding without noticeable speed degradation, even with a vast number of key phrases (up to 20K items). The obtained results showed high efficiency of the proposed method, surpassing the considered open-source context-biasing approaches in accuracy and decoding speed. Our context-biasing framework is open-sourced as a part of the NeMo toolkit.
【16】Head-steered channel selection method for hearing aid applications using remote microphones
标题:使用远程麦克风的助听器应用的头控通道选择方法
链接:https://arxiv.org/abs/2508.06928
摘要:我们提出了一种使用远程麦克风的助听器应用的通道选择方法,在存在多个竞争的谈话者。所提出的信道选择方法使用助听器用户的头部转向方向来识别源自助听器用户的正面方向的远程信道,其捕获目标说话者信号。我们提出的信道选择任务作为一个多假设检验问题,并得出一个最大似然解。在现实的、简化的假设下,该解决方案选择与头部操纵助听器波束形成器的输出具有最高加权平方绝对相关系数的远程信道。我们分析了所提出的信道选择方法,使用近距离说话的远程麦克风和表麦克风阵列的性能。通过模拟使用现实的声学场景,我们表明,所提出的信道选择方法始终优于现有的方法,在准确地找到捕获目标说话者信号的远程信道,在存在多个竞争的说话者,而不使用任何额外的传感器。
摘要:We propose a channel selection method for hearing aid applications using remote microphones, in the presence of multiple competing talkers. The proposed channel selection method uses the hearing aid user's head-steering direction to identify the remote channel originating from the frontal direction of the hearing aid user, which captures the target talker signal. We pose the channel selection task as a multiple hypothesis testing problem, and derive a maximum likelihood solution. Under realistic, simplifying assumptions, the solution selects the remote channel which has the highest weighted squared absolute correlation coefficient with the output of the head-steered hearing aid beamformer. We analyze the performance of the proposed channel selection method using close-talking remote microphones and table microphone arrays. Through simulations using realistic acoustic scenes, we show that the proposed channel selection method consistently outperforms existing methods in accurately finding the remote channel that captures the target talker signal, in the presence of multiple competing talkers, without the use of any additional sensors.
【17】Speech Enhancement based on cascaded two flow
标题:基于级联两流的语音增强
链接:https://arxiv.org/abs/2508.06842
摘要:基于扩散概率模型的语音增强(SE)表现出令人印象深刻的性能,同时需要相对较高数量的函数评估(NFE)。最近,已经提出了基于流匹配的SE,其显示出具有小NFE的竞争性能。早期的方法采用嘈杂的语音作为唯一的条件变量。已经有其他方法,其利用用预测模型增强的语音作为另一个条件变量并对初始值进行采样,但是它们需要在生成SE模型之上的单独的预测模型。在这项工作中,我们建议采用一个相同的模型,基于流匹配的SE和生成增强的语音作为初始起点和条件变量。实验结果表明,所提出的方法需要相同或更少的NFE,即使与两个级联的生成方法,同时实现等效或更好的性能,以前的基线。
摘要:Speech enhancement (SE) based on diffusion probabilistic models has exhibited impressive performance, while requiring a relatively high number of function evaluations (NFE). Recently, SE based on flow matching has been proposed, which showed competitive performance with a small NFE. Early approaches adopted the noisy speech as the only conditioning variable. There have been other approaches which utilize speech enhanced with a predictive model as another conditioning variable and to sample an initial value, but they require a separate predictive model on top of the generative SE model. In this work, we propose to employ an identical model based on flow matching for both SE and generating enhanced speech used as an initial starting point and a conditioning variable. Experimental results showed that the proposed method required the same or fewer NFEs even with two cascaded generative methods while achieving equivalent or better performances to the previous baselines.
【18】FlowSE: Flow Matching-based Speech Enhancement
标题:FlowSE:基于流匹配的语音增强
链接:https://arxiv.org/abs/2508.06840
摘要:扩散概率模型在语音增强方面表现出了令人印象深刻的性能,但它们通常需要在推理阶段进行25到60次函数评估,从而导致计算复杂度很高。最近,提出了一种微调方法来纠正反向过程,这大大降低了函数评估(NFE)的数量。流匹配是一种训练连续归一化流的方法,该方法对从已知分布到未知分布的概率路径进行建模,包括由扩散过程描述的概率路径。本文提出了一种基于条件流匹配的语音增强方法。所提出的方法实现的性能与基于扩散的语音增强与NFE为60时,NFE为5,并显示出类似的性能与扩散模型校正的逆过程在相同的NFE从1到5,而无需额外的微调过程。我们还表明,相应的扩散模型来自条件概率路径与修改后的最佳运输条件向量场表现出类似的性能与NFE 5没有任何微调程序。
摘要:Diffusion probabilistic models have shown impressive performance for speech enhancement, but they typically require 25 to 60 function evaluations in the inference phase, resulting in heavy computational complexity. Recently, a fine-tuning method was proposed to correct the reverse process, which significantly lowered the number of function evaluations (NFE). Flow matching is a method to train continuous normalizing flows which model probability paths from known distributions to unknown distributions including those described by diffusion processes. In this paper, we propose a speech enhancement based on conditional flow matching. The proposed method achieved the performance comparable to those for the diffusion-based speech enhancement with the NFE of 60 when the NFE was 5, and showed similar performance with the diffusion model correcting the reverse process at the same NFE from 1 to 5 without additional fine tuning procedure. We also have shown that the corresponding diffusion model derived from the conditional probability path with a modified optimal transport conditional vector field demonstrated similar performances with the NFE of 5 without any fine-tuning procedure.
【19】Differentiable Grouped Feedback Delay Networks for Learning Coupled Volume Acoustics
标题:用于学习耦合体声学的可区分分组反馈延迟网络
链接:https://arxiv.org/abs/2508.06686
摘要:Rendering dynamic reverberation in a complicated acoustic space for moving sources and listeners is challenging but crucial for enhancing user immersion in extended-reality (XR) applications. Capturing spatially varying room impulse responses (RIRs) is costly and often impractical. Moreover, dynamic convolution with measured RIRs is computationally expensive with high memory demands, typically not available on wearable computing devices. Grouped Feedback Delay Networks (GFDNs), on the other hand, allow efficient rendering of coupled room acoustics. However, its parameters need to be tuned to match the reverberation profile of a coupled space. In this work, we propose the concept of Differentiable GFDNs (DiffGFDNs), which have tunable parameters that are optimised to match the late reverberation profile of a set of RIRs captured from a space that exhibits multi-slope decay. Once trained on a finite set of measurements, the DiffGFDN generalises to unmeasured locations in the space. We propose a parallel processing pipeline that has multiple DiffGFDNs with frequency-independent parameters processing each octave band. The parameters of the DiffGFDN can be updated rapidly during inferencing as sources and listeners move. We evaluate the proposed architecture against the Common Slopes (CS) model on a dataset of RIRs for three coupled rooms. The proposed architecture generates multi-slope late reverberation with low memory and computational requirements, achieving better energy decay relief (EDR) error and slightly worse octave-band energy decay curve (EDC) errors compared to the CS model. Furthermore, DiffGFDN requires an order of magnitude fewer floating-point operations per sample than the CS renderer.
【20】Pindrop it! Audio and Visual Deepfake Countermeasures for Robust Detection and Fine Grained-Localization
标题:别碰它!用于鲁棒检测和细粒度定位的音频和视觉Deepfake对策
链接:https://arxiv.org/abs/2508.08141
摘要:The field of visual and audio generation is burgeoning with new state-of-the-art methods. This rapid proliferation of new techniques underscores the need for robust solutions for detecting synthetic content in videos. In particular, when fine-grained alterations via localized manipulations are performed in visual, audio, or both domains, these subtle modifications add challenges to the detection algorithms. This paper presents solutions for the problems of deepfake video classification and localization. The methods were submitted to the ACM 1M Deepfakes Detection Challenge, achieving the best performance in the temporal localization task and a top four ranking in the classification task for the TestA split of the evaluation dataset.
【21】Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
标题:为结构障碍性语音识别弥合ASC和LLM:自我监督和生成方法的基准
链接:https://arxiv.org/abs/2508.08027
摘要:Speech Recognition (ASR) due to phoneme distortions and high variability. While self-supervised ASR models like Wav2Vec, HuBERT, and Whisper have shown promise, their effectiveness in dysarthric speech remains unclear. This study systematically benchmarks these models with different decoding strategies, including CTC, seq2seq, and LLM-enhanced decoding (BART,GPT-2, Vicuna). Our contributions include (1) benchmarking ASR architectures for dysarthric speech, (2) introducing LLM-based decoding to improve intelligibility, (3) analyzing generalization across datasets, and (4) providing insights into recognition errors across severity levels. Findings highlight that LLM-enhanced decoding improves dysarthric ASR by leveraging linguistic constraints for phoneme restoration and grammatical correction.
【22】Exploring Procedural Data Generation for Automatic Acoustic Guitar Fingerpicking Transcription
标题:探索自动声学吉他指纹识别转录的程序数据生成
链接:https://arxiv.org/abs/2508.07987
摘要:原声吉他指弹演奏的自动转录仍然是一项具有挑战性的任务,由于缺乏标记的训练数据和与音乐录音相关的法律约束。这项工作研究了一个程序化的数据生成管道,作为训练转录模型的真实音频记录的替代方案。我们的方法通过四个阶段来合成训练数据:基于知识的指法谱组成,可调性能渲染,使用扩展的Karplus-Strong算法进行物理建模,以及包括混响和失真在内的音频增强。我们训练和评估基于CRNN的笔记跟踪模型的真实和合成数据集,证明程序数据可以用来实现合理的笔记跟踪结果。使用少量的真实数据进行微调进一步提高了转录的准确性,比只在真实记录上训练的模型有所改进。这些结果突出了程序生成的音频数据稀缺的音乐信息检索任务的潜力。
摘要:Automatic transcription of acoustic guitar fingerpicking performances remains a challenging task due to the scarcity of labeled training data and legal constraints connected with musical recordings. This work investigates a procedural data generation pipeline as an alternative to real audio recordings for training transcription models. Our approach synthesizes training data through four stages: knowledge-based fingerpicking tablature composition, MIDI performance rendering, physical modeling using an extended Karplus-Strong algorithm, and audio augmentation including reverb and distortion. We train and evaluate a CRNN-based note-tracking model on both real and synthetic datasets, demonstrating that procedural data can be used to achieve reasonable note-tracking results. Finetuning with a small amount of real data further enhances transcription accuracy, improving over models trained exclusively on real recordings. These results highlight the potential of procedurally generated audio for data-scarce music information retrieval tasks.
【23】Joint Transcription of Acoustic Guitar Strumming Directions and Chords
标题:原声吉他弹奏方向和和弦的联合抄写
链接:https://arxiv.org/abs/2508.07973
摘要:吉他弹奏的自动转录是音乐信息检索(MIR)中一个代表性不足且具有挑战性的任务,特别是用于从音频信号中提取弹奏方向和和弦进行。虽然现有的方法显示出希望,但它们的有效性往往受到有限数据集的阻碍。在这项工作中,我们通过引入一个新的数据集和一个基于深度学习的转录模型,将多模式方法扩展到吉他弹奏转录。我们使用ESP32智能手表运动传感器和结构化录音协议收集了90分钟的真实吉他录音,并辅以4小时标记的弹奏音频的合成数据集。卷积递归神经网络(CRNN)模型被训练用于检测弹奏事件,对其方向进行分类,并仅使用麦克风音频识别相应的和弦。我们的评估表明,与基线发作检测算法相比,该算法有了显着的改进,结合合成和真实数据的混合方法在弹奏动作检测和和弦分类方面都达到了最高的准确度。这些结果突出了深度学习在强大的吉他弹奏转录方面的潜力,并为自动节奏吉他分析开辟了新的途径。
摘要:Automatic transcription of guitar strumming is an underrepresented and challenging task in Music Information Retrieval (MIR), particularly for extracting both strumming directions and chord progressions from audio signals. While existing methods show promise, their effectiveness is often hindered by limited datasets. In this work, we extend a multimodal approach to guitar strumming transcription by introducing a novel dataset and a deep learning-based transcription model. We collect 90 min of real-world guitar recordings using an ESP32 smartwatch motion sensor and a structured recording protocol, complemented by a synthetic dataset of 4h of labeled strumming audio. A Convolutional Recurrent Neural Network (CRNN) model is trained to detect strumming events, classify their direction, and identify the corresponding chords using only microphone audio. Our evaluation demonstrates significant improvements over baseline onset detection algorithms, with a hybrid method combining synthetic and real-world data achieving the highest accuracy for both strumming action detection and chord classification. These results highlight the potential of deep learning for robust guitar strumming transcription and open new avenues for automatic rhythm guitar analysis.
【24】Filling MIDI Velocity using U-Net Image Colorizer
标题:使用U-Net图像着色器填充收件箱速度
链接:https://arxiv.org/abs/2508.07751
摘要:现代音乐制作者通常使用乐器数字接口(Musical Instrument Digital Interface)来存储他们的音乐作品。然而,使用数字软件创建的MIDI文件可能缺乏人类表演的表现力特征,基本上留下了速度参数-音符响度的控制-未定义,默认值为平坦值。填充音频速度的任务被称为音频速度预测,其使用回归模型通过仅调整该参数来增强音乐表现力。在本文中,我们介绍了U-Net,一个广泛采用的架构,在图像彩色化,这项任务。通过将MIDI数据概念化为图像,我们采用窗口注意力并开发了一个自定义损失函数来解决MIDI转换图像的稀疏性。目前的数据集可用性限制了我们的实验钢琴数据。在MAESTRO v3和SMD数据集上进行评估,我们提出的填充速度的方法在定量指标和定性听力测试方面都优于以前的方法。
摘要:Modern music producers commonly use MIDI (Musical Instrument Digital Interface) to store their musical compositions. However, MIDI files created with digital software may lack the expressive characteristics of human performances, essentially leaving the velocity parameter - a control for note loudness - undefined, which defaults to a flat value. The task of filling MIDI velocity is termed MIDI velocity prediction, which uses regression models to enhance music expressiveness by adjusting only this parameter. In this paper, we introduce the U-Net, a widely adopted architecture in image colorization, to this task. By conceptualizing MIDI data as images, we adopt window attention and develop a custom loss function to address the sparsity of MIDI-converted images. Current dataset availability restricts our experiments to piano data. Evaluated on the MAESTRO v3 and SMD datasets, our proposed method for filling MIDI velocity outperforms previous approaches in both quantitative metrics and qualitative listening tests.
【25】Think Before You Talk: Enhancing Meaningful Dialogue Generation in Full-Duplex Speech Language Models with Planning-Inspired Text Guidance
标题:说话前先思考:利用受规划启发的文本指导增强全复式语音语言模型中有意义的对话生成
链接:https://arxiv.org/abs/2508.07375
摘要:Full-Duplex Speech Language Models(FD-SLM)是专门的基础模型,旨在通过对诸如中断、反向通道和重叠语音等复杂会话动态进行建模来实现自然、实时的口语交互,而端到端(e2 e)FD-SLM利用真实世界的双通道会话数据来捕捉细微差别的两个说话者对话模式,以实现类似人类的交互。然而,他们面临着一个关键的挑战-由于长时间的语音序列和有限的高质量口语对话数据,与纯文本对话相比,他们的会话能力往往会下降。虽然文本引导的语音生成可以缓解这些问题,但当将文本引导集成到双通道音频流中时,它会遇到时间和长度问题,从而破坏自然交互所必需的精确时间对齐。为了解决这些挑战,我们提出了TurnGuide,这是一种新颖的规划启发方法,它通过动态地将助理语音分割成对话回合并在语音输出之前生成回合级文本指导来模仿人类会话规划,从而有效地解决了插入时间和长度的挑战。大量的实验表明,我们的方法显着提高e2 e FD-SLM的会话能力,使他们能够生成语义有意义的和连贯的语音,同时保持自然的会话流。演示可在https://dreamtheater123.github.io/TurnGuide-Demo/上获得。代码将在https://github.com/dreamtheater123/TurnGuide上提供。
摘要:Full-Duplex Speech Language Models (FD-SLMs) are specialized foundation models designed to enable natural, real-time spoken interactions by modeling complex conversational dynamics such as interruptions, backchannels, and overlapping speech, and End-to-end (e2e) FD-SLMs leverage real-world double-channel conversational data to capture nuanced two-speaker dialogue patterns for human-like interactions. However, they face a critical challenge -- their conversational abilities often degrade compared to pure-text conversation due to prolonged speech sequences and limited high-quality spoken dialogue data. While text-guided speech generation could mitigate these issues, it suffers from timing and length issues when integrating textual guidance into double-channel audio streams, disrupting the precise time alignment essential for natural interactions. To address these challenges, we propose TurnGuide, a novel planning-inspired approach that mimics human conversational planning by dynamically segmenting assistant speech into dialogue turns and generating turn-level text guidance before speech output, which effectively resolves both insertion timing and length challenges. Extensive experiments demonstrate our approach significantly improves e2e FD-SLMs' conversational abilities, enabling them to generate semantically meaningful and coherent speech while maintaining natural conversational flow. Demos are available at https://dreamtheater123.github.io/TurnGuide-Demo/. Code will be available at https://github.com/dreamtheater123/TurnGuide.
【26】Keyword Mamba: Spoken Keyword Spotting with State Space Models
标题:关键词曼巴:使用状态空间模型进行口语关键词定位
链接:https://arxiv.org/abs/2508.07363
摘要:关键词识别是语音处理中的一项重要任务。它广泛应用于语音助手和智能设备。CNN、RNN和Transformers等深度学习模型在KWS中表现良好。然而,他们往往难以处理长期模式,同时保持效率。在这项工作中,我们提出了关键字曼巴,一个新的架构KWS。它使用了一个叫做Mamba的神经状态空间模型(SSM)。我们沿着时间轴应用Mamba,并探索它如何取代Transformer模型中的自我注意部分。我们在Google Speech Commands数据集上测试了我们的模型。结果表明,Keyword Mamba算法具有较强的识别精度,且参数少,计算量小.据我们所知,这是第一次状态空间模型已被用于KWS。这些结果表明,曼巴在语音相关的任务有很强的潜力。
摘要:Keyword spotting (KWS) is an essential task in speech processing. It is widely used in voice assistants and smart devices. Deep learning models like CNNs, RNNs, and Transformers have performed well in KWS. However, they often struggle to handle long-term patterns and stay efficient at the same time. In this work, we present Keyword Mamba, a new architecture for KWS. It uses a neural state space model (SSM) called Mamba. We apply Mamba along the time axis and also explore how it can replace the self-attention part in Transformer models. We test our model on the Google Speech Commands datasets. The results show that Keyword Mamba reaches strong accuracy with fewer parameters and lower computational cost. To our knowledge, this is the first time a state space model has been used for KWS. These results suggest that Mamba has strong potential in speech-related tasks.
【27】Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
标题:在大型语音语言模型中融入上下文副语言理解
链接:https://arxiv.org/abs/2508.07273
摘要:目前的大型语音语言模型(Speech-LLM)通常表现出移情推理的局限性,主要是由于缺乏整合上下文内容和非语言线索的训练数据集。在这项工作中,我们提出了两种方法来将上下文语言信息纳入模型训练:(1)一种明确的方法,提供语言元数据(例如,情感注释)直接到LLM,以及(2)使用分类和维度情感注释以及语音转换自动生成新颖的训练问答(QA)对的隐式方法。我们的隐式方法在人类注释的QA基准测试中将性能(LLM判断)提高了38.41%,与显式方法相结合时达到46.02%,显示了上下文语言理解的有效性。我们还验证了LLM判断,展示其相关性与分类指标,提供支持,其可靠性。
摘要:Current large speech language models (Speech-LLMs) often exhibit limitations in empathetic reasoning, primarily due to the absence of training datasets that integrate both contextual content and paralinguistic cues. In this work, we propose two approaches to incorporate contextual paralinguistic information into model training: (1) an explicit method that provides paralinguistic metadata (e.g., emotion annotations) directly to the LLM, and (2) an implicit method that automatically generates novel training question-answer (QA) pairs using both categorical and dimensional emotion annotations alongside speech transcriptions. Our implicit method boosts performance (LLM-judged) by 38.41% on a human-annotated QA benchmark, reaching 46.02% when combined with the explicit approach, showing effectiveness in contextual paralinguistic understanding. We also validate the LLM judge by demonstrating its correlation with classification metrics, providing support for its reliability.
【28】How Does a Deep Neural Network Look at Lexical Stress?
标题:深度神经网络如何看待词汇压力?
链接:https://arxiv.org/abs/2508.07229
摘要:尽管神经网络在语音处理方面取得了成功,但它们通常像黑匣子一样运行,这就引发了一个问题:是什么影响了它们的决策,以及我们如何解释它们?本文从词汇重音的角度来探讨这一问题。从阅读和自发语音中自动构建了英语双音节词数据集。训练了几种卷积神经网络(CNN)架构,以从缺乏最小重音对的双音节词(例如,初始应力Wallet,最终应力exTEND),在保持测试数据上达到高达92%的准确度。分层相关传播(LRP),一种用于CNN可解释性分析的技术,揭示了对保留最小对(PROtest与proTEST)的预测受到重读音节与非重读音节信息的最强烈影响,特别是重读元音的频谱特性。然而,分类器也关注整个单词的信息。提出了一个特定功能的相关性分析,其结果表明,我们表现最好的分类器的强烈影响,强调元音的第一和第二共振峰,有一些证据表明,其音高和第三共振峰也作出贡献。这些结果揭示了深度学习从自然发生的数据中获取分布式线索的能力,扩展了基于高度控制刺激的传统语音工作。
摘要:Despite their success in speech processing, neural networks often operate as black boxes, prompting the question: what informs their decisions, and how can we interpret them? This work examines this issue in the context of lexical stress. A dataset of English disyllabic words was automatically constructed from read and spontaneous speech. Several Convolutional Neural Network (CNN) architectures were trained to predict stress position from a spectrographic representation of disyllabic words lacking minimal stress pairs (e.g., initial stress WAllet, final stress exTEND), achieving up to 92% accuracy on held-out test data. Layerwise Relevance Propagation (LRP), a technique for CNN interpretability analysis, revealed that predictions for held-out minimal pairs (PROtest vs. proTEST ) were most strongly influenced by information in stressed versus unstressed syllables, particularly the spectral properties of stressed vowels. However, the classifiers also attended to information throughout the word. A feature-specific relevance analysis is proposed, and its results suggest that our best-performing classifier is strongly influenced by the stressed vowel's first and second formants, with some evidence that its pitch and third formant also contribute. These results reveal deep learning's ability to acquire distributed cues to stress from naturally occurring data, extending traditional phonetic work based around highly controlled stimuli.
机器翻译由腾讯交互翻译提供,仅供参考
