微信公众号:arXiv_Daily
cs.SD语音
【1】Multichannel Keyword Spotting for Noisy Conditions
标题:多通道关键字定位以应对噪音条件
链接:http://arxiv.org/pdf/2507.15558v1
备注:Accepted to Interspeech 2025
摘要:本文提出了一种在噪声环境下改进关键词检测算法的方法。虽然波束形成(BF)和自适应噪声消除(ANC)技术在某些条件下是鲁棒的,但是它们可能通过使有用信号失真或抑制而使激活系统的性能降级。作者提出了一种神经网络架构,该架构使用多个输入通道和一个注意力机制,允许网络确定最有用的通道或它们的组合。在两个数据集上证明了该算法的改进质量:来自受控条件下的实验室和来自自然条件下的智能扬声器。与现有的解决方案相比,所提出的算法进行了比较,对几个基线的降噪指标,KWS指标和计算资源的质量。
摘要:This article presents a method for improving a keyword spotter (KWS) algorithm in noisy environments. Although beamforming (BF) and adaptive noise cancellation (ANC) techniques are robust in some conditions, they may degrade the performance of the activation system by distorting or suppressing useful signals. The authors propose a neural network architecture that uses several input channels and an attention mechanism that allows the network to determine the most useful channel or their combination. The improved quality of the algorithm was demonstrated on two datasets: from a laboratory with controlled conditions and from smart speakers in natural conditions. The proposed algorithm was compared against several baselines in terms of the quality of noise reduction metrics, KWS metrics, and computing resources in comparison with existing solutions.
【2】An Investigation of Test-time Adaptation for Audio Classification under Background Noise
标题:背景噪音下音频分类测试时间自适应研究
链接:http://arxiv.org/pdf/2507.15523v1
摘要:域偏移是深度学习中的一个突出问题,导致在源数据集上预训练的模型在测试数据集上的性能显著下降。本研究旨在使用测试时自适应(TTA)解决背景噪声引起的域偏移下的音频分类问题,TTA是一种在测试期间仅使用未标记的测试数据进行预测之前调整预训练模型的技术。我们采用了两种常见的TTA方法,TTT和TENT,以及最先进的方法CoNMix,并研究了它们在两个流行的音频分类数据集AudioMNIST(AM)和SpeechCommands V1(SC)上的各自性能,以对抗不同类型的背景噪声和噪声严重程度。实验结果表明,我们提出的修改版本的CoNMix产生了最高的分类准确率下域移位(5.31%的错误率下10 dB的健身自行车背景噪声和12.75%的错误率下3 dB的运行抽头背景噪声为AM)相比,TTT和TENT。文献检索没有提供类似作品的证据,从而激发了这里报告的工作作为第一项研究,利用TTA技术进行域转移下的音频分类。
摘要:Domain shift is a prominent problem in Deep Learning, causing a model pre-trained on a source dataset to suffer significant performance degradation on test datasets. This research aims to address the issue of audio classification under domain shift caused by background noise using Test-Time Adaptation (TTA), a technique that adapts a pre-trained model during testing using only unlabelled test data before making predictions. We adopt two common TTA methods, TTT and TENT, and a state-of-the-art method CoNMix, and investigate their respective performance on two popular audio classification datasets, AudioMNIST (AM) and SpeechCommands V1 (SC), against different types of background noise and noise severity levels. The experimental results reveal that our proposed modified version of CoNMix produced the highest classification accuracy under domain shift (5.31% error rate under 10 dB exercise bike background noise and 12.75% error rate under 3 dB running tap background noise for AM) compared to TTT and TENT. The literature search provided no evidence of similar works, thereby motivating the work reported here as the first study to leverage TTA techniques for audio classification under domain shift.
【3】Neuro-MSBG: An End-to-End Neural Model for Hearing Loss Simulation
标题:Neuro-MSBG:用于听力损失模拟的端到端神经模型
链接:http://arxiv.org/pdf/2507.15396v1
作者:Hui-Guan Yuan, Ryandhimas E. Zezario , Shafique Ahmed, Hsin-Min Wang, Kai-Lung Hua, Yu Tsao
摘要:听力损失仿真模型对于助听器的部署至关重要。然而,现有的模型具有较高的计算复杂度和延迟,这限制了实时应用,并且缺乏与语音处理系统的直接集成。为了解决这些问题,我们提出了Neuro-MSBG,一个轻量级的端到端模型,具有个性化的听力图编码器,用于有效的时频建模。实验表明,神经MSBG支持并行推理,并保留了原始MSBG的可懂度和感知质量,与斯皮尔曼的等级相关系数(SRCC)为0.9247的短时客观可懂度(STOI)和0.8671的语音质量的感知评价(PESQ)。Neuro-MSBG将模拟运行时间减少了46倍(对于1秒输入,从0.970秒减少到0.021秒),进一步证明了其效率和实用性。
摘要:Hearing loss simulation models are essential for hearing aid deployment. However, existing models have high computational complexity and latency, which limits real-time applications and lack direct integration with speech processing systems. To address these issues, we propose Neuro-MSBG, a lightweight end-to-end model with a personalized audiogram encoder for effective time-frequency modeling. Experiments show that Neuro-MSBG supports parallel inference and retains the intelligibility and perceptual quality of the original MSBG, with a Spearman's rank correlation coefficient (SRCC) of 0.9247 for Short-Time Objective Intelligibility (STOI) and 0.8671 for Perceptual Evaluation of Speech Quality (PESQ). Neuro-MSBG reduces simulation runtime by a factor of 46 (from 0.970 seconds to 0.021 seconds for a 1 second input), further demonstrating its efficiency and practicality.
【4】MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions
标题:MeMo:视觉受损条件下实时视听说话人提取的注意动量
链接:http://arxiv.org/pdf/2507.15294v1
作者:Junjie Li, Wenxuan Wu, Shuai Wang, Zexu Pan, Kong Aik Lee, Helen Meng, Haizhou Li
摘要:视听目标说话人提取(AV-TSE)的目的是利用视觉线索作为指导,从多说话人环境中分离出目标说话人的声音。然而,AV-TSE系统的性能在很大程度上依赖于这些视觉线索的质量。在视觉线索丢失或严重退化的极端情况下,系统可能无法准确地提取目标说话者。相比之下,即使在没有明确的辅助信息的情况下,人类也可以保持对目标说话者的注意力。受这种人类认知能力的启发,我们提出了一个新的框架,称为MeMo,它包含两个自适应记忆库来存储注意力相关的信息。MeMo是专门为实时场景设计的:一旦建立了初始注意力,系统就会随着时间的推移保持注意力的动力,即使视觉提示变得不可用。我们进行了全面的实验,以验证MeMo的有效性。实验结果表明,我们提出的框架实现了至少2 dB的SI-SNR的改善,在相应的基线。
摘要:Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate a target speaker's voice from multi-speaker environments by leveraging visual cues as guidance. However, the performance of AV-TSE systems heavily relies on the quality of these visual cues. In extreme scenarios where visual cues are missing or severely degraded, the system may fail to accurately extract the target speaker. In contrast, humans can maintain attention on a target speaker even in the absence of explicit auxiliary information. Motivated by such human cognitive ability, we propose a novel framework called MeMo, which incorporates two adaptive memory banks to store attention-related information. MeMo is specifically designed for real-time scenarios: once initial attention is established, the system maintains attentional momentum over time, even when visual cues become unavailable. We conduct comprehensive experiments to verify the effectiveness of MeMo. Experimental results demonstrate that our proposed framework achieves SI-SNR improvements of at least 2 dB over the corresponding baseline.
【5】A2TTS: TTS for Low Resource Indian Languages
标题:A2 TTC:低资源印度语言的TTC
链接:http://arxiv.org/pdf/2507.15272v1
作者:Ayush Singh Bhadoriya, Abhishek Nikunj Shinde, Isha Pandey, Ganesh Ramakrishnan
摘要:我们提出了一个扬声器条件的文本到语音(TTS)系统,旨在解决在生成语音看不见的扬声器和支持不同的印度语言的挑战。我们的方法利用基于扩散的TTS架构,其中扬声器编码器从短参考音频样本中提取嵌入以调节DDPM解码器用于多扬声器生成。为了进一步增强韵律和自然性,我们采用了基于交叉注意的持续时间预测机制,该机制利用参考音频,实现更准确和扬声器一致的定时。这导致语音与目标说话者非常相似,同时改善了持续时间建模和整体表现力。此外,为了改进zero-shot生成,我们采用了分类器自由指导,允许系统为未知说话者生成更接近语音的语音。使用这种方法,我们训练了特定于语言的说话者条件模型。使用IndicSUPERB数据集为多种印度语言,如孟加拉语,古吉拉特语,印地语,马拉地语,马拉雅拉姆语,旁遮普语和泰米尔语。
摘要:We present a speaker conditioned text-to-speech (TTS) system aimed at addressing challenges in generating speech for unseen speakers and supporting diverse Indian languages. Our method leverages a diffusion-based TTS architecture, where a speaker encoder extracts embeddings from short reference audio samples to condition the DDPM decoder for multispeaker generation. To further enhance prosody and naturalness, we employ a cross-attention based duration prediction mechanism that utilizes reference audio, enabling more accurate and speaker consistent timing. This results in speech that closely resembles the target speaker while improving duration modeling and overall expressiveness. Additionally, to improve zero-shot generation, we employed classifier free guidance, allowing the system to generate speech more near speech for unknown speakers. Using this approach, we trained language-specific speaker-conditioned models. Using the IndicSUPERB dataset for multiple Indian languages such as Bengali, Gujarati, Hindi, Marathi, Malayalam, Punjabi and Tamil.
【6】EchoVoices: Preserving Generational Voices and Memories for Seniors and Children
标题:EchoVoices:为老年人和儿童保留代际声音和记忆
链接:http://arxiv.org/pdf/2507.15221v1
摘要:智能语音和数字人类技术的最新突破主要针对主流成人用户,往往忽视了老年人和儿童独特的声音模式和互动风格。这些人口统计数据具有独特的声音特征,语言风格和交互模式,挑战传统的ASR,TTS和LLM系统。为了解决这个问题,我们推出了EchoVoices,这是一个端到端的数字人管道,致力于为老年人和儿童创建持久的数字角色,确保他们的声音和记忆为后代保留。我们的系统集成了三个核心创新:一个k-NN增强的Whisper模型,用于非典型语音的鲁棒语音识别;一个年龄自适应的VITS模型,用于高保真,说话者感知的语音合成;和一个LLM驱动的代理,自动生成人物卡,并利用基于RAG的记忆系统进行会话一致性。我们在SeniorTalk和ChildMandarin数据集上进行的实验表明,识别准确率,合成质量和说话人相似性都有显着提高。EchoVoices提供了一个全面的框架来保护代际声音,提供了一种新的代际联系方式,并创造了持久的数字遗产。
摘要:Recent breakthroughs in intelligent speech and digital human technologies have primarily targeted mainstream adult users, often overlooking the distinct vocal patterns and interaction styles of seniors and children. These demographics possess distinct vocal characteristics, linguistic styles, and interaction patterns that challenge conventional ASR, TTS, and LLM systems. To address this, we introduce EchoVoices, an end-to-end digital human pipeline dedicated to creating persistent digital personas for seniors and children, ensuring their voices and memories are preserved for future generations. Our system integrates three core innovations: a k-NN-enhanced Whisper model for robust speech recognition of atypical speech; an age-adaptive VITS model for high-fidelity, speaker-aware speech synthesis; and an LLM-driven agent that automatically generates persona cards and leverages a RAG-based memory system for conversational consistency. Our experiments, conducted on the SeniorTalk and ChildMandarin datasets, demonstrate significant improvements in recognition accuracy, synthesis quality, and speaker similarity. EchoVoices provides a comprehensive framework for preserving generational voices, offering a new means of intergenerational connection and the creation of lasting digital legacies.
【7】Exploiting Context-dependent Duration Features for Voice Anonymization Attack Systems
标题:利用语音匿名攻击系统的上下文相关持续时间特征
链接:http://arxiv.org/pdf/2507.15214v1
备注:Accepted at Interspeech-2025
摘要:语音的时间动态,包括节奏,语调和语速的变化,包含了关于说话人身份的重要和独特的信息。本文提出了一种新的方法来表示说话人特征提取上下文相关的持续时间嵌入语音时间动态。我们开发了新的攻击模型,使用这些表示和分析潜在的漏洞,在说话人验证和语音匿名化systems.Experimental结果表明,开发的攻击模型提供了显着改善说话人验证性能的原始和匿名数据相比,在文献中报道的语音时间动态的简单表示。
摘要:The temporal dynamics of speech, encompassing variations in rhythm, intonation, and speaking rate, contain important and unique information about speaker identity. This paper proposes a new method for representing speaker characteristics by extracting context-dependent duration embeddings from speech temporal dynamics. We develop novel attack models using these representations and analyze the potential vulnerabilities in speaker verification and voice anonymization systems.The experimental results show that the developed attack models provide a significant improvement in speaker verification performance for both original and anonymized data in comparison with simpler representations of speech temporal dynamics reported in the literature.
【8】Frame-level Temporal Difference Learning for Partial Deepfake Speech Detection
标题:用于部分Deepfake语音检测的帧级时间差异学习
链接:http://arxiv.org/pdf/2507.15101v1
备注:5 pages, 4 figures, 4 tables. Accepted to IEEE SPL
摘要:检测部分deepfake语音是必不可少的,因为它可能会产生微妙的错误信息。然而,现有方法在训练期间依赖于昂贵的帧级注释,限制了现实世界的可扩展性。此外,他们专注于检测bonafide和deepfake片段之间的过渡伪影。随着deepfake生成技术越来越平滑这些过渡,检测变得更具挑战性。为了解决这个问题,我们的工作通过分析帧级时间差异引入了一个新的视角,并揭示了与真正的语音相比,deepfake语音表现出不稳定的方向变化和不自然的局部过渡。基于这一发现,我们提出了一个时间差异注意力模块(TDAM),它将部分深度伪造检测重新定义为识别不自然的时间变化,而不依赖于显式的边界注释。一个双层的层次差异表示捕捉时间的不规则性,在精细和粗糙的尺度,而自适应平均池保存跨可变长度的输入,以尽量减少信息丢失的基本模式。我们的TDAM-AvgPool模型实现了最先进的性能,在PartialSpoof数据集上的EER为0.59%,在HAD数据集上为0.03%,显著优于现有方法,而无需帧级监督。摘要:Detecting partial deepfake speech is essential due to its potential for subtle misinformation. However, existing methods depend on costly frame-level annotations during training, limiting real-world scalability. Also, they focus on detecting transition artifacts between bonafide and deepfake segments. As deepfake generation techniques increasingly smooth these transitions, detection has become more challenging. To address this, our work introduces a new perspective by analyzing frame-level temporal differences and reveals that deepfake speech exhibits erratic directional changes and unnatural local transitions compared to bonafide speech. Based on this finding, we propose a Temporal Difference Attention Module (TDAM) that redefines partial deepfake detection as identifying unnatural temporal variations, without relying on explicit boundary annotations. A dual-level hierarchical difference representation captures temporal irregularities at both fine and coarse scales, while adaptive average pooling preserves essential patterns across variable-length inputs to minimize information loss. Our TDAM-AvgPool model achieves state-of-the-art performance, with an EER of 0.59% on the PartialSpoof dataset and 0.03% on the HAD dataset, which significantly outperforms the existing methods without requiring frame-level supervision.
【9】Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion Modeling
标题:通过分层运动建模实现音乐对齐的整体3D舞蹈生成
链接:http://arxiv.org/pdf/2507.14915v1
摘要:协调良好,音乐一致的整体舞蹈增强了情感表达和观众的参与。然而,生成这样的舞蹈仍然具有挑战性,因为缺乏完整的3D舞蹈数据集,难以实现音乐和舞蹈之间的跨模态对齐,以及对身体,手和面部的相互依赖运动建模的复杂性。为了应对这些挑战,我们引入了SoulDance,这是一个通过专业动作捕捉系统捕捉的高精度音乐舞蹈配对数据集,具有精心注释的整体舞蹈动作。在此数据集的基础上,我们提出了SoulNet,一个旨在生成音乐对齐,运动学协调的整体舞蹈序列的框架。SoulNet由三个主要组件组成:(1)分层残差矢量量化,它对身体,手和面部的复杂,细粒度的运动依赖关系进行建模;(2)音乐对齐生成模型,它将这些分层运动单元组成表达和协调的整体舞蹈;(3)音乐-动作检索模块,一个预先训练的跨模态模型,用作音乐-舞蹈对齐先验,确保在整个生成过程中生成的舞蹈和输入音乐之间的时间同步和语义一致性。大量的实验表明,SoulNet在生成高质量、音乐协调和对齐良好的整体3D舞蹈序列方面明显优于现有方法。
摘要:Well-coordinated, music-aligned holistic dance enhances emotional expressiveness and audience engagement. However, generating such dances remains challenging due to the scarcity of holistic 3D dance datasets, the difficulty of achieving cross-modal alignment between music and dance, and the complexity of modeling interdependent motion across the body, hands, and face. To address these challenges, we introduce SoulDance, a high-precision music-dance paired dataset captured via professional motion capture systems, featuring meticulously annotated holistic dance movements. Building on this dataset, we propose SoulNet, a framework designed to generate music-aligned, kinematically coordinated holistic dance sequences. SoulNet consists of three principal components: (1) Hierarchical Residual Vector Quantization, which models complex, fine-grained motion dependencies across the body, hands, and face; (2) Music-Aligned Generative Model, which composes these hierarchical motion units into expressive and coordinated holistic dance; (3) Music-Motion Retrieval Module, a pre-trained cross-modal model that functions as a music-dance alignment prior, ensuring temporal synchronization and semantic coherence between generated dance and input music throughout the generation process. Extensive experiments demonstrate that SoulNet significantly surpasses existing approaches in generating high-quality, music-coordinated, and well-aligned holistic 3D dance sequences.
【10】Multi-Sampling-Frequency Naturalness MOS Prediction Using Self-Supervised Learning Model with Sampling-Frequency-Independent Layer
标题:使用具有采样频率独立层的自监督学习模型的多采样频率自然性MOS预测
链接:http://arxiv.org/pdf/2507.14647v1
备注:4 pages, 2 figures
摘要:我们介绍我们提交的AudioMOS挑战赛(AMC)2025轨道3:多采样频率(SF)语音的平均意见得分(MOS)预测。我们提交的模型将SF独立(SFI)卷积层集成到自监督学习(SSL)模型中,以实现用于MOS预测的SFI语音特征提取。我们提出了一些策略来提高模型的MOS预测性能:从预训练的非SFI-SSL模型中提取知识,并使用大规模MOS数据集进行预训练。我们提交给AMC 2025 Track 3的作品在一个评估指标中排名第一,在最终排名中排名第四。我们还报告了我们的消融研究的结果,以调查我们的模型的基本因素。
摘要:We introduce our submission to the AudioMOS Challenge (AMC) 2025 Track 3: mean opinion score (MOS) prediction for speech with multiple sampling frequencies (SFs). Our submitted model integrates an SF-independent (SFI) convolutional layer into a self-supervised learning (SSL) model to achieve SFI speech feature extraction for MOS prediction. We present some strategies to improve the MOS prediction performance of our model: distilling knowledge from a pretrained non-SFI-SSL model and pretraining with a large-scale MOS dataset. Our submission to the AMC 2025 Track 3 ranked the first in one evaluation metric and the fourth in the final ranking. We also report the results of our ablation study to investigate essential factors of our model.
【11】The Rest is Silence: Leveraging Unseen Species Models for Computational Musicology
标题:剩下的就是沉默:利用看不见的物种模型进行计算音乐学
链接:http://arxiv.org/pdf/2507.14638v1
摘要:几十年来,音乐学家一直致力于创建大型数据库,为音乐学研究和学术研究提供不同的目的。随着音乐信息检索和数字音乐学等领域的兴起,音乐学相关数据集和语料库不断涌入。然而,在历史或观察背景下,这些数据集必然是不完整的,感兴趣的收集的真实程度仍然未知-沉默。在这里,我们第一次将所谓的“看不见的物种”模型(USMs)从生态学应用到音乐活动领域。在正式介绍模型后,我们在四个案例研究中展示了如何将USMs应用于音乐学数据,以解决定量问题,如:我们在RISM中缺少多少作曲家?我们已经编目的中世纪格里高利圣咏的来源有多少?我们期望在不同的版本中找到多少音乐拷贝的差异?民间音乐传统流派的歌曲覆盖面有多大?最后,我们对大量作曲家的和声词汇量的估计有多接近?
摘要:For many decades, musicologists have engaged in creating large databases serving different purposes for musicological research and scholarship. With the rise of fields like music information retrieval and digital musicology, there is now a constant and growing influx of musicologically relevant datasets and corpora. In historical or observational settings, however, these datasets are necessarily incomplete, and the true extent of a collection of interest remains unknown -- silent. Here, we apply, for the first time, so-called Unseen Species models (USMs) from ecology to areas of musicological activity. After introducing the models formally, we show in four case studies how USMs can be applied to musicological data to address quantitative questions like: How many composers are we missing in RISM? What percentage of medieval sources of Gregorian chant have we already cataloged? How many differences in music prints do we expect to find between editions? How large is the coverage of songs from genres of a folk music tradition? And, finally, how close are we in estimating the size of the harmonic vocabulary of a large number of composers?
【12】U-DREAM: Unsupervised Dereverberation guided by a Reverberation Model
标题:U-DREAM:由混响模型引导的无监督去混响
链接:http://arxiv.org/pdf/2507.14237v1
备注:Submitted to IEEE Transactions on Audio, Speech and Language Processing (TASLPRO)
摘要:本文探讨了训练最先进的去混响模型的结果,其监督设置范围从弱监督到完全无监督,仅依赖于混响信号和声学模型进行训练。大多数现有的深度学习方法通常需要成对的干数据和混响数据,这在实践中很难获得。相反,我们开发了一种顺序学习策略,其动机是去混响问题的bavonet公式,其中声学参数和干信号是使用深度神经网络从混响输入中估计的,由混响匹配损失指导。我们的数据效率最高的变体只需要100个混响参数标记的样本就可以超越无监督基线,证明了所提出的方法在低资源场景中的有效性和实用性。摘要:This paper explores the outcome of training state-ofthe-art dereverberation models with supervision settings ranging from weakly-supervised to fully unsupervised, relying solely on reverberant signals and an acoustic model for training. Most of the existing deep learning approaches typically require paired dry and reverberant data, which are difficult to obtain in practice. We develop instead a sequential learning strategy motivated by a bayesian formulation of the dereverberation problem, wherein acoustic parameters and dry signals are estimated from reverberant inputs using deep neural networks, guided by a reverberation matching loss. Our most data-efficient variant requires only 100 reverberation-parameter-labelled samples to outperform an unsupervised baseline, demonstrating the effectiveness and practicality of the proposed method in low-resource scenarios.
【13】Developing an AI-Guided Assistant Device for the Deaf and Hearing Impaired
标题:开发一种用于聋人和听力障碍者的AI引导辅助设备
链接:http://arxiv.org/pdf/2507.14215v1
摘要:本研究旨在为聋人或听力受损者开发一种无障碍设备的深度学习系统。该设备将实时准确定位和识别声源。这项研究将填补当前研究的一个重要空白,利用机器学习技术来针对贫困社区。该系统包括三个主要组成部分。1. JerryNet:一种定制设计的CNN架构,可以确定九个可能方向的到达方向(DoA)。2.音频分类:该模型基于对对比音频预训练(CLAP)模型的微调,以仅基于音频来识别确切的声音类别。3.多式联运模式:这是一个精确的声音定位模型,它结合了音频,视觉和文本数据,以定位图像中的确切声源。该部分由两个模块组成,一个是使用Yolov9生成对象的所有边界框的对象检测模块,另一个是使用完整的Intersection over Union(CIoU)识别最佳边界框的视听定位模型。硬件包括一个四麦克风的矩形结构和一个安装在眼镜上的摄像头,眼镜上有一个腕带,用于显示必要的信息,如方向。在自定义收集的数据集上,JerryNet的精度达到了91。1%的声音方向,优于所有的基线模型。CLAP模型在自定义和AudioSet数据集上分别达到了98.5%和95%的准确率。组件3中的视听定位模型产生了0.892的cIoU和0.658的AUC,超过了其他类似模型。这项研究有许多未来的潜力,为创造新一代的无障碍设备铺平了道路。
摘要:This study aims to develop a deep learning system for an accessibility device for the deaf or hearing impaired. The device will accurately localize and identify sound sources in real time. This study will fill an important gap in current research by leveraging machine learning techniques to target the underprivileged community. The system includes three main components. 1. JerryNet: A custom designed CNN architecture that determines the direction of arrival (DoA) for nine possible directions. 2. Audio Classification: This model is based on fine-tuning the Contrastive Language-Audio Pretraining (CLAP) model to identify the exact sound classes only based on audio. 3. Multimodal integration model: This is an accurate sound localization model that combines audio, visual, and text data to locate the exact sound sources in the images. The part consists of two modules, one object detection using Yolov9 to generate all the bounding boxes of the objects, and an audio visual localization model to identify the optimal bounding box using complete Intersection over Union (CIoU). The hardware consists of a four-microphone rectangular formation and a camera mounted on glasses with a wristband for displaying necessary information like direction. On a custom collected data set, JerryNet achieved a precision of 91. 1% for the sound direction, outperforming all the baseline models. The CLAP model achieved 98.5% and 95% accuracy on custom and AudioSet datasets, respectively. The audio-visual localization model within component 3 yielded a cIoU of 0.892 and an AUC of 0.658, surpassing other similar models. There are many future potentials to this study, paving the way to creating a new generation of accessibility devices.
【14】Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion
标题:柯南:Zero-Shot自适应语音转换的Chunkwise在线网络
链接:http://arxiv.org/pdf/2507.14534v1
摘要:Zero-shot在线语音转换(VC)为实时通信和娱乐带来了巨大的希望。然而,目前的VC模型难以在实时约束下保持语义保真度,提供自然的声音转换,并有效地适应看不见的说话者特征。为了解决这些挑战,我们介绍了柯南,一个chunkwise在线zero-shot语音转换模型,保留了源的内容,同时匹配的语音音色和风格的参考语音。Conan包括三个核心组件:1)流内容提取器,利用Emformer进行低延迟流内容编码; 2)自适应风格编码器,从参考语音中提取细粒度的风格特征,以增强风格适应; 3)因果洗牌声码器,使用像素洗牌机制实现完全因果HiFiGAN。实验评估表明,柯南优于基线模型的主观和客观指标。音频样本可以在https: aaronz345.github.io ConanDemo上找到。
摘要:Zero-shot online voice conversion (VC) holds significant promise for real-time communications and entertainment. However, current VC models struggle to preserve semantic fidelity under real-time constraints, deliver natural-sounding conversions, and adapt effectively to unseen speaker characteristics. To address these challenges, we introduce Conan, a chunkwise online zero-shot voice conversion model that preserves the content of the source while matching the voice timbre and styles of reference speech. Conan comprises three core components: 1) a Stream Content Extractor that leverages Emformer for low-latency streaming content encoding; 2) an Adaptive Style Encoder that extracts fine-grained stylistic features from reference speech for enhanced style adaptation; 3) a Causal Shuffle Vocoder that implements a fully causal HiFiGAN using a pixel-shuffle mechanism. Experimental evaluations demonstrate that Conan outperforms baseline models in subjective and objective metrics. Audio samples can be found at https: aaronz345.github.io ConanDemo.
【15】Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
标题:将Whisper调整为设备上边缘应用程序的轻量级、高效的儿童自动语音识别
链接:http://arxiv.org/pdf/2507.14451v1
备注:5 pages, 5 figures, accepted for presentation at the 2025 Workshop on Child Computer Interaction (WOCCI 2025), a Satellite Workshop of the 2025 Interspeech Conference
摘要:由于监管和隐私挑战,云提供商用于支持以儿童为中心的基于语音的应用程序的ASR推理的可靠性变得越来越具有挑战性。受隐私保护设计的启发,本研究旨在开发一个能够在Raspberry Pi上运行的轻量级高效Whisper ASR系统。在评估MyST语料库并通过检查各种过滤策略来微调“tiny.en”模型后,实现了15.9%的单词错误率(WER)(11.8%过滤)。低秩压缩将编码器大小减少了0.51M,GPU中的推理速度提高了1.26倍,相对WER增加了11%。在对Pi进行推理时,压缩版本需要的计算量减少了约2 GFLOPS。对于各种输入音频持续时间,两种模型的RTF范围在[0.23-0.41]之间。分析RAM使用率和CPU温度表明,PI能够处理这两种小型模型,但注意到小型模型会引发额外的开销 热量调节。摘要:Reliability on cloud providers for ASR inference to support child-centered voice-based applications is becoming challenging due to regulatory and privacy challenges. Motivated by a privacy-preserving design, this study aims to develop a lightweight & efficient Whisper ASR system capable of running on a Raspberry Pi. Upon evaluation of the MyST corpus and by examining various filtering strategies to fine-tune the tiny.en' model, a Word Error Rate (WER) of 15.9% was achieved (11.8% filtered). A low-rank compression reduces the encoder size by 0.51M with 1.26x faster inference in GPU, with 11% relative WER increase. During inference on Pi, the compressed version required ~2 GFLOPS fewer computations. The RTF for both the models ranged between [0.23-0.41] for various input audio durations. Analyzing the RAM usage and CPU temperature showed that the PI was capable of handling both the tiny models, however it was noticed that small models initiated additional overhead thermal throttling.
【16】Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling
标题:通过音素相似度建模实现准确的语音错误检测
链接:http://arxiv.org/pdf/2507.14346v1
备注:2025 Interspeech
摘要:语音错误检测是语音自动评估的核心子任务,它在音素水平上识别发音偏差。来自口音和不流利的语音变化挑战准确的音素识别,当前的模型未能有效地捕捉这些差异。我们提出了一个逐字音素识别框架,使用多任务训练与新的音素相似性建模,转录扬声器实际上说什么,而不是他们应该说什么。我们开发和开源 textit{VCTK-accent},一个包含语音错误的模拟数据集,并提出了两个新的度量标准来评估发音差异。我们的工作为语音错误检测建立了新的基准。摘要:Phonetic error detection, a core subtask of automatic pronunciation assessment, identifies pronunciation deviations at the phoneme level. Speech variability from accents and dysfluencies challenges accurate phoneme recognition, with current models failing to capture these discrepancies effectively. We propose a verbatim phoneme recognition framework using multi-task training with novel phoneme similarity modeling that transcribes what speakers actually say rather than what they're supposed to say. We develop and open-source textit{VCTK-accent}, a simulated dataset containing phonetic errors, and propose two novel metrics for assessing pronunciation differences. Our work establishes a new benchmark for phonetic error detection.
【1】Binaural Signal Matching with Wearable Arrays for Near-Field Sources
标题:近场声源双耳信号的可穿戴阵列匹配
链接:http://arxiv.org/pdf/2507.15517v1
备注:Published at Forum Acusticum 2025
摘要:双耳再现方法旨在通过耳机为收听者重新创建声学场景,在诸如虚拟现实(VR)和电话会议之类的应用中提供沉浸式体验。在现有的方法中,双耳信号匹配(BSM)算法已经证明了高质量的再现,由于其信号独立的配方和灵活的无约束阵列几何形状。然而,这种方法假设远场源,尚未对近场情景进行研究。这项研究评估了近场源的BSM的性能。围绕刚性球体的半圆形阵列的分析,建模头戴式设备,表明远场BSM执行足够的源高达约几十厘米的阵列。然而,对于比这个范围更近的源,双耳误差显著增加。考虑源距离的近场BSM设计显著降低了误差,特别是对于这些非常近的距离,突出了近场建模在提高再现精度方面的优势。摘要:Binaural reproduction methods aim to recreate an acoustic scene for a listener over headphones, offering immersive experiences in applications such as Virtual Reality (VR) and teleconferencing. Among the existing approaches, the Binaural Signal Matching (BSM) algorithm has demonstrated high quality reproduction due to its signal-independent formulation and the flexibility of unconstrained array geometry. However, this method assumes far-field sources and has not yet been investigated for near-field scenarios. This study evaluates the performance of BSM for near-field sources. Analysis of a semi-circular array around a rigid sphere, modeling head-mounted devices, show that far-field BSM performs adequately for sources up to approximately tens of centimeters from the array. However, for sources closer than this range, the binaural error increases significantly. Incorporating a near-field BSM design, which accounts for the source distance, significantly reduces the error, particularly for these very-close distances, highlighting the benefits of near-field modeling in improving reproduction accuracy.
【2】Mixture to Beamformed Mixture: Leveraging Beamformed Mixture as Weak-Supervision for Speech Enhancement and Noise-Robust ASR
标题:混合到束成形混合:利用束成形混合作为语音增强和噪音稳健的ASB的弱监督
链接:http://arxiv.org/pdf/2507.15229v1
备注:in submission
摘要:在多通道语音增强和鲁棒自动语音识别(ASR)中,波束赋形通常可以提高目标说话人的信噪比(SNR),并在对目标语音失真很小的情况下产生可靠的增强效果。有了这个观察结果,我们建议利用波束成形的混合,它具有比输入混合更高的目标扬声器SNR,作为一个弱监督训练深度神经网络(DNN),以增强输入混合。通过这种方式,我们可以使用真实记录的混合物及其波束形成的混合物对来训练增强模型,并且与仅在模拟混合物上训练模型相比,可能实现对真实混合物的更好的泛化,这通常与真实混合物不匹配。在CHiME-4数据集上的实验结果表明了该算法的有效性。
摘要:In multi-channel speech enhancement and robust automatic speech recognition (ASR), beamforming can typically improve the signal-to-noise ratio (SNR) of the target speaker and produce reliable enhancement with little distortion to target speech. With this observation, we propose to leverage beamformed mixture, which has a higher SNR of the target speaker than the input mixture, as a weak supervision to train deep neural networks (DNNs) to enhance the input mixture. This way, we can train enhancement models using pairs of real-recorded mixture and its beamformed mixture, and potentially realize better generalization to real mixtures, compared with only training the models on simulated mixtures, which usually mismatch real mixtures. Evaluation results on the real-recorded CHiME-4 dataset show the effectiveness of the proposed algorithm.
【3】DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis
标题:DMOSpeech 2:用于度量优化语音合成中持续时间预测的强化学习
链接:http://arxiv.org/pdf/2507.14988v1
摘要:基于扩散的文语转换(TTS)系统在zero-shot语音合成方面取得了显著的进展,然而针对感知度量优化所有组件仍然具有挑战性。以前的工作与DMOSpeech演示语音生成组件的直接度量优化,但持续时间预测仍然没有优化。本文介绍了DMOSpeech 2,它通过强化学习方法将度量优化扩展到持续时间预测器。该系统实现了一种新的持续时间策略框架,使用组相对偏好优化(GRPO)与说话人相似性和单词错误率作为奖励信号。通过优化这个以前未优化的组件,DMOSpeech 2创建了一个更完整的度量优化合成管道。此外,本文还介绍了教师指导采样,这是一种混合方法,在过渡到学生模型之前,利用教师模型进行初始去噪步骤,在保持效率的同时显着提高输出多样性。全面的评估表明,与以前的系统相比,在所有指标上都具有卓越的性能,同时将采样步骤减少了一半,而不会降低质量。这些进步代表了语音合成系统跨多个组件的度量优化的重要一步。音频样本、代码和预训练模型可在https: dmospeech2.github.io 上获得。
摘要:Diffusion-based text-to-speech (TTS) systems have made remarkable progress in zero-shot speech synthesis, yet optimizing all components for perceptual metrics remains challenging. Prior work with DMOSpeech demonstrated direct metric optimization for speech generation components, but duration prediction remained unoptimized. This paper presents DMOSpeech 2, which extends metric optimization to the duration predictor through a reinforcement learning approach. The proposed system implements a novel duration policy framework using group relative preference optimization (GRPO) with speaker similarity and word error rate as reward signals. By optimizing this previously unoptimized component, DMOSpeech 2 creates a more complete metric-optimized synthesis pipeline. Additionally, this paper introduces teacher-guided sampling, a hybrid approach leveraging a teacher model for initial denoising steps before transitioning to the student model, significantly improving output diversity while maintaining efficiency. Comprehensive evaluations demonstrate superior performance across all metrics compared to previous systems, while reducing sampling steps by half without quality degradation. These advances represent a significant step toward speech synthesis systems with metric optimization across multiple components. The audio samples, code and pre-trained models are available at https: dmospeech2.github.io .
【4】Parameter-Efficient Fine-Tuning of Foundation Models for CLP Speech Classification
标题:MPP语音分类基础模型的参数高效微调
链接:http://arxiv.org/pdf/2507.14898v1
备注:6 pages, 5 figures, conference
摘要: 我们建议使用参数有效的微调(PEFT)的基础模型唇腭裂(CLP)检测和严重程度分类。在CLP中,鼻化随着严重程度的增加而增加,这是由于口道和鼻道之间的异常通道;这导致口塞音被声门塞音取代,并改变共振峰轨迹和元音空间。由于基础模型是针对字素预测或长期量化表示预测进行训练的,因此当对特定于域的数据进行微调时,它们可以更好地区分CLP严重性。我们在两个数据集上进行实验:英语(NMCPC)和卡纳达语(AIISH)。我们使用来自自监督模型Wav 2 Vec 2和WavLM以及弱监督Whisper的嵌入进行比较分析,每个模型都与SVM分类器配对,并将其与传统的手工特征eGeMAPS和ComparE进行比较。最后,我们使用PEFT技术微调性能最好的Whisper模型:低秩适配器(LoRA)和分解低秩适配器(DoRA)。我们的研究结果表明,所提出的方法在NMCPC数据集上的最佳基础模型和手工特征基线上的宏观平均F1得分分别提高了26.4%和63.4%,在AIISH数据集上分别提高了6.1%和52.9%。摘要:We propose the use of parameter-efficient fine-tuning (PEFT) of foundation models for cleft lip and palate (CLP) detection and severity classification. In CLP, nasalization increases with severity due to the abnormal passage between the oral and nasal tracts; this causes oral stops to be replaced by glottal stops and alters formant trajectories and vowel space. Since foundation models are trained for grapheme prediction or long-term quantized representation prediction, they may better discriminate CLP severity when fine-tuned on domain-specific data. We conduct experiments on two datasets: English (NMCPC) and Kannada (AIISH). We perform a comparative analysis using embeddings from self-supervised models Wav2Vec2 and WavLM, and the weakly supervised Whisper, each paired with SVM classifiers, and compare them with traditional handcrafted features eGeMAPS and ComParE. Finally, we fine-tune the best-performing Whisper model using PEFT techniques: Low-Rank Adapter (LoRA) and Decomposed Low-Rank Adapter (DoRA). Our results demonstrate that the proposed approach achieves relative improvements of 26.4% and 63.4% in macro-average F1 score over the best foundation model and handcrafted feature baselines on the NMCPC dataset, and improvements of 6.1% and 52.9% on the AIISH dataset, respectively.
【5】Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion
标题:柯南:Zero-Shot自适应语音转换的Chunkwise在线网络
链接:http://arxiv.org/pdf/2507.14534v1
摘要:Zero-shot在线语音转换(VC)为实时通信和娱乐带来了巨大的希望。然而,目前的VC模型难以在实时约束下保持语义保真度,提供自然的声音转换,并有效地适应看不见的说话者特征。为了解决这些挑战,我们介绍了柯南,一个chunkwise在线zero-shot语音转换模型,保留了源的内容,同时匹配的语音音色和风格的参考语音。Conan包括三个核心组件:1)流内容提取器,利用Emformer进行低延迟流内容编码; 2)自适应风格编码器,从参考语音中提取细粒度的风格特征,以增强风格适应; 3)因果洗牌声码器,使用像素洗牌机制实现完全因果HiFiGAN。实验评估表明,柯南优于基线模型的主观和客观指标。音频样本可以在https: aaronz345.github.io ConanDemo上找到。
摘要:Zero-shot online voice conversion (VC) holds significant promise for real-time communications and entertainment. However, current VC models struggle to preserve semantic fidelity under real-time constraints, deliver natural-sounding conversions, and adapt effectively to unseen speaker characteristics. To address these challenges, we introduce Conan, a chunkwise online zero-shot voice conversion model that preserves the content of the source while matching the voice timbre and styles of reference speech. Conan comprises three core components: 1) a Stream Content Extractor that leverages Emformer for low-latency streaming content encoding; 2) an Adaptive Style Encoder that extracts fine-grained stylistic features from reference speech for enhanced style adaptation; 3) a Causal Shuffle Vocoder that implements a fully causal HiFiGAN using a pixel-shuffle mechanism. Experimental evaluations demonstrate that Conan outperforms baseline models in subjective and objective metrics. Audio samples can be found at https: aaronz345.github.io ConanDemo.
【6】Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
标题:将Whisper调整为设备上边缘应用程序的轻量级、高效的儿童自动语音识别
链接:http://arxiv.org/pdf/2507.14451v1
备注:5 pages, 5 figures, accepted for presentation at the 2025 Workshop on Child Computer Interaction (WOCCI 2025), a Satellite Workshop of the 2025 Interspeech Conference
摘要:由于监管和隐私挑战,云提供商用于支持以儿童为中心的基于语音的应用程序的ASR推理的可靠性变得越来越具有挑战性。受隐私保护设计的启发,本研究旨在开发一个能够在Raspberry Pi上运行的轻量级高效Whisper ASR系统。在评估MyST语料库并通过检查各种过滤策略来微调“tiny.en”模型后,实现了15.9%的单词错误率(WER)(11.8%过滤)。低秩压缩将编码器大小减少了0.51M,GPU中的推理速度提高了1.26倍,相对WER增加了11%。在对Pi进行推理时,压缩版本需要的计算量减少了约2 GFLOPS。对于各种输入音频持续时间,两种模型的RTF范围在[0.23-0.41]之间。分析RAM使用率和CPU温度表明,PI能够处理这两种小型模型,但注意到小型模型会引发额外的开销 热量调节。摘要:Reliability on cloud providers for ASR inference to support child-centered voice-based applications is becoming challenging due to regulatory and privacy challenges. Motivated by a privacy-preserving design, this study aims to develop a lightweight & efficient Whisper ASR system capable of running on a Raspberry Pi. Upon evaluation of the MyST corpus and by examining various filtering strategies to fine-tune the tiny.en' model, a Word Error Rate (WER) of 15.9% was achieved (11.8% filtered). A low-rank compression reduces the encoder size by 0.51M with 1.26x faster inference in GPU, with 11% relative WER increase. During inference on Pi, the compressed version required ~2 GFLOPS fewer computations. The RTF for both the models ranged between [0.23-0.41] for various input audio durations. Analyzing the RAM usage and CPU temperature showed that the PI was capable of handling both the tiny models, however it was noticed that small models initiated additional overhead thermal throttling.
【7】Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling
标题:通过音素相似度建模实现准确的语音错误检测
链接:http://arxiv.org/pdf/2507.14346v1
备注:2025 Interspeech
摘要:语音错误检测是语音自动评估的核心子任务,它在音素水平上识别发音偏差。来自口音和不流利的语音变化挑战准确的音素识别,当前的模型未能有效地捕捉这些差异。我们提出了一个逐字音素识别框架,使用多任务训练与新的音素相似性建模,转录扬声器实际上说什么,而不是他们应该说什么。我们开发和开源 textit{VCTK-accent},一个包含语音错误的模拟数据集,并提出了两个新的度量标准来评估发音差异。我们的工作为语音错误检测建立了新的基准。摘要:Phonetic error detection, a core subtask of automatic pronunciation assessment, identifies pronunciation deviations at the phoneme level. Speech variability from accents and dysfluencies challenges accurate phoneme recognition, with current models failing to capture these discrepancies effectively. We propose a verbatim phoneme recognition framework using multi-task training with novel phoneme similarity modeling that transcribes what speakers actually say rather than what they're supposed to say. We develop and open-source textit{VCTK-accent}, a simulated dataset containing phonetic errors, and propose two novel metrics for assessing pronunciation differences. Our work establishes a new benchmark for phonetic error detection.
【8】Multichannel Keyword Spotting for Noisy Conditions
标题:多通道关键字定位以应对噪音条件
链接:http://arxiv.org/pdf/2507.15558v1
备注:Accepted to Interspeech 2025
摘要:本文提出了一种在噪声环境下改进关键词检测算法的方法。虽然波束形成(BF)和自适应噪声消除(ANC)技术在某些条件下是鲁棒的,但是它们可能通过使有用信号失真或抑制而使激活系统的性能降级。作者提出了一种神经网络架构,该架构使用多个输入通道和一个注意力机制,允许网络确定最有用的通道或它们的组合。在两个数据集上证明了该算法的改进质量:来自受控条件下的实验室和来自自然条件下的智能扬声器。与现有的解决方案相比,所提出的算法进行了比较,对几个基线的降噪指标,KWS指标和计算资源的质量。
摘要:This article presents a method for improving a keyword spotter (KWS) algorithm in noisy environments. Although beamforming (BF) and adaptive noise cancellation (ANC) techniques are robust in some conditions, they may degrade the performance of the activation system by distorting or suppressing useful signals. The authors propose a neural network architecture that uses several input channels and an attention mechanism that allows the network to determine the most useful channel or their combination. The improved quality of the algorithm was demonstrated on two datasets: from a laboratory with controlled conditions and from smart speakers in natural conditions. The proposed algorithm was compared against several baselines in terms of the quality of noise reduction metrics, KWS metrics, and computing resources in comparison with existing solutions.
【9】An Investigation of Test-time Adaptation for Audio Classification under Background Noise
标题:背景噪音下音频分类测试时间自适应研究
链接:http://arxiv.org/pdf/2507.15523v1
摘要:域偏移是深度学习中的一个突出问题,导致在源数据集上预训练的模型在测试数据集上的性能显著下降。本研究旨在使用测试时自适应(TTA)解决背景噪声引起的域偏移下的音频分类问题,TTA是一种在测试期间仅使用未标记的测试数据进行预测之前调整预训练模型的技术。我们采用了两种常见的TTA方法,TTT和TENT,以及最先进的方法CoNMix,并研究了它们在两个流行的音频分类数据集AudioMNIST(AM)和SpeechCommands V1(SC)上的各自性能,以对抗不同类型的背景噪声和噪声严重程度。实验结果表明,我们提出的修改版本的CoNMix产生了最高的分类准确率下域移位(5.31%的错误率下10 dB的健身自行车背景噪声和12.75%的错误率下3 dB的运行抽头背景噪声为AM)相比,TTT和TENT。文献检索没有提供类似作品的证据,从而激发了这里报告的工作作为第一项研究,利用TTA技术进行域转移下的音频分类。
摘要:Domain shift is a prominent problem in Deep Learning, causing a model pre-trained on a source dataset to suffer significant performance degradation on test datasets. This research aims to address the issue of audio classification under domain shift caused by background noise using Test-Time Adaptation (TTA), a technique that adapts a pre-trained model during testing using only unlabelled test data before making predictions. We adopt two common TTA methods, TTT and TENT, and a state-of-the-art method CoNMix, and investigate their respective performance on two popular audio classification datasets, AudioMNIST (AM) and SpeechCommands V1 (SC), against different types of background noise and noise severity levels. The experimental results reveal that our proposed modified version of CoNMix produced the highest classification accuracy under domain shift (5.31% error rate under 10 dB exercise bike background noise and 12.75% error rate under 3 dB running tap background noise for AM) compared to TTT and TENT. The literature search provided no evidence of similar works, thereby motivating the work reported here as the first study to leverage TTA techniques for audio classification under domain shift.
【10】Neuro-MSBG: An End-to-End Neural Model for Hearing Loss Simulation
标题:Neuro-MSBG:用于听力损失模拟的端到端神经模型
链接:http://arxiv.org/pdf/2507.15396v1
作者:Hui-Guan Yuan, Ryandhimas E. Zezario , Shafique Ahmed, Hsin-Min Wang, Kai-Lung Hua, Yu Tsao
摘要:听力损失仿真模型对于助听器的部署至关重要。然而,现有的模型具有较高的计算复杂度和延迟,这限制了实时应用,并且缺乏与语音处理系统的直接集成。为了解决这些问题,我们提出了Neuro-MSBG,一个轻量级的端到端模型,具有个性化的听力图编码器,用于有效的时频建模。实验表明,神经MSBG支持并行推理,并保留了原始MSBG的可懂度和感知质量,与斯皮尔曼的等级相关系数(SRCC)为0.9247的短时客观可懂度(STOI)和0.8671的语音质量的感知评价(PESQ)。Neuro-MSBG将模拟运行时间减少了46倍(对于1秒输入,从0.970秒减少到0.021秒),进一步证明了其效率和实用性。
摘要:Hearing loss simulation models are essential for hearing aid deployment. However, existing models have high computational complexity and latency, which limits real-time applications and lack direct integration with speech processing systems. To address these issues, we propose Neuro-MSBG, a lightweight end-to-end model with a personalized audiogram encoder for effective time-frequency modeling. Experiments show that Neuro-MSBG supports parallel inference and retains the intelligibility and perceptual quality of the original MSBG, with a Spearman's rank correlation coefficient (SRCC) of 0.9247 for Short-Time Objective Intelligibility (STOI) and 0.8671 for Perceptual Evaluation of Speech Quality (PESQ). Neuro-MSBG reduces simulation runtime by a factor of 46 (from 0.970 seconds to 0.021 seconds for a 1 second input), further demonstrating its efficiency and practicality.
【11】STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models
标题:STITCH:口语模型的同时思维和分块推理说话
链接:http://arxiv.org/pdf/2507.15375v1
备注:Work in progress. Project page: this https URL
摘要:口语语言模型(Spoken Language Models,SLM)被设计为接受语音输入并产生口语响应。然而,当前的slm缺乏在做出反应之前进行内部、不言而喻的思考过程的能力。相比之下,人类通常在内部进行复杂的心理推理,使他们能够清晰简洁地传达想法。因此,将潜意识的思维过程整合到SLM中是非常可取的。虽然在开始谈话之前天真地生成完整的思维链(CoT)推理可以实现对SLM的思考,但这会导致语音响应的额外延迟,因为CoT推理可以任意长。为了解决这个问题,我们提出了缝合,一种新的生成方法,交替之间的生成未说出的推理块和口头响应块。由于语音响应块的音频持续时间比语音响应块中生成令牌的时间长得多,因此我们使用剩余的空闲时间来生成未说出的推理令牌。当向用户播放一段音频时,模型会继续生成下一个未说出的推理块,从而实现同时思考和说话。值得注意的是,Stitch匹配了无法通过设计生成未说出的CoT的基线的延迟,同时在数学推理数据集上的表现比这些基线高出15%; Stitch在非推理数据集上的表现也与这些基线模型一样好。一些动画和演示在项目页面上:https: d223302.github.io STITCH。
摘要:Spoken Language Models (SLMs) are designed to take speech inputs and produce spoken responses. However, current SLMs lack the ability to perform an internal, unspoken thinking process before responding. In contrast, humans typically engage in complex mental reasoning internally, enabling them to communicate ideas clearly and concisely. Thus, integrating an unspoken thought process into SLMs is highly desirable. While naively generating a complete chain-of-thought (CoT) reasoning before starting to talk can enable thinking for SLMs, this induces additional latency for the speech response, as the CoT reasoning can be arbitrarily long. To solve this issue, we propose Stitch, a novel generation method that alternates between the generation of unspoken reasoning chunks and spoken response chunks. Since the audio duration of a chunk of spoken response is much longer than the time to generate the tokens in a chunk of spoken response, we use the remaining free time to generate the unspoken reasoning tokens. When a chunk of audio is played to the user, the model continues to generate the next unspoken reasoning chunk, achieving simultaneous thinking and talking. Remarkably, Stitch matches the latency of baselines that cannot generate unspoken CoT by design while outperforming those baselines by 15% on math reasoning datasets; Stitch also performs equally well on non-reasoning datasets as those baseline models. Some animations and demonstrations are on the project page: https: d223302.github.io STITCH.
【12】A2TTS: TTS for Low Resource Indian Languages
标题:A2 TTC:低资源印度语言的TTC
链接:http://arxiv.org/pdf/2507.15272v1
作者:Ayush Singh Bhadoriya, Abhishek Nikunj Shinde, Isha Pandey, Ganesh Ramakrishnan
摘要:我们提出了一个扬声器条件的文本到语音(TTS)系统,旨在解决在生成语音看不见的扬声器和支持不同的印度语言的挑战。我们的方法利用基于扩散的TTS架构,其中扬声器编码器从短参考音频样本中提取嵌入以调节DDPM解码器用于多扬声器生成。为了进一步增强韵律和自然性,我们采用了基于交叉注意的持续时间预测机制,该机制利用参考音频,实现更准确和扬声器一致的定时。这导致语音与目标说话者非常相似,同时改善了持续时间建模和整体表现力。此外,为了改进zero-shot生成,我们采用了分类器自由指导,允许系统为未知说话者生成更接近语音的语音。使用这种方法,我们训练了特定于语言的说话者条件模型。使用IndicSUPERB数据集为多种印度语言,如孟加拉语,古吉拉特语,印地语,马拉地语,马拉雅拉姆语,旁遮普语和泰米尔语。
摘要:We present a speaker conditioned text-to-speech (TTS) system aimed at addressing challenges in generating speech for unseen speakers and supporting diverse Indian languages. Our method leverages a diffusion-based TTS architecture, where a speaker encoder extracts embeddings from short reference audio samples to condition the DDPM decoder for multispeaker generation. To further enhance prosody and naturalness, we employ a cross-attention based duration prediction mechanism that utilizes reference audio, enabling more accurate and speaker consistent timing. This results in speech that closely resembles the target speaker while improving duration modeling and overall expressiveness. Additionally, to improve zero-shot generation, we employed classifier free guidance, allowing the system to generate speech more near speech for unknown speakers. Using this approach, we trained language-specific speaker-conditioned models. Using the IndicSUPERB dataset for multiple Indian languages such as Bengali, Gujarati, Hindi, Marathi, Malayalam, Punjabi and Tamil.
【13】EchoVoices: Preserving Generational Voices and Memories for Seniors and Children
标题:EchoVoices:为老年人和儿童保留代际声音和记忆
链接:http://arxiv.org/pdf/2507.15221v1
摘要:智能语音和数字人类技术的最新突破主要针对主流成人用户,往往忽视了老年人和儿童独特的声音模式和互动风格。这些人口统计数据具有独特的声音特征,语言风格和交互模式,挑战传统的ASR,TTS和LLM系统。为了解决这个问题,我们推出了EchoVoices,这是一个端到端的数字人管道,致力于为老年人和儿童创建持久的数字角色,确保他们的声音和记忆为后代保留。我们的系统集成了三个核心创新:一个k-NN增强的Whisper模型,用于非典型语音的鲁棒语音识别;一个年龄自适应的VITS模型,用于高保真,说话者感知的语音合成;和一个LLM驱动的代理,自动生成人物卡,并利用基于RAG的记忆系统进行会话一致性。我们在SeniorTalk和ChildMandarin数据集上进行的实验表明,识别准确率,合成质量和说话人相似性都有显着提高。EchoVoices提供了一个全面的框架来保护代际声音,提供了一种新的代际联系方式,并创造了持久的数字遗产。
摘要:Recent breakthroughs in intelligent speech and digital human technologies have primarily targeted mainstream adult users, often overlooking the distinct vocal patterns and interaction styles of seniors and children. These demographics possess distinct vocal characteristics, linguistic styles, and interaction patterns that challenge conventional ASR, TTS, and LLM systems. To address this, we introduce EchoVoices, an end-to-end digital human pipeline dedicated to creating persistent digital personas for seniors and children, ensuring their voices and memories are preserved for future generations. Our system integrates three core innovations: a k-NN-enhanced Whisper model for robust speech recognition of atypical speech; an age-adaptive VITS model for high-fidelity, speaker-aware speech synthesis; and an LLM-driven agent that automatically generates persona cards and leverages a RAG-based memory system for conversational consistency. Our experiments, conducted on the SeniorTalk and ChildMandarin datasets, demonstrate significant improvements in recognition accuracy, synthesis quality, and speaker similarity. EchoVoices provides a comprehensive framework for preserving generational voices, offering a new means of intergenerational connection and the creation of lasting digital legacies.
【14】Exploiting Context-dependent Duration Features for Voice Anonymization Attack Systems
标题:利用语音匿名攻击系统的上下文相关持续时间特征
链接:http://arxiv.org/pdf/2507.15214v1
备注:Accepted at Interspeech-2025
摘要:语音的时间动态,包括节奏,语调和语速的变化,包含了关于说话人身份的重要和独特的信息。本文提出了一种新的方法来表示说话人特征提取上下文相关的持续时间嵌入语音时间动态。我们开发了新的攻击模型,使用这些表示和分析潜在的漏洞,在说话人验证和语音匿名化systems.Experimental结果表明,开发的攻击模型提供了显着改善说话人验证性能的原始和匿名数据相比,在文献中报道的语音时间动态的简单表示。
摘要:The temporal dynamics of speech, encompassing variations in rhythm, intonation, and speaking rate, contain important and unique information about speaker identity. This paper proposes a new method for representing speaker characteristics by extracting context-dependent duration embeddings from speech temporal dynamics. We develop novel attack models using these representations and analyze the potential vulnerabilities in speaker verification and voice anonymization systems.The experimental results show that the developed attack models provide a significant improvement in speaker verification performance for both original and anonymized data in comparison with simpler representations of speech temporal dynamics reported in the literature.
【15】Frame-level Temporal Difference Learning for Partial Deepfake Speech Detection
标题:用于部分Deepfake语音检测的帧级时间差异学习
链接:http://arxiv.org/pdf/2507.15101v1
备注:5 pages, 4 figures, 4 tables. Accepted to IEEE SPL
摘要:检测部分deepfake语音是必不可少的,因为它可能会产生微妙的错误信息。然而,现有方法在训练期间依赖于昂贵的帧级注释,限制了现实世界的可扩展性。此外,他们专注于检测bonafide和deepfake片段之间的过渡伪影。随着deepfake生成技术越来越平滑这些过渡,检测变得更具挑战性。为了解决这个问题,我们的工作通过分析帧级时间差异引入了一个新的视角,并揭示了与真正的语音相比,deepfake语音表现出不稳定的方向变化和不自然的局部过渡。基于这一发现,我们提出了一个时间差异注意力模块(TDAM),它将部分深度伪造检测重新定义为识别不自然的时间变化,而不依赖于显式的边界注释。一个双层的层次差异表示捕捉时间的不规则性,在精细和粗糙的尺度,而自适应平均池保存跨可变长度的输入,以尽量减少信息丢失的基本模式。我们的TDAM-AvgPool模型实现了最先进的性能,在PartialSpoof数据集上的EER为0.59%,在HAD数据集上为0.03%,显著优于现有方法,而无需帧级监督。摘要:Detecting partial deepfake speech is essential due to its potential for subtle misinformation. However, existing methods depend on costly frame-level annotations during training, limiting real-world scalability. Also, they focus on detecting transition artifacts between bonafide and deepfake segments. As deepfake generation techniques increasingly smooth these transitions, detection has become more challenging. To address this, our work introduces a new perspective by analyzing frame-level temporal differences and reveals that deepfake speech exhibits erratic directional changes and unnatural local transitions compared to bonafide speech. Based on this finding, we propose a Temporal Difference Attention Module (TDAM) that redefines partial deepfake detection as identifying unnatural temporal variations, without relying on explicit boundary annotations. A dual-level hierarchical difference representation captures temporal irregularities at both fine and coarse scales, while adaptive average pooling preserves essential patterns across variable-length inputs to minimize information loss. Our TDAM-AvgPool model achieves state-of-the-art performance, with an EER of 0.59% on the PartialSpoof dataset and 0.03% on the HAD dataset, which significantly outperforms the existing methods without requiring frame-level supervision.
【16】Music-Aligned Holistic 3D Dance Generation via Hierarchical Motion Modeling
标题:通过分层运动建模实现音乐对齐的整体3D舞蹈生成
链接:http://arxiv.org/pdf/2507.14915v1
摘要:协调良好,音乐一致的整体舞蹈增强了情感表达和观众的参与。然而,生成这样的舞蹈仍然具有挑战性,因为缺乏完整的3D舞蹈数据集,难以实现音乐和舞蹈之间的跨模态对齐,以及对身体,手和面部的相互依赖运动建模的复杂性。为了应对这些挑战,我们引入了SoulDance,这是一个通过专业动作捕捉系统捕捉的高精度音乐舞蹈配对数据集,具有精心注释的整体舞蹈动作。在此数据集的基础上,我们提出了SoulNet,一个旨在生成音乐对齐,运动学协调的整体舞蹈序列的框架。SoulNet由三个主要组件组成:(1)分层残差矢量量化,它对身体,手和面部的复杂,细粒度的运动依赖关系进行建模;(2)音乐对齐生成模型,它将这些分层运动单元组成表达和协调的整体舞蹈;(3)音乐-动作检索模块,一个预先训练的跨模态模型,用作音乐-舞蹈对齐先验,确保在整个生成过程中生成的舞蹈和输入音乐之间的时间同步和语义一致性。大量的实验表明,SoulNet在生成高质量、音乐协调和对齐良好的整体3D舞蹈序列方面明显优于现有方法。
摘要:Well-coordinated, music-aligned holistic dance enhances emotional expressiveness and audience engagement. However, generating such dances remains challenging due to the scarcity of holistic 3D dance datasets, the difficulty of achieving cross-modal alignment between music and dance, and the complexity of modeling interdependent motion across the body, hands, and face. To address these challenges, we introduce SoulDance, a high-precision music-dance paired dataset captured via professional motion capture systems, featuring meticulously annotated holistic dance movements. Building on this dataset, we propose SoulNet, a framework designed to generate music-aligned, kinematically coordinated holistic dance sequences. SoulNet consists of three principal components: (1) Hierarchical Residual Vector Quantization, which models complex, fine-grained motion dependencies across the body, hands, and face; (2) Music-Aligned Generative Model, which composes these hierarchical motion units into expressive and coordinated holistic dance; (3) Music-Motion Retrieval Module, a pre-trained cross-modal model that functions as a music-dance alignment prior, ensuring temporal synchronization and semantic coherence between generated dance and input music throughout the generation process. Extensive experiments demonstrate that SoulNet significantly surpasses existing approaches in generating high-quality, music-coordinated, and well-aligned holistic 3D dance sequences.
【17】Multi-Sampling-Frequency Naturalness MOS Prediction Using Self-Supervised Learning Model with Sampling-Frequency-Independent Layer
标题:使用具有采样频率独立层的自监督学习模型的多采样频率自然性MOS预测
链接:http://arxiv.org/pdf/2507.14647v1
备注:4 pages, 2 figures
摘要:我们介绍我们提交的AudioMOS挑战赛(AMC)2025轨道3:多采样频率(SF)语音的平均意见得分(MOS)预测。我们提交的模型将SF独立(SFI)卷积层集成到自监督学习(SSL)模型中,以实现用于MOS预测的SFI语音特征提取。我们提出了一些策略来提高模型的MOS预测性能:从预训练的非SFI-SSL模型中提取知识,并使用大规模MOS数据集进行预训练。我们提交给AMC 2025 Track 3的作品在一个评估指标中排名第一,在最终排名中排名第四。我们还报告了我们的消融研究的结果,以调查我们的模型的基本因素。
摘要:We introduce our submission to the AudioMOS Challenge (AMC) 2025 Track 3: mean opinion score (MOS) prediction for speech with multiple sampling frequencies (SFs). Our submitted model integrates an SF-independent (SFI) convolutional layer into a self-supervised learning (SSL) model to achieve SFI speech feature extraction for MOS prediction. We present some strategies to improve the MOS prediction performance of our model: distilling knowledge from a pretrained non-SFI-SSL model and pretraining with a large-scale MOS dataset. Our submission to the AMC 2025 Track 3 ranked the first in one evaluation metric and the fourth in the final ranking. We also report the results of our ablation study to investigate essential factors of our model.
【18】The Rest is Silence: Leveraging Unseen Species Models for Computational Musicology
标题:剩下的就是沉默:利用看不见的物种模型进行计算音乐学
链接:http://arxiv.org/pdf/2507.14638v1
摘要:几十年来,音乐学家一直致力于创建大型数据库,为音乐学研究和学术研究提供不同的目的。随着音乐信息检索和数字音乐学等领域的兴起,音乐学相关数据集和语料库不断涌入。然而,在历史或观察背景下,这些数据集必然是不完整的,感兴趣的收集的真实程度仍然未知-沉默。在这里,我们第一次将所谓的“看不见的物种”模型(USMs)从生态学应用到音乐活动领域。在正式介绍模型后,我们在四个案例研究中展示了如何将USMs应用于音乐学数据,以解决定量问题,如:我们在RISM中缺少多少作曲家?我们已经编目的中世纪格里高利圣咏的来源有多少?我们期望在不同的版本中找到多少音乐拷贝的差异?民间音乐传统流派的歌曲覆盖面有多大?最后,我们对大量作曲家的和声词汇量的估计有多接近?
摘要:For many decades, musicologists have engaged in creating large databases serving different purposes for musicological research and scholarship. With the rise of fields like music information retrieval and digital musicology, there is now a constant and growing influx of musicologically relevant datasets and corpora. In historical or observational settings, however, these datasets are necessarily incomplete, and the true extent of a collection of interest remains unknown -- silent. Here, we apply, for the first time, so-called Unseen Species models (USMs) from ecology to areas of musicological activity. After introducing the models formally, we show in four case studies how USMs can be applied to musicological data to address quantitative questions like: How many composers are we missing in RISM? What percentage of medieval sources of Gregorian chant have we already cataloged? How many differences in music prints do we expect to find between editions? How large is the coverage of songs from genres of a folk music tradition? And, finally, how close are we in estimating the size of the harmonic vocabulary of a large number of composers?
【19】U-DREAM: Unsupervised Dereverberation guided by a Reverberation Model
标题:U-DREAM:由混响模型引导的无监督去混响
链接:http://arxiv.org/pdf/2507.14237v1
备注:Submitted to IEEE Transactions on Audio, Speech and Language Processing (TASLPRO)
摘要:本文探讨了训练最先进的去混响模型的结果,其监督设置范围从弱监督到完全无监督,仅依赖于混响信号和声学模型进行训练。大多数现有的深度学习方法通常需要成对的干数据和混响数据,这在实践中很难获得。相反,我们开发了一种顺序学习策略,其动机是去混响问题的bavonet公式,其中声学参数和干信号是使用深度神经网络从混响输入中估计的,由混响匹配损失指导。我们的数据效率最高的变体只需要100个混响参数标记的样本就可以超越无监督基线,证明了所提出的方法在低资源场景中的有效性和实用性。摘要:This paper explores the outcome of training state-ofthe-art dereverberation models with supervision settings ranging from weakly-supervised to fully unsupervised, relying solely on reverberant signals and an acoustic model for training. Most of the existing deep learning approaches typically require paired dry and reverberant data, which are difficult to obtain in practice. We develop instead a sequential learning strategy motivated by a bayesian formulation of the dereverberation problem, wherein acoustic parameters and dry signals are estimated from reverberant inputs using deep neural networks, guided by a reverberation matching loss. Our most data-efficient variant requires only 100 reverberation-parameter-labelled samples to outperform an unsupervised baseline, demonstrating the effectiveness and practicality of the proposed method in low-resource scenarios.
【20】Developing an AI-Guided Assistant Device for the Deaf and Hearing Impaired
标题:开发一种用于聋人和听力障碍者的AI引导辅助设备
链接:http://arxiv.org/pdf/2507.14215v1
摘要:本研究旨在为聋人或听力受损者开发一种无障碍设备的深度学习系统。该设备将实时准确定位和识别声源。这项研究将填补当前研究的一个重要空白,利用机器学习技术来针对贫困社区。该系统包括三个主要组成部分。1. JerryNet:一种定制设计的CNN架构,可以确定九个可能方向的到达方向(DoA)。2.音频分类:该模型基于对对比音频预训练(CLAP)模型的微调,以仅基于音频来识别确切的声音类别。3.多式联运模式:这是一个精确的声音定位模型,它结合了音频,视觉和文本数据,以定位图像中的确切声源。该部分由两个模块组成,一个是使用Yolov9生成对象的所有边界框的对象检测模块,另一个是使用完整的Intersection over Union(CIoU)识别最佳边界框的视听定位模型。硬件包括一个四麦克风的矩形结构和一个安装在眼镜上的摄像头,眼镜上有一个腕带,用于显示必要的信息,如方向。在自定义收集的数据集上,JerryNet的精度达到了91。1%的声音方向,优于所有的基线模型。CLAP模型在自定义和AudioSet数据集上分别达到了98.5%和95%的准确率。组件3中的视听定位模型产生了0.892的cIoU和0.658的AUC,超过了其他类似模型。这项研究有许多未来的潜力,为创造新一代的无障碍设备铺平了道路。
摘要:This study aims to develop a deep learning system for an accessibility device for the deaf or hearing impaired. The device will accurately localize and identify sound sources in real time. This study will fill an important gap in current research by leveraging machine learning techniques to target the underprivileged community. The system includes three main components. 1. JerryNet: A custom designed CNN architecture that determines the direction of arrival (DoA) for nine possible directions. 2. Audio Classification: This model is based on fine-tuning the Contrastive Language-Audio Pretraining (CLAP) model to identify the exact sound classes only based on audio. 3. Multimodal integration model: This is an accurate sound localization model that combines audio, visual, and text data to locate the exact sound sources in the images. The part consists of two modules, one object detection using Yolov9 to generate all the bounding boxes of the objects, and an audio visual localization model to identify the optimal bounding box using complete Intersection over Union (CIoU). The hardware consists of a four-microphone rectangular formation and a camera mounted on glasses with a wristband for displaying necessary information like direction. On a custom collected data set, JerryNet achieved a precision of 91. 1% for the sound direction, outperforming all the baseline models. The CLAP model achieved 98.5% and 95% accuracy on custom and AudioSet datasets, respectively. The audio-visual localization model within component 3 yielded a cIoU of 0.892 and an AUC of 0.658, surpassing other similar models. There are many future potentials to this study, paving the way to creating a new generation of accessibility devices.
机器翻译由腾讯交互翻译提供,仅供参考
