微信公众号:arXiv_Daily
cs.SD语音
【1】Improving Audio Event Recognition with Consistency Regularization
标题:通过一致性规范化改进音频事件识别
链接:http://arxiv.org/pdf/2509.10391v1
摘要:一致性正则化(CR),它强制增强视图上的模型预测之间的一致性,最近在自动语音识别中发现了好处[1]。在本文中,我们提出了使用一致性正则化的音频事件识别,并证明其有效性AudioSet。通过对小型(2万美元)和大型(180万美元)监督训练集进行广泛的消融研究,我们表明CR比已经大量使用数据增强的监督基线带来了一致的改善,并且使用更强增强和多次增强的CR会为小型训练集带来额外的增益。此外,我们将CR的使用扩展到具有20K标记样本和1.8M未标记样本的半监督设置中,并获得了在小集合上训练的最佳模型的性能改进。摘要:Consistency regularization (CR), which enforces agreement between model predictions on augmented views, has found recent benefits in automatic speech recognition [1]. In this paper, we propose the use of consistency regularization for audio event recognition, and demonstrate its effectiveness on AudioSet. With extensive ablation studies for both small ($ sim$20k) and large ($ sim$1.8M) supervised training sets, we show that CR brings consistent improvement over supervised baselines which already heavily utilize data augmentation, and CR using stronger augmentation and multiple augmentations leads to additional gain for the small training set. Furthermore, we extend the use of CR into the semi-supervised setup with 20K labeled samples and 1.8M unlabeled samples, and obtain performance improvement over our best model trained on the small set.
【2】Data-independent Beamforming for End-to-end Multichannel Multi-speaker ASR
标题:用于端到端多通道多扬声器ASB的数据独立的射束形成
链接:http://arxiv.org/pdf/2509.10234v1
摘要:由于环境噪声、混响和重叠扬声器的存在,多通道、多说话人场景中的自动语音识别(ASR)仍然具有挑战性。在本文中,我们提出了一种波束成形方法,在应用端到端多通道,多扬声器ASR系统之前,根据其球面极坐标处理特定的角度扇区。该方法具有数据无关性和无需训练的特点。我们证明,使用一组波束形成的信号提高ASR性能相比,使用相同数量的原始麦克风信号。此外,增加用于波束成形的信号的数量进一步增强了识别准确性,从而导致更有效地使用多通道信号,同时降低了ASR系统的总输入负载。我们进行实验的AMI会议语料库,其中所提出的方法减少了高达11%的单词错误率和提高扬声器计数精度高达27%,相对于一个多通道ASR基线系统,不利用波束成形。摘要:Automatic speech recognition (ASR) in multichannel, multi-speaker scenarios remains challenging due to ambient noise, reverberation and overlapping speakers. In this paper, we propose a beamforming approach that processes specific angular sectors based on their spherical polar coordinates before applying an end-to-end multichannel, multi-speaker ASR system. This method is data-independent and training-free. We demonstrate that using a group of beamformed signals improves ASR performance compared to using the same number of raw microphone signals. Moreover, increasing the number of signals used for beamforming further enhances recognition accuracy, leading to a more efficient use of multichannel signals while reducing the overall input load for the ASR system. We conduct experiments on the AMI meeting corpus, where the proposed method reduces word error rate by up to 11% and improves speaker counting accuracy by up to 27% relative compared to a multichannel ASR baseline system that does not exploit beamforming.
【3】Prototypical Contrastive Learning For Improved Few-Shot Audio Classification
标题:用于改进Few-Shot音频分类的原型对比学习
链接:http://arxiv.org/pdf/2509.10074v1
摘要:Few-Shot学习已经成为训练具有有限标记数据的模型的强大范例,解决了大规模注释不切实际的情况下的挑战。虽然在图像领域已经进行了广泛的研究,但音频分类中的Few-Shot学习仍然相对不足。在这项工作中,我们研究了将监督对比度损失集成到用于音频分类的原型Few Shot训练中的效果。详细地,我们证明了角损耗进一步提高了性能相比,标准的对比度损失。我们的方法利用SpecAugment,其次是一个自我注意机制,以封装不同的信息增强输入版本到一个统一的嵌入。我们在MetaAudio上评估我们的方法,MetaAudio是一个基准测试,包括五个具有预定义分割的数据集,标准化预处理和一组全面的Few-Shot学习模型进行比较。所提出的方法在5路、5次拍摄设置中实现了最先进的性能。摘要:Few-shot learning has emerged as a powerful paradigm for training models with limited labeled data, addressing challenges in scenarios where large-scale annotation is impractical. While extensive research has been conducted in the image domain, few-shot learning in audio classification remains relatively underexplored. In this work, we investigate the effect of integrating supervised contrastive loss into prototypical few shot training for audio classification. In detail, we demonstrate that angular loss further improves the performance compared to the standard contrastive loss. Our method leverages SpecAugment followed by a self-attention mechanism to encapsulate diverse information of augmented input versions into one unified embedding. We evaluate our approach on MetaAudio, a benchmark including five datasets with predefined splits, standardized preprocessing, and a comprehensive set of few-shot learning models for comparison. The proposed approach achieves state-of-the-art performance in a 5-way, 5-shot setting.
【4】CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio
标题:CoDiCodec:统一音频的连续和离散压缩表示
链接:http://arxiv.org/pdf/2509.09836v1
摘要:在压缩的潜在空间中有效地表示音频信号对于潜在生成建模是至关重要的。然而,现有的自动编码器通常强制在连续嵌入和离散标记之间进行选择。此外,在保持音频保真度的同时实现高压缩比仍然是一个挑战。我们介绍了一种新型的音频自动编码器,它克服了这些限制,既通过摘要嵌入有效地编码全局特征,又通过从相同的训练模型中以2.38 kbps的速率产生约11 Hz的压缩连续嵌入和离散令牌,为不同的下游生成任务提供了前所未有的灵活性。这是通过有限标量量化(FSQ)和一种新的FSQ-丢弃技术实现的,并且除了用于端到端训练的单一一致性损失之外,不需要额外的损失项。CoDiCodec支持自回归解码和新型并行解码策略,后者可实现卓越的音频质量和更快的解码速度。在重建音频质量方面,CoDiCodec在类似比特率下优于现有的连续和离散自动编码器。我们的工作实现了音频压缩的统一方法,弥合了连续和离散生成建模范式之间的差距。摘要:Efficiently representing audio signals in a compressed latent space is critical for latent generative modelling. However, existing autoencoders often force a choice between continuous embeddings and discrete tokens. Furthermore, achieving high compression ratios while maintaining audio fidelity remains a challenge. We introduce CoDiCodec, a novel audio autoencoder that overcomes these limitations by both efficiently encoding global features via summary embeddings, and by producing both compressed continuous embeddings at ~ 11 Hz and discrete tokens at a rate of 2.38 kbps from the same trained model, offering unprecedented flexibility for different downstream generative tasks. This is achieved through Finite Scalar Quantization (FSQ) and a novel FSQ-dropout technique, and does not require additional loss terms beyond the single consistency loss used for end-to-end training. CoDiCodec supports both autoregressive decoding and a novel parallel decoding strategy, with the latter achieving superior audio quality and faster decoding. CoDiCodec outperforms existing continuous and discrete autoencoders at similar bitrates in terms of reconstruction audio quality. Our work enables a unified approach to audio compression, bridging the gap between continuous and discrete generative modelling paradigms.
【5】SoilSound: Smartphone-based Soil Moisture Estimation
标题:SoilSound:基于智能手机的土壤湿度估计
链接:http://arxiv.org/pdf/2509.09823v1
摘要:土壤湿度监测对农业和环境管理至关重要,但现有方法要么需要侵入式探头干扰土壤,要么需要专门设备,限制了公众的使用。我们提出了SoilSound,一个无处不在的可访问的基于智能手机的声学传感系统,可以测量土壤水分,而不会干扰土壤。我们利用内置扬声器和麦克风执行垂直扫描机制,无需任何校准即可准确测量水分。与现有的使用透射特性的工作不同,我们提出了一种替代模型,土壤中的声反射的基础上的表面粗糙度效应,使水分传感不干扰土壤。该系统的工作原理是向土壤发送声学啁啾,并在垂直扫描期间记录反射,然后将其处理并馈送到卷积神经网络,用于设备上的土壤湿度估计,计算,内存或功率开销可以忽略不计。我们通过在实验室的盒子中使用精选土壤进行训练并在室外场地进行测试来评估该系统,结果表明SoilSound在10个不同位置的平均绝对误差(MAE)为2.39%。总体而言,评估表明SoilSound可以在多种土壤类型,环境和用户中准确跟踪土壤水分水平,范围从15.9%到34.0%;无需任何校准或干扰土壤,从而为家庭园丁,城市农民,公民科学家和资源有限环境中的农业社区提供广泛的水分监测。摘要:Soil moisture monitoring is essential for agriculture and environmental management, yet existing methods require either invasive probes disturbing the soil or specialized equipment, limiting access to the public. We present SoilSound, an ubiquitous accessible smartphone-based acoustic sensing system that can measure soil moisture without disturbing the soil. We leverage the built-in speaker and microphone to perform a vertical scan mechanism to accurately measure moisture without any calibration. Unlike existing work that use transmissive properties, we propose an alternate model for acoustic reflections in soil based on the surface roughness effect to enable moisture sensing without disturbing the soil. The system works by sending acoustic chirps towards the soil and recording the reflections during a vertical scan, which are then processed and fed to a convolutional neural network for on-device soil moisture estimation with negligible computational, memory, or power overhead. We evaluated the system by training with curated soils in boxes in the lab and testing in the outdoor fields and show that SoilSound achieves a mean absolute error (MAE) of 2.39% across 10 different locations. Overall, the evaluation shows that SoilSound can accurately track soil moisture levels ranging from 15.9% to 34.0% across multiple soil types, environments, and users; without requiring any calibration or disturbing the soil, enabling widespread moisture monitoring for home gardeners, urban farmers, citizen scientists, and agricultural communities in resource-limited settings.
【6】Combining Textual and Spectral Features for Robust Classification of Pilot Communications
标题:结合文本和频谱特征对导频通信进行稳健分类
链接:http://arxiv.org/pdf/2509.09752v1
【7】DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
标题:DiTReducio:通过渐进式校准为基于DiT的TTC提供免训练加速
链接:http://arxiv.org/pdf/2509.09748v1
【8】AI-enabled tuberculosis screening in a high-burden setting using cough sound analysis and speech foundation models
标题:使用咳嗽声音分析和言语基础模型在高负担环境中进行人工智能支持的结核病筛查
链接:http://arxiv.org/pdf/2509.09746v1
摘要:背景 人工智能(AI)可以检测咳嗽声中与疾病相关的声学模式,为高负担、低资源环境中的结核病(TB)筛查提供了一种可扩展的方法。以前的研究受到小数据集、有症状的非结核病患者代表性不足、依赖简单模型以及在理想条件下收集的记录的限制。 方法 我们在赞比亚的两家医院招募了512名参与者,分为细菌学确诊的结核病(TB+),其他呼吸道疾病(OR)的症状患者和健康对照(HC)。从500名参与者中获得可用的咳嗽记录以及人口统计学和临床数据。基于语音基础模型的深度学习分类器在咳嗽记录上进行训练。在3秒片段上训练的表现最好的模型,进一步评估了人口统计学和临床特征。 结果 最好的仅音频分类器在区分TB+与所有其他(TB + Rest)方面实现了85.2%的AUROC,在区分TB+与OR方面实现了80.1%的AUROC。添加人口统计学和临床特征将性能提高到92.1%(TB + Rest)和84.2%(TB + OR)。在阈值为0.38时,多模式模型对TB + Rest的敏感性和特异性分别达到90.3%和73.1%,对TB + OR的敏感性和特异性分别达到80.6%和73.1%。 解释 使用语音基础模型进行咳嗽分析,特别是结合人口统计和临床数据时,显示出作为结核病分诊工具的强大潜力,符合世卫组织目标产品概况基准。该模型对混杂因素(包括背景噪声、记录时间和器械变异性)具有鲁棒性,表明检测到了真正的疾病相关声学模式。在临床使用之前,需要在不同地区和病例定义(包括亚临床TB)中进行进一步验证。摘要:Background Artificial intelligence (AI) can detect disease-related acoustic patterns in cough sounds, offering a scalable approach to tuberculosis (TB) screening in high-burden, low-resource settings. Previous studies have been limited by small datasets, under-representation of symptomatic non-TB patients, reliance on simple models, and recordings collected under idealised conditions. Methods We enrolled 512 participants at two hospitals in Zambia, grouped as bacteriologically confirmed TB (TB+), symptomatic patients with other respiratory diseases (OR), and healthy controls (HC). Usable cough recordings plus demographic and clinical data were obtained from 500 participants. Deep learning classifiers based on speech foundation models were trained on cough recordings. The best-performing model, trained on 3-second segments, was further evaluated with demographic and clinical features. Findings The best audio-only classifier achieved an AUROC of 85.2% for distinguishing TB+ from all others (TB+ Rest) and 80.1% for TB+ versus OR. Adding demographic and clinical features improved performance to 92.1% (TB+ Rest) and 84.2% (TB+ OR). At a threshold of 0.38, the multimodal model reached 90.3% sensitivity and 73.1% specificity for TB+ Rest, and 80.6% and 73.1% for TB+ OR. Interpretation Cough analysis using speech foundation models, especially when combined with demographic and clinical data, showed strong potential as a TB triage tool, meeting WHO target product profile benchmarks. The model was robust to confounding factors including background noise, recording time, and device variability, indicating detection of genuine disease-related acoustic patterns. Further validation across diverse regions and case definitions, including subclinical TB, is required before clinical use.
【9】Testing chatbots on the creation of encoders for audio conditioned image generation
标题:测试聊天机器人创建音频调节图像生成编码器
链接:http://arxiv.org/pdf/2509.09717v1
【10】VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions
标题:VStyle:语音风格改编的基准
链接:http://arxiv.org/pdf/2509.09716v1
【11】TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation
标题:TalkPlayData 2:用于多模式对话音乐推荐的大型合成数据管道
链接:http://arxiv.org/pdf/2509.09685v1
【12】Error Analysis in a Modular Meeting Transcription System
标题:模块化会议转录系统中的错误分析
链接:http://arxiv.org/pdf/2509.10143v1
摘要:会议记录是近年来发展迅速的一个研究领域。然而,挑战依然存在,限制了其业绩。在这项工作中,我们扩展了以前提出的框架,用于分析泄漏语音分离与适当的敏感性时间局部性。我们发现,有显着的泄漏到交叉通道的地区,只有主扬声器是活跃的。同时,结果表明,这不会影响最终的性能,因为这些泄漏的部分在很大程度上被语音活动检测(VAD)忽略。此外,不同的分割进行了比较,表明先进的日记化方法能够减少的差距,甲骨文分割的三分之一相比,一个简单的基于能量的VAD。我们还揭示了哪些因素导致了剩余的差异。这些结果代表了LibriCSS在仅对LibriSpeech数据进行训练识别模块的系统中的最新性能。摘要:Meeting transcription is a field of high relevance and remarkable progress in recent years. Still, challenges remain that limit its performance. In this work, we extend a previously proposed framework for analyzing leakage in speech separation with proper sensitivity to temporal locality. We show that there is significant leakage to the cross channel in areas where only the primary speaker is active. At the same time, the results demonstrate that this does not affect the final performance much as these leaked parts are largely ignored by the voice activity detection (VAD). Furthermore, different segmentations are compared showing that advanced diarization approaches are able to reduce the gap to oracle segmentation by a third compared to a simple energy-based VAD. We additionally reveal what factors contribute to the remaining difference. The results represent state-of-the-art performance on LibriCSS among systems that train the recognition module on LibriSpeech data only.
【13】Unified Learnable 2D Convolutional Feature Extraction for ASR
标题:用于ASB的统一可学习2D卷积特征提取
链接:http://arxiv.org/pdf/2509.10031v1
摘要:神经前端代表了自动语音识别(ASR)系统特征提取的一种有前途的方法,因为它们能够为不同的任务学习专门定制的特征。然而,许多现有的技术仍然受到经典方法的严重影响。虽然这种归纳偏差可能会简化系统设计,但我们的工作旨在开发一个更通用的特征提取前端。此外,我们寻求统一前端架构,与应用源自不同来源的多层拓扑组合的现有方法形成鲜明对比。实验系统地展示了如何减少现有技术的影响,以实现通用的前端。由此产生的2D卷积前端是参数高效的,适用于计算资源有限的场景,而不像在未标记音频上预先训练的大型模型。结果表明,这种通用的统一方法不仅是可行的,而且与现有的监督可学习特征提取器的性能相匹配。摘要:Neural front-ends represent a promising approach to feature extraction for automatic speech recognition (ASR) systems as they enable to learn specifically tailored features for different tasks. Yet, many of the existing techniques remain heavily influenced by classical methods. While this inductive bias may ease the system design, our work aims to develop a more generic front-end for feature extraction. Furthermore, we seek to unify the front-end architecture contrasting with existing approaches that apply a composition of several layer topologies originating from different sources. The experiments systematically show how to reduce the influence of existing techniques to achieve a generic front-end. The resulting 2D convolutional front-end is parameter-efficient and suitable for a scenario with limited computational resources unlike large models pre-trained on unlabeled audio. The results demonstrate that this generic unified approach is not only feasible but also matches the performance of existing supervised learnable feature extractors.
【14】The MSP-Podcast Corpus
标题:MSP-播客素材库
链接:http://arxiv.org/pdf/2509.09791v1
摘要:大的,高质量的情感语音数据库的可用性是在现实世界中推进语音情感识别(SER)的必要条件。然而,许多现有的数据库面临着规模,情感平衡和扬声器多样性的限制。本研究描述了MSP播客语料库,总结了我们十年的努力。该语料库由来自各种音频共享网站的400多个小时的不同音频样本组成,所有这些网站都具有允许分发语料库的通用许可证。我们用丰富的情感标签注释语料库,包括主要(单一主导情感)和次要(音频中感知的多种情感)情感类别,以及效价,唤醒和主导的情感属性。至少有五位评分员对这些情感标签进行了注释。语料库还具有针对大多数样本的说话人识别,以及针对整个语料库的句子的词汇内容的人工翻译。数据收集协议包括一个机器学习驱动的管道,用于选择情感多样的录音,确保在说话者和环境中平衡和多样化地表达情感。由此产生的数据库提供了一个全面的,高质量的资源,更适合于推进SER系统在实际,现实世界的情况。摘要:The availability of large, high-quality emotional speech databases is essential for advancing speech emotion recognition (SER) in real-world scenarios. However, many existing databases face limitations in size, emotional balance, and speaker diversity. This study describes the MSP-Podcast corpus, summarizing our ten-year effort. The corpus consists of over 400 hours of diverse audio samples from various audio-sharing websites, all of which have Common Licenses that permit the distribution of the corpus. We annotate the corpus with rich emotional labels, including primary (single dominant emotion) and secondary (multiple emotions perceived in the audio) emotional categories, as well as emotional attributes for valence, arousal, and dominance. At least five raters annotate these emotional labels. The corpus also has speaker identification for most samples, and human transcriptions of the lexical content of the sentences for the entire corpus. The data collection protocol includes a machine learning-driven pipeline for selecting emotionally diverse recordings, ensuring a balanced and varied representation of emotions across speakers and environments. The resulting database provides a comprehensive, high-quality resource, better suited for advancing SER systems in practical, real-world scenarios.
【15】Spectral Bottleneck in Deep Neural Networks: Noise is All You Need
标题:深度神经网络中的频谱瓶颈:噪音就是你所需要的一切
链接:http://arxiv.org/pdf/2509.09719v1
【1】Low-latency Assistive Audio Enhancement for Neurodivergent People
标题:针对神经分歧患者的低延迟辅助音频增强
链接:http://arxiv.org/pdf/2509.10202v1
【2】Error Analysis in a Modular Meeting Transcription System
标题:模块化会议转录系统中的错误分析
链接:http://arxiv.org/pdf/2509.10143v1
摘要:会议记录是近年来发展迅速的一个研究领域。然而,挑战依然存在,限制了其业绩。在这项工作中,我们扩展了以前提出的框架,用于分析泄漏语音分离与适当的敏感性时间局部性。我们发现,有显着的泄漏到交叉通道的地区,只有主扬声器是活跃的。同时,结果表明,这不会影响最终的性能,因为这些泄漏的部分在很大程度上被语音活动检测(VAD)忽略。此外,不同的分割进行了比较,表明先进的日记化方法能够减少的差距,甲骨文分割的三分之一相比,一个简单的基于能量的VAD。我们还揭示了哪些因素导致了剩余的差异。这些结果代表了LibriCSS在仅对LibriSpeech数据进行训练识别模块的系统中的最新性能。摘要:Meeting transcription is a field of high relevance and remarkable progress in recent years. Still, challenges remain that limit its performance. In this work, we extend a previously proposed framework for analyzing leakage in speech separation with proper sensitivity to temporal locality. We show that there is significant leakage to the cross channel in areas where only the primary speaker is active. At the same time, the results demonstrate that this does not affect the final performance much as these leaked parts are largely ignored by the voice activity detection (VAD). Furthermore, different segmentations are compared showing that advanced diarization approaches are able to reduce the gap to oracle segmentation by a third compared to a simple energy-based VAD. We additionally reveal what factors contribute to the remaining difference. The results represent state-of-the-art performance on LibriCSS among systems that train the recognition module on LibriSpeech data only.
【3】Towards Data Drift Monitoring for Speech Deepfake Detection in the context of MLOps
标题:MLOps背景下语音Deepfake检测的数据漂移监控
链接:http://arxiv.org/pdf/2509.10086v1
摘要:当在云上的应用程序或服务中交付时,未更新的静态语音deepfake检测器将容易受到新创建的语音deepfake攻击。从机器学习操作(MLOps)的角度来看,本文试图回答我们是否可以监控新的和看不见的语音deepfake数据,这些数据偏离了可见的参考数据集。我们进一步询问,如果检测到漂移,我们是否可以使用类似的漂移数据微调检测器,减少漂移,并提高检测性能。在玩具数据集和大规模MLAAD数据集上,我们证明了新的文本到语音(TTS)攻击引起的漂移可以使用新数据和参考数据的分布之间的距离进行监测。此外,我们证明了使用新TTS deepfakes生成的数据微调检测器可以减少漂移和检测错误率。摘要:When being delivered in applications or services on the cloud, static speech deepfake detectors that are not updated will become vulnerable to newly created speech deepfake attacks. From the perspective of machine learning operations (MLOps), this paper tries to answer whether we can monitor new and unseen speech deepfake data that drifts away from a seen reference data set. We further ask, if drift is detected, whether we can fine-tune the detector using similarly drifted data, reduce the drift, and improve the detection performance. On a toy dataset and the large-scale MLAAD dataset, we show that the drift caused by new text-to-speech (TTS) attacks can be monitored using distances between the distributions of the new data and reference data. Furthermore, we demonstrate that fine-tuning the detector using data generated by the new TTS deepfakes can reduce the drift and the detection error rates.
【4】Unified Learnable 2D Convolutional Feature Extraction for ASR
标题:用于ASB的统一可学习2D卷积特征提取
链接:http://arxiv.org/pdf/2509.10031v1
摘要:神经前端代表了自动语音识别(ASR)系统特征提取的一种有前途的方法,因为它们能够为不同的任务学习专门定制的特征。然而,许多现有的技术仍然受到经典方法的严重影响。虽然这种归纳偏差可能会简化系统设计,但我们的工作旨在开发一个更通用的特征提取前端。此外,我们寻求统一前端架构,与应用源自不同来源的多层拓扑组合的现有方法形成鲜明对比。实验系统地展示了如何减少现有技术的影响,以实现通用的前端。由此产生的2D卷积前端是参数高效的,适用于计算资源有限的场景,而不像在未标记音频上预先训练的大型模型。结果表明,这种通用的统一方法不仅是可行的,而且与现有的监督可学习特征提取器的性能相匹配。摘要:Neural front-ends represent a promising approach to feature extraction for automatic speech recognition (ASR) systems as they enable to learn specifically tailored features for different tasks. Yet, many of the existing techniques remain heavily influenced by classical methods. While this inductive bias may ease the system design, our work aims to develop a more generic front-end for feature extraction. Furthermore, we seek to unify the front-end architecture contrasting with existing approaches that apply a composition of several layer topologies originating from different sources. The experiments systematically show how to reduce the influence of existing techniques to achieve a generic front-end. The resulting 2D convolutional front-end is parameter-efficient and suitable for a scenario with limited computational resources unlike large models pre-trained on unlabeled audio. The results demonstrate that this generic unified approach is not only feasible but also matches the performance of existing supervised learnable feature extractors.
【5】Whisper Has an Internal Word Aligner
标题:Whisper有一个内部单词对齐器
链接:http://arxiv.org/pdf/2509.09987v1
摘要:从强大的自动语音识别器(特别是Whisper)中获得准确的单词级时间戳的兴趣越来越大。现有的方法要么需要额外的培训,要么根本没有竞争力。在以前的工作中的评估也是相对宽松的,通常使用超过200毫秒的公差。在这项工作中,我们发现注意头耳语捕捉准确的单词对齐,是明显不同于那些不。此外,我们发现,使用字符产生更精细,更准确的路线比使用单词。基于这些发现,我们提出了一种无监督的方法来提取词对齐过滤注意头,而教师强迫耳语字符。我们的方法不仅不需要训练,而且在20 ms和100 ms之间的更严格的公差下产生比先前工作更准确的单词对齐。摘要:There is an increasing interest in obtaining accurate word-level timestamps from strong automatic speech recognizers, in particular Whisper. Existing approaches either require additional training or are simply not competitive. The evaluation in prior work is also relatively loose, typically using a tolerance of more than 200 ms. In this work, we discover attention heads in Whisper that capture accurate word alignments and are distinctively different from those that do not. Moreover, we find that using characters produces finer and more accurate alignments than using wordpieces. Based on these findings, we propose an unsupervised approach to extracting word alignments by filtering attention heads while teacher forcing Whisper with characters. Our approach not only does not require training but also produces word alignments that are more accurate than prior work under a stricter tolerance between 20 ms and 100 ms.
【6】Effective Modeling of Critical Contextual Information for TDNN-based Speaker Verification
标题:基于TDNN的说话人验证关键上下文信息的有效建模
链接:http://arxiv.org/pdf/2509.09932v1
摘要:目前,时延神经网络(TDNN)已成为说话人确认的主流结构,其中ECAPA-TDNN是最先进的模型之一。目前致力于改进TDNN的工作主要是解决TDNN在建模全局信息方面的局限性,并弥合TDNN和二维卷积之间的差距。然而,ECAPA-TDNN提出的SE-Res 2Block中的分层卷积结构不能充分利用上下文信息,导致ECAPA-TDNN对有效上下文依赖建模的能力较弱。为此,提出了三种基于ECAPA-TDNN的改进结构,以充分有效地提取具有上下文依赖性的多尺度特征,然后将这些特征聚合。在VoxCeleb和CN-Celeb上的实验结果验证了三种结构的有效性。其中一种架构在VoxCeleb 1-O数据集上实现了比ECAPA-TDNN低近23%的等误率,证明了在可比参数计数下当前TDNN架构中可实现的竞争性能。摘要:Today, Time Delay Neural Network (TDNN) has become the mainstream architecture for speaker verification task, in which the ECAPA-TDNN is one of the state-of-the-art models. The current works that focus on improving TDNN primarily address the limitations of TDNN in modeling global information and bridge the gap between TDNN and 2-Dimensional convolutions. However, the hierarchical convolutional structure in the SE-Res2Block proposed by ECAPA-TDNN cannot make full use of the contextual information, resulting in the weak ability of ECAPA-TDNN to model effective context dependencies. To this end, three improved architectures based on ECAPA-TDNN are proposed to fully and effectively extract multi-scale features with context dependence and then aggregate these features. The experimental results on VoxCeleb and CN-Celeb verify the effectiveness of the three proposed architectures. One of these architectures achieves nearly a 23% lower Equal Error Rate compared to that of ECAPA-TDNN on VoxCeleb1-O dataset, demonstrating the competitive performance achievable among the current TDNN architectures under the comparable parameter count.
【7】Acoustic Scene Classification Using CNN-GRU Model Without Knowledge Distillation
标题:无需知识提炼的CNN-GRU模型声学场景分类
链接:http://arxiv.org/pdf/2509.09931v1
摘要:在本技术报告中,我们介绍了SNTL-NTU团队为低复杂度声学场景和事件(DCASE)2025挑战提交的任务1。这个提交偏离了从教师到学生模型的知识蒸馏的典型应用,旨在以有限的复杂性实现高性能。所提出的模型基于CNN-GRU模型,仅使用TAU Urban Acoustic Scene 2022 Mobile开发数据集进行训练,除了用于设备脉冲响应(ERF)增强的MicIRP之外,不使用任何外部数据集。该模型的内存使用量为114.2KB,需要10.9M乘法累加(MAC)操作。使用开发数据集,该模型实现了60.25%的准确率。摘要:In this technical report, we present the SNTL-NTU team's Task 1 submission for the Low-Complexity Acoustic Scenes and Events (DCASE) 2025 challenge. This submission departs from the typical application of knowledge distillation from a teacher to a student model, aiming to achieve high performance with limited complexity. The proposed model is based on a CNN-GRU model and is trained solely using the TAU Urban Acoustic Scene 2022 Mobile development dataset, without utilizing any external datasets, except for MicIRP, which is used for device impulse response (DIR) augmentation. The proposed model has a memory usage of 114.2KB and requires 10.9M muliply-and-accumulate (MAC) operations. Using the development dataset, the proposed model achieved an accuracy of 60.25%.
【8】The MSP-Podcast Corpus
标题:MSP-播客素材库
链接:http://arxiv.org/pdf/2509.09791v1
摘要:大的,高质量的情感语音数据库的可用性是在现实世界中推进语音情感识别(SER)的必要条件。然而,许多现有的数据库面临着规模,情感平衡和扬声器多样性的限制。本研究描述了MSP播客语料库,总结了我们十年的努力。该语料库由来自各种音频共享网站的400多个小时的不同音频样本组成,所有这些网站都具有允许分发语料库的通用许可证。我们用丰富的情感标签注释语料库,包括主要(单一主导情感)和次要(音频中感知的多种情感)情感类别,以及效价,唤醒和主导的情感属性。至少有五位评分员对这些情感标签进行了注释。语料库还具有针对大多数样本的说话人识别,以及针对整个语料库的句子的词汇内容的人工翻译。数据收集协议包括一个机器学习驱动的管道,用于选择情感多样的录音,确保在说话者和环境中平衡和多样化地表达情感。由此产生的数据库提供了一个全面的,高质量的资源,更适合于推进SER系统在实际,现实世界的情况。摘要:The availability of large, high-quality emotional speech databases is essential for advancing speech emotion recognition (SER) in real-world scenarios. However, many existing databases face limitations in size, emotional balance, and speaker diversity. This study describes the MSP-Podcast corpus, summarizing our ten-year effort. The corpus consists of over 400 hours of diverse audio samples from various audio-sharing websites, all of which have Common Licenses that permit the distribution of the corpus. We annotate the corpus with rich emotional labels, including primary (single dominant emotion) and secondary (multiple emotions perceived in the audio) emotional categories, as well as emotional attributes for valence, arousal, and dominance. At least five raters annotate these emotional labels. The corpus also has speaker identification for most samples, and human transcriptions of the lexical content of the sentences for the entire corpus. The data collection protocol includes a machine learning-driven pipeline for selecting emotionally diverse recordings, ensuring a balanced and varied representation of emotions across speakers and environments. The resulting database provides a comprehensive, high-quality resource, better suited for advancing SER systems in practical, real-world scenarios.
【9】Spectral Bottleneck in Deep Neural Networks: Noise is All You Need
标题:深度神经网络中的频谱瓶颈:噪音就是你所需要的一切
链接:http://arxiv.org/pdf/2509.09719v1
【10】Prominence-aware automatic speech recognition for conversational speech
标题:会话语音的突出感知自动语音识别
链接:http://arxiv.org/pdf/2509.10116v1
【11】CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio
标题:CoDiCodec:统一音频的连续和离散压缩表示
链接:http://arxiv.org/pdf/2509.09836v1
摘要:在压缩的潜在空间中有效地表示音频信号对于潜在生成建模是至关重要的。然而,现有的自动编码器通常强制在连续嵌入和离散标记之间进行选择。此外,在保持音频保真度的同时实现高压缩比仍然是一个挑战。我们介绍了一种新型的音频自动编码器,它克服了这些限制,既通过摘要嵌入有效地编码全局特征,又通过从相同的训练模型中以2.38 kbps的速率产生约11 Hz的压缩连续嵌入和离散令牌,为不同的下游生成任务提供了前所未有的灵活性。这是通过有限标量量化(FSQ)和一种新的FSQ-丢弃技术实现的,并且除了用于端到端训练的单一一致性损失之外,不需要额外的损失项。CoDiCodec支持自回归解码和新型并行解码策略,后者可实现卓越的音频质量和更快的解码速度。在重建音频质量方面,CoDiCodec在类似比特率下优于现有的连续和离散自动编码器。我们的工作实现了音频压缩的统一方法,弥合了连续和离散生成建模范式之间的差距。摘要:Efficiently representing audio signals in a compressed latent space is critical for latent generative modelling. However, existing autoencoders often force a choice between continuous embeddings and discrete tokens. Furthermore, achieving high compression ratios while maintaining audio fidelity remains a challenge. We introduce CoDiCodec, a novel audio autoencoder that overcomes these limitations by both efficiently encoding global features via summary embeddings, and by producing both compressed continuous embeddings at ~ 11 Hz and discrete tokens at a rate of 2.38 kbps from the same trained model, offering unprecedented flexibility for different downstream generative tasks. This is achieved through Finite Scalar Quantization (FSQ) and a novel FSQ-dropout technique, and does not require additional loss terms beyond the single consistency loss used for end-to-end training. CoDiCodec supports both autoregressive decoding and a novel parallel decoding strategy, with the latter achieving superior audio quality and faster decoding. CoDiCodec outperforms existing continuous and discrete autoencoders at similar bitrates in terms of reconstruction audio quality. Our work enables a unified approach to audio compression, bridging the gap between continuous and discrete generative modelling paradigms.
【12】Combining Textual and Spectral Features for Robust Classification of Pilot Communications
标题:结合文本和频谱特征对导频通信进行稳健分类
链接:http://arxiv.org/pdf/2509.09752v1
【13】DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
标题:DiTReducio:通过渐进式校准为基于DiT的TTC提供免训练加速
链接:http://arxiv.org/pdf/2509.09748v1
【14】AI-enabled tuberculosis screening in a high-burden setting using cough sound analysis and speech foundation models
标题:使用咳嗽声音分析和言语基础模型在高负担环境中进行人工智能支持的结核病筛查
链接:http://arxiv.org/pdf/2509.09746v1
摘要:背景 人工智能(AI)可以检测咳嗽声中与疾病相关的声学模式,为高负担、低资源环境中的结核病(TB)筛查提供了一种可扩展的方法。以前的研究受到小数据集、有症状的非结核病患者代表性不足、依赖简单模型以及在理想条件下收集的记录的限制。 方法 我们在赞比亚的两家医院招募了512名参与者,分为细菌学确诊的结核病(TB+),其他呼吸道疾病(OR)的症状患者和健康对照(HC)。从500名参与者中获得可用的咳嗽记录以及人口统计学和临床数据。基于语音基础模型的深度学习分类器在咳嗽记录上进行训练。在3秒片段上训练的表现最好的模型,进一步评估了人口统计学和临床特征。 结果 最好的仅音频分类器在区分TB+与所有其他(TB + Rest)方面实现了85.2%的AUROC,在区分TB+与OR方面实现了80.1%的AUROC。添加人口统计学和临床特征将性能提高到92.1%(TB + Rest)和84.2%(TB + OR)。在阈值为0.38时,多模式模型对TB + Rest的敏感性和特异性分别达到90.3%和73.1%,对TB + OR的敏感性和特异性分别达到80.6%和73.1%。 解释 使用语音基础模型进行咳嗽分析,特别是结合人口统计和临床数据时,显示出作为结核病分诊工具的强大潜力,符合世卫组织目标产品概况基准。该模型对混杂因素(包括背景噪声、记录时间和器械变异性)具有鲁棒性,表明检测到了真正的疾病相关声学模式。在临床使用之前,需要在不同地区和病例定义(包括亚临床TB)中进行进一步验证。摘要:Background Artificial intelligence (AI) can detect disease-related acoustic patterns in cough sounds, offering a scalable approach to tuberculosis (TB) screening in high-burden, low-resource settings. Previous studies have been limited by small datasets, under-representation of symptomatic non-TB patients, reliance on simple models, and recordings collected under idealised conditions. Methods We enrolled 512 participants at two hospitals in Zambia, grouped as bacteriologically confirmed TB (TB+), symptomatic patients with other respiratory diseases (OR), and healthy controls (HC). Usable cough recordings plus demographic and clinical data were obtained from 500 participants. Deep learning classifiers based on speech foundation models were trained on cough recordings. The best-performing model, trained on 3-second segments, was further evaluated with demographic and clinical features. Findings The best audio-only classifier achieved an AUROC of 85.2% for distinguishing TB+ from all others (TB+ Rest) and 80.1% for TB+ versus OR. Adding demographic and clinical features improved performance to 92.1% (TB+ Rest) and 84.2% (TB+ OR). At a threshold of 0.38, the multimodal model reached 90.3% sensitivity and 73.1% specificity for TB+ Rest, and 80.6% and 73.1% for TB+ OR. Interpretation Cough analysis using speech foundation models, especially when combined with demographic and clinical data, showed strong potential as a TB triage tool, meeting WHO target product profile benchmarks. The model was robust to confounding factors including background noise, recording time, and device variability, indicating detection of genuine disease-related acoustic patterns. Further validation across diverse regions and case definitions, including subclinical TB, is required before clinical use.
【15】Testing chatbots on the creation of encoders for audio conditioned image generation
标题:测试聊天机器人创建音频调节图像生成编码器
链接:http://arxiv.org/pdf/2509.09717v1
【16】VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions
标题:VStyle:语音风格改编的基准
链接:http://arxiv.org/pdf/2509.09716v1
【17】TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation
标题:TalkPlayData 2:用于多模式对话音乐推荐的大型合成数据管道
链接:http://arxiv.org/pdf/2509.09685v1
机器翻译由腾讯交互翻译提供,仅供参考
