今日论文合集:cs.SD语音15篇,eess.AS音频处理17篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Improving Audio Event Recognition with Consistency Regularization
标题:通过一致性规范化改进音频事件识别
链接:http://arxiv.org/pdf/2509.10391v1

备注:Under Review
摘要:一致性正则化(CR),它强制增强视图上的模型预测之间的一致性,最近在自动语音识别中发现了好处[1]。在本文中,我们提出了使用一致性正则化的音频事件识别,并证明其有效性AudioSet。通过对小型(2万美元)和大型(180万美元)监督训练集进行广泛的消融研究,我们表明CR比已经大量使用数据增强的监督基线带来了一致的改善,并且使用更强增强和多次增强的CR会为小型训练集带来额外的增益。此外,我们将CR的使用扩展到具有20K标记样本和1.8M未标记样本的半监督设置中,并获得了在小集合上训练的最佳模型的性能改进。摘要:Consistency regularization (CR), which enforces agreement between model predictions on augmented views, has found recent benefits in automatic speech recognition [1]. In this paper, we propose the use of consistency regularization for audio event recognition, and demonstrate its effectiveness on AudioSet. With extensive ablation studies for both small ($ sim$20k) and large ($ sim$1.8M) supervised training sets, we show that CR brings consistent improvement over supervised baselines which already heavily utilize data augmentation, and CR using stronger augmentation and multiple augmentations leads to additional gain for the small training set. Furthermore, we extend the use of CR into the semi-supervised setup with 20K labeled samples and 1.8M unlabeled samples, and obtain performance improvement over our best model trained on the small set.


【2】Data-independent Beamforming for End-to-end Multichannel Multi-speaker ASR
标题:用于端到端多通道多扬声器ASB的数据独立的射束形成
链接:http://arxiv.org/pdf/2509.10234v1

备注:Published in the IEEE 26th International Workshop on Multimedia Signal Processing (MMSP 2025)
摘要:由于环境噪声、混响和重叠扬声器的存在,多通道、多说话人场景中的自动语音识别(ASR)仍然具有挑战性。在本文中,我们提出了一种波束成形方法,在应用端到端多通道,多扬声器ASR系统之前,根据其球面极坐标处理特定的角度扇区。该方法具有数据无关性和无需训练的特点。我们证明,使用一组波束形成的信号提高ASR性能相比,使用相同数量的原始麦克风信号。此外,增加用于波束成形的信号的数量进一步增强了识别准确性,从而导致更有效地使用多通道信号,同时降低了ASR系统的总输入负载。我们进行实验的AMI会议语料库,其中所提出的方法减少了高达11%的单词错误率和提高扬声器计数精度高达27%,相对于一个多通道ASR基线系统,不利用波束成形。摘要:Automatic speech recognition (ASR) in multichannel, multi-speaker scenarios remains challenging due to ambient noise, reverberation and overlapping speakers. In this paper, we propose a beamforming approach that processes specific angular sectors based on their spherical polar coordinates before applying an end-to-end multichannel, multi-speaker ASR system. This method is data-independent and training-free. We demonstrate that using a group of beamformed signals improves ASR performance compared to using the same number of raw microphone signals. Moreover, increasing the number of signals used for beamforming further enhances recognition accuracy, leading to a more efficient use of multichannel signals while reducing the overall input load for the ASR system. We conduct experiments on the AMI meeting corpus, where the proposed method reduces word error rate by up to 11% and improves speaker counting accuracy by up to 27% relative compared to a multichannel ASR baseline system that does not exploit beamforming.


【3】Prototypical Contrastive Learning For Improved Few-Shot Audio Classification
标题:用于改进Few-Shot音频分类的原型对比学习
链接:http://arxiv.org/pdf/2509.10074v1

备注:Accepted and Presented at IEEE International Workshop on Machine Learning for Signal Processing, Aug. 31-- Sep. 3, 2025, Istanbul, Turkey , 6 pages, 2 figures, 1 table
摘要:Few-Shot学习已经成为训练具有有限标记数据的模型的强大范例,解决了大规模注释不切实际的情况下的挑战。虽然在图像领域已经进行了广泛的研究,但音频分类中的Few-Shot学习仍然相对不足。在这项工作中,我们研究了将监督对比度损失集成到用于音频分类的原型Few Shot训练中的效果。详细地,我们证明了角损耗进一步提高了性能相比,标准的对比度损失。我们的方法利用SpecAugment,其次是一个自我注意机制,以封装不同的信息增强输入版本到一个统一的嵌入。我们在MetaAudio上评估我们的方法,MetaAudio是一个基准测试,包括五个具有预定义分割的数据集,标准化预处理和一组全面的Few-Shot学习模型进行比较。所提出的方法在5路、5次拍摄设置中实现了最先进的性能。摘要:Few-shot learning has emerged as a powerful paradigm for training models with limited labeled data, addressing challenges in scenarios where large-scale annotation is impractical. While extensive research has been conducted in the image domain, few-shot learning in audio classification remains relatively underexplored. In this work, we investigate the effect of integrating supervised contrastive loss into prototypical few shot training for audio classification. In detail, we demonstrate that angular loss further improves the performance compared to the standard contrastive loss. Our method leverages SpecAugment followed by a self-attention mechanism to encapsulate diverse information of augmented input versions into one unified embedding. We evaluate our approach on MetaAudio, a benchmark including five datasets with predefined splits, standardized preprocessing, and a comprehensive set of few-shot learning models for comparison. The proposed approach achieves state-of-the-art performance in a 5-way, 5-shot setting.


【4】CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio
标题:CoDiCodec:统一音频的连续和离散压缩表示
链接:http://arxiv.org/pdf/2509.09836v1

备注:Accepted to ISMIR 2025
摘要:在压缩的潜在空间中有效地表示音频信号对于潜在生成建模是至关重要的。然而,现有的自动编码器通常强制在连续嵌入和离散标记之间进行选择。此外,在保持音频保真度的同时实现高压缩比仍然是一个挑战。我们介绍了一种新型的音频自动编码器,它克服了这些限制,既通过摘要嵌入有效地编码全局特征,又通过从相同的训练模型中以2.38 kbps的速率产生约11 Hz的压缩连续嵌入和离散令牌,为不同的下游生成任务提供了前所未有的灵活性。这是通过有限标量量化(FSQ)和一种新的FSQ-丢弃技术实现的,并且除了用于端到端训练的单一一致性损失之外,不需要额外的损失项。CoDiCodec支持自回归解码和新型并行解码策略,后者可实现卓越的音频质量和更快的解码速度。在重建音频质量方面,CoDiCodec在类似比特率下优于现有的连续和离散自动编码器。我们的工作实现了音频压缩的统一方法,弥合了连续和离散生成建模范式之间的差距。摘要:Efficiently representing audio signals in a compressed latent space is critical for latent generative modelling. However, existing autoencoders often force a choice between continuous embeddings and discrete tokens. Furthermore, achieving high compression ratios while maintaining audio fidelity remains a challenge. We introduce CoDiCodec, a novel audio autoencoder that overcomes these limitations by both efficiently encoding global features via summary embeddings, and by producing both compressed continuous embeddings at ~ 11 Hz and discrete tokens at a rate of 2.38 kbps from the same trained model, offering unprecedented flexibility for different downstream generative tasks. This is achieved through Finite Scalar Quantization (FSQ) and a novel FSQ-dropout technique, and does not require additional loss terms beyond the single consistency loss used for end-to-end training. CoDiCodec supports both autoregressive decoding and a novel parallel decoding strategy, with the latter achieving superior audio quality and faster decoding. CoDiCodec outperforms existing continuous and discrete autoencoders at similar bitrates in terms of reconstruction audio quality. Our work enables a unified approach to audio compression, bridging the gap between continuous and discrete generative modelling paradigms.


【5】SoilSound: Smartphone-based Soil Moisture Estimation
标题:SoilSound:基于智能手机的土壤湿度估计
链接:http://arxiv.org/pdf/2509.09823v1

备注:12 pages, 8 figures
摘要:土壤湿度监测对农业和环境管理至关重要,但现有方法要么需要侵入式探头干扰土壤,要么需要专门设备,限制了公众的使用。我们提出了SoilSound,一个无处不在的可访问的基于智能手机的声学传感系统,可以测量土壤水分,而不会干扰土壤。我们利用内置扬声器和麦克风执行垂直扫描机制,无需任何校准即可准确测量水分。与现有的使用透射特性的工作不同,我们提出了一种替代模型,土壤中的声反射的基础上的表面粗糙度效应,使水分传感不干扰土壤。该系统的工作原理是向土壤发送声学啁啾,并在垂直扫描期间记录反射,然后将其处理并馈送到卷积神经网络,用于设备上的土壤湿度估计,计算,内存或功率开销可以忽略不计。我们通过在实验室的盒子中使用精选土壤进行训练并在室外场地进行测试来评估该系统,结果表明SoilSound在10个不同位置的平均绝对误差(MAE)为2.39%。总体而言,评估表明SoilSound可以在多种土壤类型,环境和用户中准确跟踪土壤水分水平,范围从15.9%到34.0%;无需任何校准或干扰土壤,从而为家庭园丁,城市农民,公民科学家和资源有限环境中的农业社区提供广泛的水分监测。摘要:Soil moisture monitoring is essential for agriculture and environmental management, yet existing methods require either invasive probes disturbing the soil or specialized equipment, limiting access to the public. We present SoilSound, an ubiquitous accessible smartphone-based acoustic sensing system that can measure soil moisture without disturbing the soil. We leverage the built-in speaker and microphone to perform a vertical scan mechanism to accurately measure moisture without any calibration. Unlike existing work that use transmissive properties, we propose an alternate model for acoustic reflections in soil based on the surface roughness effect to enable moisture sensing without disturbing the soil. The system works by sending acoustic chirps towards the soil and recording the reflections during a vertical scan, which are then processed and fed to a convolutional neural network for on-device soil moisture estimation with negligible computational, memory, or power overhead. We evaluated the system by training with curated soils in boxes in the lab and testing in the outdoor fields and show that SoilSound achieves a mean absolute error (MAE) of 2.39% across 10 different locations. Overall, the evaluation shows that SoilSound can accurately track soil moisture levels ranging from 15.9% to 34.0% across multiple soil types, environments, and users; without requiring any calibration or disturbing the soil, enabling widespread moisture monitoring for home gardeners, urban farmers, citizen scientists, and agricultural communities in resource-limited settings.


【6】Combining Textual and Spectral Features for Robust Classification of Pilot Communications
标题:结合文本和频谱特征对导频通信进行稳健分类
链接:http://arxiv.org/pdf/2509.09752v1

摘要:准确估计飞机的起飞和降落等业务对于有效的机场管理至关重要,但仍然具有挑战性,特别是在缺乏专用监视基础设施的非塔楼设施中。本文提出了一种新的双管道机器学习框架,使用文本和频谱特征对导频无线电通信进行分类。从美国一个没有塔楼的机场收集的音频数据由经过认证的飞行员使用操作意图标签进行注释,并通过自动语音识别和Mel频谱图提取进行预处理。我们评估了广泛的传统分类器和深度学习模型,包括集成方法,LSTM和跨两个管道的CNN。据我们所知,这是第一个使用双管道ML框架对真实世界空中交通音频进行操作飞机意图分类的系统。我们的研究结果表明,光谱特征与深度架构相结合,始终产生卓越的分类性能,F1分数超过91%。数据增强进一步提高了对真实世界音频变化的鲁棒性。所提出的方法是可扩展的,具有成本效益的,无需额外的基础设施部署,在通用航空机场的空中交通监控提供了一个实用的解决方案。摘要:Accurate estimation of aircraft operations, such as takeoffs and landings, is critical for effective airport management, yet remains challenging, especially at non-towered facilities lacking dedicated surveillance infrastructure. This paper presents a novel dual pipeline machine learning framework that classifies pilot radio communications using both textual and spectral features. Audio data collected from a non-towered U.S. airport was annotated by certified pilots with operational intent labels and preprocessed through automatic speech recognition and Mel-spectrogram extraction. We evaluate a wide range of traditional classifiers and deep learning models, including ensemble methods, LSTM, and CNN across both pipelines. To our knowledge, this is the first system to classify operational aircraft intent using a dual-pipeline ML framework on real-world air traffic audio. Our results demonstrate that spectral features combined with deep architectures consistently yield superior classification performance, with F1-scores exceeding 91%. Data augmentation further improves robustness to real-world audio variability. The proposed approach is scalable, cost-effective, and deployable without additional infrastructure, offering a practical solution for air traffic monitoring at general aviation airports.


【7】DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
标题:DiTReducio:通过渐进式校准为基于DiT的TTC提供免训练加速
链接:http://arxiv.org/pdf/2509.09748v1

摘要:虽然扩散Transformers(DiT)具有先进的非自回归(NAR)语音合成,但是它们的高计算需求仍然是一个限制。现有的基于DiT的文本到语音(TTS)模型加速方法主要集中在通过蒸馏技术减少采样步骤,但它们仍然受到训练成本的限制。我们介绍了DiTReducio,这是一个无需训练的加速框架,它通过渐进式校准来压缩基于DiT的TTS模型中的计算。我们提出了两种压缩方法,时间跳跃和分支跳跃,以消除冗余计算过程中的推理。此外,基于DiT层内识别的两个特征注意模式,我们设计了一种模式引导策略来选择性地应用压缩方法。我们的方法允许通过可调的压缩阈值生成质量和计算效率之间的灵活调制。在F5-TTS和MegaTTS 3上进行的实验评估表明,DiTReducio实现了75.4%的FLOPs减少,并将实时因子(RTF)提高了37.1%,同时保持了生成质量。摘要:While Diffusion Transformers (DiT) have advanced non-autoregressive (NAR) speech synthesis, their high computational demands remain an limitation. Existing DiT-based text-to-speech (TTS) model acceleration approaches mainly focus on reducing sampling steps through distillation techniques, yet they remain constrained by training costs. We introduce DiTReducio, a training-free acceleration framework that compresses computations in DiT-based TTS models via progressive calibration. We propose two compression methods, Temporal Skipping and Branch Skipping, to eliminate redundant computations during inference. Moreover, based on two characteristic attention patterns identified within DiT layers, we devise a pattern-guided strategy to selectively apply the compression methods. Our method allows flexible modulation between generation quality and computational efficiency through adjustable compression thresholds. Experimental evaluations conducted on F5-TTS and MegaTTS 3 demonstrate that DiTReducio achieves a 75.4% reduction in FLOPs and improves the Real-Time Factor (RTF) by 37.1%, while preserving generation quality.


【8】AI-enabled tuberculosis screening in a high-burden setting using cough sound analysis and speech foundation models
标题:使用咳嗽声音分析和言语基础模型在高负担环境中进行人工智能支持的结核病筛查
链接:http://arxiv.org/pdf/2509.09746v1

备注:submitted to The Lancet Digital Health
摘要:背景 人工智能(AI)可以检测咳嗽声中与疾病相关的声学模式,为高负担、低资源环境中的结核病(TB)筛查提供了一种可扩展的方法。以前的研究受到小数据集、有症状的非结核病患者代表性不足、依赖简单模型以及在理想条件下收集的记录的限制。 方法 我们在赞比亚的两家医院招募了512名参与者,分为细菌学确诊的结核病(TB+),其他呼吸道疾病(OR)的症状患者和健康对照(HC)。从500名参与者中获得可用的咳嗽记录以及人口统计学和临床数据。基于语音基础模型的深度学习分类器在咳嗽记录上进行训练。在3秒片段上训练的表现最好的模型,进一步评估了人口统计学和临床特征。 结果 最好的仅音频分类器在区分TB+与所有其他(TB + Rest)方面实现了85.2%的AUROC,在区分TB+与OR方面实现了80.1%的AUROC。添加人口统计学和临床特征将性能提高到92.1%(TB + Rest)和84.2%(TB + OR)。在阈值为0.38时,多模式模型对TB + Rest的敏感性和特异性分别达到90.3%和73.1%,对TB + OR的敏感性和特异性分别达到80.6%和73.1%。 解释 使用语音基础模型进行咳嗽分析,特别是结合人口统计和临床数据时,显示出作为结核病分诊工具的强大潜力,符合世卫组织目标产品概况基准。该模型对混杂因素(包括背景噪声、记录时间和器械变异性)具有鲁棒性,表明检测到了真正的疾病相关声学模式。在临床使用之前,需要在不同地区和病例定义(包括亚临床TB)中进行进一步验证。摘要:Background Artificial intelligence (AI) can detect disease-related acoustic patterns in cough sounds, offering a scalable approach to tuberculosis (TB) screening in high-burden, low-resource settings. Previous studies have been limited by small datasets, under-representation of symptomatic non-TB patients, reliance on simple models, and recordings collected under idealised conditions. Methods We enrolled 512 participants at two hospitals in Zambia, grouped as bacteriologically confirmed TB (TB+), symptomatic patients with other respiratory diseases (OR), and healthy controls (HC). Usable cough recordings plus demographic and clinical data were obtained from 500 participants. Deep learning classifiers based on speech foundation models were trained on cough recordings. The best-performing model, trained on 3-second segments, was further evaluated with demographic and clinical features. Findings The best audio-only classifier achieved an AUROC of 85.2% for distinguishing TB+ from all others (TB+ Rest) and 80.1% for TB+ versus OR. Adding demographic and clinical features improved performance to 92.1% (TB+ Rest) and 84.2% (TB+ OR). At a threshold of 0.38, the multimodal model reached 90.3% sensitivity and 73.1% specificity for TB+ Rest, and 80.6% and 73.1% for TB+ OR. Interpretation Cough analysis using speech foundation models, especially when combined with demographic and clinical data, showed strong potential as a TB triage tool, meeting WHO target product profile benchmarks. The model was robust to confounding factors including background noise, recording time, and device variability, indicating detection of genuine disease-related acoustic patterns. Further validation across diverse regions and case definitions, including subclinical TB, is required before clinical use.


【9】Testing chatbots on the creation of encoders for audio conditioned image generation
标题:测试聊天机器人创建音频调节图像生成编码器
链接:http://arxiv.org/pdf/2509.09717v1

摘要:一方面,聊天机器人的最新进展导致使用这些模型进行编码任务越来越受欢迎。另一方面,现代生成图像模型主要依赖于文本编码器将语义概念转换为视觉表示,即使有明确的证据表明音频也可以用作输入。考虑到前面的内容,在这项工作中,我们将探索最先进的会话代理是否可以设计有效的音频编码器来取代Stable Diffusion 1.5中的CLIP文本编码器,从而直接从声音合成图像。我们促使五个公开可用的聊天机器人提出神经架构作为这些音频编码器,并提供一组解释良好的共享条件。每个有效的建议编码器都经过了超过200万个与上下文相关的音频-图像-文本观察的训练,并使用各种度量标准在验证和测试集上进行了评估,同时对其生成的图像进行了定性分析。尽管几乎所有聊天机器人都生成了有效的模型设计,但没有一个取得了令人满意的结果,这表明它们的音频嵌入未能与原始文本编码器的音频嵌入可靠地对齐。在这些提案中,Gemini音频编码器显示了最好的量化指标,而Grok音频编码器产生了更连贯的图像(特别是当与文本编码器配对时)。我们的研究结果揭示了聊天机器人之间共享的架构偏见,并强调了需要在这些模型的未来版本中弥合的剩余编码差距。我们还创建了一个公开演示,以便每个人都可以学习和尝试这些音频编码器。最后,我们提出了未来应该解决的研究问题,并鼓励其他研究人员执行像这样更集中和高度专业化的任务,因此相应的聊天机器人无法使用众所周知的解决方案,并且他们的创造力 推理得到了充分的测试。摘要:On one hand, recent advances in chatbots has led to a rising popularity in using these models for coding tasks. On the other hand, modern generative image models primarily rely on text encoders to translate semantic concepts into visual representations, even when there is clear evidence that audio can be employed as input as well. Given the previous, in this work, we explore whether state-of-the-art conversational agents can design effective audio encoders to replace the CLIP text encoder from Stable Diffusion 1.5, enabling image synthesis directly from sound. We prompted five publicly available chatbots to propose neural architectures to work as these audio encoders, with a set of well-explained shared conditions. Each valid suggested encoder was trained on over two million context related audio-image-text observations, and evaluated on held-out validation and test sets using various metrics, together with a qualitative analysis of their generated images. Although almost all chatbots generated valid model designs, none achieved satisfactory results, indicating that their audio embeddings failed to align reliably with those of the original text encoder. Among the proposals, the Gemini audio encoder showed the best quantitative metrics, while the Grok audio encoder produced more coherent images (particularly, when paired with the text encoder). Our findings reveal a shared architectural bias across chatbots and underscore the remaining coding gap that needs to be bridged in future versions of these models. We also created a public demo so everyone could study and try out these audio encoders. Finally, we propose research questions that should be tackled in the future, and encourage other researchers to perform more focused and highly specialized tasks like this one, so the respective chatbots cannot make use of well-known solutions and their creativity reasoning is fully tested.Furthermore, DreamAudio offers comparable performance in general text-to-audio tasks. We also provide a human-involved dataset containing audio events from real-world CTTA cases as the benchmark for customized generation tasks.


【10】VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions
标题:VStyle:语音风格改编的基准
链接:http://arxiv.org/pdf/2509.09716v1

摘要:口语模型(SLMs)已经成为语音理解和生成的统一范式,使自然的人机交互成为可能。然而,虽然大多数进展都集中在语义准确性和指令遵循,SLMs的能力,以适应他们的说话风格的基础上口头指示受到了有限的关注。我们介绍了语音风格适应(VSA),一个新的任务,检查是否SLM可以修改他们的说话风格,如音色,韵律,或人物以下的自然语言口语命令。为了研究这个任务,我们提出了VStyle,一个双语(中文和英文)基准,涵盖了四个类别的语音生成:声学属性,自然语言教学,角色扮演和内隐移情。我们还引入了大型音频语言模型作为法官(LALM作为法官)框架,该框架逐步评估文本忠实性,风格坚持性和自然性的输出,确保可再现和客观的评估。在商业系统和开源SLM上的实验表明,当前模型在可控风格适应方面面临明显的局限性,突出了这项任务的新颖性和挑战性。通过发布VStyle及其评估工具包,我们的目标是为社区提供一个推进以人为本的口语交互的基础。数据集和代码可在 href{https: junzhan2000.github.io VStyle.github.io }{project's homepage}上公开获取。摘要:Spoken language models (SLMs) have emerged as a unified paradigm for speech understanding and generation, enabling natural human machine interaction. However, while most progress has focused on semantic accuracy and instruction following, the ability of SLMs to adapt their speaking style based on spoken instructions has received limited attention. We introduce Voice Style Adaptation (VSA), a new task that examines whether SLMs can modify their speaking style, such as timbre, prosody, or persona following natural language spoken commands. To study this task, we present VStyle, a bilingual (Chinese & English) benchmark covering four categories of speech generation: acoustic attributes, natural language instruction, role play, and implicit empathy. We also introduce the Large Audio Language Model as a Judge (LALM as a Judge) framework, which progressively evaluates outputs along textual faithfulness, style adherence, and naturalness, ensuring reproducible and objective assessment. Experiments on commercial systems and open source SLMs demonstrate that current models face clear limitations in controllable style adaptation, highlighting both the novelty and challenge of this task. By releasing VStyle and its evaluation toolkit, we aim to provide the community with a foundation for advancing human centered spoken interaction. The dataset and code are publicly available at href{https: junzhan2000.github.io VStyle.github.io }{project's homepage}.


【11】TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation
标题:TalkPlayData 2:用于多模式对话音乐推荐的大型合成数据管道
链接:http://arxiv.org/pdf/2509.09685v1

摘要:我们提出了TalkPlayData 2,一个由代理数据管道生成的多模态会话音乐推荐的合成数据集。在TalkPlayData 2管道中,在各种角色下创建多个大语言模型(LLM)代理,具有专门的提示和对不同部分信息的访问,并且通过记录LLM和Recsys LLM之间的对话来获取聊天数据。为了覆盖各种会话场景,对于每个会话,将以微调的会话目标为条件来调整该会话LLM。最后,所有LLM都是多模态的,具有音频和图像,允许模拟多模态推荐和会话。在LLM-as-a-judge和主观评价实验中,TalkPlayData 2在与训练生成式音乐推荐模型相关的各个方面实现了所提出的目标。TalkPlayData 2及其生成代码在https: talkpl.ai talkplaydata2.html上开源。摘要:We present TalkPlayData 2, a synthetic dataset for multimodal conversational music recommendation generated by an agentic data pipeline. In TalkPlayData 2 pipeline, multiple large language model (LLM) agents are created under various roles with specialized prompts and access to different parts of information, and the chat data is acquired by logging the conversation between the Listener LLM and the Recsys LLM. To cover various conversation scenarios, for each conversation, the Listener LLM is conditioned on a finetuned conversation goal. Finally, all the LLMs are multimodal with audio and images, allowing a simulation of multimodal recommendation and conversation. In the LLM-as-a-judge and subjective evaluation experiments, TalkPlayData 2 achieved the proposed goal in various aspects related to training a generative recommendation model for music. TalkPlayData 2 and its generation code are open-sourced at https: talkpl.ai talkplaydata2.html.


【12】Error Analysis in a Modular Meeting Transcription System
标题:模块化会议转录系统中的错误分析
链接:http://arxiv.org/pdf/2509.10143v1

备注:Accepted at ITG Conference on Speech Communication 2025
摘要:会议记录是近年来发展迅速的一个研究领域。然而,挑战依然存在,限制了其业绩。在这项工作中,我们扩展了以前提出的框架,用于分析泄漏语音分离与适当的敏感性时间局部性。我们发现,有显着的泄漏到交叉通道的地区,只有主扬声器是活跃的。同时,结果表明,这不会影响最终的性能,因为这些泄漏的部分在很大程度上被语音活动检测(VAD)忽略。此外,不同的分割进行了比较,表明先进的日记化方法能够减少的差距,甲骨文分割的三分之一相比,一个简单的基于能量的VAD。我们还揭示了哪些因素导致了剩余的差异。这些结果代表了LibriCSS在仅对LibriSpeech数据进行训练识别模块的系统中的最新性能。摘要:Meeting transcription is a field of high relevance and remarkable progress in recent years. Still, challenges remain that limit its performance. In this work, we extend a previously proposed framework for analyzing leakage in speech separation with proper sensitivity to temporal locality. We show that there is significant leakage to the cross channel in areas where only the primary speaker is active. At the same time, the results demonstrate that this does not affect the final performance much as these leaked parts are largely ignored by the voice activity detection (VAD). Furthermore, different segmentations are compared showing that advanced diarization approaches are able to reduce the gap to oracle segmentation by a third compared to a simple energy-based VAD. We additionally reveal what factors contribute to the remaining difference. The results represent state-of-the-art performance on LibriCSS among systems that train the recognition module on LibriSpeech data only.


【13】Unified Learnable 2D Convolutional Feature Extraction for ASR
标题:用于ASB的统一可学习2D卷积特征提取
链接:http://arxiv.org/pdf/2509.10031v1

备注:Accepted at ITG Conference on Speech Communication 2025
摘要:神经前端代表了自动语音识别(ASR)系统特征提取的一种有前途的方法,因为它们能够为不同的任务学习专门定制的特征。然而,许多现有的技术仍然受到经典方法的严重影响。虽然这种归纳偏差可能会简化系统设计,但我们的工作旨在开发一个更通用的特征提取前端。此外,我们寻求统一前端架构,与应用源自不同来源的多层拓扑组合的现有方法形成鲜明对比。实验系统地展示了如何减少现有技术的影响,以实现通用的前端。由此产生的2D卷积前端是参数高效的,适用于计算资源有限的场景,而不像在未标记音频上预先训练的大型模型。结果表明,这种通用的统一方法不仅是可行的,而且与现有的监督可学习特征提取器的性能相匹配。摘要:Neural front-ends represent a promising approach to feature extraction for automatic speech recognition (ASR) systems as they enable to learn specifically tailored features for different tasks. Yet, many of the existing techniques remain heavily influenced by classical methods. While this inductive bias may ease the system design, our work aims to develop a more generic front-end for feature extraction. Furthermore, we seek to unify the front-end architecture contrasting with existing approaches that apply a composition of several layer topologies originating from different sources. The experiments systematically show how to reduce the influence of existing techniques to achieve a generic front-end. The resulting 2D convolutional front-end is parameter-efficient and suitable for a scenario with limited computational resources unlike large models pre-trained on unlabeled audio. The results demonstrate that this generic unified approach is not only feasible but also matches the performance of existing supervised learnable feature extractors.


【14】The MSP-Podcast Corpus
标题:MSP-播客素材库
链接:http://arxiv.org/pdf/2509.09791v1

备注:IEEE Transactions on Affective Computing submission
摘要:大的,高质量的情感语音数据库的可用性是在现实世界中推进语音情感识别(SER)的必要条件。然而,许多现有的数据库面临着规模,情感平衡和扬声器多样性的限制。本研究描述了MSP播客语料库,总结了我们十年的努力。该语料库由来自各种音频共享网站的400多个小时的不同音频样本组成,所有这些网站都具有允许分发语料库的通用许可证。我们用丰富的情感标签注释语料库,包括主要(单一主导情感)和次要(音频中感知的多种情感)情感类别,以及效价,唤醒和主导的情感属性。至少有五位评分员对这些情感标签进行了注释。语料库还具有针对大多数样本的说话人识别,以及针对整个语料库的句子的词汇内容的人工翻译。数据收集协议包括一个机器学习驱动的管道,用于选择情感多样的录音,确保在说话者和环境中平衡和多样化地表达情感。由此产生的数据库提供了一个全面的,高质量的资源,更适合于推进SER系统在实际,现实世界的情况。摘要:The availability of large, high-quality emotional speech databases is essential for advancing speech emotion recognition (SER) in real-world scenarios. However, many existing databases face limitations in size, emotional balance, and speaker diversity. This study describes the MSP-Podcast corpus, summarizing our ten-year effort. The corpus consists of over 400 hours of diverse audio samples from various audio-sharing websites, all of which have Common Licenses that permit the distribution of the corpus. We annotate the corpus with rich emotional labels, including primary (single dominant emotion) and secondary (multiple emotions perceived in the audio) emotional categories, as well as emotional attributes for valence, arousal, and dominance. At least five raters annotate these emotional labels. The corpus also has speaker identification for most samples, and human transcriptions of the lexical content of the sentences for the entire corpus. The data collection protocol includes a machine learning-driven pipeline for selecting emotionally diverse recordings, ensuring a balanced and varied representation of emotions across speakers and environments. The resulting database provides a comprehensive, high-quality resource, better suited for advancing SER systems in practical, real-world scenarios.


【15】Spectral Bottleneck in Deep Neural Networks: Noise is All You Need
标题:深度神经网络中的频谱瓶颈:噪音就是你所需要的一切
链接:http://arxiv.org/pdf/2509.09719v1

摘要:众所周知,深度神经网络表现出频谱学习偏差,其中低频分量在训练早期学习,而高频模式在后期逐渐出现。然而,当目标信号缺乏低频分量并且由宽带高频占主导地位时,训练会遭受“频谱检查”,并且模型无法重建整个信号,包括位于网络表示能力内的频率分量。我们研究这样的情况下,隐式神经表示(INR)与正弦表示网络(SIREN)的背景下,专注于拟合高频占主导地位的信号,容易受到频谱瓶颈的挑战。为了有效地适应任何目标信号,无论它的频率内容,我们提出了一个广义的目标感知的“权重扰动方案”(WINNER -权重初始化与噪声的神经表示)的网络初始化。该方案用高斯噪声扰动均匀初始化的权值,其中噪声尺度由目标信号的谱质心自适应确定。我们表明,噪声尺度可以提供控制网络激活的频谱和经验神经正切内核的本征基。这种方法不仅解决了频谱瓶颈,而且收敛速度更快,表示精度更高,在音频拟合方面优于最先进的方法,并在图像拟合和去噪任务中取得了显着的收益。除了信号重建之外,我们的方法还为计算机视觉和科学机器学习中的自适应权重初始化策略开辟了新的方向。摘要:Deep neural networks are known to exhibit a spectral learning bias, wherein low-frequency components are learned early in training, while high-frequency modes emerge more gradually in later epochs. However, when the target signal lacks low-frequency components and is dominated by broadband high frequencies, training suffers from a 'spectral bottleneck', and the model fails to reconstruct the entire signal, including the frequency components that lie within the network's representational capacity. We examine such a scenario in the context of implicit neural representations (INRs) with sinusoidal representation networks (SIRENs), focusing on the challenge of fitting high-frequency-dominant signals that are susceptible to spectral bottleneck. To effectively fit any target signal irrespective of it's frequency content, we propose a generalized target-aware 'weight perturbation scheme' (WINNER - weight initialization with noise for neural representations) for network initialization. The scheme perturbs uniformly initialized weights with Gaussian noise, where the noise scales are adaptively determined by the spectral centroid of the target signal. We show that the noise scales can provide control over the spectra of network activations and the eigenbasis of the empirical neural tangent kernel. This method not only addresses the spectral bottleneck but also yields faster convergence and with improved representation accuracy, outperforming state-of-the-art approaches in audio fitting and achieving notable gains in image fitting and denoising tasks. Beyond signal reconstruction, our approach opens new directions for adaptive weight initialization strategies in computer vision and scientific machine learning.


eess.AS音频处理


【1】Low-latency Assistive Audio Enhancement for Neurodivergent People
标题:针对神经分歧患者的低延迟辅助音频增强
链接:http://arxiv.org/pdf/2509.10202v1

摘要:神经发散型的人经常会经历声音耐受力下降,估计有50-70%的人会受到影响。这种提高的灵敏度可以引起从轻微不适到严重痛苦的反应,突出了辅助音频增强技术的迫切需要。在本文中,我们提出了几种辅助音频增强算法,旨在选择性地过滤令人痛苦的声音。为了解决这个问题,我们通过分析Reddit等平台上以神经分歧者为中心的社区,策划了一系列潜在的触发声音。使用此列表,从公开可用的来源(包括FSD 50 K和ESC 50)编译了触发声音样本的数据集。然后,这些样本用于训练和评估各种数字信号处理(DSP)和机器学习(ML)音频增强算法。在探索的方法中,动态范围压缩(DRC)被证明是最有效的,成功地衰减了触发声音,减少了神经分歧的听众的听觉困扰。摘要:Neurodivergent people frequently experience decreased sound tolerance, with estimates suggesting it affects 50-70% of this population. This heightened sensitivity can provoke reactions ranging from mild discomfort to severe distress, highlighting the critical need for assistive audio enhancement technologies In this paper, we propose several assistive audio enhancement algorithms designed to selectively filter distressing sounds. To address this, we curated a list of potential trigger sounds by analyzing neurodivergent-focused communities on platforms such as Reddit. Using this list, a dataset of trigger sound samples was compiled from publicly available sources, including FSD50K and ESC50. These samples were then used to train and evaluate various Digital Signal Processing (DSP) and Machine Learning (ML) audio enhancement algorithms. Among the approaches explored, Dynamic Range Compression (DRC) proved the most effective, successfully attenuating trigger sounds and reducing auditory distress for neurodivergent listeners.


【2】Error Analysis in a Modular Meeting Transcription System
标题:模块化会议转录系统中的错误分析
链接:http://arxiv.org/pdf/2509.10143v1

备注:Accepted at ITG Conference on Speech Communication 2025
摘要:会议记录是近年来发展迅速的一个研究领域。然而,挑战依然存在,限制了其业绩。在这项工作中,我们扩展了以前提出的框架,用于分析泄漏语音分离与适当的敏感性时间局部性。我们发现,有显着的泄漏到交叉通道的地区,只有主扬声器是活跃的。同时,结果表明,这不会影响最终的性能,因为这些泄漏的部分在很大程度上被语音活动检测(VAD)忽略。此外,不同的分割进行了比较,表明先进的日记化方法能够减少的差距,甲骨文分割的三分之一相比,一个简单的基于能量的VAD。我们还揭示了哪些因素导致了剩余的差异。这些结果代表了LibriCSS在仅对LibriSpeech数据进行训练识别模块的系统中的最新性能。摘要:Meeting transcription is a field of high relevance and remarkable progress in recent years. Still, challenges remain that limit its performance. In this work, we extend a previously proposed framework for analyzing leakage in speech separation with proper sensitivity to temporal locality. We show that there is significant leakage to the cross channel in areas where only the primary speaker is active. At the same time, the results demonstrate that this does not affect the final performance much as these leaked parts are largely ignored by the voice activity detection (VAD). Furthermore, different segmentations are compared showing that advanced diarization approaches are able to reduce the gap to oracle segmentation by a third compared to a simple energy-based VAD. We additionally reveal what factors contribute to the remaining difference. The results represent state-of-the-art performance on LibriCSS among systems that train the recognition module on LibriSpeech data only.


【3】Towards Data Drift Monitoring for Speech Deepfake Detection in the context of MLOps
标题:MLOps背景下语音Deepfake检测的数据漂移监控
链接:http://arxiv.org/pdf/2509.10086v1

备注:code to be pushed to this https URL
摘要:当在云上的应用程序或服务中交付时,未更新的静态语音deepfake检测器将容易受到新创建的语音deepfake攻击。从机器学习操作(MLOps)的角度来看,本文试图回答我们是否可以监控新的和看不见的语音deepfake数据,这些数据偏离了可见的参考数据集。我们进一步询问,如果检测到漂移,我们是否可以使用类似的漂移数据微调检测器,减少漂移,并提高检测性能。在玩具数据集和大规模MLAAD数据集上,我们证明了新的文本到语音(TTS)攻击引起的漂移可以使用新数据和参考数据的分布之间的距离进行监测。此外,我们证明了使用新TTS deepfakes生成的数据微调检测器可以减少漂移和检测错误率。摘要:When being delivered in applications or services on the cloud, static speech deepfake detectors that are not updated will become vulnerable to newly created speech deepfake attacks. From the perspective of machine learning operations (MLOps), this paper tries to answer whether we can monitor new and unseen speech deepfake data that drifts away from a seen reference data set. We further ask, if drift is detected, whether we can fine-tune the detector using similarly drifted data, reduce the drift, and improve the detection performance. On a toy dataset and the large-scale MLAAD dataset, we show that the drift caused by new text-to-speech (TTS) attacks can be monitored using distances between the distributions of the new data and reference data. Furthermore, we demonstrate that fine-tuning the detector using data generated by the new TTS deepfakes can reduce the drift and the detection error rates.


【4】Unified Learnable 2D Convolutional Feature Extraction for ASR
标题:用于ASB的统一可学习2D卷积特征提取
链接:http://arxiv.org/pdf/2509.10031v1

备注:Accepted at ITG Conference on Speech Communication 2025
摘要:神经前端代表了自动语音识别(ASR)系统特征提取的一种有前途的方法,因为它们能够为不同的任务学习专门定制的特征。然而,许多现有的技术仍然受到经典方法的严重影响。虽然这种归纳偏差可能会简化系统设计,但我们的工作旨在开发一个更通用的特征提取前端。此外,我们寻求统一前端架构,与应用源自不同来源的多层拓扑组合的现有方法形成鲜明对比。实验系统地展示了如何减少现有技术的影响,以实现通用的前端。由此产生的2D卷积前端是参数高效的,适用于计算资源有限的场景,而不像在未标记音频上预先训练的大型模型。结果表明,这种通用的统一方法不仅是可行的,而且与现有的监督可学习特征提取器的性能相匹配。摘要:Neural front-ends represent a promising approach to feature extraction for automatic speech recognition (ASR) systems as they enable to learn specifically tailored features for different tasks. Yet, many of the existing techniques remain heavily influenced by classical methods. While this inductive bias may ease the system design, our work aims to develop a more generic front-end for feature extraction. Furthermore, we seek to unify the front-end architecture contrasting with existing approaches that apply a composition of several layer topologies originating from different sources. The experiments systematically show how to reduce the influence of existing techniques to achieve a generic front-end. The resulting 2D convolutional front-end is parameter-efficient and suitable for a scenario with limited computational resources unlike large models pre-trained on unlabeled audio. The results demonstrate that this generic unified approach is not only feasible but also matches the performance of existing supervised learnable feature extractors.


【5】Whisper Has an Internal Word Aligner
标题:Whisper有一个内部单词对齐器
链接:http://arxiv.org/pdf/2509.09987v1

备注:ASRU 2025
摘要:从强大的自动语音识别器(特别是Whisper)中获得准确的单词级时间戳的兴趣越来越大。现有的方法要么需要额外的培训,要么根本没有竞争力。在以前的工作中的评估也是相对宽松的,通常使用超过200毫秒的公差。在这项工作中,我们发现注意头耳语捕捉准确的单词对齐,是明显不同于那些不。此外,我们发现,使用字符产生更精细,更准确的路线比使用单词。基于这些发现,我们提出了一种无监督的方法来提取词对齐过滤注意头,而教师强迫耳语字符。我们的方法不仅不需要训练,而且在20 ms和100 ms之间的更严格的公差下产生比先前工作更准确的单词对齐。摘要:There is an increasing interest in obtaining accurate word-level timestamps from strong automatic speech recognizers, in particular Whisper. Existing approaches either require additional training or are simply not competitive. The evaluation in prior work is also relatively loose, typically using a tolerance of more than 200 ms. In this work, we discover attention heads in Whisper that capture accurate word alignments and are distinctively different from those that do not. Moreover, we find that using characters produces finer and more accurate alignments than using wordpieces. Based on these findings, we propose an unsupervised approach to extracting word alignments by filtering attention heads while teacher forcing Whisper with characters. Our approach not only does not require training but also produces word alignments that are more accurate than prior work under a stricter tolerance between 20 ms and 100 ms.


【6】Effective Modeling of Critical Contextual Information for TDNN-based Speaker Verification
标题:基于TDNN的说话人验证关键上下文信息的有效建模
链接:http://arxiv.org/pdf/2509.09932v1

备注:5 pages, 3 figures
摘要:目前,时延神经网络(TDNN)已成为说话人确认的主流结构,其中ECAPA-TDNN是最先进的模型之一。目前致力于改进TDNN的工作主要是解决TDNN在建模全局信息方面的局限性,并弥合TDNN和二维卷积之间的差距。然而,ECAPA-TDNN提出的SE-Res 2Block中的分层卷积结构不能充分利用上下文信息,导致ECAPA-TDNN对有效上下文依赖建模的能力较弱。为此,提出了三种基于ECAPA-TDNN的改进结构,以充分有效地提取具有上下文依赖性的多尺度特征,然后将这些特征聚合。在VoxCeleb和CN-Celeb上的实验结果验证了三种结构的有效性。其中一种架构在VoxCeleb 1-O数据集上实现了比ECAPA-TDNN低近23%的等误率,证明了在可比参数计数下当前TDNN架构中可实现的竞争性能。摘要:Today, Time Delay Neural Network (TDNN) has become the mainstream architecture for speaker verification task, in which the ECAPA-TDNN is one of the state-of-the-art models. The current works that focus on improving TDNN primarily address the limitations of TDNN in modeling global information and bridge the gap between TDNN and 2-Dimensional convolutions. However, the hierarchical convolutional structure in the SE-Res2Block proposed by ECAPA-TDNN cannot make full use of the contextual information, resulting in the weak ability of ECAPA-TDNN to model effective context dependencies. To this end, three improved architectures based on ECAPA-TDNN are proposed to fully and effectively extract multi-scale features with context dependence and then aggregate these features. The experimental results on VoxCeleb and CN-Celeb verify the effectiveness of the three proposed architectures. One of these architectures achieves nearly a 23% lower Equal Error Rate compared to that of ECAPA-TDNN on VoxCeleb1-O dataset, demonstrating the competitive performance achievable among the current TDNN architectures under the comparable parameter count.


【7】Acoustic Scene Classification Using CNN-GRU Model Without Knowledge Distillation
标题:无需知识提炼的CNN-GRU模型声学场景分类
链接:http://arxiv.org/pdf/2509.09931v1

备注:3 pages, 2 figures, 2 tables
摘要:在本技术报告中,我们介绍了SNTL-NTU团队为低复杂度声学场景和事件(DCASE)2025挑战提交的任务1。这个提交偏离了从教师到学生模型的知识蒸馏的典型应用,旨在以有限的复杂性实现高性能。所提出的模型基于CNN-GRU模型,仅使用TAU Urban Acoustic Scene 2022 Mobile开发数据集进行训练,除了用于设备脉冲响应(ERF)增强的MicIRP之外,不使用任何外部数据集。该模型的内存使用量为114.2KB,需要10.9M乘法累加(MAC)操作。使用开发数据集,该模型实现了60.25%的准确率。摘要:In this technical report, we present the SNTL-NTU team's Task 1 submission for the Low-Complexity Acoustic Scenes and Events (DCASE) 2025 challenge. This submission departs from the typical application of knowledge distillation from a teacher to a student model, aiming to achieve high performance with limited complexity. The proposed model is based on a CNN-GRU model and is trained solely using the TAU Urban Acoustic Scene 2022 Mobile development dataset, without utilizing any external datasets, except for MicIRP, which is used for device impulse response (DIR) augmentation. The proposed model has a memory usage of 114.2KB and requires 10.9M muliply-and-accumulate (MAC) operations. Using the development dataset, the proposed model achieved an accuracy of 60.25%.


【8】The MSP-Podcast Corpus
标题:MSP-播客素材库
链接:http://arxiv.org/pdf/2509.09791v1

备注:IEEE Transactions on Affective Computing submission
摘要:大的,高质量的情感语音数据库的可用性是在现实世界中推进语音情感识别(SER)的必要条件。然而,许多现有的数据库面临着规模,情感平衡和扬声器多样性的限制。本研究描述了MSP播客语料库,总结了我们十年的努力。该语料库由来自各种音频共享网站的400多个小时的不同音频样本组成,所有这些网站都具有允许分发语料库的通用许可证。我们用丰富的情感标签注释语料库,包括主要(单一主导情感)和次要(音频中感知的多种情感)情感类别,以及效价,唤醒和主导的情感属性。至少有五位评分员对这些情感标签进行了注释。语料库还具有针对大多数样本的说话人识别,以及针对整个语料库的句子的词汇内容的人工翻译。数据收集协议包括一个机器学习驱动的管道,用于选择情感多样的录音,确保在说话者和环境中平衡和多样化地表达情感。由此产生的数据库提供了一个全面的,高质量的资源,更适合于推进SER系统在实际,现实世界的情况。摘要:The availability of large, high-quality emotional speech databases is essential for advancing speech emotion recognition (SER) in real-world scenarios. However, many existing databases face limitations in size, emotional balance, and speaker diversity. This study describes the MSP-Podcast corpus, summarizing our ten-year effort. The corpus consists of over 400 hours of diverse audio samples from various audio-sharing websites, all of which have Common Licenses that permit the distribution of the corpus. We annotate the corpus with rich emotional labels, including primary (single dominant emotion) and secondary (multiple emotions perceived in the audio) emotional categories, as well as emotional attributes for valence, arousal, and dominance. At least five raters annotate these emotional labels. The corpus also has speaker identification for most samples, and human transcriptions of the lexical content of the sentences for the entire corpus. The data collection protocol includes a machine learning-driven pipeline for selecting emotionally diverse recordings, ensuring a balanced and varied representation of emotions across speakers and environments. The resulting database provides a comprehensive, high-quality resource, better suited for advancing SER systems in practical, real-world scenarios.


【9】Spectral Bottleneck in Deep Neural Networks: Noise is All You Need
标题:深度神经网络中的频谱瓶颈:噪音就是你所需要的一切
链接:http://arxiv.org/pdf/2509.09719v1

摘要:众所周知,深度神经网络表现出频谱学习偏差,其中低频分量在训练早期学习,而高频模式在后期逐渐出现。然而,当目标信号缺乏低频分量并且由宽带高频占主导地位时,训练会遭受“频谱检查”,并且模型无法重建整个信号,包括位于网络表示能力内的频率分量。我们研究这样的情况下,隐式神经表示(INR)与正弦表示网络(SIREN)的背景下,专注于拟合高频占主导地位的信号,容易受到频谱瓶颈的挑战。为了有效地适应任何目标信号,无论它的频率内容,我们提出了一个广义的目标感知的“权重扰动方案”(WINNER -权重初始化与噪声的神经表示)的网络初始化。该方案用高斯噪声扰动均匀初始化的权值,其中噪声尺度由目标信号的谱质心自适应确定。我们表明,噪声尺度可以提供控制网络激活的频谱和经验神经正切内核的本征基。这种方法不仅解决了频谱瓶颈,而且收敛速度更快,表示精度更高,在音频拟合方面优于最先进的方法,并在图像拟合和去噪任务中取得了显着的收益。除了信号重建之外,我们的方法还为计算机视觉和科学机器学习中的自适应权重初始化策略开辟了新的方向。摘要:Deep neural networks are known to exhibit a spectral learning bias, wherein low-frequency components are learned early in training, while high-frequency modes emerge more gradually in later epochs. However, when the target signal lacks low-frequency components and is dominated by broadband high frequencies, training suffers from a 'spectral bottleneck', and the model fails to reconstruct the entire signal, including the frequency components that lie within the network's representational capacity. We examine such a scenario in the context of implicit neural representations (INRs) with sinusoidal representation networks (SIRENs), focusing on the challenge of fitting high-frequency-dominant signals that are susceptible to spectral bottleneck. To effectively fit any target signal irrespective of it's frequency content, we propose a generalized target-aware 'weight perturbation scheme' (WINNER - weight initialization with noise for neural representations) for network initialization. The scheme perturbs uniformly initialized weights with Gaussian noise, where the noise scales are adaptively determined by the spectral centroid of the target signal. We show that the noise scales can provide control over the spectra of network activations and the eigenbasis of the empirical neural tangent kernel. This method not only addresses the spectral bottleneck but also yields faster convergence and with improved representation accuracy, outperforming state-of-the-art approaches in audio fitting and achieving notable gains in image fitting and denoising tasks. Beyond signal reconstruction, our approach opens new directions for adaptive weight initialization strategies in computer vision and scientific machine learning.


【10】Prominence-aware automatic speech recognition for conversational speech
标题:会话语音的突出感知自动语音识别
链接:http://arxiv.org/pdf/2509.10116v1

摘要:本文研究了将显著性检测和语音识别相结合的会话型奥地利德语语音识别算法。首先,突出检测器的微调wav2vec2模型分类单词级的突出。然后使用该检测器在大型语料库中自动标注韵律突显。基于这些注释,我们训练了新的识别性ASR系统,该系统同时转录单词及其显著性水平。与我们的基线ASR系统相比,突出信息的集成没有改变性能,同时对于识别的单词序列正确的话语达到85.53%的突出检测准确率。本文表明,基于transformer的模型可以有效地编码韵律信息,并代表了一个新的贡献韵律增强ASR,具有潜在的应用语言学研究和韵律知情的对话系统。摘要:This paper investigates prominence-aware automatic speech recognition (ASR) by combining prominence detection and speech recognition for conversational Austrian German. First, prominence detectors were developed by fine-tuning wav2vec2 models to classify word-level prominence. The detector was then used to automatically annotate prosodic prominence in a large corpus. Based on those annotations, we trained novel prominence-aware ASR systems that simultaneously transcribe words and their prominence levels. The integration of prominence information did not change performance compared to our baseline ASR system, while reaching a prominence detection accuracy of 85.53% for utterances where the recognized word sequence was correct. This paper shows that transformer-based models can effectively encode prosodic information and represents a novel contribution to prosody-enhanced ASR, with potential applications for linguistic research and prosody-informed dialogue systems.


【11】CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio
标题:CoDiCodec:统一音频的连续和离散压缩表示
链接:http://arxiv.org/pdf/2509.09836v1

备注:Accepted to ISMIR 2025
摘要:在压缩的潜在空间中有效地表示音频信号对于潜在生成建模是至关重要的。然而,现有的自动编码器通常强制在连续嵌入和离散标记之间进行选择。此外,在保持音频保真度的同时实现高压缩比仍然是一个挑战。我们介绍了一种新型的音频自动编码器,它克服了这些限制,既通过摘要嵌入有效地编码全局特征,又通过从相同的训练模型中以2.38 kbps的速率产生约11 Hz的压缩连续嵌入和离散令牌,为不同的下游生成任务提供了前所未有的灵活性。这是通过有限标量量化(FSQ)和一种新的FSQ-丢弃技术实现的,并且除了用于端到端训练的单一一致性损失之外,不需要额外的损失项。CoDiCodec支持自回归解码和新型并行解码策略,后者可实现卓越的音频质量和更快的解码速度。在重建音频质量方面,CoDiCodec在类似比特率下优于现有的连续和离散自动编码器。我们的工作实现了音频压缩的统一方法,弥合了连续和离散生成建模范式之间的差距。摘要:Efficiently representing audio signals in a compressed latent space is critical for latent generative modelling. However, existing autoencoders often force a choice between continuous embeddings and discrete tokens. Furthermore, achieving high compression ratios while maintaining audio fidelity remains a challenge. We introduce CoDiCodec, a novel audio autoencoder that overcomes these limitations by both efficiently encoding global features via summary embeddings, and by producing both compressed continuous embeddings at ~ 11 Hz and discrete tokens at a rate of 2.38 kbps from the same trained model, offering unprecedented flexibility for different downstream generative tasks. This is achieved through Finite Scalar Quantization (FSQ) and a novel FSQ-dropout technique, and does not require additional loss terms beyond the single consistency loss used for end-to-end training. CoDiCodec supports both autoregressive decoding and a novel parallel decoding strategy, with the latter achieving superior audio quality and faster decoding. CoDiCodec outperforms existing continuous and discrete autoencoders at similar bitrates in terms of reconstruction audio quality. Our work enables a unified approach to audio compression, bridging the gap between continuous and discrete generative modelling paradigms.


【12】Combining Textual and Spectral Features for Robust Classification of Pilot Communications
标题:结合文本和频谱特征对导频通信进行稳健分类
链接:http://arxiv.org/pdf/2509.09752v1

摘要:准确估计飞机的起飞和降落等业务对于有效的机场管理至关重要,但仍然具有挑战性,特别是在缺乏专用监视基础设施的非塔楼设施中。本文提出了一种新的双管道机器学习框架,使用文本和频谱特征对导频无线电通信进行分类。从美国一个没有塔楼的机场收集的音频数据由经过认证的飞行员使用操作意图标签进行注释,并通过自动语音识别和Mel频谱图提取进行预处理。我们评估了广泛的传统分类器和深度学习模型,包括集成方法,LSTM和跨两个管道的CNN。据我们所知,这是第一个使用双管道ML框架对真实世界空中交通音频进行操作飞机意图分类的系统。我们的研究结果表明,光谱特征与深度架构相结合,始终产生卓越的分类性能,F1分数超过91%。数据增强进一步提高了对真实世界音频变化的鲁棒性。所提出的方法是可扩展的,具有成本效益的,无需额外的基础设施部署,在通用航空机场的空中交通监控提供了一个实用的解决方案。摘要:Accurate estimation of aircraft operations, such as takeoffs and landings, is critical for effective airport management, yet remains challenging, especially at non-towered facilities lacking dedicated surveillance infrastructure. This paper presents a novel dual pipeline machine learning framework that classifies pilot radio communications using both textual and spectral features. Audio data collected from a non-towered U.S. airport was annotated by certified pilots with operational intent labels and preprocessed through automatic speech recognition and Mel-spectrogram extraction. We evaluate a wide range of traditional classifiers and deep learning models, including ensemble methods, LSTM, and CNN across both pipelines. To our knowledge, this is the first system to classify operational aircraft intent using a dual-pipeline ML framework on real-world air traffic audio. Our results demonstrate that spectral features combined with deep architectures consistently yield superior classification performance, with F1-scores exceeding 91%. Data augmentation further improves robustness to real-world audio variability. The proposed approach is scalable, cost-effective, and deployable without additional infrastructure, offering a practical solution for air traffic monitoring at general aviation airports.


【13】DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration
标题:DiTReducio:通过渐进式校准为基于DiT的TTC提供免训练加速
链接:http://arxiv.org/pdf/2509.09748v1

摘要:虽然扩散Transformers(DiT)具有先进的非自回归(NAR)语音合成,但是它们的高计算需求仍然是一个限制。现有的基于DiT的文本到语音(TTS)模型加速方法主要集中在通过蒸馏技术减少采样步骤,但它们仍然受到训练成本的限制。我们介绍了DiTReducio,这是一个无需训练的加速框架,它通过渐进式校准来压缩基于DiT的TTS模型中的计算。我们提出了两种压缩方法,时间跳跃和分支跳跃,以消除冗余计算过程中的推理。此外,基于DiT层内识别的两个特征注意模式,我们设计了一种模式引导策略来选择性地应用压缩方法。我们的方法允许通过可调的压缩阈值生成质量和计算效率之间的灵活调制。在F5-TTS和MegaTTS 3上进行的实验评估表明,DiTReducio实现了75.4%的FLOPs减少,并将实时因子(RTF)提高了37.1%,同时保持了生成质量。摘要:While Diffusion Transformers (DiT) have advanced non-autoregressive (NAR) speech synthesis, their high computational demands remain an limitation. Existing DiT-based text-to-speech (TTS) model acceleration approaches mainly focus on reducing sampling steps through distillation techniques, yet they remain constrained by training costs. We introduce DiTReducio, a training-free acceleration framework that compresses computations in DiT-based TTS models via progressive calibration. We propose two compression methods, Temporal Skipping and Branch Skipping, to eliminate redundant computations during inference. Moreover, based on two characteristic attention patterns identified within DiT layers, we devise a pattern-guided strategy to selectively apply the compression methods. Our method allows flexible modulation between generation quality and computational efficiency through adjustable compression thresholds. Experimental evaluations conducted on F5-TTS and MegaTTS 3 demonstrate that DiTReducio achieves a 75.4% reduction in FLOPs and improves the Real-Time Factor (RTF) by 37.1%, while preserving generation quality.


【14】AI-enabled tuberculosis screening in a high-burden setting using cough sound analysis and speech foundation models
标题:使用咳嗽声音分析和言语基础模型在高负担环境中进行人工智能支持的结核病筛查
链接:http://arxiv.org/pdf/2509.09746v1

备注:submitted to The Lancet Digital Health
摘要:背景 人工智能(AI)可以检测咳嗽声中与疾病相关的声学模式,为高负担、低资源环境中的结核病(TB)筛查提供了一种可扩展的方法。以前的研究受到小数据集、有症状的非结核病患者代表性不足、依赖简单模型以及在理想条件下收集的记录的限制。 方法 我们在赞比亚的两家医院招募了512名参与者,分为细菌学确诊的结核病(TB+),其他呼吸道疾病(OR)的症状患者和健康对照(HC)。从500名参与者中获得可用的咳嗽记录以及人口统计学和临床数据。基于语音基础模型的深度学习分类器在咳嗽记录上进行训练。在3秒片段上训练的表现最好的模型,进一步评估了人口统计学和临床特征。 结果 最好的仅音频分类器在区分TB+与所有其他(TB + Rest)方面实现了85.2%的AUROC,在区分TB+与OR方面实现了80.1%的AUROC。添加人口统计学和临床特征将性能提高到92.1%(TB + Rest)和84.2%(TB + OR)。在阈值为0.38时,多模式模型对TB + Rest的敏感性和特异性分别达到90.3%和73.1%,对TB + OR的敏感性和特异性分别达到80.6%和73.1%。 解释 使用语音基础模型进行咳嗽分析,特别是结合人口统计和临床数据时,显示出作为结核病分诊工具的强大潜力,符合世卫组织目标产品概况基准。该模型对混杂因素(包括背景噪声、记录时间和器械变异性)具有鲁棒性,表明检测到了真正的疾病相关声学模式。在临床使用之前,需要在不同地区和病例定义(包括亚临床TB)中进行进一步验证。摘要:Background Artificial intelligence (AI) can detect disease-related acoustic patterns in cough sounds, offering a scalable approach to tuberculosis (TB) screening in high-burden, low-resource settings. Previous studies have been limited by small datasets, under-representation of symptomatic non-TB patients, reliance on simple models, and recordings collected under idealised conditions. Methods We enrolled 512 participants at two hospitals in Zambia, grouped as bacteriologically confirmed TB (TB+), symptomatic patients with other respiratory diseases (OR), and healthy controls (HC). Usable cough recordings plus demographic and clinical data were obtained from 500 participants. Deep learning classifiers based on speech foundation models were trained on cough recordings. The best-performing model, trained on 3-second segments, was further evaluated with demographic and clinical features. Findings The best audio-only classifier achieved an AUROC of 85.2% for distinguishing TB+ from all others (TB+ Rest) and 80.1% for TB+ versus OR. Adding demographic and clinical features improved performance to 92.1% (TB+ Rest) and 84.2% (TB+ OR). At a threshold of 0.38, the multimodal model reached 90.3% sensitivity and 73.1% specificity for TB+ Rest, and 80.6% and 73.1% for TB+ OR. Interpretation Cough analysis using speech foundation models, especially when combined with demographic and clinical data, showed strong potential as a TB triage tool, meeting WHO target product profile benchmarks. The model was robust to confounding factors including background noise, recording time, and device variability, indicating detection of genuine disease-related acoustic patterns. Further validation across diverse regions and case definitions, including subclinical TB, is required before clinical use.


【15】Testing chatbots on the creation of encoders for audio conditioned image generation
标题:测试聊天机器人创建音频调节图像生成编码器
链接:http://arxiv.org/pdf/2509.09717v1

摘要:一方面,聊天机器人的最新进展导致使用这些模型进行编码任务越来越受欢迎。另一方面,现代生成图像模型主要依赖于文本编码器将语义概念转换为视觉表示,即使有明确的证据表明音频也可以用作输入。考虑到前面的内容,在这项工作中,我们将探索最先进的会话代理是否可以设计有效的音频编码器来取代Stable Diffusion 1.5中的CLIP文本编码器,从而直接从声音合成图像。我们促使五个公开可用的聊天机器人提出神经架构作为这些音频编码器,并提供一组解释良好的共享条件。每个有效的建议编码器都经过了超过200万个与上下文相关的音频-图像-文本观察的训练,并使用各种度量标准在验证和测试集上进行了评估,同时对其生成的图像进行了定性分析。尽管几乎所有聊天机器人都生成了有效的模型设计,但没有一个取得了令人满意的结果,这表明它们的音频嵌入未能与原始文本编码器的音频嵌入可靠地对齐。在这些提案中,Gemini音频编码器显示了最好的量化指标,而Grok音频编码器产生了更连贯的图像(特别是当与文本编码器配对时)。我们的研究结果揭示了聊天机器人之间共享的架构偏见,并强调了需要在这些模型的未来版本中弥合的剩余编码差距。我们还创建了一个公开演示,以便每个人都可以学习和尝试这些音频编码器。最后,我们提出了未来应该解决的研究问题,并鼓励其他研究人员执行像这样更集中和高度专业化的任务,因此相应的聊天机器人无法使用众所周知的解决方案,并且他们的创造力 推理得到了充分的测试。摘要:On one hand, recent advances in chatbots has led to a rising popularity in using these models for coding tasks. On the other hand, modern generative image models primarily rely on text encoders to translate semantic concepts into visual representations, even when there is clear evidence that audio can be employed as input as well. Given the previous, in this work, we explore whether state-of-the-art conversational agents can design effective audio encoders to replace the CLIP text encoder from Stable Diffusion 1.5, enabling image synthesis directly from sound. We prompted five publicly available chatbots to propose neural architectures to work as these audio encoders, with a set of well-explained shared conditions. Each valid suggested encoder was trained on over two million context related audio-image-text observations, and evaluated on held-out validation and test sets using various metrics, together with a qualitative analysis of their generated images. Although almost all chatbots generated valid model designs, none achieved satisfactory results, indicating that their audio embeddings failed to align reliably with those of the original text encoder. Among the proposals, the Gemini audio encoder showed the best quantitative metrics, while the Grok audio encoder produced more coherent images (particularly, when paired with the text encoder). Our findings reveal a shared architectural bias across chatbots and underscore the remaining coding gap that needs to be bridged in future versions of these models. We also created a public demo so everyone could study and try out these audio encoders. Finally, we propose research questions that should be tackled in the future, and encourage other researchers to perform more focused and highly specialized tasks like this one, so the respective chatbots cannot make use of well-known solutions and their creativity reasoning is fully tested.Furthermore, DreamAudio offers comparable performance in general text-to-audio tasks. We also provide a human-involved dataset containing audio events from real-world CTTA cases as the benchmark for customized generation tasks.


【16】VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions
标题:VStyle:语音风格改编的基准
链接:http://arxiv.org/pdf/2509.09716v1

摘要:口语模型(SLMs)已经成为语音理解和生成的统一范式,使自然的人机交互成为可能。然而,虽然大多数进展都集中在语义准确性和指令遵循,SLMs的能力,以适应他们的说话风格的基础上口头指示受到了有限的关注。我们介绍了语音风格适应(VSA),一个新的任务,检查是否SLM可以修改他们的说话风格,如音色,韵律,或人物以下的自然语言口语命令。为了研究这个任务,我们提出了VStyle,一个双语(中文和英文)基准,涵盖了四个类别的语音生成:声学属性,自然语言教学,角色扮演和内隐移情。我们还引入了大型音频语言模型作为法官(LALM作为法官)框架,该框架逐步评估文本忠实性,风格坚持性和自然性的输出,确保可再现和客观的评估。在商业系统和开源SLM上的实验表明,当前模型在可控风格适应方面面临明显的局限性,突出了这项任务的新颖性和挑战性。通过发布VStyle及其评估工具包,我们的目标是为社区提供一个推进以人为本的口语交互的基础。数据集和代码可在 href{https: junzhan2000.github.io VStyle.github.io }{project's homepage}上公开获取。摘要:Spoken language models (SLMs) have emerged as a unified paradigm for speech understanding and generation, enabling natural human machine interaction. However, while most progress has focused on semantic accuracy and instruction following, the ability of SLMs to adapt their speaking style based on spoken instructions has received limited attention. We introduce Voice Style Adaptation (VSA), a new task that examines whether SLMs can modify their speaking style, such as timbre, prosody, or persona following natural language spoken commands. To study this task, we present VStyle, a bilingual (Chinese & English) benchmark covering four categories of speech generation: acoustic attributes, natural language instruction, role play, and implicit empathy. We also introduce the Large Audio Language Model as a Judge (LALM as a Judge) framework, which progressively evaluates outputs along textual faithfulness, style adherence, and naturalness, ensuring reproducible and objective assessment. Experiments on commercial systems and open source SLMs demonstrate that current models face clear limitations in controllable style adaptation, highlighting both the novelty and challenge of this task. By releasing VStyle and its evaluation toolkit, we aim to provide the community with a foundation for advancing human centered spoken interaction. The dataset and code are publicly available at href{https: junzhan2000.github.io VStyle.github.io }{project's homepage}.


【17】TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation
标题:TalkPlayData 2:用于多模式对话音乐推荐的大型合成数据管道
链接:http://arxiv.org/pdf/2509.09685v1

摘要:我们提出了TalkPlayData 2,一个由代理数据管道生成的多模态会话音乐推荐的合成数据集。在TalkPlayData 2管道中,在各种角色下创建多个大语言模型(LLM)代理,具有专门的提示和对不同部分信息的访问,并且通过记录LLM和Recsys LLM之间的对话来获取聊天数据。为了覆盖各种会话场景,对于每个会话,将以微调的会话目标为条件来调整该会话LLM。最后,所有LLM都是多模态的,具有音频和图像,允许模拟多模态推荐和会话。在LLM-as-a-judge和主观评价实验中,TalkPlayData 2在与训练生成式音乐推荐模型相关的各个方面实现了所提出的目标。TalkPlayData 2及其生成代码在https: talkpl.ai talkplaydata2.html上开源。摘要:We present TalkPlayData 2, a synthetic dataset for multimodal conversational music recommendation generated by an agentic data pipeline. In TalkPlayData 2 pipeline, multiple large language model (LLM) agents are created under various roles with specialized prompts and access to different parts of information, and the chat data is acquired by logging the conversation between the Listener LLM and the Recsys LLM. To cover various conversation scenarios, for each conversation, the Listener LLM is conditioned on a finetuned conversation goal. Finally, all the LLMs are multimodal with audio and images, allowing a simulation of multimodal recommendation and conversation. In the LLM-as-a-judge and subjective evaluation experiments, TalkPlayData 2 achieved the proposed goal in various aspects related to training a generative recommendation model for music. TalkPlayData 2 and its generation code are open-sourced at https: talkpl.ai talkplaydata2.html.


机器翻译由腾讯交互翻译提供,仅供参考