今日论文合集:cs.SD语音13篇,eess.AS音频处理14篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Input Conditioned Layer Dropping in Speech Foundation Models
标题:语音基础模型中的输入条件层下降
链接:http://arxiv.org/pdf/2507.07954v1
作者:nan, Daniele Falavigna, Alessio Brutti
备注:Accepted at IEEE MLSP 2025
摘要:为计算资源随时间变化的边缘和物联网环境管理基础语音模型,需要具有自适应缩减策略的动态架构。一种新兴的方法是层丢弃($ mathcal{LD}$),它在推理过程中跳过骨干网络的一部分层,以减少计算负载。这允许将静态模型转换为动态模型。然而,现有的方法表现出的局限性,无论是在选择层的模式,或通过显着修改的神经架构。为此,我们提出了输入驱动的$ mathcal{LD}$,它采用网络的输入功能和一个轻量级的层选择网络来确定处理层的最佳组合。使用两种不同的预训练基础模型,在4个语音和音频公共基准上进行了广泛的实验,证明了我们方法的有效性,完全优于随机丢弃,并产生了与标准相当(或更好)的结果,以提前退出。
摘要:Curating foundation speech models for edge and IoT settings, where computational resources vary over time, requires dynamic architectures featuring adaptable reduction strategies. One emerging approach is layer dropping ($ mathcal{LD}$) which skips fraction of the layers of a backbone network during inference to reduce the computational load. This allows transforming static models into dynamic ones. However, existing approaches exhibit limitations either in the mode of selecting layers or by significantly modifying the neural architecture. To this end, we propose input-driven $ mathcal{LD}$ that employs the network's input features and a lightweight layer selecting network to determine the optimum combination of processing layers. Extensive experimentation on 4 speech and audio public benchmarks, using two different pre-trained foundation models, demonstrates the effectiveness of our approach, thoroughly outperforming random dropping and producing on-par (or better) results to early exit.


【2】LISTEN: Lightweight Industrial Sound-representable Transformer for Edge Notification

标题:LISTEN:用于边缘通知的轻型工业声音表示Transformer器
链接:http://arxiv.org/pdf/2507.07879v1
作者:Han, Yun Seok Kang, Yuseop Sim, Martin Byung-Guk Jun, Hyung Wook Park
备注:7 pages,6 figures, Accepted to Speech Synthesis Workshop 2025 (SSW13)
摘要:基于深度学习的机器监听正在扩大工业声学分析的范围,用于异常检测和预测性维护等应用,从而提高制造效率和可靠性。然而,它对每个新任务的大型特定任务注释数据集的依赖限制了车间的广泛实施。虽然新兴的良好基础模型旨在减轻数据依赖性,但它们过于庞大且计算成本高昂,需要云基础设施或高端硬件,而这对于现场实时部署来说是不切实际的。我们通过LISTEN(用于边缘通知的轻量级工业声音表示Transformer)解决了这一差距,LISTEN是一个微型工业声音基础模型。使用知识蒸馏,LISTEN在低成本边缘设备上实时运行。在基准下游任务上,它的性能几乎与其更大的父模型相同,即使在使用最少的数据集和训练资源进行微调时也是如此。除了模型本身之外,我们还通过将LISTEN集成到具有工业物联网(IIoT)传感器和系统的边缘设备上的完整机器监控框架中来展示其现实效用,并在现场制造车间验证其性能和泛化能力。
摘要:Deep learning-based machine listening is broadening the scope of industrial acoustic analysis for applications like anomaly detection and predictive maintenance, thereby improving manufacturing efficiency and reliability. Nevertheless, its reliance on large, task-specific annotated datasets for every new task limits widespread implementation on shop floors. While emerging sound foundation models aim to alleviate data dependency, they are too large and computationally expensive, requiring cloud infrastructure or high-end hardware that is impractical for on-site, real-time deployment. We address this gap with LISTEN (Lightweight Industrial Sound-representable Transformer for Edge Notification), a kilobyte-sized industrial sound foundation model. Using knowledge distillation, LISTEN runs in real-time on low-cost edge devices. On benchmark downstream tasks, it performs nearly identically to its much larger parent model, even when fine-tuned with minimal datasets and training resource. Beyond the model itself, we demonstrate its real-world utility by integrating LISTEN into a complete machine monitoring framework on an edge device with an Industrial Internet of Things (IIoT) sensor and system, validating its performance and generalization capabilities on a live manufacturing shop floor.


【3】Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models

标题:Edge-ASB:迈向自动语音识别模型的低位量化
链接:http://arxiv.org/pdf/2507.07877v1
作者:Yicheng Lin, Shaojie Zhuo, Chenzheng Su, Ramchalam Kinattinkara Ramakrishnan, Zhaocong Yuan, Xiaopeng Zhang
备注:Accepted to Interspeech 2025
摘要:自动语音识别(ASR)的最新进展已经在各种音频应用(如现场转录和语音命令处理)中表现出显着的准确性和鲁棒性。然而,在资源受限的边缘设备(例如,由于内存、计算和功耗的严格限制,物联网设备、可穿戴设备)仍然面临着巨大的挑战。量化,特别是训练后量化(PTQ),提供了一种有效的方法来减少模型大小和推理成本,而无需重新训练。尽管其重要性,各种先进的量化方法和位宽配置的ASR模型的性能影响仍然不清楚。在这项工作中,我们提出了一个全面的基准八个国家的最先进的(SOTA)PTQ方法应用于两个领先的边缘ASR模型的家庭,耳语和月光。我们系统地评估模型性能(即,准确性、内存I O和位操作),分析量化和各种配置对权重和激活的影响。我们的框架建立在LLM压缩工具包的扩展之上,集成了边缘ASR模型、各种先进的量化算法、统一的校准和评估数据管道以及详细的分析工具。我们的研究结果的特点效率和准确性之间的权衡,表明即使3位量化可以成功的高容量模型时,使用先进的PTQ技术。这些发现为在低功耗、始终在线的边缘设备上优化ASR模型提供了有价值的见解。
摘要:Recent advances in Automatic Speech Recognition (ASR) have demonstrated remarkable accuracy and robustness in diverse audio applications, such as live transcription and voice command processing. However, deploying these models on resource constrained edge devices (e.g., IoT device, wearables) still presents substantial challenges due to strict limits on memory, compute and power. Quantization, particularly Post-Training Quantization (PTQ), offers an effective way to reduce model size and inference cost without retraining. Despite its importance, the performance implications of various advanced quantization methods and bit-width configurations on ASR models remain unclear. In this work, we present a comprehensive benchmark of eight state-of-the-art (SOTA) PTQ methods applied to two leading edge-ASR model families, Whisper and Moonshine. We systematically evaluate model performances (i.e., accuracy, memory I O and bit operations) across seven diverse datasets from the open ASR leaderboard, analyzing the impact of quantization and various configurations on both weights and activations. Built on an extension of the LLM compression toolkit, our framework integrates edge-ASR models, diverse advanced quantization algorithms, a unified calibration and evaluation data pipeline, and detailed analysis tools. Our results characterize the trade-offs between efficiency and accuracy, demonstrating that even 3-bit quantization can succeed on high capacity models when using advanced PTQ techniques. These findings provide valuable insights for optimizing ASR models on low-power, always-on edge devices.


【4】Re-Bottleneck: Latent Re-Structuring for Neural Audio Autoencoders

标题:重新瓶颈:神经音频自动编码器的潜在重组
链接:http://arxiv.org/pdf/2507.07867v1
作者:Bralios, Jonah Casebeer, Paris Smaragdis

备注:Accepted at IEEE MLSP 2025

摘要:神经音频编解码器和自动编码器已经成为音频压缩、传输、特征提取和潜在空间生成的通用模型。然而,一个关键的限制是,大多数都是为了最大限度地提高重建保真度而训练的,通常忽略了在各种下游应用中实现最佳性能所必需的特定潜在结构。我们提出了一个简单的事后框架,通过修改预先训练的自动编码器的瓶颈来解决这个问题。我们的方法引入了一个“重新瓶颈”,一个内部瓶颈专门通过潜在的空间损失来训练,以灌输用户定义的结构。我们在三个实验中证明了该框架的有效性。首先,我们在不牺牲重建质量的情况下对潜在通道进行排序。其次,我们将潜在的语义嵌入,分析下游扩散建模的影响。第三,我们引入等方差,确保对输入波形的滤波操作直接对应于潜在空间中的特定变换。最终,我们的Re-Bottleneck框架提供了一种灵活有效的方法来定制神经音频模型的表示,使它们能够以最少的额外训练无缝地满足不同应用的各种需求。
摘要:Neural audio codecs and autoencoders have emerged as versatile models for audio compression, transmission, feature-extraction, and latent-space generation. However, a key limitation is that most are trained to maximize reconstruction fidelity, often neglecting the specific latent structure necessary for optimal performance in diverse downstream applications. We propose a simple, post-hoc framework to address this by modifying the bottleneck of a pre-trained autoencoder. Our method introduces a "Re-Bottleneck", an inner bottleneck trained exclusively through latent space losses to instill user-defined structure. We demonstrate the framework's effectiveness in three experiments. First, we enforce an ordering on latent channels without sacrificing reconstruction quality. Second, we align latents with semantic embeddings, analyzing the impact on downstream diffusion modeling. Third, we introduce equivariance, ensuring that a filtering operation on the input waveform directly corresponds to a specific transformation in the latent space. Ultimately, our Re-Bottleneck framework offers a flexible and efficient way to tailor representations of neural audio models, enabling them to seamlessly meet the varied demands of different applications with minimal additional training.


【5】End-to-end Acoustic-linguistic Emotion and Intent Recognition Enhanced by Semi-supervised Learning

标题:半监督学习增强端到端声学语言情感和意图识别
链接:http://arxiv.org/pdf/2507.07806v1
作者:Rathi Adarshi Rammohan, Kevin Scheck, Sheng Li, Tanja Schultz

备注:Accepted by EMBC 2025

摘要:从语音中识别情感和意图是至关重要的,并且在人机交互中得到了广泛的研究。社交媒体平台、聊天机器人和其他技术的快速发展导致了来自用户的大量语音数据流。然而,手动注释这些数据是昂贵的,使得训练机器学习模型以用于识别目的具有挑战性。为此,我们建议应用半监督学习将大规模的未标记数据与相对较小的标记数据集结合起来。我们训练端到端的声学和语言模型,每个模型都采用多任务学习进行情感和意图识别。比较了固定匹配学习和完全匹配学习两种半监督学习方法。实验结果表明,半监督学习方法提高了语音情感和意图识别的声学和文本数据的模型性能。最佳模型的后期融合分别通过12.3%和10.4%的联合识别平衡度量优于声学和文本基线。
摘要:Emotion and intent recognition from speech is essential and has been widely investigated in human-computer interaction. The rapid development of social media platforms, chatbots, and other technologies has led to a large volume of speech data streaming from users. Nevertheless, annotating such data manually is expensive, making it challenging to train machine learning models for recognition purposes. To this end, we propose applying semi-supervised learning to incorporate a large scale of unlabelled data alongside a relatively smaller set of labelled data. We train end-to-end acoustic and linguistic models, each employing multi-task learning for emotion and intent recognition. Two semi-supervised learning approaches, including fix-match learning and full-match learning, are compared. The experimental results demonstrate that the semi-supervised learning approaches improve model performance in speech emotion and intent recognition from both acoustic and text data. The late fusion of the best models outperforms the acoustic and text baselines by joint recognition balance metrics of 12.3% and 10.4%, respectively.


【6】SecureSpeech: Prompt-based Speaker and Content Protection

标题:SecureSpeech:基于预算的演讲者和内容保护
链接:http://arxiv.org/pdf/2507.07799v1
作者:oh Hui Hui, Xiaoxiao Miao, Xin Wang
备注:Accepted by IEEE International Joint Conference on Biometrics (IJCB) 2025
摘要:鉴于越来越多的隐私问题,从身份盗窃和重新识别扬声器通过在语音领域的内容,本文提出了一种基于匿名的语音生成管道,确保双重匿名的扬声器身份和口语内容。这是通过以下方式解决的:1)生成由描述符控制的不受源说话者约束的说话者身份,以及2)使用名称实体识别模型和大型语言模型替换原始文本中的敏感内容。该流水线利用匿名化的说话者身份和文本来经由文本到语音合成模型生成高保真、隐私友好的语音。实验结果表明,在保持相当水平的内容保留和音频质量的同时,实现了显着的隐私保护。本文还研究了不同的扬声器描述的效用和隐私生成的语音,以确定潜在的偏见的影响。
摘要:Given the increasing privacy concerns from identity theft and the re-identification of speakers through content in the speech field, this paper proposes a prompt-based speech generation pipeline that ensures dual anonymization of both speaker identity and spoken content. This is addressed through 1) generating a speaker identity unlinkable to the source speaker, controlled by descriptors, and 2) replacing sensitive content within the original text using a name entity recognition model and a large language model. The pipeline utilizes the anonymized speaker identity and text to generate high-fidelity, privacy-friendly speech via a text-to-speech synthesis model. Experimental results demonstrate an achievement of significant privacy protection while maintaining a decent level of content retention and audio quality. This paper also investigates the impact of varying speaker descriptions on the utility and privacy of generated speech to determine potential biases.


【7】Assessing the Alignment of Audio Representations with Timbre Similarity Ratings

标题:评估音频表示与音色相似度评级的一致性
链接:http://arxiv.org/pdf/2507.07764v1
作者:an, Stefan Lattner, Charalampos Saitis

备注:Accepted to ISMIR 2025

摘要:心理声学所谓的“音色空间”通过多维缩放将乐器声音的感知相似性评级映射到低维嵌入上,但存在可扩展性问题并且无法泛化。最近的音频(音乐和语音)质量评估以及图像相似性的结果表明,深度学习能够产生与人类感知良好一致的嵌入,同时在很大程度上不受这些限制。虽然现有的人类评级音色相似性数据不足以训练深度神经网络(对334个音频样本进行2,614个成对评级),但它可以作为音频模型的测试数据。在本文中,我们引入指标来评估对齐的不同的音频表示与人类的判断音色相似性,通过比较的绝对值和排名的嵌入距离人类相似性评级。我们的评估涉及三个基于信号处理的表示,从预训练模型中提取的十二个表示,以及从一个新的声音匹配模型中提取的三个表示。其中,CLAP模型和声音匹配模型提取的受图像风格迁移启发的风格嵌入效果明显优于其他模型,显示了它们在建模音色相似性方面的潜力。
摘要:Psychoacoustical so-called "timbre spaces" map perceptual similarity ratings of instrument sounds onto low-dimensional embeddings via multidimensional scaling, but suffer from scalability issues and are incapable of generalization. Recent results from audio (music and speech) quality assessment as well as image similarity have shown that deep learning is able to produce embeddings that align well with human perception while being largely free from these constraints. Although the existing human-rated timbre similarity data is not large enough to train deep neural networks (2,614 pairwise ratings on 334 audio samples), it can serve as test-only data for audio models. In this paper, we introduce metrics to assess the alignment of diverse audio representations with human judgments of timbre similarity by comparing both the absolute values and the rankings of embedding distances to human similarity ratings. Our evaluation involves three signal-processing-based representations, twelve representations extracted from pre-trained models, and three representations extracted from a novel sound matching model. Among them, the style embeddings inspired by image style transfer, extracted from the CLAP model and the sound matching model, remarkably outperform the others, showing their potential in modeling timbre similarity.


【8】DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction

标题:DMF2Mel:一种用于脑电驱动Mel谱图重建的动态多尺度融合网络
链接:http://arxiv.org/pdf/2507.07526v1
作者:an, Sheng Zhang, Jingjing Zhang, Enrui Liu, Xinhui Li, Minggang Zhao, Zhao Lv
备注:Accepted by ACM MM 2025
摘要:从大脑信号中解码语音是一个具有挑战性的研究问题。尽管现有技术在单词或字母水平上重建听觉刺激的mel谱图方面取得了进展,但在精确重建分钟水平的连续想象语音方面仍存在核心挑战:传统模型难以平衡长序列解码中的时间依赖性建模效率和信息保留。为了解决这个问题,本文提出了动态多尺度融合网络(DMF 2 Mel),它由四个核心组件组成:动态对比特征聚合模块(DC-FAM),分层注意引导多尺度网络(HAMS-Net),SplineMap注意机制和双向状态空间模块(convMamba)。具体地说,DC-FAM通过局部卷积和全局注意机制将语音相关的“前景特征”与噪声“背景特征”分离,有效地抑制了干扰并增强了瞬态信号的表示。HAMS-Net基于U-Net框架,实现了高层语义和底层细节的跨尺度融合。SplineMap注意力机制集成了自适应门控Kolmogorov-Arnold网络(AGKAN),将全局上下文建模与基于样条的局部拟合相结合。convMamba以线性复杂度捕获长范围的时间依赖性,并增强了非线性动态建模能力。SparrKULee数据集上的结果显示,DMF 2 Mel在已知受试者的mel频谱图重建中实现了0.074的Pearson相关系数(相对于基线提高了48%),并且对于未知受试者实现了0.048的Pearson相关系数(相对于基线提高了35%)。https: github.com fchest DMF2Mel
摘要:Decoding speech from brain signals is a challenging research problem. Although existing technologies have made progress in reconstructing the mel spectrograms of auditory stimuli at the word or letter level, there remain core challenges in the precise reconstruction of minute-level continuous imagined speech: traditional models struggle to balance the efficiency of temporal dependency modeling and information retention in long-sequence decoding. To address this issue, this paper proposes the Dynamic Multiscale Fusion Network (DMF2Mel), which consists of four core components: the Dynamic Contrastive Feature Aggregation Module (DC-FAM), the Hierarchical Attention-Guided Multi-Scale Network (HAMS-Net), the SplineMap attention mechanism, and the bidirectional state space module (convMamba). Specifically, the DC-FAM separates speech-related "foreground features" from noisy "background features" through local convolution and global attention mechanisms, effectively suppressing interference and enhancing the representation of transient signals. HAMS-Net, based on the U-Net framework,achieves cross-scale fusion of high-level semantics and low-level details. The SplineMap attention mechanism integrates the Adaptive Gated Kolmogorov-Arnold Network (AGKAN) to combine global context modeling with spline-based local fitting. The convMamba captures long-range temporal dependencies with linear complexity and enhances nonlinear dynamic modeling capabilities. Results on the SparrKULee dataset show that DMF2Mel achieves a Pearson correlation coefficient of 0.074 in mel spectrogram reconstruction for known subjects (a 48% improvement over the baseline) and 0.048 for unknown subjects (a 35% improvement over the baseline).Code is available at: https: github.com fchest DMF2Mel.


【9】IML-Spikeformer: Input-aware Multi-Level Spiking Transformer for Speech Processing

标题:IML-Spikeformer:用于语音处理的输入感知多层Spikeformer
链接:http://arxiv.org/pdf/2507.07396v1
作者:ng, Shimin Zhang, Yuhong Chou, Jibin Wu, Haizhou Li
备注:Under review of TNNLS
摘要:尖峰神经网络(SNN),灵感来自生物神经机制,代表了一个有前途的神经形态计算范式,提供了节能的替代传统的人工神经网络(ANN)。尽管SNN架构的有效性已得到证明,但它在大规模语音处理任务上仍难以实现有竞争力的性能。两个关键挑战阻碍了进展:(1)多时间步尖峰放电导致的训练期间的高计算开销,以及(2)缺乏针对语音处理任务的大规模SNN架构。为了克服这些问题,我们引入了输入感知的多级Spikeformer,即IML-Spikeformer,一种专门为大规模语音处理设计的尖峰Transformer架构。我们设计的核心是输入感知多级尖峰(IMLS)机制,它使用自适应输入感知阈值方案在单个时间步内模拟多时间步尖峰发射。IML-Spikeformer进一步将重新参数化的尖峰自我注意(RepSSA)模块与分层衰减掩模(HDM)集成在一起,形成HD-RepSSA模块。该模块提高了注意力地图的精度,并能够对语音信号中的多尺度时间依赖性进行建模。实验表明,IML-Spikeformer在AiShell-1和Librispeech-960上分别实现了6.0%和3.4%的字错误率,与传统的ANN Transformers相当,同时分别减少了4.64times $和4.32times $的理论推理能耗。IML-Spikeformer标志着可扩展SNN架构在任务性能和能源效率方面的大规模语音处理的进步。
摘要:Spiking Neural Networks (SNNs), inspired by biological neural mechanisms, represent a promising neuromorphic computing paradigm that offers energy-efficient alternatives to traditional Artificial Neural Networks (ANNs). Despite proven effectiveness, SNN architectures have struggled to achieve competitive performance on large-scale speech processing task. Two key challenges hinder progress: (1) the high computational overhead during training caused by multi-timestep spike firing, and (2) the absence of large-scale SNN architectures tailored to speech processing tasks. To overcome the issues, we introduce Input-aware Multi-Level Spikeformer, i.e. IML-Spikeformer, a spiking Transformer architecture specifically designed for large-scale speech processing. Central to our design is the Input-aware Multi-Level Spike (IMLS) mechanism, which simulate multi-timestep spike firing within a single timestep using an adaptive, input-aware thresholding scheme. IML-Spikeformer further integrates a Reparameterized Spiking Self-Attention (RepSSA) module with a Hierarchical Decay Mask (HDM), forming the HD-RepSSA module. This module enhances the precision of attention maps and enables modeling of multi-scale temporal dependencies in speech signals. Experiments demonstrate that IML-Spikeformer achieves word error rates of 6.0 % on AiShell-1 and 3.4 % on Librispeech-960, comparable to conventional ANN transformers while reducing theoretical inference energy consumption by 4.64$ times$ and 4.32$ times$ respectively. IML-Spikeformer marks an advance of scalable SNN architectures for large-scale speech processing in both task performance and energy efficiency.


【10】VP-SelDoA: Visual-prompted Selective DoA Estimation of Target Sound via Semantic-Spatial Matching

标题:VP-SelDoA:通过语义空间匹配的视觉提示选择性DoA目标声音估计
链接:http://arxiv.org/pdf/2507.07384v1
作者:Xinyuan Qian, Hongxu Zhu, Jiadong Wang, Kainan Chen, Haizhou Li
备注:Under Review
摘要:视听声源定位(AV-SSL)通过利用听觉和视觉信号的互补强度来识别声源的位置。然而,现有的AV-SSL方法遇到了三个主要挑战:1)无法在多源场景中选择性地隔离目标声源,2)语义视觉特征和空间声学特征之间的不匹配,以及3)过度依赖成对的视听数据。为了克服这些限制,我们引入了跨实例视听定位(CI-AVL),这是一种新的任务,它利用来自同一声音事件类别的不同实例的图像来定位目标声源,从而减少对配对数据的依赖,同时增强泛化能力。我们提出的VP-SelDoA通过语义级模态融合来解决这一具有挑战性的任务,并采用频率-时间ConMamba架构来生成用于声音隔离的目标选择性掩模。我们进一步开发了一个语义空间匹配机制,通过集成的交叉和自我注意机制,对齐异构的语义和空间特征。为了便于CI-AVL研究,我们构建了一个名为VGG-SSL的大规模数据集,包括296个声音事件类别的13,981个空间音频片段。大量的实验表明,我们提出的方法优于国家的最先进的视听定位方法,实现了平均绝对误差(MAE)的12.04和准确度(ACC)的78.23%。
摘要:Audio-visual sound source localization (AV-SSL) identifies the position of a sound source by exploiting the complementary strengths of auditory and visual signals. However, existing AV-SSL methods encounter three major challenges: 1) inability to selectively isolate the target sound source in multi-source scenarios, 2) misalignment between semantic visual features and spatial acoustic features, and 3) overreliance on paired audio-visual data. To overcome these limitations, we introduce Cross-Instance Audio-Visual Localization (CI-AVL), a novel task that leverages images from different instances of the same sound event category to localize target sound sources, thereby reducing dependence on paired data while enhancing generalization capabilities. Our proposed VP-SelDoA tackles this challenging task through a semantic-level modality fusion and employs a Frequency-Temporal ConMamba architecture to generate target-selective masks for sound isolation. We further develop a Semantic-Spatial Matching mechanism that aligns the heterogeneous semantic and spatial features via integrated cross- and self-attention mechanisms. To facilitate the CI-AVL research, we construct a large-scale dataset named VGG-SSL, comprising 13,981 spatial audio clips across 296 sound event categories. Extensive experiments show that our proposed method outperforms state-of-the-art audio-visual localization methods, achieving a mean absolute error (MAE) of 12.04 and an accuracy (ACC) of 78.23%.


【11】SonicMotion: Dynamic Spatial Audio Soundscapes with Latent Diffusion Models

标题:SonicMotion:具有潜在扩散模型的动态空间音频声景
链接:http://arxiv.org/pdf/2507.07318v1
作者:Templin, Yanda Zhu, Hao Wang
摘要:空间音频是沉浸式娱乐(如VR AR)不可或缺的一部分,在电影和音乐中也越来越受欢迎。空间音频的最常见格式被描述为一阶高保真度立体声(FOA)。我们寻求扩展FOA生成AI模型的最新进展,以生成具有动态声源的3D场景。我们提出的端到端模型SonicMotion有两种变体,它们在用户输入和声源定位的精确度方面有所不同。除了我们的模型,我们还提出了一个新的模拟空间音频字幕对数据集。我们的模型的评估表明,他们能够匹配的语义对齐和音频质量的最先进的模型,同时捕捉所需的空间属性。
摘要:Spatial audio is an integral part of immersive entertainment, such as VR AR, and has seen increasing popularity in cinema and music as well. The most common format of spatial audio is described as first-order Ambisonics (FOA). We seek to extend recent advancements in FOA generative AI models to enable the generation of 3D scenes with dynamic sound sources. Our proposed end-to-end model, SonicMotion, comes in two variations which vary in their user input and level of precision in sound source localization. In addition to our model, we also present a new dataset of simulated spatial audio-caption pairs. Evaluation of our models demonstrate that they are capable of matching the semantic alignment and audio quality of state of the art models while capturing the desired spatial attributes.


【12】Audio-Visual Speech Separation via Bottleneck Iterative Network

标题:瓶颈迭代网络实现视听语音分离
链接:http://arxiv.org/pdf/2507.07270v1
作者:ang, Shiv Shankar, Trang Nguyen, Andrea Fanelli, Madalina Fiterau
备注:Accepted to the 42nd International Conference on Machine Learning Workshop on Machine Learning for Audio
摘要:从非听觉线索的信息的整合可以显着提高语音分离模型的性能。通常,此类模型使用深度特定于模态的网络来获取单模态特征,并且存在成本太高或重量太轻但容量不足的风险。在这项工作中,我们提出了一种迭代表示细化方法,称为瓶颈迭代网络(BIN),一种通过轻量级融合块重复进行的技术,同时通过融合令牌检查融合表示。这有助于提高模型的容量,同时避免模型大小的大幅增加以及模型性能和训练成本之间的平衡。我们在具有挑战性的嘈杂视听语音分离任务上测试BIN,并表明我们的方法在NTCD-TIMIT和LRS 3 +WHAM上的SI-SDRi方面始终优于最先进的基准模型!数据集,同时在几乎所有设置中实现训练和GPU推理时间减少50%以上。
摘要:Integration of information from non-auditory cues can significantly improve the performance of speech-separation models. Often such models use deep modality-specific networks to obtain unimodal features, and risk being too costly or lightweight but lacking capacity. In this work, we present an iterative representation refinement approach called Bottleneck Iterative Network (BIN), a technique that repeatedly progresses through a lightweight fusion block, while bottlenecking fusion representations by fusion tokens. This helps improve the capacity of the model, while avoiding major increase in model size and balancing between the model performance and training cost. We test BIN on challenging noisy audio-visual speech separation tasks, and show that our approach consistently outperforms state-of-the-art benchmark models with respect to SI-SDRi on NTCD-TIMIT and LRS3+WHAM! datasets, while simultaneously achieving a reduction of more than 50% in training and GPU inference time across nearly all settings.


【13】Generic Speech Enhancement with Self-Supervised Representation Space Loss

标题:具有自我监督表示空间损失的通用语音增强
链接:http://arxiv.org/pdf/2507.07631v1
作者:ato, Tsubasa Ochiai, Marc Delcroix, Takafumi Moriya, Takanori Ashihara, Ryo Masumura
备注:22 pages, 3 figures. Accepted for Frontiers in signal processing
摘要:单通道语音增强被用于各种任务中,以减轻干扰信号的影响。常规地,为了确保语音增强最佳地执行,需要针对每个任务来调谐语音增强。因此,将语音增强模型推广到未知的下游任务一直具有挑战性。本研究的目的是建立一个通用的语音增强前端,可以提高后端的性能,以解决多个下游任务。为此,我们提出了一种新的训练标准,该标准可以最大限度地减少自监督学习模型的特征表示域中增强的和真实的干净信号之间的距离。由于自监督学习特征表示有效地表达了对解决各种下游任务有用的高级语音信息,因此该提案有望使语音增强模型保留这些信息。实验验证表明,该建议提高了多个语音任务的性能,同时保持增强信号的感知质量。
摘要:Single-channel speech enhancement is utilized in various tasks to mitigate the effect of interfering signals. Conventionally, to ensure the speech enhancement performs optimally, the speech enhancement has needed to be tuned for each task. Thus, generalizing speech enhancement models to unknown downstream tasks has been challenging. This study aims to construct a generic speech enhancement front-end that can improve the performance of back-ends to solve multiple downstream tasks. To this end, we propose a novel training criterion that minimizes the distance between the enhanced and the ground truth clean signal in the feature representation domain of self-supervised learning models. Since self-supervised learning feature representations effectively express high-level speech information useful for solving various downstream tasks, the proposal is expected to make speech enhancement models preserve such information. Experimental validation demonstrates that the proposal improves the performance of multiple speech tasks while maintaining the perceptual quality of the enhanced signal.


eess.AS音频处理


【1】Generic Speech Enhancement with Self-Supervised Representation Space Loss

标题:具有自我监督表示空间损失的通用语音增强
链接:http://arxiv.org/pdf/2507.07631v1
作者:ato, Tsubasa Ochiai, Marc Delcroix, Takafumi Moriya, Takanori Ashihara, Ryo Masumura
备注:22 pages, 3 figures. Accepted for Frontiers in signal processing
摘要:单通道语音增强被用于各种任务中,以减轻干扰信号的影响。常规地,为了确保语音增强最佳地执行,需要针对每个任务来调谐语音增强。因此,将语音增强模型推广到未知的下游任务一直具有挑战性。本研究的目的是建立一个通用的语音增强前端,可以提高后端的性能,以解决多个下游任务。为此,我们提出了一种新的训练标准,该标准可以最大限度地减少自监督学习模型的特征表示域中增强的和真实的干净信号之间的距离。由于自监督学习特征表示有效地表达了对解决各种下游任务有用的高级语音信息,因此该提案有望使语音增强模型保留这些信息。实验验证表明,该建议提高了多个语音任务的性能,同时保持增强信号的感知质量。
摘要:Single-channel speech enhancement is utilized in various tasks to mitigate the effect of interfering signals. Conventionally, to ensure the speech enhancement performs optimally, the speech enhancement has needed to be tuned for each task. Thus, generalizing speech enhancement models to unknown downstream tasks has been challenging. This study aims to construct a generic speech enhancement front-end that can improve the performance of back-ends to solve multiple downstream tasks. To this end, we propose a novel training criterion that minimizes the distance between the enhanced and the ground truth clean signal in the feature representation domain of self-supervised learning models. Since self-supervised learning feature representations effectively express high-level speech information useful for solving various downstream tasks, the proposal is expected to make speech enhancement models preserve such information. Experimental validation demonstrates that the proposal improves the performance of multiple speech tasks while maintaining the perceptual quality of the enhanced signal.


【2】Input Conditioned Layer Dropping in Speech Foundation Models
标题:语音基础模型中的输入条件层下降
链接:http://arxiv.org/pdf/2507.07954v1
作者:nan, Daniele Falavigna, Alessio Brutti
备注:Accepted at IEEE MLSP 2025
摘要:为计算资源随时间变化的边缘和物联网环境管理基础语音模型,需要具有自适应缩减策略的动态架构。一种新兴的方法是层丢弃($ mathcal{LD}$),它在推理过程中跳过骨干网络的一部分层,以减少计算负载。这允许将静态模型转换为动态模型。然而,现有的方法表现出的局限性,无论是在选择层的模式,或通过显着修改的神经架构。为此,我们提出了输入驱动的$ mathcal{LD}$,它采用网络的输入功能和一个轻量级的层选择网络来确定处理层的最佳组合。使用两种不同的预训练基础模型,在4个语音和音频公共基准上进行了广泛的实验,证明了我们方法的有效性,完全优于随机丢弃,并产生了与标准相当(或更好)的结果,以提前退出。
摘要:Curating foundation speech models for edge and IoT settings, where computational resources vary over time, requires dynamic architectures featuring adaptable reduction strategies. One emerging approach is layer dropping ($ mathcal{LD}$) which skips fraction of the layers of a backbone network during inference to reduce the computational load. This allows transforming static models into dynamic ones. However, existing approaches exhibit limitations either in the mode of selecting layers or by significantly modifying the neural architecture. To this end, we propose input-driven $ mathcal{LD}$ that employs the network's input features and a lightweight layer selecting network to determine the optimum combination of processing layers. Extensive experimentation on 4 speech and audio public benchmarks, using two different pre-trained foundation models, demonstrates the effectiveness of our approach, thoroughly outperforming random dropping and producing on-par (or better) results to early exit.


【3】LISTEN: Lightweight Industrial Sound-representable Transformer for Edge Notification

标题:LISTEN:用于边缘通知的轻型工业声音表示Transformer器
链接:http://arxiv.org/pdf/2507.07879v1
作者:Han, Yun Seok Kang, Yuseop Sim, Martin Byung-Guk Jun, Hyung Wook Park
备注:7 pages,6 figures, Accepted to Speech Synthesis Workshop 2025 (SSW13)
摘要:基于深度学习的机器监听正在扩大工业声学分析的范围,用于异常检测和预测性维护等应用,从而提高制造效率和可靠性。然而,它对每个新任务的大型特定任务注释数据集的依赖限制了车间的广泛实施。虽然新兴的良好基础模型旨在减轻数据依赖性,但它们过于庞大且计算成本高昂,需要云基础设施或高端硬件,而这对于现场实时部署来说是不切实际的。我们通过LISTEN(用于边缘通知的轻量级工业声音表示Transformer)解决了这一差距,LISTEN是一个微型工业声音基础模型。使用知识蒸馏,LISTEN在低成本边缘设备上实时运行。在基准下游任务上,它的性能几乎与其更大的父模型相同,即使在使用最少的数据集和训练资源进行微调时也是如此。除了模型本身之外,我们还通过将LISTEN集成到具有工业物联网(IIoT)传感器和系统的边缘设备上的完整机器监控框架中来展示其现实效用,并在现场制造车间验证其性能和泛化能力。
摘要:Deep learning-based machine listening is broadening the scope of industrial acoustic analysis for applications like anomaly detection and predictive maintenance, thereby improving manufacturing efficiency and reliability. Nevertheless, its reliance on large, task-specific annotated datasets for every new task limits widespread implementation on shop floors. While emerging sound foundation models aim to alleviate data dependency, they are too large and computationally expensive, requiring cloud infrastructure or high-end hardware that is impractical for on-site, real-time deployment. We address this gap with LISTEN (Lightweight Industrial Sound-representable Transformer for Edge Notification), a kilobyte-sized industrial sound foundation model. Using knowledge distillation, LISTEN runs in real-time on low-cost edge devices. On benchmark downstream tasks, it performs nearly identically to its much larger parent model, even when fine-tuned with minimal datasets and training resource. Beyond the model itself, we demonstrate its real-world utility by integrating LISTEN into a complete machine monitoring framework on an edge device with an Industrial Internet of Things (IIoT) sensor and system, validating its performance and generalization capabilities on a live manufacturing shop floor.


【4】Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models

标题:Edge-ASB:迈向自动语音识别模型的低位量化
链接:http://arxiv.org/pdf/2507.07877v1
作者:Yicheng Lin, Shaojie Zhuo, Chenzheng Su, Ramchalam Kinattinkara Ramakrishnan, Zhaocong Yuan, Xiaopeng Zhang
备注:Accepted to Interspeech 2025
摘要:自动语音识别(ASR)的最新进展已经在各种音频应用(如现场转录和语音命令处理)中表现出显着的准确性和鲁棒性。然而,在资源受限的边缘设备(例如,由于内存、计算和功耗的严格限制,物联网设备、可穿戴设备)仍然面临着巨大的挑战。量化,特别是训练后量化(PTQ),提供了一种有效的方法来减少模型大小和推理成本,而无需重新训练。尽管其重要性,各种先进的量化方法和位宽配置的ASR模型的性能影响仍然不清楚。在这项工作中,我们提出了一个全面的基准八个国家的最先进的(SOTA)PTQ方法应用于两个领先的边缘ASR模型的家庭,耳语和月光。我们系统地评估模型性能(即,准确性、内存I O和位操作),分析量化和各种配置对权重和激活的影响。我们的框架建立在LLM压缩工具包的扩展之上,集成了边缘ASR模型、各种先进的量化算法、统一的校准和评估数据管道以及详细的分析工具。我们的研究结果的特点效率和准确性之间的权衡,表明即使3位量化可以成功的高容量模型时,使用先进的PTQ技术。这些发现为在低功耗、始终在线的边缘设备上优化ASR模型提供了有价值的见解。
摘要:Recent advances in Automatic Speech Recognition (ASR) have demonstrated remarkable accuracy and robustness in diverse audio applications, such as live transcription and voice command processing. However, deploying these models on resource constrained edge devices (e.g., IoT device, wearables) still presents substantial challenges due to strict limits on memory, compute and power. Quantization, particularly Post-Training Quantization (PTQ), offers an effective way to reduce model size and inference cost without retraining. Despite its importance, the performance implications of various advanced quantization methods and bit-width configurations on ASR models remain unclear. In this work, we present a comprehensive benchmark of eight state-of-the-art (SOTA) PTQ methods applied to two leading edge-ASR model families, Whisper and Moonshine. We systematically evaluate model performances (i.e., accuracy, memory I O and bit operations) across seven diverse datasets from the open ASR leaderboard, analyzing the impact of quantization and various configurations on both weights and activations. Built on an extension of the LLM compression toolkit, our framework integrates edge-ASR models, diverse advanced quantization algorithms, a unified calibration and evaluation data pipeline, and detailed analysis tools. Our results characterize the trade-offs between efficiency and accuracy, demonstrating that even 3-bit quantization can succeed on high capacity models when using advanced PTQ techniques. These findings provide valuable insights for optimizing ASR models on low-power, always-on edge devices.


【5】Re-Bottleneck: Latent Re-Structuring for Neural Audio Autoencoders

标题:重新瓶颈:神经音频自动编码器的潜在重组
链接:http://arxiv.org/pdf/2507.07867v1
作者:Bralios, Jonah Casebeer, Paris Smaragdis

备注:Accepted at IEEE MLSP 2025

摘要:神经音频编解码器和自动编码器已经成为音频压缩、传输、特征提取和潜在空间生成的通用模型。然而,一个关键的限制是,大多数都是为了最大限度地提高重建保真度而训练的,通常忽略了在各种下游应用中实现最佳性能所必需的特定潜在结构。我们提出了一个简单的事后框架,通过修改预先训练的自动编码器的瓶颈来解决这个问题。我们的方法引入了一个“重新瓶颈”,一个内部瓶颈专门通过潜在的空间损失来训练,以灌输用户定义的结构。我们在三个实验中证明了该框架的有效性。首先,我们在不牺牲重建质量的情况下对潜在通道进行排序。其次,我们将潜在的语义嵌入,分析下游扩散建模的影响。第三,我们引入等方差,确保对输入波形的滤波操作直接对应于潜在空间中的特定变换。最终,我们的Re-Bottleneck框架提供了一种灵活有效的方法来定制神经音频模型的表示,使它们能够以最少的额外训练无缝地满足不同应用的各种需求。
摘要:Neural audio codecs and autoencoders have emerged as versatile models for audio compression, transmission, feature-extraction, and latent-space generation. However, a key limitation is that most are trained to maximize reconstruction fidelity, often neglecting the specific latent structure necessary for optimal performance in diverse downstream applications. We propose a simple, post-hoc framework to address this by modifying the bottleneck of a pre-trained autoencoder. Our method introduces a "Re-Bottleneck", an inner bottleneck trained exclusively through latent space losses to instill user-defined structure. We demonstrate the framework's effectiveness in three experiments. First, we enforce an ordering on latent channels without sacrificing reconstruction quality. Second, we align latents with semantic embeddings, analyzing the impact on downstream diffusion modeling. Third, we introduce equivariance, ensuring that a filtering operation on the input waveform directly corresponds to a specific transformation in the latent space. Ultimately, our Re-Bottleneck framework offers a flexible and efficient way to tailor representations of neural audio models, enabling them to seamlessly meet the varied demands of different applications with minimal additional training.


【6】End-to-end Acoustic-linguistic Emotion and Intent Recognition Enhanced by Semi-supervised Learning

标题:半监督学习增强端到端声学语言情感和意图识别
链接:http://arxiv.org/pdf/2507.07806v1
作者:Rathi Adarshi Rammohan, Kevin Scheck, Sheng Li, Tanja Schultz

备注:Accepted by EMBC 2025

摘要:从语音中识别情感和意图是至关重要的,并且在人机交互中得到了广泛的研究。社交媒体平台、聊天机器人和其他技术的快速发展导致了来自用户的大量语音数据流。然而,手动注释这些数据是昂贵的,使得训练机器学习模型以用于识别目的具有挑战性。为此,我们建议应用半监督学习将大规模的未标记数据与相对较小的标记数据集结合起来。我们训练端到端的声学和语言模型,每个模型都采用多任务学习进行情感和意图识别。比较了固定匹配学习和完全匹配学习两种半监督学习方法。实验结果表明,半监督学习方法提高了语音情感和意图识别的声学和文本数据的模型性能。最佳模型的后期融合分别通过12.3%和10.4%的联合识别平衡度量优于声学和文本基线。
摘要:Emotion and intent recognition from speech is essential and has been widely investigated in human-computer interaction. The rapid development of social media platforms, chatbots, and other technologies has led to a large volume of speech data streaming from users. Nevertheless, annotating such data manually is expensive, making it challenging to train machine learning models for recognition purposes. To this end, we propose applying semi-supervised learning to incorporate a large scale of unlabelled data alongside a relatively smaller set of labelled data. We train end-to-end acoustic and linguistic models, each employing multi-task learning for emotion and intent recognition. Two semi-supervised learning approaches, including fix-match learning and full-match learning, are compared. The experimental results demonstrate that the semi-supervised learning approaches improve model performance in speech emotion and intent recognition from both acoustic and text data. The late fusion of the best models outperforms the acoustic and text baselines by joint recognition balance metrics of 12.3% and 10.4%, respectively.


【7】SecureSpeech: Prompt-based Speaker and Content Protection

标题:SecureSpeech:基于预算的演讲者和内容保护
链接:http://arxiv.org/pdf/2507.07799v1
作者:oh Hui Hui, Xiaoxiao Miao, Xin Wang
备注:Accepted by IEEE International Joint Conference on Biometrics (IJCB) 2025
摘要:鉴于越来越多的隐私问题,从身份盗窃和重新识别扬声器通过在语音领域的内容,本文提出了一种基于匿名的语音生成管道,确保双重匿名的扬声器身份和口语内容。这是通过以下方式解决的:1)生成由描述符控制的不受源说话者约束的说话者身份,以及2)使用名称实体识别模型和大型语言模型替换原始文本中的敏感内容。该流水线利用匿名化的说话者身份和文本来经由文本到语音合成模型生成高保真、隐私友好的语音。实验结果表明,在保持相当水平的内容保留和音频质量的同时,实现了显着的隐私保护。本文还研究了不同的扬声器描述的效用和隐私生成的语音,以确定潜在的偏见的影响。
摘要:Given the increasing privacy concerns from identity theft and the re-identification of speakers through content in the speech field, this paper proposes a prompt-based speech generation pipeline that ensures dual anonymization of both speaker identity and spoken content. This is addressed through 1) generating a speaker identity unlinkable to the source speaker, controlled by descriptors, and 2) replacing sensitive content within the original text using a name entity recognition model and a large language model. The pipeline utilizes the anonymized speaker identity and text to generate high-fidelity, privacy-friendly speech via a text-to-speech synthesis model. Experimental results demonstrate an achievement of significant privacy protection while maintaining a decent level of content retention and audio quality. This paper also investigates the impact of varying speaker descriptions on the utility and privacy of generated speech to determine potential biases.


【8】Assessing the Alignment of Audio Representations with Timbre Similarity Ratings

标题:评估音频表示与音色相似度评级的一致性
链接:http://arxiv.org/pdf/2507.07764v1
作者:an, Stefan Lattner, Charalampos Saitis

备注:Accepted to ISMIR 2025

摘要:心理声学所谓的“音色空间”通过多维缩放将乐器声音的感知相似性评级映射到低维嵌入上,但存在可扩展性问题并且无法泛化。最近的音频(音乐和语音)质量评估以及图像相似性的结果表明,深度学习能够产生与人类感知良好一致的嵌入,同时在很大程度上不受这些限制。虽然现有的人类评级音色相似性数据不足以训练深度神经网络(对334个音频样本进行2,614个成对评级),但它可以作为音频模型的测试数据。在本文中,我们引入指标来评估对齐的不同的音频表示与人类的判断音色相似性,通过比较的绝对值和排名的嵌入距离人类相似性评级。我们的评估涉及三个基于信号处理的表示,从预训练模型中提取的十二个表示,以及从一个新的声音匹配模型中提取的三个表示。其中,CLAP模型和声音匹配模型提取的受图像风格迁移启发的风格嵌入效果明显优于其他模型,显示了它们在建模音色相似性方面的潜力。
摘要:Psychoacoustical so-called "timbre spaces" map perceptual similarity ratings of instrument sounds onto low-dimensional embeddings via multidimensional scaling, but suffer from scalability issues and are incapable of generalization. Recent results from audio (music and speech) quality assessment as well as image similarity have shown that deep learning is able to produce embeddings that align well with human perception while being largely free from these constraints. Although the existing human-rated timbre similarity data is not large enough to train deep neural networks (2,614 pairwise ratings on 334 audio samples), it can serve as test-only data for audio models. In this paper, we introduce metrics to assess the alignment of diverse audio representations with human judgments of timbre similarity by comparing both the absolute values and the rankings of embedding distances to human similarity ratings. Our evaluation involves three signal-processing-based representations, twelve representations extracted from pre-trained models, and three representations extracted from a novel sound matching model. Among them, the style embeddings inspired by image style transfer, extracted from the CLAP model and the sound matching model, remarkably outperform the others, showing their potential in modeling timbre similarity.


【9】DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction

标题:DMF2Mel:一种用于脑电驱动Mel谱图重建的动态多尺度融合网络
链接:http://arxiv.org/pdf/2507.07526v1
作者:an, Sheng Zhang, Jingjing Zhang, Enrui Liu, Xinhui Li, Minggang Zhao, Zhao Lv
备注:Accepted by ACM MM 2025
摘要:从大脑信号中解码语音是一个具有挑战性的研究问题。尽管现有技术在单词或字母水平上重建听觉刺激的mel谱图方面取得了进展,但在精确重建分钟水平的连续想象语音方面仍存在核心挑战:传统模型难以平衡长序列解码中的时间依赖性建模效率和信息保留。为了解决这个问题,本文提出了动态多尺度融合网络(DMF 2 Mel),它由四个核心组件组成:动态对比特征聚合模块(DC-FAM),分层注意引导多尺度网络(HAMS-Net),SplineMap注意机制和双向状态空间模块(convMamba)。具体地说,DC-FAM通过局部卷积和全局注意机制将语音相关的“前景特征”与噪声“背景特征”分离,有效地抑制了干扰并增强了瞬态信号的表示。HAMS-Net基于U-Net框架,实现了高层语义和底层细节的跨尺度融合。SplineMap注意力机制集成了自适应门控Kolmogorov-Arnold网络(AGKAN),将全局上下文建模与基于样条的局部拟合相结合。convMamba以线性复杂度捕获长范围的时间依赖性,并增强了非线性动态建模能力。SparrKULee数据集上的结果显示,DMF 2 Mel在已知受试者的mel频谱图重建中实现了0.074的Pearson相关系数(相对于基线提高了48%),并且对于未知受试者实现了0.048的Pearson相关系数(相对于基线提高了35%)。https: github.com fchest DMF2Mel
摘要:Decoding speech from brain signals is a challenging research problem. Although existing technologies have made progress in reconstructing the mel spectrograms of auditory stimuli at the word or letter level, there remain core challenges in the precise reconstruction of minute-level continuous imagined speech: traditional models struggle to balance the efficiency of temporal dependency modeling and information retention in long-sequence decoding. To address this issue, this paper proposes the Dynamic Multiscale Fusion Network (DMF2Mel), which consists of four core components: the Dynamic Contrastive Feature Aggregation Module (DC-FAM), the Hierarchical Attention-Guided Multi-Scale Network (HAMS-Net), the SplineMap attention mechanism, and the bidirectional state space module (convMamba). Specifically, the DC-FAM separates speech-related "foreground features" from noisy "background features" through local convolution and global attention mechanisms, effectively suppressing interference and enhancing the representation of transient signals. HAMS-Net, based on the U-Net framework,achieves cross-scale fusion of high-level semantics and low-level details. The SplineMap attention mechanism integrates the Adaptive Gated Kolmogorov-Arnold Network (AGKAN) to combine global context modeling with spline-based local fitting. The convMamba captures long-range temporal dependencies with linear complexity and enhances nonlinear dynamic modeling capabilities. Results on the SparrKULee dataset show that DMF2Mel achieves a Pearson correlation coefficient of 0.074 in mel spectrogram reconstruction for known subjects (a 48% improvement over the baseline) and 0.048 for unknown subjects (a 35% improvement over the baseline).Code is available at: https: github.com fchest DMF2Mel.


【10】IML-Spikeformer: Input-aware Multi-Level Spiking Transformer for Speech Processing

标题:IML-Spikeformer:用于语音处理的输入感知多层Spikeformer
链接:http://arxiv.org/pdf/2507.07396v1
作者:ng, Shimin Zhang, Yuhong Chou, Jibin Wu, Haizhou Li
备注:Under review of TNNLS
摘要:尖峰神经网络(SNN),灵感来自生物神经机制,代表了一个有前途的神经形态计算范式,提供了节能的替代传统的人工神经网络(ANN)。尽管SNN架构的有效性已得到证明,但它在大规模语音处理任务上仍难以实现有竞争力的性能。两个关键挑战阻碍了进展:(1)多时间步尖峰放电导致的训练期间的高计算开销,以及(2)缺乏针对语音处理任务的大规模SNN架构。为了克服这些问题,我们引入了输入感知的多级Spikeformer,即IML-Spikeformer,一种专门为大规模语音处理设计的尖峰Transformer架构。我们设计的核心是输入感知多级尖峰(IMLS)机制,它使用自适应输入感知阈值方案在单个时间步内模拟多时间步尖峰发射。IML-Spikeformer进一步将重新参数化的尖峰自我注意(RepSSA)模块与分层衰减掩模(HDM)集成在一起,形成HD-RepSSA模块。该模块提高了注意力地图的精度,并能够对语音信号中的多尺度时间依赖性进行建模。实验表明,IML-Spikeformer在AiShell-1和Librispeech-960上分别实现了6.0%和3.4%的字错误率,与传统的ANN Transformers相当,同时分别减少了4.64times $和4.32times $的理论推理能耗。IML-Spikeformer标志着可扩展SNN架构在任务性能和能源效率方面的大规模语音处理的进步。
摘要:Spiking Neural Networks (SNNs), inspired by biological neural mechanisms, represent a promising neuromorphic computing paradigm that offers energy-efficient alternatives to traditional Artificial Neural Networks (ANNs). Despite proven effectiveness, SNN architectures have struggled to achieve competitive performance on large-scale speech processing task. Two key challenges hinder progress: (1) the high computational overhead during training caused by multi-timestep spike firing, and (2) the absence of large-scale SNN architectures tailored to speech processing tasks. To overcome the issues, we introduce Input-aware Multi-Level Spikeformer, i.e. IML-Spikeformer, a spiking Transformer architecture specifically designed for large-scale speech processing. Central to our design is the Input-aware Multi-Level Spike (IMLS) mechanism, which simulate multi-timestep spike firing within a single timestep using an adaptive, input-aware thresholding scheme. IML-Spikeformer further integrates a Reparameterized Spiking Self-Attention (RepSSA) module with a Hierarchical Decay Mask (HDM), forming the HD-RepSSA module. This module enhances the precision of attention maps and enables modeling of multi-scale temporal dependencies in speech signals. Experiments demonstrate that IML-Spikeformer achieves word error rates of 6.0 % on AiShell-1 and 3.4 % on Librispeech-960, comparable to conventional ANN transformers while reducing theoretical inference energy consumption by 4.64$ times$ and 4.32$ times$ respectively. IML-Spikeformer marks an advance of scalable SNN architectures for large-scale speech processing in both task performance and energy efficiency.


【11】VP-SelDoA: Visual-prompted Selective DoA Estimation of Target Sound via Semantic-Spatial Matching

标题:VP-SelDoA:通过语义空间匹配的视觉提示选择性DoA目标声音估计
链接:http://arxiv.org/pdf/2507.07384v1
作者:Xinyuan Qian, Hongxu Zhu, Jiadong Wang, Kainan Chen, Haizhou Li
备注:Under Review
摘要:视听声源定位(AV-SSL)通过利用听觉和视觉信号的互补强度来识别声源的位置。然而,现有的AV-SSL方法遇到了三个主要挑战:1)无法在多源场景中选择性地隔离目标声源,2)语义视觉特征和空间声学特征之间的不匹配,以及3)过度依赖成对的视听数据。为了克服这些限制,我们引入了跨实例视听定位(CI-AVL),这是一种新的任务,它利用来自同一声音事件类别的不同实例的图像来定位目标声源,从而减少对配对数据的依赖,同时增强泛化能力。我们提出的VP-SelDoA通过语义级模态融合来解决这一具有挑战性的任务,并采用频率-时间ConMamba架构来生成用于声音隔离的目标选择性掩模。我们进一步开发了一个语义空间匹配机制,通过集成的交叉和自我注意机制,对齐异构的语义和空间特征。为了便于CI-AVL研究,我们构建了一个名为VGG-SSL的大规模数据集,包括296个声音事件类别的13,981个空间音频片段。大量的实验表明,我们提出的方法优于国家的最先进的视听定位方法,实现了平均绝对误差(MAE)的12.04和准确度(ACC)的78.23%。
摘要:Audio-visual sound source localization (AV-SSL) identifies the position of a sound source by exploiting the complementary strengths of auditory and visual signals. However, existing AV-SSL methods encounter three major challenges: 1) inability to selectively isolate the target sound source in multi-source scenarios, 2) misalignment between semantic visual features and spatial acoustic features, and 3) overreliance on paired audio-visual data. To overcome these limitations, we introduce Cross-Instance Audio-Visual Localization (CI-AVL), a novel task that leverages images from different instances of the same sound event category to localize target sound sources, thereby reducing dependence on paired data while enhancing generalization capabilities. Our proposed VP-SelDoA tackles this challenging task through a semantic-level modality fusion and employs a Frequency-Temporal ConMamba architecture to generate target-selective masks for sound isolation. We further develop a Semantic-Spatial Matching mechanism that aligns the heterogeneous semantic and spatial features via integrated cross- and self-attention mechanisms. To facilitate the CI-AVL research, we construct a large-scale dataset named VGG-SSL, comprising 13,981 spatial audio clips across 296 sound event categories. Extensive experiments show that our proposed method outperforms state-of-the-art audio-visual localization methods, achieving a mean absolute error (MAE) of 12.04 and an accuracy (ACC) of 78.23%.


【12】SonicMotion: Dynamic Spatial Audio Soundscapes with Latent Diffusion Models

标题:SonicMotion:具有潜在扩散模型的动态空间音频声景
链接:http://arxiv.org/pdf/2507.07318v1
作者:Templin, Yanda Zhu, Hao Wang
摘要:空间音频是沉浸式娱乐(如VR AR)不可或缺的一部分,在电影和音乐中也越来越受欢迎。空间音频的最常见格式被描述为一阶高保真度立体声(FOA)。我们寻求扩展FOA生成AI模型的最新进展,以生成具有动态声源的3D场景。我们提出的端到端模型SonicMotion有两种变体,它们在用户输入和声源定位的精确度方面有所不同。除了我们的模型,我们还提出了一个新的模拟空间音频字幕对数据集。我们的模型的评估表明,他们能够匹配的语义对齐和音频质量的最先进的模型,同时捕捉所需的空间属性。
摘要:Spatial audio is an integral part of immersive entertainment, such as VR AR, and has seen increasing popularity in cinema and music as well. The most common format of spatial audio is described as first-order Ambisonics (FOA). We seek to extend recent advancements in FOA generative AI models to enable the generation of 3D scenes with dynamic sound sources. Our proposed end-to-end model, SonicMotion, comes in two variations which vary in their user input and level of precision in sound source localization. In addition to our model, we also present a new dataset of simulated spatial audio-caption pairs. Evaluation of our models demonstrate that they are capable of matching the semantic alignment and audio quality of state of the art models while capturing the desired spatial attributes.


【13】ViDove: A Translation Agent System with Multimodal Context and Memory-Augmented Reasoning

标题:ViDove:一个具有多模式上下文和记忆增强推理的翻译代理系统
链接:http://arxiv.org/pdf/2507.07306v1
作者:Wei Dai, Jiaen Liu, Ching Wing Kwok, Zongheng Wu, Xudong Xiao, Ao Sun, Sheng Fu, Jianyuan Zhan, Yian Wang, Takatomo Saito, Sicheng Lai
摘要:基于LLM的翻译代理已经实现了高度人性化的翻译结果,并且能够以更高的效率处理更长,更复杂的上下文。然而,它们通常限于纯文本输入。在本文中,我们介绍ViDove,翻译代理系统设计的多模态输入。受人工翻译工作流程的启发,ViDove利用视觉和上下文背景信息来增强翻译过程。此外,我们集成了一个多模态记忆系统和长期短期记忆模块丰富的特定领域的知识,使代理在现实世界的情况下,执行更准确和自适应。因此,ViDove在字幕生成和一般翻译任务中实现了显著更高的翻译质量,与之前的最先进基线相比,BLEU分数提高了28%,SubER提高了15%。此外,我们还介绍了DoveBench,这是一个用于长格式自动视频字幕和翻译的新基准,具有17小时的高质量人工注释数据。我们的代码可从以下网址获得:https: github.com pigeonai-org ViDove
摘要:LLM-based translation agents have achieved highly human-like translation results and are capable of handling longer and more complex contexts with greater efficiency. However, they are typically limited to text-only inputs. In this paper, we introduce ViDove, a translation agent system designed for multimodal input. Inspired by the workflow of human translators, ViDove leverages visual and contextual background information to enhance the translation process. Additionally, we integrate a multimodal memory system and long-short term memory modules enriched with domain-specific knowledge, enabling the agent to perform more accurately and adaptively in real-world scenarios. As a result, ViDove achieves significantly higher translation quality in both subtitle generation and general translation tasks, with a 28% improvement in BLEU scores and a 15% improvement in SubER compared to previous state-of-the-art baselines. Moreover, we introduce DoveBench, a new benchmark for long-form automatic video subtitling and translation, featuring 17 hours of high-quality, human-annotated data. Our code is available here: https: github.com pigeonai-org ViDove


【14】Audio-Visual Speech Separation via Bottleneck Iterative Network

标题:瓶颈迭代网络实现视听语音分离
链接:http://arxiv.org/pdf/2507.07270v1
作者:ang, Shiv Shankar, Trang Nguyen, Andrea Fanelli, Madalina Fiterau
备注:Accepted to the 42nd International Conference on Machine Learning Workshop on Machine Learning for Audio
摘要:从非听觉线索的信息的整合可以显着提高语音分离模型的性能。通常,此类模型使用深度特定于模态的网络来获取单模态特征,并且存在成本太高或重量太轻但容量不足的风险。在这项工作中,我们提出了一种迭代表示细化方法,称为瓶颈迭代网络(BIN),一种通过轻量级融合块重复进行的技术,同时通过融合令牌检查融合表示。这有助于提高模型的容量,同时避免模型大小的大幅增加以及模型性能和训练成本之间的平衡。我们在具有挑战性的嘈杂视听语音分离任务上测试BIN,并表明我们的方法在NTCD-TIMIT和LRS 3 +WHAM上的SI-SDRi方面始终优于最先进的基准模型!数据集,同时在几乎所有设置中实现训练和GPU推理时间减少50%以上。
摘要:Integration of information from non-auditory cues can significantly improve the performance of speech-separation models. Often such models use deep modality-specific networks to obtain unimodal features, and risk being too costly or lightweight but lacking capacity. In this work, we present an iterative representation refinement approach called Bottleneck Iterative Network (BIN), a technique that repeatedly progresses through a lightweight fusion block, while bottlenecking fusion representations by fusion tokens. This helps improve the capacity of the model, while avoiding major increase in model size and balancing between the model performance and training cost. We test BIN on challenging noisy audio-visual speech separation tasks, and show that our approach consistently outperforms state-of-the-art benchmark models with respect to SI-SDRi on NTCD-TIMIT and LRS3+WHAM! datasets, while simultaneously achieving a reduction of more than 50% in training and GPU inference time across nearly all settings.


机器翻译由腾讯交互翻译提供,仅供参考