今日论文合集:cs.SD语音12篇,eess.AS音频处理11篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】P2Mark: Plug-and-play Parameter-intrinsic Watermarking for Neural Speech  Generation
标题: P2Mark:即插即用的参数内在水印算法
链接:https://arxiv.org/abs/2504.05197

作者: Yong Ren,  Jiangyan Yi,  Tao Wang,  Jianhua Tao,  Zhengqi Wen,  Chenxing Li,  Zheng Lian,  Ruibo Fu,  Ye Bai,  Xiaohui Zhang 
摘要:最近,开源社区中出现了大量先进的神经语音生成方法。这虽然促进了技术的应用和发展,但也增加了防止滥用生成语音和保护版权的难度。音频水印技术是一种主动保护语音的有效方法,但当神经语音生成方法的源代码和模型权值是开源的时,基于先前水印方法的音频水印很容易被去除或篡改。提出了一种用于神经语音生成系统保护的即插即用参数内在水印(P2Mark)方法。P2Mark的主要优点是通过训练水印适配器将水印信息以参数的形式灵活地集成到神经语音生成模型中,而不是将水印以特征的形式注入模型中。在具有水印嵌入的水印适配器与预训练的生成模型合并之后,水印信息不能被容易地移除或操纵。因此,P2Mark将是在开源白盒场景中主动追踪和保护神经语音生成模型版权的可靠选择。我们在神经语音生成中的两种主要解码器上验证了P2Mark:声码器和编解码器。实验结果表明,P2Mark在水印提取准确性、水印不可感知性和鲁棒性等方面达到了与不能用于开源白盒保护场景的最新音频水印方法相当的性能。
摘要:Recently, a large number of advanced neural speech generation methods have emerged in the open-source community. Although this has facilitated the application and development of technology, it has also increased the difficulty of preventing the abuse of generated speech and protecting copyrights. Audio watermarking technology is an effective method for proactively protecting generated speech, but when the source codes and model weights of the neural speech generation methods are open-sourced, audio watermarks based on previous watermarking methods can be easily removed or manipulated. This paper proposes a Plug-and-play Parameter-intrinsic WaterMarking (P2Mark) method for neural speech generation system protection. The main advantage of P2Mark is that the watermark information is flexibly integrated into the neural speech generation model in the form of parameters by training a watermark adapter rather than injecting the watermark into the model in the form of features. After the watermark adapter with the watermark embedding is merged with the pre-trained generation model, the watermark information cannot be easily removed or manipulated. Therefore, P2Mark will be a reliable choice for proactively tracing and protecting the copyrights of neural speech generation models in open-source white-box scenarios. We validated P2Mark on two main types of decoders in neural speech generation: vocoder and codec. Experimental results show that P2Mark achieves performance comparable to state-of-the-art audio watermarking methods that cannot be used for open-source white-box protection scenarios in terms of watermark extraction accuracy, watermark imperceptibility, and robustness.


【2】 Leveraging Label Potential for Enhanced Multimodal Emotion Recognition
标题: 利用标签潜力增强多模式情感识别
链接:https://arxiv.org/abs/2504.05158

作者: Xuechun Shao,  Yinfeng Yu,  Liejun Wang 
备注:Main paper (8 pages). Accepted for publication by IJCNN 2025
摘要:多模态情感识别(MER)试图整合各种模态来准确地预测情感状态。然而,目前的研究大多集中在音频和文本特征的融合,忽略了情感标签中有价值的信息。这种疏忽可能会阻碍现有方法的性能,因为情感标签包含丰富的,有见地的信息,可以显着帮助MER。我们引入了一种名为标签信号引导多模式情感识别(LSGMER)的新模型来克服这一限制。该模型旨在充分利用情感标签信息的力量,以提高MER的分类精度和稳定性。具体来说,LSGMER采用了一个标签信号增强模块,通过标签嵌入与音频和文本特征交互来优化模态特征的表示,使其能够精确地捕捉情感的细微差别。此外,我们提出了一个联合目标优化(JOO)的方法,以提高分类精度,通过引入属性预测一致性约束(APC),加强融合功能和情感类别之间的对齐。在IEMOCAP和MELD数据集上进行的大量实验证明了我们提出的LSGMER模型的有效性。
摘要:Multimodal emotion recognition (MER) seeks to integrate various modalities to predict emotional states accurately. However, most current research focuses solely on the fusion of audio and text features, overlooking the valuable information in emotion labels. This oversight could potentially hinder the performance of existing methods, as emotion labels harbor rich, insightful information that could significantly aid MER. We introduce a novel model called Label Signal-Guided Multimodal Emotion Recognition (LSGMER) to overcome this limitation. This model aims to fully harness the power of emotion label information to boost the classification accuracy and stability of MER. Specifically, LSGMER employs a Label Signal Enhancement module that optimizes the representation of modality features by interacting with audio and text features through label embeddings, enabling it to capture the nuances of emotions precisely. Furthermore, we propose a Joint Objective Optimization(JOO) approach to enhance classification accuracy by introducing the Attribution-Prediction Consistency Constraint (APC), which strengthens the alignment between fused features and emotion categories. Extensive experiments conducted on the IEMOCAP and MELD datasets have demonstrated the effectiveness of our proposed LSGMER model.


【3】 Deconstructing Jazz Piano Style Using Machine Learning
标题: 使用机器学习解构爵士钢琴风格
链接:https://arxiv.org/abs/2504.05009

作者: Huw Cheston,  Reuben Bance,  Peter M. C. Harrison 
备注:Paper: 40 pages, 11 figures, 1 table. Supplementary material: 33 pages, 48 figures, 6 tables
摘要:艺术风格已经被研究了几个世纪,机器学习的最新进展为通过计算来理解它创造了新的可能性。然而,确保机器学习模型产生符合实践者和批评者利益的见解仍然是一个重大挑战。在这里,我们专注于音乐风格,这得益于丰富的理论和数学分析传统。我们训练了各种监督学习模型,以在精心策划的84小时录音数据集中识别20位标志性爵士音乐家,并解释他们的决策过程。我们的模型包括一个新颖的多输入架构,使四个音乐领域(旋律,和声,节奏和动态)进行单独分析。这些模型使我们能够解决音乐理论中的基本问题,并推进音乐表演者识别的最新技术(20个类别的准确率为94%)。我们发布了我们的模型的开源实现和一个用于探索音乐风格的附带Web应用程序。
摘要:Artistic style has been studied for centuries, and recent advances in machine learning create new possibilities for understanding it computationally. However, ensuring that machine-learning models produce insights aligned with the interests of practitioners and critics remains a significant challenge. Here, we focus on musical style, which benefits from a rich theoretical and mathematical analysis tradition. We train a variety of supervised-learning models to identify 20 iconic jazz musicians across a carefully curated dataset of 84 hours of recordings, and interpret their decision-making processes. Our models include a novel multi-input architecture that enables four musical domains (melody, harmony, rhythm, and dynamics) to be analysed separately. These models enable us to address fundamental questions in music theory and also advance the state-of-the-art in music performer identification (94% accuracy across 20 classes). We release open-source implementations of our models and an accompanying web application for exploring musical styles.


【4】 One Quantizer is Enough: Toward a Lightweight Audio Codec
标题: 一个量化器就足够了:迈向轻量级音频编解码器
链接:https://arxiv.org/abs/2504.04949

作者: Linwei Zhai,  Han Ding,  Cui Zhao,  fei wang,  Ge Wang,  Wang Zhi,  Wei Xi 
摘要:神经音频编解码器最近因其压缩高保真音频并生成可用于下游生成建模任务的离散令牌的能力而受到关注。然而,领先的方法往往依赖于资源密集型模型和多量化器架构,导致相当大的计算开销和现实世界的适用性受到限制。在本文中,我们介绍了SQCodec,一种轻量级的神经音频编解码器,它利用单个量化器来解决这些限制。SQCodec探索了精简的卷积网络和本地Transformer模块,以及TConv,这是一种旨在捕获多个时间尺度上的声学变化的新颖机制,从而提高重建保真度,同时降低模型复杂性。在不同数据集上进行的大量实验表明,SQCodec实现了与多量化器基线相当的音频质量,而其单量化器设计提供了增强的适应性,其轻量级架构将资源消耗降低了一个数量级。源代码可在https://github.com/zhai-lw/SQCodec上公开获得。
摘要:Neural audio codecs have recently gained traction for their ability to compress high-fidelity audio and generate discrete tokens that can be utilized in downstream generative modeling tasks. However, leading approaches often rely on resource-intensive models and multi-quantizer architectures, resulting in considerable computational overhead and constrained real-world applicability. In this paper, we present SQCodec, a lightweight neural audio codec that leverages a single quantizer to address these limitations. SQCodec explores streamlined convolutional networks and local Transformer modules, alongside TConv, a novel mechanism designed to capture acoustic variations across multiple temporal scales, thereby enhancing reconstruction fidelity while reducing model complexity. Extensive experiments across diverse datasets show that SQCodec achieves audio quality comparable to multi-quantizer baselines, while its single-quantizer design offers enhanced adaptability and its lightweight architecture reduces resource consumption by an order of magnitude. The source code is publicly available at https://github.com/zhai-lw/SQCodec.


【5】 Diff-SSL-G-Comp: Towards a Large-Scale and Diverse Dataset for Virtual  Analog Modeling
标题: 迪夫-SSL-G-Comp:迈向大规模、多样化的虚拟模拟建模数据集
链接:https://arxiv.org/abs/2504.04589

作者: Yicheng Gu,  Runsong Zhang,  Lauri Juvela,  Zhizheng Wu 
备注:Submitted to DAFx 2025
摘要:虚拟模拟(VA)建模旨在通过算法模拟硬件电路的行为,以数字方式复制其音调。动态范围压缩器(DRC)是一个音频处理模块,通过减少和放大响亮和安静声音的音量来控制音轨的动态,这在音乐制作中至关重要。近年来,基于神经网络的VA建模在产生高保真模型方面显示出巨大的潜力。但由于数据量和多样性的不足,它们在不同参数设置和输入声音下的泛化能力仍然有限。为了解决这个问题,我们提出了Diff-SSL-G-Comp,这是第一个用于建模SSL 500 G-Bus Compressor的大规模和多样化数据集。具体来说,我们从剑桥多轨图书馆手动收集了175首未掌握的歌曲。我们记录了220个参数组合的压缩音频,产生了一个广泛的2528小时的数据集,包括不同的流派,乐器,节奏和键。此外,为了便于使用我们提出的数据集,我们在各种开源的黑盒和灰盒模型以及白盒插件中进行了基准实验。我们还在不同的数据子集中进行了消融研究,以说明提高数据多样性和数量的有效性。数据集和演示在我们的项目页面上:http://www.yichenggu.com/DiffSSLGComp/。
摘要:Virtual Analog (VA) modeling aims to simulate the behavior of hardware circuits via algorithms to replicate their tone digitally. Dynamic Range Compressor (DRC) is an audio processing module that controls the dynamics of a track by reducing and amplifying the volumes of loud and quiet sounds, which is essential in music production. In recent years, neural-network-based VA modeling has shown great potential in producing high-fidelity models. However, due to the lack of data quantity and diversity, their generalization ability in different parameter settings and input sounds is still limited. To tackle this problem, we present Diff-SSL-G-Comp, the first large-scale and diverse dataset for modeling the SSL 500 G-Bus Compressor. Specifically, we manually collected 175 unmastered songs from the Cambridge Multitrack Library. We recorded the compressed audio in 220 parameter combinations, resulting in an extensive 2528-hour dataset with diverse genres, instruments, tempos, and keys. Moreover, to facilitate the use of our proposed dataset, we conducted benchmark experiments in various open-sourced black-box and grey-box models, as well as white-box plugins. We also conducted ablation studies in different data subsets to illustrate the effectiveness of improved data diversity and quantity. The dataset and demos are on our project page: http://www.yichenggu.com/DiffSSLGComp/.


【6】 Activation Patching for Interpretable Steering in Music Generation
标题: 音乐生成中的可解释引导的激活修补
链接:https://arxiv.org/abs/2504.04479

作者: Simone Facchiano,  Giorgio Strano,  Donato Crisostomi,  Irene Tallini,  Tommaso Mencattini,  Fabio Galasso,  Emanuele Rodolà 
摘要:了解大音频模型如何表示音乐,并使用这种理解来指导生成,既具有挑战性,又未充分探索。受语言模型中的机械可解释性的启发,Transformer残余流中的方向向量是模型分析和控制的关键,我们研究了音频领域中的类似技术。本文首次研究了大型音频模型中的潜在方向向量,并将其用于文本到音乐生成中的音乐属性的连续控制。专注于二进制概念,如节奏(快与慢)和音色(亮与暗),我们使用差异均值方法在策划提示集上计算导向向量。这些向量通过系数缩放并注入到中间激活中,允许对特定音乐特征进行细粒度调制,同时保持整体音频质量。我们分析了转向强度的影响,比较了注入策略,并确定了影响最大的层。我们的研究结果强调了基于方向的转向作为可控音乐生成的一种更机械和可解释的方法的前景。
摘要:Understanding how large audio models represent music, and using that understanding to steer generation, is both challenging and underexplored. Inspired by mechanistic interpretability in language models, where direction vectors in transformer residual streams are key to model analysis and control, we investigate similar techniques in the audio domain. This paper presents the first study of latent direction vectors in large audio models and their use for continuous control of musical attributes in text-to-music generation. Focusing on binary concepts like tempo (fast vs. slow) and timbre (bright vs. dark), we compute steering vectors using the difference-in-means method on curated prompt sets. These vectors, scaled by a coefficient and injected into intermediate activations, allow fine-grained modulation of specific musical traits while preserving overall audio quality. We analyze the effect of steering strength, compare injection strategies, and identify layers with the greatest influence. Our findings highlight the promise of direction-based steering as a more mechanistic and interpretable approach to controllable music generation.


【7】 LoopGen: Training-Free Loopable Music Generation
标题: LoopGen:免训练的Loopable Music Generation
链接:https://arxiv.org/abs/2504.04466

作者: Davide Marincione,  Giorgio Strano,  Donato Crisostomi,  Roberto Ribuoli,  Emanuele Rodolà 
摘要:循环--为无缝重复而设计的短音频片段--是许多音乐流派的核心,特别是那些植根于舞蹈和电子风格的音乐。然而,目前的生成音乐模型很难产生真正可循环的音频,因为仅生成短波形并不能保证从端点平滑过渡到起点,通常会导致可听不连续。循环-为无缝重复而设计的短音频片段-是许多音乐流派的核心,特别是那些植根于舞蹈和电子风格的音乐。然而,当前的生成音乐模型很难产生真正可循环的音频,因为单独生成短波形并不能保证从其端点平滑过渡到其开始,通常会导致可听不连续性。我们通过修改非自回归模型(MAGNeT)来解决这个问题,以生成圆形模式的令牌,让模型在创建其结尾时关注音频的开始。这种仅推理的方法导致了能够意识到未来上下文并自然循环的一代,而不需要任何额外的训练或数据。我们通过计算循环接缝周围的令牌困惑来评估循环转换的一致性,观察到55%的改善。盲听测试进一步证实了基线方法的显著感知增益,平均评级提高了70%。总之,这些结果突出了仅推理方法在改进生成模型方面的有效性,并强调了非自回归方法在上下文感知音乐生成中的优势。
摘要:Loops--short audio segments designed for seamless repetition--are central to many music genres, particularly those rooted in dance and electronic styles. However, current generative music models struggle to produce truly loopable audio, as generating a short waveform alone does not guarantee a smooth transition from its endpoint back to its start, often resulting in audible discontinuities.Loops--short audio segments designed for seamless repetition--are central to many music genres, particularly those rooted in dance and electronic styles. However, current generative music models struggle to produce truly loopable audio, as generating a short waveform alone does not guarantee a smooth transition from its endpoint back to its start, often resulting in audible discontinuities.We address this gap by modifying a non-autoregressive model (MAGNeT) to generate tokens in a circular pattern, letting the model attend to the beginning of the audio when creating its ending. This inference-only approach results in generations that are aware of future context and loop naturally, without the need for any additional training or data. We evaluate the consistency of loop transitions by computing token perplexity around the seam of the loop, observing a 55% improvement. Blind listening tests further confirm significant perceptual gains over baseline methods, improving mean ratings by 70%. Taken together, these results highlight the effectiveness of inference-only approaches in improving generative models and underscore the advantages of non-autoregressive methods for context-aware music generation.


【8】 Formula-Supervised Sound Event Detection: Pre-Training Without Real Data
标题: 公式监督的声音事件检测:没有真实数据的预训练
链接:https://arxiv.org/abs/2504.04428

作者: Yuto Shibata,  Keitaro Tanaka,  Yoshiaki Bando,  Keisuke Imoto,  Hirokatsu Kataoka,  Yoshimitsu Aoki 
备注:Accepted by ICASSP 2025
摘要:在本文中,我们提出了一种新的公式驱动的监督学习(FDSL)框架预训练环境声音分析模型,利用声学信号参数合成通过公式驱动的方法。具体而言,我们概述了详细的程序,并评估其有效性的声音事件检测(SED)。SED任务涉及估计声音事件的类型和时间,特别是由于难以获得足够数量的准确标记的训练数据而受到挑战。此外,众所周知,手动标注的标签通常包含噪声,并且受到标注者主观判断的显著影响。为了解决这些挑战,我们提出了一种新的预训练方法,该方法利用合成数据集Formula-SED,其中声学数据仅基于数学公式生成。所提出的方法通过使用在每个时间步应用的合成参数作为地面真值标签来实现大规模的预训练,从而消除标签噪声和偏差。我们证明,使用Formula-SED进行大规模预训练可显著提高模型准确性并加快训练,DCASE 2023挑战任务4所用DESED数据集的结果证明了这一点。项目页面位于https://yutoshibata07.github.io/Formula-SED/
摘要:In this paper, we propose a novel formula-driven supervised learning (FDSL) framework for pre-training an environmental sound analysis model by leveraging acoustic signals parametrically synthesized through formula-driven methods. Specifically, we outline detailed procedures and evaluate their effectiveness for sound event detection (SED). The SED task, which involves estimating the types and timings of sound events, is particularly challenged by the difficulty of acquiring a sufficient quantity of accurately labeled training data. Moreover, it is well known that manually annotated labels often contain noises and are significantly influenced by the subjective judgment of annotators. To address these challenges, we propose a novel pre-training method that utilizes a synthetic dataset, Formula-SED, where acoustic data are generated solely based on mathematical formulas. The proposed method enables large-scale pre-training by using the synthesis parameters applied at each time step as ground truth labels, thereby eliminating label noise and bias. We demonstrate that large-scale pre-training with Formula-SED significantly enhances model accuracy and accelerates training, as evidenced by our results in the DESED dataset used for DCASE2023 Challenge Task 4. The project page is at https://yutoshibata07.github.io/Formula-SED/


【9】 Selective Masking Adversarial Attack on Automatic Speech Recognition  Systems
标题: 自动语音识别系统的选择性掩蔽对抗攻击
链接:https://arxiv.org/abs/2504.04394

作者: Zheng Fang,  Shenyi Zhang,  Tao Wang,  Bowen Li,  Lingchen Zhao,  Zhangyi Wang 
摘要:广泛的研究表明,自动语音识别(ASR)系统容易受到音频对抗性攻击。目前的攻击主要集中在单源场景,忽略了两个人同时说话的双源场景。为了弥补这一差距,我们提出了一种选择性掩蔽对抗攻击,即SMA攻击,它确保在双源场景中选择一个音频源进行识别,而另一个音频源被静音。为了更好地适应双源场景,我们的SMA攻击从静音音频和选定音频构建正常的双源音频。SMA攻击利用小的高斯噪声对对抗扰动进行滤波,并使用选择性掩蔽优化算法对其进行迭代优化。大量实验表明,SMA攻击可以在双源场景中生成有效且不可感知的音频对抗示例,在Conformer-CTC上实现了100%的平均攻击成功率和37.15dB的信噪比,优于基线。
摘要:Extensive research has shown that Automatic Speech Recognition (ASR) systems are vulnerable to audio adversarial attacks. Current attacks mainly focus on single-source scenarios, ignoring dual-source scenarios where two people are speaking simultaneously. To bridge the gap, we propose a Selective Masking Adversarial attack, namely SMA attack, which ensures that one audio source is selected for recognition while the other audio source is muted in dual-source scenarios. To better adapt to the dual-source scenario, our SMA attack constructs the normal dual-source audio from the muted audio and selected audio. SMA attack initializes the adversarial perturbation with a small Gaussian noise and iteratively optimizes it using a selective masking optimization algorithm. Extensive experiments demonstrate that the SMA attack can generate effective and imperceptible audio adversarial examples in the dual-source scenario, achieving an average success rate of attack of 100% and signal-to-noise ratio of 37.15dB on Conformer-CTC, outperforming the baselines.


【10】 VocalNet: Speech LLM with Multi-Token Prediction for Faster and  High-Quality Generation
标题: VocalNet:具有多令牌预测的语音LLM,用于更快、更高质量的生成
链接:https://arxiv.org/abs/2504.04060

作者: Yuhao Wang,  Heyang Liu,  Ziyang Cheng,  Ronghua Wu,  Qunshan Gu,  Yanfeng Wang,  Yu Wang 
摘要:语音大语言模型(LLM)已经成为语音处理领域的一个重要研究热点。我们提出了VocalNet-1B和VocalNet-8B,这是一系列高性能,低延迟的语音LLM,由可扩展和模型无关的实时语音交互训练框架实现。从传统的下一个令牌预测(NTP)出发,我们引入了多令牌预测(MTP),一种新的方法优化语音LLM,同时提高生成速度和质量。实验表明,VocalNet的性能优于主流的Omni LLM,尽管它使用的训练数据要少得多,同时也大大超过了现有的开源语音LLM。为了支持可重复性和社区发展,我们将在发布时开源所有模型权重,推理代码,训练数据和框架实现。
摘要:Speech large language models (LLMs) have emerged as a prominent research focus in speech processing. We propose VocalNet-1B and VocalNet-8B, a series of high-performance, low-latency speech LLMs enabled by a scalable and model-agnostic training framework for real-time voice interaction. Departing from the conventional next-token prediction (NTP), we introduce multi-token prediction (MTP), a novel approach optimized for speech LLMs that simultaneously improves generation speed and quality. Experiments show that VocalNet outperforms mainstream Omni LLMs despite using significantly less training data, while also surpassing existing open-source speech LLMs by a substantial margin. To support reproducibility and community advancement, we will open-source all model weights, inference code, training data, and framework implementations upon publication.


【11】 Determined blind source separation via modeling adjacent frequency band  correlations in speech signals
标题: 通过建模语音信号中的相邻频段相关性确定盲源分离
链接:https://arxiv.org/abs/2504.03998

作者: Jianyu Wang,  Shanzheng Guan,  Zhengqiao Zhao,  Nicolas Dobigeon,  Jingdong Chen 
摘要:多通道盲源分离(MBSS)是从混合观测信号中分离出感兴趣信号的一种方法,在声学和语音处理领域得到了广泛的研究。现有的MBSS算法,如独立低秩矩阵分析(ILRMA)和多通道非负矩阵分解(MNMF),利用源模型的低秩结构,但假设频率点是独立的。相反,独立向量分析(IVA)不依赖于低秩源模型,而是基于均匀相关假设捕获频率依赖性。在这项工作中,我们证明了相邻的频率箱之间的依赖性显着强于那些在典型的语音信号中相距较远的箱之间。为了解决这个问题,我们引入了一个加权的基于Sinkhorn发散的ILRMA(wsILRMA),同时捕获这些频率间的依赖关系和模型的联合概率分布。我们的方法采用了频率间的相关性约束,从而提高了源分离性能相比,现有的方法,更高的信号失真比(SDR)和源干扰比(SIR)证明。
摘要:Multichannel blind source separation (MBSS), which focuses on separating signals of interest from mixed observations, has been extensively studied in acoustic and speech processing. Existing MBSS algorithms, such as independent low-rank matrix analysis (ILRMA) and multichannel nonnegative matrix factorization (MNMF), utilize the low-rank structure of source models but assume that frequency bins are independent. In contrast, independent vector analysis (IVA) does not rely on a low-rank source model but rather captures frequency dependencies based on a uniform correlation assumption. In this work, we demonstrate that dependencies between adjacent frequency bins are significantly stronger than those between bins that are farther apart in typical speech signals. To address this, we introduce a weighted Sinkhorn divergence-based ILRMA (wsILRMA) that simultaneously captures these inter-frequency dependencies and models joint probability distributions. Our approach incorporates an inter-frequency correlation constraint, leading to improved source separation performance compared to existing methods, as evidenced by higher Signal-to-Distortion Ratios (SDRs) and Source-to-Interference Ratios (SIRs).


【12】 Continuous Boostlet Transform and Associated Uncertainty Principles
标题: 连续Boostlet变换和相关的不确定性原则
链接:https://arxiv.org/abs/2504.03679

作者: Owais Ahmad,  Jasifa Fayaz 
备注:28pages,6 figures
摘要:连续Boostlet变换(CBT)是分析时空信号,特别是声波场的有力工具。CBT克服了经典小波的局限性,利用庞加莱群和各向同性膨胀来捕获自然声场的稀疏特征。本文介绍了CBT的数学框架,包括它的定义,基本性质,以及相关的不确定性原理,如海森堡的,对数,皮特的,和Nazarov的不等式。这些结果阐明了时间和频率的本地化之间的权衡在boostlet域。常数函数和指数函数的实际例子突出了CBT的适应性。CBT在雷达、通信、音频处理和地震分析中有着广泛的应用,它提供了灵活的时频分辨率,是非平稳和瞬态信号的理想选择,也是现代信号处理的宝贵工具。
摘要:The Continuous Boostlet Transform (CBT) is introduced as a powerful tool for analyzing spatiotemporal signals, particularly acoustic wavefields. Overcoming the limitations of classical wavelets, the CBT leverages the Poincar\'e group and isotropic dilations to capture sparse features of natural acoustic fields. This paper presents the mathematical framework of the CBT, including its definition, fundamental properties, and associated uncertainty principles, such as Heisenberg's, logarithmic, Pitt's, and Nazarov's inequalities. These results illuminate the trade-offs between time and frequency localization in the boostlet domain. Practical examples with constant and exponential functions highlight the CBT's adaptability. With applications in radar, communications, audio processing, and seismic analysis, the CBT offers flexible time-frequency resolution, making it ideal for non-stationary and transient signals, and a valuable tool for modern signal processing.


eess.AS音频处理


【1】 Unsupervised Estimation of Nonlinear Audio Effects: Comparing  Diffusion-Based and Adversarial approaches
标题: 非线性音频效果的无监督估计:比较基于扩散的方法和对抗的方法
链接:https://arxiv.org/abs/2504.04751

作者: Eloi Moliner,  Michal Švento,  Alec Wright,  Lauri Juvela,  Pavel Rajmic,  Vesa Välimäki 
备注:Submitted to the 28th International Conference on Digital Audio Effects (DAFx25)
摘要:在没有成对输入输出信号的情况下精确估计非线性音频效果仍然是一个具有挑战性的问题,本文研究了解决这一问题的无监督概率方法。我们介绍了一种方法,新的这种应用,基于扩散生成模型的盲系统识别,使未知的非线性效应的估计,使用黑盒和灰盒模型。本研究将该方法与先前提出的对抗性方法进行了比较,分析了两种方法在不同效果算子参数化和不同可用效果录音长度下的性能。通过对吉他失真效果的实验,我们表明基于扩散的方法提供了更稳定的结果,并且对数据可用性不太敏感,而对抗方法在估计更明显的失真效应方面更优越。我们的研究结果有助于音频效果的鲁棒无监督盲估计,展示了音乐技术中系统识别的扩散模型的潜力。
摘要:Accurately estimating nonlinear audio effects without access to paired input-output signals remains a challenging problem.This work studies unsupervised probabilistic approaches for solving this task. We introduce a method, novel for this application, based on diffusion generative models for blind system identification, enabling the estimation of unknown nonlinear effects using black- and gray-box models. This study compares this method with a previously proposed adversarial approach, analyzing the performance of both methods under different parameterizations of the effect operator and varying lengths of available effected recordings.Through experiments on guitar distortion effects, we show that the diffusion-based approach provides more stable results and is less sensitive to data availability, while the adversarial approach is superior at estimating more pronounced distortion effects. Our findings contribute to the robust unsupervised blind estimation of audio effects, demonstrating the potential of diffusion models for system identification in music technology.


【2】 Bridging the Gap between Continuous and Informative Discrete  Representations by Random Product Quantization
标题: 通过随机积量化弥合连续和信息离散表示之间的差距
链接:https://arxiv.org/abs/2504.04721

作者: Xueqing Li,  Zehan Li,  Boyu Zhu,  Ruihao Jing,  Jian Kang,  Jie Li,  Xiao-Lei Zhang,  Xuelong Li 
摘要:自监督学习已成为语音处理的核心技术,但其表示的高维性使得离散化对于提高效率至关重要。然而,现有的离散化方法仍然遭受显着的信息丢失,导致一个显着的性能差距相比,连续表示。为了克服这些局限性,我们提出了两种基于量化的离散化方法:乘积量化(PQ)和随机乘积量化(RPQ)。PQ将原始特征空间划分为多个子空间,并独立量化每个子向量,产生一组融合的离散单元,保留来自不同子空间的不同信息,从而减轻与单簇量化相关的损失。RPQ通过多次随机采样固定比例的特征维度来构建子向量,从而更好地捕获数据分布中的可变性,从而进一步增强了表示多样性。理论分析表明,RPQ降低了子量化器之间的相关系数rho(其中0 <= rho <= 1)。其量化误差由rho和epsilon-kms的乘积下界,其中epsilon-kms表示单个K均值量化器的量化误差。在LibriSpeech和ML-SUPERB构建的组合数据集上的实验结果表明,PQ和RPQ优于标准K均值离散化,在LibriSpeech上的WER中分别实现了21.8%和20.0%的相对改进,在ML-SUPERB上的CER中分别实现了24.1%和19.6%的相对改进。此外,它们的性能与连续SSL表示具有竞争力,在某些情况下甚至超过连续SSL表示。
摘要:Self-supervised learning has become a core technique in speech processing, but the high dimensionality of its representations makes discretization essential for improving efficiency. However, existing discretization methods still suffer from significant information loss, resulting in a notable performance gap compared to continuous representations. To overcome these limitations, we propose two quantization-based discretization methods: Product Quantization (PQ) and Random Product Quantization (RPQ). PQ partitions the original feature space into multiple subspaces and independently quantizes each sub-vector, producing a fused set of discrete units that retain diverse information from different subspaces, thus mitigating the loss associated with single-cluster quantization. RPQ further enhances representation diversity by randomly sampling a fixed proportion of feature dimensions multiple times to construct sub-vectors, thereby better capturing the variability in the data distribution. Theoretical analysis shows that RPQ reduces the correlation coefficient rho (where 0 <= rho <= 1) between sub-quantizers. Its quantization error is lower-bounded by the product of rho and epsilon-kms, where epsilon-kms denotes the quantization error of a single K-means quantizer. Experimental results on a combined dataset built from LibriSpeech and ML-SUPERB show that PQ and RPQ outperform standard K-means discretization, achieving relative improvements of 21.8 percent and 20.0 percent in WER on LibriSpeech, and 24.1 percent and 19.6 percent in CER on ML-SUPERB, respectively. Moreover, their performance is competitive with, and in some cases even surpasses, that of continuous SSL representations.


【3】 Trainable Adaptive Score Normalization for Automatic Speaker  Verification
标题: 用于自动说话人验证的可训练自适应分数标准化
链接:https://arxiv.org/abs/2504.04512

作者: Jeong-Hwan Choi,  Ju-Seok Seong,  Ye-Rin Jeoung,  Joon-Hyuk Chang 
备注:Accepted at ICASSP'25
摘要:自适应S范数(AS-norm)利用与输入说话人相似的冒名顶替者的语音分数对语音分数进行归一化,从而校正自动说话人确认(ASV)的语音分数。然而,AS-norm不涉及任何学习过程,限制了其为各种评价话语提供适当正则化强度的能力。为了解决这个限制,我们提出了一个可训练的AS-范数(TAS-范数),它利用了可学习的冒名顶替者嵌入(LIE),用于组成队列。这些LIE被初始化以表示由冒名顶替者扬声器组成的训练数据集中的每个扬声器。随后,通过模拟ASV评估来微调LIE。我们在最高得分的IE选择过程中利用边际惩罚进行微调,以防止非冒名顶替者被选中。在我们的ECAPA-TDNN实验中,与不使用建议的LIE的标准AS范数相比,在VoxCeleb 1-O试验中,建议的TAS范数在相等错误率和最小检测成本函数方面分别观察到4.11%和10.62%的相对改善。我们进一步验证了包括波斯语和汉语在内的其他ASV数据集上的TAS规范的有效性,证明了其在不同语言中的鲁棒性。
摘要:Adaptive S-norm (AS-norm) calibrates automatic speaker verification (ASV) scores by normalizing them utilize the scores of impostors which are similar to the input speaker. However, AS-norm does not involve any learning process, limiting its ability to provide appropriate regularization strength for various evaluation utterances. To address this limitation, we propose a trainable AS-norm (TAS-norm) that leverages learnable impostor embeddings (LIEs), which are used to compose the cohort. These LIEs are initialized to represent each speaker in a training dataset consisting of impostor speakers. Subsequently, LIEs are fine-tuned by simulating an ASV evaluation. We utilize a margin penalty during top-scoring IEs selection in fine-tuning to prevent non-impostor speakers from being selected. In our experiments with ECAPA-TDNN, the proposed TAS-norm observed 4.11% and 10.62% relative improvement in equal error rate and minimum detection cost function, respectively, on VoxCeleb1-O trial compared with standard AS-norm without using proposed LIEs. We further validated the effectiveness of the TAS-norm on additional ASV datasets comprising Persian and Chinese, demonstrating its robustness across different languages.


【4】 WaveNet-Volterra Neural Networks for Active Noise Control: A Fully  Causal Approach
标题: 用于主动噪音控制的WaveNet-Volterra神经网络:全因果方法
链接:https://arxiv.org/abs/2504.04450

作者: Lu Bai,  Mengtong Li,  Siyuan Lian,  Kai Chen,  Jing Lu 
摘要:有源噪声控制(ANC)系统受到非线性失真的挑战,它降低了传统自适应滤波器的性能。虽然已经出现了基于深度学习的ANC算法来解决非线性问题,但现有方法往往忽略了关键限制:(1)端到端深度神经网络(DNN)模型经常违反实时ANC应用固有的因果关系约束;(2)许多研究将基于DNN的方法与简化或低阶自适应滤波器进行比较,而不是完全优化的高阶对应方法。在这封信中,我们提出了一个保持容量的时域ANC框架,该框架将WaveNet与Volterra神经网络(VNN)协同作用,明确解决系统非线性问题,同时确保严格的因果运算。与之前基于DNN的方法不同,我们的方法以最先进的深度学习架构和严格优化的高阶自适应滤波器(包括Wiener解决方案)为基准。仿真结果表明,该框架实现了优于现有DNN方法和传统算法的性能,揭示了DNN优越性的先前声明源于与次优传统基线的不完全比较。源代码可在https://github.com/Lu-Baihh/WaveNet-VNNs-for-ANC.git上获得。
摘要:Active Noise Control (ANC) systems are challenged by nonlinear distortions, which degrade the performance of traditional adaptive filters. While deep learning-based ANC algorithms have emerged to address nonlinearity, existing approaches often overlook critical limitations: (1) end-to-end Deep Neural Network (DNN) models frequently violate causality constraints inherent to real-time ANC applications; (2) many studies compare DNN-based methods against simplified or low-order adaptive filters rather than fully optimized high-order counterparts. In this letter, we propose a causality-preserving time-domain ANC framework that synergizes WaveNet with Volterra Neural Networks (VNNs), explicitly addressing system nonlinearity while ensuring strict causal operation. Unlike prior DNN-based approaches, our method is benchmarked against both state-of-the-art deep learning architectures and rigorously optimized high-order adaptive filters, including Wiener solutions. Simulations demonstrate that the proposed framework achieves superior performance over existing DNN methods and traditional algorithms, revealing that prior claims of DNN superiority stem from incomplete comparisons with suboptimal traditional baselines. Source code is available at https://github.com/Lu-Baihh/WaveNet-VNNs-for-ANC.git.


【5】 Real-Time Auralization for First-Person Vocal Interaction in Immersive  Virtual Environments
标题: 沉浸式虚拟环境中第一人称语音交互的实时可听化
链接:https://arxiv.org/abs/2504.04075

作者: Mauricio Flores-Vargas,  Enda Bates,  Rachel McDonnell 
摘要:随着虚拟现实(VR)技术集成了不同的感官反馈,使人们能够在视听环境中再现真实空间,多模态研究和应用变得越来越普遍。在VR体验中,许多应用依赖于用户的声音作为交互的关键元素,包括音乐表演和公共演讲应用。我们对声音的自我感知在发声中起着至关重要的作用。当唱歌或说话时,我们的声音与环境的声学特性相互作用,根据空间的感知特征塑造声音参数的调整。本技术报告介绍了一种实时可听化管道,该管道利用三维空间脉冲响应(SIR)用于需要第一人称语音交互的VR中的多模态研究应用。它描述了脉冲响应创建和渲染工作流、视听集成,并解决了延迟和计算方面的考虑。该系统使用户能够从预定义区域内的各种位置和方向探索声学空间,支持视听多模态感知中的三个和五个自由度(3Dof和5DoF),用于VR中的研究和创意应用。
摘要:Multimodal research and applications are becoming more commonplace as Virtual Reality (VR) technology integrates different sensory feedback, enabling the recreation of real spaces in an audio-visual context. Within VR experiences, numerous applications rely on the user's voice as a key element of interaction, including music performances and public speaking applications. Self-perception of our voice plays a crucial role in vocal production. When singing or speaking, our voice interacts with the acoustic properties of the environment, shaping the adjustment of vocal parameters in response to the perceived characteristics of the space. This technical report presents a real-time auralization pipeline that leverages three-dimensional Spatial Impulse Responses (SIRs) for multimodal research applications in VR requiring first-person vocal interaction. It describes the impulse response creation and rendering workflow, the audio-visual integration, and addresses latency and computational considerations. The system enables users to explore acoustic spaces from various positions and orientations within a predefined area, supporting three and five Degrees of Freedom (3Dof and 5DoF) in audio-visual multimodal perception for both research and creative applications in VR.


【6】 Continuous Boostlet Transform and Associated Uncertainty Principles
标题: 连续Boostlet变换和相关的不确定性原则
链接:https://arxiv.org/abs/2504.03679

作者: Owais Ahmad,  Jasifa Fayaz 
备注:28pages,6 figures
摘要:连续Boostlet变换(CBT)是分析时空信号,特别是声波场的有力工具。CBT克服了经典小波的局限性,利用庞加莱群和各向同性膨胀来捕获自然声场的稀疏特征。本文介绍了CBT的数学框架,包括它的定义,基本性质,以及相关的不确定性原理,如海森堡的,对数,皮特的,和Nazarov的不等式。这些结果阐明了时间和频率的本地化之间的权衡在boostlet域。常数函数和指数函数的实际例子突出了CBT的适应性。CBT在雷达、通信、音频处理和地震分析中有着广泛的应用,它提供了灵活的时频分辨率,是非平稳和瞬态信号的理想选择,也是现代信号处理的宝贵工具。
摘要:The Continuous Boostlet Transform (CBT) is introduced as a powerful tool for analyzing spatiotemporal signals, particularly acoustic wavefields. Overcoming the limitations of classical wavelets, the CBT leverages the Poincar\'e group and isotropic dilations to capture sparse features of natural acoustic fields. This paper presents the mathematical framework of the CBT, including its definition, fundamental properties, and associated uncertainty principles, such as Heisenberg's, logarithmic, Pitt's, and Nazarov's inequalities. These results illuminate the trade-offs between time and frequency localization in the boostlet domain. Practical examples with constant and exponential functions highlight the CBT's adaptability. With applications in radar, communications, audio processing, and seismic analysis, the CBT offers flexible time-frequency resolution, making it ideal for non-stationary and transient signals, and a valuable tool for modern signal processing.


【7】 Leveraging Label Potential for Enhanced Multimodal Emotion Recognition
标题: 利用标签潜力增强多模式情感识别
链接:https://arxiv.org/abs/2504.05158

作者: Xuechun Shao,  Yinfeng Yu,  Liejun Wang 
备注:Main paper (8 pages). Accepted for publication by IJCNN 2025
摘要:多模态情感识别(MER)试图整合各种模态来准确地预测情感状态。然而,目前的研究大多集中在音频和文本特征的融合,忽略了情感标签中有价值的信息。这种疏忽可能会阻碍现有方法的性能,因为情感标签包含丰富的,有见地的信息,可以显着帮助MER。我们引入了一种新的模型,称为标签信号引导的多模态情感识别(LSGMER),以克服这一限制。该模型旨在充分利用情感标签信息的力量,以提高MER的分类精度和稳定性。具体来说,LSGMER采用了一个标签信号增强模块,通过标签嵌入与音频和文本特征交互来优化模态特征的表示,使其能够精确地捕捉情感的细微差别。此外,我们提出了一个联合目标优化(JOO)的方法,以提高分类精度,通过引入属性预测一致性约束(APC),加强融合功能和情感类别之间的对齐。在IEMOCAP和MELD数据集上进行的大量实验证明了我们提出的LSGMER模型的有效性。
摘要:Multimodal emotion recognition (MER) seeks to integrate various modalities to predict emotional states accurately. However, most current research focuses solely on the fusion of audio and text features, overlooking the valuable information in emotion labels. This oversight could potentially hinder the performance of existing methods, as emotion labels harbor rich, insightful information that could significantly aid MER. We introduce a novel model called Label Signal-Guided Multimodal Emotion Recognition (LSGMER) to overcome this limitation. This model aims to fully harness the power of emotion label information to boost the classification accuracy and stability of MER. Specifically, LSGMER employs a Label Signal Enhancement module that optimizes the representation of modality features by interacting with audio and text features through label embeddings, enabling it to capture the nuances of emotions precisely. Furthermore, we propose a Joint Objective Optimization(JOO) approach to enhance classification accuracy by introducing the Attribution-Prediction Consistency Constraint (APC), which strengthens the alignment between fused features and emotion categories. Extensive experiments conducted on the IEMOCAP and MELD datasets have demonstrated the effectiveness of our proposed LSGMER model.


【8】 Deconstructing Jazz Piano Style Using Machine Learning
标题: 使用机器学习解构爵士钢琴风格
链接:https://arxiv.org/abs/2504.05009

作者: Huw Cheston,  Reuben Bance,  Peter M. C. Harrison 
备注:Paper: 40 pages, 11 figures, 1 table. Supplementary material: 33 pages, 48 figures, 6 tables
摘要:艺术风格已经被研究了几个世纪,机器学习的最新进展为通过计算来理解它创造了新的可能性。然而,确保机器学习模型产生符合实践者和批评者利益的见解仍然是一个重大挑战。在这里,我们专注于音乐风格,这得益于丰富的理论和数学分析传统。我们训练了各种监督学习模型,以在精心策划的84小时录音数据集中识别20位标志性爵士音乐家,并解释他们的决策过程。我们的模型包括一个新颖的多输入架构,使四个音乐领域(旋律,和声,节奏和动态)进行单独分析。这些模型使我们能够解决音乐理论中的基本问题,并推进音乐表演者识别的最新技术(20个类别的准确率为94%)。我们发布了我们的模型的开源实现和一个用于探索音乐风格的附带Web应用程序。
摘要:Artistic style has been studied for centuries, and recent advances in machine learning create new possibilities for understanding it computationally. However, ensuring that machine-learning models produce insights aligned with the interests of practitioners and critics remains a significant challenge. Here, we focus on musical style, which benefits from a rich theoretical and mathematical analysis tradition. We train a variety of supervised-learning models to identify 20 iconic jazz musicians across a carefully curated dataset of 84 hours of recordings, and interpret their decision-making processes. Our models include a novel multi-input architecture that enables four musical domains (melody, harmony, rhythm, and dynamics) to be analysed separately. These models enable us to address fundamental questions in music theory and also advance the state-of-the-art in music performer identification (94% accuracy across 20 classes). We release open-source implementations of our models and an accompanying web application for exploring musical styles.


【9】 Diff-SSL-G-Comp: Towards a Large-Scale and Diverse Dataset for Virtual  Analog Modeling
标题: 迪夫-SSL-G-Comp:迈向大规模、多样化的虚拟模拟建模数据集
链接:https://arxiv.org/abs/2504.04589

作者: Yicheng Gu,  Runsong Zhang,  Lauri Juvela,  Zhizheng Wu 
备注:Submitted to DAFx 2025
摘要:虚拟模拟(VA)建模旨在通过算法模拟硬件电路的行为,以数字方式复制其音调。动态范围压缩器(DRC)是一个音频处理模块,通过减少和放大响亮和安静声音的音量来控制音轨的动态,这在音乐制作中至关重要。近年来,基于神经网络的VA建模在产生高保真模型方面显示出巨大的潜力。但由于数据量和多样性的不足,它们在不同参数设置和输入声音下的泛化能力仍然有限。为了解决这个问题,我们提出了Diff-SSL-G-Comp,这是第一个用于建模SSL 500 G-Bus Compressor的大规模和多样化数据集。具体来说,我们从剑桥多轨图书馆手动收集了175首未掌握的歌曲。我们记录了220个参数组合的压缩音频,产生了一个广泛的2528小时的数据集,包括不同的流派,乐器,节奏和键。此外,为了便于使用我们提出的数据集,我们在各种开源的黑盒和灰盒模型以及白盒插件中进行了基准实验。我们还在不同的数据子集中进行了消融研究,以说明提高数据多样性和数量的有效性。数据集和演示在我们的项目页面上:http://www.yichenggu.com/DiffSSLGComp/。
摘要:Virtual Analog (VA) modeling aims to simulate the behavior of hardware circuits via algorithms to replicate their tone digitally. Dynamic Range Compressor (DRC) is an audio processing module that controls the dynamics of a track by reducing and amplifying the volumes of loud and quiet sounds, which is essential in music production. In recent years, neural-network-based VA modeling has shown great potential in producing high-fidelity models. However, due to the lack of data quantity and diversity, their generalization ability in different parameter settings and input sounds is still limited. To tackle this problem, we present Diff-SSL-G-Comp, the first large-scale and diverse dataset for modeling the SSL 500 G-Bus Compressor. Specifically, we manually collected 175 unmastered songs from the Cambridge Multitrack Library. We recorded the compressed audio in 220 parameter combinations, resulting in an extensive 2528-hour dataset with diverse genres, instruments, tempos, and keys. Moreover, to facilitate the use of our proposed dataset, we conducted benchmark experiments in various open-sourced black-box and grey-box models, as well as white-box plugins. We also conducted ablation studies in different data subsets to illustrate the effectiveness of improved data diversity and quantity. The dataset and demos are on our project page: http://www.yichenggu.com/DiffSSLGComp/.


【10】 VocalNet: Speech LLM with Multi-Token Prediction for Faster and  High-Quality Generation
标题: VocalNet:具有多令牌预测的语音LLM,用于更快、更高质量的生成
链接:https://arxiv.org/abs/2504.04060

作者: Yuhao Wang,  Heyang Liu,  Ziyang Cheng,  Ronghua Wu,  Qunshan Gu,  Yanfeng Wang,  Yu Wang 
摘要:语音大语言模型(LLM)已经成为语音处理领域的一个重要研究热点。我们提出了VocalNet-1B和VocalNet-8B,这是一系列高性能,低延迟的语音LLM,由可扩展和模型无关的实时语音交互训练框架实现。从传统的下一个令牌预测(NTP)出发,我们引入了多令牌预测(MTP),一种新的方法优化语音LLM,同时提高生成速度和质量。实验表明,VocalNet的性能优于主流的Omni LLM,尽管它使用的训练数据要少得多,同时也大大超过了现有的开源语音LLM。为了支持可重复性和社区发展,我们将在发布时开源所有模型权重,推理代码,训练数据和框架实现。
摘要:Speech large language models (LLMs) have emerged as a prominent research focus in speech processing. We propose VocalNet-1B and VocalNet-8B, a series of high-performance, low-latency speech LLMs enabled by a scalable and model-agnostic training framework for real-time voice interaction. Departing from the conventional next-token prediction (NTP), we introduce multi-token prediction (MTP), a novel approach optimized for speech LLMs that simultaneously improves generation speed and quality. Experiments show that VocalNet outperforms mainstream Omni LLMs despite using significantly less training data, while also surpassing existing open-source speech LLMs by a substantial margin. To support reproducibility and community advancement, we will open-source all model weights, inference code, training data, and framework implementations upon publication.


【11】 Determined blind source separation via modeling adjacent frequency band  correlations in speech signals
标题: 通过建模语音信号中的相邻频段相关性确定盲源分离
链接:https://arxiv.org/abs/2504.03998

作者: Jianyu Wang,  Shanzheng Guan,  Zhengqiao Zhao,  Nicolas Dobigeon,  Jingdong Chen 
摘要:多通道盲源分离(MBSS)是从混合观测信号中分离出感兴趣信号的一种方法,在声学和语音处理领域得到了广泛的研究。现有的MBSS算法,如独立低秩矩阵分析(ILRMA)和多通道非负矩阵分解(MNMF),利用源模型的低秩结构,但假设频率点是独立的。相反,独立向量分析(IVA)不依赖于低秩源模型,而是基于均匀相关假设捕获频率依赖性。在这项工作中,我们证明了相邻的频率箱之间的依赖性显着强于那些在典型的语音信号中相距较远的箱之间。为了解决这个问题,我们引入了一个加权的基于Sinkhorn发散的ILRMA(wsILRMA),同时捕获这些频率间的依赖关系和模型的联合概率分布。我们的方法采用了频率间的相关性约束,从而提高了源分离性能相比,现有的方法,更高的信号失真比(SDR)和源干扰比(SIR)证明。
摘要:Multichannel blind source separation (MBSS), which focuses on separating signals of interest from mixed observations, has been extensively studied in acoustic and speech processing. Existing MBSS algorithms, such as independent low-rank matrix analysis (ILRMA) and multichannel nonnegative matrix factorization (MNMF), utilize the low-rank structure of source models but assume that frequency bins are independent. In contrast, independent vector analysis (IVA) does not rely on a low-rank source model but rather captures frequency dependencies based on a uniform correlation assumption. In this work, we demonstrate that dependencies between adjacent frequency bins are significantly stronger than those between bins that are farther apart in typical speech signals. To address this, we introduce a weighted Sinkhorn divergence-based ILRMA (wsILRMA) that simultaneously captures these inter-frequency dependencies and models joint probability distributions. Our approach incorporates an inter-frequency correlation constraint, leading to improved source separation performance compared to existing methods, as evidenced by higher Signal-to-Distortion Ratios (SDRs) and Source-to-Interference Ratios (SIRs).


机器翻译由腾讯交互翻译提供,仅供参考