微信公众号:arXiv_Daily
cs.SD语音
标题: UVAR:不确定性加权弱监督视听视频解析
链接:https://arxiv.org/abs/2505.09615
备注:CVPR 2025
摘要:视听视频解析(AVVP)需要定位两个单模态事件(即,那些仅以视频的视觉或声学模态出现的事件)和多模态事件(即,在两种模式中同时发生的那些)。此外,使用所有这些事件的类标签以及它们的开始和结束时间来注释训练数据的高昂成本对AVVP技术的可扩展性施加了限制,除非它们可以在弱监督设置中进行训练,其中只有模态不可知的视频级标签在训练数据中可用。为此,最近提出的方法试图生成段级伪标签,以更好地指导模型训练。然而,在生成这些伪标签时缺乏段间依赖性以及对预测段中不存在的标签的一般偏见限制了它们的性能。这项工作提出了一种新的方法来克服这些弱点称为不确定加权弱监督视听视频解析(UWAV)。此外,我们的创新方法考虑了与这些估计的伪标签相关的不确定性,并结合了基于特征混合的训练正则化以改进训练。实证结果表明,UWAV在两个不同的数据集上,在多个指标上优于AVVP任务的最新方法,证明了其有效性和通用性。
摘要:Audio-Visual Video Parsing (AVVP) entails the challenging task of localizing both uni-modal events (i.e., those occurring exclusively in either the visual or acoustic modality of a video) and multi-modal events (i.e., those occurring in both modalities concurrently). Moreover, the prohibitive cost of annotating training data with the class labels of all these events, along with their start and end times, imposes constraints on the scalability of AVVP techniques unless they can be trained in a weakly-supervised setting, where only modality-agnostic, video-level labels are available in the training data. To this end, recently proposed approaches seek to generate segment-level pseudo-labels to better guide model training. However, the absence of inter-segment dependencies when generating these pseudo-labels and the general bias towards predicting labels that are absent in a segment limit their performance. This work proposes a novel approach towards overcoming these weaknesses called Uncertainty-weighted Weakly-supervised Audio-visual Video Parsing (UWAV). Additionally, our innovative approach factors in the uncertainty associated with these estimated pseudo-labels and incorporates a feature mixup based training regularization for improved training. Empirical results show that UWAV outperforms state-of-the-art methods for the AVVP task on multiple metrics, across two different datasets, attesting to its effectiveness and generalizability.
【2】 The Voice Timbre Attribute Detection 2025 Challenge Evaluation Plan
标题: 语音音色属性检测2025年挑战评估计划链接:https://arxiv.org/abs/2505.09382
摘要:声音音色是指一个人的声音的独特品质或特征,使其与人类听觉感知的其他声音区分开来。语音音色属性检测(VtaD)2025挑战赛的重点是以比较的方式解释语音音色属性。在这个挑战中,人类对声音音色的印象被用一组感官描述符来描述,包括明亮、粗糙、柔软、磁性等。音色是通过两个声音在特定描述符维度内的强度比较来解释的。VtaD 2025挑战赛将于5月开始,最终将于2025年10月在中国镇江举行的NCMMSC 2025会议上提出一项特别提案。
摘要:Voice timbre refers to the unique quality or character of a person's voice that distinguishes it from others as perceived by human hearing. The Voice Timbre Attribute Detection (VtaD) 2025 challenge focuses on explaining the voice timbre attribute in a comparative manner. In this challenge, the human impression of voice timbre is verbalized with a set of sensory descriptors, including bright, coarse, soft, magnetic, and so on. The timbre is explained from the comparison between two voices in their intensity within a specific descriptor dimension. The VtaD 2025 challenge starts in May and culminates in a special proposal at the NCMMSC2025 conference in October 2025 in Zhenjiang, China.
【3】 SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset
标题: SingNet:迈向大规模、多样化和野外歌唱声音数据集链接:https://arxiv.org/abs/2505.09325
摘要:长期以来,缺乏公开的大规模和多样化的数据集一直是歌唱语音应用程序(如歌唱语音合成(SVS)和歌唱语音转换(SVC))的重要瓶颈。为了解决这个问题,我们提出了SingNet,这是一个广泛的,多样化的,野外的歌唱声音数据集。具体来说,我们提出了一个数据处理管道,从互联网上的样本包和歌曲中提取现成的训练数据,形成3000小时的各种语言和风格的歌声。此外,为了方便使用和展示SingNet的有效性,我们根据收集到的歌声数据,在Wav 2 vec 2、BigVGAN和NSF-HiFiGAN上预训练和开源了各种最先进的(SOTA)模型。我们还对自动歌词转录(ALT),神经声码器和歌唱语音转换(SVC)进行了基准实验。音频演示可在https://singnet-dataset.github.io/上获得。
摘要:The lack of a publicly-available large-scale and diverse dataset has long been a significant bottleneck for singing voice applications like Singing Voice Synthesis (SVS) and Singing Voice Conversion (SVC). To tackle this problem, we present SingNet, an extensive, diverse, and in-the-wild singing voice dataset. Specifically, we propose a data processing pipeline to extract ready-to-use training data from sample packs and songs on the internet, forming 3000 hours of singing voices in various languages and styles. Furthermore, to facilitate the use and demonstrate the effectiveness of SingNet, we pre-train and open-source various state-of-the-art (SOTA) models on Wav2vec2, BigVGAN, and NSF-HiFiGAN based on our collected singing voice data. We also conduct benchmark experiments on Automatic Lyric Transcription (ALT), Neural Vocoder, and Singing Voice Conversion (SVC). Audio demos are available at: https://singnet-dataset.github.io/.
【4】 Adaptive Noise Resilient Keyword Spotting Using One-Shot Learning
标题: 使用一次学习的自适应抗噪关键词发现链接:https://arxiv.org/abs/2505.09304
备注:Preprint submitted to the IEEE 11th World Forum on Internet of Things
摘要:关键词定位(KWS)是智能设备的关键组件,可实现高效直观的音频交互。然而,在嵌入式设备上部署的标准KWS系统在现实世界的操作条件下经常遭受性能下降。弹性KWS系统通过启用动态适应来解决这个问题,例如添加或替换关键字,调整特定用户以及提高噪声鲁棒性。然而,由于内存和计算资源有限,在资源受限的设备上部署具有低延迟的弹性独立KWS系统仍然具有挑战性。这项研究提出了一种低计算方法,用于KWS分类的预训练神经网络的连续噪声适应,只需要1次学习和一个时期。该方法使用两个预训练模型和三个真实世界的噪声源进行评估,信噪比(SNR)范围从24到-3 dB。在所有情况下,自适应模型的性能始终优于预训练模型,特别是在SNR $\leq $18 dB时,准确度提高了4.9%至46.0%。这些结果突出了所提出的方法的有效性,同时足够轻巧,可用于资源受限的设备上的部署。
摘要:Keyword spotting (KWS) is a key component of smart devices, enabling efficient and intuitive audio interaction. However, standard KWS systems deployed on embedded devices often suffer performance degradation under real-world operating conditions. Resilient KWS systems address this issue by enabling dynamic adaptation, with applications such as adding or replacing keywords, adjusting to specific users, and improving noise robustness. However, deploying resilient, standalone KWS systems with low latency on resource-constrained devices remains challenging due to limited memory and computational resources. This study proposes a low computational approach for continuous noise adaptation of pretrained neural networks used for KWS classification, requiring only 1-shot learning and one epoch. The proposed method was assessed using two pretrained models and three real-world noise sources at signal-to-noise ratios (SNRs) ranging from 24 to -3 dB. The adapted models consistently outperformed the pretrained models across all scenarios, especially at SNR $\leq$ 18 dB, achieving accuracy improvements of 4.9% to 46.0%. These results highlight the efficacy of the proposed methodology while being lightweight enough for deployment on resource-constrained devices.
【5】 DPN-GAN: Inducing Periodic Activations in Generative Adversarial Networks for High-Fidelity Audio Synthesis
标题: DPFN-GAN:在生成对抗网络中诱导周期性激活以实现高保真音频合成链接:https://arxiv.org/abs/2505.09091
备注:None
摘要:近年来,生成对抗网络(GANs)在生成音频序列方面取得了重大进展。然而,这些模型通常依赖于带宽受限的梅尔频谱图,这限制了所生成的音频序列的分辨率,并导致条件生成期间的模式崩溃。为了解决这个问题,我们提出了基于可变形周期性网络的GAN(DPN-GAN),这是一种新型的GAN架构,它采用了基于内核的周期性ReLU激活函数来诱导音频生成中的周期性偏差。这种创新的方法增强了模型捕捉和再现复杂音频模式的能力。特别是,我们提出的模型具有一个DPN模块,用于利用可变形卷积操作的多分辨率生成,允许自适应感受野,提高合成音频的质量和保真度。此外,我们使用可变形卷积来增强网络,以更好地区分真实和生成的样本,进一步改善音频质量。我们训练了两个版本的模型:DPN-GAN小(38.67 M参数)和DPN-GAN大(124 M参数)。为了进行评估,我们使用了五个不同的数据集,涵盖了语音合成和音乐生成任务,以证明DPN-GAN的效率。实验结果表明,DPN-GAN在分布外和噪声数据上都具有优异的性能,展示了其鲁棒性和适应性。经过各种数据集的训练,DPN-GAN在标准评估指标上优于最先进的GAN架构,并在合成音频中表现出更高的鲁棒性。
摘要:In recent years, generative adversarial networks (GANs) have made significant progress in generating audio sequences. However, these models typically rely on bandwidth-limited mel-spectrograms, which constrain the resolution of generated audio sequences, and lead to mode collapse during conditional generation. To address this issue, we propose Deformable Periodic Network based GAN (DPN-GAN), a novel GAN architecture that incorporates a kernel-based periodic ReLU activation function to induce periodic bias in audio generation. This innovative approach enhances the model's ability to capture and reproduce intricate audio patterns. In particular, our proposed model features a DPN module for multi-resolution generation utilizing deformable convolution operations, allowing for adaptive receptive fields that improve the quality and fidelity of the synthetic audio. Additionally, we enhance the discriminator network using deformable convolution to better distinguish between real and generated samples, further refining the audio quality. We trained two versions of the model: DPN-GAN small (38.67M parameters) and DPN-GAN large (124M parameters). For evaluation, we use five different datasets, covering both speech synthesis and music generation tasks, to demonstrate the efficiency of the DPN-GAN. The experimental results demonstrate that DPN-GAN delivers superior performance on both out-of-distribution and noisy data, showcasing its robustness and adaptability. Trained across various datasets, DPN-GAN outperforms state-of-the-art GAN architectures on standard evaluation metrics, and exhibits increased robustness in synthesized audio.
【6】 Inference Attacks for X-Vector Speaker Anonymization
标题: X向量说话人自动化中的推理攻击链接:https://arxiv.org/abs/2505.08978
摘要:我们重新审视x向量说话人匿名化的隐私效用权衡。现有的方法通过训练复杂的说话人验证或识别模型来量化隐私,这些模型后来被用作攻击。相反,我们提出了一种新的推理攻击去匿名化。我们的攻击是简单的,ML免费的,但我们的实验表明,它优于现有的方法。
摘要:We revisit the privacy-utility tradeoff of x-vector speaker anonymization. Existing approaches quantify privacy through training complex speaker verification or identification models that are later used as attacks. Instead, we propose a novel inference attack for de-anonymization. Our attack is simple and ML-free yet we show experimentally that it outperforms existing approaches.
【7】 WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
标题: WavReward:与多面手奖励评估者的口头对话模型链接:https://arxiv.org/abs/2505.09558
摘要:诸如GPT-4 o-audio之类的端到端口语对话模型最近在语音领域获得了极大的关注。然而,口语对话模型的会话性能的评估在很大程度上被忽视了。这主要是由于智能聊天机器人传达了大量的非文本信息,这些信息无法使用ChatGPT等基于文本的语言模型轻松测量。为了解决这个问题,我们提出了WavReward,一个基于音频语言模型的奖励反馈模型,可以评估语音输入的口语对话系统的IQ和EQ。具体而言,1)基于音频语言模型,WavReward结合了深度推理过程和非线性奖励机制进行后期训练。通过利用多样本反馈通过强化学习算法,我们构建了一个专门的评估器定制的口语对话模型。2)我们引入ChatReward-30 K,这是一个用于训练WavReward的偏好数据集。ChatReward-30 K包括口语对话模型的理解和生成方面。这些场景跨越各种任务,例如基于文本的聊天、指令聊天的九个声学属性以及隐式聊天。WavReward在多个口语对话场景中的表现优于之前最先进的评估模型,实现了Qwen2.5-Omni在客观准确性上的大幅提升,从55.1$\%$到91.5$\%$。在主观A/B测试中,WavReward也领先83$\%$。全面的消融研究证实了WavReward每个组件的必要性。所有数据和代码将在论文被接受后在https://github.com/jishengpeng/WavReward上公开。
摘要:End-to-end spoken dialogue models such as GPT-4o-audio have recently garnered significant attention in the speech domain. However, the evaluation of spoken dialogue models' conversational performance has largely been overlooked. This is primarily due to the intelligent chatbots convey a wealth of non-textual information which cannot be easily measured using text-based language models like ChatGPT. To address this gap, we propose WavReward, a reward feedback model based on audio language models that can evaluate both the IQ and EQ of spoken dialogue systems with speech input. Specifically, 1) based on audio language models, WavReward incorporates the deep reasoning process and the nonlinear reward mechanism for post-training. By utilizing multi-sample feedback via the reinforcement learning algorithm, we construct a specialized evaluator tailored to spoken dialogue models. 2) We introduce ChatReward-30K, a preference dataset used to train WavReward. ChatReward-30K includes both comprehension and generation aspects of spoken dialogue models. These scenarios span various tasks, such as text-based chats, nine acoustic attributes of instruction chats, and implicit chats. WavReward outperforms previous state-of-the-art evaluation models across multiple spoken dialogue scenarios, achieving a substantial improvement about Qwen2.5-Omni in objective accuracy from 55.1$\%$ to 91.5$\%$. In subjective A/B testing, WavReward also leads by a margin of 83$\%$. Comprehensive ablation studies confirm the necessity of each component of WavReward. All data and code will be publicly at https://github.com/jishengpeng/WavReward after the paper is accepted.
【8】 Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?
标题: Omni-R1:您真的需要音频来微调音频LLM吗?链接:https://arxiv.org/abs/2505.09439
摘要:我们提出了Omni-R1,它使用强化学习方法GRPO在音频问答数据集上微调了最近的多模态LLM Qwen2.5-Omni。这导致新的国家的最先进的性能在最近的MMAU基准。Omni-R1在声音、音乐、语音和总体平均类别(无论是Test-mini还是Test-full分割)上都实现了最高的准确度。为了理解性能的提高,我们测试了有音频和没有音频的模型,发现GRPO的大部分性能提高可以归因于更好的基于文本的推理。我们还发现了一个令人惊讶的发现,在纯文本数据集上进行无音频微调可以有效地提高基于音频的性能。
摘要:We propose Omni-R1 which fine-tunes a recent multi-modal LLM, Qwen2.5-Omni, on an audio question answering dataset with the reinforcement learning method GRPO. This leads to new State-of-the-Art performance on the recent MMAU benchmark. Omni-R1 achieves the highest accuracies on the sounds, music, speech, and overall average categories, both on the Test-mini and Test-full splits. To understand the performance improvement, we tested models both with and without audio and found that much of the performance improvement from GRPO could be attributed to better text-based reasoning. We also made a surprising discovery that fine-tuning without audio on a text-only dataset was effective at improving the audio-based performance.
标题: WavReward:与多面手奖励评估者的口头对话模型
链接:https://arxiv.org/abs/2505.09558
摘要:诸如GPT-4o-audio之类的端到端口语对话模型最近在语音领域获得了极大的关注。然而,口语对话模型的会话性能的评估在很大程度上被忽视了。这主要是由于智能聊天机器人传达了大量的非文本信息,这些信息无法使用ChatGPT等基于文本的语言模型轻松测量。为了解决这个问题,我们提出了WavReward,一个基于音频语言模型的奖励反馈模型,可以评估语音输入的口语对话系统的IQ和EQ。具体而言,1)基于音频语言模型,WavReward结合了深度推理过程和非线性奖励机制进行后期训练。通过利用多样本反馈通过强化学习算法,我们构建了一个专门的评估器定制的口语对话模型。2)我们引入ChatReward-30K,这是一个用于训练WavReward的偏好数据集。ChatReward-30K包括口语对话模型的理解和生成方面。这些场景跨越各种任务,例如基于文本的聊天、指令聊天的九个声学属性以及隐式聊天。WavReward在多个口语对话场景中的表现优于之前最先进的评估模型,实现了Qwen2.5-Omni在客观准确性上的大幅提升,从55.1 $\ %$到91.5 $\ %$。在主观A/B测试中,WavReward也领先83 $\ %$。全面的消融研究证实了WavReward每个组件的必要性。所有数据和代码将在论文被接受后在www.example.com上公开。
摘要:End-to-end spoken dialogue models such as GPT-4o-audio have recently garnered significant attention in the speech domain. However, the evaluation of spoken dialogue models' conversational performance has largely been overlooked. This is primarily due to the intelligent chatbots convey a wealth of non-textual information which cannot be easily measured using text-based language models like ChatGPT. To address this gap, we propose WavReward, a reward feedback model based on audio language models that can evaluate both the IQ and EQ of spoken dialogue systems with speech input. Specifically, 1) based on audio language models, WavReward incorporates the deep reasoning process and the nonlinear reward mechanism for post-training. By utilizing multi-sample feedback via the reinforcement learning algorithm, we construct a specialized evaluator tailored to spoken dialogue models. 2) We introduce ChatReward-30K, a preference dataset used to train WavReward. ChatReward-30K includes both comprehension and generation aspects of spoken dialogue models. These scenarios span various tasks, such as text-based chats, nine acoustic attributes of instruction chats, and implicit chats. WavReward outperforms previous state-of-the-art evaluation models across multiple spoken dialogue scenarios, achieving a substantial improvement about Qwen2.5-Omni in objective accuracy from 55.1$\%$ to 91.5$\%$. In subjective A/B testing, WavReward also leads by a margin of 83$\%$. Comprehensive ablation studies confirm the necessity of each component of WavReward. All data and code will be publicly at https://github.com/jishengpeng/WavReward after the paper is accepted.
【2】 Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM?
标题: Omni-R1:您真的需要音频来微调音频LLM吗?链接:https://arxiv.org/abs/2505.09439
摘要:我们提出了Omni-R1,它使用强化学习方法GRPO在音频问答数据集上微调了最近的多模态LLM Qwen2.5-Omni。这导致新的国家的最先进的性能在最近的MMAU基准。Omni-R1在声音、音乐、语音和总体平均类别(无论是Test-mini还是Test-full分割)上都实现了最高的准确度。为了理解性能的提高,我们测试了有音频和没有音频的模型,发现GRPO的大部分性能提高可以归因于更好的基于文本的推理。我们还发现了一个令人惊讶的发现,在纯文本数据集上进行无音频微调可以有效地提高基于音频的性能。
摘要:We propose Omni-R1 which fine-tunes a recent multi-modal LLM, Qwen2.5-Omni, on an audio question answering dataset with the reinforcement learning method GRPO. This leads to new State-of-the-Art performance on the recent MMAU benchmark. Omni-R1 achieves the highest accuracies on the sounds, music, speech, and overall average categories, both on the Test-mini and Test-full splits. To understand the performance improvement, we tested models both with and without audio and found that much of the performance improvement from GRPO could be attributed to better text-based reasoning. We also made a surprising discovery that fine-tuning without audio on a text-only dataset was effective at improving the audio-based performance.
【3】 UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing
标题: UVAR:不确定性加权弱监督视听视频解析链接:https://arxiv.org/abs/2505.09615
备注:CVPR 2025
摘要:视听视频解析(AVVP)需要定位两个单模态事件(即,那些仅以视频的视觉或声学模态出现的事件)和多模态事件(即,在两种模式中同时发生的那些)。此外,使用所有这些事件的类标签以及它们的开始和结束时间来注释训练数据的高昂成本对AVVP技术的可扩展性施加了限制,除非它们可以在弱监督设置中进行训练,其中只有模态不可知的视频级标签在训练数据中可用。为此,最近提出的方法试图生成段级伪标签,以更好地指导模型训练。然而,在生成这些伪标签时缺乏段间依赖性以及对预测段中不存在的标签的一般偏见限制了它们的性能。这项工作提出了一种新的方法来克服这些弱点称为不确定加权弱监督视听视频解析(UWAV)。此外,我们的创新方法考虑了与这些估计的伪标签相关的不确定性,并结合了基于特征混合的训练正则化以改进训练。实证结果表明,UWAV在两个不同的数据集上,在多个指标上优于AVVP任务的最新方法,证明了其有效性和通用性。
摘要:Audio-Visual Video Parsing (AVVP) entails the challenging task of localizing both uni-modal events (i.e., those occurring exclusively in either the visual or acoustic modality of a video) and multi-modal events (i.e., those occurring in both modalities concurrently). Moreover, the prohibitive cost of annotating training data with the class labels of all these events, along with their start and end times, imposes constraints on the scalability of AVVP techniques unless they can be trained in a weakly-supervised setting, where only modality-agnostic, video-level labels are available in the training data. To this end, recently proposed approaches seek to generate segment-level pseudo-labels to better guide model training. However, the absence of inter-segment dependencies when generating these pseudo-labels and the general bias towards predicting labels that are absent in a segment limit their performance. This work proposes a novel approach towards overcoming these weaknesses called Uncertainty-weighted Weakly-supervised Audio-visual Video Parsing (UWAV). Additionally, our innovative approach factors in the uncertainty associated with these estimated pseudo-labels and incorporates a feature mixup based training regularization for improved training. Empirical results show that UWAV outperforms state-of-the-art methods for the AVVP task on multiple metrics, across two different datasets, attesting to its effectiveness and generalizability.
【4】 The Voice Timbre Attribute Detection 2025 Challenge Evaluation Plan
标题: 语音音色属性检测2025年挑战评估计划链接:https://arxiv.org/abs/2505.09382
摘要:声音音色是指一个人的声音的独特品质或特征,使其与人类听觉感知的其他声音区分开来。语音音色属性检测(VtaD)2025挑战赛的重点是以比较的方式解释语音音色属性。在这个挑战中,人类对声音音色的印象被用一组感官描述符来描述,包括明亮、粗糙、柔软、磁性等。音色是通过两个声音在特定描述符维度内的强度比较来解释的。VtaD 2025挑战赛将于5月开始,最终将于2025年10月在中国镇江举行的NCMMSC 2025会议上提出一项特别提案。
摘要:Voice timbre refers to the unique quality or character of a person's voice that distinguishes it from others as perceived by human hearing. The Voice Timbre Attribute Detection (VtaD) 2025 challenge focuses on explaining the voice timbre attribute in a comparative manner. In this challenge, the human impression of voice timbre is verbalized with a set of sensory descriptors, including bright, coarse, soft, magnetic, and so on. The timbre is explained from the comparison between two voices in their intensity within a specific descriptor dimension. The VtaD 2025 challenge starts in May and culminates in a special proposal at the NCMMSC2025 conference in October 2025 in Zhenjiang, China.
【5】 SingNet: Towards a Large-Scale, Diverse, and In-the-Wild Singing Voice Dataset
标题: SingNet:迈向大规模、多样化和野外歌唱声音数据集链接:https://arxiv.org/abs/2505.09325
摘要:长期以来,缺乏公开的大规模和多样化的数据集一直是歌唱语音应用程序(如歌唱语音合成(SVS)和歌唱语音转换(SVC))的重要瓶颈。为了解决这个问题,我们提出了SingNet,这是一个广泛的,多样化的,野外的歌唱声音数据集。具体来说,我们提出了一个数据处理管道,从互联网上的样本包和歌曲中提取现成的训练数据,形成3000小时的各种语言和风格的歌声。此外,为了方便使用和展示SingNet的有效性,我们根据收集到的歌声数据,在Wav 2 vec 2、BigVGAN和NSF-HiFiGAN上预训练和开源了各种最先进的(SOTA)模型。我们还对自动歌词转录(ALT),神经声码器和歌唱语音转换(SVC)进行了基准实验。音频演示可在https://singnet-dataset.github.io/上获得。
摘要:The lack of a publicly-available large-scale and diverse dataset has long been a significant bottleneck for singing voice applications like Singing Voice Synthesis (SVS) and Singing Voice Conversion (SVC). To tackle this problem, we present SingNet, an extensive, diverse, and in-the-wild singing voice dataset. Specifically, we propose a data processing pipeline to extract ready-to-use training data from sample packs and songs on the internet, forming 3000 hours of singing voices in various languages and styles. Furthermore, to facilitate the use and demonstrate the effectiveness of SingNet, we pre-train and open-source various state-of-the-art (SOTA) models on Wav2vec2, BigVGAN, and NSF-HiFiGAN based on our collected singing voice data. We also conduct benchmark experiments on Automatic Lyric Transcription (ALT), Neural Vocoder, and Singing Voice Conversion (SVC). Audio demos are available at: https://singnet-dataset.github.io/.
【6】 Adaptive Noise Resilient Keyword Spotting Using One-Shot Learning
标题: 使用一次学习的自适应抗噪关键词发现链接:https://arxiv.org/abs/2505.09304
备注:Preprint submitted to the IEEE 11th World Forum on Internet of Things
摘要:关键词定位(KWS)是智能设备的关键组件,可实现高效直观的音频交互。然而,在嵌入式设备上部署的标准KWS系统在现实世界的操作条件下经常遭受性能下降。弹性KWS系统通过启用动态适应来解决这个问题,例如添加或替换关键字,调整特定用户以及提高噪声鲁棒性。然而,由于内存和计算资源有限,在资源受限的设备上部署具有低延迟的弹性独立KWS系统仍然具有挑战性。这项研究提出了一种低计算方法,用于KWS分类的预训练神经网络的连续噪声适应,只需要1次学习和一个时期。该方法使用两个预训练模型和三个真实世界的噪声源进行评估,信噪比(SNR)范围从24到-3 dB。在所有情况下,自适应模型的性能始终优于预训练模型,特别是在SNR $\leq $18 dB时,准确度提高了4.9%至46.0%。这些结果突出了所提出的方法的有效性,同时足够轻巧,可用于资源受限的设备上的部署。
摘要:Keyword spotting (KWS) is a key component of smart devices, enabling efficient and intuitive audio interaction. However, standard KWS systems deployed on embedded devices often suffer performance degradation under real-world operating conditions. Resilient KWS systems address this issue by enabling dynamic adaptation, with applications such as adding or replacing keywords, adjusting to specific users, and improving noise robustness. However, deploying resilient, standalone KWS systems with low latency on resource-constrained devices remains challenging due to limited memory and computational resources. This study proposes a low computational approach for continuous noise adaptation of pretrained neural networks used for KWS classification, requiring only 1-shot learning and one epoch. The proposed method was assessed using two pretrained models and three real-world noise sources at signal-to-noise ratios (SNRs) ranging from 24 to -3 dB. The adapted models consistently outperformed the pretrained models across all scenarios, especially at SNR $\leq$ 18 dB, achieving accuracy improvements of 4.9% to 46.0%. These results highlight the efficacy of the proposed methodology while being lightweight enough for deployment on resource-constrained devices.
【7】 DPN-GAN: Inducing Periodic Activations in Generative Adversarial Networks for High-Fidelity Audio Synthesis
标题: DPFN-GAN:在生成对抗网络中诱导周期性激活以实现高保真音频合成链接:https://arxiv.org/abs/2505.09091
备注:None
摘要:近年来,生成对抗网络(GANs)在生成音频序列方面取得了重大进展。然而,这些模型通常依赖于带宽受限的梅尔频谱图,这限制了所生成的音频序列的分辨率,并导致条件生成期间的模式崩溃。为了解决这个问题,我们提出了基于可变形周期性网络的GAN(DPN-GAN),这是一种新型的GAN架构,它采用了基于内核的周期性ReLU激活函数来诱导音频生成中的周期性偏差。这种创新的方法增强了模型捕捉和再现复杂音频模式的能力。特别是,我们提出的模型具有一个DPN模块,用于利用可变形卷积操作的多分辨率生成,允许自适应感受野,提高合成音频的质量和保真度。此外,我们使用可变形卷积来增强网络,以更好地区分真实和生成的样本,进一步改善音频质量。我们训练了两个版本的模型:DPN-GAN小(38.67 M参数)和DPN-GAN大(124 M参数)。为了进行评估,我们使用了五个不同的数据集,涵盖了语音合成和音乐生成任务,以证明DPN-GAN的效率。实验结果表明,DPN-GAN在分布外和噪声数据上都具有优异的性能,展示了其鲁棒性和适应性。经过各种数据集的训练,DPN-GAN在标准评估指标上优于最先进的GAN架构,并在合成音频中表现出更高的鲁棒性。
摘要:In recent years, generative adversarial networks (GANs) have made significant progress in generating audio sequences. However, these models typically rely on bandwidth-limited mel-spectrograms, which constrain the resolution of generated audio sequences, and lead to mode collapse during conditional generation. To address this issue, we propose Deformable Periodic Network based GAN (DPN-GAN), a novel GAN architecture that incorporates a kernel-based periodic ReLU activation function to induce periodic bias in audio generation. This innovative approach enhances the model's ability to capture and reproduce intricate audio patterns. In particular, our proposed model features a DPN module for multi-resolution generation utilizing deformable convolution operations, allowing for adaptive receptive fields that improve the quality and fidelity of the synthetic audio. Additionally, we enhance the discriminator network using deformable convolution to better distinguish between real and generated samples, further refining the audio quality. We trained two versions of the model: DPN-GAN small (38.67M parameters) and DPN-GAN large (124M parameters). For evaluation, we use five different datasets, covering both speech synthesis and music generation tasks, to demonstrate the efficiency of the DPN-GAN. The experimental results demonstrate that DPN-GAN delivers superior performance on both out-of-distribution and noisy data, showcasing its robustness and adaptability. Trained across various datasets, DPN-GAN outperforms state-of-the-art GAN architectures on standard evaluation metrics, and exhibits increased robustness in synthesized audio.
【8】 Inference Attacks for X-Vector Speaker Anonymization
标题: X向量说话人自动化中的推理攻击链接:https://arxiv.org/abs/2505.08978
摘要:我们重新审视x向量说话人匿名化的隐私效用权衡。现有的方法通过训练复杂的说话人验证或识别模型来量化隐私,这些模型后来被用作攻击。相反,我们提出了一种新的推理攻击去匿名化。我们的攻击是简单的,ML免费的,但我们的实验表明,它优于现有的方法。
摘要:We revisit the privacy-utility tradeoff of x-vector speaker anonymization. Existing approaches quantify privacy through training complex speaker verification or identification models that are later used as attacks. Instead, we propose a novel inference attack for de-anonymization. Our attack is simple and ML-free yet we show experimentally that it outperforms existing approaches.
【9】 SaFARi: State-Space Models for Frame-Agnostic Representation
标题: SaFari:框架不可知表示的状态空间模型链接:https://arxiv.org/abs/2505.08977
备注:13 pages, 5 figures
摘要:状态空间模型(State-Space Models,SSM)已经重新成为在线函数逼近的强大工具,并成为长距离相关数据的机器学习模型的支柱。然而,到目前为止,只有少数多项式基已被探索用于此目的,最先进的实现是建立在最好的几个有限的选项。在本文中,我们提出了一个广义的方法来建立一个SSM与任何框架或基础,而不是被限制到多项式。这个框架包含了被称为HiPPO的方法,但也允许SSM架构中其他可能的“物种”的无限多样性。我们称这种方法为SaFARI:用于帧不可知表示的SSM。
摘要:State-Space Models (SSMs) have re-emerged as a powerful tool for online function approximation, and as the backbone of machine learning models for long-range dependent data. However, to date, only a few polynomial bases have been explored for this purpose, and the state-of-the-art implementations were built upon the best of a few limited options. In this paper, we present a generalized method for building an SSM with any frame or basis, rather than being restricted to polynomials. This framework encompasses the approach known as HiPPO, but also permits an infinite diversity of other possible "species" within the SSM architecture. We dub this approach SaFARi: SSMs for Frame-Agnostic Representation.
机器翻译由腾讯交互翻译提供,仅供参考
