微信公众号:arXiv_Daily
cs.SD语音
【1】Controllable Embedding Transformation for Mood-Guided Music Retrieval
标题:用于情绪引导音乐检索的可控嵌入变换
链接:https://arxiv.org/abs/2510.20759
备注:Preprint; under review
摘要:音乐表示是现代推荐系统的支柱,支持播放列表生成,相似性搜索和个性化发现。然而,大多数嵌入几乎没有提供调整单个音乐属性的控制,例如,在保留音乐类型或乐器的同时只改变音乐的情绪。在这项工作中,我们解决的问题,可控的音乐检索,通过基于嵌入的变换,其目标是检索歌曲,保持类似的种子轨道,但沿一个选定的尺寸进行修改。我们提出了一种新的框架,用于情绪引导的音乐嵌入变换,该框架学习从种子音频嵌入到情绪标签引导的目标嵌入的映射,同时保留其他音乐属性。由于情绪不能直接改变种子音频,我们引入了一个采样机制,检索代理目标,以平衡多样性与相似性的种子。我们使用这种采样策略训练一个轻量级的翻译模型,并引入一个新的联合目标,鼓励转换和信息保存。在两个数据集上进行的大量实验显示了强大的情绪转换性能,同时保留了远优于无训练基线的流派和仪器,建立了可控的嵌入转换作为个性化音乐检索的一个有前途的范例。
摘要:Music representations are the backbone of modern recommendation systems, powering playlist generation, similarity search, and personalized discovery. Yet most embeddings offer little control for adjusting a single musical attribute, e.g., changing only the mood of a track while preserving its genre or instrumentation. In this work, we address the problem of controllable music retrieval through embedding-based transformation, where the objective is to retrieve songs that remain similar to a seed track but are modified along one chosen dimension. We propose a novel framework for mood-guided music embedding transformation, which learns a mapping from a seed audio embedding to a target embedding guided by mood labels, while preserving other musical attributes. Because mood cannot be directly altered in the seed audio, we introduce a sampling mechanism that retrieves proxy targets to balance diversity with similarity to the seed. We train a lightweight translation model using this sampling strategy and introduce a novel joint objective that encourages transformation and information preservation. Extensive experiments on two datasets show strong mood transformation performance while retaining genre and instrumentation far better than training-free baselines, establishing controllable embedding transformation as a promising paradigm for personalized music retrieval.
【2】R2-SVC: Towards Real-World Robust and Expressive Zero-shot Singing Voice Conversion
标题:R2-SRC:迈向现实世界稳健且表现力强的Zero-Shot歌唱声音转换
链接:https://arxiv.org/abs/2510.20677
备注:5 pages, 2 figures
摘要:在现实世界的歌声转换(SVC)应用中,环境噪声和表达输出的需求提出了重大挑战。然而,传统方法通常在设计时没有考虑实际部署场景,因为训练和推理通常都依赖于干净的数据。这种不匹配阻碍了实际应用,因为音乐分离不可避免地存在各种噪声源和伪像。为了解决这些问题,我们提出了R2-SVC,一个强大的和富有表现力的SVC框架。首先,我们通过随机基频($F_0$)扰动和音乐分离伪影模拟(例如,混响、回声),从而大大提高了在噪声条件下的性能。其次,我们使用特定领域的歌唱数据丰富扬声器表示:除了干净的人声外,我们还结合了DNSMOS过滤的分离人声和公共歌唱语料库,使模型能够在捕捉歌唱风格细微差别的同时保留扬声器音色。第三,我们整合神经源滤波器(NSF)模型,明确表示谐波和噪声成分,提高自然和可控性转换歌唱。R2-SVC在清洁和噪声条件下的多个SVC基准测试中获得了最先进的结果。
摘要:In real-world singing voice conversion (SVC) applications, environmental noise and the demand for expressive output pose significant challenges. Conventional methods, however, are typically designed without accounting for real deployment scenarios, as both training and inference usually rely on clean data. This mismatch hinders practical use, given the inevitable presence of diverse noise sources and artifacts from music separation. To tackle these issues, we propose R2-SVC, a robust and expressive SVC framework. First, we introduce simulation-based robustness enhancement through random fundamental frequency ($F_0$) perturbations and music separation artifact simulations (e.g., reverberation, echo), substantially improving performance under noisy conditions. Second, we enrich speaker representation using domain-specific singing data: alongside clean vocals, we incorporate DNSMOS-filtered separated vocals and public singing corpora, enabling the model to preserve speaker timbre while capturing singing style nuances. Third, we integrate the Neural Source-Filter (NSF) model to explicitly represent harmonic and noise components, enhancing the naturalness and controllability of converted singing. R2-SVC achieves state-of-the-art results on multiple SVC benchmarks under both clean and noisy conditions.
【3】Resounding Acoustic Fields with Reciprocity
标题:具有互惠性的响亮声学场
链接:https://arxiv.org/abs/2510.20602
备注:NeurIPS 2025
摘要:在虚拟环境中实现沉浸式听觉体验需要支持动态源位置的灵活声音建模。在本文中,我们介绍了一个任务,称为回响,其目的是估计房间的脉冲响应在任意的发射器位置从一组稀疏的测量发射器的位置,类似于视觉中的重新照明问题。我们利用互易性,并引入Versa,一种物理启发的方法来促进声场学习。我们的方法通过交换发射器和收听者姿势来创建具有密集虚拟发射器位置的物理有效样本。我们还确定了由于发射器/监听器增益模式部署互惠性的挑战,并提出了一种自监督学习方法来解决这些问题。结果表明,Versa在不同的指标上大大提高了模拟和真实数据集上的声场学习性能。感知用户研究表明,Versa可以极大地改善沉浸式空间声音体验。代码、数据集和演示视频可在项目网站上获得:https://waves.seas.upenn.edu/projects/versa。
摘要:Achieving immersive auditory experiences in virtual environments requires flexible sound modeling that supports dynamic source positions. In this paper, we introduce a task called resounding, which aims to estimate room impulse responses at arbitrary emitter location from a sparse set of measured emitter positions, analogous to the relighting problem in vision. We leverage the reciprocity property and introduce Versa, a physics-inspired approach to facilitating acoustic field learning. Our method creates physically valid samples with dense virtual emitter positions by exchanging emitter and listener poses. We also identify challenges in deploying reciprocity due to emitter/listener gain patterns and propose a self-supervised learning approach to address them. Results show that Versa substantially improve the performance of acoustic field learning on both simulated and real-world datasets across different metrics. Perceptual user studies show that Versa can greatly improve the immersive spatial sound experience. Code, dataset and demo videos are available on the project website: https://waves.seas.upenn.edu/projects/versa.
【4】Decoding the Ear: A Framework for Objectifying Expressiveness from Human Preference Through Efficient Alignment
标题:解码耳朵:通过有效对齐将表现力客观化于人类偏好的框架
链接:https://arxiv.org/abs/2510.20513
备注:Submitted to ICASSP 2026. Demos and codes are available at this https URL
摘要:最近的语音到语音(S2S)模型生成可理解的语音,但仍然缺乏自然的表现力,主要是由于缺乏一个可靠的评估指标。现有的方法,如主观MOS评级,低层次的声学特征,情感识别是昂贵的,有限的,或不完整的。为了解决这个问题,我们提出了Deadly(解码的表达偏好的eAR),一个框架,将人类的偏好语音表达到一个客观的分数。以语音学和心理学为基础,Deepness在三个维度上评估语音:情感,韵律和自发性,使用不到500个注释样本实现与人类感知的高度一致(斯皮尔曼等级相关系数,SRCC = 0.86)。除了可靠的评分外,Deborge还支持公平的基准测试和有针对性的数据管理。它不仅区分了S2S模型之间的表达能力差距,还选择了14K表达性话语来形成ExpressiveSpeech,从而提高了S2S模型的表达能力得分(在100分制上从2.0提高到23.4)。演示和代码可在https://github.com/FreedomIntelligence/ExpressiveSpeech上获得
摘要:Recent speech-to-speech (S2S) models generate intelligible speech but still lack natural expressiveness, largely due to the absence of a reliable evaluation metric. Existing approaches, such as subjective MOS ratings, low-level acoustic features, and emotion recognition are costly, limited, or incomplete. To address this, we present DeEAR (Decoding the Expressive Preference of eAR), a framework that converts human preference for speech expressiveness into an objective score. Grounded in phonetics and psychology, DeEAR evaluates speech across three dimensions: Emotion, Prosody, and Spontaneity, achieving strong alignment with human perception (Spearman's Rank Correlation Coefficient, SRCC = 0.86) using fewer than 500 annotated samples. Beyond reliable scoring, DeEAR enables fair benchmarking and targeted data curation. It not only distinguishes expressiveness gaps across S2S models but also selects 14K expressive utterances to form ExpressiveSpeech, which improves the expressive score (from 2.0 to 23.4 on a 100-point scale) of S2S models. Demos and codes are available at https://github.com/FreedomIntelligence/ExpressiveSpeech
【5】Speaking Clearly: A Simplified Whisper-Based Codec for Low-Bitrate Speech Coding
标题:清晰地说话:用于低比特率语音编码的简化的基于耳语的编解码器
链接:https://arxiv.org/abs/2510.20504
备注:5 pages, 3 figures, 2 tables
摘要:语音编解码器作为连续语音信号和大型语言模型之间的桥梁,但面临着声学保真度和语义保持之间的固有冲突。为了缓解这种冲突,流行的方法增加了声学编解码器与复杂的语义监督。我们探索相反的方向:语义优先的方法,从语义能力的模型开始,并将其用于高保真声学重建。通过实证分析,我们发现,有针对性的架构简化可以释放的声音建模潜力的耳语,文本对齐的自动语音识别(ASR)模型。基于这一发现,我们提出了SimWhisper-Codec,这是一种新颖的编解码器,它通过利用冻结的简化Whisper编码器来平衡语义和声学保存,而无需外部监督。实验结果表明,SimWhisper-Codec在语义保持和声学质量方面都取得了优异的性能,与语义监督的编解码器(如Mimi Codec和SpeechTokenizer)相比,在相似的比特率下,验证了我们的语义优先方法的有效性。代码可在https://github.com/ZhangXinWhut/SimWhisper-Codec上获得。
摘要:Speech codecs serve as bridges between continuous speech signals and large language models, yet face an inherent conflict between acoustic fidelity and semantic preservation. To mitigate this conflict, prevailing methods augment acoustic codecs with complex semantic supervision. We explore the opposite direction: a semantic-first approach that starts from a semantically-capable model and adapts it for high-fidelity acoustic reconstruction. Through empirical analysis, we discover that targeted architectural simplification can unlock the acoustic modeling potential of Whisper, a text-aligned Automatic Speech Recognition (ASR) model. Based on this finding, we propose SimWhisper-Codec, a novel codec that balances the semantic and acoustic preservation by leveraging a frozen, simplified Whisper encoder without requiring external supervision. Experimental results demonstrate that SimWhisper-Codec achieves superior performance in both semantic preservation and acoustic quality compared to semantically-supervised codecs such as Mimi Codec and SpeechTokenizer at similar bitrates, validating the effectiveness of our semantic-first approach. Code is available at https://github.com/ZhangXinWhut/SimWhisper-Codec.
【6】UniSE: A Unified Framework for Decoder-only Autoregressive LM-based Speech Enhancement
标题:UniSE:仅解码器基于自回归LM的语音增强的统一框架
链接:https://arxiv.org/abs/2510.20441
备注:5 pages, submitted to ICASSP 2026
摘要:神经音频编解码器的发展极大地促进了语言模型在语音处理和理解中的应用。然而,缺乏对基于自回归(AR)LM模型的有效性的验证,统一的语音增强(SE)的不同子任务。在这项工作中,我们提出了UniSE,一个统一的解码器只有LM为基础的框架来处理不同的SE任务,包括语音恢复,目标说话人提取和语音分离。它以输入语音特征为条件,利用AR建模生成目标语音的离散标记,这有利于多个任务的不同学习模式之间的兼容性。几个基准测试的实验表明,建议的UniSE可以实现有竞争力的性能相比,歧视性和生成基线,显示能力的LM在统一SE任务。演示页面可以在这里找到:https://github.com/hyyan2k/UniSE。
摘要:The development of neural audio codecs (NACs) has largely promoted applications of language models (LMs) to speech processing and understanding. However, there lacks the verification on the effectiveness of autoregressive (AR) LMbased models in unifying different sub-tasks of speech enhancement (SE). In this work, we propose UniSE, a unified decoder-only LM-based framework to handle different SE tasks including speech restoration, target speaker extraction and speech separation. It takes input speech features as conditions and generates discrete tokens of the target speech using AR modeling, which facilitates a compatibility between distinct learning patterns of multiple tasks. Experiments on several benchmarks indicate the proposed UniSE can achieve competitive performance compared to discriminative and generative baselines, showing the capacity of LMs in unifying SE tasks. The demo page is available here: https://github.com/hyyan2k/UniSE.
【7】From Generation to Attribution: Music AI Agent Architectures for the Post-Streaming Era
标题:从一代到归属:后流媒体时代的音乐AI代理架构
链接:https://arxiv.org/abs/2510.20276
备注:Accepted to the NeurIPS 2025 AI4Music Workshop
摘要:生成式人工智能正在重塑音乐创作,但其快速增长暴露了归属、版权管理和经济模式方面的结构性差距。与过去的媒体转移不同,从现场表演到录音,下载和流媒体,人工智能改变了音乐的整个生命周期,打破了创作,分发和货币化之间的界限。然而,现有的流媒体系统具有不透明和集中的版税流,无法处理人工智能驱动的生产的规模和复杂性。我们提出了一个基于内容的音乐AI代理架构,通过块级检索和代理编排将属性直接嵌入到创作工作流中。该系统专为迭代的、基于会话的交互而设计,将音乐组织成存储在BlockDB中的粒度组件(块);每次使用都会触发一个归因层事件,以实现透明的出处和实时结算。该框架将AI从生成工具重新构建为公平AI媒体平台的基础设施。通过实现细粒度的归属、公平的补偿和参与性的参与,它指向了一个后流媒体模式,在这个模式中,音乐不是作为一个静态的目录,而是作为一个协作和适应性的生态系统。
摘要:Generative AI is reshaping music creation, but its rapid growth exposes structural gaps in attribution, rights management, and economic models. Unlike past media shifts, from live performance to recordings, downloads, and streaming, AI transforms the entire lifecycle of music, collapsing boundaries between creation, distribution, and monetization. However, existing streaming systems, with opaque and concentrated royalty flows, are ill-equipped to handle the scale and complexity of AI-driven production. We propose a content-based Music AI Agent architecture that embeds attribution directly into the creative workflow through block-level retrieval and agentic orchestration. Designed for iterative, session-based interaction, the system organizes music into granular components (Blocks) stored in BlockDB; each use triggers an Attribution Layer event for transparent provenance and real-time settlement. This framework reframes AI from a generative tool into infrastructure for a Fair AI Media Platform. By enabling fine-grained attribution, equitable compensation, and participatory engagement, it points toward a post-streaming paradigm where music functions not as a static catalog but as a collaborative and adaptive ecosystem.
【8】Vox-Evaluator: Enhancing Stability and Fidelity for Zero-shot TTS with A Multi-Level Evaluator
标题:Vox评估器:通过多级别评估器增强Zero-ShotTTC的稳定性和保真度
链接:https://arxiv.org/abs/2510.20210
备注:10 pages, 5 figures
摘要:zero-shot文语转换(TTS)技术在语言模型、扩散模型和掩蔽生成等方面取得了巨大的进展,在语音合成中取得了令人印象深刻的自然度。然而,稳定性和保真度仍然是关键的挑战,表现为发音错误,可听见的噪音和质量下降。为了解决这些问题,我们引入了Vox-Evaluator,一个多层次的评估,旨在指导错误的语音段和偏好对齐TTS系统的纠正。它能够识别错误段的时间边界,并提供生成的语音的整体质量评估。具体而言,完善错误的部分,并提高鲁棒性的zero-shot TTS模型,我们建议自动识别声学错误的评估,掩盖错误的部分,并最终重新生成语音调节的正确部分。此外,Vox-Evaluator获得的精细信息可以指导TTS模型的偏好对齐,从而减少语音合成中的不良情况。由于Vox-Evaluator缺乏合适的训练数据集,我们还构建了一个带有细粒度发音错误或音频质量问题的合成文本语音数据集。实验结果表明,该Vox-Evaluator通过语音纠错机制和偏好优化,有效地提高了TTS系统的稳定性和保真度。演示已显示。
摘要:Recent advances in zero-shot text-to-speech (TTS), driven by language models, diffusion models and masked generation, have achieved impressive naturalness in speech synthesis. Nevertheless, stability and fidelity remain key challenges, manifesting as mispronunciations, audible noise, and quality degradation. To address these issues, we introduce Vox-Evaluator, a multi-level evaluator designed to guide the correction of erroneous speech segments and preference alignment for TTS systems. It is capable of identifying the temporal boundaries of erroneous segments and providing a holistic quality assessment of the generated speech. Specifically, to refine erroneous segments and enhance the robustness of the zero-shot TTS model, we propose to automatically identify acoustic errors with the evaluator, mask the erroneous segments, and finally regenerate speech conditioning on the correct portions. In addition, the fine-gained information obtained from Vox-Evaluator can guide the preference alignment for TTS model, thereby reducing the bad cases in speech synthesis. Due to the lack of suitable training datasets for the Vox-Evaluator, we also constructed a synthesized text-speech dataset annotated with fine-grained pronunciation errors or audio quality issues. The experimental results demonstrate the effectiveness of the proposed Vox-Evaluator in enhancing the stability and fidelity of TTS systems through the speech correction mechanism and preference optimization. The demos are shown.
【9】SpeechAgent: An End-to-End Mobile Infrastructure for Speech Impairment Assistance
标题:SpeechAgent:一个端到端的语音障碍辅助移动基础设施
链接:https://arxiv.org/abs/2510.20113
摘要:言语对于人类交流至关重要,但数百万人面临构音障碍、口吃和失语症等障碍,这些障碍往往导致社会孤立和参与减少。尽管最近在自动语音识别(ASR)和文本到语音(TTS)技术方面取得了进展,但语音受损用户的可访问Web和移动基础设施仍然有限,阻碍了这些进步在日常通信中的实际应用。为了弥合这一差距,我们提出了SpeechAgent,一个移动SpeechAgent,旨在方便人们在日常交流中的语音障碍。该系统集成了大型语言模型(LLM)驱动的推理与先进的语音处理模块,提供自适应支持,适合不同的障碍类型。为了确保现实世界的实用性,我们开发了一个结构化的部署管道,可以在移动和边缘设备上实现实时语音处理,实现难以察觉的延迟,同时保持高准确性和语音质量。对真实世界受损语音数据集和边缘设备延迟分析的评估证实,SpeechAgent提供了有效和用户友好的性能,证明了其个性化日常辅助通信的可行性。
摘要:Speech is essential for human communication, yet millions of people face impairments such as dysarthria, stuttering, and aphasia conditions that often lead to social isolation and reduced participation. Despite recent progress in automatic speech recognition (ASR) and text-to-speech (TTS) technologies, accessible web and mobile infrastructures for users with impaired speech remain limited, hindering the practical adoption of these advances in daily communication. To bridge this gap, we present SpeechAgent, a mobile SpeechAgent designed to facilitate people with speech impairments in everyday communication. The system integrates large language model (LLM)- driven reasoning with advanced speech processing modules, providing adaptive support tailored to diverse impairment types. To ensure real-world practicality, we develop a structured deployment pipeline that enables real-time speech processing on mobile and edge devices, achieving imperceptible latency while maintaining high accuracy and speech quality. Evaluation on real-world impaired speech datasets and edge-device latency profiling confirms that SpeechAgent delivers both effective and user-friendly performance, demonstrating its feasibility for personalized, day-to-day assistive communication.
【1】Time-series Random Process Complexity Ranking Using a Bound on Conditional Differential Entropy
标题:使用条件差熵界的时间序列随机过程复杂性排名
链接:https://arxiv.org/abs/2510.20551
备注:7 pages, 4 figures
摘要:条件微分熵通过量化给定过去背景的未来观测中的不确定性,为时间序列复杂性的相对排名提供了一种直观的衡量标准。然而,它对未知分布的高维过程的直接计算往往很棘手。本文建立在Fang et al. \cite{fang 2019 generic}建立的信息论预测误差界的基础上,该误差界证明了条件微分熵\textbf{$h(X_k \mid X_{k-1},. X_{k-m})$}的上界是下一步预测误差协方差矩阵行列式的函数。我们添加到这个理论框架,通过利用Hadamard不等式和协方差矩阵的半正定性质进一步增加这个界限。 为了看看这些界限是否可以用来对时间序列的复杂性进行排名,我们进行了两个合成实验:(1)具有加性高斯噪声的受控线性自回归过程,其中我们将普通最小二乘预测误差熵代理与各种加性噪声的真实熵进行比较,以及(2)具有未知熵的生物启发合成音频数据的复杂度排名任务,其中神经网络预测误差用于恢复已知的复杂度排序。 该框架提供了一种计算上易于处理的方法,用于使用来自下一步预测模型的预测误差进行时间序列复杂度排名,该方法保持了信息论的理论基础。
摘要:Conditional differential entropy provides an intuitive measure for relatively ranking time-series complexity by quantifying uncertainty in future observations given past context. However, its direct computation for high-dimensional processes from unknown distributions is often intractable. This paper builds on the information theoretic prediction error bounds established by Fang et al. \cite{fang2019generic}, which demonstrate that the conditional differential entropy \textbf{$h(X_k \mid X_{k-1},...,X_{k-m})$} is upper bounded by a function of the determinant of the covariance matrix of next-step prediction errors for any next step prediction model. We add to this theoretical framework by further increasing this bound by leveraging Hadamard's inequality and the positive semi-definite property of covariance matrices. To see if these bounds can be used to rank the complexity of time series, we conducted two synthetic experiments: (1) controlled linear autoregressive processes with additive Gaussian noise, where we compare ordinary least squares prediction error entropy proxies to the true entropies of various additive noises, and (2) a complexity ranking task of bio-inspired synthetic audio data with unknown entropy, where neural network prediction errors are used to recover the known complexity ordering. This framework provides a computationally tractable method for time-series complexity ranking using prediction errors from next-step prediction models, that maintains a theoretical foundation in information theory.
【2】Neural Directional Filtering with Configurable Directivity Pattern at Inference
标题:具有可配置推理方向性模式的神经方向过滤
链接:https://arxiv.org/abs/2510.20253
摘要:具有期望的方向性图案的空间滤波对于许多音频应用是有利的。在这项工作中,我们提出了神经方向过滤与用户定义的方向性模式(UNDF),使空间过滤的基础上,用户可以定义在推理过程中的方向性模式。为了实现这一目标,我们提出了一种DNN架构,该架构集成了特征线性调制(FILM),允许用户定义的模式作为条件输入。通过分析,我们证明了基于FiLM的体系结构使UNDF能够在具有更高方向性、缩放变化和不同转向方向的干扰期间推广到看不见的用户定义模式。此外,我们逐步完善训练策略,以提高模式近似,使UNDF近似不规则形状。最后,实验比较表明,UNDF优于传统的方法。
摘要:Spatial filtering with a desired directivity pattern is advantageous for many audio applications. In this work, we propose neural directional filtering with user-defined directivity patterns (UNDF), which enables spatial filtering based on directivity patterns that users can define during inference. To achieve this, we propose a DNN architecture that integrates feature-wise linear modulation (FiLM), allowing user-defined patterns to serve as conditioning inputs. Through analysis, we demonstrate that the FiLM-based architecture enables the UNDF to generalize to unseen user-defined patterns during interference with higher directivities, scaling variations, and different steering directions. Furthermore, we progressively refine training strategies to enhance pattern approximation and enable UNDF to approximate irregular shapes. Lastly, experimental comparisons show that UNDF outperforms conventional methods.
【3】R2-SVC: Towards Real-World Robust and Expressive Zero-shot Singing Voice Conversion
标题:R2-SRC:迈向现实世界稳健且表现力强的Zero-Shot歌唱声音转换
链接:https://arxiv.org/abs/2510.20677
备注:5 pages, 2 figures
摘要:在现实世界的歌声转换(SVC)应用中,环境噪声和表达输出的需求提出了重大挑战。然而,传统方法通常在设计时没有考虑实际部署场景,因为训练和推理通常都依赖于干净的数据。考虑到音乐分离不可避免地存在不同的噪声源和伪影,这种不匹配阻碍了实际使用。为了解决这些问题,我们提出了R2-SVC,一个强大的和富有表现力的SVC框架。首先,我们通过随机基频($F_0$)扰动和音乐分离伪影模拟(例如,混响、回声),从而大大提高了在噪声条件下的性能。其次,我们使用特定领域的歌唱数据丰富了扬声器表示:除了干净的人声外,我们还结合了DNSMOS过滤的分离人声和公共歌唱语料库,使模型能够保留扬声器音色,同时捕捉歌唱风格的细微差别。第三,我们整合神经源滤波器(NSF)模型,明确表示谐波和噪声成分,提高自然和可控性转换歌唱。R2-SVC在清洁和噪声条件下的多个SVC基准测试中获得了最先进的结果。
摘要:In real-world singing voice conversion (SVC) applications, environmental noise and the demand for expressive output pose significant challenges. Conventional methods, however, are typically designed without accounting for real deployment scenarios, as both training and inference usually rely on clean data. This mismatch hinders practical use, given the inevitable presence of diverse noise sources and artifacts from music separation. To tackle these issues, we propose R2-SVC, a robust and expressive SVC framework. First, we introduce simulation-based robustness enhancement through random fundamental frequency ($F_0$) perturbations and music separation artifact simulations (e.g., reverberation, echo), substantially improving performance under noisy conditions. Second, we enrich speaker representation using domain-specific singing data: alongside clean vocals, we incorporate DNSMOS-filtered separated vocals and public singing corpora, enabling the model to preserve speaker timbre while capturing singing style nuances. Third, we integrate the Neural Source-Filter (NSF) model to explicitly represent harmonic and noise components, enhancing the naturalness and controllability of converted singing. R2-SVC achieves state-of-the-art results on multiple SVC benchmarks under both clean and noisy conditions.
【4】Resounding Acoustic Fields with Reciprocity
标题:具有互惠性的响亮声学场
链接:https://arxiv.org/abs/2510.20602
备注:NeurIPS 2025
摘要:在虚拟环境中实现沉浸式听觉体验需要支持动态源位置的灵活声音建模。在本文中,我们介绍了一个任务,称为回响,其目的是估计房间的脉冲响应在任意的发射器位置从一组稀疏的测量发射器的位置,类似于视觉中的重新照明问题。我们利用互易性,并引入Versa,一种物理启发的方法来促进声场学习。我们的方法通过交换发射器和收听者姿势来创建具有密集虚拟发射器位置的物理有效样本。我们还确定了由于发射器/监听器增益模式部署互惠性的挑战,并提出了一种自监督学习方法来解决这些问题。结果表明,Versa在不同的指标上大大提高了模拟和真实数据集上的声场学习性能。感知用户研究表明,Versa可以极大地改善沉浸式空间声音体验。代码、数据集和演示视频可在项目网站上获得:https://waves.seas.upenn.edu/projects/versa。
摘要:Achieving immersive auditory experiences in virtual environments requires flexible sound modeling that supports dynamic source positions. In this paper, we introduce a task called resounding, which aims to estimate room impulse responses at arbitrary emitter location from a sparse set of measured emitter positions, analogous to the relighting problem in vision. We leverage the reciprocity property and introduce Versa, a physics-inspired approach to facilitating acoustic field learning. Our method creates physically valid samples with dense virtual emitter positions by exchanging emitter and listener poses. We also identify challenges in deploying reciprocity due to emitter/listener gain patterns and propose a self-supervised learning approach to address them. Results show that Versa substantially improve the performance of acoustic field learning on both simulated and real-world datasets across different metrics. Perceptual user studies show that Versa can greatly improve the immersive spatial sound experience. Code, dataset and demo videos are available on the project website: https://waves.seas.upenn.edu/projects/versa.
机器翻译由腾讯交互翻译提供,仅供参考
