今日论文合集:cs.SD语音14篇,eess.AS音频处理14篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】Align Your Rhythm: Generating Highly Aligned Dance Poses with  Gating-Enhanced Rhythm-Aware Feature Representation

标题:对齐节奏:使用门控增强节奏感知特征表示生成高度对齐的舞蹈姿势
链接:https://arxiv.org/abs/2503.17340
作者:Congyi Fan,  Jian Guan,  Xuanjia Zhao,  Dongli Xu,  Youtian Lin,  Tong Ye,  Pengming Feng,  Haiwei Pan
备注:10 pages, 6 figures
摘要:自动生成由音乐驱动的自然,多样和有节奏的人类舞蹈动作对于虚拟现实和电影行业至关重要。然而,生成自然地跟随音乐的舞蹈仍然是一个挑战,因为现有的方法缺乏适当的节拍对齐并且表现出不自然的运动动态。在本文中,我们提出了Danceba,一个新的框架,利用门控机制,以提高节奏感知功能表示音乐驱动的舞蹈生成,实现高度对齐的舞蹈姿势与增强的节奏敏感性。具体来说,我们引入基于相位的节奏提取(PRE),精确地提取节奏信息的音乐相位数据,利用音乐的内在周期性和时间结构。此外,我们提出了时间门控因果注意(TGCA),专注于全球的节奏特征,确保舞蹈动作紧密跟随音乐节奏。我们还引入了并行曼巴运动建模(PMMM)架构,分别模拟上,下半身的运动以及音乐功能,从而提高了自然性和多样性的舞蹈动作。大量的实验证实,Danceba优于最先进的方法,实现了更好的节奏对齐和运动多样性。项目页面:https://danceba.github.io/。
摘要:Automatically generating natural, diverse and rhythmic human dance movementsdriven by music is vital for virtual reality and film industries. However,generating dance that naturally follows music remains a challenge, as existingmethods lack proper beat alignment and exhibit unnatural motion dynamics. Inthis paper, we propose Danceba, a novel framework that leverages gatingmechanism to enhance rhythm-aware feature representation for music-driven dancegeneration, which achieves highly aligned dance poses with enhanced rhythmicsensitivity. Specifically, we introduce Phase-Based Rhythm Extraction (PRE) toprecisely extract rhythmic information from musical phase data, capitalizing onthe intrinsic periodicity and temporal structures of music. Additionally, wepropose Temporal-Gated Causal Attention (TGCA) to focus on global rhythmicfeatures, ensuring that dance movements closely follow the musical rhythm. Wealso introduce Parallel Mamba Motion Modeling (PMMM) architecture to separatelymodel upper and lower body motions along with musical features, therebyimproving the naturalness and diversity of generated dance movements. Extensiveexperiments confirm that Danceba outperforms state-of-the-art methods,achieving significantly better rhythmic alignment and motion diversity. Projectpage: https://danceba.github.io/ .

【2】 Learning disentangled representations for instrument-based music  similarity
标题:学习基于乐器的音乐相似性的解开表示
链接:https://arxiv.org/abs/2503.17281
作者:Yuka Hashizume,  Li Li,  Atsushi Miyashita,  Tomoki Toda
备注:arXiv admin note: text overlap with arXiv:2404.06682
摘要:灵活的推荐和检索系统需要音乐片段的多个部分元素方面的音乐相似性,以允许用户选择他们想要关注的元素。使用具有单独乐器信号的多个网络进行音乐相似性学习的方法是有效的,但面临的问题是,使用每个干净的乐器信号作为检索系统的查询是不切实际的,并且使用分离的乐器声音降低了准确性因为艺术品。在本文中,我们提出了基于乐器部分的音乐相似性学习与一个单一的网络,以混合的声音作为输入,而不是个别乐器的声音。具体来说,我们为每个仪器设计了一个具有解纠缠维度的单一相似性嵌入空间,由条件相似性网络提取,该网络使用带有掩码的三重丢失进行训练。实验结果表明:(1)在对乐器进行评价时,与使用单独的声音作为输入的网络相比,该方法可以获得更准确的特征表示;(2)每个子嵌入空间都可以保持相应乐器的特征,以及(3)通过所提出的方法选择关注每个乐器声音的相似音乐作品可以获得人的接受,特别是当关注音色时。
摘要:A flexible recommendation and retrieval system requires music similarity interms of multiple partial elements of musical pieces to allow users to selectthe element they want to focus on. A method for music similarity learning usingmultiple networks with individual instrumental signals is effective but facesthe problem that using each clean instrumental signal as a query is impracticalfor retrieval systems and using separated instrumental sounds reduces accuracyowing to artifacts. In this paper, we present instrumental-part-based musicsimilarity learning with a single network that takes mixed sounds as inputinstead of individual instrumental sounds. Specifically, we designed a singlesimilarity embedding space with disentangled dimensions for each instrument,extracted by Conditional Similarity Networks, which are trained using thetriplet loss with masks. Experimental results showed that (1) the proposedmethod can obtain more accurate feature representation than using individualnetworks using separated sounds as input in the evaluation of an instrumentthat had low accuracy, (2) each sub-embedding space can hold thecharacteristics of the corresponding instrument, and (3) the selection ofsimilar musical pieces focusing on each instrumental sound by the proposedmethod can obtain human acceptance, especially when focusing on timbre.

【3】 HiFi-Stream: Streaming Speech Enhancement with Generative Adversarial  Networks
标题:HiFi-Stream:使用生成对抗网络的流媒体语音增强
链接:https://arxiv.org/abs/2503.17141
作者:Ekaterina Dmitrieva,  Maksim Kaledin
备注:5 pages (4 content pages + 1 page of references)
摘要:语音增强技术已经成为移动设备和语音软件中简化下游语音任务的核心技术。尽管如此,现代深度学习(DL)解决方案通常需要大量的计算资源,这使得它们在低资源设备上的使用具有挑战性。我们提出了HiFi-Stream,最近发布的HiFi++模型的优化版本。我们的实验表明,HiFiStream保留了原始模型的大部分质量,尽管它的大小和计算复杂性:最轻的版本只有大约490 k个参数,与原始HiFi++相比减少了3.5倍,使其成为最小和最快的模型之一。该模型在流媒体设置中进行评估,与现代基线相比,它表现出卓越的性能。
摘要:Speech Enhancement techniques have become core technologies in mobile devicesand voice software simplifying downstream speech tasks. Still, modern DeepLearning (DL) solutions often require high amount of computational resourceswhat makes their usage on low-resource devices challenging. We presentHiFi-Stream, an optimized version of recently published HiFi++ model. Ourexperiments demonstrate that HiFiStream saves most of the qualities of theoriginal model despite its size and computational complexity: the lightestversion has only around 490k parameters which is 3.5x reduction in comparisonto the original HiFi++ making it one of the smallest and fastest modelsavailable. The model is evaluated in streaming setting where it demonstratesits superior performance in comparison to modern baselines.

【4】 DIDiffGes: Decoupled Semi-Implicit Diffusion Models for Real-time  Gesture Generation from Speech
标题:DIDiffGes:用于从语音实时生成手势的去耦合半隐式扩散模型
链接:https://arxiv.org/abs/2503.17059
作者:Yongkang Cheng,  Shaoli Huang,  Xuelin Chen,  Jifeng Ning,  Mingming Gong
备注:Accepted by AAAI 2025
摘要:扩散模型在产生共同语音手势方面表现出显着的合成质量和多样性。然而,与扩散模型相关的计算密集型采样步骤阻碍了其在现实世界应用中的实用性。因此,我们提出了DIDiffGes,一个解耦的半隐式扩散模型为基础的框架,可以合成高品质的,表达的手势,从语音只使用几个采样步骤。我们的方法利用生成对抗网络(GANs)来实现扩散模型的大步长采样。我们将手势数据解耦为身体和手的分布,并进一步将其分解为边缘和条件分布。GANs隐式地对边缘分布进行建模,而L2重建损失则对条件分布进行了精确的学习。该策略增强了GAN训练的稳定性,并确保生成的全身姿势的表现力。我们的框架还学习去噪以局部身体表示为条件的根噪声,保证稳定性和真实性。DIDiffGes只需10个采样步骤就可以从语音中生成手势,而不会影响质量和表现力,与现有方法相比,采样步骤的数量减少了100倍。我们的用户研究表明,我们的方法优于国家的最先进的方法在人类的相似性,适当性和风格的正确性。项目是https://cyk990422.github.io/DIDiffGes。
摘要:Diffusion models have demonstrated remarkable synthesis quality and diversityin generating co-speech gestures. However, the computationally intensivesampling steps associated with diffusion models hinder their practicality inreal-world applications. Hence, we present DIDiffGes, for a DecoupledSemi-Implicit Diffusion model-based framework, that can synthesizehigh-quality, expressive gestures from speech using only a few sampling steps.Our approach leverages Generative Adversarial Networks (GANs) to enablelarge-step sampling for diffusion model. We decouple gesture data into body andhands distributions and further decompose them into marginal and conditionaldistributions. GANs model the marginal distribution implicitly, while L2reconstruction loss learns the conditional distributions exciplictly. Thisstrategy enhances GAN training stability and ensures expressiveness ofgenerated full-body gestures. Our framework also learns to denoise root noiseconditioned on local body representation, guaranteeing stability and realism.DIDiffGes can generate gestures from speech with just 10 sampling steps,without compromising quality and expressiveness, reducing the number ofsampling steps by a factor of 100 compared to existing methods. Our user studyreveals that our method outperforms state-of-the-art approaches in humanlikeness, appropriateness, and style correctness. Project ishttps://cyk990422.github.io/DIDiffGes.

【5】 Symbolic Audio Classification via Modal Decision Tree Learning
标题:通过模式决策树学习进行符号音频分类
链接:https://arxiv.org/abs/2503.17018
作者:Enrico Marzano,  Giovanni Pagliarini,  Riccardo Pasini,  Guido Sciavicco,  Ionel Eduard Stan
摘要:声学分析的潜在应用范围很广。特别是声音分类,是近年来受到广泛关注的典型机器学习任务。最常见的声音分类方法是基于神经网络的子符号方法,并且导致具有高性能但透明度非常低的黑箱模型。在这项工作中,我们考虑了几个音频任务,即年龄和性别识别,情感分类和呼吸系统疾病诊断,我们接近他们的象征性技术,即(模态)决策树学习。我们证明了这些任务可以使用相同的符号管道来解决,该符号管道允许以非常高的准确性和低复杂性提取简单的规则。原则上,所有这些任务都可以与自主对话系统相关联,这在不同的环境中可能是有用的,例如医院或诊所的自动预约代理。
摘要:The range of potential applications of acoustic analysis is wide.Classification of sounds, in particular, is a typical machine learning taskthat received a lot of attention in recent years. The most common approaches tosound classification are sub-symbolic, typically based on neural networks, andresult in black-box models with high performances but very low transparency. Inthis work, we consider several audio tasks, namely, age and gender recognition,emotion classification, and respiratory disease diagnosis, and we approach themwith a symbolic technique, that is, (modal) decision tree learning. We provethat such tasks can be solved using the same symbolic pipeline, that allows toextract simple rules with very high accuracy and low complexity. In principle,all such tasks could be associated to an autonomous conversation system, whichcould be useful in different contexts, such as an automatic reservation agentfor an hospital or a clinic.

【6】 STFTCodec: High-Fidelity Audio Compression through Time-Frequency Domain  Representation
标题:STFTCodec:通过时频域表示的高保真音频压缩
链接:https://arxiv.org/abs/2503.16989
作者:Tao Feng,  Zhiyuan Zhao,  Yifan Xie,  Yuqi Ye,  Xiangyang Luo,  Xun Guan,  Yu Li
备注:7 pages, 2 figures, accepted by ICME 2025
摘要:我们提出了STFTCodec,一种新的基于频谱的神经音频编解码器,使用短时傅立叶变换(STFT)有效地压缩音频。与基于波形的方法,需要大的模型容量和大量的内存消耗,该方法利用STFT紧凑的频谱表示,并引入展开相位导数作为辅助功能。我们的架构采用并行的幅度和相位处理分支增强先进的特征提取机制。通过放松严格的相位重建约束,同时保持相位感知处理,我们实现了卓越的感知质量。实验结果表明,STFTCodec优于基于波形和基于频谱的方法在多个比特率,同时提供独特的灵活性,在压缩比调整,通过STFT参数修改,而无需架构的变化。
摘要:We present STFTCodec, a novel spectral-based neural audio codec thatefficiently compresses audio using Short-Time Fourier Transform (STFT). Unlikewaveform-based approaches that require large model capacity and substantialmemory consumption, this method leverages STFT for compact spectralrepresentation and introduces unwrapped phase derivatives as auxiliaryfeatures. Our architecture employs parallel magnitude and phase processingbranches enhanced by advanced feature extraction mechanisms. By relaxing strictphase reconstruction constraints while maintaining phase-aware processing, weachieve superior perceptual quality. Experimental results demonstrate thatSTFTCodec outperforms both waveform-based and spectral-based approaches acrossmultiple bitrates, while offering unique flexibility in compression ratioadjustment through STFT parameter modification without architectural changes.

【7】 City2Scene: Improving Acoustic Scene Classification with City Features
标题:City2 Scene:利用城市特色改进声学场景分类
链接:https://arxiv.org/abs/2503.16862
作者:Yiqiang Cai,  Yizhou Tan,  Peihong Zhang,  Yuxuan Liu,  Shengchen Li,  Xi Shao,  Mark D. Plumbley
摘要:声学现场录音通常从不同的城市收集。大多数现有的声学场景分类(ASC)方法侧重于识别城市之间的共同声学场景模式,以提高泛化能力。相比之下,我们假设城市特定的环境和文化差异的声学特征是有益的ASC任务。在本文中,我们介绍了City2Scene,一个新的框架,利用城市功能,以提高ASC。City2Scene利用知识蒸馏将城市分类模型中的城市特定知识转换为场景分类模型。我们在DCASE挑战任务1数据集上评估了City2Scene,其中每个音频片段都用场景和城市标签进行了注释。实验结果表明,城市特征为场景分类提供了有价值的信息。通过提取城市特定的知识,City2Scene有效地提高了各种最先进的ASC骨干模型的准确性,包括CNN和Transformers。
摘要:Acoustic scene recordings are often collected from a diverse range of cities.Most existing acoustic scene classification (ASC) approaches focus onidentifying common acoustic scene patterns across cities to enhancegeneralization. In contrast, we hypothesize that city-specific environmentaland cultural differences in acoustic features are beneficial for the ASC task.In this paper, we introduce City2Scene, a novel framework that leverages cityfeatures to improve ASC. City2Scene transfers the city-specific knowledge fromcity classification models to a scene classification model using knowledgedistillation. We evaluated City2Scene on the DCASE Challenge Task 1 datasets,where each audio clip is annotated with both scene and city labels.Experimental results demonstrate that city features provide valuableinformation for classifying scenes. By distilling the city-specific knowledge,City2Scene effectively improves accuracy for various state-of-the-art ASCbackbone models, including both CNNs and Transformers.

【8】 Imagine to Hear: Auditory Knowledge Generation can be an Effective  Assistant for Language Models
标题:想象聆听:听觉知识生成可以成为语言模型的有效助手
链接:https://arxiv.org/abs/2503.16853
作者:Suho Yoo,  Hyunjong Ok,  Jaeho Lee
备注:Preprint
摘要:在纯文本语料库上预训练的语言模型通常难以完成需要听觉常识知识的任务。以前的工作解决了这个问题,通过增强语言模型从外部音频数据库中检索知识。这种方法有几个局限性,如数据库中可能缺乏相关音频,以及与构建和查询数据库相关的高成本。为了解决这些问题,我们提出了想象听到,一种新的方法,动态生成听觉知识生成模型。我们的框架检测多个音频相关的文本跨度从给定的提示,并生成相应的听觉知识。我们开发了几种机制来有效地处理多种听觉知识,包括基于CLAP的拒绝采样器和语言音频融合模块。我们的实验表明,我们的方法在AuditoryBench上实现了最先进的性能,而不依赖于外部数据库,突出了我们基于生成的方法的有效性。
摘要:Language models pretrained on text-only corpora often struggle with tasksthat require auditory commonsense knowledge. Previous work addresses thisproblem by augmenting the language model to retrieve knowledge from externalaudio databases. This approach has several limitations, such as the potentiallack of relevant audio in databases and the high costs associated withconstructing and querying the databases. To address these issues, we proposeImagine to Hear, a novel approach that dynamically generates auditory knowledgeusing generative models. Our framework detects multiple audio-related textualspans from the given prompt and generates corresponding auditory knowledge. Wedevelop several mechanisms to efficiently process multiple auditory knowledge,including a CLAP-based rejection sampler and a language-audio fusion module.Our experiments show that our method achieves state-of-the-art performance onAuditoryBench without relying on external databases, highlighting theeffectiveness of our generation-based approach.

【9】 The Deployment of End-to-End Audio Language Models Should Take into  Account the Principle of Least Privilege
标题:端到端音频语言模型的部署应考虑最小特权原则
链接:https://arxiv.org/abs/2503.16833
作者:Luxi He,  Xiangyu Qi,  Michel Liao,  Inyoung Cheong,  Prateek Mittal,  Danqi Chen,  Peter Henderson
摘要:我们正处于接受音频输入的语言模型的转折点。最新的端到端音频语言模型(Audio LM)直接处理语音,而不是依赖于单独的转录步骤。这种转换保留了详细的信息,如语调或多个发言者的存在,否则这些信息将在转录中丢失。然而,它也引入了新的安全风险,包括可能滥用说话者身份线索和其他敏感的声音属性,这可能会产生法律影响。在这份立场文件中,我们敦促更仔细地审查如何建立和部署这些模型。我们认为,最小特权的原则应该指导决定是否部署级联或端到端的模型。具体来说,评估应该评估(1)端到端建模对于给定的应用程序是否必要;以及(2)信息访问的适当范围。最后,我们强调了当前音频LM基准测试中的相关差距,并确定了关键的开放研究问题,包括技术和政策相关的问题,这些问题必须得到解决,才能实现负责任的端到端音频LM部署。
摘要:We are at a turning point for language models that accept audio input. Thelatest end-to-end audio language models (Audio LMs) process speech directlyinstead of relying on a separate transcription step. This shift preservesdetailed information, such as intonation or the presence of multiple speakers,that would otherwise be lost in transcription. However, it also introduces newsafety risks, including the potential misuse of speaker identity cues and othersensitive vocal attributes, which could have legal implications. In thisposition paper, we urge a closer examination of how these models are built anddeployed. We argue that the principle of least privilege should guide decisionson whether to deploy cascaded or end-to-end models. Specifically, evaluationsshould assess (1) whether end-to-end modeling is necessary for a givenapplication; and (2), the appropriate scope of information access. Finally, Wehighlight related gaps in current audio LM benchmarks and identify key openresearch questions, both technical and policy-related, that must be addressedto enable the responsible deployment of end-to-end Audio LMs.

【10】 CAARMA: Class Augmentation with Adversarial Mixup Regularization
标题:CAARMA:具有对抗混淆正规化的类扩充
链接:https://arxiv.org/abs/2503.16718
作者:Massa Baali,  Xiang Li,  Hao Chen,  Rita Singh,  Bhiksha Raj
摘要:说话人确认是一个典型的zero-shot学习任务,其中未见过的类的推理是通过将测试实例的嵌入与已知实例进行比较来执行的。因此,执行推理的模型必须自然地生成嵌入,这些嵌入将同类实例聚集在一起,同时保持类之间的分离。为了学会这样做,他们通常在大量的类(扬声器)上进行训练,通常使用专门的损失。然而,现实世界的说话者数据集通常缺乏以可推广的方式有效学习所需的类别多样性。我们引入CAARMA,一个类增强框架,通过在嵌入空间中混合数据生成合成类,扩展训练类的数量来解决这个问题。为了确保合成类的真实性,我们采用了一种新的对抗性细化机制,最大限度地减少了合成类和真实类之间的分类差异。我们评估CAARMA对多个说话人验证任务,以及其他代表性的zero-shot比较为基础的语音分析任务,并获得一致的改进:我们的框架表现出了显着的改善8%以上的所有基线模型。CAARMA的代码将被释放。
摘要:Speaker verification is a typical zero-shot learning task, where inference ofunseen classes is performed by comparing embeddings of test instances to knownexamples. The models performing inference must hence naturally generateembeddings that cluster same-class instances compactly, while maintainingseparation across classes. In order to learn to do so, they are typicallytrained on a large number of classes (speakers), often using specializedlosses. However real-world speaker datasets often lack the class diversityneeded to effectively learn this in a generalizable manner. We introduceCAARMA, a class augmentation framework that addresses this problem bygenerating synthetic classes through data mixing in the embedding space,expanding the number of training classes. To ensure the authenticity of thesynthetic classes we adopt a novel adversarial refinement mechanism thatminimizes categorical distinctions between synthetic and real classes. Weevaluate CAARMA on multiple speaker verification tasks, as well as otherrepresentative zero-shot comparison-based speech analysis tasks and obtainconsistent improvements: our framework demonstrates a significant improvementof 8\% over all baseline models. Code for CAARMA will be released.

【11】 WaveFM: A High-Fidelity and Efficient Vocoder Based on Flow Matching
标题:WaveFM:一种基于流匹配的高保真高效声码器
链接:https://arxiv.org/abs/2503.16689
作者:Tianze Luo,  Xingchen Miao,  Wenbo Duan
备注:Accepted to the main conference of NAACL 2025. The codes are available at this https URL
摘要:流匹配为训练扩散模型提供了一种强大而稳定的方法。然而,直接将流匹配应用于神经声码器可能导致低于标准的音频质量。在这项工作中,我们提出了WaveFM,一个重新参数化的流量匹配模型梅尔频谱图条件的语音合成,旨在提高样本质量和生成速度的扩散声码器。由于梅尔频谱图表示波形的能量分布,WaveFM采用梅尔条件先验分布而不是标准高斯先验,以最大限度地减少合成过程中不必要的传输成本。此外,虽然大多数扩散声码器依赖于一个单一的损失函数,我们认为,纳入辅助损失,包括一个完善的多分辨率STFT损失,可以进一步提高音频质量。为了在不显著降低样本质量的情况下加快推理速度,我们为WaveFM引入了一种定制的一致性蒸馏方法。实验结果表明,我们的模型实现了优越的性能,在质量和效率相比,以前的扩散声码器,同时使波形生成在一个单一的推理步骤。
摘要:Flow matching offers a robust and stable approach to training diffusionmodels. However, directly applying flow matching to neural vocoders can resultin subpar audio quality. In this work, we present WaveFM, a reparameterizedflow matching model for mel-spectrogram conditioned speech synthesis, designedto enhance both sample quality and generation speed for diffusion vocoders.Since mel-spectrograms represent the energy distribution of waveforms, WaveFMadopts a mel-conditioned prior distribution instead of a standard Gaussianprior to minimize unnecessary transportation costs during synthesis. Moreover,while most diffusion vocoders rely on a single loss function, we argue thatincorporating auxiliary losses, including a refined multi-resolution STFT loss,can further improve audio quality. To speed up inference without degradingsample quality significantly, we introduce a tailored consistency distillationmethod for WaveFM. Experiment results demonstrate that our model achievessuperior performance in both quality and efficiency compared to previousdiffusion vocoders, while enabling waveform generation in a single inferencestep.

【12】 Aligning Text-to-Music Evaluation with Human Preferences
标题:将文本与音乐评估与人类偏好保持一致
链接:https://arxiv.org/abs/2503.16669
作者:Yichen Huang,  Zachary Novack,  Koichi Saito,  Jiatong Shi,  Shinji Watanabe,  Yuki Mitsufuji,  John Thickstun,  Chris Donahue
摘要:尽管最近在生成声学文本到音乐(TTM)建模方面取得了重大进展,但对这些模型的鲁棒评估仍然落后,特别是依赖于流行的Fr 'echet音频距离(FAD)。在这项工作中,我们严格研究了用于评估TTM模型的基于参考的分歧度量的设计空间,通过(1)设计四个合成元评估来测量对特定音乐需求的敏感性,以及(2)收集和评估MusicPrefs,这是TTM系统人类偏好的第一个开源数据集。我们发现,不仅是标准的FAD设置不一致的合成和人类的偏好数据,但几乎所有现有的指标未能有效地捕捉desiderata,只有弱相关的人类感知。我们提出了一个新的度量,MAUVE音频发散(MAD),计算表示从一个自我监督的音频嵌入模型。我们发现,这个指标有效地捕捉到了不同的音乐需求(MAD的平均等级相关性为0.84,FAD的平均等级相关性为0.49,并且与MusicPrefs的相关性更强(0.62对0.14)。
摘要:Despite significant recent advances in generative acoustic text-to-music(TTM) modeling, robust evaluation of these models lags behind, relying inparticular on the popular Fr\'echet Audio Distance (FAD). In this work, werigorously study the design space of reference-based divergence metrics forevaluating TTM models through (1) designing four synthetic meta-evaluations tomeasure sensitivity to particular musical desiderata, and (2) collecting andevaluating on MusicPrefs, the first open-source dataset of human preferencesfor TTM systems. We find that not only is the standard FAD setup inconsistenton both synthetic and human preference data, but that nearly all existingmetrics fail to effectively capture desiderata, and are only weakly correlatedwith human perception. We propose a new metric, the MAUVE Audio Divergence(MAD), computed on representations from a self-supervised audio embeddingmodel. We find that this metric effectively captures diverse musical desiderata(average rank correlation 0.84 for MAD vs. 0.49 for FAD and also correlatesmore strongly with MusicPrefs (0.62 vs. 0.14).

【13】 SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for  Super-Aged Seniors
标题:SeniorTalk:针对超龄老年人的具有丰富注释的中国对话数据集
链接:https://arxiv.org/abs/2503.16578
作者:Yang Chen,  Hui Wang,  Shiyao Wang,  Junyang Chen,  Jiabei He,  Jiaming Zhou,  Xi Yang,  Yequan Wang,  Yonghua Lin,  Yong Qin
摘要:虽然语音技术越来越多地服务于老龄化人口,但由于缺乏足够的训练数据来捕捉老年人特有的声音特征,如老视和方言变化,目前的系统表现出显着的性能差距。在现有的老年人语音数据集中,关于超老年人的数据有限,再加上过于简单的记录风格和注释维度,加剧了这个问题。为了解决75岁及以上人群语音数据严重不足的问题,我们引入了SeniorTalk,这是一个经过仔细注释的中文口语对话数据集。该数据集包含来自101个自然对话的55.53小时的演讲,涉及202名参与者,确保了性别,地区和年龄的战略平衡。通过跨多个维度的详细注释,它可以支持广泛的语音任务。我们对说话人验证,说话人日记,语音识别和语音编辑任务进行了广泛的实验,为针对该年龄组的语音技术的发展提供了重要的见解。
摘要:While voice technologies increasingly serve aging populations, currentsystems exhibit significant performance gaps due to inadequate training datacapturing elderly-specific vocal characteristics like presbyphonia anddialectal variations. The limited data available on super-aged individuals inexisting elderly speech datasets, coupled with overly simple recording stylesand annotation dimensions, exacerbates this issue. To address the criticalscarcity of speech data from individuals aged 75 and above, we introduceSeniorTalk, a carefully annotated Chinese spoken dialogue dataset. This datasetcontains 55.53 hours of speech from 101 natural conversations involving 202participants, ensuring a strategic balance across gender, region, and age.Through detailed annotation across multiple dimensions, it can support a widerange of speech tasks. We perform extensive experiments on speakerverification, speaker diarization, speech recognition, and speech editingtasks, offering crucial insights for the development of speech technologiestargeting this age group.

【14】 From Faces to Voices: Learning Hierarchical Representations for  High-quality Video-to-Speech
标题:从面孔到声音:学习高质量视频到语音的分层表示
链接:https://arxiv.org/abs/2503.16956
作者:Ji-Hoon Kim,  Jeongsoo Choi,  Jaehun Kim,  Chaeyoung Jung,  Joon Son Chung
备注:CVPR 2025, demo page: this https URL
摘要:本研究的目的是从无声的说话人脸视频中生成高质量的语音,这一任务也被称为视频到语音合成。视频到语音合成的一个重大挑战在于无声视频和多面语音之间的巨大模态差距。在本文中,我们提出了一种新的视频到语音系统,有效地弥合这一模式的差距,显着提高合成语音的质量。这是通过学习从视频到语音的分层表示来实现的。具体来说,我们逐步将无声视频转换为声学特征空间,通过三个连续的阶段-内容,音色和韵律建模。在每个阶段中,我们将视觉因素-嘴唇运动,面部身份和面部表情-与相应的声学对应物对齐,以确保无缝转换。此外,为了从视觉表示生成逼真和连贯的语音,我们采用了流匹配模型,该模型估计从简单的先验分布到目标语音分布的直接轨迹。大量的实验表明,我们的方法实现了特殊的生成质量与真实的话语,优于现有的方法的显着保证金。
摘要:The objective of this study is to generate high-quality speech from silenttalking face videos, a task also known as video-to-speech synthesis. Asignificant challenge in video-to-speech synthesis lies in the substantialmodality gap between silent video and multi-faceted speech. In this paper, wepropose a novel video-to-speech system that effectively bridges this modalitygap, significantly enhancing the quality of synthesized speech. This isachieved by learning of hierarchical representations from video to speech.Specifically, we gradually transform silent video into acoustic feature spacesthrough three sequential stages -- content, timbre, and prosody modeling. Ineach stage, we align visual factors -- lip movements, face identity, and facialexpressions -- with corresponding acoustic counterparts to ensure the seamlesstransformation. Additionally, to generate realistic and coherent speech fromthe visual representations, we employ a flow matching model that estimatesdirect trajectories from a simple prior distribution to the target speechdistribution. Extensive experiments demonstrate that our method achievesexceptional generation quality comparable to real utterances, outperformingexisting methods by a significant margin.

eess.AS音频处理

【1】 From Faces to Voices: Learning Hierarchical Representations for  High-quality Video-to-Speech
标题:从面孔到声音:学习高质量视频到语音的分层表示
链接:https://arxiv.org/abs/2503.16956
作者:Ji-Hoon Kim,  Jeongsoo Choi,  Jaehun Kim,  Chaeyoung Jung,  Joon Son Chung
备注:CVPR 2025, demo page: this https URL
摘要:本研究的目的是从无声的说话人脸视频中生成高质量的语音,这一任务也被称为视频到语音合成。视频到语音合成中的一个重大挑战在于无声视频和多面语音之间的实质性模态差距。在本文中,我们提出了一种新的视频到语音系统,有效地弥合这一模式的差距,显着提高合成语音的质量。这是通过学习从视频到语音的分层表示来实现的。具体来说,我们逐步将无声视频转换为声学特征空间,通过三个连续的阶段-内容,音色和韵律建模。在每个阶段中,我们将视觉因素-嘴唇运动,面部身份和面部表情-与相应的声学对应物对齐,以确保无缝转换。此外,为了从视觉表示生成逼真和连贯的语音,我们采用了流匹配模型,该模型估计从简单的先验分布到目标语音分布的直接轨迹。大量的实验表明,我们的方法实现了特殊的生成质量与真实的话语,优于现有的方法的显着保证金。
摘要:The objective of this study is to generate high-quality speech from silenttalking face videos, a task also known as video-to-speech synthesis. Asignificant challenge in video-to-speech synthesis lies in the substantialmodality gap between silent video and multi-faceted speech. In this paper, wepropose a novel video-to-speech system that effectively bridges this modalitygap, significantly enhancing the quality of synthesized speech. This isachieved by learning of hierarchical representations from video to speech.Specifically, we gradually transform silent video into acoustic feature spacesthrough three sequential stages -- content, timbre, and prosody modeling. Ineach stage, we align visual factors -- lip movements, face identity, and facialexpressions -- with corresponding acoustic counterparts to ensure the seamlesstransformation. Additionally, to generate realistic and coherent speech fromthe visual representations, we employ a flow matching model that estimatesdirect trajectories from a simple prior distribution to the target speechdistribution. Extensive experiments demonstrate that our method achievesexceptional generation quality comparable to real utterances, outperformingexisting methods by a significant margin.

【2】 Learning disentangled representations for instrument-based music  similarity
标题:学习基于乐器的音乐相似性的解开表示
链接:https://arxiv.org/abs/2503.17281
作者:Yuka Hashizume,  Li Li,  Atsushi Miyashita,  Tomoki Toda
备注:arXiv admin note: text overlap with arXiv:2404.06682
摘要:灵活的推荐和检索系统需要音乐片段的多个部分元素方面的音乐相似性,以允许用户选择他们想要关注的元素。使用具有单独乐器信号的多个网络进行音乐相似性学习的方法是有效的,但面临的问题是,使用每个干净的乐器信号作为检索系统的查询是不切实际的,并且使用分离的乐器声音降低了准确性因为艺术品。在本文中,我们提出了基于乐器部分的音乐相似性学习与一个单一的网络,以混合的声音作为输入,而不是个别乐器的声音。具体来说,我们为每个仪器设计了一个具有解纠缠维度的单一相似性嵌入空间,由条件相似性网络提取,该网络使用带有掩码的三重丢失进行训练。实验结果表明:(1)在对乐器进行评价时,与使用单独的声音作为输入的网络相比,该方法可以获得更准确的特征表示;(2)每个子嵌入空间都可以保持相应乐器的特征,以及(3)通过所提出的方法选择关注每个乐器声音的相似音乐作品可以获得人的接受,特别是当关注音色时。
摘要:A flexible recommendation and retrieval system requires music similarity interms of multiple partial elements of musical pieces to allow users to selectthe element they want to focus on. A method for music similarity learning usingmultiple networks with individual instrumental signals is effective but facesthe problem that using each clean instrumental signal as a query is impracticalfor retrieval systems and using separated instrumental sounds reduces accuracyowing to artifacts. In this paper, we present instrumental-part-based musicsimilarity learning with a single network that takes mixed sounds as inputinstead of individual instrumental sounds. Specifically, we designed a singlesimilarity embedding space with disentangled dimensions for each instrument,extracted by Conditional Similarity Networks, which are trained using thetriplet loss with masks. Experimental results showed that (1) the proposedmethod can obtain more accurate feature representation than using individualnetworks using separated sounds as input in the evaluation of an instrumentthat had low accuracy, (2) each sub-embedding space can hold thecharacteristics of the corresponding instrument, and (3) the selection ofsimilar musical pieces focusing on each instrumental sound by the proposedmethod can obtain human acceptance, especially when focusing on timbre.

【3】 Imagine to Hear: Auditory Knowledge Generation can be an Effective  Assistant for Language Models
标题:想象聆听:听觉知识生成可以成为语言模型的有效助手
链接:https://arxiv.org/abs/2503.16853
作者:Suho Yoo,  Hyunjong Ok,  Jaeho Lee
备注:Preprint
摘要:在纯文本语料库上预训练的语言模型通常难以完成需要听觉常识知识的任务。以前的工作解决了这个问题,通过增强语言模型从外部音频数据库中检索知识。这种方法有几个局限性,如数据库中可能缺乏相关音频,以及与构建和查询数据库相关的高成本。为了解决这些问题,我们提出了想象听到,一种新的方法,动态生成听觉知识生成模型。我们的框架检测多个音频相关的文本跨度从给定的提示,并生成相应的听觉知识。我们开发了几种机制来有效地处理多种听觉知识,包括基于CLAP的拒绝采样器和语言音频融合模块。我们的实验表明,我们的方法在AuditoryBench上实现了最先进的性能,而不依赖于外部数据库,突出了我们基于生成的方法的有效性。
摘要:Language models pretrained on text-only corpora often struggle with tasksthat require auditory commonsense knowledge. Previous work addresses thisproblem by augmenting the language model to retrieve knowledge from externalaudio databases. This approach has several limitations, such as the potentiallack of relevant audio in databases and the high costs associated withconstructing and querying the databases. To address these issues, we proposeImagine to Hear, a novel approach that dynamically generates auditory knowledgeusing generative models. Our framework detects multiple audio-related textualspans from the given prompt and generates corresponding auditory knowledge. Wedevelop several mechanisms to efficiently process multiple auditory knowledge,including a CLAP-based rejection sampler and a language-audio fusion module.Our experiments show that our method achieves state-of-the-art performance onAuditoryBench without relying on external databases, highlighting theeffectiveness of our generation-based approach.

【4】 SeniorTalk: A Chinese Conversation Dataset with Rich Annotations for  Super-Aged Seniors
标题:SeniorTalk:针对超龄老年人的具有丰富注释的中国对话数据集
链接:https://arxiv.org/abs/2503.16578
作者:Yang Chen,  Hui Wang,  Shiyao Wang,  Junyang Chen,  Jiabei He,  Jiaming Zhou,  Xi Yang,  Yequan Wang,  Yonghua Lin,  Yong Qin
摘要:虽然语音技术越来越多地服务于老龄化人口,但由于缺乏足够的训练数据来捕捉老年人特有的声音特征,如老视和方言变化,目前的系统表现出显着的性能差距。在现有的老年人语音数据集中,关于超老年人的数据有限,再加上过于简单的记录风格和注释维度,加剧了这个问题。为了解决75岁及以上人群语音数据严重不足的问题,我们引入了SeniorTalk,这是一个经过仔细注释的中文口语对话数据集。该数据集包含来自101个自然对话的55.53小时的演讲,涉及202名参与者,确保了性别,地区和年龄的战略平衡。通过跨多个维度的详细注释,它可以支持广泛的语音任务。我们对说话人验证,说话人日记,语音识别和语音编辑任务进行了广泛的实验,为针对该年龄组的语音技术的发展提供了重要的见解。
摘要:While voice technologies increasingly serve aging populations, currentsystems exhibit significant performance gaps due to inadequate training datacapturing elderly-specific vocal characteristics like presbyphonia anddialectal variations. The limited data available on super-aged individuals inexisting elderly speech datasets, coupled with overly simple recording stylesand annotation dimensions, exacerbates this issue. To address the criticalscarcity of speech data from individuals aged 75 and above, we introduceSeniorTalk, a carefully annotated Chinese spoken dialogue dataset. This datasetcontains 55.53 hours of speech from 101 natural conversations involving 202participants, ensuring a strategic balance across gender, region, and age.Through detailed annotation across multiple dimensions, it can support a widerange of speech tasks. We perform extensive experiments on speakerverification, speaker diarization, speech recognition, and speech editingtasks, offering crucial insights for the development of speech technologiestargeting this age group.

机器翻译由腾讯交互翻译提供,仅供参考