今日论文合集:cs.SD语音8篇,eess.AS音频处理10篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for  Speech, Music, and Sound

标题:Meta音频盒美学:语音、音乐和声音的统一自动质量评估
链接:https://arxiv.org/abs/2502.05139
作者:Andros Tjandra,  Yi-Chiao Wu,  Baishan Guo,  John Hoffman,  Brian Ellis,  Apoorv Vyas,  Bowen Shi,  Sanyuan Chen,  Matt Le,  Nick Zacharov,  Carleigh Wood,  Ann Lee,  Wei-Ning Hsu
备注:Repository: this https URL Website: this https URL
摘要:音频美学的量化仍然是音频处理中的一个复杂挑战,主要是由于其主观性,这是受人类感知和文化背景的影响。传统的方法通常依赖于人类听众进行评估,导致不一致和高资源需求。本文解决了对能够在无需人为干预的情况下预测音频美学的自动化系统日益增长的需求。这样的系统对于数据过滤、伪标记大型数据集和评估生成音频模型等应用至关重要,特别是当这些模型变得越来越复杂时。在这项工作中,我们引入了一种新的方法,音频美学评价提出了新的注释准则,分解成四个不同的轴人类听力的角度。我们开发并训练无参考的单项预测模型,提供更细致入微的音频质量评估。我们的模型对人类平均意见评分(MOS)和现有的方法进行了评估,表现出相当或优越的性能。这项研究不仅推进了音频美学领域,还提供了开源模型和数据集,以促进未来的工作和基准测试。我们发布我们的代码和预训练模型:https://github.com/facebookresearch/audiobox-aesthetics
摘要:The quantification of audio aesthetics remains a complex challenge in audioprocessing, primarily due to its subjective nature, which is influenced byhuman perception and cultural context. Traditional methods often depend onhuman listeners for evaluation, leading to inconsistencies and high resourcedemands. This paper addresses the growing need for automated systems capable ofpredicting audio aesthetics without human intervention. Such systems arecrucial for applications like data filtering, pseudo-labeling large datasets,and evaluating generative audio models, especially as these models become moresophisticated. In this work, we introduce a novel approach to audio aestheticevaluation by proposing new annotation guidelines that decompose humanlistening perspectives into four distinct axes. We develop and trainno-reference, per-item prediction models that offer a more nuanced assessmentof audio quality. Our models are evaluated against human mean opinion scores(MOS) and existing methods, demonstrating comparable or superior performance.This research not only advances the field of audio aesthetics but also providesopen-source models and datasets to facilitate future work and benchmarking. Werelease our code and pre-trained model at:https://github.com/facebookresearch/audiobox-aesthetics

【2】 Latent Swap Joint Diffusion for Long-Form Audio Generation
标题:用于长格式音频生成的潜在交换联合扩散
链接:https://arxiv.org/abs/2502.05130
作者:Yusheng Dai,  Chenxi Wang,  Chang Li,  Chen Wang,  Jun Du,  Kewei Li,  Ruoyu Wang,  Jiefeng Ma,  Lei Sun,  Jianqing Gao
摘要:以前使用全局视图扩散或迭代生成的长格式音频生成工作需要大量的训练或推理成本。虽然用于全景生成的多视图联合扩散的最新进展提供了一种有效的选择,但它们难以产生具有严重重叠失真和高交叉视图一致性成本的频谱。我们最初通过潜在地图的连接继承来探索这种现象,并发现平均操作过度平滑潜在地图的高频分量。为了解决这些问题,我们提出了Swap Forward(SaFa),这是一种帧级潜在交换框架,它以仅向前的方式对多个扩散进行优化,以产生具有更多频谱细节的全局相干长音频。其核心是在相邻视图之间应用双向Self-Loop Latent Swap,利用逐步扩散轨迹自适应地增强高频分量,而不干扰低频分量。此外,为了确保跨视图的一致性,在早期阶段,在每个子视图的参考区域和非重叠区域之间应用单向参考引导的潜在交换,提供集中的轨迹引导。定量和定性实验表明,SaFa显着优于现有的联合扩散方法,甚至基于训练的长音频生成模型。此外,我们发现它也很好地适应全景生成,实现了可比的最先进的性能,更高的效率和模型的泛化能力。项目页面可在https://swapforward.github.io/上找到。
摘要:Previous work on long-form audio generation using global-view diffusion oriterative generation demands significant training or inference costs. Whilerecent advancements in multi-view joint diffusion for panoramic generationprovide an efficient option, they struggle with spectrum generation with severeoverlap distortions and high cross-view consistency costs. We initially explorethis phenomenon through the connectivity inheritance of latent maps and uncoverthat averaging operations excessively smooth the high-frequency components ofthe latent map. To address these issues, we propose Swap Forward (SaFa), aframe-level latent swap framework that synchronizes multiple diffusions toproduce a globally coherent long audio with more spectrum details in aforward-only manner. At its core, the bidirectional Self-Loop Latent Swap isapplied between adjacent views, leveraging stepwise diffusion trajectory toadaptively enhance high-frequency components without disrupting low-frequencycomponents. Furthermore, to ensure cross-view consistency, the unidirectionalReference-Guided Latent Swap is applied between the reference and thenon-overlap regions of each subview during the early stages, providingcentralized trajectory guidance. Quantitative and qualitative experimentsdemonstrate that SaFa significantly outperforms existing joint diffusionmethods and even training-based long audio generation models. Moreover, we findthat it also adapts well to panoramic generation, achieving comparablestate-of-the-art performance with greater efficiency and modelgeneralizability. Project page is available at https://swapforward.github.io/.

【3】 Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning  and Language Identification for Improved Low-resource Performance
标题:评估标准和方言弗里斯兰ASB:多语言微调和语言识别以提高低资源性能
链接:https://arxiv.org/abs/2502.04883
作者:Reihaneh Amooie,  Wietse de Vries,  Yun Hao,  Jelske Dijkstra,  Matt Coler,  Martijn Wieling
摘要:由于缺乏足够的标记数据,低资源语言的自动语音识别(ASR)性能仍然远远落后于英语等高资源语言。最先进的方法部署了自监督迁移学习,其中在大量数据上预训练的模型使用目标低资源语言中的少量标记数据进行微调。在本文中,我们提出并研究了一种微调基于SSL的模型的方法,以提高弗里斯兰语及其地区方言(粘土弗里斯兰语,木材弗里斯兰语和南弗里斯兰语)的性能。我们表明,弗里斯兰语ASR性能可以提高使用多语言(弗里斯兰语,荷兰语,英语和德语)微调数据和辅助语言识别任务。此外,我们的研究结果表明,方言语音的性能受到很大的影响,而且,重要的是,这种影响是缓和的启发式方法用于收集方言数据。我们的研究结果还特别表明,仅仅依靠标准语言数据进行ASR评估可能会低估现实世界的表现,特别是在方言差异很大的语言中。
摘要:Automatic Speech Recognition (ASR) performance for low-resource languages isstill far behind that of higher-resource languages such as English, due to alack of sufficient labeled data. State-of-the-art methods deployself-supervised transfer learning where a model pre-trained on large amounts ofdata is fine-tuned using little labeled data in a target low-resource language.In this paper, we present and examine a method for fine-tuning an SSL-basedmodel in order to improve the performance for Frisian and its regional dialects(Clay Frisian, Wood Frisian, and South Frisian). We show that Frisian ASRperformance can be improved by using multilingual (Frisian, Dutch, English andGerman) fine-tuning data and an auxiliary language identification task. Inaddition, our findings show that performance on dialectal speech sufferssubstantially, and, importantly, that this effect is moderated by theelicitation approach used to collect the dialectal data. Our findings alsoparticularly suggest that relying solely on standard language data for ASRevaluation may underestimate real-world performance, particularly in languageswith substantial dialectal variation.

【4】 Singing Voice Conversion with Accompaniment Using Self-Supervised  Representation-Based Melody Features
标题:使用自我监督的基于表达的旋律特征进行伴奏演唱声音转换
链接:https://arxiv.org/abs/2502.04722
作者:Wei Chen,  Binzhu Sha,  Jing Yang,  Zhuo Wang,  Fan Fan,  Zhiyong Wu
备注:Accepted by ICASSP2025
摘要:在歌唱嗓音转换中,旋律的保持是至关重要的。然而,在许多场景中,音频通常伴随着背景音乐(BGM),这会导致音频失真并干扰旋律和其他关键特征的提取,从而显著降低SVC性能。以前的方法试图通过使用更强大的基于神经网络的旋律提取器来解决这个问题,但是在复杂伴奏的情况下,它们的性能急剧下降。其他方法涉及在转换之前执行源分离,但这通常会引入明显的伪像,导致转换质量显著下降并增加用户的运营成本。为了解决这些问题,我们引入了一种新的SVC方法,该方法使用基于自我监督表示的旋律特征,以提高旋律建模的准确性存在的BGM。在我们的实验中,我们比较了不同的自我监督学习(SSL)模型的旋律提取的有效性,并首次探索SSL如何有利于旋律提取的任务。实验结果表明,我们提出的SVC模型显着优于现有的基线方法在旋律的准确性,并显示出更高的相似性和自然度在主观和客观评价在嘈杂和干净的音频环境。
摘要:Melody preservation is crucial in singing voice conversion (SVC). However, inmany scenarios, audio is often accompanied with background music (BGM), whichcan cause audio distortion and interfere with the extraction of melody andother key features, significantly degrading SVC performance. Previous methodshave attempted to address this by using more robust neural network-based melodyextractors, but their performance drops sharply in the presence of complexaccompaniment. Other approaches involve performing source separation beforeconversion, but this often introduces noticeable artifacts, leading to asignificant drop in conversion quality and increasing the user's operationalcosts. To address these issues, we introduce a novel SVC method that usesself-supervised representation-based melody features to improve melody modelingaccuracy in the presence of BGM. In our experiments, we compare theeffectiveness of different self-supervised learning (SSL) models for melodyextraction and explore for the first time how SSL benefits the task of melodyextraction. The experimental results demonstrate that our proposed SVC modelsignificantly outperforms existing baseline methods in terms of melody accuracyand shows higher similarity and naturalness in both subjective and objectiveevaluations across noisy and clean audio environments.

【5】 Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement
标题:语音增强的动态频率自适应知识提炼
链接:https://arxiv.org/abs/2502.04711
作者:Xihao Yuan,  Siqi Liu,  Hanting Chen,  Lu Zhou,  Jian Li,  Jie Hu
备注:5 pages, 2 figures, accepted by ICASSP2025
摘要:基于深度学习的语音增强(SE)模型最近的表现优于传统技术,但由于计算和内存需求高,它们在资源受限设备上的部署仍然具有挑战性。本文提出了一种新的动态频率自适应知识提取(DFKD)方法,有效地压缩SE模型。我们的方法动态评估模型的输出,区分高频和低频分量,并调整学习目标以满足不同频带的独特要求,利用SE任务的固有特性。为了评估DFKD的有效性,我们在三个最先进的模型上进行了实验:DCCRN,ConTasNet和DPTNet。结果表明,我们的方法不仅显着提高了压缩模型(学生模型)的性能,但也优于其他基于逻辑的知识蒸馏方法,专门为SE任务。
摘要:Deep learning-based speech enhancement (SE) models have recently outperformedtraditional techniques, yet their deployment on resource-constrained devicesremains challenging due to high computational and memory demands. This paperintroduces a novel dynamic frequency-adaptive knowledge distillation (DFKD)approach to effectively compress SE models. Our method dynamically assesses themodel's output, distinguishing between high and low-frequency components, andadapts the learning objectives to meet the unique requirements of differentfrequency bands, capitalizing on the SE task's inherent characteristics. Toevaluate the DFKD's efficacy, we conducted experiments on threestate-of-the-art models: DCCRN, ConTasNet, and DPTNet. The results demonstratethat our method not only significantly enhances the performance of thecompressed model (student model) but also surpasses other logit-based knowledgedistillation methods specifically for SE tasks.

【6】 ImprovNet: Generating Controllable Musical Improvisations with Iterative  Corruption Refinement
标题:DeliverNet:通过迭代腐败精炼生成可控的音乐即兴表演
链接:https://arxiv.org/abs/2502.04522
作者:Keshav Bhandari,  Sungkyun Chang,  Tongyu Lu,  Fareza R. Enus,  Louis B. Bradshaw,  Dorien Herremans,  Simon Colton
备注:10 pages, 6 figures
摘要:深度学习在各个领域的风格转移方面取得了显着进展,为创造性内容的生成提供了新的可能性。然而,在符号音乐领域,由于数据集有限,特别是对于爵士乐等流派,以及缺乏可以处理多个音乐生成任务的统一模型,为完整的音乐作品生成可控且富有表现力的表演级风格转移仍然具有挑战性。本文介绍了一种基于transformer的架构,它通过自我监督的腐败细化训练策略生成富有表现力和可控的音乐即兴创作。CNET在一个模型中统一了多种功能:它可以执行跨流派和流派内的即兴创作,协调旋律与特定流派的风格,并执行简短的提示延续和填充任务。该模型的迭代生成框架允许用户控制风格转换的程度和与原始组成的结构相似性。客观和主观的评估表明,在生成音乐连贯的即兴,同时保持与原始作品的结构关系的有效性。该模型优于预期音乐Transformer在短期延续和填充任务,并成功地实现了可识别的流派转换,与79%的参与者正确识别爵士风格的即兴。我们的代码和演示页面可以在https://github.com/keshavbhandari/improvnet上找到。
摘要:Deep learning has enabled remarkable advances in style transfer acrossvarious domains, offering new possibilities for creative content generation.However, in the realm of symbolic music, generating controllable and expressiveperformance-level style transfers for complete musical works remainschallenging due to limited datasets, especially for genres such as jazz, andthe lack of unified models that can handle multiple music generation tasks.This paper presents ImprovNet, a transformer-based architecture that generatesexpressive and controllable musical improvisations through a self-supervisedcorruption-refinement training strategy. ImprovNet unifies multiplecapabilities within a single model: it can perform cross-genre and intra-genreimprovisations, harmonize melodies with genre-specific styles, and executeshort prompt continuation and infilling tasks. The model's iterative generationframework allows users to control the degree of style transfer and structuralsimilarity to the original composition. Objective and subjective evaluationsdemonstrate ImprovNet's effectiveness in generating musically coherentimprovisations while maintaining structural relationships with the originalpieces. The model outperforms Anticipatory Music Transformer in shortcontinuation and infilling tasks and successfully achieves recognizable genreconversion, with 79\% of participants correctly identifying jazz-styleimprovisations. Our code and demo page can be found athttps://github.com/keshavbhandari/improvnet.

【7】 ADIFF: Explaining audio difference using natural language
标题:ADiff:使用自然语言解释音频差异
链接:https://arxiv.org/abs/2502.04476
作者:Soham Deshmukh,  Shuo Han,  Rita Singh,  Bhiksha Raj
备注:Accepted at ICLR 2025. Dataset and checkpoints are available at: this https URL
摘要:理解和解释录音之间的差异对于音频取证、质量评估和音频生成等领域至关重要。这涉及识别和描述音频事件,声学场景,信号特征及其对听众的情感影响。本文首先全面研究了解释音频差异的任务,然后提出了任务的基准,基线。首先,我们提出了两个新的数据集,用于音频差异解释,这些数据集源自AudioCaps和Clotho音频字幕数据集。使用大语言模型(LLM),我们生成三个层次的差异解释:(1)音频事件和对象的简洁描述,(2)关于音频事件,声学场景和信号属性的简短句子,以及(3)包括语义和听众情绪的综合解释。对于基线,我们使用前缀调优,其中来自两个音频文件的音频嵌入用于提示冻结的语言模型。我们的实证分析和消融研究表明,天真的基线努力区分感知相似的声音,并产生详细的第3层解释。为了解决这些局限性,我们提出了ADIFF,它引入了交叉投影模块,位置字幕和三步训练过程,以增强模型产生详细解释的能力。我们使用客观指标和人工评估来评估我们的模型,并显示我们的模型增强导致了朴素基线和SoTA音频语言模型(ALM)Qwen Audio性能的显着改善。最后,我们进行了多个消融研究,以研究交叉投影,语言模型参数,位置字幕,第三阶段微调的影响,并提出我们的研究结果。我们的基准,发现和强大的基线为音频差异的细微和人性化解释铺平了道路。
摘要:Understanding and explaining differences between audio recordings is crucialfor fields like audio forensics, quality assessment, and audio generation. Thisinvolves identifying and describing audio events, acoustic scenes, signalcharacteristics, and their emotional impact on listeners. This paper stands outas the first work to comprehensively study the task of explaining audiodifferences and then propose benchmark, baselines for the task. First, wepresent two new datasets for audio difference explanation derived from theAudioCaps and Clotho audio captioning datasets. Using Large Language Models(LLMs), we generate three levels of difference explanations: (1) concisedescriptions of audio events and objects, (2) brief sentences about audioevents, acoustic scenes, and signal properties, and (3) comprehensiveexplanations that include semantics and listener emotions. For the baseline, weuse prefix tuning where audio embeddings from two audio files are used toprompt a frozen language model. Our empirical analysis and ablation studiesreveal that the naive baseline struggles to distinguish perceptually similarsounds and generate detailed tier 3 explanations. To address these limitations,we propose ADIFF, which introduces a cross-projection module, positioncaptioning, and a three-step training process to enhance the model's ability toproduce detailed explanations. We evaluate our model using objective metricsand human evaluation and show our model enhancements lead to significantimprovements in performance over naive baseline and SoTA Audio-Language Model(ALM) Qwen Audio. Lastly, we conduct multiple ablation studies to study theeffects of cross-projection, language model parameters, position captioning,third stage fine-tuning, and present our findings. Our benchmarks, findings,and strong baseline pave the way for nuanced and human-like explanations ofaudio differences.

【8】 FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks
标题:FocalCodec:通过Focal调制网络进行低比特率语音编码
链接:https://arxiv.org/abs/2502.04465
作者:Luca Della Libera,  Francesco Paissan,  Cem Subakan,  Mirco Ravanelli
备注:18 pages
摘要:大型语言模型通过在大规模数据集上进行自我监督预训练,彻底改变了自然语言处理。受这一成功的启发,研究人员探索了通过使用神经音频编解码器将连续音频离散化为令牌来使这些方法适应语音。然而,现有的方法面临着限制,包括高比特率,语义或声学信息的丢失,以及在试图捕获两者时对多码本设计的依赖,这增加了下游任务的架构复杂性。为了解决这些挑战,我们介绍了FocalCodec,一个有效的低比特率编解码器的基础上,利用一个单一的二进制码本压缩语音0.16和0.65 kbps的焦点调制。FocalCodec在语音再合成和语音转换方面提供了具有竞争力的性能,其比特率低于当前最先进的技术水平,同时有效地处理多语言语音和嘈杂环境。对下游任务的评估表明,FocalCodec成功地保留了足够的语义和声学信息,同时也非常适合生成式建模。演示示例、代码和检查点可在https://lucadellalib.github.io/focalcodec-web/上获得。
摘要:Large language models have revolutionized natural language processing throughself-supervised pretraining on massive datasets. Inspired by this success,researchers have explored adapting these methods to speech by discretizingcontinuous audio into tokens using neural audio codecs. However, existingapproaches face limitations, including high bitrates, the loss of eithersemantic or acoustic information, and the reliance on multi-codebook designswhen trying to capture both, which increases architectural complexity fordownstream tasks. To address these challenges, we introduce FocalCodec, anefficient low-bitrate codec based on focal modulation that utilizes a singlebinary codebook to compress speech between 0.16 and 0.65 kbps. FocalCodecdelivers competitive performance in speech resynthesis and voice conversion atlower bitrates than the current state-of-the-art, while effectively handlingmultilingual speech and noisy environments. Evaluation on downstream tasksshows that FocalCodec successfully preserves sufficient semantic and acousticinformation, while also being well-suited for generative modeling. Demosamples, code and checkpoints are available athttps://lucadellalib.github.io/focalcodec-web/.

eess.AS音频处理

【1】 Efficient Evaluation of Quantization-Effects in Neural Codecs
标题:神经编解码器量化效应的有效评估
链接:https://arxiv.org/abs/2502.04770
作者:Wolfgang Mack,  Ahmed Mustafa,  Rafał Łaganowski,  Samer Hijazy
摘要:神经编解码器包括编码器、量化器和解码器,能够以极低的比特率进行信号传输。训练这些系统需要像直通估计器、软到硬退火或统计量化器仿真这样的技术,以允许量化器上的非零梯度。评估神经编解码器中量化的效果,就像梯度传递技术对整个系统的影响一样,由于训练需求和缺乏负担得起的可靠指标,通常是昂贵和耗时的。本文提出了一种有效的评估框架,神经编解码器使用模拟数据与一个定义的比特数和低复杂度的神经编码器/解码器,以模拟在较大的网络中的非线性行为。我们的系统在训练时间、计算和硬件要求方面非常高效,使我们能够发现神经编解码器中的不同行为。根据我们的研究结果,我们提出了一个修改,以稳定直通估计器的训练。我们验证我们的研究结果对内部神经音频编解码器和对国家的最先进的音频编解码器。
摘要:Neural codecs, comprising an encoder, quantizer, and decoder, enable signaltransmission at exceptionally low bitrates. Training these systems requirestechniques like the straight-through estimator, soft-to-hard annealing, orstatistical quantizer emulation to allow a non-zero gradient across thequantizer. Evaluating the effect of quantization in neural codecs, like theinfluence of gradient passing techniques on the whole system, is often costlyand time-consuming due to training demands and the lack of affordable andreliable metrics. This paper proposes an efficient evaluation framework forneural codecs using simulated data with a defined number of bits andlow-complexity neural encoders/decoders to emulate the non-linear behavior inlarger networks. Our system is highly efficient in terms of training time andcomputational and hardware requirements, allowing us to uncover distinctbehaviors in neural codecs. We propose a modification to stabilize trainingwith the straight-through estimator based on our findings. We validate ourfindings against an internal neural audio codec and against thestate-of-the-art descript-audio-codec.

【2】 GenVC: Self-Supervised Zero-Shot Voice Conversion
标题:GenVC:自我监督的Zero-Shot语音转换
链接:https://arxiv.org/abs/2502.04519
作者:Zexin Cai,  Henry Li Xinyuan,  Ashi Garg,  Leibny Paola García-Perera,  Kevin Duh,  Sanjeev Khudanpur,  Matthew Wiesner,  Nicholas Andrews
摘要:最近,Zero-shot语音转换取得了实质性的进展,但许多模型仍然依赖于外部监督系统来解开说话人身份和语言内容。此外,目前的方法通常使用并行转换,其中转换后的语音继承源话语的时间结构,限制说话人的相似性和隐私。为了克服这些限制,我们引入了GenVC,一个生成的zero-shot语音转换模型。GenVC学习以自我监督的方式理清语言内容和说话者风格,无需外部模型,并在大型无标签数据集上进行有效训练。实验结果表明,GenVC实现了国家的最先进的说话人相似性,同时保持与领先的方法竞争的自然。它的自回归生成也允许转换后的语音偏离源话语的时间结构。此功能使GenVC在语音匿名化方面非常有效,因为它最大限度地减少了对源韵律和说话者特征的保留,从而增强了隐私保护。
摘要:Zero-shot voice conversion has recently made substantial progress, but manymodels still depend on external supervised systems to disentangle speakeridentity and linguistic content. Furthermore, current methods often useparallel conversion, where the converted speech inherits the source utterance'stemporal structure, restricting speaker similarity and privacy. To overcomethese limitations, we introduce GenVC, a generative zero-shot voice conversionmodel. GenVC learns to disentangle linguistic content and speaker style in aself-supervised manner, eliminating the need for external models and enablingefficient training on large, unlabeled datasets. Experimental results show thatGenVC achieves state-of-the-art speaker similarity while maintainingnaturalness competitive with leading approaches. Its autoregressive generationalso allows the converted speech to deviate from the source utterance'stemporal structure. This feature makes GenVC highly effective for voiceanonymization, as it minimizes the preservation of source prosody and speakercharacteristics, enhancing privacy protection.

【3】 Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for  Speech, Music, and Sound
标题:Meta音频盒美学:语音、音乐和声音的统一自动质量评估
链接:https://arxiv.org/abs/2502.05139
作者:Andros Tjandra,  Yi-Chiao Wu,  Baishan Guo,  John Hoffman,  Brian Ellis,  Apoorv Vyas,  Bowen Shi,  Sanyuan Chen,  Matt Le,  Nick Zacharov,  Carleigh Wood,  Ann Lee,  Wei-Ning Hsu
备注:Repository: this https URL Website: this https URL
摘要:音频美学的量化仍然是音频处理中的一个复杂挑战,主要是由于其主观性,这是受人类感知和文化背景的影响。传统的方法通常依赖于人类听众进行评估,导致不一致和高资源需求。本文解决了对能够在无需人为干预的情况下预测音频美学的自动化系统日益增长的需求。这样的系统对于数据过滤、伪标记大型数据集和评估生成音频模型等应用至关重要,特别是当这些模型变得越来越复杂时。在这项工作中,我们引入了一种新的方法,音频美学评价提出了新的注释准则,分解成四个不同的轴人类听力的角度。我们开发并训练无参考的单项预测模型,提供更细致入微的音频质量评估。我们的模型对人类平均意见评分(MOS)和现有的方法进行了评估,表现出相当或优越的性能。这项研究不仅推进了音频美学领域,还提供了开源模型和数据集,以促进未来的工作和基准测试。我们发布我们的代码和预训练模型:https://github.com/facebookresearch/audiobox-aesthetics
摘要:The quantification of audio aesthetics remains a complex challenge in audioprocessing, primarily due to its subjective nature, which is influenced byhuman perception and cultural context. Traditional methods often depend onhuman listeners for evaluation, leading to inconsistencies and high resourcedemands. This paper addresses the growing need for automated systems capable ofpredicting audio aesthetics without human intervention. Such systems arecrucial for applications like data filtering, pseudo-labeling large datasets,and evaluating generative audio models, especially as these models become moresophisticated. In this work, we introduce a novel approach to audio aestheticevaluation by proposing new annotation guidelines that decompose humanlistening perspectives into four distinct axes. We develop and trainno-reference, per-item prediction models that offer a more nuanced assessmentof audio quality. Our models are evaluated against human mean opinion scores(MOS) and existing methods, demonstrating comparable or superior performance.This research not only advances the field of audio aesthetics but also providesopen-source models and datasets to facilitate future work and benchmarking. Werelease our code and pre-trained model at:https://github.com/facebookresearch/audiobox-aesthetics

【4】 Latent Swap Joint Diffusion for Long-Form Audio Generation
标题:用于长格式音频生成的潜在交换联合扩散
链接:https://arxiv.org/abs/2502.05130
作者:Yusheng Dai,  Chenxi Wang,  Chang Li,  Chen Wang,  Jun Du,  Kewei Li,  Ruoyu Wang,  Jiefeng Ma,  Lei Sun,  Jianqing Gao
摘要:以前使用全局视图扩散或迭代生成的长格式音频生成工作需要大量的训练或推理成本。虽然用于全景生成的多视图联合扩散的最新进展提供了一种有效的选择,但它们难以产生具有严重重叠失真和高交叉视图一致性成本的频谱。我们最初通过潜在地图的连接继承来探索这种现象,并发现平均操作过度平滑潜在地图的高频分量。为了解决这些问题,我们提出了Swap Forward(SaFa),这是一种帧级潜在交换框架,它以仅向前的方式对多个扩散进行优化,以产生具有更多频谱细节的全局相干长音频。其核心是在相邻视图之间应用双向Self-Loop Latent Swap,利用逐步扩散轨迹自适应地增强高频分量,而不干扰低频分量。此外,为了确保跨视图的一致性,在早期阶段,在每个子视图的参考区域和非重叠区域之间应用单向参考引导的潜在交换,提供集中的轨迹引导。定量和定性实验表明,SaFa显着优于现有的联合扩散方法,甚至基于训练的长音频生成模型。此外,我们发现它也很好地适应全景生成,实现了可比的最先进的性能,更高的效率和模型的泛化能力。项目"页可在https://swapforward.github.io/上找到。
摘要:Previous work on long-form audio generation using global-view diffusion oriterative generation demands significant training or inference costs. Whilerecent advancements in multi-view joint diffusion for panoramic generationprovide an efficient option, they struggle with spectrum generation with severeoverlap distortions and high cross-view consistency costs. We initially explorethis phenomenon through the connectivity inheritance of latent maps and uncoverthat averaging operations excessively smooth the high-frequency components ofthe latent map. To address these issues, we propose Swap Forward (SaFa), aframe-level latent swap framework that synchronizes multiple diffusions toproduce a globally coherent long audio with more spectrum details in aforward-only manner. At its core, the bidirectional Self-Loop Latent Swap isapplied between adjacent views, leveraging stepwise diffusion trajectory toadaptively enhance high-frequency components without disrupting low-frequencycomponents. Furthermore, to ensure cross-view consistency, the unidirectionalReference-Guided Latent Swap is applied between the reference and thenon-overlap regions of each subview during the early stages, providingcentralized trajectory guidance. Quantitative and qualitative experimentsdemonstrate that SaFa significantly outperforms existing joint diffusionmethods and even training-based long audio generation models. Moreover, we findthat it also adapts well to panoramic generation, achieving comparablestate-of-the-art performance with greater efficiency and modelgeneralizability. Project page is available at https://swapforward.github.io/.

【5】 Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning  and Language Identification for Improved Low-resource Performance
标题:评估标准和方言弗里斯兰ASB:多语言微调和语言识别以提高低资源性能
链接:https://arxiv.org/abs/2502.04883
作者:Reihaneh Amooie,  Wietse de Vries,  Yun Hao,  Jelske Dijkstra,  Matt Coler,  Martijn Wieling
摘要:由于缺乏足够的标记数据,低资源语言的自动语音识别(ASR)性能仍然远远落后于英语等高资源语言。最先进的方法部署了自监督迁移学习,其中在大量数据上预训练的模型使用目标低资源语言中的少量标记数据进行微调。在本文中,我们提出并研究了一种微调基于SSL的模型的方法,以提高弗里斯兰语及其地区方言(粘土弗里斯兰语,木材弗里斯兰语和南弗里斯兰语)的性能。我们表明,弗里斯兰语ASR性能可以提高使用多语言(弗里斯兰语,荷兰语,英语和德语)微调数据和辅助语言识别任务。此外,我们的研究结果表明,方言语音的性能受到很大的影响,而且,重要的是,这种影响是缓和的启发式方法用于收集方言数据。我们的研究结果还特别表明,仅仅依靠标准语言数据进行ASR评估可能会低估现实世界的表现,特别是在方言差异很大的语言中。
摘要:Automatic Speech Recognition (ASR) performance for low-resource languages isstill far behind that of higher-resource languages such as English, due to alack of sufficient labeled data. State-of-the-art methods deployself-supervised transfer learning where a model pre-trained on large amounts ofdata is fine-tuned using little labeled data in a target low-resource language.In this paper, we present and examine a method for fine-tuning an SSL-basedmodel in order to improve the performance for Frisian and its regional dialects(Clay Frisian, Wood Frisian, and South Frisian). We show that Frisian ASRperformance can be improved by using multilingual (Frisian, Dutch, English andGerman) fine-tuning data and an auxiliary language identification task. Inaddition, our findings show that performance on dialectal speech sufferssubstantially, and, importantly, that this effect is moderated by theelicitation approach used to collect the dialectal data. Our findings alsoparticularly suggest that relying solely on standard language data for ASRevaluation may underestimate real-world performance, particularly in languageswith substantial dialectal variation.

【6】 Singing Voice Conversion with Accompaniment Using Self-Supervised  Representation-Based Melody Features
标题:使用自我监督的基于表达的旋律特征进行伴奏演唱声音转换
链接:https://arxiv.org/abs/2502.04722
作者:Wei Chen,  Binzhu Sha,  Jing Yang,  Zhuo Wang,  Fan Fan,  Zhiyong Wu
备注:Accepted by ICASSP2025
摘要:在歌唱嗓音转换中,旋律的保持是至关重要的。然而,在许多场景中,音频通常伴随着背景音乐(BGM),这会导致音频失真并干扰旋律和其他关键特征的提取,从而显著降低SVC性能。以前的方法试图通过使用更强大的基于神经网络的旋律提取器来解决这个问题,但是在复杂伴奏的情况下,它们的性能急剧下降。其他方法涉及在转换之前执行源分离,但这通常会引入明显的伪影,导致转换质量显着下降并增加用户的运营成本。为了解决这些问题,我们引入了一种新的SVC方法,该方法使用基于自我监督表示的旋律特征,以提高旋律建模的准确性存在的BGM。在我们的实验中,我们比较了不同的自我监督学习(SSL)模型的旋律提取的有效性,并首次探索SSL如何有利于旋律提取的任务。实验结果表明,我们提出的SVC模型显着优于现有的基线方法在旋律的准确性,并显示出更高的相似性和自然度在主观和客观评价在嘈杂和干净的音频环境。
摘要:Melody preservation is crucial in singing voice conversion (SVC). However, inmany scenarios, audio is often accompanied with background music (BGM), whichcan cause audio distortion and interfere with the extraction of melody andother key features, significantly degrading SVC performance. Previous methodshave attempted to address this by using more robust neural network-based melodyextractors, but their performance drops sharply in the presence of complexaccompaniment. Other approaches involve performing source separation beforeconversion, but this often introduces noticeable artifacts, leading to asignificant drop in conversion quality and increasing the user's operationalcosts. To address these issues, we introduce a novel SVC method that usesself-supervised representation-based melody features to improve melody modelingaccuracy in the presence of BGM. In our experiments, we compare theeffectiveness of different self-supervised learning (SSL) models for melodyextraction and explore for the first time how SSL benefits the task of melodyextraction. The experimental results demonstrate that our proposed SVC modelsignificantly outperforms existing baseline methods in terms of melody accuracyand shows higher similarity and naturalness in both subjective and objectiveevaluations across noisy and clean audio environments.

【7】 Dynamic Frequency-Adaptive Knowledge Distillation for Speech Enhancement
标题:语音增强的动态频率自适应知识提炼
链接:https://arxiv.org/abs/2502.04711
作者:Xihao Yuan,  Siqi Liu,  Hanting Chen,  Lu Zhou,  Jian Li,  Jie Hu
备注:5 pages, 2 figures, accepted by ICASSP2025
摘要:基于深度学习的语音增强(SE)模型最近的表现优于传统技术,但由于计算和内存需求高,它们在资源受限设备上的部署仍然具有挑战性。本文提出了一种新的动态频率自适应知识提取(DFKD)方法,有效地压缩SE模型。我们的方法动态评估模型的输出,区分高频和低频分量,并调整学习目标以满足不同频带的独特要求,利用SE任务的固有特性。为了评估DFKD的有效性,我们在三个最先进的模型上进行了实验:DCCRN,ConTasNet和DPTNet。结果表明,我们的方法不仅显着提高了压缩模型(学生模型)的性能,但也优于其他基于逻辑的知识蒸馏方法,专门为SE任务。
摘要:Deep learning-based speech enhancement (SE) models have recently outperformedtraditional techniques, yet their deployment on resource-constrained devicesremains challenging due to high computational and memory demands. This paperintroduces a novel dynamic frequency-adaptive knowledge distillation (DFKD)approach to effectively compress SE models. Our method dynamically assesses themodel's output, distinguishing between high and low-frequency components, andadapts the learning objectives to meet the unique requirements of differentfrequency bands, capitalizing on the SE task's inherent characteristics. Toevaluate the DFKD's efficacy, we conducted experiments on threestate-of-the-art models: DCCRN, ConTasNet, and DPTNet. The results demonstratethat our method not only significantly enhances the performance of thecompressed model (student model) but also surpasses other logit-based knowledgedistillation methods specifically for SE tasks.

【8】 ImprovNet: Generating Controllable Musical Improvisations with Iterative  Corruption Refinement
标题:DeliverNet:通过迭代腐败精炼生成可控的音乐即兴表演
链接:https://arxiv.org/abs/2502.04522
作者:Keshav Bhandari,  Sungkyun Chang,  Tongyu Lu,  Fareza R. Enus,  Louis B. Bradshaw,  Dorien Herremans,  Simon Colton
备注:10 pages, 6 figures
摘要:深度学习在各个领域的风格转移方面取得了显着进展,为创造性内容的生成提供了新的可能性。然而,在符号音乐领域,由于数据集有限,特别是对于爵士乐等流派,以及缺乏可以处理多个音乐生成任务的统一模型,为完整的音乐作品生成可控且富有表现力的表演级风格转移仍然具有挑战性。本文介绍了一种基于transformer的架构,它通过自我监督的腐败细化训练策略生成富有表现力和可控的音乐即兴创作。CNET在一个模型中统一了多种功能:它可以执行跨流派和流派内的即兴创作,协调旋律与特定流派的风格,并执行简短的提示延续和填充任务。该模型的迭代生成框架允许用户控制风格转换的程度和与原始组成的结构相似性。客观和主观的评估表明,在生成音乐连贯的即兴,同时保持与原始作品的结构关系的有效性。该模型优于预期音乐Transformer在短期延续和填充任务,并成功地实现了可识别的流派转换,与79%的参与者正确识别爵士风格的即兴。我们的代码和演示页面可在https://github.com/keshavbhandari/improvnet上找到。
摘要:Deep learning has enabled remarkable advances in style transfer acrossvarious domains, offering new possibilities for creative content generation.However, in the realm of symbolic music, generating controllable and expressiveperformance-level style transfers for complete musical works remainschallenging due to limited datasets, especially for genres such as jazz, andthe lack of unified models that can handle multiple music generation tasks.This paper presents ImprovNet, a transformer-based architecture that generatesexpressive and controllable musical improvisations through a self-supervisedcorruption-refinement training strategy. ImprovNet unifies multiplecapabilities within a single model: it can perform cross-genre and intra-genreimprovisations, harmonize melodies with genre-specific styles, and executeshort prompt continuation and infilling tasks. The model's iterative generationframework allows users to control the degree of style transfer and structuralsimilarity to the original composition. Objective and subjective evaluationsdemonstrate ImprovNet's effectiveness in generating musically coherentimprovisations while maintaining structural relationships with the originalpieces. The model outperforms Anticipatory Music Transformer in shortcontinuation and infilling tasks and successfully achieves recognizable genreconversion, with 79\% of participants correctly identifying jazz-styleimprovisations. Our code and demo page can be found athttps://github.com/keshavbhandari/improvnet.

【9】 ADIFF: Explaining audio difference using natural language
标题:ADiff:使用自然语言解释音频差异
链接:https://arxiv.org/abs/2502.04476
作者:Soham Deshmukh,  Shuo Han,  Rita Singh,  Bhiksha Raj
备注:Accepted at ICLR 2025. Dataset and checkpoints are available at: this https URL
摘要:理解和解释录音之间的差异对于音频取证、质量评估和音频生成等领域至关重要。这涉及识别和描述音频事件,声学场景,信号特征及其对听众的情感影响。本文首先全面研究了解释音频差异的任务,然后提出了任务的基准,基线。首先,我们提出了两个新的数据集,用于音频差异解释,这些数据集源自AudioCaps和Clotho音频字幕数据集。使用大语言模型(LLM),我们生成三个层次的差异解释:(1)音频事件和对象的简洁描述,(2)关于音频事件,声学场景和信号属性的简短句子,以及(3)包括语义和听众情绪的综合解释。对于基线,我们使用前缀调优,其中来自两个音频文件的音频嵌入用于提示冻结的语言模型。我们的实证分析和消融研究表明,天真的基线努力区分感知相似的声音,并产生详细的第3层解释。为了解决这些局限性,我们提出了ADIFF,它引入了交叉投影模块,位置字幕和三步训练过程,以增强模型产生详细解释的能力。我们使用客观指标和人工评估来评估我们的模型,并显示我们的模型增强导致了朴素基线和SoTA音频语言模型(ALM)Qwen Audio性能的显着改善。最后,我们进行了多个消融研究,以研究交叉投影,语言模型参数,位置字幕,第三阶段微调的影响,并提出我们的研究结果。我们的基准,发现和强大的基线为音频差异的细微和人性化解释铺平了道路。
摘要:Understanding and explaining differences between audio recordings is crucialfor fields like audio forensics, quality assessment, and audio generation. Thisinvolves identifying and describing audio events, acoustic scenes, signalcharacteristics, and their emotional impact on listeners. This paper stands outas the first work to comprehensively study the task of explaining audiodifferences and then propose benchmark, baselines for the task. First, wepresent two new datasets for audio difference explanation derived from theAudioCaps and Clotho audio captioning datasets. Using Large Language Models(LLMs), we generate three levels of difference explanations: (1) concisedescriptions of audio events and objects, (2) brief sentences about audioevents, acoustic scenes, and signal properties, and (3) comprehensiveexplanations that include semantics and listener emotions. For the baseline, weuse prefix tuning where audio embeddings from two audio files are used toprompt a frozen language model. Our empirical analysis and ablation studiesreveal that the naive baseline struggles to distinguish perceptually similarsounds and generate detailed tier 3 explanations. To address these limitations,we propose ADIFF, which introduces a cross-projection module, positioncaptioning, and a three-step training process to enhance the model's ability toproduce detailed explanations. We evaluate our model using objective metricsand human evaluation and show our model enhancements lead to significantimprovements in performance over naive baseline and SoTA Audio-Language Model(ALM) Qwen Audio. Lastly, we conduct multiple ablation studies to study theeffects of cross-projection, language model parameters, position captioning,third stage fine-tuning, and present our findings. Our benchmarks, findings,and strong baseline pave the way for nuanced and human-like explanations ofaudio differences.

【10】 FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks
标题:FocalCodec:通过Focal调制网络进行低比特率语音编码
链接:https://arxiv.org/abs/2502.04465
作者:Luca Della Libera,  Francesco Paissan,  Cem Subakan,  Mirco Ravanelli
备注:18 pages
摘要:大型语言模型通过在大规模数据集上进行自我监督预训练,彻底改变了自然语言处理。受这一成功的启发,研究人员探索了通过使用神经音频编解码器将连续音频离散化为令牌来使这些方法适应语音。然而,现有的方法面临着限制,包括高比特率,语义或声学信息的丢失,以及在试图捕获两者时对多码本设计的依赖,这增加了下游任务的架构复杂性。为了解决这些挑战,我们介绍了FocalCodec,一个有效的低比特率编解码器的基础上,利用一个单一的二进制码本压缩语音0.16和0.65 kbps的焦点调制。FocalCodec在语音再合成和语音转换方面提供了具有竞争力的性能,其比特率低于当前最先进的技术水平,同时有效地处理多语言语音和嘈杂环境。对下游任务的评估表明,FocalCodec成功地保留了足够的语义和声学信息,同时也非常适合生成式建模。演示示例、代码和检查点可在https://lucadellalib.github.io/focalcodec-web/上获得。
摘要:Large language models have revolutionized natural language processing throughself-supervised pretraining on massive datasets. Inspired by this success,researchers have explored adapting these methods to speech by discretizingcontinuous audio into tokens using neural audio codecs. However, existingapproaches face limitations, including high bitrates, the loss of eithersemantic or acoustic information, and the reliance on multi-codebook designswhen trying to capture both, which increases architectural complexity fordownstream tasks. To address these challenges, we introduce FocalCodec, anefficient low-bitrate codec based on focal modulation that utilizes a singlebinary codebook to compress speech between 0.16 and 0.65 kbps. FocalCodecdelivers competitive performance in speech resynthesis and voice conversion atlower bitrates than the current state-of-the-art, while effectively handlingmultilingual speech and noisy environments. Evaluation on downstream tasksshows that FocalCodec successfully preserves sufficient semantic and acousticinformation, while also being well-suited for generative modeling. Demosamples, code and checkpoints are available athttps://lucadellalib.github.io/focalcodec-web/.

机器翻译由腾讯交互翻译提供,仅供参考