本文经arXiv每日学术速递授权转载
微信公众号:arXiv_Daily
链接:https://arxiv.org/abs/2502.07562
摘要:语音合成模型将书面文本转换为自然的声音。虽然早期的模型仅限于单个扬声器,但最近的进步已经导致zero-shot系统的开发,该系统使用他们的声音作为额外的提示从广泛的扬声器生成逼真的语音。然而,他们仍然难以模仿与训练数据集显著不同的非工作室质量样本。在这项工作中,我们证明,利用低秩自适应(LoRA),使我们能够成功地使用即使是单记录的自发语音在嘈杂的环境中作为提示。这种方法将说话人相似度提高了30 pp $,同时保留了内容和自然度。它代表了朝着创建真正多样化的语音语料库迈出的重要一步,这在所有与语音相关的任务中至关重要。
摘要:Speech synthesis models convert written text into natural-sounding audio.While earlier models were limited to a single speaker, recent advancements haveled to the development of zero-shot systems that generate realistic speech froma wide range of speakers using their voices as additional prompts. However,they still struggle with imitating non-studio-quality samples that differsignificantly from the training datasets. In this work, we demonstrate thatutilizing Low-Rank Adaptation (LoRA) allows us to successfully use even singlerecordings of spontaneous speech in noisy environments as prompts. Thisapproach enhances speaker similarity by up to $30pp$ while preserving contentand naturalness. It represents a significant step toward creating truly diversespeech corpora, that is crucial in all speech-related tasks.
标题:用于多扬声器环境的基于视觉的空间音频生成系统
链接:https://arxiv.org/abs/2502.07538
摘要:在电影和视频游戏等多媒体应用中,空间音频技术被广泛用于通过模拟3D声音来增强用户体验:将单声道音频转换为双耳格式。然而,这个过程对于声音设计师来说通常是复杂和劳动密集型的,需要音频与视觉组件的空间位置精确同步。为了解决这些挑战,我们提出了一个基于视觉的空间音频生成系统-一个自动化系统,集成了人脸检测YOLOv 8的对象检测,单目深度估计和空间音频技术。值得注意的是,该系统在不需要额外的双耳数据集训练的情况下操作。所提出的系统进行评估,对现有的空间音频生成系统,使用客观的指标。实验结果表明,我们的方法显着提高音频和视频之间的空间一致性,提高语音质量,并在多说话人的情况下表现出鲁棒性。通过简化视听对齐过程,该系统使音响工程师能够有效地实现高质量的结果,使其成为多媒体制作专业人员的宝贵工具。
摘要:In multimedia applications such as films and video games, spatial audiotechniques are widely employed to enhance user experiences by simulating 3Dsound: transforming mono audio into binaural formats. However, this process isoften complex and labor-intensive for sound designers, requiring precisesynchronization of audio with the spatial positions of visual components. Toaddress these challenges, we propose a visual-based spatial audio generationsystem - an automated system that integrates face detection YOLOv8 for objectdetection, monocular depth estimation, and spatial audio techniques. Notably,the system operates without requiring additional binaural dataset training. Theproposed system is evaluated against existing Spatial Audio generation systemusing objective metrics. Experimental results demonstrate that our methodsignificantly improves spatial consistency between audio and video, enhancesspeech quality, and performs robustly in multi-speaker scenarios. Bystreamlining the audio-visual alignment process, the proposed system enablessound engineers to achieve high-quality results efficiently, making it avaluable tool for professionals in multimedia production.
标题:使用Roland TR-808低音鼓产生的和声和调换限制
链接:https://arxiv.org/abs/2502.07524
备注:Proc. of the 25th Int. Society for Music Information Retrieval Conf., San Francisco, United States, 2024. 8 pages, 9 figures
摘要:这项研究调查了嘻哈音乐制作人Scott Storch的调性方法,其中歌曲的基调被调换以适应Roland TR-808低音鼓,而不是将鼓调整到歌曲的基调。这一过程涉及到除低音鼓外的所有音轨的调整,表明了重要的生产动机。主要的限制源于TR-808低音鼓的可用音高范围有限,如果要保留其特有的声音。该研究考察了鼓的调音实践,罗兰TR-808在音乐中的作用,以及低音鼓的低音质量。对TR-808样本的分析揭示了它们的特征以及它们与陷阱和嘻哈等现代流派的融合。该研究还考虑了扬声器频率响应和人耳灵敏度对低音鼓感知的影响。研究结果表明,Storch的方法优先考虑低音鼓的频谱特性,而不是传统的音高值,以增强低音响应。保持TR-808低音鼓独特声音的需要强调了频谱共振峰和音域在当代流行音乐制作中的重要性。
摘要:The study investigates hip-hop music producer Scott Storch's approach totonality, where the song's key is transposed to fit the Roland TR-808 bass druminstead of tuning the drums to the song's key. This process, involving theadjustment of all tracks except the bass drum, suggests significant productionmotives. The primary constraint stems from the limited usable pitch range ofthe TR-808 bass drum if its characteristic sound is to be preserved. Theresearch examines drum tuning practices, the role of the Roland TR-808 inmusic, and the sub-bass qualities of its bass drum. Analysis of TR-808 samplesreveals their characteristics and their integration into modern genres liketrap and hip-hop. The study also considers the impact of loudspeaker frequencyresponse and human ear sensitivity on bass drum perception. The findingssuggest that Storch's method prioritizes the spectral properties of the bassdrum over traditional pitch values to enhance the bass response. The need tomaintain the unique sound of the TR-808 bass drum underscores the importance ofspectral formants and register in contemporary popular music production.
标题:JamendoMaxCaps:具有输入元数据的大规模音乐字幕数据集
链接:https://arxiv.org/abs/2502.07461
备注:8 pages, 5 figures
摘要:我们介绍JamendoMaxCaps,这是一个大规模的音乐字幕数据集,包含来自著名的Jamendo平台的20多万首免费许可的器乐曲目。该数据集包括由最先进的字幕模型生成的字幕,并通过估算元数据进行增强。我们还介绍了一个检索系统,利用音乐功能和元数据来识别类似的歌曲,然后使用本地大型语言模型(LLLM)来填充缺失的元数据。这种方法使我们能够为从事音乐语言理解任务的研究人员提供更全面和信息丰富的数据集。我们用五种不同的测量方法定量地验证了这种方法。通过公开JamendoMaxCaps数据集,我们提供了一个高质量的资源来推进音乐语言理解任务的研究,如音乐检索,多模态表示学习和生成音乐模型。
摘要:We introduce JamendoMaxCaps, a large-scale music-caption dataset featuringover 200,000 freely licensed instrumental tracks from the renowned Jamendoplatform. The dataset includes captions generated by a state-of-the-artcaptioning model, enhanced with imputed metadata. We also introduce a retrievalsystem that leverages both musical features and metadata to identify similarsongs, which are then used to fill in missing metadata using a local largelanguage model (LLLM). This approach allows us to provide a more comprehensiveand informative dataset for researchers working on music-language understandingtasks. We validate this approach quantitatively with five differentmeasurements. By making the JamendoMaxCaps dataset publicly available, weprovide a high-quality resource to advance research in music-languageunderstanding tasks such as music retrieval, multimodal representationlearning, and generative music models.
标题:先进的Zero-Shot文本到语音技术,通过可控的掩蔽语音预测进行背景去除和保留
链接:https://arxiv.org/abs/2502.07345
备注:Accepted by ICASSP 2025
摘要:声学背景在自然对话中起着至关重要的作用。它提供了上下文,帮助听众理解环境,但强烈的背景使听众难以理解口语。对这些背景的适当处理取决于情况:虽然可能有必要去除背景以确保语音清晰,但保留背景有时对保持语音的上下文完整性至关重要。尽管最近在zero-shot文本到语音技术方面取得了进步,但是当前的系统经常与包含背景的语音提示作斗争。为了解决这些挑战,我们提出了一个可控的掩蔽语音预测策略,再加上一个双扬声器编码器,利用任务相关的控制信号来指导双背景去除和保留目标的预测。实验结果表明,我们的方法可以精确控制在各种声学条件下的背景的去除或保留,并表现出很强的泛化能力,在看不见的场景。
摘要:The acoustic background plays a crucial role in natural conversation. Itprovides context and helps listeners understand the environment, but a strongbackground makes it difficult for listeners to understand spoken words. Theappropriate handling of these backgrounds is situation-dependent: Although itmay be necessary to remove background to ensure speech clarity, preserving thebackground is sometimes crucial to maintaining the contextual integrity of thespeech. Despite recent advancements in zero-shot Text-to-Speech technologies,current systems often struggle with speech prompts containing backgrounds. Toaddress these challenges, we propose a Controllable Masked Speech Predictionstrategy coupled with a dual-speaker encoder, utilizing a task-related controlsignal to guide the prediction of dual background removal and preservationtargets. Experimental results demonstrate that our approach enables precisecontrol over the removal or preservation of background across various acousticconditions and exhibits strong generalization capabilities in unseen scenarios.
标题:全民音乐:探索音乐生成模型中的多元文化表达(相机准备)
链接:https://arxiv.org/abs/2502.07328
备注:17 pages, 5 figures, accepted to NAACL'25
摘要:音乐语言模型的出现极大地增强了人工智能系统的自动音乐生成能力,但它们在覆盖世界音乐流派和文化方面也受到限制。我们提出了一项研究的数据集和研究论文的音乐生成和量化的偏见和代表性不足的流派。我们发现,现有音乐数据集的总时长中只有5.7%来自非西方流派,这自然会导致不同流派的模型表现不一。然后,我们研究了参数高效微调(PEFT)技术在减轻这种偏差方面的功效。我们用两个流行的模型进行了实验- MusicGen和Mustango,两个代表性不足的非西方音乐传统-印度斯坦古典音乐和土耳其Makam音乐,突出了通过小数据集对音乐进行跨流派改编的承诺和非平凡性,这意味着需要更公平的基线音乐语言模型,旨在进行跨文化迁移学习。
摘要:The advent of Music-Language Models has greatly enhanced the automatic musicgeneration capability of AI systems, but they are also limited in theircoverage of the musical genres and cultures of the world. We present a study ofthe datasets and research papers for music generation and quantify the bias andunder-representation of genres. We find that only 5.7% of the total hours ofexisting music datasets come from non-Western genres, which naturally leads todisparate performance of the models across genres. We then investigate theefficacy of Parameter-Efficient Fine-Tuning (PEFT) techniques in mitigatingthis bias. Our experiments with two popular models -- MusicGen and Mustango,for two underrepresented non-Western music traditions -- Hindustani Classicaland Turkish Makam music, highlight the promises as well as the non-trivialityof cross-genre adaptation of music through small datasets, implying the needfor more equitable baseline music-language models that are designed forcross-cultural transfer learning.
标题:Vevo:具有自我监督解开的可控Zero-Shot语音模仿
链接:https://arxiv.org/abs/2502.07243
备注:Accepted by ICLR 2025
摘要:语音的模仿是语音生成的关键,它针对语音的特定属性,如音色、说话风格等。然而,现有的方法严重依赖于注释数据,并且难以有效地解开音色和风格,导致在实现可控生成方面的挑战,特别是在zero-shot场景中。为了解决这些问题,我们提出了Vevo,一个多功能的zero-shot语音模仿框架,具有可控的音色和风格。Vevo在两个核心阶段进行操作:(1)内容风格建模:给定文本或语音的内容标记作为输入,我们利用自回归Transformer来生成内容风格标记,这是由风格参考提示的;(2)声学建模:给定内容风格标记作为输入,我们采用流匹配Transformer来生成声学表示,这是由音色参考提示的。为了获得语音的内容和内容风格标记,我们设计了一种完全自我监督的方法,该方法可以逐步识别语音的音色,风格和语言内容。具体来说,我们采用VQ-VAE作为HuBERT的连续隐藏特征的标记器。我们把VQ-VAE码书的词汇量作为信息瓶颈,并仔细调整它以获得解纠缠的语音表示。Vevo仅在6万小时的有声读物语音数据上进行了自我监督训练,而没有对特定风格的语料库进行任何微调,在口音和情感转换任务中达到或超过了现有的方法。此外,Vevo在zero-shot语音转换和文本到语音任务中的有效性进一步证明了其强大的通用性和多功能性。音频样本可在https://versavoice.github.io上获得。
摘要:The imitation of voice, targeted on specific speech attributes such as timbreand speaking style, is crucial in speech generation. However, existing methodsrely heavily on annotated data, and struggle with effectively disentanglingtimbre and style, leading to challenges in achieving controllable generation,especially in zero-shot scenarios. To address these issues, we propose Vevo, aversatile zero-shot voice imitation framework with controllable timbre andstyle. Vevo operates in two core stages: (1) Content-Style Modeling: Giveneither text or speech's content tokens as input, we utilize an autoregressivetransformer to generate the content-style tokens, which is prompted by a stylereference; (2) Acoustic Modeling: Given the content-style tokens as input, weemploy a flow-matching transformer to produce acoustic representations, whichis prompted by a timbre reference. To obtain the content and content-styletokens of speech, we design a fully self-supervised approach that progressivelydecouples the timbre, style, and linguistic content of speech. Specifically, weadopt VQ-VAE as the tokenizer for the continuous hidden features of HuBERT. Wetreat the vocabulary size of the VQ-VAE codebook as the information bottleneck,and adjust it carefully to obtain the disentangled speech representations.Solely self-supervised trained on 60K hours of audiobook speech data, withoutany fine-tuning on style-specific corpora, Vevo matches or surpasses existingmethods in accent and emotion conversion tasks. Additionally, Vevo'seffectiveness in zero-shot voice conversion and text-to-speech tasks furtherdemonstrates its strong generalization and versatility. Audio samples areavailable at https://versavoice.github.io.
标题:自适应中心频率局部竞争语音算法
链接:https://arxiv.org/abs/2502.06989
备注:(C) 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. arXiv admin note: text overlap with arXiv:2409.08188
摘要:受神经系统启发的神经形态计算以其对效率和低功耗的关注彻底改变了信息处理。使用稀疏编码,这种模式提高了处理效率,这对于具有功率限制的边缘设备至关重要。局部竞争算法(LCA),适用于音频与Gammatone和Gammachirp滤波器组,提供了一个有效的稀疏编码方法的神经形态语音处理。自适应LCA(ALCA)通过动态调整调制参数进一步改进了该方法,从而提高了重建质量和稀疏性。本文介绍了一种增强的ALCA版本,ALCA中心频率(ALCA-CF),它动态地适应调制参数和中心频率,优化语音表示。评估表明,这种方法提高了重建质量和稀疏性,同时显着降低语音分类的功耗,而不影响分类精度,特别是在英特尔的Loihi 2神经形态芯片。
摘要:Neuromorphic computing, inspired by nervous systems, revolutionizesinformation processing with its focus on efficiency and low power consumption.Using sparse coding, this paradigm enhances processing efficiency, which iscrucial for edge devices with power constraints. The Locally CompetitiveAlgorithm (LCA), adapted for audio with Gammatone and Gammachirp filter banks,provides an efficient sparse coding method for neuromorphic speech processing.Adaptive LCA (ALCA) further refines this method by dynamically adjustingmodulation parameters, thereby improving reconstruction quality and sparsity.This paper introduces an enhanced ALCA version, the ALCA Central Frequency(ALCA-CF), which dynamically adapts both modulation parameters and centralfrequencies, optimizing the speech representation. Evaluations show that thisapproach improves reconstruction quality and sparsity while significantlyreducing the power consumption of speech classification, without compromisingclassification accuracy, particularly on Intel's Loihi 2 neuromorphic chip.
标题:合成音频有助于认知状态任务
链接:https://arxiv.org/abs/2502.06922
备注:John Murzaku and Adil Soubki contributed equally to this work
摘要:NLP社区广泛关注认知状态任务的纯文本方法,但音频可以通过韵律提供重要的缺失线索。我们认为,文本到语音模型学习跟踪认知状态的各个方面,以产生自然的音频,信号音频模型隐含地识别是正交的语言模型利用的信息。我们提出了合成音频数据微调(SAD),一个框架,在这个框架中,我们展示了与认知状态建模相关的7个任务,这些任务受益于来自现成TTS系统的文本和zero-shot合成音频数据的多模态训练。我们显示了改进的纯文本模态时,添加合成音频数据的纯文本语料库。此外,在包含黄金音频的任务和语料库上,我们展示了我们的SAD框架与文本和黄金音频相比,在文本和合成音频方面实现了具有竞争力的性能。
摘要:The NLP community has broadly focused on text-only approaches of cognitivestate tasks, but audio can provide vital missing cues through prosody. We positthat text-to-speech models learn to track aspects of cognitive state in orderto produce naturalistic audio, and that the signal audio models implicitlyidentify is orthogonal to the information that language models exploit. Wepresent Synthetic Audio Data fine-tuning (SAD), a framework where we show that7 tasks related to cognitive state modeling benefit from multimodal training onboth text and zero-shot synthetic audio data from an off-the-shelf TTS system.We show an improvement over the text-only modality when adding synthetic audiodata to text-only corpora. Furthermore, on tasks and corpora that do containgold audio, we show our SAD framework achieves competitive performance withtext and synthetic audio compared to text and gold audio.
标题:了解声音事件检测的频率依赖性
链接:https://arxiv.org/abs/2502.07208
摘要:在这项工作中,各种分析方法进行了频率相关的方法在SED进一步深入研究其详细的特点和行为。虽然SED通过采用来自其他模式识别领域的各种深度学习技术而迅速发展,但这些技术通常不适合SED。为了解决这个问题,以前提出了两种频率相关的SED方法:FilterAugment,一种随机加权频带的数据增强,以及频率动态卷积(FDY Conv),一种应用频率自适应卷积核的架构。这些方法在SED中表现出了优越的性能,我们的目标是进一步分析它们在SED中的详细有效性和特点。我们比较类的性能,找出具体的优点和缺点的FilterAugment和FDY Conv。我们应用加权类激活映射(Grad-CAM),它突出了模型推断的时频区域,在有和没有频率掩蔽和两种类型的FilterAugment的SED模型上观察它们的详细特性。我们提出了更简单的频率相关卷积方法,并将其与FDY Conv进行比较,以进一步了解FDY Conv的哪些组件会影响SED性能。最后,我们应用PCA来展示FDY Conv如何在不同的声音事件类上跨频率维度调整动态内核。结果表明,频率相关性在声音事件检测中起着重要的作用,并进一步证实了频率相关方法对SED的有效性。
摘要:In this work, various analysis methods are conducted on frequency-dependentmethods on SED to further delve into their detailed characteristics andbehaviors on SED. While SED has been rapidly advancing through the adoption ofvarious deep learning techniques from other pattern recognition fields, thesetechniques are often not suitable for SED. To address this issue, twofrequency-dependent SED methods were previously proposed: FilterAugment, a dataaugmentation randomly weighting frequency bands, and frequency dynamicconvolution (FDY Conv), an architecture applying frequency adaptive convolutionkernels. These methods have demonstrated superior performance in SED, and weaim to further analyze their detailed effectiveness and characteristics in SED.We compare class-wise performance to find out specific pros and cons ofFilterAugment and FDY Conv. We apply Gradient-weighted Class Activation Mapping(Grad-CAM), which highlights time-frequency region that is more inferred by themodel, on SED models with and without frequency masking and two types ofFilterAugment to observe their detailed characteristics. We propose simplerfrequency dependent convolution methods and compare them with FDY Conv tofurther understand which components of FDY Conv affects SED performance.Lastly, we apply PCA to show how FDY Conv adapts dynamic kernel acrossfrequency dimensions on different sound event classes. The results anddiscussions demonstrate that frequency dependency plays a significant role insound event detection and further confirms the effectiveness of frequencydependent methods on SED.
标题:弱监督语音去回响的混合模型
链接:https://arxiv.org/abs/2502.06839
备注:None
摘要:本文介绍了一种新的训练策略,以改善语音去混响系统使用最少的声学信息和混响(湿)的语音。大多数现有的算法依赖于成对的干/湿数据,这是很难获得的,或对目标度量,可能无法充分捕捉混响特性,并可能导致非目标度量的结果不佳。我们的方法使用有限的声学信息,如混响时间(RT 60),训练去混响系统。该系统的输出重新合成使用生成的房间脉冲响应,并与原始混响语音进行比较,提供了一种新的混响匹配损失取代标准的目标指标。在推理过程中,只使用经过训练的去混响模型。实验结果表明,与最先进的方法相比,我们的方法在语音去混响中使用的各种客观指标上实现了更一致的性能。
摘要:This paper introduces a new training strategy to improve speechdereverberation systems using minimal acoustic information and reverberant(wet) speech. Most existing algorithms rely on paired dry/wet data, which isdifficult to obtain, or on target metrics that may not adequately capturereverberation characteristics and can lead to poor results on non-targetmetrics. Our approach uses limited acoustic information, like the reverberationtime (RT60), to train a dereverberation system. The system's output isresynthesized using a generated room impulse response and compared with theoriginal reverberant speech, providing a novel reverberation matching lossreplacing the standard target metrics. During inference, only the traineddereverberation model is used. Experimental results demonstrate that our methodachieves more consistent performance across various objective metrics used inspeech dereverberation than the state-of-the-art.
标题:RenderBox:具有文本控制的表现力渲染
链接:https://arxiv.org/abs/2502.07711
摘要:表现性的音乐表演渲染涉及用定时、动态、清晰度和乐器特定技术的变化来解释符号乐谱,从而产生捕获音乐情感意图的表演。我们介绍RenderBox,一个统一的框架,用于跨多种乐器的文本和乐谱控制的音频性能生成,通过自然语言描述和使用乐谱的粒度级控制应用粗糙级控制。基于扩散Transformer架构和交叉注意联合条件反射,我们提出了一个基于训练的范式,从简单的合成到表达性能,逐渐纳入可控因素,如速度,错误和风格的多样性。 与基准模型相比,RenderBox在关键指标(如FAD和CLAP)以及不同提示任务下的节奏和音高准确性方面实现了高性能。主观评估进一步表明,RenderBox能够生成可控的表现性表演,听起来自然,音乐上引人入胜,与提示和意图保持良好的一致。
摘要:Expressive music performance rendering involves interpreting symbolic scoreswith variations in timing, dynamics, articulation, and instrument-specifictechniques, resulting in performances that capture musical can emotionalintent. We introduce RenderBox, a unified framework for text-and-scorecontrolled audio performance generation across multiple instruments, applyingcoarse-level controls through natural language descriptions and granular-levelcontrols using music scores. Based on a diffusion transformer architecture andcross-attention joint conditioning, we propose a curriculum-based paradigm thattrains from plain synthesis to expressive performance, gradually incorporatingcontrollable factors such as speed, mistakes, and style diversity. RenderBox achieves high performance compared to baseline models across keymetrics such as FAD and CLAP, and also tempo and pitch accuracy under differentprompting tasks. Subjective evaluation further demonstrates that RenderBox isable to generate controllable expressive performances that sound natural andmusically engaging, aligning well with prompts and intent.
标题:利用分层选择性状态空间模型和去耦合交叉信息损失实现高效和多方面的计算机辅助发音训练
链接:https://arxiv.org/abs/2502.07575
备注:Accepted to NAACL 2025 Main Conference
摘要:在构建计算机辅助发音训练(CAPT)系统方面的先前努力通常将自动发音评估(APA)和发音错误检测和诊断(MDD)作为单独的战线:前者旨在提供跨不同语言水平的多个发音方面分数,而后者则专注于精确定位非母语学习者所犯的精确语音发音错误。然而,人们普遍认为,一个成熟的CAPT系统应该同时有效地执行这两种功能。为了应对这种激增的需求,我们在这项工作中首先提出了HMamba,一种新颖的CAPT方法,无缝集成APA和MDD任务并行。此外,我们引入了一种新的损失函数,解耦交叉熵损失(deXent),专门为MDD定制,以促进更好的监督学习,用于检测发音错误的电话,从而提高整体性能。在speechocean762基准数据集上的一组全面的实证结果证明了我们的APA方法的有效性。值得注意的是,我们提出的方法在MDD性能上也有了相当大的改进,达到了63.85%的F1分数。我们的代码可在https://github.com/Fuann/hmamba上获得
摘要:Prior efforts in building computer-assisted pronunciation training (CAPT)systems often treat automatic pronunciation assessment (APA) andmispronunciation detection and diagnosis (MDD) as separate fronts: the formeraims to provide multiple pronunciation aspect scores across diverse linguisticlevels, while the latter focuses instead on pinpointing the precise phoneticpronunciation errors made by non-native language learners. However, it isgenerally expected that a full-fledged CAPT system should perform bothfunctionalities simultaneously and efficiently. In response to this surgingdemand, we in this work first propose HMamba, a novel CAPT approach thatseamlessly integrates APA and MDD tasks in parallel. In addition, we introducea novel loss function, decoupled cross-entropy loss (deXent), specificallytailored for MDD to facilitate better-supervised learning for detectingmispronounced phones, thereby enhancing overall performance. A comprehensiveset of empirical results on the speechocean762 benchmark dataset demonstratesthe effectiveness of our approach on APA. Notably, our proposed approach alsoyields a considerable improvement in MDD performance over a strong baseline,achieving an F1-score of 63.85%. Our codes are made available athttps://github.com/Fuann/hmamba
标题:了解声音事件检测的频率依赖性
链接:https://arxiv.org/abs/2502.07208
摘要:在这项工作中,各种分析方法进行了频率相关的方法在SED进一步深入研究其详细的特点和行为。虽然SED通过采用来自其他模式识别领域的各种深度学习技术而迅速发展,但这些技术通常不适合SED。为了解决这个问题,以前提出了两种频率相关的SED方法:FilterAugment,一种随机加权频带的数据增强,以及频率动态卷积(FDY Conv),一种应用频率自适应卷积核的架构。这些方法在SED中表现出了优越的性能,我们的目标是进一步分析它们在SED中的详细有效性和特点。我们比较类的性能,找出具体的优点和缺点的FilterAugment和FDY Conv。我们应用加权类激活映射(Grad-CAM),它突出了模型推断的时频区域,在有和没有频率掩蔽和两种类型的FilterAugment的SED模型上观察它们的详细特性。我们提出了更简单的频率相关卷积方法,并将其与FDY Conv进行比较,以进一步了解FDY Conv的哪些组件会影响SED性能。最后,我们应用PCA来展示FDY Conv如何在不同的声音事件类上跨频率维度调整动态内核。结果表明,频率相关性在声音事件检测中起着重要的作用,并进一步证实了频率相关方法对SED的有效性。
摘要:In this work, various analysis methods are conducted on frequency-dependentmethods on SED to further delve into their detailed characteristics andbehaviors on SED. While SED has been rapidly advancing through the adoption ofvarious deep learning techniques from other pattern recognition fields, thesetechniques are often not suitable for SED. To address this issue, twofrequency-dependent SED methods were previously proposed: FilterAugment, a dataaugmentation randomly weighting frequency bands, and frequency dynamicconvolution (FDY Conv), an architecture applying frequency adaptive convolutionkernels. These methods have demonstrated superior performance in SED, and weaim to further analyze their detailed effectiveness and characteristics in SED.We compare class-wise performance to find out specific pros and cons ofFilterAugment and FDY Conv. We apply Gradient-weighted Class Activation Mapping(Grad-CAM), which highlights time-frequency region that is more inferred by themodel, on SED models with and without frequency masking and two types ofFilterAugment to observe their detailed characteristics. We propose simplerfrequency dependent convolution methods and compare them with FDY Conv tofurther understand which components of FDY Conv affects SED performance.Lastly, we apply PCA to show how FDY Conv adapts dynamic kernel acrossfrequency dimensions on different sound event classes. The results anddiscussions demonstrate that frequency dependency plays a significant role insound event detection and further confirms the effectiveness of frequencydependent methods on SED.
标题:VINP:结合神经语音先验的变分Bayesian推理,用于联合ASB有效语音去回响和盲RIR识别
链接:https://arxiv.org/abs/2502.07205
备注:Submitted to IEEE/ACM Trans. on TASLP
摘要:混响语音是指在混响过程中产生的语音信号,它包含了消声源语音和房间脉冲响应(RIR)的重要知识。本文提出了一种基于神经语音先验的变分贝叶斯推理(VBI)框架,用于联合语音去混响和盲RIR识别。在VINP中,基于卷积传递函数(CTF)近似在时频域(T-F)中构造概率信号模型。我们首次提出使用任意判别去混响深度神经网络(DNN)来预测概率模型内的消声语音的先验分布。通过整合混响语音和消声语音先验,VINP分别产生消声语音频谱和CTF滤波器的最大后验(MAP)和最大似然(ML)估计。经过简单的变换,估计出了无回声语音和RIR的波形。此外,VINP对于自动语音识别(ASR)系统是有效的,这使其与大多数基于深度学习(DL)的单通道去混响方法不同。单通道语音去混响的实验表明,VINP达到了先进的水平,在大多数指标与人类的感知和显示毫无疑问的最先进的(SOTA)性能在ASR相关的指标。对于RIR盲辨识,实验表明VINP在60 dB混响时间(RT 60)和直达混响比(DRR)盲估计方面达到了SOTA水平。代码和音频样本可在线获取。
摘要:Reverberant speech, denoting the speech signal degraded by the process ofreverberation, contains crucial knowledge of both anechoic source speech androom impulse response (RIR). This work proposes a variational Bayesianinference (VBI) framework with neural speech prior (VINP) for joint speechdereverberation and blind RIR identification. In VINP, a probabilistic signalmodel is constructed in the time-frequency (T-F) domain based on convolutiontransfer function (CTF) approximation. For the first time, we propose using anarbitrary discriminative dereverberation deep neural network (DNN) to predictthe prior distribution of anechoic speech within a probabilistic model. Byintegrating both reverberant speech and the anechoic speech prior, VINP yieldsthe maximum a posteriori (MAP) and maximum likelihood (ML) estimations of theanechoic speech spectrum and CTF filter, respectively. After simpletransformations, the waveforms of anechoic speech and RIR are estimated.Moreover, VINP is effective for automatic speech recognition (ASR) systems,which sets it apart from most deep learning (DL)-based single-channeldereverberation approaches. Experiments on single-channel speechdereverberation demonstrate that VINP reaches an advanced level in most metricsrelated to human perception and displays unquestionable state-of-the-art (SOTA)performance in ASR-related metrics. For blind RIR identification, experimentsindicate that VINP attains the SOTA level in blind estimation of reverberationtime at 60 dB (RT60) and direct-to-reverberation ratio (DRR). Codes and audiosamples are available online.
标题:弱监督语音去回响的混合模型
链接:https://arxiv.org/abs/2502.06839
备注:None
摘要:本文介绍了一种新的训练策略,以改善语音去混响系统使用最少的声学信息和混响(湿)的语音。大多数现有的算法依赖于成对的干/湿数据,这是很难获得的,或对目标度量,可能无法充分捕捉混响特性,并可能导致非目标度量的结果不佳。我们的方法使用有限的声学信息,如混响时间(RT 60),训练去混响系统。该系统的输出重新合成使用生成的房间脉冲响应,并与原始混响语音进行比较,提供了一种新的混响匹配损失取代标准的目标指标。在推理过程中,只使用经过训练的去混响模型。实验结果表明,我们的方法在语音去混响中使用的各种客观指标比最先进的实现更一致的性能。
摘要:This paper introduces a new training strategy to improve speechdereverberation systems using minimal acoustic information and reverberant(wet) speech. Most existing algorithms rely on paired dry/wet data, which isdifficult to obtain, or on target metrics that may not adequately capturereverberation characteristics and can lead to poor results on non-targetmetrics. Our approach uses limited acoustic information, like the reverberationtime (RT60), to train a dereverberation system. The system's output isresynthesized using a generated room impulse response and compared with theoriginal reverberant speech, providing a novel reverberation matching lossreplacing the standard target metrics. During inference, only the traineddereverberation model is used. Experimental results demonstrate that our methodachieves more consistent performance across various objective metrics used inspeech dereverberation than the state-of-the-art.
标题:LoRP-TTC:低级个性化文本转语音
链接:https://arxiv.org/abs/2502.07562
摘要:语音合成模型将书面文本转换为自然的声音。虽然早期的模型仅限于单个扬声器,但最近的进步已经导致zero-shot系统的开发,该系统使用他们的声音作为额外的提示从广泛的扬声器生成逼真的语音。然而,他们仍然难以模仿与训练数据集显著不同的非工作室质量样本。在这项工作中,我们证明,利用低秩自适应(LoRA),使我们能够成功地使用即使是单记录的自发语音在嘈杂的环境中作为提示。这种方法将说话人相似度提高了30 pp $,同时保留了内容和自然度。它代表了朝着创建真正多样化的语音语料库迈出的重要一步,这在所有与语音相关的任务中至关重要。
摘要:Speech synthesis models convert written text into natural-sounding audio.While earlier models were limited to a single speaker, recent advancements haveled to the development of zero-shot systems that generate realistic speech froma wide range of speakers using their voices as additional prompts. However,they still struggle with imitating non-studio-quality samples that differsignificantly from the training datasets. In this work, we demonstrate thatutilizing Low-Rank Adaptation (LoRA) allows us to successfully use even singlerecordings of spontaneous speech in noisy environments as prompts. Thisapproach enhances speaker similarity by up to $30pp$ while preserving contentand naturalness. It represents a significant step toward creating truly diversespeech corpora, that is crucial in all speech-related tasks.
标题:先进的Zero-Shot文本到语音技术,通过可控的掩蔽语音预测进行背景去除和保留
链接:https://arxiv.org/abs/2502.07345
备注:Accepted by ICASSP 2025
摘要:声学背景在自然对话中起着至关重要的作用。它提供了上下文,帮助听众理解环境,但强烈的背景使听众难以理解口语。对这些背景的适当处理取决于情况:虽然可能有必要去除背景以确保语音清晰,但保留背景有时对保持语音的上下文完整性至关重要。尽管最近在zero-shot文本到语音技术方面取得了进步,但是当前的系统经常与包含背景的语音提示作斗争。为了解决这些挑战,我们提出了一个可控的掩蔽语音预测策略,再加上一个双扬声器编码器,利用任务相关的控制信号来指导双背景去除和保留目标的预测。实验结果表明,我们的方法可以精确控制在各种声学条件下的背景的去除或保留,并表现出很强的泛化能力,在看不见的场景。
摘要:The acoustic background plays a crucial role in natural conversation. Itprovides context and helps listeners understand the environment, but a strongbackground makes it difficult for listeners to understand spoken words. Theappropriate handling of these backgrounds is situation-dependent: Although itmay be necessary to remove background to ensure speech clarity, preserving thebackground is sometimes crucial to maintaining the contextual integrity of thespeech. Despite recent advancements in zero-shot Text-to-Speech technologies,current systems often struggle with speech prompts containing backgrounds. Toaddress these challenges, we propose a Controllable Masked Speech Predictionstrategy coupled with a dual-speaker encoder, utilizing a task-related controlsignal to guide the prediction of dual background removal and preservationtargets. Experimental results demonstrate that our approach enables precisecontrol over the removal or preservation of background across various acousticconditions and exhibits strong generalization capabilities in unseen scenarios.
标题:利用自我监督语音模型中的同名音进行非典型发音评估
链接:https://arxiv.org/abs/2502.07029
备注:Accepted to NAACL 2025. Codebase available at this https URL
摘要:音位变体是指音位在语音环境的基础上在语音实现上的变化。非典型发音评估涉及到区分典型发音和非典型发音,因此对非典型发音建模是非常重要的。然而,最近的音素分类器为基础的方法往往简化这一处理作为一个单一的音素的各种实现,绕过了复杂的模拟变音。受冻结自监督语音模型(S3M)特征的声学建模能力的启发,我们提出了MixGoP,这是一种利用高斯混合模型对具有多个子簇的音素分布进行建模的新方法。我们的实验表明,MixGoP在五个数据集中的四个数据集上实现了最先进的性能,包括构音障碍和非母语语音。我们的分析进一步表明,S3M功能比MFCC和Mel声谱图更有效地捕获了变音,突出了将MixGoP与S3M功能集成的好处。
摘要:Allophony refers to the variation in the phonetic realization of a phonemebased on its phonetic environment. Modeling allophones is crucial for atypicalpronunciation assessment, which involves distinguishing atypical from typicalpronunciations. However, recent phoneme classifier-based approaches oftensimplify this by treating various realizations as a single phoneme, bypassingthe complexity of modeling allophonic variation. Motivated by the acousticmodeling capabilities of frozen self-supervised speech model (S3M) features, wepropose MixGoP, a novel approach that leverages Gaussian mixture models tomodel phoneme distributions with multiple subclusters. Our experiments showthat MixGoP achieves state-of-the-art performance across four out of fivedatasets, including dysarthric and non-native speech. Our analysis furthersuggests that S3M features capture allophonic variation more effectively thanMFCCs and Mel spectrograms, highlighting the benefits of integrating MixGoPwith S3M features.
标题:自适应中心频率局部竞争语音算法
链接:https://arxiv.org/abs/2502.06989
备注:(C) 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. arXiv admin note: text overlap with arXiv:2409.08188
摘要:受神经系统启发的神经形态计算以其对效率和低功耗的关注彻底改变了信息处理。使用稀疏编码,这种模式提高了处理效率,这对于具有功率限制的边缘设备至关重要。局部竞争算法(LCA),适用于音频与Gammatone和Gammachirp滤波器组,提供了一个有效的稀疏编码方法的神经形态语音处理。自适应LCA(ALCA)通过动态调整调制参数进一步改进了该方法,从而提高了重建质量和稀疏性。本文介绍了一种增强的ALCA版本,ALCA中心频率(ALCA-CF),它动态地适应调制参数和中心频率,优化语音表示。评估表明,这种方法提高了重建质量和稀疏性,同时显着降低语音分类的功耗,而不影响分类精度,特别是在英特尔的Loihi 2神经形态芯片。
摘要:Neuromorphic computing, inspired by nervous systems, revolutionizesinformation processing with its focus on efficiency and low power consumption.Using sparse coding, this paradigm enhances processing efficiency, which iscrucial for edge devices with power constraints. The Locally CompetitiveAlgorithm (LCA), adapted for audio with Gammatone and Gammachirp filter banks,provides an efficient sparse coding method for neuromorphic speech processing.Adaptive LCA (ALCA) further refines this method by dynamically adjustingmodulation parameters, thereby improving reconstruction quality and sparsity.This paper introduces an enhanced ALCA version, the ALCA Central Frequency(ALCA-CF), which dynamically adapts both modulation parameters and centralfrequencies, optimizing the speech representation. Evaluations show that thisapproach improves reconstruction quality and sparsity while significantlyreducing the power consumption of speech classification, without compromisingclassification accuracy, particularly on Intel's Loihi 2 neuromorphic chip.
