本文经arXiv每日学术速递授权转载
链接:https://arxiv.org/abs/2501.03183
摘要:大多数当前的字幕系统使用基于特定设置数据的语言模型,例如通过Amazon Mechanical Turk进行的基于图像的字幕,限制了它们推广到其他模态分布和上下文的能力。这种限制阻碍了音频或视频字幕等任务的性能,这些任务需要不同的语义提示。解决这一挑战对于创建适用于各种现实环境的更具适应性和通用性的字幕框架至关重要。在这项工作中,我们介绍了一种方法,以适应字幕网络的语义的替代设置,如捕获可听度的音频字幕,它是至关重要的描述声音和它们的来源。我们的框架由两个主要组件组成:(i)包含语言模型(LM)的冻结字幕系统,以及(ii)引导字幕系统的文本分类器。分类器是在GPT-4自动生成的数据集上训练的,使用专门设计的定制提示来增强生成的字幕的关键方面。重要的是,该框架仅在推理过程中运行,无需进一步训练底层字幕模型。我们评估框架的各种模型和方式,重点是音频字幕,并报告有希望的结果。值得注意的是,当与现有的zero-shot音频字幕系统相结合时,我们的框架提高了其质量,并在zero-shot音频字幕中设置了最先进的性能。
摘要:Most current captioning systems use language models trained on data fromspecific settings, such as image-based captioning via Amazon Mechanical Turk,limiting their ability to generalize to other modality distributions andcontexts. This limitation hinders performance in tasks like audio or videocaptioning, where different semantic cues are needed. Addressing this challengeis crucial for creating more adaptable and versatile captioning frameworksapplicable across diverse real-world contexts. In this work, we introduce amethod to adapt captioning networks to the semantics of alternative settings,such as capturing audibility in audio captioning, where it is crucial todescribe sounds and their sources. Our framework consists of two maincomponents: (i) a frozen captioning system incorporating a language model (LM),and (ii) a text classifier that guides the captioning system. The classifier istrained on a dataset automatically generated by GPT-4, using tailored promptsspecifically designed to enhance key aspects of the generated captions.Importantly, the framework operates solely during inference, eliminating theneed for further training of the underlying captioning model. We evaluate theframework on various models and modalities, with a focus on audio captioning,and report promising results. Notably, when combined with an existing zero-shotaudio captioning system, our framework improves its quality and setsstate-of-the-art performance in zero-shot audio captioning.
标题:FaceSpeak:从不同风格的人体肖像中进行富有表现力且高质量的语音合成
链接:https://arxiv.org/abs/2501.03181
摘要:人类可以感知说话者的特征(例如,身份、性别、个性和情感),这通常与他们的声音风格一致。最近,视觉驱动的文本到语音(TTS)的学者们将他们的研究建立在真实人脸的基础上,从而限制了有效的语音合成应用于具有不同字符和图像风格的巨大潜在使用场景。为了解决这个问题,我们引入了一种新的FaceSpeak方法。它从各种各样的图像风格中提取突出的身份特征和情感表征。同时,它减轻了无关信息(例如,背景、服装和头发颜色等),从而产生与人物角色紧密一致的合成语音。此外,为了克服多模态TTS数据的稀缺性,我们设计了一个创新的数据集,即表达性多模态TTS,它经过精心策划和注释,以促进这一领域的研究。实验结果表明,我们提出的FaceSpeak可以产生令人满意的自然度和质量的纵向对齐的声音。
摘要:Humans can perceive speakers' characteristics (e.g., identity, gender,personality and emotion) by their appearance, which are generally aligned totheir voice style. Recently, vision-driven Text-to-speech (TTS) scholarsgrounded their investigations on real-person faces, thereby restrictingeffective speech synthesis from applying to vast potential usage scenarios withdiverse characters and image styles. To solve this issue, we introduce a novelFaceSpeak approach. It extracts salient identity characteristics and emotionalrepresentations from a wide variety of image styles. Meanwhile, it mitigatesthe extraneous information (e.g., background, clothing, and hair color, etc.),resulting in synthesized speech closely aligned with a character's persona.Furthermore, to overcome the scarcity of multi-modal TTS data, we have devisedan innovative dataset, namely Expressive Multi-Modal TTS, which is diligentlycurated and annotated to facilitate research in this domain. The experimentalresults demonstrate our proposed FaceSpeak can generate portrait-aligned voicewith satisfactory naturalness and quality.
标题:通过分层语言建模和预先训练的基于滚动的编码器进行钢琴抄写
链接:https://arxiv.org/abs/2501.03038
备注:Accepted by ICASSP 2025
摘要:自动音乐转录(AMT)旨在从原始音频中获取音符,通常使用具有钢琴滚动输出的帧级系统或具有音符级预测的基于语言模型(LM)的系统。然而,帧级系统需要手动阈值化,而基于LM的系统则难以处理长序列。在本文中,我们提出了一种混合方法结合预先训练的滚动编码器与LM解码器,以利用这两种方法的优势。此外,我们的方法采用了分层预测策略,首先预测的开始和音高,然后速度,最后偏移。分层预测策略通过将长序列分解为不同的层次来降低计算成本。在两个基准的基于滚动的编码器上进行评估,我们的方法在起始偏移速度F1分数上优于传统的钢琴滚动输出0.01和0.022,证明了其作为任意基于滚动的音乐转录编码器的性能增强插件的潜力。我们在https://github.com/yongyizang/AMT_train上发布了这项工作的代码。
摘要:Automatic Music Transcription (AMT), aiming to get musical notes from rawaudio, typically uses frame-level systems with piano-roll outputs or languagemodel (LM)-based systems with note-level predictions. However, frame-levelsystems require manual thresholding, while the LM-based systems struggle withlong sequences. In this paper, we propose a hybrid method combining pre-trainedroll-based encoders with an LM decoder to leverage the strengths of bothmethods. Besides, our approach employs a hierarchical prediction strategy,first predicting onset and pitch, then velocity, and finally offset. Thehierarchical prediction strategy reduces computational costs by breaking downlong sequences into different hierarchies. Evaluated on two benchmarkroll-based encoders, our method outperforms traditional piano-roll outputs 0.01and 0.022 in onset-offset-velocity F1 score, demonstrating its potential as aperformance-enhancing plug-in for arbitrary roll-based music transcriptionencoder. We release the code of this work athttps://github.com/yongyizang/AMT_train.
标题:SYKI-SRC:通过后处理创新和开源专业测试集推进歌唱声音转换
链接:https://arxiv.org/abs/2501.02953
备注:Accepted by ICASSP 2025
摘要:歌唱声音转换的目的是将原歌唱声音转换为目标歌唱者的声音,同时保留原歌词、旋律和各种声乐技巧。在本文中,我们提出了一个高保真的歌声转换系统。我们的系统构建在SVCC T02框架之上,由三个关键组件组成:特征提取器、语音转换器和后处理器。特征提取器利用ContentVec和Whisper模型来导出F0轮廓,并从输入的歌声中提取与说话者无关的语言特征。然后,语音转换器将提取的音色、F0和语言内容进行整合,以合成目标说话者的波形。后处理器通过简单而有效的信号处理直接从源中增加高频信息,以提高音频质量。由于缺乏一个标准化的专业数据集来评估表达歌唱转换系统,我们已经创建并公开了一个专门的测试集。比较评估表明,我们的系统实现了非常高的自然度,进一步的分析证实了我们提出的系统设计的有效性。
摘要:Singing voice conversion aims to transform a source singing voice into thatof a target singer while preserving the original lyrics, melody, and variousvocal techniques. In this paper, we propose a high-fidelity singing voiceconversion system. Our system builds upon the SVCC T02 framework and consistsof three key components: a feature extractor, a voice converter, and apost-processor. The feature extractor utilizes the ContentVec and Whispermodels to derive F0 contours and extract speaker-independent linguisticfeatures from the input singing voice. The voice converter then integrates theextracted timbre, F0, and linguistic content to synthesize the target speaker'swaveform. The post-processor augments high-frequency information directly fromthe source through simple and effective signal processing to enhance audioquality. Due to the lack of a standardized professional dataset for evaluatingexpressive singing conversion systems, we have created and made publiclyavailable a specialized test set. Comparative evaluations demonstrate that oursystem achieves a remarkably high level of naturalness, and further analysisconfirms the efficacy of our proposed system design.
标题:使用去噪扩散模型实现HRTI个性化
链接:https://arxiv.org/abs/2501.02871
备注:to appear in ICASSP 2025
摘要:头部相关传递函数(HRTF)在沉浸式音频场景中具有用于真实感渲染的基本应用。然而,它们是强烈的主体依赖性,因为它们根据耳朵,头部和躯干的形状而变化很大。因此,个性化程序需要准确的双耳渲染。近年来,去噪扩散概率模型(DDPMs),一类生成学习技术,已被应用于解决各种信号处理相关的问题。在本文中,我们提出了第一种方法,使用DDPM条件下的人体测量,以产生个性化的头部相关的脉冲响应(HRIR),时域表示的HRTF。结果表明,DDPM的HRTF个性化获得性能符合国家的最先进的模型的可行性。
摘要:Head-Related Transfer Functions (HRTFs) have fundamental applications forrealistic rendering in immersive audio scenarios. However, they are stronglysubject-dependent as they vary considerably depending on the shape of the ears,head and torso. Thus, personalization procedures are required for accuratebinaural rendering. Recently, Denoising Diffusion Probabilistic Models (DDPMs),a class of generative learning techniques, have been applied to solve a varietyof signal processing-related problems. In this paper, we propose a firstapproach for using DDPM conditioned on anthropometric measurements to generatepersonalized Head-Related Impulse Response (HRIR), the time-domainrepresentation of HRTF. The results show the feasibility of DDPMs for HRTFpersonalization obtaining performance in line with state-of-the-art models.
标题:Samba-asr利用结构化状态空间模型的最先进语音识别
链接:https://arxiv.org/abs/2501.02832
摘要:我们提出了Samba ASR,第一个国家的最先进的自动语音识别(ASR)模型,利用新的曼巴架构作为编码器和解码器,建立在状态空间模型(SSM)的基础上。与基于transformer的ASR模型(依赖于自注意机制来捕获依赖关系)不同,Samba ASR使用高效的状态空间动态有效地对局部和全局时间依赖关系进行建模,从而实现显着的性能提升。通过解决Transformers的局限性,例如输入长度的二次缩放和处理长范围依赖关系的困难,Samba ASR实现了卓越的准确性和效率。 实验结果表明,Samba ASR在各种标准基准测试中超越了现有的基于开源transformer的ASR模型,将其确立为ASR的最新技术水平。对基准数据集的广泛评估显示,字错误率(WER)有显着改善,即使在低资源场景下也具有竞争力的性能。此外,Mamba架构的计算效率和参数优化使Samba ASR成为各种ASR任务的可扩展和强大的解决方案。 我们的贡献包括: 一种新的Samba ASR架构,展示了SSM在语音序列处理中优于基于变换的模型。对公共基准进行全面评估,展示最先进的性能。分析计算效率、对噪声的鲁棒性和序列泛化。这项工作突出了Mamba SSM作为高效准确ASR的无变压器替代品的可行性。通过利用状态空间建模的进步,Samba ASR为ASR性能和未来研究设定了新的基准。
摘要:We propose Samba ASR, the first state-of-the-art Automatic Speech Recognition(ASR) model leveraging the novel Mamba architecture as both encoder anddecoder, built on the foundation of state-space models (SSMs). Unliketransformer-based ASR models, which rely on self-attention mechanisms tocapture dependencies, Samba ASR effectively models both local and globaltemporal dependencies using efficient state-space dynamics, achievingremarkable performance gains. By addressing the limitations of transformers,such as quadratic scaling with input length and difficulty in handlinglong-range dependencies, Samba ASR achieves superior accuracy and efficiency. Experimental results demonstrate that Samba ASR surpasses existingopen-source transformer-based ASR models across various standard benchmarks,establishing it as the new state of the art in ASR. Extensive evaluations onbenchmark datasets show significant improvements in Word Error Rate (WER), withcompetitive performance even in low-resource scenarios. Furthermore, thecomputational efficiency and parameter optimization of the Mamba architecturemake Samba ASR a scalable and robust solution for diverse ASR tasks. Our contributions include: A new Samba ASR architecture demonstrating the superiority of SSMs overtransformer-based models for speech sequence processing. A comprehensiveevaluation on public benchmarks showcasing state-of-the-art performance. Ananalysis of computational efficiency, robustness to noise, and sequencegeneralization. This work highlights the viability of Mamba SSMs as atransformer-free alternative for efficient and accurate ASR. By leveragingstate-space modeling advancements, Samba ASR sets a new benchmark for ASRperformance and future research.
标题:CCStereo:用于双耳音频生成的视听上下文和对比学习
链接:https://arxiv.org/abs/2501.02786
摘要:双耳音频生成(BAG)的目的是使用视觉提示将单声道音频转换为立体声音频,需要对空间和语义信息有深入的理解。然而,目前的模型有过度拟合房间环境的风险,并失去了细粒度的空间细节。在本文中,我们提出了一种新的视听双耳生成模型,其中包含一个视听条件归一化层,该层使用视觉上下文动态对齐目标差异音频特征的均值和方差,以及一种新的对比学习方法,通过从混洗的视觉特征中挖掘负样本来增强空间灵敏度。我们还引入了一种具有成本效益的方式来利用测试时增强视频数据,以提高性能。我们的方法实现了最先进的生成精度的公平播放和音乐立体声基准。
摘要:Binaural audio generation (BAG) aims to convert monaural audio to stereoaudio using visual prompts, requiring a deep understanding of spatial andsemantic information. However, current models risk overfitting to roomenvironments and lose fine-grained spatial details. In this paper, we propose anew audio-visual binaural generation model incorporating an audio-visualconditional normalisation layer that dynamically aligns the mean and varianceof the target difference audio features using visual context, along with a newcontrastive learning method to enhance spatial sensitivity by mining negativesamples from shuffled visual features. We also introduce a cost-efficient wayto utilise test-time augmentation in video data to enhance performance. Ourapproach achieves state-of-the-art generation accuracy on the FAIR-Play andMUSIC-Stereo benchmarks.
标题:使用勋伯格区域、巨型台阶和教堂模式的旋律协调系统
链接:https://arxiv.org/abs/2501.02642
摘要:像Microsoft Songsmith这样的系统通过最小化所有和弦变化中的不和谐来自动为旋律分配和弦和和声。虽然这产生和谐的音乐,但这不是练习音乐家所做的。在本文中,我描述了Harmonizer,一个旋律协调的原型系统。Harmonizer使用Schoenberg的区域图作为基础数据结构,允许使用几种不同的方法进行协调。因为图表揭示了弦间关系,所以和声可以被编程以强调期望的关系。在Harmonizer原型中,我还探索了最近的信号处理方法,使词曲作者能够通过唱歌或演奏乐器轻松输入旋律。原型Harmonizer可以在GitHub上找到,YouTube上有一段视频展示了它独特的和声效果,详见本文的结果部分。
摘要:Systems such as Microsoft Songsmith automatically assign chords and harmonyto a melody by minimizing the dissonance across all chord changes. Althoughthis produces harmonious music, it is not what practicing musicians do. In thispaper, I describe Harmonizer, a prototype system for melodic harmonization.Harmonizer uses Schoenberg's chart of regions as the underlying data structurethat allows harmonization using several different methods. Because the chartreveals inter-chordal relationships, the harmonizations may be programmed toemphasize desired relationships. In the prototype Harmonizer, I also explorerecent signal-processing methods that enable songwriters to easily input amelody by singing or by playing a musical instrument. The prototype Harmonizeris available on GitHub and a video demonstrating its distinctive harmonizationsis on YouTube as explained in the Results section of the paper.
标题:可以从缩略图中提取音乐的印象吗?
链接:https://arxiv.org/abs/2501.02511
备注:Accepted at NLP4MusA 2024
摘要:近年来,针对能够将自然语言句子作为输入的音乐检索和生成系统的机器学习模型的研究显著增加。然而,缺乏大规模的公开数据集,包括音乐数据及其相应的自然语言描述(称为音乐字幕)。特别是,非音乐信息,如适合听一首曲目的情况和听时引起的情绪是至关重要的描述音乐。这种类型的信息在现有的音乐字幕数据集中表现不足,这是因为直接从音乐数据中提取它存在挑战。为了解决这个问题,我们提出了一种方法,用于生成音乐字幕数据,其中包含从音乐缩略图推断的非音乐方面,并通过人工评估验证了我们的方法的有效性。此外,我们还创建了一个数据集,其中包含约360,000个包含非音乐方面的字幕。利用这个数据集,我们训练了一个音乐检索模型,并通过评估证明了它在音乐检索任务中的有效性。
摘要:In recent years, there has been a notable increase in research on machinelearning models for music retrieval and generation systems that are capable oftaking natural language sentences as inputs. However, there is a scarcity oflarge-scale publicly available datasets, consisting of music data and theircorresponding natural language descriptions known as music captions. Inparticular, non-musical information such as suitable situations for listeningto a track and the emotions elicited upon listening is crucial for describingmusic. This type of information is underrepresented in existing music captiondatasets due to the challenges associated with extracting it directly frommusic data. To address this issue, we propose a method for generating musiccaption data that incorporates non-musical aspects inferred from musicthumbnail images, and validated the effectiveness of our approach through humanevaluations. Additionally, we created a dataset with approximately 360,000captions containing non-musical aspects. Leveraging this dataset, we trained amusic retrieval model and demonstrated its effectiveness in music retrievaltasks through evaluation.
标题:使用真实语音训练的桥梁模块缩小预训练的语音增强和识别模型之间的差距
链接:https://arxiv.org/abs/2501.02452
备注:None
摘要:单通道语音增强(SE)造成的信息丢失或失真会影响自动语音识别(ASR)的性能。观测值相加(OA)是一种有效的后处理方法,通过平衡噪声和增强语音来提高ASR性能。确定OA系数至关重要。然而,目前监督的OA系数模块,称为桥接模块,仅利用模拟的带噪语音进行训练,这与真实的带噪语音有严重的失配。在本文中,我们提出的训练策略,训练的桥接模块与真正的嘈杂的语音。首先,选择DNSMOS来评估真实噪声语音的感知质量,而不需要相应的干净标签来训练桥接模块。在训练过程中引入了额外的约束,以进一步提高桥接模块的鲁棒性。ASR后端使用各种OA系数评估每个话语,以获得单词错误率(WER)。WER用于构建多维向量。该向量被引入到具有多任务学习的桥接模块中,并用于确定最佳OA系数。在CHiME-4数据集上的实验结果表明,与模拟数据训练的桥接模块相比,本文提出的方法都有显著的改进,尤其是在真实评估集下。
摘要:The information loss or distortion caused by single-channel speechenhancement (SE) harms the performance of automatic speech recognition (ASR).Observation addition (OA) is an effective post-processing method to improve ASRperformance by balancing noisy and enhanced speech. Determining the OAcoefficient is crucial. However, the currently supervised OA coefficientmodule, called the bridging module, only utilizes simulated noisy speech fortraining, which has a severe mismatch with real noisy speech. In this paper, wepropose training strategies to train the bridging module with real noisyspeech. First, DNSMOS is selected to evaluate the perceptual quality of realnoisy speech with no need for the corresponding clean label to train thebridging module. Additional constraints during training are introduced toenhance the robustness of the bridging module further. Each utterance isevaluated by the ASR back-end using various OA coefficients to obtain the worderror rates (WERs). The WERs are used to construct a multidimensional vector.This vector is introduced into the bridging module with multi-task learning andis used to determine the optimal OA coefficients. The experimental results onthe CHiME-4 dataset show that the proposed methods all had significantimprovement compared with the simulated data trained bridging module,especially under real evaluation sets.
标题:语音转文本的优先关注还是交叉关注?了实证研究
链接:https://arxiv.org/abs/2501.02370
备注:Submitted to ARR October 2024
摘要:随着大型语言模型(LLM)在NLP任务中取得的巨大成功,人们越来越有兴趣将其功能扩展到语音-最常见的交流形式。为了将语音集成到LLM中,一种有前途的方法是密集特征前置(DFP),其将投影的语音表示前置到文本表示,从而允许使用语音编码器进行端到端训练。然而,DFP通常需要将文本解码器连接到语音编码器。这就提出了一个问题,即拥有一个复杂的语音编码器对DFP的重要性,以及它的性能与标准的编码器-解码器(即交叉注意)架构相比如何。为了执行受控的架构比较,我们从头开始训练所有模型,而不是使用大型预训练模型,并使用可比数据和参数设置,在MuST-C v1.0和CoVoST 2数据集上测试语音到文本识别(ASR)和翻译(ST)。我们研究了语音编码器对DFP的影响。更重要的是,我们比较了各种配置下的DFP和交叉注意力,例如CTC压缩,序列级知识蒸馏,生成速度和单语,双语和多语言模型上的GPU内存占用。尽管DFP比交叉注意更普遍,但我们的总体结果并没有表明DFP有明显的优势。
摘要:Following the remarkable success of Large Language Models (LLMs) in NLPtasks, there is increasing interest in extending their capabilities to speech-- the most common form in communication. To integrate speech into LLMs, onepromising approach is dense feature prepending (DFP) which prepends theprojected speech representations to the textual representations, allowingend-to-end training with the speech encoder. However, DFP typically requiresconnecting a text decoder to a speech encoder. This raises questions about theimportance of having a sophisticated speech encoder for DFP, and how itsperformance compares with a standard encoder-decoder (i.e. cross-attention)architecture. In order to perform a controlled architectural comparison, wetrain all models from scratch, rather than using large pretrained models, anduse comparable data and parameter settings, testing speech-to-text recognition(ASR) and translation (ST) on MuST-C v1.0 and CoVoST2 datasets. We study theinfluence of a speech encoder in DFP. More importantly, we compare DFP andcross-attention under a variety of configurations, such as CTC compression,sequence-level knowledge distillation, generation speed and GPU memoryfootprint on monolingual, bilingual and multilingual models. Despite theprevalence of DFP over cross-attention, our overall results do not indicate aclear advantage of DFP.
标题:使用Transformer检测音乐表演错误
链接:https://arxiv.org/abs/2501.02030
备注:AAAI 2025
摘要:初学者经常很难找出他们表演中的具体错误,比如演奏不正确的音符或节奏。现有的音乐错误检测工具存在两个局限性:(1)现有方法依赖于自动对齐;因此,它们容易因对齐目标之间的小偏差而导致错误。(2)缺乏足够的数据来训练音乐错误检测模型,导致过度依赖于统计学。为了解决(1),我们提出了一种新的Transformer模型Polytune,它接受音频输入并输出带注释的乐谱。该模型可以进行端到端的训练,以通过潜在空间表示隐式地将表演音频与乐谱进行对齐和比较。为了解决(2),我们提出了一种新的数据生成技术,能够创建大规模的合成音乐错误数据集。我们的方法实现了64.1%的平均错误检测F1分数,在14种仪器上比之前的工作提高了40个百分点。此外,与现有的用于音乐错误检测的转录方法相比,我们的模型可以处理多种乐器。我们的源代码和数据集可以在https://github.com/ben2002chou/Polytune上找到。
摘要:Beginner musicians often struggle to identify specific errors in theirperformances, such as playing incorrect notes or rhythms. There are twolimitations in existing tools for music error detection: (1) Existingapproaches rely on automatic alignment; therefore, they are prone to errorscaused by small deviations between alignment targets.; (2) There is a lack ofsufficient data to train music error detection models, resulting inover-reliance on heuristics. To address (1), we propose a novel transformermodel, Polytune, that takes audio inputs and outputs annotated music scores.This model can be trained end-to-end to implicitly align and compareperformance audio with music scores through latent space representations. Toaddress (2), we present a novel data generation technique capable of creatinglarge-scale synthetic music error datasets. Our approach achieves a 64.1%average Error Detection F1 score, improving upon prior work by 40 percentagepoints across 14 instruments. Additionally, compared with existingtranscription methods repurposed for music error detection, our model canhandle multiple instruments. Our source code and datasets are available athttps://github.com/ben2002chou/Polytune.
标题:通过自我监督预训练进行噪音稳健的目标说话者语音活动检测
链接:https://arxiv.org/abs/2501.03184
备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing for possible publication. 12 pages, 4 figures, 5 tables
摘要:目标说话人语音活动检测(TS-VAD)是检测音频帧中来自已知目标说话人的语音的存在的任务。最近,基于深度神经网络的模型在这项任务中表现出良好的性能。然而,训练这些模型需要大量的标记数据,这是昂贵和耗时的,特别是如果泛化到看不见的环境是至关重要的。为了缓解这一问题,我们提出了一个因果关系的自监督学习(SSL)预训练框架,称为去噪自回归预测编码(DN-APC),以提高TS-VAD在噪声条件下的性能。我们还探讨了各种扬声器条件反射方法,并评估其性能在不同的噪声条件下。我们的实验表明,DN-APC可以提高噪音条件下的性能,总体提高约10%。2%,可见和不可见的噪音。此外,我们发现薄膜空调提供了最佳的整体性能。通过tSNE图的表示分析揭示了来自预训练的语音和非语音的鲁棒初始表示。这强调了SSL预训练在提高噪声环境中TS-VAD模型的鲁棒性和性能方面的有效性。
摘要:Target-Speaker Voice Activity Detection (TS-VAD) is the task of detecting thepresence of speech from a known target-speaker in an audio frame. Recently,deep neural network-based models have shown good performance in this task.However, training these models requires extensive labelled data, which iscostly and time-consuming to obtain, particularly if generalization to unseenenvironments is crucial. To mitigate this, we propose a causal, Self-SupervisedLearning (SSL) pretraining framework, called Denoising AutoregressivePredictive Coding (DN-APC), to enhance TS-VAD performance in noisy conditions.We also explore various speaker conditioning methods and evaluate theirperformance under different noisy conditions. Our experiments show that DN-APCimproves performance in noisy conditions, with a general improvement of approx.2% in both seen and unseen noise. Additionally, we find that FiLM conditioningprovides the best overall performance. Representation analysis via tSNE plotsreveals robust initial representations of speech and non-speech frompretraining. This underscores the effectiveness of SSL pretraining in improvingthe robustness and performance of TS-VAD models in noisy environments.
标题:用于音频精神障碍评估的频率感知增强网络
链接:https://arxiv.org/abs/2501.02516
摘要:抑郁症和注意力缺陷多动障碍(ADHD)是当今常见的心理健康挑战。在情感计算中,语音信号是评估精神障碍的有效生物标志物。目前的研究,依赖于劳动密集型手工制作的功能或简单的时间-频率表示,往往忽略了关键的细节,没有考虑到不同的频带和时间波动的差异影响。因此,我们提出了一种具有动态卷积的频率感知增强网络,用于抑郁症和多动症评估。在该方法中,频谱图被用作输入特征,并采用多尺度卷积,以帮助网络专注于与精神障碍相关的判别频带。动态卷积也被设计成基于多个卷积核的注意力动态地聚合它们,这些注意力与输入无关以捕获动态信息。最后,提出了一种特征增强块,以增强特征表示能力,并充分利用捕获的信息。在AVEC 2014和自录ADHD数据集上的实验结果证明了该方法的鲁棒性,对抑郁症严重程度的估计RMSE为9.23,对ADHD的检测准确率为89.8%。
摘要:Depression and Attention Deficit Hyperactivity Disorder (ADHD) stand out asthe common mental health challenges today. In affective computing, speechsignals serve as effective biomarkers for mental disorder assessment. Currentresearch, relying on labor-intensive hand-crafted features or simplistictime-frequency representations, often overlooks critical details by notaccounting for the differential impacts of various frequency bands and temporalfluctuations. Therefore, we propose a frequency-aware augmentation network withdynamic convolution for depression and ADHD assessment. In the proposed method,the spectrogram is used as the input feature and adopts a multi-scaleconvolution to help the network focus on discriminative frequency bands relatedto mental disorders. A dynamic convolution is also designed to aggregatemultiple convolution kernels dynamically based upon their attentions which areinput-independent to capture dynamic information. Finally, a featureaugmentation block is proposed to enhance the feature representation abilityand make full use of the captured information. Experimental results on AVEC2014 and self-recorded ADHD dataset prove the robustness of our method, an RMSEof 9.23 was attained for estimating depression severity, along with an accuracyof 89.8\% in detecting ADHD.
标题:有效的长语音序列建模用于时间域抑郁程度估计
链接:https://arxiv.org/abs/2501.02512
摘要:抑郁症严重影响情绪,思想和日常活动。最近的研究表明,语音信号包含有关抑郁症的重要线索,引发了人们对基于音频的深度学习方法来估计其严重程度的兴趣。然而,大多数方法都依赖于语音的时频表示,由于在执行时频投影(例如傅里叶变换和梅尔尺度变换)时丢失信息,这些方法最近因其局限性而受到批评。此外,将真实世界的语音分割成简短的间隔可能会丢失录音之间的关键互连。此外,这种方法可能无法充分反映现实世界的情况,因为抑郁症患者在对话和互动中经常会暂停和放慢速度。基于这些观察,我们提出了一种有效的方法,抑郁水平估计使用长语音信号在时域中。所提出的方法利用状态空间模型与基于双路径结构的长序列建模模块和时间外部注意模块相结合,以重建和增强隐藏在原始音频波形中的抑郁相关线索的检测。在AVEC2013和AVEC2014数据集上的实验结果显示,在捕获后续长序列抑郁线索方面取得了令人鼓舞的结果,并在最先进的水平上表现出出色的性能。
摘要:Depression significantly affects emotions, thoughts, and daily activities.Recent research indicates that speech signals contain vital cues aboutdepression, sparking interest in audio-based deep-learning methods forestimating its severity. However, most methods rely on time-frequencyrepresentations of speech which have recently been criticized for theirlimitations due to the loss of information when performing time-frequencyprojections, e.g. Fourier transform, and Mel-scale transformation. Furthermore,segmenting real-world speech into brief intervals risks losing criticalinterconnections between recordings. Additionally, such an approach may notadequately reflect real-world scenarios, as individuals with depression oftenpause and slow down in their conversations and interactions. Building on theseobservations, we present an efficient method for depression level estimationusing long speech signals in the time domain. The proposed method leverages astate space model coupled with the dual-path structure-based long sequencemodelling module and temporal external attention module to reconstruct andenhance the detection of depression-related cues hidden in the raw audiowaveforms. Experimental results on the AVEC2013 and AVEC2014 datasets showpromising results in capturing consequential long-sequence depression cues anddemonstrate outstanding performance over the state-of-the-art.
标题:通过熵控制抖动优化音频压缩
链接:https://arxiv.org/abs/2501.02293
备注:19 pages, 15 figures
摘要:本文探讨了音频压缩中的熵控制抖动技术,研究了标准和修改后的TPDF的应用,结合噪声整形和熵控制参数,在各种音频环境中,包括音高,响度,节奏和乐器的变化。感知质量指标,如VISQOL和STOI被用来评估性能。结果表明,基于TPDF的抖动始终优于RPDF,特别是在最佳alpha条件下,同时突出基于信号特性的性能变化。这些研究结果表明,使用各种TPDF分布的情境适当性。这项工作强调熵和感知保真度之间的权衡,熵控制抖动作为增强的音频压缩算法的基础的潜力提供见解。作为数字音频工作站插件的实际实现引入了可定制的抖动控制,为音频压缩算法的未来发展奠定了基础。
摘要:This paper explores entropy-controlled dithering techniques in audiocompression, examining the application of standard and modified TPDFs, combinedwith noise shaping and entropy-controlled parameters, across various audiocontexts, including pitch, loudness, rhythm, and instrumentation variations.Perceptual quality metrics such as VISQOL and STOI were used to evaluateperformance. The results demonstrate that TPDF-based dithering consistentlyoutperforms RPDF, particularly under optimal alpha conditions, whilehighlighting performance variability based on signal characteristics. Thesefindings suggest the situational appropriateness of using various TPDFdistributions. This work emphasizes the trade-off between entropy andperceptual fidelity, offering insights into the potential of entropy-controlleddithering as a foundation for enhanced audio compression algorithms. Apractical implementation as a Digital Audio Workstation plugin introducescustomizable dithering controls, laying the groundwork for future advancementsin audio compression algorithms.
标题:通过自我监督预训练进行噪音稳健的目标说话者语音活动检测
链接:https://arxiv.org/abs/2501.03184
备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing for possible publication. 12 pages, 4 figures, 5 tables
摘要:目标说话人语音活动检测(TS-VAD)是检测音频帧中来自已知目标说话人的语音的存在的任务。最近,基于深度神经网络的模型在这项任务中表现出良好的性能。然而,训练这些模型需要大量的标记数据,这是昂贵和耗时的,特别是如果泛化到看不见的环境是至关重要的。为了缓解这一问题,我们提出了一个因果关系的自监督学习(SSL)预训练框架,称为去噪自回归预测编码(DN-APC),以提高TS-VAD在噪声条件下的性能。我们还探讨了各种扬声器条件反射方法,并评估其性能在不同的噪声条件下。我们的实验表明,DN-APC提高了性能在嘈杂的条件下,与一般的改善约。2%,可见和不可见的噪音。此外,我们发现薄膜空调提供了最佳的整体性能。通过tSNE图的表示分析揭示了来自预训练的语音和非语音的鲁棒初始表示。这强调了SSL预训练在提高噪声环境中TS-VAD模型的鲁棒性和性能方面的有效性。
摘要:Target-Speaker Voice Activity Detection (TS-VAD) is the task of detecting thepresence of speech from a known target-speaker in an audio frame. Recently,deep neural network-based models have shown good performance in this task.However, training these models requires extensive labelled data, which iscostly and time-consuming to obtain, particularly if generalization to unseenenvironments is crucial. To mitigate this, we propose a causal, Self-SupervisedLearning (SSL) pretraining framework, called Denoising AutoregressivePredictive Coding (DN-APC), to enhance TS-VAD performance in noisy conditions.We also explore various speaker conditioning methods and evaluate theirperformance under different noisy conditions. Our experiments show that DN-APCimproves performance in noisy conditions, with a general improvement of approx.2% in both seen and unseen noise. Additionally, we find that FiLM conditioningprovides the best overall performance. Representation analysis via tSNE plotsreveals robust initial representations of speech and non-speech frompretraining. This underscores the effectiveness of SSL pretraining in improvingthe robustness and performance of TS-VAD models in noisy environments.
标题:户外和室内环境中移动图形处理器的单通道基于距离的源分离
链接:https://arxiv.org/abs/2501.03045
备注:Accepted by ICASSP2025. \c{opyright} 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component
摘要:本研究强调了在室外环境中探索基于距离的源分离(DSS)的意义。与现有的研究主要集中在室内环境不同,所提出的模型旨在捕捉室外音频源的独特特性。它采用了先进的技术,包括两阶段一致性块、线性关系感知自注意力(RSA)和TensorFlow Lite GPU代表。虽然线性RSA可能不会像二次RSA那样明确地捕获物理线索,但线性RSA增强了模型的上下文感知,从而提高了DSS的性能,这需要了解室外和室内环境中的物理线索。实验结果表明,该模型克服了现有方法的局限性,大大提高了能源效率和实时推理速度的移动设备上。
摘要:This study emphasizes the significance of exploring distance-based sourceseparation (DSS) in outdoor environments. Unlike existing studies thatprimarily focus on indoor settings, the proposed model is designed to capturethe unique characteristics of outdoor audio sources. It incorporates advancedtechniques, including a two-stage conformer block, a linear relation-awareself-attention (RSA), and a TensorFlow Lite GPU delegate. While the linear RSAmay not capture physical cues as explicitly as the quadratic RSA, the linearRSA enhances the model's context awareness, leading to improved performance onthe DSS that requires an understanding of physical cues in outdoor and indoorenvironments. The experimental results demonstrated that the proposed modelovercomes the limitations of existing approaches and considerably enhancesenergy efficiency and real-time inference speed on mobile devices.
标题:用于音频精神障碍评估的频率感知增强网络
链接:https://arxiv.org/abs/2501.02516
摘要:抑郁症和注意力缺陷多动障碍(ADHD)是当今常见的心理健康挑战。在情感计算中,语音信号是评估精神障碍的有效生物标志物。目前的研究,依赖于劳动密集型手工制作的功能或简单的时间-频率表示,往往忽略了关键的细节,没有考虑到不同的频带和时间波动的差异影响。因此,我们提出了一个具有动态卷积的频率感知增强网络,用于抑郁症和ADHD评估。在该方法中,频谱图被用作输入特征,并采用多尺度卷积,以帮助网络专注于与精神障碍相关的判别频带。动态卷积也被设计成基于多个卷积核的注意力动态地聚合它们,这些注意力与输入无关以捕获动态信息。最后,提出了一种特征增强块,以增强特征表示能力,并充分利用捕获的信息。在AVEC 2014和自录ADHD数据集上的实验结果证明了该方法的鲁棒性,对抑郁症严重程度的估计RMSE为9.23,对ADHD的检测准确率为89.8%。
摘要:Depression and Attention Deficit Hyperactivity Disorder (ADHD) stand out asthe common mental health challenges today. In affective computing, speechsignals serve as effective biomarkers for mental disorder assessment. Currentresearch, relying on labor-intensive hand-crafted features or simplistictime-frequency representations, often overlooks critical details by notaccounting for the differential impacts of various frequency bands and temporalfluctuations. Therefore, we propose a frequency-aware augmentation network withdynamic convolution for depression and ADHD assessment. In the proposed method,the spectrogram is used as the input feature and adopts a multi-scaleconvolution to help the network focus on discriminative frequency bands relatedto mental disorders. A dynamic convolution is also designed to aggregatemultiple convolution kernels dynamically based upon their attentions which areinput-independent to capture dynamic information. Finally, a featureaugmentation block is proposed to enhance the feature representation abilityand make full use of the captured information. Experimental results on AVEC2014 and self-recorded ADHD dataset prove the robustness of our method, an RMSEof 9.23 was attained for estimating depression severity, along with an accuracyof 89.8\% in detecting ADHD.
标题:有效的长语音序列建模用于时间域抑郁程度估计
链接:https://arxiv.org/abs/2501.02512
摘要:抑郁症严重影响情绪,思想和日常活动。最近的研究表明,语音信号包含有关抑郁症的重要线索,引发了人们对基于音频的深度学习方法来估计其严重程度的兴趣。然而,大多数方法都依赖于语音的时频表示,由于在执行时频投影(例如傅里叶变换和梅尔尺度变换)时丢失信息,这些方法最近因其局限性而受到批评。此外,将真实世界的语音分割成简短的间隔可能会丢失录音之间的关键互连。此外,这种方法可能无法充分反映现实世界的情况,因为抑郁症患者在对话和互动中经常会暂停和放慢速度。基于这些观察,我们提出了一种有效的方法,抑郁水平估计使用长语音信号在时域中。所提出的方法利用状态空间模型与基于双路径结构的长序列建模模块和时间外部注意模块相结合,以重建和增强隐藏在原始音频波形中的抑郁相关线索的检测。在AVEC2013和AVEC2014数据集上的实验结果显示,在捕获后续长序列抑郁线索方面取得了令人鼓舞的结果,并在最先进的水平上表现出出色的性能。
摘要:Depression significantly affects emotions, thoughts, and daily activities.Recent research indicates that speech signals contain vital cues aboutdepression, sparking interest in audio-based deep-learning methods forestimating its severity. However, most methods rely on time-frequencyrepresentations of speech which have recently been criticized for theirlimitations due to the loss of information when performing time-frequencyprojections, e.g. Fourier transform, and Mel-scale transformation. Furthermore,segmenting real-world speech into brief intervals risks losing criticalinterconnections between recordings. Additionally, such an approach may notadequately reflect real-world scenarios, as individuals with depression oftenpause and slow down in their conversations and interactions. Building on theseobservations, we present an efficient method for depression level estimationusing long speech signals in the time domain. The proposed method leverages astate space model coupled with the dual-path structure-based long sequencemodelling module and temporal external attention module to reconstruct andenhance the detection of depression-related cues hidden in the raw audiowaveforms. Experimental results on the AVEC2013 and AVEC2014 datasets showpromising results in capturing consequential long-sequence depression cues anddemonstrate outstanding performance over the state-of-the-art.
标题:通过熵控制抖动优化音频压缩
链接:https://arxiv.org/abs/2501.02293
备注:19 pages, 15 figures
摘要:本文探讨了音频压缩中的熵控制抖动技术,研究了标准和修改后的TPDF的应用,结合噪声整形和熵控制参数,在各种音频环境中,包括音高,响度,节奏和乐器的变化。感知质量指标,如VISQOL和STOI被用来评估性能。结果表明,基于TPDF的抖动始终优于RPDF,特别是在最佳alpha条件下,同时突出基于信号特性的性能变化。这些研究结果表明,使用各种TPDF分布的情境适当性。这项工作强调熵和感知保真度之间的权衡,熵控制抖动作为增强的音频压缩算法的基础的潜力提供见解。作为数字音频工作站插件的实际实现引入了可定制的抖动控制,为音频压缩算法的未来发展奠定了基础。
摘要:This paper explores entropy-controlled dithering techniques in audiocompression, examining the application of standard and modified TPDFs, combinedwith noise shaping and entropy-controlled parameters, across various audiocontexts, including pitch, loudness, rhythm, and instrumentation variations.Perceptual quality metrics such as VISQOL and STOI were used to evaluateperformance. The results demonstrate that TPDF-based dithering consistentlyoutperforms RPDF, particularly under optimal alpha conditions, whilehighlighting performance variability based on signal characteristics. Thesefindings suggest the situational appropriateness of using various TPDFdistributions. This work emphasizes the trade-off between entropy andperceptual fidelity, offering insights into the potential of entropy-controlleddithering as a foundation for enhanced audio compression algorithms. Apractical implementation as a Digital Audio Workstation plugin introducescustomizable dithering controls, laying the groundwork for future advancementsin audio compression algorithms.
标题:多模式机器学习可以预测视频会议的流畅性和乐趣
链接:https://arxiv.org/abs/2501.03190
备注:ICASSP 2025
摘要:视频会议现在是专业和非正式环境中常用的沟通方式,但它往往缺乏面对面交谈的流畅性和乐趣。这项研究利用多模态机器学习来预测视频会议中的负面体验。我们从RoomReader语料库中采样了数千个短片段,提取音频嵌入,面部动作和身体运动特征,以训练模型来识别低会话流动性,低享受,并对会话事件进行分类(反向引导,中断或间隙)。我们最好的模型在保持视频会议会话上实现了高达0.87的ROC-AUC,其中领域通用音频功能证明是最关键的。这项工作表明,多模态音视频信号可以有效地预测高层次的主观会话结果。此外,这是对视频会议用户体验研究的贡献,表明多模态机器学习可用于识别罕见的负面用户体验时刻,以供进一步研究或缓解。
摘要:Videoconferencing is now a frequent mode of communication in bothprofessional and informal settings, yet it often lacks the fluidity andenjoyment of in-person conversation. This study leverages multimodal machinelearning to predict moments of negative experience in videoconferencing. Wesampled thousands of short clips from the RoomReader corpus, extracting audioembeddings, facial actions, and body motion features to train models foridentifying low conversational fluidity, low enjoyment, and classifyingconversational events (backchanneling, interruption, or gap). Our best modelsachieved an ROC-AUC of up to 0.87 on hold-out videoconference sessions, withdomain-general audio features proving most critical. This work demonstratesthat multimodal audio-video signals can effectively predict high-levelsubjective conversational outcomes. In addition, this is a contribution toresearch on videoconferencing user experience by showing that multimodalmachine learning can be used to identify rare moments of negative userexperience for further study or mitigation.
标题:跨模式的分类器引导字幕
链接:https://arxiv.org/abs/2501.03183
摘要:大多数当前的字幕系统使用基于特定设置数据的语言模型,例如通过Amazon Mechanical Turk进行的基于图像的字幕,限制了它们推广到其他模态分布和上下文的能力。这种限制阻碍了音频或视频字幕等任务的性能,这些任务需要不同的语义提示。解决这一挑战对于创建适用于各种现实环境的更具适应性和通用性的字幕框架至关重要。在这项工作中,我们介绍了一种方法,以适应字幕网络的语义的替代设置,如捕获可听度的音频字幕,它是至关重要的描述声音和它们的来源。我们的框架包括两个主要组成部分:(i)一个冻结的字幕系统,包括一个语言模型(LM),和(ii)一个文本分类器,指导字幕系统。分类器是在GPT-4自动生成的数据集上训练的,使用专门设计的定制提示来增强生成的字幕的关键方面。重要的是,该框架仅在推理过程中运行,无需进一步训练底层字幕模型。我们评估框架的各种模型和方式,重点是音频字幕,并报告有希望的结果。值得注意的是,当与现有的zero-shot音频字幕系统相结合时,我们的框架提高了其质量,并在zero-shot音频字幕中设置了最先进的性能。
摘要:Most current captioning systems use language models trained on data fromspecific settings, such as image-based captioning via Amazon Mechanical Turk,limiting their ability to generalize to other modality distributions andcontexts. This limitation hinders performance in tasks like audio or videocaptioning, where different semantic cues are needed. Addressing this challengeis crucial for creating more adaptable and versatile captioning frameworksapplicable across diverse real-world contexts. In this work, we introduce amethod to adapt captioning networks to the semantics of alternative settings,such as capturing audibility in audio captioning, where it is crucial todescribe sounds and their sources. Our framework consists of two maincomponents: (i) a frozen captioning system incorporating a language model (LM),and (ii) a text classifier that guides the captioning system. The classifier istrained on a dataset automatically generated by GPT-4, using tailored promptsspecifically designed to enhance key aspects of the generated captions.Importantly, the framework operates solely during inference, eliminating theneed for further training of the underlying captioning model. We evaluate theframework on various models and modalities, with a focus on audio captioning,and report promising results. Notably, when combined with an existing zero-shotaudio captioning system, our framework improves its quality and setsstate-of-the-art performance in zero-shot audio captioning.
标题:FaceSpeak:从不同风格的人体肖像中进行富有表现力且高质量的语音合成
链接:https://arxiv.org/abs/2501.03181
摘要:人类可以感知说话者的特征(例如,身份、性别、个性和情感),这通常与他们的声音风格一致。最近,视觉驱动的文本到语音(TTS)的学者们将他们的研究建立在真实人脸的基础上,从而限制了有效的语音合成应用于具有不同字符和图像风格的巨大潜在使用场景。为了解决这个问题,我们引入了一种新的FaceSpeak方法。它从各种各样的图像风格中提取突出的身份特征和情感表征。同时,它减轻了无关信息(例如,背景、服装和头发颜色等),从而产生与人物角色紧密一致的合成语音。此外,为了克服多模态TTS数据的稀缺性,我们设计了一个创新的数据集,即表达性多模态TTS,它经过精心策划和注释,以促进这一领域的研究。实验结果表明,我们提出的FaceSpeak可以产生令人满意的自然度和质量的纵向对齐的声音。
摘要:Humans can perceive speakers' characteristics (e.g., identity, gender,personality and emotion) by their appearance, which are generally aligned totheir voice style. Recently, vision-driven Text-to-speech (TTS) scholarsgrounded their investigations on real-person faces, thereby restrictingeffective speech synthesis from applying to vast potential usage scenarios withdiverse characters and image styles. To solve this issue, we introduce a novelFaceSpeak approach. It extracts salient identity characteristics and emotionalrepresentations from a wide variety of image styles. Meanwhile, it mitigatesthe extraneous information (e.g., background, clothing, and hair color, etc.),resulting in synthesized speech closely aligned with a character's persona.Furthermore, to overcome the scarcity of multi-modal TTS data, we have devisedan innovative dataset, namely Expressive Multi-Modal TTS, which is diligentlycurated and annotated to facilitate research in this domain. The experimentalresults demonstrate our proposed FaceSpeak can generate portrait-aligned voicewith satisfactory naturalness and quality.
标题:通过分层语言建模和预先训练的基于滚动的编码器进行钢琴抄写
链接:https://arxiv.org/abs/2501.03038
备注:Accepted by ICASSP 2025
摘要:自动音乐转录(AMT)旨在从原始音频中获取音符,通常使用具有钢琴滚动输出的帧级系统或具有音符级预测的基于语言模型(LM)的系统。然而,帧级系统需要手动阈值化,而基于LM的系统则难以处理长序列。在本文中,我们提出了一种混合方法结合预先训练的滚动编码器与LM解码器,以利用这两种方法的优势。此外,我们的方法采用分层预测策略,首先预测起点和音高,然后预测速度,最后预测偏移。分层预测策略通过将长序列分解为不同的层次来降低计算成本。在两个基准的基于滚动的编码器上进行评估,我们的方法在起始偏移速度F1分数上优于传统的钢琴滚动输出0.01和0.022,证明了其作为任意基于滚动的音乐转录编码器的性能增强插件的潜力。我们在https://github.com/yongyizang/AMT_train上发布了这项工作的代码。
摘要:Automatic Music Transcription (AMT), aiming to get musical notes from rawaudio, typically uses frame-level systems with piano-roll outputs or languagemodel (LM)-based systems with note-level predictions. However, frame-levelsystems require manual thresholding, while the LM-based systems struggle withlong sequences. In this paper, we propose a hybrid method combining pre-trainedroll-based encoders with an LM decoder to leverage the strengths of bothmethods. Besides, our approach employs a hierarchical prediction strategy,first predicting onset and pitch, then velocity, and finally offset. Thehierarchical prediction strategy reduces computational costs by breaking downlong sequences into different hierarchies. Evaluated on two benchmarkroll-based encoders, our method outperforms traditional piano-roll outputs 0.01and 0.022 in onset-offset-velocity F1 score, demonstrating its potential as aperformance-enhancing plug-in for arbitrary roll-based music transcriptionencoder. We release the code of this work athttps://github.com/yongyizang/AMT_train.
标题:SYKI-SRC:通过后处理创新和开源专业测试集推进歌唱声音转换
链接:https://arxiv.org/abs/2501.02953
备注:Accepted by ICASSP 2025
摘要:歌唱声音转换的目的是将原歌唱声音转换为目标歌唱声音,同时保留原歌词、旋律和各种声乐技巧。在本文中,我们提出了一个高保真的歌声转换系统。我们的系统建立在SVCC T02框架上,由三个关键组件组成:特征提取器,语音转换器和后处理器。特征提取器利用ContentVec和Whisper模型来导出F0轮廓,并从输入的歌声中提取与说话者无关的语言特征。然后,语音转换器将提取的音色、F0和语言内容进行整合,以合成目标说话者的波形。后处理器通过简单而有效的信号处理直接从源中增加高频信息,以提高音频质量。由于缺乏一个标准化的专业数据集来评估表达歌唱转换系统,我们已经创建并公开了一个专门的测试集。比较评估表明,我们的系统实现了非常高的自然度,进一步的分析证实了我们提出的系统设计的有效性。
摘要:Singing voice conversion aims to transform a source singing voice into thatof a target singer while preserving the original lyrics, melody, and variousvocal techniques. In this paper, we propose a high-fidelity singing voiceconversion system. Our system builds upon the SVCC T02 framework and consistsof three key components: a feature extractor, a voice converter, and apost-processor. The feature extractor utilizes the ContentVec and Whispermodels to derive F0 contours and extract speaker-independent linguisticfeatures from the input singing voice. The voice converter then integrates theextracted timbre, F0, and linguistic content to synthesize the target speaker'swaveform. The post-processor augments high-frequency information directly fromthe source through simple and effective signal processing to enhance audioquality. Due to the lack of a standardized professional dataset for evaluatingexpressive singing conversion systems, we have created and made publiclyavailable a specialized test set. Comparative evaluations demonstrate that oursystem achieves a remarkably high level of naturalness, and further analysisconfirms the efficacy of our proposed system design.
标题:使用去噪扩散模型实现HRTI个性化
链接:https://arxiv.org/abs/2501.02871
备注:to appear in ICASSP 2025
摘要:头部相关传递函数(HRTF)在沉浸式音频场景中具有用于真实感渲染的基本应用。然而,它们是强烈的主体依赖性,因为它们根据耳朵,头部和躯干的形状而变化很大。因此,个性化程序需要准确的双耳渲染。近年来,去噪扩散概率模型(DDPMs),一类生成学习技术,已被应用于解决各种信号处理相关的问题。在本文中,我们提出了第一种方法,使用DDPM条件下的人体测量,以产生个性化的头部相关的脉冲响应(HRIR),时域表示的HRTF。结果表明,DDPM的HRTF个性化获得性能符合国家的最先进的模型的可行性。
摘要:Head-Related Transfer Functions (HRTFs) have fundamental applications forrealistic rendering in immersive audio scenarios. However, they are stronglysubject-dependent as they vary considerably depending on the shape of the ears,head and torso. Thus, personalization procedures are required for accuratebinaural rendering. Recently, Denoising Diffusion Probabilistic Models (DDPMs),a class of generative learning techniques, have been applied to solve a varietyof signal processing-related problems. In this paper, we propose a firstapproach for using DDPM conditioned on anthropometric measurements to generatepersonalized Head-Related Impulse Response (HRIR), the time-domainrepresentation of HRTF. The results show the feasibility of DDPMs for HRTFpersonalization obtaining performance in line with state-of-the-art models.
标题:Samba-asr利用结构化状态空间模型的最先进语音识别
链接:https://arxiv.org/abs/2501.02832
摘要:我们提出了Samba ASR,第一个国家的最先进的自动语音识别(ASR)模型,利用新的曼巴架构作为编码器和解码器,建立在状态空间模型(SSM)的基础上。与基于transformer的ASR模型(依赖于自注意机制来捕获依赖关系)不同,Samba ASR使用高效的状态空间动态有效地对局部和全局时间依赖关系进行建模,从而实现显着的性能提升。通过解决Transformers的局限性,例如输入长度的二次缩放和处理长范围依赖关系的困难,Samba ASR实现了卓越的准确性和效率。 实验结果表明,Samba ASR在各种标准基准测试中超越了现有的基于开源transformer的ASR模型,将其确立为ASR的最新技术水平。对基准数据集的广泛评估显示,字错误率(WER)有显着改善,即使在低资源场景下也具有竞争力的性能。此外,Mamba架构的计算效率和参数优化使Samba ASR成为各种ASR任务的可扩展和强大的解决方案。 我们的贡献包括: 一种新的Samba ASR架构,展示了SSM在语音序列处理方面优于基于变换器的模型。对公共基准进行全面评估,展示最先进的性能。分析计算效率、对噪声的鲁棒性和序列泛化。这项工作突出了Mamba SSM作为高效准确ASR的无变压器替代品的可行性。通过利用状态空间建模的进步,Samba ASR为ASR性能和未来研究设定了新的基准。
摘要:We propose Samba ASR, the first state-of-the-art Automatic Speech Recognition(ASR) model leveraging the novel Mamba architecture as both encoder anddecoder, built on the foundation of state-space models (SSMs). Unliketransformer-based ASR models, which rely on self-attention mechanisms tocapture dependencies, Samba ASR effectively models both local and globaltemporal dependencies using efficient state-space dynamics, achievingremarkable performance gains. By addressing the limitations of transformers,such as quadratic scaling with input length and difficulty in handlinglong-range dependencies, Samba ASR achieves superior accuracy and efficiency. Experimental results demonstrate that Samba ASR surpasses existingopen-source transformer-based ASR models across various standard benchmarks,establishing it as the new state of the art in ASR. Extensive evaluations onbenchmark datasets show significant improvements in Word Error Rate (WER), withcompetitive performance even in low-resource scenarios. Furthermore, thecomputational efficiency and parameter optimization of the Mamba architecturemake Samba ASR a scalable and robust solution for diverse ASR tasks. Our contributions include: A new Samba ASR architecture demonstrating the superiority of SSMs overtransformer-based models for speech sequence processing. A comprehensiveevaluation on public benchmarks showcasing state-of-the-art performance. Ananalysis of computational efficiency, robustness to noise, and sequencegeneralization. This work highlights the viability of Mamba SSMs as atransformer-free alternative for efficient and accurate ASR. By leveragingstate-space modeling advancements, Samba ASR sets a new benchmark for ASRperformance and future research.
标题:CCStereo:用于双耳音频生成的视听上下文和对比学习
链接:https://arxiv.org/abs/2501.02786
摘要:双耳音频生成(BAG)的目的是使用视觉提示将单声道音频转换为立体声音频,需要对空间和语义信息有深入的理解。然而,目前的模型有过度拟合房间环境的风险,并失去了细粒度的空间细节。在本文中,我们提出了一种新的视听双耳生成模型,其中包含一个视听条件归一化层,该层使用视觉上下文动态对齐目标差异音频特征的均值和方差,以及一种新的对比学习方法,通过从混洗的视觉特征中挖掘负样本来增强空间灵敏度。我们还引入了一种具有成本效益的方式来利用测试时增强视频数据,以提高性能。我们的方法实现了最先进的生成精度的公平播放和音乐立体声基准。
摘要:Binaural audio generation (BAG) aims to convert monaural audio to stereoaudio using visual prompts, requiring a deep understanding of spatial andsemantic information. However, current models risk overfitting to roomenvironments and lose fine-grained spatial details. In this paper, we propose anew audio-visual binaural generation model incorporating an audio-visualconditional normalisation layer that dynamically aligns the mean and varianceof the target difference audio features using visual context, along with a newcontrastive learning method to enhance spatial sensitivity by mining negativesamples from shuffled visual features. We also introduce a cost-efficient wayto utilise test-time augmentation in video data to enhance performance. Ourapproach achieves state-of-the-art generation accuracy on the FAIR-Play andMUSIC-Stereo benchmarks.
标题:使用勋伯格区域、巨型台阶和教堂模式的旋律协调系统
链接:https://arxiv.org/abs/2501.02642
摘要:像Microsoft Songsmith这样的系统通过最小化所有和弦变化中的不和谐来自动为旋律分配和弦和和声。虽然这产生和谐的音乐,但这不是练习音乐家所做的。在本文中,我描述了Harmonizer,一个旋律协调的原型系统。Harmonizer使用Schoenberg的区域图作为基础数据结构,允许使用几种不同的方法进行协调。因为图表揭示了弦间关系,所以和声可以被编程以强调期望的关系。在Harmonizer原型中,我还探索了最近的信号处理方法,使词曲作者能够通过唱歌或演奏乐器轻松输入旋律。原型Harmonizer可以在GitHub上找到,YouTube上有一段视频展示了它独特的和声效果,详见本文的结果部分。
摘要:Systems such as Microsoft Songsmith automatically assign chords and harmonyto a melody by minimizing the dissonance across all chord changes. Althoughthis produces harmonious music, it is not what practicing musicians do. In thispaper, I describe Harmonizer, a prototype system for melodic harmonization.Harmonizer uses Schoenberg's chart of regions as the underlying data structurethat allows harmonization using several different methods. Because the chartreveals inter-chordal relationships, the harmonizations may be programmed toemphasize desired relationships. In the prototype Harmonizer, I also explorerecent signal-processing methods that enable songwriters to easily input amelody by singing or by playing a musical instrument. The prototype Harmonizeris available on GitHub and a video demonstrating its distinctive harmonizationsis on YouTube as explained in the Results section of the paper.
标题:可以从缩略图中提取音乐的印象吗?
链接:https://arxiv.org/abs/2501.02511
备注:Accepted at NLP4MusA 2024
摘要:近年来,针对能够将自然语言句子作为输入的音乐检索和生成系统的机器学习模型的研究显著增加。然而,缺乏大规模的公开数据集,包括音乐数据及其相应的自然语言描述(称为音乐字幕)。特别是,非音乐信息,如适合听一首曲目的情况和听时引起的情绪是至关重要的描述音乐。这种类型的信息在现有的音乐字幕数据集中表现不足,这是因为直接从音乐数据中提取它存在挑战。为了解决这个问题,我们提出了一种方法,用于生成音乐字幕数据,其中包含从音乐缩略图推断的非音乐方面,并通过人工评估验证了我们的方法的有效性。此外,我们还创建了一个数据集,其中包含约360,000个包含非音乐方面的字幕。利用这个数据集,我们训练了一个音乐检索模型,并通过评估证明了它在音乐检索任务中的有效性。
摘要:In recent years, there has been a notable increase in research on machinelearning models for music retrieval and generation systems that are capable oftaking natural language sentences as inputs. However, there is a scarcity oflarge-scale publicly available datasets, consisting of music data and theircorresponding natural language descriptions known as music captions. Inparticular, non-musical information such as suitable situations for listeningto a track and the emotions elicited upon listening is crucial for describingmusic. This type of information is underrepresented in existing music captiondatasets due to the challenges associated with extracting it directly frommusic data. To address this issue, we propose a method for generating musiccaption data that incorporates non-musical aspects inferred from musicthumbnail images, and validated the effectiveness of our approach through humanevaluations. Additionally, we created a dataset with approximately 360,000captions containing non-musical aspects. Leveraging this dataset, we trained amusic retrieval model and demonstrated its effectiveness in music retrievaltasks through evaluation.
标题:使用真实语音训练的桥梁模块缩小预训练的语音增强和识别模型之间的差距
链接:https://arxiv.org/abs/2501.02452
备注:None
摘要:单通道语音增强(SE)造成的信息丢失或失真会影响自动语音识别(ASR)的性能。观察添加(OA)是一种有效的后处理方法,可以通过平衡带噪语音和增强语音来提高ASR性能。确定OA系数至关重要。然而,目前监督的OA系数模块,称为桥接模块,仅利用模拟的带噪语音进行训练,这与真实的带噪语音有严重的失配。在本文中,我们提出的训练策略,训练的桥接模块与真正的嘈杂的语音。首先,选择DNSMOS来评估真实噪声语音的感知质量,而不需要相应的干净标签来训练桥接模块。在训练过程中引入了额外的约束,以进一步提高桥接模块的鲁棒性。ASR后端使用各种OA系数评估每个话语,以获得单词错误率(WER)。WER用于构建多维向量。该向量被引入到具有多任务学习的桥接模块中,并用于确定最佳OA系数。在CHiME-4数据集上的实验结果表明,与模拟数据训练的桥接模块相比,本文提出的方法都有显著的改进,尤其是在真实评估集下。
摘要:The information loss or distortion caused by single-channel speechenhancement (SE) harms the performance of automatic speech recognition (ASR).Observation addition (OA) is an effective post-processing method to improve ASRperformance by balancing noisy and enhanced speech. Determining the OAcoefficient is crucial. However, the currently supervised OA coefficientmodule, called the bridging module, only utilizes simulated noisy speech fortraining, which has a severe mismatch with real noisy speech. In this paper, wepropose training strategies to train the bridging module with real noisyspeech. First, DNSMOS is selected to evaluate the perceptual quality of realnoisy speech with no need for the corresponding clean label to train thebridging module. Additional constraints during training are introduced toenhance the robustness of the bridging module further. Each utterance isevaluated by the ASR back-end using various OA coefficients to obtain the worderror rates (WERs). The WERs are used to construct a multidimensional vector.This vector is introduced into the bridging module with multi-task learning andis used to determine the optimal OA coefficients. The experimental results onthe CHiME-4 dataset show that the proposed methods all had significantimprovement compared with the simulated data trained bridging module,especially under real evaluation sets.
标题:语音转文本的优先关注还是交叉关注?了实证研究
链接:https://arxiv.org/abs/2501.02370
备注:Submitted to ARR October 2024
摘要:随着大型语言模型(LLM)在NLP任务中取得的巨大成功,人们越来越有兴趣将其功能扩展到语音-最常见的交流形式。为了将语音集成到LLM中,一种有前途的方法是密集特征前置(DFP),其将投影的语音表示前置到文本表示,从而允许使用语音编码器进行端到端训练。然而,DFP通常需要将文本解码器连接到语音编码器。这就提出了一个问题,即拥有一个复杂的语音编码器对DFP的重要性,以及它的性能与标准的编码器-解码器(即交叉注意)架构相比如何。为了执行受控的架构比较,我们从头开始训练所有模型,而不是使用大型预训练模型,并使用可比数据和参数设置,在MuST-C v1.0和CoVoST 2数据集上测试语音到文本识别(ASR)和翻译(ST)。我们研究了语音编码器对DFP的影响。更重要的是,我们比较了各种配置下的DFP和交叉注意力,例如CTC压缩,序列级知识蒸馏,生成速度和单语,双语和多语言模型上的GPU内存占用。尽管DFP比交叉注意更普遍,但我们的总体结果并没有表明DFP有明显的优势。
摘要:Following the remarkable success of Large Language Models (LLMs) in NLPtasks, there is increasing interest in extending their capabilities to speech-- the most common form in communication. To integrate speech into LLMs, onepromising approach is dense feature prepending (DFP) which prepends theprojected speech representations to the textual representations, allowingend-to-end training with the speech encoder. However, DFP typically requiresconnecting a text decoder to a speech encoder. This raises questions about theimportance of having a sophisticated speech encoder for DFP, and how itsperformance compares with a standard encoder-decoder (i.e. cross-attention)architecture. In order to perform a controlled architectural comparison, wetrain all models from scratch, rather than using large pretrained models, anduse comparable data and parameter settings, testing speech-to-text recognition(ASR) and translation (ST) on MuST-C v1.0 and CoVoST2 datasets. We study theinfluence of a speech encoder in DFP. More importantly, we compare DFP andcross-attention under a variety of configurations, such as CTC compression,sequence-level knowledge distillation, generation speed and GPU memoryfootprint on monolingual, bilingual and multilingual models. Despite theprevalence of DFP over cross-attention, our overall results do not indicate aclear advantage of DFP.
标题:使用Transformer检测音乐表演错误
链接:https://arxiv.org/abs/2501.02030
备注:AAAI 2025
摘要:初学者经常很难找出他们表演中的具体错误,比如演奏不正确的音符或节奏。现有的音乐错误检测工具存在两个局限性:(1)现有方法依赖于自动对齐;因此,它们容易因对齐目标之间的小偏差而导致错误。(2)缺乏足够的数据来训练音乐错误检测模型,导致过度依赖于统计学。为了解决(1),我们提出了一种新的Transformer模型Polytune,它接受音频输入并输出带注释的乐谱。该模型可以进行端到端的训练,以通过潜在空间表示隐式地将表演音频与乐谱进行对齐和比较。为了解决(2),我们提出了一种新的数据生成技术,能够创建大规模的合成音乐错误数据集。我们的方法实现了64.1%的平均错误检测F1分数,在14种仪器上比之前的工作提高了40个百分点。此外,与现有的用于音乐错误检测的转录方法相比,我们的模型可以处理多种乐器。我们的源代码和数据集可以在https://github.com/ben2002chou/Polytune上找到。
摘要:Beginner musicians often struggle to identify specific errors in theirperformances, such as playing incorrect notes or rhythms. There are twolimitations in existing tools for music error detection: (1) Existingapproaches rely on automatic alignment; therefore, they are prone to errorscaused by small deviations between alignment targets.; (2) There is a lack ofsufficient data to train music error detection models, resulting inover-reliance on heuristics. To address (1), we propose a novel transformermodel, Polytune, that takes audio inputs and outputs annotated music scores.This model can be trained end-to-end to implicitly align and compareperformance audio with music scores through latent space representations. Toaddress (2), we present a novel data generation technique capable of creatinglarge-scale synthetic music error datasets. Our approach achieves a 64.1%average Error Detection F1 score, improving upon prior work by 40 percentagepoints across 14 instruments. Additionally, compared with existingtranscription methods repurposed for music error detection, our model canhandle multiple instruments. Our source code and datasets are available athttps://github.com/ben2002chou/Polytune.
