本文经arXiv每日学术速递授权转载
【1】Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation
链接:https://arxiv.org/abs/2501.05413
摘要:训练音频到图像生成模型需要大量在语义上对齐的多样化视听对。这些数据几乎总是从野外视频中挑选出来的,因为它们固有的跨模态语义对应。在这项工作中,我们假设坚持对地面真实视听对应的绝对需要不仅是不必要的,而且会导致数据的规模,质量和多样性受到严重限制,最终损害其在现代生成模型中的使用。也就是说,我们提出了一个可扩展的图像声化框架,在该框架中,来自各种高质量但不相交的单峰起源的实例可以通过现代视觉语言模型的推理能力所赋予的检索过程来人工配对。为了证明这种方法的有效性,我们使用我们的声化图像来训练音频到图像生成模型,该模型的性能与最先进的模型具有竞争力。最后,通过一系列的消融研究,我们展示了几个有趣的听觉功能,如语义混合和插值,响度校准和声学空间建模,通过混响,我们的模型已经隐含地开发来指导图像生成过程。
摘要:Training audio-to-image generative models requires an abundance of diverseaudio-visual pairs that are semantically aligned. Such data is almost alwayscurated from in-the-wild videos, given the cross-modal semantic correspondencethat is inherent to them. In this work, we hypothesize that insisting on theabsolute need for ground truth audio-visual correspondence, is not onlyunnecessary, but also leads to severe restrictions in scale, quality, anddiversity of the data, ultimately impairing its use in the modern generativemodels. That is, we propose a scalable image sonification framework whereinstances from a variety of high-quality yet disjoint uni-modal origins can beartificially paired through a retrieval process that is empowered by reasoningcapabilities of modern vision-language models. To demonstrate the efficacy ofthis approach, we use our sonified images to train an audio-to-image generativemodel that performs competitively against state-of-the-art. Finally, through aseries of ablation studies, we exhibit several intriguing auditory capabilitieslike semantic mixing and interpolation, loudness calibration and acoustic spacemodeling through reverberation that our model has implicitly developed to guidethe image generation process.
标题:AnCoGen:使用掩蔽自动编码器分析、控制和生成语音
链接:https://arxiv.org/abs/2501.05332
备注:5 pages, this https URL
摘要:本文介绍了AnCoGen,这是一种新的方法,它利用掩码自动编码器将语音信号的分析、控制和生成统一在一个模型中。AnCoGen可以通过估计关键属性来分析语音,例如说话者身份,音高,内容,响度,信噪比和清晰度指数。此外,它可以从这些属性生成语音,并允许通过修改它们来精确控制合成语音。大量的实验证明了AnCoGen在语音分析-再合成、基音估计、基音修改和语音增强方面的有效性。
摘要:This article introduces AnCoGen, a novel method that leverages a maskedautoencoder to unify the analysis, control, and generation of speech signalswithin a single model. AnCoGen can analyze speech by estimating key attributes,such as speaker identity, pitch, content, loudness, signal-to-noise ratio, andclarity index. In addition, it can generate speech from these attributes andallow precise control of the synthesized speech by modifying them. Extensiveexperiments demonstrated the effectiveness of AnCoGen across speechanalysis-resynthesis, pitch estimation, pitch modification, and speechenhancement.
标题:ZipEnhancer:基于双向上下采样的Zipformer,用于单耳语音增强
链接:https://arxiv.org/abs/2501.05183
备注:Accepted by ICASSP 2025
摘要:与其他用三轴建模隐藏层特征的序列任务相比,双路径时域和时频域语音增强模型是有效的,并且具有低参数,但由于其隐藏层特征具有四轴,因此计算要求很高。我们提出了ZipEnhancer,这是双路径上下采样为基础的单声道语音增强Zipformer,结合时域和频域上下采样,以减少计算成本。我们引入了ZipformerBlock作为核心模块,并提出了对称缩放和缩放的双路径DownSampleStacks的设计。此外,我们还引入了ScaleAdam优化器和Eden学习率调度器来进一步提高性能。我们的模型在DNS 2020 Challenge和Voicebank+DEMAND数据集上获得了最新的结果,语音质量感知评估(PESQ)为3.69和3.63,使用2.04M参数和62.41G FLOPS,优于具有类似复杂度的其他方法。
摘要:In contrast to other sequence tasks modeling hidden layer features with threeaxes, Dual-Path time and time-frequency domain speech enhancement models areeffective and have low parameters but are computationally demanding due totheir hidden layer features with four axes. We propose ZipEnhancer, which isDual-Path Down-Up Sampling-based Zipformer for Monaural Speech Enhancement,incorporating time and frequency domain Down-Up sampling to reducecomputational costs. We introduce the ZipformerBlock as the core block andpropose the design of the Dual-Path DownSampleStacks that symmetrically scaledown and scale up. Also, we introduce the ScaleAdam optimizer and Eden learningrate scheduler to improve the performance further. Our model achieves newstate-of-the-art results on the DNS 2020 Challenge and Voicebank+DEMANDdatasets, with a perceptual evaluation of speech quality (PESQ) of 3.69 and3.63, using 2.04M parameters and 62.41G FLOPS, outperforming other methods withsimilar complexity levels.
标题:DistAttack:说话人识别中基于扩散的Timbre-Reserved对抗攻击
链接:https://arxiv.org/abs/2501.05127
备注:5 pages,4 figures, accepted by ICASSP 2025
摘要:说话人识别系统作为生物特征识别的一种形式,其安全性至关重要。为了更好地理解SID系统的鲁棒性,我们的目标是在SID中执行更真实的攻击,这对人类和机器都具有挑战性。在这项研究中,我们提出了DiffAttack,这是一种新的保留音色的对抗性攻击方法,它利用基于扩散的语音转换(DiffVC)模型的能力来生成具有不同目标扬声器属性的对抗性假音频。通过在基于扩散的语音转换模型的生成过程中引入对抗性约束,我们制作了假样本,有效地误导了目标模型,同时保留了说话人的特征。具体来说,受传统对抗攻击和扩散过程中随机采样高斯噪声的启发,我们将对抗约束纳入反向扩散过程。这些约束巧妙地引导反向扩散过程与目标扬声器分布对准。我们在LibriTTS数据集上的实验表明,与普通DiffVC和其他方法相比,DiffAttack显著提高了攻击成功率。此外,客观和主观的评价表明,引入对抗性约束不会损害DiffVC模型产生的语音质量。
摘要:Being a form of biometric identification, the security of the speakeridentification (SID) system is of utmost importance. To better understand therobustness of SID systems, we aim to perform more realistic attacks in SID,which are challenging for both humans and machines to detect. In this study, wepropose DiffAttack, a novel timbre-reserved adversarial attack approach thatexploits the capability of a diffusion-based voice conversion (DiffVC) model togenerate adversarial fake audio with distinct target speaker attribution. Byintroducing adversarial constraints into the generative process of thediffusion-based voice conversion model, we craft fake samples that effectivelymislead target models while preserving speaker-wise characteristics.Specifically, inspired by the use of randomly sampled Gaussian noise inconventional adversarial attacks and diffusion processes, we incorporateadversarial constraints into the reverse diffusion process. These constraintssubtly guide the reverse diffusion process toward aligning with the targetspeaker distribution. Our experiments on the LibriTTS dataset indicate thatDiffAttack significantly improves the attack success rate compared to vanillaDiffVC and other methods. Moreover, objective and subjective evaluationsdemonstrate that introducing adversarial constraints does not compromise thespeech quality generated by the DiffVC model.
标题:D3 RM:钢琴抄写的离散去噪扩散细化模型
链接:https://arxiv.org/abs/2501.05068
备注:Accepted to ICASSP 2025
摘要:扩散模型由于其在复杂数据分布建模方面的良好表现,在生成领域得到了广泛的应用。此外,它们在区分性任务(如图像分割)上显示出具有竞争力的结果。虽然扩散模型也被用于自动音乐转录,但它们的性能尚未达到竞争水平。在本文中,我们专注于离散扩散模型的细化能力,并提出了一种新的架构钢琴转录。我们的模型利用邻域注意层作为去噪模块,以预训练声学模型的微调特征为条件,逐步预测目标高分辨率钢琴滚音。为了进一步提高细化,我们设计了一种新的策略,在离散扩散模型的训练和推理阶段应用不同的过渡状态。在MAESTRO数据集上的实验表明,我们的方法在F1分数方面优于以前的基于扩散的钢琴转录模型和基线模型。我们的代码可在https://github.com/hanshounsu/d3rm上获得。
摘要:Diffusion models have been widely used in the generative domain due to theirconvincing performance in modeling complex data distributions. Moreover, theyhave shown competitive results on discriminative tasks, such as imagesegmentation. While diffusion models have also been explored for automaticmusic transcription, their performance has yet to reach a competitive level. Inthis paper, we focus on discrete diffusion model's refinement capabilities andpresent a novel architecture for piano transcription. Our model utilizesNeighborhood Attention layers as the denoising module, gradually predicting thetarget high-resolution piano roll, conditioned on the finetuned features of apretrained acoustic model. To further enhance refinement, we devise a novelstrategy which applies distinct transition states during training and inferencestage of discrete diffusion models. Experiments on the MAESTRO dataset showthat our approach outperforms previous diffusion-based piano transcriptionmodels and the baseline model in terms of F1 score. Our code is available inhttps://github.com/hanshounsu/d3rm.
标题:使用分类器组链标记音乐
链接:https://arxiv.org/abs/2501.05050
备注:5 pages, 2 figures
摘要:我们提出了音乐标签与分类器链,模型的相互作用的音乐标签。大多数传统方法通过将它们视为多个独立的二进制分类问题来独立地估计多个标签。这种处理忽略了音乐标签之间的条件依赖性,导致次优的标记性能。与大多数音乐标签,所提出的方法顺序估计每个标签的基础上的分类器链的想法。除了朴素的分类器链之外,所提出的方法按类别(如流派)对多个标签进行分组,并按组的单位执行链,我们称之为\textit{classifier group chains}。我们的方法允许标记组之间的依赖性建模。我们通过使用MTG-Jamendo数据集的音乐标记实验来评估所提出的方法对音乐标记性能的有效性。此外,我们研究了有效的顺序链的音乐标记。
摘要:We propose music tagging with classifier chains that model the interplay ofmusic tags. Most conventional methods estimate multiple tags independently bytreating them as multiple independent binary classification problems. Thistreatment overlooks the conditional dependencies among music tags, leading tosuboptimal tagging performance. Unlike most music taggers, the proposed methodsequentially estimates each tag based on the idea of the classifier chains.Beyond the naive classifier chains, the proposed method groups the multipletags by category, such as genre, and performs chains by unit of groups, whichwe call \textit{classifier group chains}. Our method allows the modeling of thedependence between tag groups. We evaluate the effectiveness of the proposedmethod for music tagging performance through music tagging experiments usingthe MTG-Jamendo dataset. Furthermore, we investigate the effective order ofchains for music tagging.
标题:VoxEval:对端到端口语模型的知识理解能力进行基准测试
链接:https://arxiv.org/abs/2501.04962
摘要:随着对开发基于语音的交互模型的需求不断增长,端到端口语模型(SLM)已成为一种有前途的解决方案。当与人类进行对话时,这些模型必须理解广泛的世界知识。在本文中,我们介绍了VoxEval,这是一种新型的语音问答基准,专门设计用于通过纯粹基于语音的交互来评估SLM的知识理解。与现有的AudioQA基准测试不同,VoxEval保持了问题和答案的语音格式,在不同的音频条件下(不同的音色,音频质量和说话风格)评估模型的鲁棒性,并率先评估具有挑战性的领域,如以口语格式解决数学问题。我们最近的SLM使用VoxEval的全面评估揭示了当前模型的显着性能局限性,突出了未来改进的关键领域。
摘要:With the growing demand for developing speech-based interaction models,end-to-end Spoken Language Models (SLMs) have emerged as a promising solution.When engaging in conversations with humans, it is essential for these models tocomprehend a wide range of world knowledge. In this paper, we introduceVoxEval, a novel speech question-answering benchmark specifically designed toassess SLMs' knowledge understanding through purely speech-based interactions.Unlike existing AudioQA benchmarks, VoxEval maintains speech format for bothquestions and answers, evaluates model robustness across diverse audioconditions (varying timbres, audio qualities, and speaking styles), andpioneers the assessment of challenging domains like mathematicalproblem-solving in spoken format. Our comprehensive evaluation of recent SLMsusing VoxEval reveals significant performance limitations in current models,highlighting crucial areas for future improvements.
标题:用于有限标签的音频深度伪造检测的视觉图非对比学习
链接:https://arxiv.org/abs/2501.04942
摘要:音频深度伪造检测的最新进展利用图神经网络(GNN)对音频数据中的频率和时间相关性进行建模,有效地识别深度伪造伪影。然而,基于GNN的方法依赖于大量标记数据来进行图形构建和鲁棒性能,这限制了它们在标记数据有限的情况下的适用性。虽然存在大量的音频数据,但将样本标记为真品或赝品的过程仍然是劳动密集型的,成本高昂。为了应对这一挑战,我们提出了SIGNL(时空视觉图非对比学习),这是一种在低标签环境中保持高GNN性能的新框架。SIGNL通过将来自音频的视觉频谱图的补丁表示为节点来构建时空图。这些图结构使用视觉图卷积(GC)编码器进行建模,这些编码器通过图非对比学习进行预训练,这是一种无标签的方法,可以最大限度地提高正对之间的相似性。然后,预训练的编码器会进行微调以进行音频deepfake检测,从而减少对标记数据的依赖。实验表明,SIGNL在多个音频deepfake检测数据集上的表现优于最先进的基线,仅用5%的标记数据就实现了最低的等错误率(EER)。此外,SIGNL表现出强大的跨域泛化能力,在In-The-Wild数据集中涉及不同攻击类型和语言的评估中实现了最低的EER。
摘要:Recent advancements in audio deepfake detection have leveraged graph neuralnetworks (GNNs) to model frequency and temporal interdependencies in audiodata, effectively identifying deepfake artifacts. However, the reliance ofGNN-based methods on substantial labeled data for graph construction and robustperformance limits their applicability in scenarios with limited labeled data.Although vast amounts of audio data exist, the process of labeling samples asgenuine or fake remains labor-intensive and costly. To address this challenge,we propose SIGNL (Spatio-temporal vIsion Graph Non-contrastive Learning), anovel framework that maintains high GNN performance in low-label settings.SIGNL constructs spatio-temporal graphs by representing patches from theaudio's visual spectrogram as nodes. These graph structures are modeled usingvision graph convolutional (GC) encoders pre-trained through graphnon-contrastive learning, a label-free that maximizes the similarity betweenpositive pairs. The pre-trained encoders are then fine-tuned for audio deepfakedetection, reducing reliance on labeled data. Experiments demonstrate thatSIGNL outperforms state-of-the-art baselines across multiple audio deepfakedetection datasets, achieving the lowest Equal Error Rate (EER) with as littleas 5% labeled data. Additionally, SIGNL exhibits strong cross-domaingeneralization, achieving the lowest EER in evaluations involving diverseattack types and languages in the In-The-Wild dataset.
标题:JELLY:利用LLM联合情感识别和上下文推理用于对话语音合成
链接:https://arxiv.org/abs/2501.04904
备注:Accepted by ICASSP 2025
摘要:近来,对于通过考虑会话上下文来生成更自然的语音的会话语音合成(CSS)的需求不断增长。为了解决这个问题,我们介绍了JELLY,一种新的CSS框架,它集成了情感识别和上下文推理,通过微调具有多个部分LoRA模块的大型语言模型(LLM),在对话中生成适当的语音。我们提出了一个具有感知能力的Q-前编码器,它使LLM能够感知语音中的情感。编码器经过训练,利用情感语音数据集将语音情感与文本对齐。然后,整个模型与会话语音数据进行微调,以推断情感上下文,从而在会话中生成情感上合适的语音。我们的实验结果表明,JELLY在情感上下文建模方面表现出色,合成的语音自然与对话一致,同时减轻了情感对话语音数据集的稀缺性。
摘要:Recently, there has been a growing demand for conversational speech synthesis(CSS) that generates more natural speech by considering the conversationalcontext. To address this, we introduce JELLY, a novel CSS framework thatintegrates emotion recognition and context reasoning for generating appropriatespeech in conversation by fine-tuning a large language model (LLM) withmultiple partial LoRA modules. We propose an Emotion-aware Q-former encoder,which enables the LLM to perceive emotions in speech. The encoder is trained toalign speech emotions with text, utilizing datasets of emotional speech. Theentire model is then fine-tuned with conversational speech data to inferemotional context for generating emotionally appropriate speech inconversation. Our experimental results demonstrate that JELLY excels inemotional context modeling, synthesizing speech that naturally aligns withconversation, while mitigating the scarcity of emotional conversational speechdatasets.
标题:用耳朵刨削:用于工业木刨声学异常检测的卷积神经网络
链接:https://arxiv.org/abs/2501.04819
摘要:近年来,木制品行业一直面临着熟练劳动力短缺的问题。其结果是更频繁的突然故障,导致这些公司已经在竞争激烈的市场上运营的额外成本。此外,锯木厂对机械和传感器来说是一个具有挑战性的环境。鉴于有经验的机器操作员可能能够诊断缺陷或故障,一种可能的方法是通过声学监测来帮助新手操作员。作为实现木材加工设备自动化和机器操作员决策支持系统的一步,在本文中,我们探索使用深度卷积自动编码器在新的现实数据集上进行木工刨床的声学异常检测。具体来说,我们的具有跳过连接的卷积自动编码器(Skip-CAE)和我们的Skip-CAE Transformer优于DCASE自动编码器基线、单类SVM、隔离森林和已发布的卷积自动编码器架构,分别在真实工厂刨床声音数据集上获得0.846和0.875的ROC曲线下面积。此外,我们表明,添加跳过连接和注意力机制下的一个Transformer编码器-解码器的形式,有助于进一步提高异常检测能力。
摘要:In recent years, the wood product industry has been facing a skilled laborshortage. The result is more frequent sudden failures, resulting in additionalcosts for these companies already operating in a very competitive market.Moreover, sawmills are challenging environments for machinery and sensors.Given that experienced machine operators may be able to diagnose defects ormalfunctions, one possible way of assisting novice operators is throughacoustic monitoring. As a step towards the automation of wood-processingequipment and decision support systems for machine operators, in this paper, weexplore using a deep convolutional autoencoder for acoustic anomaly detectionof wood planers on a new real-life dataset. Specifically, our convolutionalautoencoder with skip connections (Skip-CAE) and our Skip-CAE transformeroutperform the DCASE autoencoder baseline, one-class SVM, isolation forest anda published convolutional autoencoder architecture, respectively obtaining anarea under the ROC curve of 0.846 and 0.875 on a dataset of real-factory planersounds. Moreover, we show that adding skip connections and attention mechanismunder the form of a transformer encoder-decoder helps to further improve theanomaly detection capabilities.
标题:探索说话者表示中说话者特定的特征
链接:https://arxiv.org/abs/2501.05310
摘要:本研究探讨了特定于说话人的特征编码在说话人嵌入和语音自监督学习(SSL)模型的中间层。通过利用探测方法,我们分析了突出说话人嵌入模型和语音SSL模型(包括HuBERT,WavLM和Wav2vec 2.0)的音高,节奏和能量等特征。结果表明,像CAM++这样的说话人嵌入在能量分类方面表现出色,而语音SSL模型由于其分层特征编码而在多个特征上表现出优异的性能。中间层有效地捕获了声学和类语言信息的混合,更深的层细化了这些表示。这项调查提供了对模型设计的见解,并强调了这些表示在下游应用中的潜力,例如说话人验证和文本到语音合成,同时为探索其他功能和高级探测方法奠定了基础。
摘要:This study explores speaker-specific features encoded in speaker embeddingsand intermediate layers of speech self-supervised learning (SSL) models. Byutilising a probing method, we analyse features such as pitch, tempo, andenergy across prominent speaker embedding models and speech SSL models,including HuBERT, WavLM, and Wav2vec 2.0. The results reveal that speakerembeddings like CAM++ excel in energy classification, while speech SSL modelsdemonstrate superior performance across multiple features due to theirhierarchical feature encoding. Intermediate layers effectively capture a mix ofacoustic and para-linguistic information, with deeper layers refining theserepresentations. This investigation provides insights into model design andhighlights the potential of these representations for downstream applications,such as speaker verification and text-to-speech synthesis, while laying thegroundwork for exploring additional features and advanced probing methods.
标题:FlowHigh:通过一步流匹配实现高效、高质量的音频超分辨率
链接:https://arxiv.org/abs/2501.04926
备注:Accepted by ICASSP 2025
摘要:音频超分辨率由于其不适定性而具有挑战性。最近,扩散模型在音频超分辨率中的应用在缓解这一挑战方面显示出有希望的结果。然而,基于扩散的模型具有局限性,主要是需要许多采样步骤,这导致在合成高质量音频样本时显著增加的延迟。在本文中,我们提出了FLowHigh,一种新的方法,集成流匹配,一个高效的生成模型,到音频超分辨率。我们还探索了专门为音频超分辨率定制的概率路径,它可以有效地捕获高分辨率音频分布,从而提高重建质量。所提出的方法通过跨各种输入采样率的单步采样过程来生成高保真、高分辨率的音频。在VCTK基准数据集上的实验结果表明,FLowHigh在音频超分辨率方面实现了最先进的性能,如对数谱距离和ViSQOL所评估的,同时仅通过单步采样过程保持计算效率。
摘要:Audio super-resolution is challenging owing to its ill-posed nature.Recently, the application of diffusion models in audio super-resolution hasshown promising results in alleviating this challenge. However, diffusion-basedmodels have limitations, primarily the necessity for numerous sampling steps,which causes significantly increased latency when synthesizing high-qualityaudio samples. In this paper, we propose FLowHigh, a novel approach thatintegrates flow matching, a highly efficient generative model, into audiosuper-resolution. We also explore probability paths specially tailored foraudio super-resolution, which effectively capture high-resolution audiodistributions, thereby enhancing reconstruction quality. The proposed methodgenerates high-fidelity, high-resolution audio through a single-step samplingprocess across various input sampling rates. The experimental results on theVCTK benchmark dataset demonstrate that FLowHigh achieves state-of-the-artperformance in audio super-resolution, as evaluated by log-spectral distanceand ViSQOL while maintaining computational efficiency with only a single-stepsampling process.
标题:基本频率估计器与次调和语音信号的比较
链接:https://arxiv.org/abs/2501.04789
备注:9 pages, 6 figures
摘要:在临床语音信号分析中,次谐波发声的错误处理可能导致声学参数发出假阴性信号。因此,基本频率估计器识别说话基本频率的能力至关重要。本文提出了一项持续元音研究,该研究使用估计质量分类来识别次谐波误差,并使用次谐波与谐波比(SHR)来测量次谐波发声的强度。使用持续元音数据集研究了五个估计量:Praat,YAAPT,Harvest,CREPE和FCN-F0。深度学习模型FCN-F0在整体准确性和正确解析次谐波信号方面表现最好。CREPE和Harvest也是持续元音分析的高能力估计器。
摘要:In clinical voice signal analysis, mishandling of subharmonic voicing maycause an acoustic parameter to signal false negatives. As such, the ability ofa fundamental frequency estimator to identify speaking fundamental frequency iscritical. This paper presents a sustained-vowel study, which used aquality-of-estimate classification to identify subharmonic errors andsubharmonics-to-harmonics ratio (SHR) to measure the strength of subharmonicvoicing. Five estimators were studied with a sustained vowel dataset: Praat,YAAPT, Harvest, CREPE, and FCN-F0. FCN-F0, a deep-learning model, performed thebest both in overall accuracy and in correctly resolving subharmonic signals.CREPE and Harvest are also highly capable estimators for sustained vowelanalysis.
标题:探索说话者表示中说话者特定的特征
链接:https://arxiv.org/abs/2501.05310
摘要:本研究探讨了特定于说话人的特征编码在说话人嵌入和语音自监督学习(SSL)模型的中间层。通过利用探测方法,我们分析了突出说话人嵌入模型和语音SSL模型(包括HuBERT,WavLM和Wav2vec 2.0)的音高,节奏和能量等特征。结果表明,像CAM++这样的说话人嵌入在能量分类方面表现出色,而语音SSL模型由于其分层特征编码而在多个特征上表现出优异的性能。中间层有效地捕获了声学和类语言信息的混合,更深的层细化了这些表示。这项调查提供了对模型设计的见解,并强调了这些表示在下游应用中的潜力,例如说话人验证和文本到语音合成,同时为探索其他功能和高级探测方法奠定了基础。
摘要:This study explores speaker-specific features encoded in speaker embeddingsand intermediate layers of speech self-supervised learning (SSL) models. Byutilising a probing method, we analyse features such as pitch, tempo, andenergy across prominent speaker embedding models and speech SSL models,including HuBERT, WavLM, and Wav2vec 2.0. The results reveal that speakerembeddings like CAM++ excel in energy classification, while speech SSL modelsdemonstrate superior performance across multiple features due to theirhierarchical feature encoding. Intermediate layers effectively capture a mix ofacoustic and para-linguistic information, with deeper layers refining theserepresentations. This investigation provides insights into model design andhighlights the potential of these representations for downstream applications,such as speaker verification and text-to-speech synthesis, while laying thegroundwork for exploring additional features and advanced probing methods.
标题:FlowHigh:通过一步流匹配实现高效、高质量的音频超分辨率
链接:https://arxiv.org/abs/2501.04926
备注:Accepted by ICASSP 2025
摘要:音频超分辨率由于其不适定性而具有挑战性。最近,扩散模型在音频超分辨率中的应用在缓解这一挑战方面显示出有希望的结果。然而,基于扩散的模型具有局限性,主要是需要许多采样步骤,这导致在合成高质量音频样本时显著增加的延迟。在本文中,我们提出了FLowHigh,一种新的方法,集成流匹配,一个高效的生成模型,到音频超分辨率。我们还探索了专门为音频超分辨率定制的概率路径,它可以有效地捕获高分辨率音频分布,从而提高重建质量。所提出的方法通过跨各种输入采样率的单步采样过程来生成高保真、高分辨率的音频。在VCTK基准数据集上的实验结果表明,FLowHigh在音频超分辨率方面实现了最先进的性能,如通过对数频谱距离和ViSQOL评估的,同时仅通过单步采样过程保持计算效率。
摘要:Audio super-resolution is challenging owing to its ill-posed nature.Recently, the application of diffusion models in audio super-resolution hasshown promising results in alleviating this challenge. However, diffusion-basedmodels have limitations, primarily the necessity for numerous sampling steps,which causes significantly increased latency when synthesizing high-qualityaudio samples. In this paper, we propose FLowHigh, a novel approach thatintegrates flow matching, a highly efficient generative model, into audiosuper-resolution. We also explore probability paths specially tailored foraudio super-resolution, which effectively capture high-resolution audiodistributions, thereby enhancing reconstruction quality. The proposed methodgenerates high-fidelity, high-resolution audio through a single-step samplingprocess across various input sampling rates. The experimental results on theVCTK benchmark dataset demonstrate that FLowHigh achieves state-of-the-artperformance in audio super-resolution, as evaluated by log-spectral distanceand ViSQOL while maintaining computational efficiency with only a single-stepsampling process.
标题:通过并行音素序列预测增强脑电收听语音解码
链接:https://arxiv.org/abs/2501.04844
备注:ICASSP 2025
摘要:脑机接口(BCI)提供了许多以人为中心的应用可能性,特别是对神经系统疾病患者的影响。从大脑活动中解码文本或语音是一个相关的领域,可以提高语音感知受损的人的生活质量。我们提出了一种新的方法来提高听语音解码从脑电图(EEG)信号,利用辅助音素预测,同时解码文本音素序列。该模型由脑电模型、语音模型和音素预测器三部分组成。EEG模块学习将EEG信号正确表示为EEG嵌入。语音模块从EEG嵌入生成语音波形。音素预测器以文本模态输出解码的音素序列。我们提出的方法允许用户同时从两种模式(语音波形和文本音素序列)的EEG信号中获得解码的收听语音,从而消除了对每种模式的级联顺序管道的需要。所提出的方法也优于以前的方法在这两种模式。源代码和语音样本是公开的。
摘要:Brain-computer interfaces (BCI) offer numerous human-centered applicationpossibilities, particularly affecting people with neurological disorders. Textor speech decoding from brain activities is a relevant domain that couldaugment the quality of life for people with impaired speech perception. Wepropose a novel approach to enhance listened speech decoding fromelectroencephalography (EEG) signals by utilizing an auxiliary phonemepredictor that simultaneously decodes textual phoneme sequences. The proposedmodel architecture consists of three main parts: EEG module, speech module, andphoneme predictor. The EEG module learns to properly represent EEG signals intoEEG embeddings. The speech module generates speech waveforms from the EEGembeddings. The phoneme predictor outputs the decoded phoneme sequences in textmodality. Our proposed approach allows users to obtain decoded listened speechfrom EEG signals in both modalities (speech waveforms and textual phonemesequences) simultaneously, eliminating the need for a concatenated sequentialpipeline for each modality. The proposed approach also outperforms previousmethods in both modalities. The source code and speech samples are publiclyavailable.
标题:基本频率估计器与次调和语音信号的比较
链接:https://arxiv.org/abs/2501.04789
备注:9 pages, 6 figures
摘要:在临床语音信号分析中,次谐波发声的错误处理可能导致声学参数发出假阴性信号。因此,基本频率估计器识别说话基本频率的能力至关重要。本文提出了一项持续元音研究,该研究使用估计质量分类来识别次谐波误差,并使用次谐波与谐波比(SHR)来测量次谐波发声的强度。使用持续元音数据集研究了五个估计量:Praat,YAAPT,Harvest,CREPE和FCN-F0。深度学习模型FCN-F0在整体准确性和正确解析次谐波信号方面表现最好。CREPE和Harvest也是持续元音分析的高能力估计器。
摘要:In clinical voice signal analysis, mishandling of subharmonic voicing maycause an acoustic parameter to signal false negatives. As such, the ability ofa fundamental frequency estimator to identify speaking fundamental frequency iscritical. This paper presents a sustained-vowel study, which used aquality-of-estimate classification to identify subharmonic errors andsubharmonics-to-harmonics ratio (SHR) to measure the strength of subharmonicvoicing. Five estimators were studied with a sustained vowel dataset: Praat,YAAPT, Harvest, CREPE, and FCN-F0. FCN-F0, a deep-learning model, performed thebest both in overall accuracy and in correctly resolving subharmonic signals.CREPE and Harvest are also highly capable estimators for sustained vowelanalysis.
标题:基于元学习的打击乐转录和$tar{a}la$来自低资源音频的识别
链接:https://arxiv.org/abs/2501.04742
备注:arXiv admin note: substantial text overlap with arXiv:2407.20935
摘要:本研究介绍了一种基于元学习的方法,用于低资源Tabla笔划转录(TST)和印度斯坦古典音乐中的$t\bar{a}la$识别。使用模型不可知元学习(MAML),我们解决了有限的注释数据集的挑战,使快速适应新的任务,最少的数据。该方法在各种数据集上进行了验证,包括Tabla独奏和音乐会录音,证明了复调音频场景的鲁棒性。我们提出了两个新的$t\bar{a}la$识别技术的基础上的笔画序列和节奏模式。此外,该方法被证明是有效的自动鼓转录(ADT),展示了其灵活性,印度和西方打击乐。实验结果表明,该方法在低资源环境下的性能优于现有技术,为音乐转录和通过计算工具研究音乐传统做出了显着贡献。
摘要:This study introduces a meta-learning-based approach for low-resource TablaStroke Transcription (TST) and $t\bar{a}la$ identification in Hindustaniclassical music. Using Model-Agnostic Meta-Learning (MAML), we address thechallenge of limited annotated datasets, enabling rapid adaptation to new taskswith minimal data. The method is validated across various datasets, includingtabla solo and concert recordings, demonstrating robustness in polyphonic audioscenarios. We propose two novel $t\bar{a}la$ identification techniques based onstroke sequences and rhythmic patterns. Additionally, the approach proveseffective for Automatic Drum Transcription (ADT), showcasing its flexibilityfor Indian and Western percussion music. Experimental results show that theproposed method outperforms existing techniques in low-resource settings,significantly contributing to music transcription and studying musicaltraditions through computational tools.
标题:看到声音:从视觉图像中收集声音以生成音频到图像
链接:https://arxiv.org/abs/2501.05413
摘要:训练音频到图像生成模型需要大量在语义上对齐的多样化视听对。这些数据几乎总是从野外视频中挑选出来的,因为它们固有的跨模态语义对应。在这项工作中,我们假设坚持对地面真实视听对应的绝对需要不仅是不必要的,而且会导致数据的规模,质量和多样性受到严重限制,最终损害其在现代生成模型中的使用。也就是说,我们提出了一个可扩展的图像声化框架,在该框架中,来自各种高质量但不相交的单峰起源的实例可以通过现代视觉语言模型的推理能力所赋予的检索过程来人工配对。为了证明这种方法的有效性,我们使用我们的声化图像来训练音频到图像生成模型,该模型的性能与最先进的模型具有竞争力。最后,通过一系列的消融研究,我们展示了几个有趣的听觉功能,如语义混合和插值,响度校准和声学空间建模,通过混响,我们的模型已经隐含地开发来指导图像生成过程。
摘要:Training audio-to-image generative models requires an abundance of diverseaudio-visual pairs that are semantically aligned. Such data is almost alwayscurated from in-the-wild videos, given the cross-modal semantic correspondencethat is inherent to them. In this work, we hypothesize that insisting on theabsolute need for ground truth audio-visual correspondence, is not onlyunnecessary, but also leads to severe restrictions in scale, quality, anddiversity of the data, ultimately impairing its use in the modern generativemodels. That is, we propose a scalable image sonification framework whereinstances from a variety of high-quality yet disjoint uni-modal origins can beartificially paired through a retrieval process that is empowered by reasoningcapabilities of modern vision-language models. To demonstrate the efficacy ofthis approach, we use our sonified images to train an audio-to-image generativemodel that performs competitively against state-of-the-art. Finally, through aseries of ablation studies, we exhibit several intriguing auditory capabilitieslike semantic mixing and interpolation, loudness calibration and acoustic spacemodeling through reverberation that our model has implicitly developed to guidethe image generation process.
标题:利用半监督学习和LLM优化爱沙尼亚电视字幕
链接:https://arxiv.org/abs/2501.05234
摘要:本文提出了一种方法,用于生成高质量的,相同的语言字幕爱沙尼亚电视内容。我们对人工生成的爱沙尼亚语字幕的Whisper模型进行了微调,并通过迭代伪标签和基于大语言模型(LLM)的后期编辑对其进行了增强。我们的实验表明,通过使用未标记的数据集进行伪标记,字幕质量得到了显着的改善。我们发现,在测试时应用基于LLM的编辑可以提高字幕的准确性,而在训练过程中使用它不会产生进一步的收益。这种方法有望创建接近人类标准的字幕质量,并可以扩展到实时应用程序。
摘要:This paper presents an approach for generating high-quality, same-languagesubtitles for Estonian TV content. We fine-tune the Whisper model onhuman-generated Estonian subtitles and enhance it with iterativepseudo-labeling and large language model (LLM) based post-editing. Ourexperiments demonstrate notable subtitle quality improvement throughpseudo-labeling with an unlabeled dataset. We find that applying LLM-basedediting at test time enhances subtitle accuracy, while its use during trainingdoes not yield further gains. This approach holds promise for creating subtitlequality close to human standard and could be extended to real-timeapplications.
标题:DistAttack:说话人识别中基于扩散的Timbre-Reserved对抗攻击
链接:https://arxiv.org/abs/2501.05127
备注:5 pages,4 figures, accepted by ICASSP 2025
摘要:说话人识别系统作为生物特征识别的一种形式,其安全性至关重要。为了更好地理解SID系统的鲁棒性,我们的目标是在SID中执行更真实的攻击,这对人类和机器都具有挑战性。在这项研究中,我们提出了DiffAttack,这是一种新的保留音色的对抗性攻击方法,它利用基于扩散的语音转换(DiffVC)模型的能力来生成具有不同目标扬声器属性的对抗性假音频。通过在基于扩散的语音转换模型的生成过程中引入对抗性约束,我们制作了假样本,有效地误导了目标模型,同时保留了说话人的特征。具体来说,受传统对抗攻击和扩散过程中随机采样高斯噪声的启发,我们将对抗约束纳入反向扩散过程。这些约束巧妙地引导反向扩散过程与目标扬声器分布对准。我们在LibriTTS数据集上的实验表明,与普通DiffVC和其他方法相比,DiffAttack显著提高了攻击成功率。此外,客观和主观的评价表明,引入对抗性约束不会损害DiffVC模型产生的语音质量。
摘要:Being a form of biometric identification, the security of the speakeridentification (SID) system is of utmost importance. To better understand therobustness of SID systems, we aim to perform more realistic attacks in SID,which are challenging for both humans and machines to detect. In this study, wepropose DiffAttack, a novel timbre-reserved adversarial attack approach thatexploits the capability of a diffusion-based voice conversion (DiffVC) model togenerate adversarial fake audio with distinct target speaker attribution. Byintroducing adversarial constraints into the generative process of thediffusion-based voice conversion model, we craft fake samples that effectivelymislead target models while preserving speaker-wise characteristics.Specifically, inspired by the use of randomly sampled Gaussian noise inconventional adversarial attacks and diffusion processes, we incorporateadversarial constraints into the reverse diffusion process. These constraintssubtly guide the reverse diffusion process toward aligning with the targetspeaker distribution. Our experiments on the LibriTTS dataset indicate thatDiffAttack significantly improves the attack success rate compared to vanillaDiffVC and other methods. Moreover, objective and subjective evaluationsdemonstrate that introducing adversarial constraints does not compromise thespeech quality generated by the DiffVC model.
标题:D3 RM:钢琴抄写的离散去噪扩散细化模型
链接:https://arxiv.org/abs/2501.05068
备注:Accepted to ICASSP 2025
摘要:扩散模型由于其在复杂数据分布建模方面的良好表现,在生成领域得到了广泛的应用。此外,它们在区分性任务(如图像分割)上显示出具有竞争力的结果。虽然扩散模型也被用于自动音乐转录,但它们的性能尚未达到竞争水平。在本文中,我们专注于离散扩散模型的细化能力,并提出了一种新的架构钢琴转录。我们的模型利用邻域注意层作为去噪模块,以预训练声学模型的微调特征为条件,逐步预测目标高分辨率钢琴滚音。为了进一步提高细化,我们设计了一种新的策略,在离散扩散模型的训练和推理阶段应用不同的过渡状态。在MAESTRO数据集上的实验表明,我们的方法在F1分数方面优于以前的基于扩散的钢琴转录模型和基线模型。我们的代码可在https://github.com/hanshounsu/d3rm上获得。
摘要:Diffusion models have been widely used in the generative domain due to theirconvincing performance in modeling complex data distributions. Moreover, theyhave shown competitive results on discriminative tasks, such as imagesegmentation. While diffusion models have also been explored for automaticmusic transcription, their performance has yet to reach a competitive level. Inthis paper, we focus on discrete diffusion model's refinement capabilities andpresent a novel architecture for piano transcription. Our model utilizesNeighborhood Attention layers as the denoising module, gradually predicting thetarget high-resolution piano roll, conditioned on the finetuned features of apretrained acoustic model. To further enhance refinement, we devise a novelstrategy which applies distinct transition states during training and inferencestage of discrete diffusion models. Experiments on the MAESTRO dataset showthat our approach outperforms previous diffusion-based piano transcriptionmodels and the baseline model in terms of F1 score. Our code is available inhttps://github.com/hanshounsu/d3rm.
标题:VoxEval:对端到端口语模型的知识理解能力进行基准测试
链接:https://arxiv.org/abs/2501.04962
摘要:随着对开发基于语音的交互模型的需求不断增长,端到端口语模型(SLM)已成为一种有前途的解决方案。当与人类进行对话时,这些模型必须理解广泛的世界知识。在本文中,我们介绍VoxEval,一种新的语音问答基准,专门设计用于评估SLM的知识理解,通过纯粹的语音为基础的互动。与现有的AudioQA基准测试不同,VoxEval保持了问题和答案的语音格式,在不同的音频条件下(不同的音色,音频质量和说话风格)评估模型的鲁棒性,并率先评估具有挑战性的领域,如以口语格式解决数学问题。我们最近的SLM使用VoxEval的全面评估揭示了当前模型的显着性能局限性,突出了未来改进的关键领域。
摘要:With the growing demand for developing speech-based interaction models,end-to-end Spoken Language Models (SLMs) have emerged as a promising solution.When engaging in conversations with humans, it is essential for these models tocomprehend a wide range of world knowledge. In this paper, we introduceVoxEval, a novel speech question-answering benchmark specifically designed toassess SLMs' knowledge understanding through purely speech-based interactions.Unlike existing AudioQA benchmarks, VoxEval maintains speech format for bothquestions and answers, evaluates model robustness across diverse audioconditions (varying timbres, audio qualities, and speaking styles), andpioneers the assessment of challenging domains like mathematicalproblem-solving in spoken format. Our comprehensive evaluation of recent SLMsusing VoxEval reveals significant performance limitations in current models,highlighting crucial areas for future improvements.
标题:用于有限标签的音频深度伪造检测的视觉图非对比学习
链接:https://arxiv.org/abs/2501.04942
摘要:音频深度伪造检测的最新进展利用图神经网络(GNN)对音频数据中的频率和时间相关性进行建模,有效地识别深度伪造伪影。然而,基于GNN的方法依赖于大量标记数据来进行图形构建和鲁棒性能,这限制了它们在标记数据有限的情况下的适用性。虽然存在大量的音频数据,但将样本标记为真品或赝品的过程仍然是劳动密集型的,成本高昂。为了应对这一挑战,我们提出了SIGNL(时空视觉图非对比学习),这是一种在低标签环境中保持高GNN性能的新框架。SIGNL通过将来自音频的视觉频谱图的补丁表示为节点来构建时空图。这些图结构使用视觉图卷积(GC)编码器进行建模,这些编码器通过图非对比学习进行预训练,这是一种无标签的方法,可以最大限度地提高正对之间的相似性。然后,预训练的编码器会进行微调以进行音频deepfake检测,从而减少对标记数据的依赖。实验表明,SIGNL在多个音频deepfake检测数据集上的表现优于最先进的基线,仅用5%的标记数据就实现了最低的等错误率(EER)。此外,SIGNL表现出强大的跨域泛化能力,在In-The-Wild数据集中涉及不同攻击类型和语言的评估中实现了最低的EER。
摘要:Recent advancements in audio deepfake detection have leveraged graph neuralnetworks (GNNs) to model frequency and temporal interdependencies in audiodata, effectively identifying deepfake artifacts. However, the reliance ofGNN-based methods on substantial labeled data for graph construction and robustperformance limits their applicability in scenarios with limited labeled data.Although vast amounts of audio data exist, the process of labeling samples asgenuine or fake remains labor-intensive and costly. To address this challenge,we propose SIGNL (Spatio-temporal vIsion Graph Non-contrastive Learning), anovel framework that maintains high GNN performance in low-label settings.SIGNL constructs spatio-temporal graphs by representing patches from theaudio's visual spectrogram as nodes. These graph structures are modeled usingvision graph convolutional (GC) encoders pre-trained through graphnon-contrastive learning, a label-free that maximizes the similarity betweenpositive pairs. The pre-trained encoders are then fine-tuned for audio deepfakedetection, reducing reliance on labeled data. Experiments demonstrate thatSIGNL outperforms state-of-the-art baselines across multiple audio deepfakedetection datasets, achieving the lowest Equal Error Rate (EER) with as littleas 5% labeled data. Additionally, SIGNL exhibits strong cross-domaingeneralization, achieving the lowest EER in evaluations involving diverseattack types and languages in the In-The-Wild dataset.
标题:具有1位测量的广义线性模型:最大似然估计的渐进性
链接:https://arxiv.org/abs/2501.04937
备注:ICASSP 2025
摘要:这项工作建立了正则性条件的一致性和渐近正态性的多参数最大似然估计(MLE)从删失数据,删失机制是在1 $位测量的形式。假设未经审查的数据的底层分布属于指数族,自然参数表示为预测变量的线性组合,称为广义线性模型(GLM)。作为分析的一部分,Fisher信息矩阵也来自删失和未删失数据,这有助于量化删失的影响和评估MLE的性能。选择GLM可以考虑各种感兴趣的1位估计的实际示例。特别是,它示出了如何导出的结果可以用来分析两个实际相关的情况:高斯模型与未知的均值和方差,和泊松模型与未知的均值。
摘要:This work establishes regularity conditions for consistency and asymptoticnormality of the multiple parameter maximum likelihood estimator(MLE) fromcensored data, where the censoring mechanism is in the form of $1$-bitmeasurements. The underlying distribution of the uncensored data is assumed tobelong to the exponential family, with natural parameters expressed as a linearcombination of the predictors, known as generalized linear model (GLM). As partof the analysis, the Fisher information matrix is also derived for bothcensored and uncensored data, which helps to quantify the impact of censoringand assess the performance of the MLE. The choice of GLM allows one to considera variety of practical examples where 1-bit estimation is of interest. Inparticular, it is shown how the derived results can be used to analyze twopractically relevant scenarios: the Gaussian model with both unknown mean andvariance, and the Poisson model with an unknown mean.
标题:JELLY:利用LLM联合情感识别和上下文推理用于对话语音合成
链接:https://arxiv.org/abs/2501.04904
备注:Accepted by ICASSP 2025
摘要:最近,人们对对话语音合成(CSS)的需求越来越大,这种合成可以通过考虑对话上下文来生成更自然的语音。为了解决这个问题,我们介绍了JELLY,一种新的CSS框架,它集成了情感识别和上下文推理,通过微调具有多个部分LoRA模块的大型语言模型(LLM),在对话中生成适当的语音。我们提出了一个具有感知能力的Q-前编码器,它使LLM能够感知语音中的情感。编码器经过训练,利用情感语音数据集将语音情感与文本对齐。然后,整个模型与会话语音数据进行微调,以推断情感上下文,从而在会话中生成情感上合适的语音。我们的实验结果表明,JELLY在情感上下文建模方面表现出色,合成的语音自然与对话一致,同时减轻了情感对话语音数据集的稀缺性。
摘要:Recently, there has been a growing demand for conversational speech synthesis(CSS) that generates more natural speech by considering the conversationalcontext. To address this, we introduce JELLY, a novel CSS framework thatintegrates emotion recognition and context reasoning for generating appropriatespeech in conversation by fine-tuning a large language model (LLM) withmultiple partial LoRA modules. We propose an Emotion-aware Q-former encoder,which enables the LLM to perceive emotions in speech. The encoder is trained toalign speech emotions with text, utilizing datasets of emotional speech. Theentire model is then fine-tuned with conversational speech data to inferemotional context for generating emotionally appropriate speech inconversation. Our experimental results demonstrate that JELLY excels inemotional context modeling, synthesizing speech that naturally aligns withconversation, while mitigating the scarcity of emotional conversational speechdatasets.
