本文经arXiv每日学术速递授权转载
微信公众号:arXiv_Daily
链接:https://arxiv.org/abs/2502.06710
备注:Accepted at EMNLP 2024
摘要:音乐表演是视听建模的典型场景。与具有稀疏音频的常见场景不同,音乐表演始终连续地涉及密集的音频信号。虽然现有的多模态学习方法的音频-视频QA表现出令人印象深刻的能力,在一般情况下,他们是无法处理的基本问题,在音乐表演:他们underexplore的多模态信号之间的相互作用的性能,并没有考虑到乐器和音乐的独特特征。因此,现有的方法倾向于不准确地回答关于音乐表演的问题。为了弥补上述研究空白,(i)考虑到音乐数据固有的复杂多模态互联性,我们的主要骨干旨在将多模态交互纳入音乐背景中;(ii)为了使模型能够学习音乐特征,我们在当前音乐数据集中注释和发布节奏和音乐源;(iii)对于时间感知的视听建模,我们将模型的音乐预测与时间维度对齐。我们的实验显示了对Music AVQA数据集的最新影响。我们的代码可在https://github.com/xid32/Amuse上获得。
摘要:Music performances are representative scenarios for audio-visual modeling.Unlike common scenarios with sparse audio, music performances continuouslyinvolve dense audio signals throughout. While existing multimodal learningmethods on the audio-video QA demonstrate impressive capabilities in generalscenarios, they are incapable of dealing with fundamental problems within themusic performances: they underexplore the interaction between the multimodalsignals in performance and fail to consider the distinctive characteristics ofinstruments and music. Therefore, existing methods tend to answer questionsregarding musical performances inaccurately. To bridge the above research gaps,(i) given the intricate multimodal interconnectivity inherent to music data,our primary backbone is designed to incorporate multimodal interactions withinthe context of music; (ii) to enable the model to learn music characteristics,we annotate and release rhythmic and music sources in the current musicdatasets; (iii) for time-aware audio-visual modeling, we align the model'smusic predictions with the temporal dimension. Our experiments showstate-of-the-art effects on the Music AVQA datasets. Our code is available athttps://github.com/xid32/Amuse.
标题:可听到的深度音频表示的评估
链接:https://arxiv.org/abs/2502.06664
备注:Accepted at International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2025)
摘要:有效地操纵可听设备需要理解用户周围的声学环境。在声音场景的计算分析中,基础模型已经成为产生高性能、鲁棒、多用途音频表示的最新技术。我们推出并发布了音频表示的深度评估(DEAR),这是第一个评估基础模型在捕获可听设备基本声学特性方面的有效性的数据集和基准。该数据集包括1,158个音轨,每个30秒长,通过空间混合专有独白与商业,高质量的日常声学场景录音创建。我们的基准测试包括八个任务,评估音频场景的一般背景、语音源和技术声学特性。通过对四种通用音频表示模型的评估,我们证明了BEATs模型明显优于其同类模型。这种优越性强调了在不同音频集合上训练的模型的优势,证实了它们适用于各种听觉任务,包括编码听觉转向所需的环境属性。DEAR数据集和相关代码可在https://dear-dataset.github.io上获得。
摘要:Effectively steering hearable devices requires understanding the acousticenvironment around the user. In the computational analysis of sound scenes,foundation models have emerged as the state of the art to producehigh-performance, robust, multi-purpose audio representations. We introduce andrelease Deep Evaluation of Audio Representations (DEAR), the first dataset andbenchmark to evaluate the efficacy of foundation models in capturing essentialacoustic properties for hearables. The dataset includes 1,158 audio tracks,each 30 seconds long, created by spatially mixing proprietary monologues withcommercial, high-quality recordings of everyday acoustic scenes. Our benchmarkencompasses eight tasks that assess the general context, speech sources, andtechnical acoustic properties of the audio scenes. Through our evaluation offour general-purpose audio representation models, we demonstrate that the BEATsmodel significantly surpasses its counterparts. This superiority underscoresthe advantage of models trained on diverse audio collections, confirming theirapplicability to a wide array of auditory tasks, including encoding theenvironment properties necessary for hearable steering. The DEAR dataset andassociated code are available at https://dear-dataset.github.io.
标题:通过多重损失训练和人工数据集自动识别嘻哈音乐中的样本
链接:https://arxiv.org/abs/2502.06364
备注:17 pages, 6 figures
摘要:采样,即在新作品中重复使用录制的音乐或来自其他来源的声音的做法,在嘻哈和说唱等流行音乐流派中很常见。已经出现了许多服务,允许用户识别样本和包含它们的歌曲之间的联系,目的是增强音乐发现。设计一个可以自动执行相同任务的系统是具有挑战性的,因为样本通常会被音调和时间拉伸等音频效果改变,并且可能只有几秒钟长。这项任务的进展甚微,培训数据有限进一步阻碍了进展。在这里,我们展示了在人工数据集上训练的卷积神经网络可以识别商业嘻哈音乐中的真实样本。我们使用音频源分离从非商业音乐录音的几个数据库中提取声乐,和声和泛音元素,并训练模型以在原始音频的转换版本中识别这些元素的子集。我们使用联合分类和度量学习损失来优化模型,并表明它在实际采样实例中的精度比使用声学地标的指纹识别系统高出13%,并且它可以识别音调偏移和时间拉伸的样本。我们还表明,对于我们测试的一半商业音乐录音,我们的模型能够在5秒内定位样本的位置。
摘要:Sampling, the practice of reusing recorded music or sounds from anothersource in a new work, is common in popular music genres like hip-hop and rap.Numerous services have emerged that allow users to identify connections betweensamples and the songs that incorporate them, with the goal of enhancing musicdiscovery. Designing a system that can perform the same task automatically ischallenging, as samples are commonly altered with audio effects like pitch- andtime-stretching and may only be seconds long. Progress on this task has beenminimal and is further blocked by the limited availability of training data.Here, we show that a convolutional neural network trained on an artificialdataset can identify real-world samples in commercial hip-hop music. We extractvocal, harmonic, and percussive elements from several databases ofnon-commercial music recordings using audio source separation, and train themodel to fingerprint a subset of these elements in transformed versions of theoriginal audio. We optimize the model using a joint classification and metriclearning loss and show that it achieves 13% greater precision on real-worldinstances of sampling than a fingerprinting system using acoustic landmarks,and that it can recognize samples that have been both pitch shifted and timestretched. We also show that, for half of the commercial music recordings wetested, our model is capable of locating the position of a sample to withinfive seconds.
标题:使用相对传递函数的端到端多麦克风说话人提取
链接:https://arxiv.org/abs/2502.06285
摘要:本文介绍了一种在混响环境中从含有多个说话人和方向性噪声的混合声中提取期望说话人的多麦克风方法。在这项工作中,我们建议利用瞬时相对传递函数(RTF),估计从参考话语记录在相同的位置作为所需的源。基于RTF的空间线索的有效性进行了比较方向的到来(DOA)为基础的空间线索和传统的频谱嵌入。在具有挑战性的声学场景中的实验结果表明,使用空间线索产生更好的性能比基于频谱的线索,瞬时RTF优于基于DOA的空间线索。
摘要:This paper introduces a multi-microphone method for extracting a desiredspeaker from a mixture involving multiple speakers and directional noise in areverberant environment. In this work, we propose leveraging the instantaneousrelative transfer function (RTF), estimated from a reference utterance recordedin the same position as the desired source. The effectiveness of the RTF-basedspatial cue is compared with direction of arrival (DOA)-based spatial cue andthe conventional spectral embedding. Experimental results in challengingacoustic scenarios demonstrate that using spatial cues yields betterperformance than the spectral-based cue and that the instantaneous RTFoutperforms the DOA-based spatial cue.
标题:通过批量优化改进声学摄像机的外部校准
链接:https://arxiv.org/abs/2502.06196
备注:This paper was accepted and is going to be presented at ICASSP 2025
摘要:声学摄像机在实践中得到了许多应用。声学摄像机中麦克风阵列和视觉传感器的准确和可靠的外部校准对于融合视觉和听觉测量至关重要。现有的校准方法要么需要麦克风阵列几何形状的先验知识,要么依赖于网格搜索,其迭代速度慢或收敛性差。为了克服这些限制,在本文中,我们提出了一种自动校准技术,使用校准板与视觉和声学标记,以确定每个麦克风的位置在摄像机帧。我们制定的外部校准问题(麦克风和视觉传感器之间)作为一个非线性最小二乘问题,并采用批量优化策略来解决相关的问题。大量的数值模拟和实际实验表明,所提出的方法提高了声学相机的外部参数校准的准确性和鲁棒性,与现有的方法相比。为了造福社区,我们在https://github.com/AISLAB-sustech/AcousticCamera上开源了所有代码和数据。
摘要:Acoustic cameras have found many applications in practice. Accurate andreliable extrinsic calibration of the microphone array and visual sensorswithin acoustic cameras is crucial for fusing visual and auditory measurements.Existing calibration methods either require prior knowledge of the microphonearray geometry or rely on grid search which suffers from slow iteration speedor poor convergence. To overcome these limitations, in this paper, we proposean automatic calibration technique using a calibration board with both visualand acoustic markers to identify each microphone position in the camera frame.We formulate the extrinsic calibration problem (between microphones and thevisual sensor) as a nonlinear least squares problem and employ a batchoptimization strategy to solve the associated problem. Extensive numericalsimulations and realworld experiments show that the proposed method improvesboth the accuracy and robustness of extrinsic parameter calibration foracoustic cameras, in comparison to existing methods. To benefit the community,we open-source all the codes and data athttps://github.com/AISLAB-sustech/AcousticCamera.
标题:使用混合时差校准多个同步麦克风阵列
链接:https://arxiv.org/abs/2502.06195
备注:This paper was accepted and is going to be presented at ICASSP 2025
摘要:由多个异步传声器阵列组成的声传感系统的准确校准是声源定位和跟踪的关键。用于这种类型的系统的现有技术的校准方法依赖于麦克风阵列之间的到达时间差和到达方向测量(分别表示为TDOA-M和DOA)。在本文中,为了提高校准精度,我们建议将相邻的声音事件(TDOAS)之间的到达测量的时间差相对于麦克风阵列。更具体地说,我们提出了一个两阶段的校准方法,包括初始值估计(IVE)的过程和最后的联合优化步骤。IVE阶段首先使用混合TDOA(即,TDOAM和TDOA-S)、来自携带扬声器的移动机器人的里程表数据和DOA。随后,麦克风方向估计通过迭代最近点法。最后的联合优化步骤同时估计多个麦克风阵列位置、方向、时间偏移、时钟漂移率和声源位置。仿真和实验结果表明,对于低或中等TDOA噪声水平的情况下,我们的方法优于现有的方法在精度方面。所有代码和数据都可以在https://github.com/AISLABsustech/Hybrid-TDOA-Multi-Calib上找到。
摘要:Accurate calibration of acoustic sensing systems made of multipleasynchronous microphone arrays is essential for satisfactory performance insound source localization and tracking. State-of-the-art calibration methodsfor this type of system rely on the time difference of arrival and direction ofarrival measurements among the microphone arrays (denoted as TDOA-M and DOA,respectively). In this paper, to enhance calibration accuracy, we propose toincorporate the time difference of arrival measurements between adjacent soundevents (TDOAS) with respect to the microphone arrays. More specifically, wepropose a two-stage calibration approach, including an initial value estimation(IVE) procedure and the final joint optimization step. The IVE stage firstinitializes all parameters except for microphone array orientations, usinghybrid TDOA (i.e., TDOAM and TDOA-S), odometer data from a moving robotcarrying a speaker, and DOA. Subsequently, microphone orientations areestimated through the iterative closest point method. The final jointoptimization step estimates multiple microphone array locations, orientations,time offsets, clock drift rates, and sound source locations simultaneously.Both simulation and experiment results show that for scenarios with low ormoderate TDOA noise levels, our approach outperforms existing methods in termsof accuracy. All code and data are available athttps://github.com/AISLABsustech/Hybrid-TDOA-Multi-Calib.
标题:基于自适应过滤器组的神经网络延迟估计和语音增强方法
链接:https://arxiv.org/abs/2502.06098
备注:audio 3A
摘要:时延估计在自适应滤波器声回波消除中起着关键作用。如果估计误差出现,将留下相当大的残余回波。在这里,在本文中,我们提出了一种自适应滤波器组基于神经网络的方法,其中的延迟是由一个银行的自适应滤波器与重叠的时间范围估计,和所有的能量滤波器的权重连接和饲料的分类网络。选择具有最大概率的指标作为估计时延。在此基础上,设计了一种基于神经网络的剩余回波和噪声抑制AEC方案,并采用最优修正对数谱幅度(OMLSA)算法来提高其鲁棒性。此外,一个强大的自动增益控制(AGC)方案与频谱平滑方法的设计,以放大语音段。性能评估表明,我们的计划可以实现更高的性能。
摘要:Time delay estimation (TDE) plays a key role in acoustic echo cancellation(AEC) using adaptive filter method. Considerable residual echo will be left ifestimation error arises. Here, in this paper, we proposed an adaptive filterbank based neural network approach where the delay is estimated by a bank ofadaptive filters with overlapped time scope, and all the energy of filterweights are concatenated and feed to a classification network. The index withmaximal probability is chosen as the estimated delay. Based on this TDE, an AECscheme is designed using a neural network for residual echo and noisesuppression, and the optimally-modified log-spectral amplitude (OMLSA)algorithm is adopted to make it robust. Also, a robust automatic gain control(AGC) scheme with spectrum smoothing method is designed to amplify speechsegments. Performance evaluations reveal that higher performance can beachieved for our scheme.
标题:时间工作记忆:查询引导的片段细化以增强多模式理解
链接:https://arxiv.org/abs/2502.06020
备注:Accepted at NAACL 2025
摘要:多模态基础模型(MFM)在视觉字幕、问答和图文检索等任务中取得了显着成功。然而,这些模型由于其有限的内部容量而面临固有的限制,这限制了它们处理扩展时间序列的能力,这是全面视频和音频分析的关键要求。为了克服这些挑战,我们引入了一个专门的认知模块,时间工作记忆(TWM),其目的是提高时间建模能力的MFM。它选择性地保留跨时间维度的任务相关信息,确保在视频和音频内容的处理过程中保留关键细节。TWM使用查询引导的注意力的方法,专注于时间序列内最丰富的多模态段。通过只保留最相关的内容,TWM优化了模型有限容量的使用,增强了其时间建模能力。该即插即用模块可以轻松集成到现有的MFM中。通过我们的TWM,九种最先进的模型在视频字幕、问答和视频文本检索等任务中表现出显着的性能改进。通过增强时态建模,TWM扩展了MFM的能力,使其能够有效地处理复杂的、对时间敏感的数据。我们的代码可在https://github.com/xid32/NAACL_2025_TWM上获得。
摘要:Multimodal foundation models (MFMs) have demonstrated significant success intasks such as visual captioning, question answering, and image-text retrieval.However, these models face inherent limitations due to their finite internalcapacity, which restricts their ability to process extended temporal sequences,a crucial requirement for comprehensive video and audio analysis. To overcomethese challenges, we introduce a specialized cognitive module, temporal workingmemory (TWM), which aims to enhance the temporal modeling capabilities of MFMs.It selectively retains task-relevant information across temporal dimensions,ensuring that critical details are preserved throughout the processing of videoand audio content. The TWM uses a query-guided attention approach to focus onthe most informative multimodal segments within temporal sequences. Byretaining only the most relevant content, TWM optimizes the use of the model'slimited capacity, enhancing its temporal modeling ability. This plug-and-playmodule can be easily integrated into existing MFMs. With our TWM, ninestate-of-the-art models exhibit significant performance improvements acrosstasks such as video captioning, question answering, and video-text retrieval.By enhancing temporal modeling, TWM extends the capability of MFMs to handlecomplex, time-sensitive data effectively. Our code is available athttps://github.com/xid32/NAACL_2025_TWM.
标题:基于大语言模型的非负矩阵分解用于呼吸声分离
链接:https://arxiv.org/abs/2502.05757
摘要:这项研究代表了大型语言模型(LLM)与非负矩阵分解(NMF)的首次集成,标志着源分离领域的新进展。LLM以两种独特的方式使用:通过为疾病预测提供详细的见解来增强分离结果,并在反馈回路中操作以优化添加到NMF成本函数的基频惩罚。我们在两个数据集上测试了该算法:100个真实测量的合成混合物,以及210个来自临床人体模型的心肺声音记录,包括使用数字听诊器捕获的单独和混合声音。该方法始终优于现有方法,证明了其在显著增强疾病诊断的医学声音分析方面的潜力。
摘要:This study represents the first integration of large language models (LLMs)with non-negative matrix factorization (NMF), marking a novel advancement inthe source separation field. The LLM is employed in two unique ways: enhancingthe separation results by providing detailed insights for disease predictionand operating in a feedback loop to optimize a fundamental frequency penaltyadded to the NMF cost function. We tested the algorithm on two datasets: 100synthesized mixtures of real measurements, and 210 recordings of heart and lungsounds from a clinical manikin including both individual and mixed sounds,captured using a digital stethoscope. The approach consistently outperformedexisting methods, demonstrating its potential to significantly enhance medicalsound analysis for disease diagnostics.
标题:IndexTTC:工业级可控且高效的Zero-Shot文本转语音系统
链接:https://arxiv.org/abs/2502.05512
摘要:近年来,基于大语言模型(LLM)的文语转换(TTS)系统以其高自然度和强大的zero-shot语音克隆能力逐渐成为业界的主流,本文介绍了基于XTTS和Toronto模型的IndexTTS系统。我们添加了一些新的改进。具体来说,在中文场景中,我们采用了一种混合建模的方法,结合字符和字符,使多音字和长尾字符的发音可控。我们还进行了比较分析的矢量量化(VQ)与标量量化(FSQ)的声学语音令牌的码本利用率。为了进一步提高语音克隆的效果和稳定性,我们引入了基于一致性的语音条件编码器,并将语音解码器替换为BigVGAN 2。与XTTS相比,在自然度、内容一致性、zero-shot语音克隆等方面都有了显著提升。对于开源中流行的TTS系统,如Fish-Speech、CosyVoice 2、FireRedTTS和F5-TTS,IndexTTS具有相对简单的训练过程,更可控的使用,更快的推理速度。此外,它的性能超过了这些系统。我们的演示可在https://index-tts.github.io上获得。
摘要:Recently, large language model (LLM) based text-to-speech (TTS) systems havegradually become the mainstream in the industry due to their high naturalnessand powerful zero-shot voice cloning capabilities.Here, we introduce theIndexTTS system, which is mainly based on the XTTS and Tortoise model. We addsome novel improvements. Specifically, in Chinese scenarios, we adopt a hybridmodeling method that combines characters and pinyin, making the pronunciationsof polyphonic characters and long-tail characters controllable. We alsoperformed a comparative analysis of the Vector Quantization (VQ) withFinite-Scalar Quantization (FSQ) for codebook utilization of acoustic speechtokens. To further enhance the effect and stability of voice cloning, weintroduce a conformer-based speech conditional encoder and replace thespeechcode decoder with BigVGAN2. Compared with XTTS, it has achievedsignificant improvements in naturalness, content consistency, and zero-shotvoice cloning. As for the popular TTS systems in the open-source, such asFish-Speech, CosyVoice2, FireRedTTS and F5-TTS, IndexTTS has a relativelysimple training process, more controllable usage, and faster inference speed.Moreover, its performance surpasses that of these systems. Our demos areavailable at https://index-tts.github.io.
标题:利用离散音调条件流匹配模型增强表达性语音转换
链接:https://arxiv.org/abs/2502.05471
备注:Accepted by ICASSP 2025
摘要:本文介绍了PFlow-VC,一个条件流匹配语音转换模型,利用细粒度的离散音高令牌和目标说话人提示信息表达语音转换(VC)。以前的VC作品主要集中在说话人转换,需要进一步探索如何增强音色转换的表现力(如韵律和情感)。与以前的方法不同,我们采用了一种简单而有效的方法来提高语音转换模型的风格表现力。具体来说,我们预训练了一个自监督的音高VQVAE模型,以离散化与说话人无关的音高信息,并利用一个掩蔽的音高条件流匹配模型进行梅尔频谱图合成,该模型为说话人转换模型提供了上下文音高建模能力,有效地提高了语音风格传输能力。此外,我们通过将全局音色嵌入与时变音色令牌相结合来提高音色相似性。在看不见的LibriTTS测试干净和情感语音数据集ESD上的实验表明,PFlow-VC模型在音色转换和风格转换方面都具有优越性。音频样本可在演示页面https://speechai-demo.github.io/PFlow-VC/上获得。
摘要:This paper introduces PFlow-VC, a conditional flow matching voice conversionmodel that leverages fine-grained discrete pitch tokens and target speakerprompt information for expressive voice conversion (VC). Previous VC worksprimarily focus on speaker conversion, with further exploration needed inenhancing expressiveness (such as prosody and emotion) for timbre conversion.Unlike previous methods, we adopt a simple and efficient approach to enhancethe style expressiveness of voice conversion models. Specifically, we pretraina self-supervised pitch VQVAE model to discretize speaker-irrelevant pitchinformation and leverage a masked pitch-conditioned flow matching model forMel-spectrogram synthesis, which provides in-context pitch modelingcapabilities for the speaker conversion model, effectively improving the voicestyle transfer capacity. Additionally, we improve timbre similarity bycombining global timbre embeddings with time-varying timbre tokens. Experimentson unseen LibriTTS test-clean and emotional speech dataset ESD show thesuperiority of the PFlow-VC model in both timbre conversion and style transfer.Audio samples are available on the demo pagehttps://speechai-demo.github.io/PFlow-VC/.
标题:Koel-TTC:通过偏好对齐和分类器自由指导增强基于LLM的语音生成
链接:https://arxiv.org/abs/2502.05236
摘要:虽然自回归语音令牌生成模型产生具有显著多样性和自然性的语音,但其固有的可控性缺乏通常会导致诸如不符合条件输入的幻觉和不期望的发声等问题。我们介绍Koel-TTS,一套增强的编码器-解码器Transformer TTS模型,通过将自动语音识别和说话人验证模型指导的偏好对齐技术,解决这些挑战。此外,我们将无分类器的指导,以进一步提高合成坚持的成绩单和参考扬声器音频。我们的实验表明,这些优化显着提高目标说话人的相似性,可懂度和合成语音的自然度。值得注意的是,Koel-TTS直接将文本和上下文音频映射到声学标记,并且在上述指标上,尽管在一个非常小的数据集上进行了训练,但其性能优于最先进的TTS模型。音频样本和演示可以在我们的网站上找到。
摘要:While autoregressive speech token generation models produce speech withremarkable variety and naturalness, their inherent lack of controllabilityoften results in issues such as hallucinations and undesired vocalizations thatdo not conform to conditioning inputs. We introduce Koel-TTS, a suite ofenhanced encoder-decoder Transformer TTS models that address these challengesby incorporating preference alignment techniques guided by automatic speechrecognition and speaker verification models. Additionally, we incorporateclassifier-free guidance to further improve synthesis adherence to thetranscript and reference speaker audio. Our experiments demonstrate that theseoptimizations significantly enhance target speaker similarity, intelligibility,and naturalness of synthesized speech. Notably, Koel-TTS directly maps text andcontext audio to acoustic tokens, and on the aforementioned metrics,outperforms state-of-the-art TTS models, despite being trained on asignificantly smaller dataset. Audio samples and demos are available on ourwebsite.
标题:校准者编码器:自我注意Transformer可以成为自我转换器
链接:https://arxiv.org/abs/2502.05232
摘要:现代的自动语音识别系统,包括RNN传感器和基于注意力的编码器-解码器(AED),其设计使得编码器不需要改变从音频序列到嵌入的信息的时间位置;在解码期间处理与最终文本输出的对齐。我们发现,近年来采用的基于变换器的编码器实际上能够在解码之前的前向传递期间内部执行对齐。这一新现象使一个更简单和更有效的模型,“对齐器编码器”。为了训练它,我们放弃了RNN-T的动态规划,以支持AED的逐帧交叉熵损失,而解码器采用了RNN-T的更轻的纯文本递归,而没有学习交叉注意-它只是从开始按顺序扫描嵌入帧,每个生成一个令牌,直到预测消息结束。我们进行的实验表明,性能非常接近最先进的,包括一个特殊的推理配置,使长形式的识别。在一个代表性的比较中,我们测量了我们模型的总推理时间比RNN-T快2倍,比AED快16倍。最后,我们发现音频-文本对齐在某一层的自我注意力权重中清晰可见,可以说是进行了“自我转换”。
摘要:Modern systems for automatic speech recognition, including the RNN-Transducerand Attention-based Encoder-Decoder (AED), are designed so that the encoder isnot required to alter the time-position of information from the audio sequenceinto the embedding; alignment to the final text output is processed duringdecoding. We discover that the transformer-based encoder adopted in recentyears is actually capable of performing the alignment internally during theforward pass, prior to decoding. This new phenomenon enables a simpler and moreefficient model, the "Aligner-Encoder". To train it, we discard the dynamicprogramming of RNN-T in favor of the frame-wise cross-entropy loss of AED,while the decoder employs the lighter text-only recurrence of RNN-T withoutlearned cross-attention -- it simply scans embedding frames in order from thebeginning, producing one token each until predicting the end-of-message. Weconduct experiments demonstrating performance remarkably close to the state ofthe art, including a special inference configuration enabling long-formrecognition. In a representative comparison, we measure the total inferencetime for our model to be 2x faster than RNN-T and 16x faster than AED. Lastly,we find that the audio-text alignment is clearly visible in the self-attentionweights of a certain layer, which could be said to perform "self-transduction".
标题:离散语音令牌的最新进展:评论
链接:https://arxiv.org/abs/2502.06490
备注:26 pages, 8 figures, 3 tables. Work in progress
摘要:大型语言模型(LLM)时代语音生成技术的快速发展已经将离散语音令牌确立为语音表示的基础范式。这些令牌的特点是离散、紧凑和简洁,不仅有利于有效的传输和存储,而且与语言建模框架具有内在的兼容性,能够将语音无缝集成到文本主导的LLM架构中。目前的研究将离散语音标记分为两大类:声学标记和语义标记,每一类都已经发展成为一个丰富的研究领域,其特点是独特的设计理念和方法。这项调查系统地综合了现有的分类和离散语音标记化的最新创新,进行了严格的检查的优势和每个范例的局限性,并提出了系统的实验比较令牌类型。此外,我们还确定了该领域持续存在的挑战,并提出了潜在的研究方向,旨在提供可操作的见解,以激励离散语音令牌开发和应用的未来发展。
摘要:The rapid advancement of speech generation technologies in the era of largelanguage models (LLMs) has established discrete speech tokens as a foundationalparadigm for speech representation. These tokens, characterized by theirdiscrete, compact, and concise nature, are not only advantageous for efficienttransmission and storage, but also inherently compatible with the languagemodeling framework, enabling seamless integration of speech into text-dominatedLLM architectures. Current research categorizes discrete speech tokens into twoprincipal classes: acoustic tokens and semantic tokens, each of which hasevolved into a rich research domain characterized by unique design philosophiesand methodological approaches. This survey systematically synthesizes theexisting taxonomy and recent innovations in discrete speech tokenization,conducts a critical examination of the strengths and limitations of eachparadigm, and presents systematic experimental comparisons across token types.Furthermore, we identify persistent challenges in the field and proposepotential research directions, aiming to offer actionable insights to inspirefuture advancements in the development and application of discrete speechtokens.
标题:论表演者和代理人注意力在口语识别中的使用
链接:https://arxiv.org/abs/2502.05841
备注:5 pages, 1 figure
摘要:语言识别(LID)的方法之一涉及使用自监督学习从预训练的模型中导出语音表示,然后微调LID任务的模型。用于LID的最先进方法使用基于注意力的统计池化层来促进跨从预训练模型提取的嵌入向量的时间帧的上下文信息的聚合。在本文中,我们深入探讨了最近提出的注意力机制,即表演者和代理注意力,结合统计池层。LID实验在三个数据集上进行:VoxPopuli,FLEURS和VoxLingua。我们将他们的表现与普通的自我关注进行比较。我们的研究结果表明,表演者的注意力优于自我注意力和代理人的注意力表现出相当或偶尔优于自我注意力的性能,同时也是计算成本较低。
摘要:One of the methods for language Identification (LID) involves deriving speechrepresentation from pre-trained models using self-supervised learning, followedby fine-tuning the model for the LID task. State-of-the-art approaches for LIDuse an attention-based statistical pooling layer to facilitate the aggregationof contextual information across time frames of the embedding vectors extractedfrom the pre-trained model. In this paper, we delve into exploring recentlyproposed attention mechanisms, namely performer and agent-attention, inconjunction with the statistical pooling layer. The LID experiments areperformed on three datasets: VoxPopuli, FLEURS, and VoxLingua. We compare theirperformance against vanilla self-attention. Our findings suggest thatperformer-attention outperforms self-attention and agent-attention exhibitscomparable or occasionally superior performance to self-attention, while alsobeing computationally less expensive.
标题:通过语音基础模型知识提炼的视听表示学习
链接:https://arxiv.org/abs/2502.05766
备注:accepted to Pattern Recognition
摘要:视听表征学习对于推进多模态语音处理任务(如唇读和视听语音识别)至关重要。最近,语音基础模型(SFM)在各种语音相关的任务中表现出显着的泛化能力。在此基础上,我们提出了一个视听表示学习模型,利用跨模态的知识蒸馏SFM。在我们的方法中,SFM充当教师,从其中使用干净的音频输入提取多层隐藏表示。我们还引入了多教师集成方法来提取学生,接收视听数据作为输入。一种新的代表性的知识蒸馏损失的训练学生在预训练,这也是在微调,以进一步提高下游任务的性能。我们的实验利用了自监督SFM,WavLM和监督SFM,iFLYTEK语音。结果表明,我们提出的方法在自动语音识别、视觉语音识别和视听语音识别任务中取得了优于或至少与以前最先进的基线相当的性能。此外,全面的消融研究和可视化的学习表示进行评估我们提出的方法的有效性。
摘要:Audio-visual representation learning is crucial for advancing multimodalspeech processing tasks, such as lipreading and audio-visual speechrecognition. Recently, speech foundation models (SFMs) have shown remarkablegeneralization capabilities across various speech-related tasks. Building onthis progress, we propose an audio-visual representation learning model thatleverages cross-modal knowledge distillation from SFMs. In our method, SFMsserve as teachers, from which multi-layer hidden representations are extractedusing clean audio inputs. We also introduce a multi-teacher ensemble method todistill the student, which receives audio-visual data as inputs. A novelrepresentational knowledge distillation loss is employed to train the studentduring pretraining, which is also applied during finetuning to further enhancethe performance on downstream tasks. Our experiments utilized both aself-supervised SFM, WavLM, and a supervised SFM, iFLYTEK-speech. The resultsdemonstrated that our proposed method achieved superior or at least comparableperformance to previous state-of-the-art baselines across automatic speechrecognition, visual speech recognition, and audio-visual speech recognitiontasks. Additionally, comprehensive ablation studies and the visualization oflearned representations were conducted to evaluate the effectiveness of ourproposed method.
标题:野外合成语音检测的少即是多
链接:https://arxiv.org/abs/2502.05674
摘要:在语音自监督学习的推动下,最先进的合成语音检测器在ASVspoof等流行的基准测试中实现了低错误率。然而,先前的基准并没有解决语音中广泛的真实世界的可变性。报告的错误率在现实条件下是否真实?为了评估检测器的故障模式和鲁棒性控制下的分布变化,我们介绍ShiftySpeech,一个基准超过3000小时的合成语音7域,6 TTS系统,12声码器,和3种语言。我们发现,所有的分布变化都会降低模型的性能,与之前的研究结果相反,对更多的声码器,扬声器或数据增强进行训练并不能保证更好的泛化。事实上,我们发现,在不太多样化的数据上进行训练可以获得更好的泛化能力,并且使用来自单个精心选择的声码器和单个扬声器的样本的检测器拟合在具有挑战性的野外基准测试中获得了最先进的结果。
摘要:Driven by advances in self-supervised learning for speech, state-of-the-artsynthetic speech detectors have achieved low error rates on popular benchmarkssuch as ASVspoof. However, prior benchmarks do not address the wide range ofreal-world variability in speech. Are reported error rates realistic inreal-world conditions? To assess detector failure modes and robustness undercontrolled distribution shifts, we introduce ShiftySpeech, a benchmark withmore than 3000 hours of synthetic speech from 7 domains, 6 TTS systems, 12vocoders, and 3 languages. We found that all distribution shifts degraded modelperformance, and contrary to prior findings, training on more vocoders,speakers, or with data augmentation did not guarantee better generalization. Infact, we found that training on less diverse data resulted in bettergeneralization, and that a detector fit using samples from a single carefullyselected vocoder and a single speaker achieved state-of-the-art results on thechallenging In-the-Wild benchmark.
标题:可扩展的基于自监督表示的语音质量评估的蒸馏和修剪
链接:https://arxiv.org/abs/2502.05356
备注:Accepted at ICASSP 2025
摘要:在本文中,我们研究蒸馏和修剪方法,以减少模型的大小,非侵入性的语音质量评估的基础上,自监督表示。我们的实验建立在XLS-R-SQA上,这是一种使用wav 2 vec 2.0 XLS-R嵌入的语音质量评估模型。我们在一个大型的平均意见评分数据集上重新训练了这个模型,其中包含超过100,000个标记的剪辑。对于蒸馏,使用这个模型作为老师,我们产生伪标签的未标记的退化语音信号和训练学生模型的大小不同。对于修剪,我们使用数据驱动策略。虽然数据驱动的修剪在更大的模型尺寸下表现更好,但对未标记数据的蒸馏对于较小的模型尺寸更有效。蒸馏可以将基线与地面实况MOS标签的相关性与基于XLS-R的教师模型的相关性之间的差距减半,同时与教师模型相比,将模型大小减少两个数量级。
摘要:In this paper, we investigate distillation and pruning methods to reducemodel size for non-intrusive speech quality assessment based on self-supervisedrepresentations. Our experiments build on XLS-R-SQA, a speech qualityassessment model using wav2vec 2.0 XLS-R embeddings. We retrain this model on alarge compilation of mean opinion score datasets, encompassing over 100,000labeled clips. For distillation, using this model as a teacher, we generatepseudo-labels on unlabeled degraded speech signals and train student models ofvarying sizes. For pruning, we use a data-driven strategy. While data-drivenpruning performs better at larger model sizes, distillation on unlabeled datais more effective for smaller model sizes. Distillation can halve the gapbetween the baseline's correlation with ground-truth MOS labels and that of theXLS-R-based teacher model, while reducing model size by two orders of magnitudecompared to the teacher model.
标题:离散语音令牌的最新进展:评论
链接:https://arxiv.org/abs/2502.06490
备注:26 pages, 8 figures, 3 tables. Work in progress
摘要:在大语言模型时代,语音生成技术的快速发展已经将离散语音标记作为语音表示的基本范式。这些令牌的特点是离散、紧凑和简洁,不仅有利于有效的传输和存储,而且与语言建模框架具有内在的兼容性,能够将语音无缝集成到文本主导的LLM架构中。目前的研究将离散语音标记分为两大类:声学标记和语义标记,每一类都已经发展成为一个丰富的研究领域,其特点是独特的设计理念和方法。这项调查系统地综合了现有的分类和离散语音标记化的最新创新,进行了严格的检查的优势和每个范例的局限性,并提出了系统的实验比较令牌类型。此外,我们还确定了该领域持续存在的挑战,并提出了潜在的研究方向,旨在提供可操作的见解,以激励离散语音令牌开发和应用的未来发展。
摘要:The rapid advancement of speech generation technologies in the era of largelanguage models (LLMs) has established discrete speech tokens as a foundationalparadigm for speech representation. These tokens, characterized by theirdiscrete, compact, and concise nature, are not only advantageous for efficienttransmission and storage, but also inherently compatible with the languagemodeling framework, enabling seamless integration of speech into text-dominatedLLM architectures. Current research categorizes discrete speech tokens into twoprincipal classes: acoustic tokens and semantic tokens, each of which hasevolved into a rich research domain characterized by unique design philosophiesand methodological approaches. This survey systematically synthesizes theexisting taxonomy and recent innovations in discrete speech tokenization,conducts a critical examination of the strengths and limitations of eachparadigm, and presents systematic experimental comparisons across token types.Furthermore, we identify persistent challenges in the field and proposepotential research directions, aiming to offer actionable insights to inspirefuture advancements in the development and application of discrete speechtokens.
标题:论表演者和代理人注意力在口语识别中的使用
链接:https://arxiv.org/abs/2502.05841
备注:5 pages, 1 figure
摘要:语言识别(LID)的方法之一涉及使用自监督学习从预训练的模型中导出语音表示,然后微调LID任务的模型。用于LID的最先进方法使用基于注意力的统计池化层来促进跨从预训练模型提取的嵌入向量的时间帧的上下文信息的聚合。在本文中,我们深入探讨了最近提出的注意力机制,即表演者和代理注意力,结合统计池层。LID实验在三个数据集上进行:VoxPopuli,FLEURS和VoxLingua。我们将他们的表现与普通的自我关注进行比较。我们的研究结果表明,表演者的注意力优于自我注意力和代理人的注意力表现出相当或偶尔优于自我注意力的性能,同时也是计算成本较低。
摘要:One of the methods for language Identification (LID) involves deriving speechrepresentation from pre-trained models using self-supervised learning, followedby fine-tuning the model for the LID task. State-of-the-art approaches for LIDuse an attention-based statistical pooling layer to facilitate the aggregationof contextual information across time frames of the embedding vectors extractedfrom the pre-trained model. In this paper, we delve into exploring recentlyproposed attention mechanisms, namely performer and agent-attention, inconjunction with the statistical pooling layer. The LID experiments areperformed on three datasets: VoxPopuli, FLEURS, and VoxLingua. We compare theirperformance against vanilla self-attention. Our findings suggest thatperformer-attention outperforms self-attention and agent-attention exhibitscomparable or occasionally superior performance to self-attention, while alsobeing computationally less expensive.
标题:知识提炼和结构化修剪对自我监督语音模型的协同效应
链接:https://arxiv.org/abs/2502.05837
备注:5 pages, 2 figures, 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing
摘要:传统上,知识蒸馏(KD)用于模型压缩,通常导致次优性能。在本文中,我们评估了在自监督学习(SSL)范式下,将KD损失与其他修剪技术(包括低秩因子分解(LRF)和l0正则化)相结合对基于一致性的预训练网络的影响。我们还提出了一种策略来联合修剪和训练基于RNN-T的ASR模型,证明了这种方法与首先修剪预训练的网络然后将其用于ASR训练相比具有更好的性能。这种方法导致字错误率的显著降低:l0和KD组合实现了最佳的非流性能,相对于基线的相对字错误率(RWER)改善了8.9%,而LRF和KD组合产生了流ASR的最佳结果,将RWER改善了13.4%。
摘要:Traditionally, Knowledge Distillation (KD) is used for model compression,often leading to suboptimal performance. In this paper, we evaluate the impactof combining KD loss with alternative pruning techniques, including Low-RankFactorization (LRF) and l0 regularization, on a conformer-based pre-trainednetwork under the paradigm of Self-Supervised Learning (SSL). We also propose astrategy to jointly prune and train an RNN-T-based ASR model, demonstratingthat this approach yields superior performance compared to pruning apre-trained network first and then using it for ASR training. This approach ledto a significant reduction in word error rate: l0 and KD combination achievesthe best non-streaming performance, with a 8.9% Relative Word Error Rate (RWER)improvement over the baseline, while LRF and KD combination yields the bestresults for streaming ASR, improving RWER by 13.4%.
标题:通过语音基础模型知识提炼的视听表示学习
链接:https://arxiv.org/abs/2502.05766
备注:accepted to Pattern Recognition
摘要:视听表征学习对于推进多模态语音处理任务(如唇读和视听语音识别)至关重要。最近,语音基础模型(SFM)在各种语音相关的任务中表现出显着的泛化能力。在此基础上,我们提出了一个视听表示学习模型,利用跨模态的知识蒸馏SFM。在我们的方法中,SFM充当教师,从其中使用干净的音频输入提取多层隐藏表示。我们还引入了多教师集成方法来提取学生,接收视听数据作为输入。一种新的代表性的知识蒸馏损失的训练学生在预训练,这也是在微调,以进一步提高下游任务的性能。我们的实验利用了自监督SFM,WavLM和监督SFM,iFLYTEK语音。结果表明,我们提出的方法在自动语音识别、视觉语音识别和视听语音识别任务中取得了优于或至少与以前最先进的基线相当的性能。此外,全面的消融研究和可视化的学习表示进行评估我们提出的方法的有效性。
摘要:Audio-visual representation learning is crucial for advancing multimodalspeech processing tasks, such as lipreading and audio-visual speechrecognition. Recently, speech foundation models (SFMs) have shown remarkablegeneralization capabilities across various speech-related tasks. Building onthis progress, we propose an audio-visual representation learning model thatleverages cross-modal knowledge distillation from SFMs. In our method, SFMsserve as teachers, from which multi-layer hidden representations are extractedusing clean audio inputs. We also introduce a multi-teacher ensemble method todistill the student, which receives audio-visual data as inputs. A novelrepresentational knowledge distillation loss is employed to train the studentduring pretraining, which is also applied during finetuning to further enhancethe performance on downstream tasks. Our experiments utilized both aself-supervised SFM, WavLM, and a supervised SFM, iFLYTEK-speech. The resultsdemonstrated that our proposed method achieved superior or at least comparableperformance to previous state-of-the-art baselines across automatic speechrecognition, visual speech recognition, and audio-visual speech recognitiontasks. Additionally, comprehensive ablation studies and the visualization oflearned representations were conducted to evaluate the effectiveness of ourproposed method.
标题:无创肌电言语神经假体:几何视角
链接:https://arxiv.org/abs/2502.05762
摘要:在这篇文章中,我们提出了一个高带宽的自我中心的神经肌肉语音接口翻译无声的语音发音成文本和音频。具体来说,我们收集肌电图(EMG)信号从多个articulatorysites在脸上和脖子上的个人发音语音在一个alaryngeal的方式来执行EMG到文本或EMG到音频的翻译。这样的接口对于恢复由于喉切除术、神经肌肉疾病、中风或创伤引起的损伤(例如,放射治疗毒性)。以前的工作集中在训练文本或语音合成模型,使用在可听语音发音期间收集的EMG或通过将音频目标从在可听发音期间收集的EMG转移到在无声发音期间收集的EMG。然而,这种模式不适合那些已经失去清晰发音能力的人。我们是第一个以开源方式仅使用在无声发音语音期间收集的EMG来呈现无障碍EMG到文本和EMG到音频转换的公司。在一个有限的词汇语料库,我们的方法实现了近2.4倍的单词错误率的改善与模型,是25倍小,利用固有的几何形状的EMG。
摘要:In this article, we present a high-bandwidth egocentric neuromuscular speechinterface for translating silently voiced speech articulations into textandaudio. Specifically, we collect electromyogram (EMG) signals from multiplearticulatorysites on the face and neck as individuals articulate speech in analaryngeal manner to perform EMG-to-text or EMG-to-audio translation. Such aninterface is useful for restoring audible speech in individuals who have lostthe ability to speak intelligibly due to laryngectomy, neuromuscular disease,stroke, or trauma-induced damage (e.g., radiotherapy toxicity) to speecharticulators. Previous works have focused on training text or speech synthesismodels using EMG collected during audible speech articulations or bytransferring audio targets from EMG collected during audible articulation toEMG collected during silent articulation. However, such paradigms are notsuited for individuals who have already lost the ability to audibly articulatespeech. We are the first to present an alignment-free EMG-to-text andEMG-to-audio conversion using only EMG collected during silently articulatedspeech in an open-sourced manner. On a limited vocabulary corpora, our approachachieves almost 2.4x improvement in word error rate with a model that is 25xsmaller by leveraging the inherent geometry of EMG.
标题:通过视听自蒸馏预训练和说话人适应实现目标说话人唇读
链接:https://arxiv.org/abs/2502.05758
备注:accepted to ESWA journal
摘要:唇读是在噪声环境中促进人机交互的重要技术。我们以前开发的自监督学习方法AV2vec利用多模态自蒸馏,在英语LRS3数据集上表现出与说话者无关的唇读性能。然而,AV2vec面临着诸如高培训成本和潜在的英语以外语言(如中文)唇读视听数据稀缺等挑战。此外,大多数研究集中在speakerindependent唇读模型,这很难占整个迪说话风格的实质性变化?演讲者为了解决这些问题,我们提出了一个全面的办法。首先,我们研究了跨语言迁移学习,从源语言中调整预训练的AV2vec模型,并针对目标语言的唇读任务进行优化。其次,我们通过说话人自适应策略来提高特定目标说话人唇读的准确性,这在以前的研究中没有得到广泛的探讨。第三,唇读与唇感兴趣区域(ROI)和人脸输入的互补性能分析后,我们引入了一个模型集成策略,集成了两个,signi?巧妙地提高模型性能。我们的方法在Chatterfly数据集的评估集上实现了77.3%的字符错误率(CER),低于2024年聊天场景中国唇读挑战赛的最高结果。
摘要:Lipreading is an important technique for facilitating human-computerinteraction in noisy environments. Our previously developed self-supervisedlearning method, AV2vec, which leverages multimodal self-distillation, hasdemonstrated promising performance in speaker-independent lipreading on theEnglish LRS3 dataset. However, AV2vec faces challenges such as high trainingcosts and a potential scarcity of audio-visual data for lipreading in languagesother than English, such as Chinese. Additionally, most studies concentrate onspeakerindependent lipreading models, which struggle to account for thesubstantial variation in speaking styles across di?erent speakers. To addressthese issues, we propose a comprehensive approach. First, we investigatecross-lingual transfer learning, adapting a pre-trained AV2vec model from asource language and optimizing it for the lipreading task in a target language.Second, we enhance the accuracy of lipreading for specific target speakersthrough a speaker adaptation strategy, which is not extensively explored inprevious research. Third, after analyzing the complementary performance oflipreading with lip region-of-interest (ROI) and face inputs, we introduce amodel ensembling strategy that integrates both, signi?cantly boosting modelperformance. Our method achieved a character error rate (CER) of 77.3% on theevaluation set of the ChatCLR dataset, which is lower than the top result fromthe 2024 Chat-scenario Chinese Lipreading Challenge.
标题:野外合成语音检测的少即是多
链接:https://arxiv.org/abs/2502.05674
摘要:在语音自监督学习的推动下,最先进的合成语音检测器在ASVspoof等流行的基准测试中实现了低错误率。然而,先前的基准并没有解决语音中广泛的真实世界的可变性。报告的错误率在现实条件下是否真实?为了评估检测器的故障模式和鲁棒性控制下的分布变化,我们介绍ShiftySpeech,一个基准超过3000小时的合成语音7域,6 TTS系统,12声码器,和3种语言。我们发现,所有的分布变化都会降低模型的性能,与之前的研究结果相反,对更多的声码器,扬声器或数据增强进行训练并不能保证更好的泛化。事实上,我们发现,在不太多样化的数据上进行训练可以获得更好的泛化能力,并且使用来自单个精心选择的声码器和单个扬声器的样本的检测器拟合在具有挑战性的野外基准测试中获得了最先进的结果。
摘要:Driven by advances in self-supervised learning for speech, state-of-the-artsynthetic speech detectors have achieved low error rates on popular benchmarkssuch as ASVspoof. However, prior benchmarks do not address the wide range ofreal-world variability in speech. Are reported error rates realistic inreal-world conditions? To assess detector failure modes and robustness undercontrolled distribution shifts, we introduce ShiftySpeech, a benchmark withmore than 3000 hours of synthetic speech from 7 domains, 6 TTS systems, 12vocoders, and 3 languages. We found that all distribution shifts degraded modelperformance, and contrary to prior findings, training on more vocoders,speakers, or with data augmentation did not guarantee better generalization. Infact, we found that training on less diverse data resulted in bettergeneralization, and that a detector fit using samples from a single carefullyselected vocoder and a single speaker achieved state-of-the-art results on thechallenging In-the-Wild benchmark.
标题:无偏见切片Wasserstein内核,用于高质量音频字幕
链接:https://arxiv.org/abs/2502.05435
备注:17 pages, 9 tables, 2 figures
摘要:教师强迫式的语音字幕训练通常会由于训练和推理的不匹配而导致暴露偏差。先前的工作提出了对比方法来处理字幕退化。然而,对比方法忽略了时间信息时,测量声学和语言模态的相似性,导致性能较差。在这项工作中,我们通过引入无偏切片Wasserstein RBF(USW-RBF)内核来开发时间相似性得分,该内核配备了旋转位置嵌入以考虑跨模态的时间信息。与传统的切片Wasserstein RBF核相比,我们可以通过Monte Carlo估计来形成USW-RBF核的无偏估计。因此,它非常适合于随机梯度优化算法,其近似误差以$\mathcal{O}(L^{-1/2})$的参数速率减小,具有$L$ Monte Carlo样本。此外,我们引入了一个音频字幕框架的基础上无偏切片Wasserstein内核,结合随机解码方法,以减轻字幕退化过程中的生成过程。我们在AudioCaps和Clotho两个数据集上进行了大量的定量和定性实验,以说明生成高质量音频字幕的能力。实验结果表明,我们的框架是能够增加字幕长度,词汇多样性和文本到音频的自检索准确率。
摘要:Teacher-forcing training for audio captioning usually leads to exposure biasdue to training and inference mismatch. Prior works propose the contrastivemethod to deal with caption degeneration. However, the contrastive methodignores the temporal information when measuring similarity across acoustic andlinguistic modalities, leading to inferior performance. In this work, wedevelop the temporal-similarity score by introducing the unbiased slicedWasserstein RBF (USW-RBF) kernel equipped with rotary positional embedding toaccount for temporal information across modalities. In contrast to theconventional sliced Wasserstein RBF kernel, we can form an unbiased estimationof USW-RBF kernel via Monte Carlo estimation. Therefore, it is well-suited tostochastic gradient optimization algorithms, and its approximation errordecreases at a parametric rate of $\mathcal{O}(L^{-1/2})$ with $L$ Monte Carlosamples. Additionally, we introduce an audio captioning framework based on theunbiased sliced Wasserstein kernel, incorporating stochastic decoding methodsto mitigate caption degeneration during the generation process. We conductextensive quantitative and qualitative experiments on two datasets, AudioCapsand Clotho, to illustrate the capability of generating high-quality audiocaptions. Experimental results show that our framework is able to increasecaption length, lexical diversity, and text-to-audio self-retrieval accuracy.
标题:可扩展的基于自监督表示的语音质量评估的蒸馏和修剪
链接:https://arxiv.org/abs/2502.05356
备注:Accepted at ICASSP 2025
摘要:在本文中,我们研究蒸馏和修剪方法,以减少模型的大小,非侵入性的语音质量评估的基础上,自监督表示。我们的实验建立在XLS-R-SQA上,这是一种使用wav 2 vec 2.0 XLS-R嵌入的语音质量评估模型。我们在一个大型的平均意见评分数据集上重新训练了这个模型,其中包含超过100,000个标记的剪辑。对于蒸馏,使用这个模型作为老师,我们产生伪标签的未标记的退化语音信号和训练学生模型的大小不同。对于修剪,我们使用数据驱动策略。虽然数据驱动的修剪在较大的模型大小下表现得更好,但对未标记数据的蒸馏对于较小的模型大小更有效。蒸馏可以将基线与地面实况MOS标签的相关性与基于XLS-R的教师模型的相关性之间的差距减半,同时与教师模型相比,将模型大小减少两个数量级。
摘要:In this paper, we investigate distillation and pruning methods to reducemodel size for non-intrusive speech quality assessment based on self-supervisedrepresentations. Our experiments build on XLS-R-SQA, a speech qualityassessment model using wav2vec 2.0 XLS-R embeddings. We retrain this model on alarge compilation of mean opinion score datasets, encompassing over 100,000labeled clips. For distillation, using this model as a teacher, we generatepseudo-labels on unlabeled degraded speech signals and train student models ofvarying sizes. For pruning, we use a data-driven strategy. While data-drivenpruning performs better at larger model sizes, distillation on unlabeled datais more effective for smaller model sizes. Distillation can halve the gapbetween the baseline's correlation with ground-truth MOS labels and that of theXLS-R-based teacher model, while reducing model size by two orders of magnitudecompared to the teacher model.
标题:为音乐表演问题解答学习音乐表达
链接:https://arxiv.org/abs/2502.06710
备注:Accepted at EMNLP 2024
摘要:音乐表演是视听建模的典型场景。与具有稀疏音频的常见场景不同,音乐表演始终连续地涉及密集的音频信号。虽然现有的多模态学习方法的音频-视频QA表现出令人印象深刻的能力,在一般情况下,他们是无法处理的基本问题,在音乐表演:他们underexplore的多模态信号之间的相互作用的性能,并没有考虑到乐器和音乐的独特特征。因此,现有的方法倾向于不准确地回答关于音乐表演的问题。为了弥补上述研究空白,(i)考虑到音乐数据固有的复杂多模态互联性,我们的主要骨干旨在将多模态交互纳入音乐背景中;(ii)为了使模型能够学习音乐特征,我们在当前音乐数据集中注释和发布节奏和音乐源;(iii)对于时间感知的视听建模,我们将模型的音乐预测与时间维度对齐。我们的实验显示了对Music AVQA数据集的最新影响。我们的代码可在https://github.com/xid32/Amuse上获得。
摘要:Music performances are representative scenarios for audio-visual modeling.Unlike common scenarios with sparse audio, music performances continuouslyinvolve dense audio signals throughout. While existing multimodal learningmethods on the audio-video QA demonstrate impressive capabilities in generalscenarios, they are incapable of dealing with fundamental problems within themusic performances: they underexplore the interaction between the multimodalsignals in performance and fail to consider the distinctive characteristics ofinstruments and music. Therefore, existing methods tend to answer questionsregarding musical performances inaccurately. To bridge the above research gaps,(i) given the intricate multimodal interconnectivity inherent to music data,our primary backbone is designed to incorporate multimodal interactions withinthe context of music; (ii) to enable the model to learn music characteristics,we annotate and release rhythmic and music sources in the current musicdatasets; (iii) for time-aware audio-visual modeling, we align the model'smusic predictions with the temporal dimension. Our experiments showstate-of-the-art effects on the Music AVQA datasets. Our code is available athttps://github.com/xid32/Amuse.
标题:通过多重损失训练和人工数据集自动识别嘻哈音乐中的样本
链接:https://arxiv.org/abs/2502.06364
备注:17 pages, 6 figures
摘要:采样,即在新作品中重复使用录制的音乐或来自其他来源的声音的做法,在嘻哈和说唱等流行音乐流派中很常见。已经出现了许多服务,允许用户识别样本和包含它们的歌曲之间的联系,目的是增强音乐发现。设计一个可以自动执行相同任务的系统是具有挑战性的,因为样本通常会被音调和时间拉伸等音频效果改变,并且可能只有几秒钟长。这项任务的进展甚微,培训数据有限进一步阻碍了进展。在这里,我们展示了在人工数据集上训练的卷积神经网络可以识别商业嘻哈音乐中的真实样本。我们使用音频源分离从多个非商业音乐录音数据库中提取声乐、和声和打击乐元素,并训练模型以在原始音频的转换版本中对这些元素的子集进行指纹识别。我们使用联合分类和度量学习损失来优化模型,并表明它在实际采样实例中的精度比使用声学地标的指纹识别系统高出13%,并且它可以识别音调偏移和时间拉伸的样本。我们还表明,对于我们测试的一半商业音乐录音,我们的模型能够在5秒内定位样本的位置。
摘要:Sampling, the practice of reusing recorded music or sounds from anothersource in a new work, is common in popular music genres like hip-hop and rap.Numerous services have emerged that allow users to identify connections betweensamples and the songs that incorporate them, with the goal of enhancing musicdiscovery. Designing a system that can perform the same task automatically ischallenging, as samples are commonly altered with audio effects like pitch- andtime-stretching and may only be seconds long. Progress on this task has beenminimal and is further blocked by the limited availability of training data.Here, we show that a convolutional neural network trained on an artificialdataset can identify real-world samples in commercial hip-hop music. We extractvocal, harmonic, and percussive elements from several databases ofnon-commercial music recordings using audio source separation, and train themodel to fingerprint a subset of these elements in transformed versions of theoriginal audio. We optimize the model using a joint classification and metriclearning loss and show that it achieves 13% greater precision on real-worldinstances of sampling than a fingerprinting system using acoustic landmarks,and that it can recognize samples that have been both pitch shifted and timestretched. We also show that, for half of the commercial music recordings wetested, our model is capable of locating the position of a sample to withinfive seconds.
标题:基于自适应过滤器组的神经网络延迟估计和语音增强方法
链接:https://arxiv.org/abs/2502.06098
备注:audio 3A
摘要:时延估计在自适应滤波器声回波消除中起着关键作用。如果估计误差出现,将留下相当大的残余回波。在这里,在本文中,我们提出了一种自适应滤波器组基于神经网络的方法,其中的延迟是由一个银行的自适应滤波器与重叠的时间范围估计,和所有的能量滤波器的权重连接和饲料的分类网络。选择具有最大概率的指标作为估计时延。在此基础上,设计了一种基于神经网络的剩余回波和噪声抑制AEC方案,并采用最优修正对数谱幅度(OMLSA)算法来提高其鲁棒性。此外,一个强大的自动增益控制(AGC)方案与频谱平滑方法的设计,以放大语音段。性能评估表明,我们的计划可以实现更高的性能。
摘要:Time delay estimation (TDE) plays a key role in acoustic echo cancellation(AEC) using adaptive filter method. Considerable residual echo will be left ifestimation error arises. Here, in this paper, we proposed an adaptive filterbank based neural network approach where the delay is estimated by a bank ofadaptive filters with overlapped time scope, and all the energy of filterweights are concatenated and feed to a classification network. The index withmaximal probability is chosen as the estimated delay. Based on this TDE, an AECscheme is designed using a neural network for residual echo and noisesuppression, and the optimally-modified log-spectral amplitude (OMLSA)algorithm is adopted to make it robust. Also, a robust automatic gain control(AGC) scheme with spectrum smoothing method is designed to amplify speechsegments. Performance evaluations reveal that higher performance can beachieved for our scheme.
标题:时间工作记忆:查询引导的片段细化以增强多模式理解
链接:https://arxiv.org/abs/2502.06020
备注:Accepted at NAACL 2025
摘要:多模态基础模型(MFM)在视觉字幕、问答和图像-文本检索等任务中取得了显著的成功。然而,这些模型由于其有限的内部容量而面临固有的限制,这限制了它们处理扩展时间序列的能力,这是全面视频和音频分析的关键要求。为了克服这些挑战,我们引入了一个专门的认知模块,时间工作记忆(TWM),其目的是提高时间建模能力的MFM。它选择性地保留跨时间维度的任务相关信息,确保在视频和音频内容的处理过程中保留关键细节。TWM使用查询引导的注意力的方法,专注于时间序列内最丰富的多模态段。通过只保留最相关的内容,TWM优化了模型有限容量的使用,增强了其时间建模能力。该即插即用模块可以轻松集成到现有的MFM中。通过我们的TWM,九种最先进的模型在视频字幕、问答和视频文本检索等任务中表现出显着的性能改进。通过增强时态建模,TWM扩展了MFM的能力,使其能够有效地处理复杂的、对时间敏感的数据。我们的代码可在https://github.com/xid32/NAACL_2025_TWM上获得。
摘要:Multimodal foundation models (MFMs) have demonstrated significant success intasks such as visual captioning, question answering, and image-text retrieval.However, these models face inherent limitations due to their finite internalcapacity, which restricts their ability to process extended temporal sequences,a crucial requirement for comprehensive video and audio analysis. To overcomethese challenges, we introduce a specialized cognitive module, temporal workingmemory (TWM), which aims to enhance the temporal modeling capabilities of MFMs.It selectively retains task-relevant information across temporal dimensions,ensuring that critical details are preserved throughout the processing of videoand audio content. The TWM uses a query-guided attention approach to focus onthe most informative multimodal segments within temporal sequences. Byretaining only the most relevant content, TWM optimizes the use of the model'slimited capacity, enhancing its temporal modeling ability. This plug-and-playmodule can be easily integrated into existing MFMs. With our TWM, ninestate-of-the-art models exhibit significant performance improvements acrosstasks such as video captioning, question answering, and video-text retrieval.By enhancing temporal modeling, TWM extends the capability of MFMs to handlecomplex, time-sensitive data effectively. Our code is available athttps://github.com/xid32/NAACL_2025_TWM.
标题:基于大语言模型的非负矩阵分解用于呼吸声分离
链接:https://arxiv.org/abs/2502.05757
摘要:这项研究代表了大型语言模型(LLM)与非负矩阵分解(NMF)的首次集成,标志着源分离领域的新进展。LLM以两种独特的方式使用:通过为疾病预测提供详细的见解来增强分离结果,并在反馈回路中操作以优化添加到NMF成本函数的基频惩罚。我们在两个数据集上测试了该算法:100个真实测量的合成混合物,以及210个来自临床人体模型的心肺声音记录,包括使用数字听诊器捕获的单独和混合声音。该方法始终优于现有方法,证明了其显着增强疾病诊断医学声音分析的潜力。
摘要:This study represents the first integration of large language models (LLMs)with non-negative matrix factorization (NMF), marking a novel advancement inthe source separation field. The LLM is employed in two unique ways: enhancingthe separation results by providing detailed insights for disease predictionand operating in a feedback loop to optimize a fundamental frequency penaltyadded to the NMF cost function. We tested the algorithm on two datasets: 100synthesized mixtures of real measurements, and 210 recordings of heart and lungsounds from a clinical manikin including both individual and mixed sounds,captured using a digital stethoscope. The approach consistently outperformedexisting methods, demonstrating its potential to significantly enhance medicalsound analysis for disease diagnostics.
标题:教学引导语音合成模型中的性别偏见
链接:https://arxiv.org/abs/2502.05649
备注:NAACL 2025 Findings
摘要:可控表达语音合成的最新进展,特别是在文本到语音(TTS)模型,允许生成语音与特定风格的文本描述,称为风格提示引导。虽然这种发展增强了合成语音的灵活性和自然性,但在理解这些模型如何处理模糊或抽象风格提示方面仍然存在重大差距。这项研究调查了模特如何解释职业相关提示的潜在性别偏见,特别是检查他们对“像护士一样行动”的指示的反应。我们探讨这些模型是否表现出放大性别刻板印象的倾向时,解释这样的提示。我们的实验结果揭示了该模型的倾向,表现出某些职业的性别偏见。此外,不同规模的模型在这些职业中显示出不同程度的这种偏见。
摘要:Recent advancements in controllable expressive speech synthesis, especiallyin text-to-speech (TTS) models, have allowed for the generation of speech withspecific styles guided by textual descriptions, known as style prompts. Whilethis development enhances the flexibility and naturalness of synthesizedspeech, there remains a significant gap in understanding how these modelshandle vague or abstract style prompts. This study investigates the potentialgender bias in how models interpret occupation-related prompts, specificallyexamining their responses to instructions like "Act like a nurse". We explorewhether these models exhibit tendencies to amplify gender stereotypes wheninterpreting such prompts. Our experimental results reveal the model's tendencyto exhibit gender bias for certain occupations. Moreover, models of differentsizes show varying degrees of this bias across these occupations.
标题:IndexTTC:工业级可控且高效的Zero-Shot文本转语音系统
链接:https://arxiv.org/abs/2502.05512
摘要:近年来,基于大语言模型(LLM)的文语转换(TTS)系统以其高自然度和强大的zero-shot语音克隆能力逐渐成为业界的主流,本文介绍了基于XTTS和Toronto模型的IndexTTS系统。我们添加了一些新的改进。具体来说,在中文场景中,我们采用了一种混合建模的方法,结合字符和字符,使多音字和长尾字符的发音可控。我们还进行了比较分析的矢量量化(VQ)与标量量化(FSQ)的声学语音令牌的码本利用率。为了进一步提高语音克隆的效果和稳定性,我们引入了基于一致性的语音条件编码器,并将语音解码器替换为BigVGAN 2。与XTTS相比,在自然度、内容一致性、zero-shot语音克隆等方面都有了显著提升。对于开源中流行的TTS系统,如Fish-Speech、CosyVoice 2、FireRedTTS和F5-TTS,IndexTTS具有相对简单的训练过程,更可控的使用,更快的推理速度。此外,它的性能超过了这些系统。我们的演示可在https://index-tts.github.io上获得。
摘要:Recently, large language model (LLM) based text-to-speech (TTS) systems havegradually become the mainstream in the industry due to their high naturalnessand powerful zero-shot voice cloning capabilities.Here, we introduce theIndexTTS system, which is mainly based on the XTTS and Tortoise model. We addsome novel improvements. Specifically, in Chinese scenarios, we adopt a hybridmodeling method that combines characters and pinyin, making the pronunciationsof polyphonic characters and long-tail characters controllable. We alsoperformed a comparative analysis of the Vector Quantization (VQ) withFinite-Scalar Quantization (FSQ) for codebook utilization of acoustic speechtokens. To further enhance the effect and stability of voice cloning, weintroduce a conformer-based speech conditional encoder and replace thespeechcode decoder with BigVGAN2. Compared with XTTS, it has achievedsignificant improvements in naturalness, content consistency, and zero-shotvoice cloning. As for the popular TTS systems in the open-source, such asFish-Speech, CosyVoice2, FireRedTTS and F5-TTS, IndexTTS has a relativelysimple training process, more controllable usage, and faster inference speed.Moreover, its performance surpasses that of these systems. Our demos areavailable at https://index-tts.github.io.
标题:利用离散音调条件流匹配模型增强表达性语音转换
链接:https://arxiv.org/abs/2502.05471
备注:Accepted by ICASSP 2025
摘要:本文介绍了PFlow-VC,这是一种条件流匹配语音转换模型,它利用细粒度离散音调令牌和目标说话人提示信息进行表达性语音转换(VC)。以前的VC作品主要集中在说话人转换,需要进一步探索如何增强音色转换的表现力(如韵律和情感)。与以前的方法不同,我们采用了一种简单而有效的方法来提高语音转换模型的风格表现力。具体来说,我们预训练了一个自监督的音高VQVAE模型,以离散化与说话人无关的音高信息,并利用一个掩蔽的音高条件流匹配模型进行梅尔频谱图合成,该模型为说话人转换模型提供了上下文音高建模能力,有效地提高了语音风格传输能力。此外,我们通过将全局音色嵌入与时变音色令牌相结合来提高音色相似性。在看不见的LibriTTS测试干净和情感语音数据集ESD上的实验表明,PFlow-VC模型在音色转换和风格转换方面都具有优越性。音频样本可在演示页面https://speechai-demo.github.io/PFlow-VC/上获得。
摘要:This paper introduces PFlow-VC, a conditional flow matching voice conversionmodel that leverages fine-grained discrete pitch tokens and target speakerprompt information for expressive voice conversion (VC). Previous VC worksprimarily focus on speaker conversion, with further exploration needed inenhancing expressiveness (such as prosody and emotion) for timbre conversion.Unlike previous methods, we adopt a simple and efficient approach to enhancethe style expressiveness of voice conversion models. Specifically, we pretraina self-supervised pitch VQVAE model to discretize speaker-irrelevant pitchinformation and leverage a masked pitch-conditioned flow matching model forMel-spectrogram synthesis, which provides in-context pitch modelingcapabilities for the speaker conversion model, effectively improving the voicestyle transfer capacity. Additionally, we improve timbre similarity bycombining global timbre embeddings with time-varying timbre tokens. Experimentson unseen LibriTTS test-clean and emotional speech dataset ESD show thesuperiority of the PFlow-VC model in both timbre conversion and style transfer.Audio samples are available on the demo pagehttps://speechai-demo.github.io/PFlow-VC/.
标题:Koel-TTC:通过偏好对齐和分类器自由指导增强基于LLM的语音生成
链接:https://arxiv.org/abs/2502.05236
摘要:虽然自回归语音令牌生成模型产生具有显著多样性和自然性的语音,但其固有的可控性缺乏通常会导致诸如不符合条件输入的幻觉和不期望的发声等问题。我们介绍Koel-TTS,一套增强的编码器-解码器Transformer TTS模型,通过将自动语音识别和说话人验证模型指导的偏好对齐技术,解决这些挑战。此外,我们将无分类器的指导,以进一步提高合成坚持的成绩单和参考扬声器音频。我们的实验表明,这些优化显着提高目标说话人的相似性,可懂度和合成语音的自然度。值得注意的是,Koel-TTS直接将文本和上下文音频映射到声学标记,并且在上述指标上,尽管在一个非常小的数据集上进行了训练,但其性能优于最先进的TTS模型。音频样本和演示可以在我们的网站上找到。
摘要:While autoregressive speech token generation models produce speech withremarkable variety and naturalness, their inherent lack of controllabilityoften results in issues such as hallucinations and undesired vocalizations thatdo not conform to conditioning inputs. We introduce Koel-TTS, a suite ofenhanced encoder-decoder Transformer TTS models that address these challengesby incorporating preference alignment techniques guided by automatic speechrecognition and speaker verification models. Additionally, we incorporateclassifier-free guidance to further improve synthesis adherence to thetranscript and reference speaker audio. Our experiments demonstrate that theseoptimizations significantly enhance target speaker similarity, intelligibility,and naturalness of synthesized speech. Notably, Koel-TTS directly maps text andcontext audio to acoustic tokens, and on the aforementioned metrics,outperforms state-of-the-art TTS models, despite being trained on asignificantly smaller dataset. Audio samples and demos are available on ourwebsite.
标题:校准者编码器:自我注意Transformer可以成为自我转换器
链接:https://arxiv.org/abs/2502.05232
摘要:现代的自动语音识别系统,包括RNN传感器和基于注意力的编码器-解码器(AED),其设计使得编码器不需要改变从音频序列到嵌入的信息的时间位置;在解码期间处理与最终文本输出的对齐。我们发现,近年来采用的基于变换器的编码器实际上能够在解码之前的前向传递期间进行内部对齐。这一新现象使一个更简单和更有效的模型,“对齐器编码器”。为了训练它,我们放弃了RNN-T的动态规划,以支持AED的逐帧交叉熵损失,而解码器采用了RNN-T的更轻的纯文本递归,而没有学习交叉注意-它只是从开始按顺序扫描嵌入帧,每个生成一个令牌,直到预测消息结束。我们进行的实验表明,性能非常接近最先进的,包括一个特殊的推理配置,使长形式的识别。在一个代表性的比较中,我们测量了我们模型的总推理时间比RNN-T快2倍,比AED快16倍。最后,我们发现音频-文本对齐在某一层的自我注意力权重中清晰可见,可以说是进行了“自我转换”。
摘要:Modern systems for automatic speech recognition, including the RNN-Transducerand Attention-based Encoder-Decoder (AED), are designed so that the encoder isnot required to alter the time-position of information from the audio sequenceinto the embedding; alignment to the final text output is processed duringdecoding. We discover that the transformer-based encoder adopted in recentyears is actually capable of performing the alignment internally during theforward pass, prior to decoding. This new phenomenon enables a simpler and moreefficient model, the "Aligner-Encoder". To train it, we discard the dynamicprogramming of RNN-T in favor of the frame-wise cross-entropy loss of AED,while the decoder employs the lighter text-only recurrence of RNN-T withoutlearned cross-attention -- it simply scans embedding frames in order from thebeginning, producing one token each until predicting the end-of-message. Weconduct experiments demonstrating performance remarkably close to the state ofthe art, including a special inference configuration enabling long-formrecognition. In a representative comparison, we measure the total inferencetime for our model to be 2x faster than RNN-T and 16x faster than AED. Lastly,we find that the audio-text alignment is clearly visible in the self-attentionweights of a certain layer, which could be said to perform "self-transduction".
