本文经arXiv每日学术速递授权转载
微信公众号:arXiv_Daily
【1】Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
链接:https://arxiv.org/abs/2503.01710
备注:Submitted to ACL 2025
摘要:大语言模型(LLM)的最新进展推动了zero-shot文本到语音(TTS)合成的重大进展。然而,现有的基础模型依赖于多阶段处理或复杂的架构来预测多个码本,限制了效率和集成灵活性。为了克服这些挑战,我们引入了Spark-TTS,这是一种由BiCodec提供支持的新型系统,BiCodec是一种单流语音编解码器,它将语音分解为两种互补的令牌类型:用于语言内容的低比特率语义令牌和用于扬声器属性的固定长度全局令牌。这种分离的表示,结合Qwen2.5 LLM和一种思想链(CoT)生成方法,可以实现粗粒度控制(例如,性别、说话风格)和细粒度调整(例如,精确的音高值、说话速率)。为了促进可控TTS的研究,我们引入了VoxBox,这是一个精心策划的10万小时数据集,具有全面的属性注释。大量的实验表明,Spark-TTS不仅实现了最先进的zero-shot语音克隆,而且还生成了高度可定制的语音,超越了基于参考的合成的限制。源代码、预训练模型和音频样本可在https://github.com/SparkAudio/Spark-TTS上获得。
摘要:Recent advancements in large language models (LLMs) have driven significantprogress in zero-shot text-to-speech (TTS) synthesis. However, existingfoundation models rely on multi-stage processing or complex architectures forpredicting multiple codebooks, limiting efficiency and integration flexibility.To overcome these challenges, we introduce Spark-TTS, a novel system powered byBiCodec, a single-stream speech codec that decomposes speech into twocomplementary token types: low-bitrate semantic tokens for linguistic contentand fixed-length global tokens for speaker attributes. This disentangledrepresentation, combined with the Qwen2.5 LLM and a chain-of-thought (CoT)generation approach, enables both coarse-grained control (e.g., gender,speaking style) and fine-grained adjustments (e.g., precise pitch values,speaking rate). To facilitate research in controllable TTS, we introduceVoxBox, a meticulously curated 100,000-hour dataset with comprehensiveattribute annotations. Extensive experiments demonstrate that Spark-TTS notonly achieves state-of-the-art zero-shot voice cloning but also generateshighly customizable voices that surpass the limitations of reference-basedsynthesis. Source code, pre-trained models, and audio samples are available athttps://github.com/SparkAudio/Spark-TTS.
标题:FlowDec:基于流的全频段通用音频编解码器,具有高感知质量
链接:https://arxiv.org/abs/2503.01485
备注:Accepted at ICLR 2025
摘要:我们提出了FlowDec,这是一种用于以48 kHz采样的一般音频的神经全频带音频编解码器,它将非对抗性编解码器训练与基于新的条件流匹配方法的随机后置滤波器相结合。与基于分数匹配的先前工作ScoreDec相比,我们从语音推广到一般音频,并从24 kbit/s移动到低至4 kbit/s,同时提高输出质量并将所需的后置滤波器DNN评估从60减少到6,而无需任何微调或蒸馏技术。我们为我们的方法提供了理论见解和几何直观,与ScoreDec以及最近使用流量匹配的另一项工作相比,并对我们提出的组件进行消融研究。我们表明,FlowDec是最近GAN主导的神经编解码器流的一个有竞争力的替代方案,其FAD得分优于已建立的基于GAN的编解码器DAC和听力测试得分,并为音乐中的语音和和声结构产生质量更自然的重建。
摘要:We propose FlowDec, a neural full-band audio codec for general audio sampledat 48 kHz that combines non-adversarial codec training with a stochasticpostfilter based on a novel conditional flow matching method. Compared to theprior work ScoreDec which is based on score matching, we generalize from speechto general audio and move from 24 kbit/s to as low as 4 kbit/s, while improvingoutput quality and reducing the required postfilter DNN evaluations from 60 to6 without any fine-tuning or distillation techniques. We provide theoreticalinsights and geometric intuitions for our approach in comparison to ScoreDec aswell as another recent work that uses flow matching, and conduct ablationstudies on our proposed components. We show that FlowDec is a competitivealternative to the recent GAN-dominated stream of neural codecs, achieving FADscores better than those of the established GAN-based codec DAC and listeningtest scores that are on par, and producing qualitatively more naturalreconstructions for speech and harmonic structures in music.
标题:基于一致起始和偏差解码和持续踏板检测的流钢琴抄写
链接:https://arxiv.org/abs/2503.01362
备注:Accepted to ISMIR 2024
摘要:本文描述了一种流音频到钢琴的转录方法,其目的是顺序地将音乐信号翻译成一系列的音符开始和偏移事件。该任务的序列到序列性质可能需要计算密集型Transformer模型以获得更好的性能,该模型最近已用于离线转录基准,并且可以扩展用于具有因果注意机制的流式转录。我们假设这种朴素方法的性能限制在于解码器。尽管用于起始点检测的时频特征与用于偏移检测的时频特征显著不同,但是训练单个解码器以输出起始点和偏移事件的混合序列,而不保证相同音符的起始点和偏移事件之间的对应性。为了克服这一限制,我们提出了一个流编码器-解码器模型,该模型使用卷积编码器聚合本地声学特征,然后由自回归Transformer解码器检测可变数量的起始事件和另一个解码器检测活动音高的偏移事件,并在每个时间帧验证维持踏板。使用MAESTRO数据集的实验表明,所提出的流式传输方法的性能与最先进的离线方法相当甚至更好,同时显着降低了计算成本。
摘要:This paper describes a streaming audio-to-MIDI piano transcription approachthat aims to sequentially translate a music signal into a sequence of noteonset and offset events. The sequence-to-sequence nature of this task may callfor the computationally-intensive transformer model for better performance,which has recently been used for offline transcription benchmarks and could beextended for streaming transcription with causal attention mechanisms. Weassume that the performance limitation of this naive approach lies in thedecoder. Although time-frequency features useful for onset detection areconsiderably different from those for offset detection, the single decoder istrained to output a mixed sequence of onset and offset events without guaranteeof the correspondence between the onset and offset events of the same note. Toovercome this limitation, we propose a streaming encoder-decoder model thatuses a convolutional encoder aggregating local acoustic features, followed byan autoregressive Transformer decoder detecting a variable number of onsetevents and another decoder detecting the offset events for the active pitcheswith validation of the sustain pedal at each time frame. Experiments using theMAESTRO dataset showed that the proposed streaming method performed comparablywith or even better than the state-of-the-art offline methods whilesignificantly reducing the computational cost.
标题:用于发音障碍语音合成的语音克隆:解决语音语言病理学中的数据稀缺问题
链接:https://arxiv.org/abs/2503.01266
摘要:本研究探讨声音克隆,以产生合成语音复制与构音障碍的个人的独特模式。使用TORGO数据集,我们解决了语音语言病理学中的数据稀缺和隐私挑战。我们的贡献包括证明语音克隆保留构音障碍的语音特征,分析真实和合成数据之间的差异,并讨论诊断,康复和通信的影响。我们使用商业平台克隆了构音障碍者的声音,并控制扬声器,确保性别匹配的合成声音。有执照的语音语言病理学家(SLP)评估构音障碍,说话者性别和综合指标的子集。SLP在所有情况下都正确识别了构音障碍,95%的说话者性别,但将30%的合成样本误分类为真实样本,表明高度真实。我们的研究结果表明,合成语音有效地捕捉了无序的特征,语音克隆已经发展到产生类似于真实语音的高质量数据,即使是受过训练的专业人员。这对医疗保健具有关键影响,合成数据可以缓解数据稀缺性、保护隐私并增强人工智能驱动的诊断。通过创建多样化的高质量语音数据集,语音克隆可以改善可推广的模型,个性化治疗,并推进构音障碍的辅助技术。 我们公开发布我们的合成数据集,以促进进一步的研究和合作,旨在开发强大的模型,改善语音语言病理学的患者结果。
摘要:This study explores voice cloning to generate synthetic speech replicatingthe unique patterns of individuals with dysarthria. Using the TORGO dataset, weaddress data scarcity and privacy challenges in speech-language pathology. Ourcontributions include demonstrating that voice cloning preserves dysarthricspeech characteristics, analyzing differences between real and synthetic data,and discussing implications for diagnostics, rehabilitation, and communication.We cloned voices from dysarthric and control speakers using a commercialplatform, ensuring gender-matched synthetic voices. A licensed speech-languagepathologist (SLP) evaluated a subset for dysarthria, speaker gender, andsynthetic indicators. The SLP correctly identified dysarthria in all cases andspeaker gender in 95% but misclassified 30% of synthetic samples as real,indicating high realism. Our results suggest synthetic speech effectivelycaptures disordered characteristics and that voice cloning has advanced toproduce high-quality data resembling real speech, even to trainedprofessionals. This has critical implications for healthcare, where syntheticdata can mitigate data scarcity, protect privacy, and enhance AI-drivendiagnostics. By enabling the creation of diverse, high-quality speech datasets,voice cloning can improve generalizable models, personalize therapy, andadvance assistive technologies for dysarthria. We publicly release our synthetic dataset to foster further research andcollaboration, aiming to develop robust models that improve patient outcomes inspeech-language pathology.
标题:说话的回合:音频基金会模型在回合转换动力学方面进行基准测试
链接:https://arxiv.org/abs/2503.01174
备注:Accepted at ICLR 2025
摘要:最近的音频基础模型(FM)浪潮可以为会话建模提供新的功能。然而,已经有有限的努力,以评估这些音频FM全面的能力,有自然和互动的对话。为了与最终用户进行有意义的对话,我们希望FM能够流畅地执行一连串的回合,而不会有太多重叠的讲话或长时间的沉默。受此启发,我们问最近提出的音频FM是否可以理解,预测和执行轮换事件?为了回答这个问题,我们提出了一种新的评估协议,可以评估口语对话系统的话轮转换能力,使用监督模型作为判断,已被训练来预测人与人的对话中的话轮转换事件。使用这个协议,我们提出了第一个全面的用户研究,评估现有的口语对话系统的能力,执行轮对事件,并揭示了许多有趣的见解,如他们有时不明白什么时候说话,可以中断太积极,很少反向通道。我们进一步评估了多个开源和专有的音频FM,这些FM可以通过Switchboard精心策划的测试基准通过API访问,以衡量它们理解和预测话轮转换事件的能力,并确定显著的改进空间。我们将开源我们的评估平台,以促进先进的对话式人工智能系统的开发。
摘要:The recent wave of audio foundation models (FMs) could provide newcapabilities for conversational modeling. However, there have been limitedefforts to evaluate these audio FMs comprehensively on their ability to havenatural and interactive conversations. To engage in meaningful conversationwith the end user, we would want the FMs to additionally perform a fluentsuccession of turns without too much overlapping speech or long stretches ofsilence. Inspired by this, we ask whether the recently proposed audio FMs canunderstand, predict, and perform turn-taking events? To answer this, we proposea novel evaluation protocol that can assess spoken dialog system's turn-takingcapabilities using a supervised model as a judge that has been trained topredict turn-taking events in human-human conversations. Using this protocol,we present the first comprehensive user study that evaluates existing spokendialogue systems on their ability to perform turn-taking events and reveal manyinteresting insights, such as they sometimes do not understand when to speakup, can interrupt too aggressively and rarely backchannel. We further evaluatemultiple open-source and proprietary audio FMs accessible through APIs oncarefully curated test benchmarks from Switchboard to measure their ability tounderstand and predict turn-taking events and identify significant room forimprovement. We will open source our evaluation platform to promote thedevelopment of advanced conversational AI systems.
标题:通过定向对抗攻击利用语音翻译系统中的漏洞
链接:https://arxiv.org/abs/2503.00957
摘要:随着语音翻译(ST)系统变得越来越普遍,了解其漏洞对于确保强大和可靠的通信至关重要。然而,深入探讨这一问题的工作有限。本文探讨了通过不可察觉的音频操作损害这些系统的方法。具体来说,我们提出了两种创新的方法:(1)将扰动注入源音频,(2)生成对抗性音乐,旨在指导有针对性的翻译,同时在物理世界中进行更实际的空中攻击。我们的实验表明,精心制作的音频扰动可以误导翻译模型产生有针对性的有害输出,而对抗性音乐则利用音乐的自然不可感知性更隐蔽地实现这一目标。事实证明,这些攻击在多种语言和翻译模型中都是有效的,突出了当前ST架构中的系统漏洞。这项研究的影响超出了直接的安全问题,揭示了神经语音处理系统的可解释性和鲁棒性。我们的研究结果强调了在音频系统领域需要先进的防御机制和更有弹性的架构。更多细节和示例可以在https://adv-st.github.io上找到。
摘要:As speech translation (ST) systems become increasingly prevalent,understanding their vulnerabilities is crucial for ensuring robust and reliablecommunication. However, limited work has explored this issue in depth. Thispaper explores methods of compromising these systems through imperceptibleaudio manipulations. Specifically, we present two innovative approaches: (1)the injection of perturbation into source audio, and (2) the generation ofadversarial music designed to guide targeted translation, while also conductingmore practical over-the-air attacks in the physical world. Our experimentsreveal that carefully crafted audio perturbations can mislead translationmodels to produce targeted, harmful outputs, while adversarial music achievethis goal more covertly, exploiting the natural imperceptibility of music.These attacks prove effective across multiple languages and translation models,highlighting a systemic vulnerability in current ST architectures. Theimplications of this research extend beyond immediate security concerns,shedding light on the interpretability and robustness of neural speechprocessing systems. Our findings underscore the need for advanced defensemechanisms and more resilient architectures in the realm of audio systems. Moredetails and samples can be found at https://adv-st.github.io.
标题:在拥抱可持续发展的同时揭露偏见:评估自动语音识别系统的双重挑战
链接:https://arxiv.org/abs/2503.00907
备注:Interspeech 2024
摘要:在本文中,我们提出了一个偏见和可持续性为重点的调查自动语音识别(ASR)系统,即耳语和大规模多语言语音(MMS),已达到国家的最先进的(SOTA)的性能。尽管它们在受控环境中的表现有所改善,但在了解它们在现实世界中的功效和公平性方面仍然存在重大差距。我们分析ASR偏差w.r.t.性别、口音和年龄组,以及它们对下游任务的影响。此外,我们还研究了ASR系统对环境的影响,仔细研究了大型声学模型对碳排放和能源消耗的影响。我们还提供了对我们的实证分析的见解,为ASR系统中有关偏见和可持续性的主张做出了宝贵的贡献。
摘要:In this paper, we present a bias and sustainability focused investigation ofAutomatic Speech Recognition (ASR) systems, namely Whisper and MassivelyMultilingual Speech (MMS), which have achieved state-of-the-art (SOTA)performances. Despite their improved performance in controlled settings, thereremains a critical gap in understanding their efficacy and equity in real-worldscenarios. We analyze ASR biases w.r.t. gender, accent, and age group, as wellas their effect on downstream tasks. In addition, we examine the environmentalimpact of ASR systems, scrutinizing the use of large acoustic models on carbonemission and energy consumption. We also provide insights into our empiricalanalyses, offering a valuable contribution to the claims surrounding bias andsustainability in ASR systems.
标题:利用无人机螺旋桨裂纹声学数据集检测UAM螺旋桨缺陷声学异常(ADCP)
链接:https://arxiv.org/abs/2503.00790
备注:25 pages
摘要:UAM即将商业化,需要稳定的基于人工智能的维护系统来确保乘客和行人的安全。提出了一种利用无人机螺旋桨声数据集无损检测无人机螺旋桨裂纹的方法。正常的操作声音被记录下来,异常的声音(分类为撕裂和破碎)通过改变麦克风螺旋桨角度和油门功率来区分。我们的新方法集成了FFT和STFT预处理技术,以捕获全局频率模式和局部时频变化,从而提高异常检测性能。所构建的无人机螺旋桨裂纹声数据集(ADCP)展示了螺旋桨裂纹检测的潜力,并为未来的无人机维修应用奠定了基础。
摘要:The imminent commercialization of UAM requires stable, AI-based maintenancesystems to ensure safety for both passengers and pedestrians. This paperpresents a methodology for non-destructively detecting cracks in UAM propellersusing drone propeller sound datasets. Normal operating sounds were recorded,and abnormal sounds (categorized as ripped and broken) were differentiated byvarying the microphone-propeller angle and throttle power. Our novel approachintegrates FFT and STFT preprocessing techniques to capture both globalfrequency patterns and local time-frequency variations, thereby enhancinganomaly detection performance. The constructed Acoustic Dataset for Crack ofDrone Propeller (ADCP) demonstrates the potential for detecting propellercracks and lays the groundwork for future UAM maintenance applications.
标题:PodAgent:播客生成的全面框架
链接:https://arxiv.org/abs/2503.00455
摘要:现有的自动音频生成方法难以有效地生成类似播客的音频节目。关键的挑战在于深入的内容生成,适当和富有表现力的语音制作。本文提出了一个音频节目制作的综合框架PodAgent。PodAgent 1)通过设计一个Host-Guest-Writer多Agent协作系统生成信息丰富的主题讨论内容; 2)构建一个语音池以实现合适的语音-角色匹配; 3)利用LLM增强的语音合成方法生成富有表达力的会话语音。由于缺乏标准化的评估标准播客类音频生成,我们制定了全面的评估准则,以有效地评估模型的性能。实验结果表明,PodAgent的有效性,显着超过直接GPT-4生成的主题讨论对话内容,实现了87.4%的语音匹配的准确性,并通过LLM指导的合成产生更具表现力的语音。演示页面:https://podcast-agent.github.io/demo/。源代码:https://github.com/yujxx/PodAgent。
摘要:Existing Existing automatic audio generation methods struggle to generatepodcast-like audio programs effectively. The key challenges lie in in-depthcontent generation, appropriate and expressive voice production. This paperproposed PodAgent, a comprehensive framework for creating audio programs.PodAgent 1) generates informative topic-discussion content by designing aHost-Guest-Writer multi-agent collaboration system, 2) builds a voice pool forsuitable voice-role matching and 3) utilizes LLM-enhanced speech synthesismethod to generate expressive conversational speech. Given the absence ofstandardized evaluation criteria for podcast-like audio generation, wedeveloped comprehensive assessment guidelines to effectively evaluate themodel's performance. Experimental results demonstrate PodAgent's effectiveness,significantly surpassing direct GPT-4 generation in topic-discussion dialoguecontent, achieving an 87.4% voice-matching accuracy, and producing moreexpressive speech through LLM-guided synthesis. Demo page:https://podcast-agent.github.io/demo/. Source code:https://github.com/yujxx/PodAgent.
标题:多模式音乐学习中的语言模型映射:一个巨大的挑战提案
链接:https://arxiv.org/abs/2503.00427
摘要:我们已经看到了使用深度神经网络的表示学习和语言模型(LM)的显著成功。许多研究旨在通过标记或嵌入级别的对齐和映射来建立不同模态之间的底层连接,但到目前为止,大多数方法都非常需要数据,限制了它们在配对数据不太丰富的音乐等领域的性能。我们认为,嵌入对齐只是在表面水平的多模态对齐。在本文中,我们提出了一个巨大的挑战\textit{语言模型映射}(LMM),即,如何在不同模态的LM跟踪相同的潜在现象的假设下,将一个域的LM中隐含的本质映射到另一个域的LM。我们首先介绍了LMM的基本设置,强调了揭示跨模态对齐的更深层次以及实现更高样本效率学习的目标。然后,我们讨论为什么音乐是一个理想的域进行LMM研究。在此之后,我们将音乐中的LMM与一个更一般和更具挑战性的科学问题\textit{学习根据感官输入和抽象符号采取行动}联系起来,最后,提出了一个高级版本的挑战问题设置。
摘要:We have seen remarkable success in representation learning and languagemodels (LMs) using deep neural networks. Many studies aim to build theunderlying connections among different modalities via the alignment andmappings at the token or embedding level, but so far, most methods are verydata-hungry, limiting their performance in domains such as music where paireddata are less abundant. We argue that the embedding alignment is only at thesurface level of multimodal alignment. In this paper, we propose a grandchallenge of \textit{language model mapping} (LMM), i.e., how to map theessence implied in the LM of one domain to the LM of another domain under theassumption that LMs of different modalities are tracking the same underlyingphenomena. We first introduce a basic setup of LMM, highlighting the goal tounveil a deeper aspect of cross-modal alignment as well as to achieve moresample-efficiency learning. We then discuss why music is an ideal domain inwhich to conduct LMM research. After that, we connect LMM in music with a moregeneral and challenging scientific problem of \textit{learning to take actionsbased on both sensory input and abstract symbols}, and in the end, present anadvanced version of the challenge problem setup.
标题:BGM 2 Pose:使用非静止声音的主动3D人体姿势估计
链接:https://arxiv.org/abs/2503.00389
摘要:我们提出了BGM2Pose,一种使用任意音乐(例如,背景音乐)作为主动感测信号。与现有的方法,显着限制实用性,采用侵入性的啁啾信号在可听范围内,我们的方法利用自然的音乐,导致人类最小的不适。从标准音乐中估计人类姿势存在重大挑战。与专门为测量而设计的声源相比,普通音乐在音量和音高上都有变化。由音乐引起的信号的这些动态变化不可避免地与由人体运动引起的声场的改变混合,使得难以提取用于姿态估计的可靠线索。为了应对这些挑战,BGM2Pose引入了对比姿势提取模块,该模块采用对比学习和硬负采样来消除记录数据中的音乐成分,从而隔离姿势信息。此外,我们提出了一个频率方面的注意力模块,使模型能够专注于微妙的声学变化归因于人类运动的动态计算注意力跨频带。实验表明,我们的方法优于现有的方法,展示了巨大的潜力,为现实世界的应用。我们的数据集和代码将公开提供。
摘要:We propose BGM2Pose, a non-invasive 3D human pose estimation method usingarbitrary music (e.g., background music) as active sensing signals. Unlikeexisting approaches that significantly limit practicality by employingintrusive chirp signals within the audible range, our method utilizes naturalmusic that causes minimal discomfort to humans. Estimating human poses fromstandard music presents significant challenges. In contrast to sound sourcesspecifically designed for measurement, regular music varies in both volume andpitch. These dynamic changes in signals caused by music are inevitably mixedwith alterations in the sound field resulting from human motion, making it hardto extract reliable cues for pose estimation. To address these challenges,BGM2Pose introduces a Contrastive Pose Extraction Module that employscontrastive learning and hard negative sampling to eliminate musical componentsfrom the recorded data, isolating the pose information. Additionally, wepropose a Frequency-wise Attention Module that enables the model to focus onsubtle acoustic variations attributable to human movement by dynamicallycomputing attention across frequency bands. Experiments suggest that our methodoutperforms the existing methods, demonstrating substantial potential forreal-world applications. Our datasets and code will be made publicly available.
标题:合成数据实现上下文感知生物声学声音事件检测
链接:https://arxiv.org/abs/2503.00296
摘要:我们提出了一种训练基础模型的方法,该方法增强了生物声学信号处理领域内的上下文学习能力。我们使用合成生成的训练数据,引入了一个基于域随机化的管道,该管道构建了具有时间强标签的各种声学场景。我们生成了超过8800小时的强标记音频和训练一个查询的例子,transformer为基础的模型来执行Few-Shot生物声学声音事件检测。我们的第二个贡献是13个不同的Few-Shot生物声学任务的公共基准。我们的模型比以前发表的方法高出49%,我们证明这是由于模型设计和数据规模。我们通过API提供我们的训练模型,为生态学家和动物行为学家提供用于生物声学声音事件检测的免训练工具。
摘要:We propose a methodology for training foundation models that enhances theirin-context learning capabilities within the domain of bioacoustic signalprocessing. We use synthetically generated training data, introducing adomain-randomization-based pipeline that constructs diverse acoustic sceneswith temporally strong labels. We generate over 8.8 thousand hours ofstrongly-labeled audio and train a query-by-example, transformer-based model toperform few-shot bioacoustic sound event detection. Our second contribution isa public benchmark of 13 diverse few-shot bioacoustics tasks. Our modeloutperforms previously published methods by 49%, and we demonstrate that thisis due to both model design and data scale. We make our trained model availablevia an API, to provide ecologists and ethologists with a training-free tool forbioacoustic sound event detection.
标题:InspireMusic:集成超分辨率和大型语言模型,实现高保真长形式音乐生成
链接:https://arxiv.org/abs/2503.00084
备注:Work in progress. Correspondence regarding this technical report should be directed to {chong.zhang, yukun.ma}@alibaba-inc.com. Online demo available on this https URL and this https URL
摘要:我们介绍InspireMusic,一个框架集成了超分辨率和大语言模型,用于高保真长格式音乐生成。一个统一的框架生成高保真音乐,歌曲和音频,它结合了一个自回归Transformer与超分辨率流匹配模型。该框架使得能够从文本和音频提示以更高的采样率可控地生成高保真长格式音乐。我们的模型与以前的方法不同,因为我们使用一个包含更丰富语义信息的码本的音频标记器,从而降低了训练成本并提高了效率。这种组合使我们能够实现高质量的音频生成,长格式相干性高达8 $分钟。然后,基于Qwen 2.5的自回归Transformer模型预测音频令牌。接下来,我们采用超分辨率流匹配模型来生成具有从声学编解码器模型学习的细粒度细节的高采样率音频。综合实验表明,InspireMusic-1. 5 B-Long模型在主观和客观评估方面与最近的顶级开源系统(包括MusicGen和Stable Audio 2.0)具有相当的性能。代码和预训练模型发布在https://github.com/FunAudioLLM/InspireMusic上。
摘要:We introduce InspireMusic, a framework integrated super resolution and largelanguage model for high-fidelity long-form music generation. A unifiedframework generates high-fidelity music, songs, and audio, which incorporatesan autoregressive transformer with a super-resolution flow-matching model. Thisframework enables the controllable generation of high-fidelity long-form musicat a higher sampling rate from both text and audio prompts. Our model differsfrom previous approaches, as we utilize an audio tokenizer with one codebookthat contains richer semantic information, thereby reducing training costs andenhancing efficiency. This combination enables us to achieve high-quality audiogeneration with long-form coherence of up to $8$ minutes. Then, anautoregressive transformer model based on Qwen 2.5 predicts audio tokens. Next,we employ a super-resolution flow-matching model to generate high-sampling rateaudio with fine-grained details learned from an acoustic codec model.Comprehensive experiments show that the InspireMusic-1.5B-Long model has acomparable performance to recent top-tier open-source systems, includingMusicGen and Stable Audio 2.0, on subjective and objective evaluations. Thecode and pre-trained models are released athttps://github.com/FunAudioLLM/InspireMusic.
标题:UniWav:实现语音表示学习和生成的统一预训练
链接:https://arxiv.org/abs/2503.00733
备注:ICLR 2025; demo page at this https URL
摘要:预训练和表征学习在现代语音处理中发挥着越来越重要的作用。然而,不同的应用程序一直依赖于不同的基础模型,因为主要的预训练技术要么是为区分性任务或生成性任务设计的。在这项工作中,我们首次尝试建立一个统一的预训练框架,这两种类型的任务在讲话。我们表明,通过适当的预训练设计选择,人们可以共同学习可应用于这两种类型任务的表示编码器和生成音频解码器。我们提出了UniWav,这是一个编码器-解码器框架,旨在统一预训练表示学习和生成任务。在语音识别、文本到语音和语音标记化方面,UniWav实现了与不同的现有基础模型相当的性能,每个模型都是在特定任务上训练的。我们的研究结果表明,可以构建一个通用的语音基础模型来取代不同的基础模型,从而减少预训练的开销和成本。
摘要:Pre-training and representation learning have been playing an increasinglyimportant role in modern speech processing. Nevertheless, differentapplications have been relying on different foundation models, sincepredominant pre-training techniques are either designed for discriminativetasks or generative tasks. In this work, we make the first attempt at buildinga unified pre-training framework for both types of tasks in speech. We showthat with the appropriate design choices for pre-training, one can jointlylearn a representation encoder and generative audio decoder that can be appliedto both types of tasks. We propose UniWav, an encoder-decoder frameworkdesigned to unify pre-training representation learning and generative tasks. Onspeech recognition, text-to-speech, and speech tokenization, UniWav achievescomparable performance to different existing foundation models, each trained ona specific task. Our findings suggest that a single general-purpose foundationmodel for speech can be built to replace different foundation models, reducingthe overhead and cost of pre-training.
标题:LLaSE-G1:激励基于LLaMA的语音增强的概括能力
链接:https://arxiv.org/abs/2503.00493
备注:13 pages, 2 figures, 8 tables
摘要:语言模型(LM)的最新进展已经显示出强大的语义理解和上下文建模能力,这在生成语音增强(SE)中蓬勃发展。然而,许多基于LM的SE方法主要集中在语义信息,往往忽略了声学信息的关键作用,这导致增强后的声学不一致性和有限的泛化在不同的SE任务。在本文中,我们介绍了LLaSE-G1,一个基于LLaMA的语言模型,激励语音增强的泛化能力。LLaSE-G1提供了以下关键贡献:首先,为了减轻声学不一致性,LLaSE-G1采用WavLM的连续表示作为输入,并从X-Codec 2预测语音令牌,最大限度地保留声学。其次,为了提高泛化能力,LLaSE-G1引入了双通道输入和输出,统一了多个SE任务,而不需要特定于任务的ID。第三,LLaSE-G1优于之前的特定任务的判别和生成SE模型,在测试时展示了缩放效应,并为看不见的SE任务提供了新兴功能。此外,我们还发布了我们的代码和模型,以支持该领域的进一步研究。
摘要:Recent advancements in language models (LMs) have demonstrated strongcapabilities in semantic understanding and contextual modeling, which haveflourished in generative speech enhancement (SE). However, many LM-based SEapproaches primarily focus on semantic information, often neglecting thecritical role of acoustic information, which leads to acoustic inconsistencyafter enhancement and limited generalization across diverse SE tasks. In thispaper, we introduce LLaSE-G1, a LLaMA-based language model that incentivizesgeneralization capabilities for speech enhancement. LLaSE-G1 offers thefollowing key contributions: First, to mitigate acoustic inconsistency,LLaSE-G1 employs continuous representations from WavLM as input and predictsspeech tokens from X-Codec2, maximizing acoustic preservation. Second, topromote generalization capability, LLaSE-G1 introduces dual-channel inputs andoutputs, unifying multiple SE tasks without requiring task-specific IDs. Third,LLaSE-G1 outperforms prior task-specific discriminative and generative SEmodels, demonstrating scaling effects at test time and emerging capabilitiesfor unseen SE tasks. Additionally, we release our code and models to supportfurther research in this area.
标题:迪夫节奏:极其快速且极其简单的具有潜在扩散的端到端全长歌曲生成
链接:https://arxiv.org/abs/2503.01183
摘要:音乐生成的最新进展已经引起了极大的关注,但现有的方法面临着严重的局限性。目前的生成模型只能合成人声音轨或伴奏音轨。虽然一些模型可以生成组合的声乐和伴奏,但它们通常依赖于精心设计的多级级联架构和复杂的数据管道,这阻碍了可扩展性。此外,大多数系统仅限于生成较短的音乐片段,而不是完整长度的歌曲。此外,广泛使用的基于语言模型的方法遭受缓慢的推理速度。为了解决这些挑战,我们提出了DiffRhythm,这是第一个基于潜在扩散的歌曲生成模型,能够在短短十秒内合成具有人声和伴奏的完整歌曲,持续时间长达4分45秒,同时保持高音乐性和可懂度。尽管其卓越的功能,DiffRhythm的设计是简单和优雅的:它消除了复杂的数据准备的需要,采用了一个简单的模型结构,并在推理过程中只需要歌词和风格提示。此外,它的非自回归结构确保了快速的推理速度。这种简单性保证了DiffRhythm的可伸缩性。此外,我们发布了完整的训练代码以及大规模数据上的预训练模型,以促进可重复性和进一步研究。
摘要:Recent advancements in music generation have garnered significant attention,yet existing approaches face critical limitations. Some current generativemodels can only synthesize either the vocal track or the accompaniment track.While some models can generate combined vocal and accompaniment, they typicallyrely on meticulously designed multi-stage cascading architectures and intricatedata pipelines, hindering scalability. Additionally, most systems arerestricted to generating short musical segments rather than full-length songs.Furthermore, widely used language model-based methods suffer from slowinference speeds. To address these challenges, we propose DiffRhythm, the firstlatent diffusion-based song generation model capable of synthesizing completesongs with both vocal and accompaniment for durations of up to 4m45s in onlyten seconds, maintaining high musicality and intelligibility. Despite itsremarkable capabilities, DiffRhythm is designed to be simple and elegant: iteliminates the need for complex data preparation, employs a straightforwardmodel structure, and requires only lyrics and a style prompt during inference.Additionally, its non-autoregressive structure ensures fast inference speeds.This simplicity guarantees the scalability of DiffRhythm. Moreover, we releasethe complete training code along with the pre-trained model on large-scale datato promote reproducibility and further research.
标题:UniWav:实现语音表示学习和生成的统一预训练
链接:https://arxiv.org/abs/2503.00733
备注:ICLR 2025; demo page at this https URL
摘要:预训练和表征学习在现代语音处理中发挥着越来越重要的作用。然而,不同的应用程序一直依赖于不同的基础模型,因为主要的预训练技术要么是为区分性任务或生成性任务设计的。在这项工作中,我们首次尝试建立一个统一的预训练框架,这两种类型的任务在讲话。我们表明,通过适当的预训练设计选择,人们可以共同学习可应用于这两种类型任务的表示编码器和生成音频解码器。我们提出了UniWav,这是一个编码器-解码器框架,旨在统一预训练表示学习和生成任务。在语音识别、文本到语音和语音标记化方面,UniWav实现了与不同的现有基础模型相当的性能,每个模型都是在特定任务上训练的。我们的研究结果表明,可以构建一个通用的语音基础模型来取代不同的基础模型,从而减少预训练的开销和成本。
摘要:Pre-training and representation learning have been playing an increasinglyimportant role in modern speech processing. Nevertheless, differentapplications have been relying on different foundation models, sincepredominant pre-training techniques are either designed for discriminativetasks or generative tasks. In this work, we make the first attempt at buildinga unified pre-training framework for both types of tasks in speech. We showthat with the appropriate design choices for pre-training, one can jointlylearn a representation encoder and generative audio decoder that can be appliedto both types of tasks. We propose UniWav, an encoder-decoder frameworkdesigned to unify pre-training representation learning and generative tasks. Onspeech recognition, text-to-speech, and speech tokenization, UniWav achievescomparable performance to different existing foundation models, each trained ona specific task. Our findings suggest that a single general-purpose foundationmodel for speech can be built to replace different foundation models, reducingthe overhead and cost of pre-training.
标题:LLaSE-G1:激励基于LLaMA的语音增强的概括能力
链接:https://arxiv.org/abs/2503.00493
备注:13 pages, 2 figures, 8 tables
摘要:语言模型(LM)的最新进展已经显示出强大的语义理解和上下文建模能力,这在生成语音增强(SE)中蓬勃发展。然而,许多基于LM的SE方法主要集中在语义信息,往往忽略了声学信息的关键作用,这导致增强后的声学不一致性和有限的泛化在不同的SE任务。在本文中,我们介绍了LLaSE-G1,一个基于LLaMA的语言模型,激励语音增强的泛化能力。LLaSE-G1提供了以下关键贡献:首先,为了减轻声学不一致性,LLaSE-G1采用WavLM的连续表示作为输入,并从X-Codec 2预测语音令牌,最大限度地保留声学。其次,为了提高泛化能力,LLaSE-G1引入了双通道输入和输出,统一了多个SE任务,而不需要特定于任务的ID。第三,LLaSE-G1优于之前的特定任务的判别和生成SE模型,在测试时展示了缩放效应,并为看不见的SE任务提供了新兴功能。此外,我们还发布了我们的代码和模型,以支持该领域的进一步研究。
摘要:Recent advancements in language models (LMs) have demonstrated strongcapabilities in semantic understanding and contextual modeling, which haveflourished in generative speech enhancement (SE). However, many LM-based SEapproaches primarily focus on semantic information, often neglecting thecritical role of acoustic information, which leads to acoustic inconsistencyafter enhancement and limited generalization across diverse SE tasks. In thispaper, we introduce LLaSE-G1, a LLaMA-based language model that incentivizesgeneralization capabilities for speech enhancement. LLaSE-G1 offers thefollowing key contributions: First, to mitigate acoustic inconsistency,LLaSE-G1 employs continuous representations from WavLM as input and predictsspeech tokens from X-Codec2, maximizing acoustic preservation. Second, topromote generalization capability, LLaSE-G1 introduces dual-channel inputs andoutputs, unifying multiple SE tasks without requiring task-specific IDs. Third,LLaSE-G1 outperforms prior task-specific discriminative and generative SEmodels, demonstrating scaling effects at test time and emerging capabilitiesfor unseen SE tasks. Additionally, we release our code and models to supportfurther research in this area.
标题:UL-UNAS:通过网络架构搜索进行实时语音增强的超轻量级U-Net
链接:https://arxiv.org/abs/2503.00340
备注:13 pages, 8 figures, submitted to Neural Networks
摘要:轻量级模型对于实时语音增强应用是必不可少的。近年来,有一个日益增长的趋势,越来越紧凑的语音增强模型的发展。在本文中,我们提出了一个超轻量的U-net优化的网络架构搜索(UL-UNAS),这是适合于实现在低占用设备。首先,我们探索了U-Net框架中各种高效卷积块的应用,以识别最有希望的候选者。其次,我们引入了两个提升组件来增强这些卷积块的容量:一个名为仿射PReLU的新激活函数和一个因果时频注意力模块。此外,我们利用神经架构搜索在我们精心设计的搜索空间中发现最佳架构。通过集成上述策略,UL-UNAS不仅在相同或更低的计算复杂度下显著优于最新的超轻量模型,而且与需要更高计算资源的最新基准模型相比,还提供了具有竞争力的性能。
摘要:Lightweight models are essential for real-time speech enhancementapplications. In recent years, there has been a growing trend toward developingincreasingly compact models for speech enhancement. In this paper, we proposean Ultra-Lightweight U-net optimized by Network Architecture Search (UL-UNAS),which is suitable for implementation in low-footprint devices. Firstly, weexplore the application of various efficient convolutional blocks within theU-Net framework to identify the most promising candidates. Secondly, weintroduce two boosting components to enhance the capacity of theseconvolutional blocks: a novel activation function named affine PReLU and acausal time-frequency attention module. Furthermore, we leverage neuralarchitecture search to discover an optimal architecture within our carefullydesigned search space. By integrating the above strategies, UL-UNAS not onlysignificantly outperforms the latest ultra-lightweight models with the same orlower computational complexity, but also delivers competitive performancecompared to recent baseline models that require substantially highercomputational resources.
标题:Spark-TTC:一种高效的基于LLM的文本到语音模型,具有单流去耦合语音令牌
链接:https://arxiv.org/abs/2503.01710
备注:Submitted to ACL 2025
摘要:大语言模型(LLM)的最新进展推动了zero-shot文本到语音(TTS)合成的重大进展。然而,现有的基础模型依赖于多阶段处理或复杂的架构来预测多个码本,限制了效率和集成灵活性。为了克服这些挑战,我们引入了Spark-TTS,这是一种由BiCodec提供支持的新型系统,BiCodec是一种单流语音编解码器,它将语音分解为两种互补的令牌类型:用于语言内容的低比特率语义令牌和用于扬声器属性的固定长度全局令牌。这种分离的表示,结合Qwen2.5 LLM和一种思想链(CoT)生成方法,可以实现粗粒度控制(例如,性别、说话风格)和细粒度调整(例如,精确的音高值、说话速率)。为了促进可控TTS的研究,我们引入了VoxBox,这是一个精心策划的10万小时数据集,具有全面的属性注释。大量的实验表明,Spark-TTS不仅实现了最先进的zero-shot语音克隆,而且还生成了高度可定制的语音,超越了基于参考的合成的限制。源代码、预训练模型和音频样本可在https://github.com/SparkAudio/Spark-TTS上获得。
摘要:Recent advancements in large language models (LLMs) have driven significantprogress in zero-shot text-to-speech (TTS) synthesis. However, existingfoundation models rely on multi-stage processing or complex architectures forpredicting multiple codebooks, limiting efficiency and integration flexibility.To overcome these challenges, we introduce Spark-TTS, a novel system powered byBiCodec, a single-stream speech codec that decomposes speech into twocomplementary token types: low-bitrate semantic tokens for linguistic contentand fixed-length global tokens for speaker attributes. This disentangledrepresentation, combined with the Qwen2.5 LLM and a chain-of-thought (CoT)generation approach, enables both coarse-grained control (e.g., gender,speaking style) and fine-grained adjustments (e.g., precise pitch values,speaking rate). To facilitate research in controllable TTS, we introduceVoxBox, a meticulously curated 100,000-hour dataset with comprehensiveattribute annotations. Extensive experiments demonstrate that Spark-TTS notonly achieves state-of-the-art zero-shot voice cloning but also generateshighly customizable voices that surpass the limitations of reference-basedsynthesis. Source code, pre-trained models, and audio samples are available athttps://github.com/SparkAudio/Spark-TTS.
标题:FlowDec:基于流的全频段通用音频编解码器,具有高感知质量
链接:https://arxiv.org/abs/2503.01485
备注:Accepted at ICLR 2025
摘要:我们提出了FlowDec,这是一种用于以48 kHz采样的一般音频的神经全频带音频编解码器,它将非对抗性编解码器训练与基于新的条件流匹配方法的随机后置滤波器相结合。与基于分数匹配的先前工作ScoreDec相比,我们从语音推广到一般音频,并从24 kbit/s移动到低至4 kbit/s,同时提高输出质量并将所需的后置滤波器DNN评估从60减少到6,而无需任何微调或蒸馏技术。我们为我们的方法提供了理论见解和几何直观,与ScoreDec以及最近使用流量匹配的另一项工作相比,并对我们提出的组件进行消融研究。我们表明,FlowDec是最近GAN主导的神经编解码器流的一个有竞争力的替代方案,其FAD得分优于已建立的基于GAN的编解码器DAC和听力测试得分,并为音乐中的语音和和声结构产生质量更自然的重建。
摘要:We propose FlowDec, a neural full-band audio codec for general audio sampledat 48 kHz that combines non-adversarial codec training with a stochasticpostfilter based on a novel conditional flow matching method. Compared to theprior work ScoreDec which is based on score matching, we generalize from speechto general audio and move from 24 kbit/s to as low as 4 kbit/s, while improvingoutput quality and reducing the required postfilter DNN evaluations from 60 to6 without any fine-tuning or distillation techniques. We provide theoreticalinsights and geometric intuitions for our approach in comparison to ScoreDec aswell as another recent work that uses flow matching, and conduct ablationstudies on our proposed components. We show that FlowDec is a competitivealternative to the recent GAN-dominated stream of neural codecs, achieving FADscores better than those of the established GAN-based codec DAC and listeningtest scores that are on par, and producing qualitatively more naturalreconstructions for speech and harmonic structures in music.
标题:基于一致起始和偏差解码和持续踏板检测的流钢琴抄写
链接:https://arxiv.org/abs/2503.01362
备注:Accepted to ISMIR 2024
摘要:本文描述了一种流音频到钢琴的转录方法,其目的是顺序地将音乐信号翻译成一系列的音符开始和偏移事件。该任务的序列到序列性质可能需要计算密集型Transformer模型以获得更好的性能,该模型最近已用于离线转录基准,并且可以扩展用于具有因果注意机制的流式转录。我们假设这种朴素方法的性能限制在于解码器。尽管用于起始点检测的时频特征与用于偏移检测的时频特征显著不同,但是训练单个解码器以输出起始点和偏移事件的混合序列,而不保证相同音符的起始点和偏移事件之间的对应性。为了克服这一限制,我们提出了一个流编码器-解码器模型,该模型使用卷积编码器聚合本地声学特征,然后由自回归Transformer解码器检测可变数量的起始事件和另一个解码器检测活动音高的偏移事件,并在每个时间帧验证维持踏板。使用MAESTRO数据集的实验表明,所提出的流式传输方法的性能与最先进的离线方法相当甚至更好,同时显着降低了计算成本。
摘要:This paper describes a streaming audio-to-MIDI piano transcription approachthat aims to sequentially translate a music signal into a sequence of noteonset and offset events. The sequence-to-sequence nature of this task may callfor the computationally-intensive transformer model for better performance,which has recently been used for offline transcription benchmarks and could beextended for streaming transcription with causal attention mechanisms. Weassume that the performance limitation of this naive approach lies in thedecoder. Although time-frequency features useful for onset detection areconsiderably different from those for offset detection, the single decoder istrained to output a mixed sequence of onset and offset events without guaranteeof the correspondence between the onset and offset events of the same note. Toovercome this limitation, we propose a streaming encoder-decoder model thatuses a convolutional encoder aggregating local acoustic features, followed byan autoregressive Transformer decoder detecting a variable number of onsetevents and another decoder detecting the offset events for the active pitcheswith validation of the sustain pedal at each time frame. Experiments using theMAESTRO dataset showed that the proposed streaming method performed comparablywith or even better than the state-of-the-art offline methods whilesignificantly reducing the computational cost.
标题:用于发音障碍语音合成的语音克隆:解决语音语言病理学中的数据稀缺问题
链接:https://arxiv.org/abs/2503.01266
摘要:本研究探讨声音克隆,以产生合成语音复制与构音障碍的个人的独特模式。使用TORGO数据集,我们解决了语音语言病理学中的数据稀缺和隐私挑战。我们的贡献包括证明语音克隆保留构音障碍的语音特征,分析真实和合成数据之间的差异,并讨论诊断,康复和通信的影响。我们使用商业平台克隆了构音障碍者的声音,并控制扬声器,确保性别匹配的合成声音。有执照的语音语言病理学家(SLP)评估构音障碍,说话者性别和综合指标的子集。SLP在所有情况下都正确识别了构音障碍,95%的说话者性别,但将30%的合成样本误分类为真实样本,表明高度真实。我们的研究结果表明,合成语音有效地捕捉了无序的特征,语音克隆已经发展到产生类似于真实语音的高质量数据,即使是受过训练的专业人员。这对医疗保健有着至关重要的影响,合成数据可以缓解数据稀缺性,保护隐私,并增强人工智能驱动的诊断。通过创建多样化的高质量语音数据集,语音克隆可以改善可推广的模型,个性化治疗,并推进构音障碍的辅助技术。 我们公开发布我们的合成数据集,以促进进一步的研究和合作,旨在开发强大的模型,改善语音语言病理学的患者结果。
摘要:This study explores voice cloning to generate synthetic speech replicatingthe unique patterns of individuals with dysarthria. Using the TORGO dataset, weaddress data scarcity and privacy challenges in speech-language pathology. Ourcontributions include demonstrating that voice cloning preserves dysarthricspeech characteristics, analyzing differences between real and synthetic data,and discussing implications for diagnostics, rehabilitation, and communication.We cloned voices from dysarthric and control speakers using a commercialplatform, ensuring gender-matched synthetic voices. A licensed speech-languagepathologist (SLP) evaluated a subset for dysarthria, speaker gender, andsynthetic indicators. The SLP correctly identified dysarthria in all cases andspeaker gender in 95% but misclassified 30% of synthetic samples as real,indicating high realism. Our results suggest synthetic speech effectivelycaptures disordered characteristics and that voice cloning has advanced toproduce high-quality data resembling real speech, even to trainedprofessionals. This has critical implications for healthcare, where syntheticdata can mitigate data scarcity, protect privacy, and enhance AI-drivendiagnostics. By enabling the creation of diverse, high-quality speech datasets,voice cloning can improve generalizable models, personalize therapy, andadvance assistive technologies for dysarthria. We publicly release our synthetic dataset to foster further research andcollaboration, aiming to develop robust models that improve patient outcomes inspeech-language pathology.
标题:说话的回合:音频基金会模型在回合转换动力学方面进行基准测试
链接:https://arxiv.org/abs/2503.01174
备注:Accepted at ICLR 2025
摘要:最近的音频基础模型(FM)浪潮可以为会话建模提供新的功能。然而,已经有有限的努力,以评估这些音频FM全面的能力,有自然和互动的对话。为了与最终用户进行有意义的对话,我们希望FM能够流畅地执行一连串的回合,而不会有太多重叠的讲话或长时间的沉默。受此启发,我们问最近提出的音频FM是否可以理解,预测和执行轮换事件?为了回答这个问题,我们提出了一种新的评估协议,可以评估口语对话系统的话轮转换能力,使用监督模型作为判断,已被训练来预测人与人的对话中的话轮转换事件。使用这个协议,我们提出了第一个全面的用户研究,评估现有的口语对话系统的能力,执行轮对事件,并揭示了许多有趣的见解,如他们有时不明白什么时候说话,可以中断太积极,很少反向通道。我们进一步评估了多个开源和专有的音频FM,这些FM可以通过Switchboard精心策划的测试基准通过API访问,以衡量它们理解和预测话轮转换事件的能力,并确定显著的改进空间。我们将开源我们的评估平台,以促进先进的对话式人工智能系统的开发。
摘要:The recent wave of audio foundation models (FMs) could provide newcapabilities for conversational modeling. However, there have been limitedefforts to evaluate these audio FMs comprehensively on their ability to havenatural and interactive conversations. To engage in meaningful conversationwith the end user, we would want the FMs to additionally perform a fluentsuccession of turns without too much overlapping speech or long stretches ofsilence. Inspired by this, we ask whether the recently proposed audio FMs canunderstand, predict, and perform turn-taking events? To answer this, we proposea novel evaluation protocol that can assess spoken dialog system's turn-takingcapabilities using a supervised model as a judge that has been trained topredict turn-taking events in human-human conversations. Using this protocol,we present the first comprehensive user study that evaluates existing spokendialogue systems on their ability to perform turn-taking events and reveal manyinteresting insights, such as they sometimes do not understand when to speakup, can interrupt too aggressively and rarely backchannel. We further evaluatemultiple open-source and proprietary audio FMs accessible through APIs oncarefully curated test benchmarks from Switchboard to measure their ability tounderstand and predict turn-taking events and identify significant room forimprovement. We will open source our evaluation platform to promote thedevelopment of advanced conversational AI systems.
标题:通过定向对抗攻击利用语音翻译系统中的漏洞
链接:https://arxiv.org/abs/2503.00957
摘要:随着语音翻译(ST)系统变得越来越普遍,了解其漏洞对于确保强大和可靠的通信至关重要。然而,深入探讨这一问题的工作有限。本文探讨了通过不可察觉的音频操作损害这些系统的方法。具体来说,我们提出了两种创新的方法:(1)将扰动注入源音频,(2)生成对抗性音乐,旨在指导有针对性的翻译,同时在物理世界中进行更实际的空中攻击。我们的实验表明,精心制作的音频扰动可以误导翻译模型产生有针对性的有害输出,而对抗性音乐则利用音乐的自然不可感知性更隐蔽地实现这一目标。事实证明,这些攻击在多种语言和翻译模型中都是有效的,突出了当前ST架构中的系统漏洞。这项研究的影响超出了直接的安全问题,揭示了神经语音处理系统的可解释性和鲁棒性。我们的研究结果强调了在音频系统领域需要先进的防御机制和更有弹性的架构。更多细节和示例可以在https://adv-st.github.io上找到。
摘要:As speech translation (ST) systems become increasingly prevalent,understanding their vulnerabilities is crucial for ensuring robust and reliablecommunication. However, limited work has explored this issue in depth. Thispaper explores methods of compromising these systems through imperceptibleaudio manipulations. Specifically, we present two innovative approaches: (1)the injection of perturbation into source audio, and (2) the generation ofadversarial music designed to guide targeted translation, while also conductingmore practical over-the-air attacks in the physical world. Our experimentsreveal that carefully crafted audio perturbations can mislead translationmodels to produce targeted, harmful outputs, while adversarial music achievethis goal more covertly, exploiting the natural imperceptibility of music.These attacks prove effective across multiple languages and translation models,highlighting a systemic vulnerability in current ST architectures. Theimplications of this research extend beyond immediate security concerns,shedding light on the interpretability and robustness of neural speechprocessing systems. Our findings underscore the need for advanced defensemechanisms and more resilient architectures in the realm of audio systems. Moredetails and samples can be found at https://adv-st.github.io.
标题:在拥抱可持续发展的同时揭露偏见:评估自动语音识别系统的双重挑战
链接:https://arxiv.org/abs/2503.00907
备注:Interspeech 2024
摘要:在本文中,我们提出了一个偏见和可持续性为重点的调查自动语音识别(ASR)系统,即耳语和大规模多语言语音(MMS),已达到国家的最先进的(SOTA)的性能。尽管它们在受控环境中的表现有所改善,但在了解它们在现实世界中的功效和公平性方面仍然存在重大差距。我们分析ASR偏差w.r.t.性别、口音和年龄组,以及它们对下游任务的影响。此外,我们还研究了ASR系统对环境的影响,仔细研究了大型声学模型对碳排放和能源消耗的影响。我们还提供了对我们的实证分析的见解,为ASR系统中有关偏见和可持续性的主张做出了宝贵的贡献。
摘要:In this paper, we present a bias and sustainability focused investigation ofAutomatic Speech Recognition (ASR) systems, namely Whisper and MassivelyMultilingual Speech (MMS), which have achieved state-of-the-art (SOTA)performances. Despite their improved performance in controlled settings, thereremains a critical gap in understanding their efficacy and equity in real-worldscenarios. We analyze ASR biases w.r.t. gender, accent, and age group, as wellas their effect on downstream tasks. In addition, we examine the environmentalimpact of ASR systems, scrutinizing the use of large acoustic models on carbonemission and energy consumption. We also provide insights into our empiricalanalyses, offering a valuable contribution to the claims surrounding bias andsustainability in ASR systems.
标题:利用无人机螺旋桨裂纹声学数据集检测UAM螺旋桨缺陷声学异常(ADCP)
链接:https://arxiv.org/abs/2503.00790
备注:25 pages
摘要:UAM即将商业化,需要稳定的基于人工智能的维护系统来确保乘客和行人的安全。提出了一种利用无人机螺旋桨声数据集无损检测无人机螺旋桨裂纹的方法。正常的操作声音被记录下来,异常的声音(分类为撕裂和破碎)通过改变麦克风螺旋桨角度和油门功率来区分。我们的新方法集成了FFT和STFT预处理技术,以捕获全局频率模式和局部时频变化,从而提高异常检测性能。所构建的无人机螺旋桨裂纹声数据集(ADCP)展示了螺旋桨裂纹检测的潜力,并为未来的无人机维修应用奠定了基础。
摘要:The imminent commercialization of UAM requires stable, AI-based maintenancesystems to ensure safety for both passengers and pedestrians. This paperpresents a methodology for non-destructively detecting cracks in UAM propellersusing drone propeller sound datasets. Normal operating sounds were recorded,and abnormal sounds (categorized as ripped and broken) were differentiated byvarying the microphone-propeller angle and throttle power. Our novel approachintegrates FFT and STFT preprocessing techniques to capture both globalfrequency patterns and local time-frequency variations, thereby enhancinganomaly detection performance. The constructed Acoustic Dataset for Crack ofDrone Propeller (ADCP) demonstrates the potential for detecting propellercracks and lays the groundwork for future UAM maintenance applications.
标题:PodAgent:播客生成的全面框架
链接:https://arxiv.org/abs/2503.00455
摘要:现有的自动音频生成方法难以有效地生成类似播客的音频节目。关键的挑战在于深入的内容生成,适当和富有表现力的语音制作。本文提出了一个音频节目制作的综合框架PodAgent。PodAgent 1)通过设计一个Host-Guest-Writer多Agent协作系统生成信息丰富的主题讨论内容; 2)构建一个语音池以实现合适的语音-角色匹配; 3)利用LLM增强的语音合成方法生成富有表达力的会话语音。由于缺乏标准化的评估标准播客类音频生成,我们制定了全面的评估准则,以有效地评估模型的性能。实验结果表明,PodAgent的有效性,显着超过直接GPT-4生成的主题讨论对话内容,实现了87.4%的语音匹配的准确性,并通过LLM指导的合成产生更具表现力的语音。演示页面:https://podcast-agent.github.io/demo/。源代码:https://github.com/yujxx/PodAgent。
摘要:Existing Existing automatic audio generation methods struggle to generatepodcast-like audio programs effectively. The key challenges lie in in-depthcontent generation, appropriate and expressive voice production. This paperproposed PodAgent, a comprehensive framework for creating audio programs.PodAgent 1) generates informative topic-discussion content by designing aHost-Guest-Writer multi-agent collaboration system, 2) builds a voice pool forsuitable voice-role matching and 3) utilizes LLM-enhanced speech synthesismethod to generate expressive conversational speech. Given the absence ofstandardized evaluation criteria for podcast-like audio generation, wedeveloped comprehensive assessment guidelines to effectively evaluate themodel's performance. Experimental results demonstrate PodAgent's effectiveness,significantly surpassing direct GPT-4 generation in topic-discussion dialoguecontent, achieving an 87.4% voice-matching accuracy, and producing moreexpressive speech through LLM-guided synthesis. Demo page:https://podcast-agent.github.io/demo/. Source code:https://github.com/yujxx/PodAgent.
标题:BGM 2 Pose:使用非静止声音的主动3D人体姿势估计
链接:https://arxiv.org/abs/2503.00389
摘要:我们提出了BGM2Pose,一种使用任意音乐(例如,背景音乐)作为主动感测信号。与现有的方法,显着限制实用性,采用侵入性的啁啾信号在可听范围内,我们的方法利用自然的音乐,导致人类最小的不适。从标准音乐中估计人类姿势存在重大挑战。与专门为测量而设计的声源相比,普通音乐在音量和音高上都有变化。由音乐引起的信号的这些动态变化不可避免地与由人体运动引起的声场的改变混合,使得难以提取用于姿态估计的可靠线索。为了应对这些挑战,BGM2Pose引入了对比姿势提取模块,该模块采用对比学习和硬负采样来消除记录数据中的音乐成分,从而隔离姿势信息。此外,我们提出了一个频率方面的注意力模块,使模型能够专注于微妙的声学变化归因于人类运动的动态计算注意力跨频带。实验表明,我们的方法优于现有的方法,展示了巨大的潜力,为现实世界的应用。我们的数据集和代码将公开提供。
摘要:We propose BGM2Pose, a non-invasive 3D human pose estimation method usingarbitrary music (e.g., background music) as active sensing signals. Unlikeexisting approaches that significantly limit practicality by employingintrusive chirp signals within the audible range, our method utilizes naturalmusic that causes minimal discomfort to humans. Estimating human poses fromstandard music presents significant challenges. In contrast to sound sourcesspecifically designed for measurement, regular music varies in both volume andpitch. These dynamic changes in signals caused by music are inevitably mixedwith alterations in the sound field resulting from human motion, making it hardto extract reliable cues for pose estimation. To address these challenges,BGM2Pose introduces a Contrastive Pose Extraction Module that employscontrastive learning and hard negative sampling to eliminate musical componentsfrom the recorded data, isolating the pose information. Additionally, wepropose a Frequency-wise Attention Module that enables the model to focus onsubtle acoustic variations attributable to human movement by dynamicallycomputing attention across frequency bands. Experiments suggest that our methodoutperforms the existing methods, demonstrating substantial potential forreal-world applications. Our datasets and code will be made publicly available.
标题:InspireMusic:集成超分辨率和大型语言模型,实现高保真长形式音乐生成
链接:https://arxiv.org/abs/2503.00084
备注:Work in progress. Correspondence regarding this technical report should be directed to {chong.zhang, yukun.ma}@alibaba-inc.com. Online demo available on this https URL and this https URL
摘要:我们介绍InspireMusic,一个框架集成了超分辨率和大语言模型,用于高保真长格式音乐生成。一个统一的框架生成高保真音乐,歌曲和音频,它结合了一个自回归Transformer与超分辨率流匹配模型。该框架使得能够从文本和音频提示以更高的采样率可控地生成高保真长格式音乐。我们的模型与之前的方法不同,因为我们使用具有包含更丰富语义信息的码本的音频标记器,从而降低训练成本并提高效率。这种组合使我们能够实现高质量的音频生成,长格式相干性高达8 $分钟。然后,基于Qwen 2.5的自回归Transformer模型预测音频令牌。接下来,我们采用超分辨率流匹配模型来生成具有从声学编解码器模型学习的细粒度细节的高采样率音频。综合实验表明,InspireMusic-1. 5 B-Long模型在主观和客观评价上与最近的顶级开源系统(包括MusicGen和Stable Audio 2. 0)具有相当的性能。代码和预训练模型发布在https://github.com/FunAudioLLM/InspireMusic上。
摘要:We introduce InspireMusic, a framework integrated super resolution and largelanguage model for high-fidelity long-form music generation. A unifiedframework generates high-fidelity music, songs, and audio, which incorporatesan autoregressive transformer with a super-resolution flow-matching model. Thisframework enables the controllable generation of high-fidelity long-form musicat a higher sampling rate from both text and audio prompts. Our model differsfrom previous approaches, as we utilize an audio tokenizer with one codebookthat contains richer semantic information, thereby reducing training costs andenhancing efficiency. This combination enables us to achieve high-quality audiogeneration with long-form coherence of up to $8$ minutes. Then, anautoregressive transformer model based on Qwen 2.5 predicts audio tokens. Next,we employ a super-resolution flow-matching model to generate high-sampling rateaudio with fine-grained details learned from an acoustic codec model.Comprehensive experiments show that the InspireMusic-1.5B-Long model has acomparable performance to recent top-tier open-source systems, includingMusicGen and Stable Audio 2.0, on subjective and objective evaluations. Thecode and pre-trained models are released athttps://github.com/FunAudioLLM/InspireMusic.
