本文经arXiv每日学术速递授权转载
微信公众号:arXiv_Daily
cs.SD语音
链接:https://arxiv.org/abs/2503.14185
备注:ACL 2021 Findings
摘要:在端到端语音翻译中,从解码器的角度来看,编码器学习的声学表示通常是固定和静态的,这对于处理语音翻译中的跨模态和跨语言挑战是不可取的。在本文中,我们根据解码器的隐藏状态不同的声学状态的好处,并提出了一个自适应语音到文本的翻译模型,能够动态适应解码器中的声学状态。我们连接声学状态和目标词嵌入序列,并将连接的序列馈送到解码器中的后续块中。为了模拟声学状态和目标隐藏状态之间的深层交互,引入语音-文本混合注意子层来代替传统的交叉注意网络。在两个广泛使用的数据集上的实验结果表明,该方法的性能显着优于最先进的神经语音翻译模型。
摘要:In end-to-end speech translation, acoustic representations learned by theencoder are usually fixed and static, from the perspective of the decoder,which is not desirable for dealing with the cross-modal and cross-lingualchallenge in speech translation. In this paper, we show the benefits of varyingacoustic states according to decoder hidden states and propose an adaptivespeech-to-text translation model that is able to dynamically adapt acousticstates in the decoder. We concatenate the acoustic state and target wordembedding sequence and feed the concatenated sequence into subsequent blocks inthe decoder. In order to model the deep interaction between acoustic states andtarget hidden states, a speech-text mixed attention sublayer is introduced toreplace the conventional cross-attention network. Experiment results on twowidely-used datasets show that the proposed method significantly outperformsstate-of-the-art neural speech translation models.
标题:MAG:无需载体量化的多模式对齐自回归共语音手势生成
链接:https://arxiv.org/abs/2503.14040
摘要:本文主要研究了全身协同语音手势的生成。现有的方法通常采用一个自回归模型伴随着矢量量化令牌的手势生成,这导致信息丢失和妥协的真实感所生成的手势。为了解决这个问题,受现实世界人类运动的自然连续性的启发,我们提出了MAG,一种新颖的多模态对齐框架,用于高质量和多样化的协同语音手势合成,而不依赖于离散标记化。具体来说,(1)我们引入了一个运动文本音频对齐的变分自动编码器(MTA-VAE),它利用预先训练的WavCaps的文本和音频嵌入来增强语义和节奏与运动的对齐,最终产生更逼真的手势。(2)在此基础上,我们提出了一个多模态掩蔽自回归模型(MMAG),使自回归建模连续运动嵌入通过扩散,而无需矢量量化。为了进一步确保多模态的一致性,MMAG采用了混合粒度的音频-文本融合块,作为扩散过程的条件。在两个基准数据集上的大量实验表明,MAG在定量和定性方面都达到了最先进的性能,产生了高度逼真和多样化的联合语音手势。
摘要:This work focuses on full-body co-speech gesture generation. Existing methodstypically employ an autoregressive model accompanied by vector-quantized tokensfor gesture generation, which results in information loss and compromises therealism of the generated gestures. To address this, inspired by the naturalcontinuity of real-world human motion, we propose MAG, a novel multi-modalaligned framework for high-quality and diverse co-speech gesture synthesiswithout relying on discrete tokenization. Specifically, (1) we introduce amotion-text-audio-aligned variational autoencoder (MTA-VAE), which leveragespre-trained WavCaps' text and audio embeddings to enhance both semantic andrhythmic alignment with motion, ultimately producing more realistic gestures.(2) Building on this, we propose a multimodal masked autoregressive model(MMAG) that enables autoregressive modeling in continuous motion embeddingsthrough diffusion without vector quantization. To further ensure multi-modalconsistency, MMAG incorporates a hybrid granularity audio-text fusion block,which serves as conditioning for diffusion process. Extensive experiments ontwo benchmark datasets demonstrate that MAG achieves stateof-the-artperformance both quantitatively and qualitatively, producing highly realisticand diverse co-speech gestures.The code will be released to facilitate futureresearch.
标题:用于水下声学目标识别的神经边缘柱状图描述符
链接:https://arxiv.org/abs/2503.13763
备注:6 pages, 5 figures. This work has been accepted to IEEE OCEANS 2025
摘要:许多海洋应用依赖于使用被动声纳识别声学目标的能力。虽然人们越来越依赖于预训练的模型来执行分类任务,但这些模型通常需要大量的计算资源,并且由于数据集的变化,在转移到新的领域时可能无法实现最佳性能。为了解决这些挑战,这项工作采用最初为图像分类开发的神经边缘直方图描述符(NEHD)方法来对被动声纳信号进行分类。我们对统计和结构纹理特征进行了全面评估,证明它们的组合与大型预训练模型相比具有竞争力的性能。建议NEHD为基础的方法提供了一个轻量级的和有效的解决方案,水下目标识别,显着降低计算成本,同时保持准确性。
摘要:Numerous maritime applications rely on the ability to recognize acoustictargets using passive sonar. While there is a growing reliance on pre-trainedmodels for classification tasks, these models often require extensivecomputational resources and may not perform optimally when transferred to newdomains due to dataset variations. To address these challenges, this workadapts the neural edge histogram descriptors (NEHD) method originally developedfor image classification, to classify passive sonar signals. We conduct acomprehensive evaluation of statistical and structural texture features,demonstrating that their combination achieves competitive performance withlarge pre-trained models. The proposed NEHD-based approach offers a lightweightand efficient solution for underwater target recognition, significantlyreducing computational costs while maintaining accuracy.
标题:MoonCast:高质量Zero-Shot播客一代
链接:https://arxiv.org/abs/2503.14345
摘要:文本到语音合成的最新进展在为单个说话者生成高质量的短话语方面取得了显著的成功。然而,这些系统在将其功能扩展到长时间、多说话者和自发对话时仍然面临挑战,这是播客等真实世界场景的典型特征。这些限制来自两个主要的挑战:1)长演讲:播客通常持续几分钟,超过了大多数现有作品的上限; 2)自发性:播客的特点是自发的,口头的性质,这与正式的,书面的背景形成鲜明对比;现有的作品往往无法捕捉这种自发性。在本文中,我们提出了MoonCast,一种用于高质量zero-shot播客生成的解决方案,旨在从纯文本源(例如,故事、技术报告、TXT、PDF或Web URL格式的新闻)。为了生成长音频,我们采用了基于长上下文语言模型的音频建模方法,利用大规模的长上下文语音数据。为了提高自发性,我们利用播客生成模块生成脚本与自发的细节,这已被经验证明是至关重要的文本到语音建模本身。实验表明,MoonCast优于基线,在自发性和一致性方面有特别显著的改善。
摘要:Recent advances in text-to-speech synthesis have achieved notable success ingenerating high-quality short utterances for individual speakers. However,these systems still face challenges when extending their capabilities to long,multi-speaker, and spontaneous dialogues, typical of real-world scenarios suchas podcasts. These limitations arise from two primary challenges: 1) longspeech: podcasts typically span several minutes, exceeding the upper limit ofmost existing work; 2) spontaneity: podcasts are marked by their spontaneous,oral nature, which sharply contrasts with formal, written contexts; existingworks often fall short in capturing this spontaneity. In this paper, we proposeMoonCast, a solution for high-quality zero-shot podcast generation, aiming tosynthesize natural podcast-style speech from text-only sources (e.g., stories,technical reports, news in TXT, PDF, or Web URL formats) using the voices ofunseen speakers. To generate long audio, we adopt a long-context languagemodel-based audio modeling approach utilizing large-scale long-context speechdata. To enhance spontaneity, we utilize a podcast generation module togenerate scripts with spontaneous details, which have been empirically shown tobe as crucial as the text-to-speech modeling itself. Experiments demonstratethat MoonCast outperforms baselines, with particularly notable improvements inspontaneity and coherence.
标题:用于个性化病理语音增强的变分自动编码器
链接:https://arxiv.org/abs/2503.14036
备注:Submitted to EUSIPCO 2025
摘要:语音增强(SE)模型在说话人条件下的普遍性仍然在很大程度上未被探索,尽管它对更广泛的适用性至关重要。本文研究了混合变分自编码器(VAE)-非负矩阵分解(NMF)模型SE的性能,主要集中在其可推广性与帕金森病的病理扬声器。我们发现,在大型神经典型数据集上训练的VAE模型在病理语音上表现不佳。虽然用病理性语音微调这些预先训练的模型可以提高性能,但神经型和病理性说话者之间仍然存在性能差距。为了解决这一差距,我们建议使用个性化的SE模型,这些模型是从微调预训练模型中获得的,每个扬声器只有几秒钟的干净数据。我们的研究结果表明,个性化的模型大大提高了所有扬声器的性能,实现神经典型和病理扬声器的可比结果。
摘要:The generalizability of speech enhancement (SE) models across speakerconditions remains largely unexplored, despite its critical importance forbroader applicability. This paper investigates the performance of the hybridvariational autoencoder (VAE)-non-negative matrix factorization (NMF) model forSE, focusing primarily on its generalizability to pathological speakers withParkinson's disease. We show that VAE models trained on large neurotypicaldatasets perform poorly on pathological speech. While fine-tuning thesepre-trained models with pathological speech improves performance, a performancegap remains between neurotypical and pathological speakers. To address thisgap, we propose using personalized SE models derived from fine-tuningpre-trained models with only a few seconds of clean data from each speaker. Ourresults demonstrate that personalized models considerably enhance performancefor all speakers, achieving comparable results for both neurotypical andpathological speakers.
标题:MoonCast:高质量Zero-Shot播客一代
链接:https://arxiv.org/abs/2503.14345
摘要:文本到语音合成的最新进展在为单个说话者生成高质量的短话语方面取得了显著的成功。然而,这些系统在将其功能扩展到长时间、多说话者和自发对话时仍然面临挑战,这是播客等真实世界场景的典型特征。这些限制来自两个主要的挑战:1)长演讲:播客通常持续几分钟,超过了大多数现有作品的上限; 2)自发性:播客的特点是自发的,口头的性质,这与正式的,书面的背景形成鲜明对比;现有的作品往往无法捕捉这种自发性。在本文中,我们提出了MoonCast,一种用于高质量zero-shot播客生成的解决方案,旨在从纯文本源(例如,故事、技术报告、TXT、PDF或Web URL格式的新闻)。为了生成长音频,我们采用了基于长上下文语言模型的音频建模方法,利用大规模的长上下文语音数据。为了提高自发性,我们利用播客生成模块生成脚本与自发的细节,这已被经验证明是至关重要的文本到语音建模本身。实验表明,MoonCast优于基线,在自发性和一致性方面有特别显著的改善。
摘要:Recent advances in text-to-speech synthesis have achieved notable success ingenerating high-quality short utterances for individual speakers. However,these systems still face challenges when extending their capabilities to long,multi-speaker, and spontaneous dialogues, typical of real-world scenarios suchas podcasts. These limitations arise from two primary challenges: 1) longspeech: podcasts typically span several minutes, exceeding the upper limit ofmost existing work; 2) spontaneity: podcasts are marked by their spontaneous,oral nature, which sharply contrasts with formal, written contexts; existingworks often fall short in capturing this spontaneity. In this paper, we proposeMoonCast, a solution for high-quality zero-shot podcast generation, aiming tosynthesize natural podcast-style speech from text-only sources (e.g., stories,technical reports, news in TXT, PDF, or Web URL formats) using the voices ofunseen speakers. To generate long audio, we adopt a long-context languagemodel-based audio modeling approach utilizing large-scale long-context speechdata. To enhance spontaneity, we utilize a podcast generation module togenerate scripts with spontaneous details, which have been empirically shown tobe as crucial as the text-to-speech modeling itself. Experiments demonstratethat MoonCast outperforms baselines, with particularly notable improvements inspontaneity and coherence.
标题:通过最佳质量运输重心估计房间脉冲响应
链接:https://arxiv.org/abs/2503.14207
备注:Submitted to EUSCIPCO 2025
摘要:在这项工作中,我们考虑的问题,联合估计一组房间脉冲响应(RIR)对应于紧密间隔的麦克风。在语音增强、噪声消除和可听化等声学应用中,RIR的准确估计是至关重要的。然而,现实世界的约束,如短激励信号,低信噪比,和差的频谱激励,往往使估计问题不适定。在本文中,我们通过最优质量传输(OMT)正则化来解决这些挑战。特别是,我们建议使用OMT重心,或广义平均,作为麦克风之间的信息共享的机制。这使我们能够量化和利用不同麦克风之间的延迟结构的相似性,而不必对房间声学施加严格的假设。由此产生的估计制定的凸优化问题的解决方案,可以使用标准的求解器。在数值例子中,我们证明了所提出的方法在解决其他病态估计的情况下的潜力。
摘要:In this work, we consider the problem of jointly estimating a set of roomimpulse responses (RIRs) corresponding to closely spaced microphones. Theaccurate estimation of RIRs is crucial in acoustic applications such as speechenhancement, noise cancellation, and auralization. However, real-worldconstraints such as short excitation signals, low signal-to-noise ratios, andpoor spectral excitation, often render the estimation problem ill-posed. Inthis paper, we address these challenges by means of optimal mass transport(OMT) regularization. In particular, we propose to use an OMT barycenter, orgeneralized mean, as a mechanism for information sharing between themicrophones. This allows us to quantify and exploit similarities in thedelay-structures between the different microphones without having to imposerigid assumptions on the room acoustics. The resulting estimator is formulatedin terms of the solution to a convex optimization problem which can beimplemented using standard solvers. In numerical examples, we demonstrate thepotential of the proposed method in addressing otherwise ill-conditionedestimation scenarios.
标题:用于个性化病理语音增强的变分自动编码器
链接:https://arxiv.org/abs/2503.14036
备注:Submitted to EUSIPCO 2025
摘要:语音增强(SE)模型在说话人条件下的普遍性仍然在很大程度上未被探索,尽管它对更广泛的适用性至关重要。本文研究了混合变分自编码器(VAE)-非负矩阵分解(NMF)模型SE的性能,主要集中在其可推广性与帕金森病的病理扬声器。我们发现,在大型神经典型数据集上训练的VAE模型在病理语音上表现不佳。虽然用病理性语音微调这些预先训练的模型可以提高性能,但神经型和病理性说话者之间仍然存在性能差距。为了解决这一差距,我们建议使用个性化的SE模型,这些模型是从微调预训练模型中获得的,每个扬声器只有几秒钟的干净数据。我们的研究结果表明,个性化的模型大大提高了所有扬声器的性能,实现神经典型和病理扬声器的可比结果。
摘要:The generalizability of speech enhancement (SE) models across speakerconditions remains largely unexplored, despite its critical importance forbroader applicability. This paper investigates the performance of the hybridvariational autoencoder (VAE)-non-negative matrix factorization (NMF) model forSE, focusing primarily on its generalizability to pathological speakers withParkinson's disease. We show that VAE models trained on large neurotypicaldatasets perform poorly on pathological speech. While fine-tuning thesepre-trained models with pathological speech improves performance, a performancegap remains between neurotypical and pathological speakers. To address thisgap, we propose using personalized SE models derived from fine-tuningpre-trained models with only a few seconds of clean data from each speaker. Ourresults demonstrate that personalized models considerably enhance performancefor all speakers, achieving comparable results for both neurotypical andpathological speakers.
标题:AdaST:在解码器中动态适应编码器状态,用于端到端语音到文本翻译
链接:https://arxiv.org/abs/2503.14185
备注:ACL 2021 Findings
摘要:在端到端语音翻译中,从解码器的角度来看,编码器学习的声学表示通常是固定和静态的,这对于处理语音翻译中的跨模态和跨语言挑战是不可取的。在本文中,我们根据解码器的隐藏状态不同的声学状态的好处,并提出了一个自适应语音到文本的翻译模型,能够动态适应解码器中的声学状态。我们连接声学状态和目标词嵌入序列,并将连接的序列馈送到解码器中的后续块中。为了模拟声学状态和目标隐藏状态之间的深层交互,引入语音-文本混合注意子层来代替传统的交叉注意网络。在两个广泛使用的数据集上的实验结果表明,该方法的性能显着优于最先进的神经语音翻译模型。
摘要:In end-to-end speech translation, acoustic representations learned by theencoder are usually fixed and static, from the perspective of the decoder,which is not desirable for dealing with the cross-modal and cross-lingualchallenge in speech translation. In this paper, we show the benefits of varyingacoustic states according to decoder hidden states and propose an adaptivespeech-to-text translation model that is able to dynamically adapt acousticstates in the decoder. We concatenate the acoustic state and target wordembedding sequence and feed the concatenated sequence into subsequent blocks inthe decoder. In order to model the deep interaction between acoustic states andtarget hidden states, a speech-text mixed attention sublayer is introduced toreplace the conventional cross-attention network. Experiment results on twowidely-used datasets show that the proposed method significantly outperformsstate-of-the-art neural speech translation models.
