今日论文合集:cs.SD语音10篇,eess.AS音频处理10篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily

【1】 Should Audio Front-ends be Adaptive? Comparing Learnable and Adaptive  Front-ends
标题:音频前端应该自适应吗?比较可学习和自适应前端
链接:https://arxiv.org/abs/2502.03260
作者:Qiquan Zhang,  Buddhi Wickramasinghe,  Eliathamby Ambikairajah,  Vidhyasaharan Sethu,  Haizhou Li
备注:Accepted by IEEE TASLP
摘要:手工制作的功能,如梅尔滤波器组,传统上是许多音频处理应用的选择。最近,人们对直接从原始音频波形中提取表示的可学习前端越来越感兴趣。\textcolor{black}{然而,手工制作的滤波器组和当前可学习的前端都导致推理时的计算图固定,无法动态适应变化的声学环境,这是人类听觉系统的一个关键特性。}为此,我们探讨音频前端是否应该是自适应的问题,通过比较Ada-FE前端(最近开发的自适应前端,采用神经自适应反馈控制器动态调整其频谱分解滤波器的Q因子)建立可学习的前端。具体来说,我们系统地研究了两个常用后端骨干和各种音频基准(包括语音、声音事件和音乐)的可学习前端和Ada-FE。综合结果表明,我们的Ada-FE优于先进的可学习前端,更重要的是,它在各种训练时期的测试样本上表现出令人印象深刻的稳定性或鲁棒性。
摘要:Hand-crafted features, such as Mel-filterbanks, have traditionally been thechoice for many audio processing applications. Recently, there has been agrowing interest in learnable front-ends that extract representations directlyfrom the raw audio waveform. \textcolor{black}{However, both hand-craftedfilterbanks and current learnable front-ends lead to fixed computation graphsat inference time, failing to dynamically adapt to varying acousticenvironments, a key feature of human auditory systems.} To this end, we explorethe question of whether audio front-ends should be adaptive by comparing theAda-FE front-end (a recently developed adaptive front-end that employs a neuraladaptive feedback controller to dynamically adjust the Q-factors of itsspectral decomposition filters) to established learnable front-ends.Specifically, we systematically investigate learnable front-ends and Ada-FEacross two commonly used back-end backbones and a wide range of audiobenchmarks including speech, sound event, and music. The comprehensive resultsshow that our Ada-FE outperforms advanced learnable front-ends, and moreimportantly, it exhibits impressive stability or robustness on test samplesover various training epochs.

【2】 Leveraging Broadcast Media Subtitle Transcripts for Automatic Speech  Recognition and Subtitling
标题:利用广播媒体字幕脚本进行自动语音识别和字幕
链接:https://arxiv.org/abs/2502.03212
作者:Jakob Poncelet,  Hugo Van hamme
备注:Preprint
摘要:语音识别技术的最新进展是由大规模数据集和基于注意力的架构驱动的,但仍然存在许多挑战,特别是对于低资源语言和方言。本文探讨了弱监督的电视字幕文本到自动语音识别(ASR)系统的集成,旨在提高逐字翻译和自动生成的字幕。为此,逐字数据和字幕被视为不同的域或语言,由于其不同的特点。我们提出并比较了几个端到端的架构,旨在共同建模与单独或共享的编码器和解码器的方式。所提出的方法能够联合生成逐字转录和字幕。对佛兰芒语(比利时荷兰语)的评价表明,具有级联编码器和单独解码器的模型可以最有效地表示两种数据类型之间的差异,同时在两个域上进行改进。尽管在领域和语言变化的差异,逐字记录与字幕数据相结合,导致显着的ASR的改善,而不需要广泛的预处理。此外,一个大规模的字幕数据集的实验表明,所提出的方法的可扩展性。这些方法不仅提高了ASR的准确性,而且还生成了与标准书面文本密切匹配的字幕,提供了几个潜在的应用。
摘要:The recent advancement of speech recognition technology has been driven bylarge-scale datasets and attention-based architectures, but many challengesstill remain, especially for low-resource languages and dialects. This paperexplores the integration of weakly supervised transcripts from TV subtitlesinto automatic speech recognition (ASR) systems, aiming to improve bothverbatim transcriptions and automatically generated subtitles. To this end,verbatim data and subtitles are regarded as different domains or languages, dueto their distinct characteristics. We propose and compare several end-to-endarchitectures that are designed to jointly model both modalities with separateor shared encoders and decoders. The proposed methods are able to jointlygenerate a verbatim transcription and a subtitle. Evaluation on Flemish(Belgian Dutch) demonstrates that a model with cascaded encoders and separatedecoders allows to represent the differences between the two data types mostefficiently while improving on both domains. Despite differences in domain andlinguistic variations, combining verbatim transcripts with subtitle data leadsto notable ASR improvements without the need for extensive preprocessing.Additionally, experiments with a large-scale subtitle dataset show thescalability of the proposed approach. The methods not only improve ASR accuracybut also generate subtitles that closely match standard written text, offeringseveral potential applications.

【3】 Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech
标题:细粒度偏好优化改进Zero-Shot文本到语音
链接:https://arxiv.org/abs/2502.02950
作者:Jixun Yao,  Yuguang Yang,  Yu Pan,  Yuan Feng,  Ziqian Ning,  Jianhao Ye,  Hongbin Zhou,  Lei Xie
备注:WIP
摘要:结合人类的反馈来调整文本到语音(TTS)系统的输出与人类的喜好已被证明是一种有效的方法,以提高基于语言模型的TTS系统的鲁棒性。目前的方法主要集中在使用偏好数据注释的话语水平。然而,影响收听体验的常见问题往往只出现在音频样本的特定片段中,而其他片段则是精心生成的。在这项研究中,我们提出了一个细粒度的偏好优化方法(FPO),以提高TTS系统的鲁棒性。FPO专注于解决生成的样本中的局部问题,而不是统一优化整个话语。具体来说,我们首先分析生成的样本中的问题类型,将其分为两组,并提出了一种选择性训练损失策略,以优化基于细粒度标签的每个问题类型的偏好。实验结果表明,FPO算法有效地解决了zero-shot TTS系统的局部问题,显著降低了系统的坏点率,提高了系统的可懂度,增强了系统的鲁棒性。此外,FPO与基线系统相比具有更高的数据效率,用更少的训练样本实现了类似的性能。
摘要:Integrating human feedback to align text-to-speech (TTS) system outputs withhuman preferences has proven to be an effective approach for enhancing therobustness of language model-based TTS systems. Current approaches primarilyfocus on using preference data annotated at the utterance level. However,frequent issues that affect the listening experience often only arise inspecific segments of audio samples, while other segments are well-generated. Inthis study, we propose a fine-grained preference optimization approach (FPO) toenhance the robustness of TTS systems. FPO focuses on addressing localizedissues in generated samples rather than uniformly optimizing the entireutterance. Specifically, we first analyze the types of issues in generatedsamples, categorize them into two groups, and propose a selective training lossstrategy to optimize preferences based on fine-grained labels for each issuetype. Experimental results show that FPO enhances the robustness of zero-shotTTS systems by effectively addressing local issues, significantly reducing thebad case ratio, and improving intelligibility. Furthermore, FPO exhibitssuperior data efficiency compared with baseline systems, achieving similarperformance with fewer training samples.

【4】 GenSE: Generative Speech Enhancement via Language Models using  Hierarchical Modeling
标题:GenSE:通过使用分层建模的语言模型进行生成语音增强
链接:https://arxiv.org/abs/2502.02942
作者:Jixun Yao,  Hexin Liu,  Chen Chen,  Yuchen Hu,  EngSiong Chng,  Lei Xie
备注:Accepted by ICLR 2025
摘要:语义信息是指在给定的语言结构中通过词、短语和上下文关系传达的意义。人类可以利用语义信息,如熟悉的语言模式和上下文线索,在嘈杂的环境中重建不完整或掩蔽的语音信号。然而,现有的语音增强(SE)方法往往忽略了丰富的语义信息嵌入在语音中,这是至关重要的,以提高可懂度,说话人的一致性,和增强的语音信号的整体质量。为了丰富SE模型的语义信息,我们采用语言模型作为一个有效的语义学习器,并提出了一个全面的框架,专门为基于语言模型的语音增强,称为\textit{GenSE}。具体来说,我们接近SE作为一个条件语言建模任务,而不是在现有的作品中定义的连续信号回归问题。这是通过使用预先训练的自监督模型将语音信号标记为语义标记,并使用定制设计的单量化器神经编解码器模型将语音信号标记为声学标记来实现的。为了提高语言模型预测的稳定性,我们提出了一种分层建模方法,该方法将清洁语义令牌和清洁声学令牌的生成分为两个不同的阶段。此外,我们引入了一个令牌链提示机制,在声学令牌生成阶段,以确保整个语音增强过程中的音色一致性。在基准数据集上的实验结果表明,我们提出的方法在语音质量和泛化能力方面优于最先进的SE系统。
摘要:Semantic information refers to the meaning conveyed through words, phrases,and contextual relationships within a given linguistic structure. Humans canleverage semantic information, such as familiar linguistic patterns andcontextual cues, to reconstruct incomplete or masked speech signals in noisyenvironments. However, existing speech enhancement (SE) approaches oftenoverlook the rich semantic information embedded in speech, which is crucial forimproving intelligibility, speaker consistency, and overall quality of enhancedspeech signals. To enrich the SE model with semantic information, we employlanguage models as an efficient semantic learner and propose a comprehensiveframework tailored for language model-based speech enhancement, called\textit{GenSE}. Specifically, we approach SE as a conditional language modelingtask rather than a continuous signal regression problem defined in existingworks. This is achieved by tokenizing speech signals into semantic tokens usinga pre-trained self-supervised model and into acoustic tokens using acustom-designed single-quantizer neural codec model. To improve the stabilityof language model predictions, we propose a hierarchical modeling method thatdecouples the generation of clean semantic tokens and clean acoustic tokensinto two distinct stages. Moreover, we introduce a token chain promptingmechanism during the acoustic token generation stage to ensure timbreconsistency throughout the speech enhancement process. Experimental results onbenchmark datasets demonstrate that our proposed approach outperformsstate-of-the-art SE systems in terms of speech quality and generalizationcapability.

【5】 SEAL: Speech Embedding Alignment Learning for Speech Large Language  Model with Retrieval-Augmented Generation
标题:SEAL:具有检索增强生成的语音大语言模型的语音嵌入对齐学习
链接:https://arxiv.org/abs/2502.02603
作者:Chunyu Sun,  Bingyu Liu,  Zhichao Cui,  Anbin Qi,  Tian-hao Zhang,  Dinghao Zhou,  Lewei Lu
摘要:基于嵌入的检索模型在文本和多模态大语言模型(LLM)应用的检索增强生成(RAG)技术方面取得了重大进展。然而,当涉及到语音大规模语言模型(SLLM),这些方法仅限于一个两阶段的过程中,自动语音识别(ASR)结合基于文本的检索。这种顺序架构具有高延迟和错误传播的问题。为了解决这些限制,我们提出了一个统一的嵌入框架,消除了中间文本表示的需要。具体来说,该框架包括单独的语音和文本编码器,然后是一个共享的缩放层,将两种模态映射到一个公共的嵌入空间。与传统的两阶段方法相比,我们的模型将管道延迟降低了50%,同时实现了更高的检索准确性。我们还提供了端到端语音检索固有的挑战的理论分析,并介绍了有效的语音文档匹配的架构原则。大量的实验表明,我们的方法在不同的声学条件和扬声器变化的鲁棒性,铺平了道路,一个新的范式在多模态SLLM检索系统。
摘要:Embedding-based retrieval models have made significant strides inretrieval-augmented generation (RAG) techniques for text and multimodal largelanguage models (LLMs) applications. However, when it comes to speech laragelanguage models (SLLMs), these methods are limited to a two-stage process,where automatic speech recognition (ASR) is combined with text-based retrieval.This sequential architecture suffers from high latency and error propagation.To address these limitations, we propose a unified embedding framework thateliminates the need for intermediate text representations. Specifically, theframework includes separate speech and text encoders, followed by a sharedscaling layer that maps both modalities into a common embedding space. Ourmodel reduces pipeline latency by 50\% while achieving higher retrievalaccuracy compared to traditional two-stage methods. We also provide atheoretical analysis of the challenges inherent in end-to-end speech retrievaland introduce architectural principles for effective speech-to-documentmatching. Extensive experiments demonstrate the robustness of our approachacross diverse acoustic conditions and speaker variations, paving the way for anew paradigm in multimodal SLLMs retrieval systems.

【6】 High-Fidelity Simultaneous Speech-To-Speech Translation
标题:高保真语音同步翻译
链接:https://arxiv.org/abs/2502.03382
作者:Tom Labiausse,  Laurent Mazaré,  Edouard Grave,  Patrick Pérez,  Alexandre Défossez,  Neil Zeghidour
摘要:我们介绍Hibiki,一个解码器只模型的同步语音翻译。Hibiki利用多流语言模型同步处理源和目标语音,并联合生成文本和音频令牌,以执行语音到文本和语音到语音的翻译。我们还解决了同声传译的根本挑战,它不像它的连续对应物,在那里人们等待源话语的结束开始翻译,适应其流程,以积累足够的上下文,以实时产生正确的翻译,逐块。为此,我们引入了一种弱监督方法,该方法利用现成文本翻译系统的复杂性来识别每个单词的最佳延迟并创建对齐的合成数据。经过监督训练后,Hibiki使用香草温度采样执行自适应同步语音翻译。在法语-英语同声翻译任务中,Hibiki在翻译质量、说话者保真度和自然度方面表现出了最先进的性能。此外,其推理过程的简单性使其与批量翻译甚至实时设备部署兼容。我们提供的例子以及模型和推理代码。
摘要:We introduce Hibiki, a decoder-only model for simultaneous speechtranslation. Hibiki leverages a multistream language model to synchronouslyprocess source and target speech, and jointly produces text and audio tokens toperform speech-to-text and speech-to-speech translation. We furthermore addressthe fundamental challenge of simultaneous interpretation, which unlike itsconsecutive counterpart, where one waits for the end of the source utterance tostart translating, adapts its flow to accumulate just enough context to producea correct translation in real-time, chunk by chunk. To do so, we introduce aweakly-supervised method that leverages the perplexity of an off-the-shelf texttranslation system to identify optimal delays on a per-word basis and createaligned synthetic data. After supervised training, Hibiki performs adaptive,simultaneous speech translation with vanilla temperature sampling. On aFrench-English simultaneous speech translation task, Hibiki demonstratesstate-of-the-art performance in translation quality, speaker fidelity andnaturalness. Moreover, the simplicity of its inference process makes itcompatible with batched translation and even real-time on-device deployment. Weprovide examples as well as models and inference code.

【7】 Metis: A Foundation Speech Generation Model with Masked Generative  Pre-training
标题:Metis:具有掩蔽生成预训练的基础语音生成模型
链接:https://arxiv.org/abs/2502.03128
作者:Yuancheng Wang,  Jiachen Zheng,  Junan Zhang,  Xueyao Zhang,  Huan Liao,  Zhizheng Wu
摘要:我们介绍Metis,统一语音生成的基础模型。与以前的特定任务或多任务模型不同,Metis遵循预训练和微调范式。它使用掩码生成模型在大规模未标记语音数据上进行预训练,然后进行微调以适应不同的语音生成任务。具体来说,1)Metis利用两种离散的语音表示:从语音自监督学习(SSL)特征导出的SSL令牌,以及从波形直接量化的声学令牌。2)Metis在SSL令牌上执行掩码生成预训练,利用30万小时的各种语音数据,没有任何附加条件。3)通过对特定任务条件的微调,Metis实现了对各种语音生成任务的有效适应,同时支持多模态输入,即使在使用有限的数据和可训练参数时也是如此。实验表明,Metis可以作为统一语音生成的基础模型:Metis在五个语音生成任务上优于最先进的特定任务或多任务系统,包括zero-shot文本到语音,语音转换,目标说话人提取,语音增强和唇到语音,即使可训练参数少于20M或训练数据少300倍。音频样本可在https://metis-demo.github.io/上获得。
摘要:We introduce Metis, a foundation model for unified speech generation. Unlikeprevious task-specific or multi-task models, Metis follows a pre-training andfine-tuning paradigm. It is pre-trained on large-scale unlabeled speech datausing masked generative modeling and then fine-tuned to adapt to diverse speechgeneration tasks. Specifically, 1) Metis utilizes two discrete speechrepresentations: SSL tokens derived from speech self-supervised learning (SSL)features, and acoustic tokens directly quantized from waveforms. 2) Metisperforms masked generative pre-training on SSL tokens, utilizing 300K hours ofdiverse speech data, without any additional condition. 3) Through fine-tuningwith task-specific conditions, Metis achieves efficient adaptation to variousspeech generation tasks while supporting multimodal input, even when usinglimited data and trainable parameters. Experiments demonstrate that Metis canserve as a foundation model for unified speech generation: Metis outperformsstate-of-the-art task-specific or multi-task systems across five speechgeneration tasks, including zero-shot text-to-speech, voice conversion, targetspeaker extraction, speech enhancement, and lip-to-speech, even with fewer than20M trainable parameters or 300 times less training data. Audio samples are areavailable at https://metis-demo.github.io/.

【8】 AudioMiXR: Spatial Audio Object Manipulation with 6DoF for Sound Design  in Augmented Reality
标题:AudioMiXR:具有6DoF的空间音频对象操纵,用于增强现实中的声音设计
链接:https://arxiv.org/abs/2502.02929
作者:Brandon Woodard,  Margarita Geleta,  Joseph J. LaViola Jr.,  Andrea Fanelli,  Rhonda Wilson
备注:34 pages, 18 Figures
摘要:AudioMiXR是一种增强现实(AR)界面,旨在评估用户如何使用部署在头戴式显示器(Apple Vision Pro)上的六个自由度(6DoF)操纵位于其物理空间中的虚拟音频对象进行3D声音设计。用于3D声音设计的现有工具通常限于桌面显示器,这可能限制执行环境内的混合的空间感知。利用XR HMD创建音景可以为3D声音设计提供实时测试环境,因为现代HMD可以通过跨模态交互提供精确的空间定位。然而,没有针对XR中六自由度(6DoF)声音设计的设计指南的研究。为了提供第一步,以确定设计相关的研究方向,在这个空间,我们进行了一项探索性研究,我们招募了27名参与者,包括专家和非专家的声音设计师。其目标是评估设计经验教训,可用于为未来的3D声音设计研究场所提供信息。我们进行了一项受试者内研究,让用户设计音乐和电影的音景。在对参与者数据进行主题分析后,我们构建了两个设计课程:1。AR声音设计的本体感受,以及2。平衡AR GUI中的视听模式。此外,我们还根据我们的结果提供了可以从6DoF声音设计中受益最多的应用领域。
摘要:We present AudioMiXR, an augmented reality (AR) interface intended to assesshow users manipulate virtual audio objects situated in their physical spaceusing six degrees of freedom (6DoF) deployed on a head-mounted display (AppleVision Pro) for 3D sound design. Existing tools for 3D sound design aretypically constrained to desktop displays, which may limit spatial awareness ofmixing within the execution environment. Utilizing an XR HMD to createsoundscapes may provide a real-time test environment for 3D sound design, asmodern HMDs can provide precise spatial localization assisted by cross-modalinteractions. However, there is no research on design guidelines specific tosound design with six degrees of freedom (6DoF) in XR. To provide a first steptoward identifying design-related research directions in this space, weconducted an exploratory study where we recruited 27 participants, consistingof expert and non-expert sound designers. The goal was to assess design lessonsthat can be used to inform future research venues in 3D sound design. We ran awithin-subjects study where users designed both a music and cinematicsoundscapes. After thematically analyzing participant data, we constructed twodesign lessons: 1. Proprioception for AR Sound Design, and 2. BalancingAudio-Visual Modalities in AR GUIs. Additionally, we provide applicationdomains that can benefit most from 6DoF sound design based on our results.

【9】 Developing multilingual speech synthesis system for Ojibwe, Mi'kmaq, and  Maliseet
标题:为Ojibwe、Mi ' kmaq和Maliseet开发多语言语音合成系统
链接:https://arxiv.org/abs/2502.02703
作者:Shenran Wang,  Changbing Yang,  Mike Parkhill,  Chad Quinn,  Christopher Hammerly,  Jian Zhu
摘要:我们提出了轻量级的流匹配的多语言文本到语音(TTS)系统Ojibwe,Mi'kmaq,和Maliseet,三个土著语言在北美。我们的研究结果表明,在三种类型相似的语言上训练多语言TTS模型可以提高单语模型的性能,特别是在数据稀缺的情况下。无注意力体系结构具有更高的存储效率,与自注意力体系结构相比具有很强的竞争力。我们的研究不仅推进了振兴低资源语言的技术发展,而且还强调了人类评估协议中的文化差距,呼吁采取更加以社区为中心的方法进行人类评估。
摘要:We present lightweight flow matching multilingual text-to-speech (TTS)systems for Ojibwe, Mi'kmaq, and Maliseet, three Indigenous languages in NorthAmerica. Our results show that training a multilingual TTS model on threetypologically similar languages can improve the performance over monolingualmodels, especially when data are scarce. Attention-free architectures arehighly competitive with self-attention architecture with higher memoryefficiency. Our research not only advances technical development for therevitalization of low-resource languages but also highlights the cultural gapin human evaluation protocols, calling for a more community-centered approachto human evaluation.

【10】 Streaming Speaker Change Detection and Gender Classification for  Transducer-Based Multi-Talker Speech Translation
标题:基于传感器的多说话者语音翻译的流说话人变化检测和性别分类
链接:https://arxiv.org/abs/2502.02683
作者:Peidong Wang,  Naoyuki Kanda,  Jian Xue,  Jinyu Li,  Xiaofei Wang,  Aswin Shanmugam Subramanian,  Junkun Chen,  Sunit Sivasankaran,  Xiong Xiao,  Yong Zhao
摘要:流式多说话者语音翻译是一项任务,不仅涉及以低延迟生成准确和流畅的翻译,而且还涉及识别说话者何时发生变化以及说话者的性别。说话者改变信息可以用于为zero-shot文本到语音系统创建音频提示,并且性别可以帮助在传统的文本到语音模型中选择说话者简档。我们建议通过将说话人嵌入到基于传感器的流式端到端语音翻译模型中来解决流式说话人变化检测和性别分类。实验结果表明,本文提出的方法在说话人变化检测和性别分类方面都具有较高的准确率。
摘要:Streaming multi-talker speech translation is a task that involves not onlygenerating accurate and fluent translations with low latency but alsorecognizing when a speaker change occurs and what the speaker's gender is.Speaker change information can be used to create audio prompts for a zero-shottext-to-speech system, and gender can help to select speaker profiles in aconventional text-to-speech model. We propose to tackle streaming speakerchange detection and gender classification by incorporating speaker embeddingsinto a transducer-based streaming end-to-end speech translation model. Ourexperiments demonstrate that the proposed methods can achieve high accuracy forboth speaker change detection and gender classification.

机器翻译由腾讯交互翻译提供,仅供参考