微信公众号:arXiv_Daily
cs.SD语音
标题: ALMTTokenizer:用于音频语言建模的低比特率和语义丰富的音频编解码器Tokenizer
链接:https://arxiv.org/abs/2504.10344
摘要:音频语言模型的最新进展强调了音频标记化的关键作用,它将音频信号转换为离散的标记,从而促进了语言模型架构在音频领域的应用。在这项研究中,我们介绍ALMTokenizer,一种新的低比特率和语义丰富的音频编解码器标记器的音频语言模型。诸如Encodec之类的现有方法通常将各个音频帧编码成离散令牌,而不考虑跨帧使用上下文信息。与这些方法不同的是,我们引入了一种新的基于查询的压缩策略,通过明确建模跨帧的上下文信息,利用一组可学习的查询令牌来捕获整体信息。这种设计不仅使编解码器模型能够捕获更多的语义信息,而且可以用更少的令牌序列对音频信号进行编码。此外,为了增强音频编解码器模型中的语义信息,我们引入了以下内容:(1)掩蔽自编码器(MAE)损失,(2)基于语义先验的矢量量化,以及(3)自回归(AR)预测损失。因此,ALMTokenizer在以较低比特率操作的同时,相对于最先进的方法实现了具有竞争力的重建性能。在相同的音频语言模型框架内,ALMTokenizer在音频理解和生成任务方面优于以前的tokenizer。
摘要:Recent advancements in audio language models have underscored the pivotal role of audio tokenization, which converts audio signals into discrete tokens, thereby facilitating the application of language model architectures to the audio domain. In this study, we introduce ALMTokenizer, a novel low-bitrate and semantically rich audio codec tokenizer for audio language models. Prior methods, such as Encodec, typically encode individual audio frames into discrete tokens without considering the use of context information across frames. Unlike these methods, we introduce a novel query-based compression strategy to capture holistic information with a set of learnable query tokens by explicitly modeling the context information across frames. This design not only enables the codec model to capture more semantic information but also encodes the audio signal with fewer token sequences. Additionally, to enhance the semantic information in audio codec models, we introduce the following: (1) A masked autoencoder (MAE) loss, (2) Vector quantization based on semantic priors, and (3) An autoregressive (AR) prediction loss. As a result, ALMTokenizer achieves competitive reconstruction performance relative to state-of-the-art approaches while operating at a lower bitrate. Within the same audio language model framework, ALMTokenizer outperforms previous tokenizers in audio understanding and generation tasks.
【2】 AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis
标题: AutoStyle-TTS:基于检索增强生成的自动风格匹配文语合成链接:https://arxiv.org/abs/2504.10309
备注:accepted by ICME25
摘要:随着语音合成技术的进步,用户对合成语音的自然度和表现力有了更高的要求。但以往的研究忽视了及时选择的重要性。提出了一种基于检索增强生成(RAG)技术的文语转换框架,该框架能够根据文本内容动态调整语音风格,以达到更加自然生动的交流效果。我们已经建立了一个语音风格知识库,包含高质量的语音样本在各种情况下,并制定了一个风格匹配计划。该方案使用Llama,PER-LLM-Embedder和Moka提取的嵌入,与知识数据库中的样本匹配,选择最合适的语音风格进行合成。此外,我们的实证研究验证了所提出的方法的有效性。我们的演示可以在https://thuhcsi.github.io/icme2025-AutoStyle-TTS上查看
摘要:With the advancement of speech synthesis technology, users have higher expectations for the naturalness and expressiveness of synthesized speech. But previous research ignores the importance of prompt selection. This study proposes a text-to-speech (TTS) framework based on Retrieval-Augmented Generation (RAG) technology, which can dynamically adjust the speech style according to the text content to achieve more natural and vivid communication effects. We have constructed a speech style knowledge database containing high-quality speech samples in various contexts and developed a style matching scheme. This scheme uses embeddings, extracted by Llama, PER-LLM-Embedder,and Moka, to match with samples in the knowledge database, selecting the most appropriate speech style for synthesis. Furthermore, our empirical research validates the effectiveness of the proposed method. Our demo can be viewed at: https://thuhcsi.github.io/icme2025-AutoStyle-TTS
【3】 Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion Synthesis
标题: 分开合作:协调钢琴手部运动合成的双流扩散模型链接:https://arxiv.org/abs/2504.09885
备注:12 pages, 4 figures
摘要:自动化协调双手钢琴演奏的合成带来了巨大的挑战,特别是在捕捉双手之间复杂的编舞,同时保留其独特的运动学特征。在本文中,我们提出了一个双流神经框架,旨在从音频输入中生成钢琴演奏的同步手势,解决了手部独立性和协调性建模的关键挑战。我们的框架引入了两个关键创新:(i)解耦的基于扩散的生成框架,其经由双噪声初始化独立地对每只手的运动进行建模,在利用共享位置条件的同时对每只手的不同潜在噪声进行采样,以及(ii)手协调非对称注意力(HCAA)机制抑制对称性运动。(共模)噪声以突出不对称的手部特定特征,同时在去噪期间自适应地增强手部间协调。该系统分层运行:它首先从音频特征预测3D手部位置,然后通过位置感知扩散模型生成关节角度,其中并行去噪流通过HCAA进行交互。综合评估表明,我们的框架在多个指标上优于现有的最先进的方法。
摘要:Automating the synthesis of coordinated bimanual piano performances poses significant challenges, particularly in capturing the intricate choreography between the hands while preserving their distinct kinematic signatures. In this paper, we propose a dual-stream neural framework designed to generate synchronized hand gestures for piano playing from audio input, addressing the critical challenge of modeling both hand independence and coordination. Our framework introduces two key innovations: (i) a decoupled diffusion-based generation framework that independently models each hand's motion via dual-noise initialization, sampling distinct latent noise for each while leveraging a shared positional condition, and (ii) a Hand-Coordinated Asymmetric Attention (HCAA) mechanism suppresses symmetric (common-mode) noise to highlight asymmetric hand-specific features, while adaptively enhancing inter-hand coordination during denoising. The system operates hierarchically: it first predicts 3D hand positions from audio features and then generates joint angles through position-aware diffusion models, where parallel denoising streams interact via HCAA. Comprehensive evaluations demonstrate that our framework outperforms existing state-of-the-art methods across multiple metrics.
【4】 SafeSpeech: Robust and Universal Voice Protection Against Malicious Speech Synthesis
标题: SafeSpeech:强大而通用的语音保护,防止恶意语音合成链接:https://arxiv.org/abs/2504.09839
备注:Accepted to USENIX Security 2025
摘要:语音合成技术带来了极大的便利,而逼真的deepfake音频的广泛使用引发了危害。恶意攻击者可能会未经授权收集受害者的语音,并克隆类似的声音进行非法利用(\textit{e.g.},电信诈骗)。然而,现有的防御方法不能有效地防止deepfake攻击,并且容易受到强大的训练技术的影响。因此,迫切需要一种更有效、更健壮的数据保护方法。作为回应,我们提出了一个防御框架,\textit{\textbf{SafeSpeech}},它通过在原始语音上嵌入不可感知的扰动来防止高质量的合成语音,从而在上传之前保护用户的音频。在SafeSpeech中,我们设计了一种鲁棒且通用的主动保护技术,\textbf{S}peech \textbf{PE}rturbative \textbf{C} oncephrase(\textbf{SPEC}),该技术利用代理模型为生成合成模型生成普遍适用的扰动。此外,我们优化了人类的感知嵌入扰动的时域和频域。为了全面评估我们的方法,我们在先进的模型和数据集上进行了大量的实验,包括主观和客观的实验。我们的实验结果表明,SafeSpeech实现了最先进的(SOTA)语音保护的有效性和可转移性,并对先进的自适应对手具有很强的鲁棒性。此外,SafeSpeech在真实世界的测试中具有实时能力。源代码可以在\href{https://github.com/wxzyd123/SafeSpeech}{https://github.com/wxzyd123/SafeSpeech}上找到。
摘要:Speech synthesis technology has brought great convenience, while the widespread usage of realistic deepfake audio has triggered hazards. Malicious adversaries may unauthorizedly collect victims' speeches and clone a similar voice for illegal exploitation (\textit{e.g.}, telecom fraud). However, the existing defense methods cannot effectively prevent deepfake exploitation and are vulnerable to robust training techniques. Therefore, a more effective and robust data protection method is urgently needed. In response, we propose a defensive framework, \textit{\textbf{SafeSpeech}}, which protects the users' audio before uploading by embedding imperceptible perturbations on original speeches to prevent high-quality synthetic speech. In SafeSpeech, we devise a robust and universal proactive protection technique, \textbf{S}peech \textbf{PE}rturbative \textbf{C}oncealment (\textbf{SPEC}), that leverages a surrogate model to generate universally applicable perturbation for generative synthetic models. Moreover, we optimize the human perception of embedded perturbation in terms of time and frequency domains. To evaluate our method comprehensively, we conduct extensive experiments across advanced models and datasets, both subjectively and objectively. Our experimental results demonstrate that SafeSpeech achieves state-of-the-art (SOTA) voice protection effectiveness and transferability and is highly robust against advanced adaptive adversaries. Moreover, SafeSpeech has real-time capability in real-world tests. The source code is available at \href{https://github.com/wxzyd123/SafeSpeech}{https://github.com/wxzyd123/SafeSpeech}.
【5】 FSSUAVL: A Discriminative Framework using Vision Models for Federated Self-Supervised Audio and Image Understanding
标题: FSSUAVL:使用视觉模型进行联邦自我监督音频和图像理解的区分性框架链接:https://arxiv.org/abs/2504.09516
备注:8 pages
摘要:最近的研究表明,视觉模型可以有效地学习多模态音频图像表示配对时。然而,使深度模型能够从未配对的模态中学习表示的挑战仍然没有解决。这个问题在像联邦学习(FL)这样的场景中尤其重要,在这些场景中,数据通常是分散的,异构的,并且缺乏配对数据的可靠保证。以前的尝试通过在本地客户端上使用辅助预训练编码器或生成模型来解决这个问题,这总是会随着模态数量的增加而增加计算成本。与这些方法不同,在本文中,我们的目标是使用\texttt{FSSUAVL}来解决未配对音频和图像识别的任务,这是一个使用自监督对比学习(SSL)在FL中预训练的单一深度模型。代替对齐音频和图像模态,\texttt{FSSUAVL}通过使用对比SSL将它们投影到公共嵌入空间中来共同区分它们。这将\texttt{FSSUAVL}的实用程序扩展到配对和未配对的音频和图像识别任务。我们使用CNN和ViT的实验表明,与为每种模态使用单独的深度模型相比,\texttt{FSSUAVL}显著提高了各种基于图像和音频的下游任务的性能。此外,FSSUAVL学习多模态特征表示的能力允许整合辅助信息(如果可用),以提高识别准确性。
摘要:Recent studies have demonstrated that vision models can effectively learn multimodal audio-image representations when paired. However, the challenge of enabling deep models to learn representations from unpaired modalities remains unresolved. This issue is especially pertinent in scenarios like Federated Learning (FL), where data is often decentralized, heterogeneous, and lacks a reliable guarantee of paired data. Previous attempts tackled this issue through the use of auxiliary pretrained encoders or generative models on local clients, which invariably raise computational cost with increasing number modalities. Unlike these approaches, in this paper, we aim to address the task of unpaired audio and image recognition using \texttt{FSSUAVL}, a single deep model pretrained in FL with self-supervised contrastive learning (SSL). Instead of aligning the audio and image modalities, \texttt{FSSUAVL} jointly discriminates them by projecting them into a common embedding space using contrastive SSL. This extends the utility of \texttt{FSSUAVL} to paired and unpaired audio and image recognition tasks. Our experiments with CNN and ViT demonstrate that \texttt{FSSUAVL} significantly improves performance across various image- and audio-based downstream tasks compared to using separate deep models for each modality. Additionally, \texttt{FSSUAVL}'s capacity to learn multimodal feature representations allows for integrating auxiliary information, if available, to enhance recognition accuracy.
【6】 AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis
标题: AMNet:增强型普通话语音合成的声学模型网络链接:https://arxiv.org/abs/2504.09225
备注:Main paper (8 pages). Accepted for publication by IJCNN 2025
摘要:本文介绍了AMNet,一个声学模型网络,旨在提高汉语语音合成的性能,通过结合短语结构注释和本地卷积模块。AMNet建立在FastSpeech 2架构的基础上,同时解决了本地上下文建模的挑战,这对于捕获复杂的语音特征(如停顿、重音和语调)至关重要。通过在模型中嵌入短语结构解析器并引入局部卷积模块,AMNet增强了模型对局部信息的敏感性。此外,AMNet从音素中提取音调特征,为音调建模提供明确的指导,从而提高音调准确性和发音。实验结果表明,AMNet优于基线模型的主观和客观评价。所提出的模型实现了优越的平均意见分数(MOS),较低的梅尔倒谱系数失真(MCD),并改善了基频拟合$F0(R^2)$,证实了其能够生成高质量,自然,和表达的普通话语音。
摘要:This paper presents AMNet, an Acoustic Model Network designed to improve the performance of Mandarin speech synthesis by incorporating phrase structure annotation and local convolution modules. AMNet builds upon the FastSpeech 2 architecture while addressing the challenge of local context modeling, which is crucial for capturing intricate speech features such as pauses, stress, and intonation. By embedding a phrase structure parser into the model and introducing a local convolution module, AMNet enhances the model's sensitivity to local information. Additionally, AMNet decouples tonal characteristics from phonemes, providing explicit guidance for tone modeling, which improves tone accuracy and pronunciation. Experimental results demonstrate that AMNet outperforms baseline models in subjective and objective evaluations. The proposed model achieves superior Mean Opinion Scores (MOS), lower Mel Cepstral Distortion (MCD), and improved fundamental frequency fitting $F0 (R^2)$, confirming its ability to generate high-quality, natural, and expressive Mandarin speech.
【7】 Generation of Musical Timbres using a Text-Guided Diffusion Model
标题: 使用文本引导扩散模型生成音乐音色链接:https://arxiv.org/abs/2504.09219
备注:10 pages, 5 figures
摘要:近年来,文本到音频系统取得了显著的成功,使得能够直接从文本描述生成完整的音频片段。虽然这些系统也促进了音乐创作,但人类创造力和刻意表达的元素往往是有限的。相比之下,目前的工作允许作曲家,作曲家和表演者创建音乐创作的基本构建块:用于电子乐器和数据采集器的单个音符的音频。通过文本提示,用户可以指定音频的音色特征。我们介绍了一个系统,结合了潜在的扩散模型和多模态对比学习生成音乐音色的文本描述的条件。通过联合生成频谱图的幅度和相位,我们的方法消除了随后运行相位恢复算法的需要,因为相关的方法。 音频示例、源代码和Web应用程序可在https://wxuanyuan.github.io/Musical-Note-Generation/上获得
摘要:In recent years, text-to-audio systems have achieved remarkable success, enabling the generation of complete audio segments directly from text descriptions. While these systems also facilitate music creation, the element of human creativity and deliberate expression is often limited. In contrast, the present work allows composers, arrangers, and performers to create the basic building blocks for music creation: audio of individual musical notes for use in electronic instruments and DAWs. Through text prompts, the user can specify the timbre characteristics of the audio. We introduce a system that combines a latent diffusion model and multi-modal contrastive learning to generate musical timbres conditioned on text descriptions. By jointly generating the magnitude and phase of the spectrogram, our method eliminates the need for subsequently running a phase retrieval algorithm, as related methods do. Audio examples, source code, and a web app are available at https://wxuanyuan.github.io/Musical-Note-Generation/
【8】 EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation
标题: EchoMass:用于整体共语音运动生成的语音查询基于注意力的面具建模链接:https://arxiv.org/abs/2504.09209
备注:12 pages, 12 figures
摘要:Masked建模框架已经在联合语音运动生成中显示出了希望。然而,它努力识别语义上重要的帧,以实现有效的运动掩蔽。在这项工作中,我们提出了一个语音查询的注意力为基础的掩模建模框架,共同语音运动生成。我们的关键见解是利用运动对齐的语音特征来指导掩蔽的运动建模过程,选择性地掩蔽节奏相关和语义表达的运动帧。具体来说,我们首先提出了一个运动音频对齐模块(MAM)来构建一个潜在的运动音频联合空间。在这个空间中,低级别和高级别的语音特征都被投影,从而使用可学习的语音查询实现运动对齐的语音表示。然后,语音查询注意力机制(SQA)被引入到计算帧级的注意力分数,通过运动键和语音查询之间的相互作用,引导选择性掩蔽运动帧与高的注意力分数。最后,运动对齐的语音特征也被注入到生成网络中,以促进共同语音运动生成。定性和定量评估证实,我们的方法优于现有的最先进的方法,成功地产生高质量的共同语音运动。
摘要:Masked modeling framework has shown promise in co-speech motion generation. However, it struggles to identify semantically significant frames for effective motion masking. In this work, we propose a speech-queried attention-based mask modeling framework for co-speech motion generation. Our key insight is to leverage motion-aligned speech features to guide the masked motion modeling process, selectively masking rhythm-related and semantically expressive motion frames. Specifically, we first propose a motion-audio alignment module (MAM) to construct a latent motion-audio joint space. In this space, both low-level and high-level speech features are projected, enabling motion-aligned speech representation using learnable speech queries. Then, a speech-queried attention mechanism (SQA) is introduced to compute frame-level attention scores through interactions between motion keys and speech queries, guiding selective masking toward motion frames with high attention scores. Finally, the motion-aligned speech features are also injected into the generation network to facilitate co-speech motion generation. Qualitative and quantitative evaluations confirm that our method outperforms existing state-of-the-art approaches, successfully producing high-quality co-speech motion.
【9】 Spatial Audio Processing with Large Language Model on Wearable Devices
标题: 可穿戴设备上使用大语言模型的空间音频处理链接:https://arxiv.org/abs/2504.08907
摘要:将空间上下文集成到大型语言模型(LLM)中有可能彻底改变人机交互,特别是在可穿戴设备中。在这项工作中,我们提出了一种新的系统架构,将空间语音理解到LLM,使上下文感知和自适应应用可穿戴技术。我们的方法利用基于微结构的空间感测来使用单声道麦克风提取精确的到达方向(DoA)信息。为了解决现有数据集缺乏微结构辅助语音记录的问题,我们通过使用LibriSpeech数据集综合创建了一个名为OmniTalk的数据集。这种空间信息与OpenAI的Whisper模型的语言嵌入相融合,允许每种模态学习互补的上下文表示。融合嵌入与LLaMA-3.2 3B模型的输入空间对齐,并使用轻量级自适应技术LoRA进行微调,以优化设备上的处理。SING支持空间感知的自动语音识别(ASR),实现了25.72美元的平均误差-与现有工作中的88.52美元的中位误差相比有了很大的改善-单词错误率(WER)为5.3。SING还支持声音景观,例如,推断有多少人在说话及其方向,最多5人,DoA错误中位数为16$^\circ$。我们的系统在空间语音理解方面表现出卓越的性能,同时解决了能效、隐私和硬件限制的挑战,为增强现实、可访问性和沉浸式体验方面的高级应用铺平了道路。
摘要:Integrating spatial context into large language models (LLMs) has the potential to revolutionize human-computer interaction, particularly in wearable devices. In this work, we present a novel system architecture that incorporates spatial speech understanding into LLMs, enabling contextually aware and adaptive applications for wearable technologies. Our approach leverages microstructure-based spatial sensing to extract precise Direction of Arrival (DoA) information using a monaural microphone. To address the lack of existing dataset for microstructure-assisted speech recordings, we synthetically create a dataset called OmniTalk by using the LibriSpeech dataset. This spatial information is fused with linguistic embeddings from OpenAI's Whisper model, allowing each modality to learn complementary contextual representations. The fused embeddings are aligned with the input space of LLaMA-3.2 3B model and fine-tuned with lightweight adaptation technique LoRA to optimize for on-device processing. SING supports spatially-aware automatic speech recognition (ASR), achieving a mean error of $25.72^\circ$-a substantial improvement compared to the 88.52$^\circ$ median error in existing work-with a word error rate (WER) of 5.3. SING also supports soundscaping, for example, inference how many people were talking and their directions, with up to 5 people and a median DoA error of 16$^\circ$. Our system demonstrates superior performance in spatial speech understanding while addressing the challenges of power efficiency, privacy, and hardware constraints, paving the way for advanced applications in augmented reality, accessibility, and immersive experiences.
【10】 DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
标题: DiPSE:通过潜在扩散变换器进行高保真生成语音增强链接:https://arxiv.org/abs/2504.09381
备注:Manuscript under review
摘要:真实世界的语音记录遭受诸如背景噪声和混响的劣化。语音增强旨在通过生成清晰的高保真信号来缓解这些问题。虽然最近的语音增强生成方法已经显示出有希望的结果,但它们仍然面临两个主要挑战:(1)内容幻觉,其中生成的看似合理的音素与原始话语不同;以及(2)不一致性,未能从输入语音中保留说话者的身份和非语言特征。在这项工作中,我们介绍了DiTSE(扩散Transformer的语音增强),它解决了质量问题的退化语音在全带宽。我们的方法采用了一个潜在的扩散Transformer模型与强大的空调功能,有效地解决这些挑战,同时保持计算效率。主观和客观评估的实验结果表明,DiTSE实现了最先进的音频质量,首次与DAPS数据集的真实录音室质量音频相匹配。此外,与最先进的增强器相比,DiTSE显着提高了说话者身份和内容保真度的保留,减少了数据集上的幻觉。音频样本可在以下网址获得:http://hguimaraes.me/DiTSE
摘要:Real-world speech recordings suffer from degradations such as background noise and reverberation. Speech enhancement aims to mitigate these issues by generating clean high-fidelity signals. While recent generative approaches for speech enhancement have shown promising results, they still face two major challenges: (1) content hallucination, where plausible phonemes generated differ from the original utterance; and (2) inconsistency, failing to preserve speaker's identity and paralinguistic features from the input speech. In this work, we introduce DiTSE (Diffusion Transformer for Speech Enhancement), which addresses quality issues of degraded speech in full bandwidth. Our approach employs a latent diffusion transformer model together with robust conditioning features, effectively addressing these challenges while remaining computationally efficient. Experimental results from both subjective and objective evaluations demonstrate that DiTSE achieves state-of-the-art audio quality that, for the first time, matches real studio-quality audio from the DAPS dataset. Furthermore, DiTSE significantly improves the preservation of speaker identity and content fidelity, reducing hallucinations across datasets compared to state-of-the-art enhancers. Audio samples are available at: http://hguimaraes.me/DiTSE
标题: 用于高效Zero-Shot文本到语音合成的伪自回归神经编解码语言模型
链接:https://arxiv.org/abs/2504.10352
备注:Submitted to ACM MM 2025
摘要:最近的zero-shot文本到语音(TTS)系统面临着一个共同的困境:自回归(AR)模型的生成速度慢,缺乏持续时间的可控性,而非自回归(NAR)模型缺乏时间建模,通常需要复杂的设计。在本文中,我们介绍了一种新的伪自回归(PAR)编解码器语言建模方法,统一AR和NAR建模。PAR将AR的显式时态建模与NAR的并行生成相结合,以固定的时间步长生成动态长度跨度。PAR的基础上,我们提出了PALLE,一个两阶段的TTS系统,利用PAR的初始生成,然后NAR细化。在第一阶段中,PAR沿着时间维度逐步生成语音标记,每一步并行预测所有位置,但只保留最左边的跨度。在第二阶段中,低置信度令牌被并行迭代地细化,利用全局上下文信息。实验表明,在LibriTTS上训练的PALLE在语音质量、说话人相似度和可懂度方面优于在LibriSpeech测试干净集上训练的最先进的大规模数据系统,包括F5-TTS、E2-TTS和MaskGCT,同时实现了高达10倍的推理速度。音频样本可在https://anonymous-palle.github.io上获得。
摘要:Recent zero-shot text-to-speech (TTS) systems face a common dilemma: autoregressive (AR) models suffer from slow generation and lack duration controllability, while non-autoregressive (NAR) models lack temporal modeling and typically require complex designs. In this paper, we introduce a novel pseudo-autoregressive (PAR) codec language modeling approach that unifies AR and NAR modeling. Combining explicit temporal modeling from AR with parallel generation from NAR, PAR generates dynamic-length spans at fixed time steps. Building on PAR, we propose PALLE, a two-stage TTS system that leverages PAR for initial generation followed by NAR refinement. In the first stage, PAR progressively generates speech tokens along the time dimension, with each step predicting all positions in parallel but only retaining the left-most span. In the second stage, low-confidence tokens are iteratively refined in parallel, leveraging the global contextual information. Experiments demonstrate that PALLE, trained on LibriTTS, outperforms state-of-the-art systems trained on large-scale data, including F5-TTS, E2-TTS, and MaskGCT, on the LibriSpeech test-clean set in terms of speech quality, speaker similarity, and intelligibility, while achieving up to ten times faster inference speed. Audio samples are available at https://anonymous-palle.github.io.
【2】 DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
标题: DiPSE:通过潜在扩散变换器进行高保真生成语音增强链接:https://arxiv.org/abs/2504.09381
备注:Manuscript under review
摘要:真实世界的语音记录遭受诸如背景噪声和混响的劣化。语音增强旨在通过生成清晰的高保真信号来缓解这些问题。虽然最近的语音增强生成方法已经显示出有希望的结果,但它们仍然面临两个主要挑战:(1)内容幻觉,其中生成的看似合理的音素与原始话语不同;以及(2)不一致性,未能从输入语音中保留说话者的身份和非语言特征。在这项工作中,我们介绍了DiTSE(扩散Transformer的语音增强),它解决了质量问题的退化语音在全带宽。我们的方法采用了一个潜在的扩散Transformer模型与强大的空调功能,有效地解决这些挑战,同时保持计算效率。主观和客观评估的实验结果表明,DiTSE实现了最先进的音频质量,首次与DAPS数据集的真实录音室质量音频相匹配。此外,与最先进的增强器相比,DiTSE显着提高了说话者身份和内容保真度的保留,减少了数据集上的幻觉。音频样本可在以下网址获得:http://hguimaraes.me/DiTSE
摘要:Real-world speech recordings suffer from degradations such as background noise and reverberation. Speech enhancement aims to mitigate these issues by generating clean high-fidelity signals. While recent generative approaches for speech enhancement have shown promising results, they still face two major challenges: (1) content hallucination, where plausible phonemes generated differ from the original utterance; and (2) inconsistency, failing to preserve speaker's identity and paralinguistic features from the input speech. In this work, we introduce DiTSE (Diffusion Transformer for Speech Enhancement), which addresses quality issues of degraded speech in full bandwidth. Our approach employs a latent diffusion transformer model together with robust conditioning features, effectively addressing these challenges while remaining computationally efficient. Experimental results from both subjective and objective evaluations demonstrate that DiTSE achieves state-of-the-art audio quality that, for the first time, matches real studio-quality audio from the DAPS dataset. Furthermore, DiTSE significantly improves the preservation of speaker identity and content fidelity, reducing hallucinations across datasets compared to state-of-the-art enhancers. Audio samples are available at: http://hguimaraes.me/DiTSE
【3】 SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning
标题: Sift-50 M:用于语音教学微调的大规模多语言数据集链接:https://arxiv.org/abs/2504.09081
摘要:我们介绍了SIFT(Speech Instruction Fine-Tuning),这是一个5000万示例数据集,旨在对语音-文本大型语言模型(LLM)进行指令微调和预训练。SIFT-50 M是从公开可用的语音语料库构建的,这些语料库总共包含14 K小时的语音,并利用LLM和现成的专家模型。该数据集涵盖五种语言,包括各种语音理解以及可控的语音生成指令。使用SIFT-50 M,我们训练SIFT-LLM,它优于现有的语音-文本LLM的后续基准,同时实现基础语音任务的竞争力的表现。为了支持进一步的研究,我们还介绍了EvalSIFT,这是一个专门用于评估语音文本LLM的解释能力的基准数据集。
摘要:We introduce SIFT (Speech Instruction Fine-Tuning), a 50M-example dataset designed for instruction fine-tuning and pre-training of speech-text large language models (LLMs). SIFT-50M is built from publicly available speech corpora, which collectively contain 14K hours of speech, and leverages LLMs along with off-the-shelf expert models. The dataset spans five languages, encompassing a diverse range of speech understanding as well as controllable speech generation instructions. Using SIFT-50M, we train SIFT-LLM, which outperforms existing speech-text LLMs on instruction-following benchmarks while achieving competitive performance on foundational speech tasks. To support further research, we also introduce EvalSIFT, a benchmark dataset specifically designed to evaluate the instruction-following capabilities of speech-text LLMs.
【4】 Beyond Global Metrics: A Fairness Analysis for Interpretable Voice Disorder Detection Systems
标题: 超越全球范围:可解释语音障碍检测系统的公平性分析链接:https://arxiv.org/abs/2504.08997
备注:34 pages, 6 figures, 2 tables
摘要:我们使用现有的语音障碍数据集和可用的人口统计元数据对自动语音障碍检测(AVDD)系统进行了全面分析。该研究涉及分析不同人口群体的系统性能,特别关注基于性别和年龄的群体。绩效评估基于多个指标,包括标准化成本和交叉熵。我们采用了针对预定义人口统计群体单独培训的校准技术,以解决群体相关的误校准问题。分析显示,尽管全球指标很强,但各组之间的业绩差异很大。该系统显示出系统性偏差,将55岁以上的健康说话者误认为有声音障碍,将14-30岁的说话者误认为健康。特定于组的校准提高了后验概率质量,减少了过度自信。对于年轻的障碍扬声器,低的严重性分数被确定为有助于系统性能差。对于年龄较大的说话者,与年龄相关的语音特征和用作特征提取器的预训练Hubert模型中的潜在限制可能会影响结果。该研究表明,全球性能指标是不够的评估AVDD系统的性能。特定于组的分析可以揭示隐藏在全局度量中的系统性能问题。此外,组相关校准策略有助于减轻偏差,从而更可靠地指示系统置信度。这些研究结果强调了在语音障碍检测系统中进行人口统计学特定评估和校准的必要性,同时提供了一个适用于人口统计学元数据的更广泛生物医学分类任务的方法框架。
摘要:We conducted a comprehensive analysis of an Automatic Voice Disorders Detection (AVDD) system using existing voice disorder datasets with available demographic metadata. The study involved analysing system performance across various demographic groups, particularly focusing on gender and age-based cohorts. Performance evaluation was based on multiple metrics, including normalised costs and cross-entropy. We employed calibration techniques trained separately on predefined demographic groups to address group-dependent miscalibration. Analysis revealed significant performance disparities across groups despite strong global metrics. The system showed systematic biases, misclassifying healthy speakers over 55 as having a voice disorder and speakers with disorders aged 14-30 as healthy. Group-specific calibration improved posterior probability quality, reducing overconfidence. For young disordered speakers, low severity scores were identified as contributing to poor system performance. For older speakers, age-related voice characteristics and potential limitations in the pretrained Hubert model used as feature extractor likely affected results. The study demonstrates that global performance metrics are insufficient for evaluating AVDD system performance. Group-specific analysis may unmask problems in system performance which are hidden within global metrics. Further, group-dependent calibration strategies help mitigate biases, resulting in a more reliable indication of system confidence. These findings emphasize the need for demographic-specific evaluation and calibration in voice disorder detection systems, while providing a methodological framework applicable to broader biomedical classification tasks where demographic metadata is available.
【5】 Turn-taking annotation for quantitative and qualitative analyses of conversation
标题: 用于对话定量和定性分析的轮流注释链接:https://arxiv.org/abs/2504.09980
备注:41 pages
摘要:本文有两个目标。首先,我们提出的话轮转换注释层创建的95分钟的会话讲话的格拉茨语料库的阅读和自发演讲(GRASS),可供科学界。其次,我们更详细地描述了注释系统和注释过程,以便其他研究人员可以将其用于自己的会话数据。注释系统的开发考虑到了跨学科的应用。根据会话分析,它应该是基于顺序的标准,适合于随后的语音分析,因此,时间对齐的注释进行了Praat,它应该是适合于自动分类,这需要连续的语音注释和标签库存,是不是太大,并导致高评分员之间的协议。话轮转换被标注在两个层面上,即停顿间单位(IPU)和潜在完成点(PCCOMP;类似于过渡相关性位置)。我们提供了一个详细的说明的注释过程和分割和标签标准。对评分者间一致性和常见混淆的详细分析表明,IPU注释的一致性接近完美,PCCOMP注释的一致性很高,而分歧往往是部分的,或者可以通过对序列的不同分析来解释,这也是有价值的。注释系统可以应用于语言学研究和技术应用的各种会话数据,我们希望注释以及注释系统将有助于这些学科之间更强的交叉。
摘要:This paper has two goals. First, we present the turn-taking annotation layers created for 95 minutes of conversational speech of the Graz Corpus of Read and Spontaneous Speech (GRASS), available to the scientific community. Second, we describe the annotation system and the annotation process in more detail, so other researchers may use it for their own conversational data. The annotation system was developed with an interdisciplinary application in mind. It should be based on sequential criteria according to Conversation Analysis, suitable for subsequent phonetic analysis, thus time-aligned annotations were made Praat, and it should be suitable for automatic classification, which required the continuous annotation of speech and a label inventory that is not too large and results in a high inter-rater agreement. Turn-taking was annotated on two layers, Inter-Pausal Units (IPU) and points of potential completion (PCOMP; similar to transition relevance places). We provide a detailed description of the annotation process and of segmentation and labelling criteria. A detailed analysis of inter-rater agreement and common confusions shows that agreement for IPU annotation is near-perfect, that agreement for PCOMP annotations is substantial, and that disagreements often are either partial or can be explained by a different analysis of a sequence which also has merit. The annotation system can be applied to a variety of conversational data for linguistic studies and technological applications, and we hope that the annotations, as well as the annotation system will contribute to a stronger cross-fertilization between these disciplines.
【6】 Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion Synthesis
标题: 分开合作:协调钢琴手部运动合成的双流扩散模型链接:https://arxiv.org/abs/2504.09885
备注:12 pages, 4 figures
摘要:自动化协调双手钢琴演奏的合成带来了巨大的挑战,特别是在捕捉双手之间复杂的编舞,同时保留其独特的运动学特征。在本文中,我们提出了一个双流神经框架,旨在从音频输入中生成钢琴演奏的同步手势,解决了手部独立性和协调性建模的关键挑战。我们的框架引入了两个关键创新:(i)解耦的基于扩散的生成框架,其经由双噪声初始化独立地对每只手的运动进行建模,在利用共享位置条件的同时对每只手的不同潜在噪声进行采样,以及(ii)手协调非对称注意力(HCAA)机制抑制对称性运动。(共模)噪声以突出不对称的手部特定特征,同时在去噪期间自适应地增强手部间协调。该系统分层运行:它首先从音频特征预测3D手部位置,然后通过位置感知扩散模型生成关节角度,其中并行去噪流通过HCAA进行交互。综合评估表明,我们的框架在多个指标上优于现有的最先进的方法。
摘要:Automating the synthesis of coordinated bimanual piano performances poses significant challenges, particularly in capturing the intricate choreography between the hands while preserving their distinct kinematic signatures. In this paper, we propose a dual-stream neural framework designed to generate synchronized hand gestures for piano playing from audio input, addressing the critical challenge of modeling both hand independence and coordination. Our framework introduces two key innovations: (i) a decoupled diffusion-based generation framework that independently models each hand's motion via dual-noise initialization, sampling distinct latent noise for each while leveraging a shared positional condition, and (ii) a Hand-Coordinated Asymmetric Attention (HCAA) mechanism suppresses symmetric (common-mode) noise to highlight asymmetric hand-specific features, while adaptively enhancing inter-hand coordination during denoising. The system operates hierarchically: it first predicts 3D hand positions from audio features and then generates joint angles through position-aware diffusion models, where parallel denoising streams interact via HCAA. Comprehensive evaluations demonstrate that our framework outperforms existing state-of-the-art methods across multiple metrics.
【7】 FSSUAVL: A Discriminative Framework using Vision Models for Federated Self-Supervised Audio and Image Understanding
标题: FSSUAVL:使用视觉模型进行联邦自我监督音频和图像理解的区分性框架链接:https://arxiv.org/abs/2504.09516
备注:8 pages
摘要:最近的研究表明,视觉模型可以有效地学习多模态音频图像表示配对时。然而,使深度模型能够从未配对的模态中学习表示的挑战仍然没有解决。这个问题在像联邦学习(FL)这样的场景中尤其重要,在这些场景中,数据通常是分散的,异构的,并且缺乏配对数据的可靠保证。以前的尝试通过在本地客户端上使用辅助预训练编码器或生成模型来解决这个问题,这总是会随着模态数量的增加而增加计算成本。与这些方法不同,在本文中,我们的目标是使用\texttt{FSSUAVL}来解决未配对音频和图像识别的任务,这是一个使用自监督对比学习(SSL)在FL中预训练的单一深度模型。代替对齐音频和图像模态,\texttt{FSSUAVL}通过使用对比SSL将它们投影到公共嵌入空间中来共同区分它们。这将\texttt{FSSUAVL}的实用程序扩展到配对和未配对的音频和图像识别任务。我们使用CNN和ViT的实验表明,与为每种模态使用单独的深度模型相比,\texttt{FSSUAVL}显著提高了各种基于图像和音频的下游任务的性能。此外,FSSUAVL学习多模态特征表示的能力允许整合辅助信息(如果可用),以提高识别准确性。
摘要:Recent studies have demonstrated that vision models can effectively learn multimodal audio-image representations when paired. However, the challenge of enabling deep models to learn representations from unpaired modalities remains unresolved. This issue is especially pertinent in scenarios like Federated Learning (FL), where data is often decentralized, heterogeneous, and lacks a reliable guarantee of paired data. Previous attempts tackled this issue through the use of auxiliary pretrained encoders or generative models on local clients, which invariably raise computational cost with increasing number modalities. Unlike these approaches, in this paper, we aim to address the task of unpaired audio and image recognition using \texttt{FSSUAVL}, a single deep model pretrained in FL with self-supervised contrastive learning (SSL). Instead of aligning the audio and image modalities, \texttt{FSSUAVL} jointly discriminates them by projecting them into a common embedding space using contrastive SSL. This extends the utility of \texttt{FSSUAVL} to paired and unpaired audio and image recognition tasks. Our experiments with CNN and ViT demonstrate that \texttt{FSSUAVL} significantly improves performance across various image- and audio-based downstream tasks compared to using separate deep models for each modality. Additionally, \texttt{FSSUAVL}'s capacity to learn multimodal feature representations allows for integrating auxiliary information, if available, to enhance recognition accuracy.
【8】 AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis
标题: AMNet:增强型普通话语音合成的声学模型网络链接:https://arxiv.org/abs/2504.09225
备注:Main paper (8 pages). Accepted for publication by IJCNN 2025
摘要:本文介绍了AMNet,一个声学模型网络,旨在提高汉语语音合成的性能,通过结合短语结构注释和本地卷积模块。AMNet建立在FastSpeech 2架构的基础上,同时解决了本地上下文建模的挑战,这对于捕获复杂的语音特征(如停顿、重音和语调)至关重要。通过在模型中嵌入短语结构解析器并引入局部卷积模块,AMNet增强了模型对局部信息的敏感性。此外,AMNet从音素中提取音调特征,为音调建模提供明确的指导,从而提高音调准确性和发音。实验结果表明,AMNet优于基线模型的主观和客观评价。该模型实现了卓越的平均意见评分(MOS)、较低的梅尔倒谱系失真(MCD)以及改进的基频拟合$F0(R^2)$,证实了其生成高质量、自然且富有表现力的普通话语音的能力。
摘要:This paper presents AMNet, an Acoustic Model Network designed to improve the performance of Mandarin speech synthesis by incorporating phrase structure annotation and local convolution modules. AMNet builds upon the FastSpeech 2 architecture while addressing the challenge of local context modeling, which is crucial for capturing intricate speech features such as pauses, stress, and intonation. By embedding a phrase structure parser into the model and introducing a local convolution module, AMNet enhances the model's sensitivity to local information. Additionally, AMNet decouples tonal characteristics from phonemes, providing explicit guidance for tone modeling, which improves tone accuracy and pronunciation. Experimental results demonstrate that AMNet outperforms baseline models in subjective and objective evaluations. The proposed model achieves superior Mean Opinion Scores (MOS), lower Mel Cepstral Distortion (MCD), and improved fundamental frequency fitting $F0 (R^2)$, confirming its ability to generate high-quality, natural, and expressive Mandarin speech.
【9】 Generation of Musical Timbres using a Text-Guided Diffusion Model
标题: 使用文本引导扩散模型生成音乐音色链接:https://arxiv.org/abs/2504.09219
备注:10 pages, 5 figures
摘要:近年来,文本到音频系统取得了显著的成功,使得能够直接从文本描述生成完整的音频片段。虽然这些系统也促进了音乐创作,但人类创造力和刻意表达的元素往往是有限的。相比之下,目前的工作允许作曲家,作曲家和表演者创建音乐创作的基本构建块:用于电子乐器和数据采集器的单个音符的音频。通过文本提示,用户可以指定音频的音色特征。我们介绍了一个系统,结合了潜在的扩散模型和多模态对比学习生成音乐音色的文本描述的条件。通过联合生成频谱图的幅度和相位,我们的方法消除了随后运行相位恢复算法的需要,因为相关的方法。 音频示例、源代码和Web应用程序可在https://wxuanyuan.github.io/Musical-Note-Generation/上获得
摘要:In recent years, text-to-audio systems have achieved remarkable success, enabling the generation of complete audio segments directly from text descriptions. While these systems also facilitate music creation, the element of human creativity and deliberate expression is often limited. In contrast, the present work allows composers, arrangers, and performers to create the basic building blocks for music creation: audio of individual musical notes for use in electronic instruments and DAWs. Through text prompts, the user can specify the timbre characteristics of the audio. We introduce a system that combines a latent diffusion model and multi-modal contrastive learning to generate musical timbres conditioned on text descriptions. By jointly generating the magnitude and phase of the spectrogram, our method eliminates the need for subsequently running a phase retrieval algorithm, as related methods do. Audio examples, source code, and a web app are available at https://wxuanyuan.github.io/Musical-Note-Generation/
【10】 Spatial Audio Processing with Large Language Model on Wearable Devices
标题: 可穿戴设备上使用大语言模型的空间音频处理链接:https://arxiv.org/abs/2504.08907
摘要:将空间上下文集成到大型语言模型(LLM)中有可能彻底改变人机交互,特别是在可穿戴设备中。在这项工作中,我们提出了一种新的系统架构,将空间语音理解到LLM,使上下文感知和自适应应用可穿戴技术。我们的方法利用基于微结构的空间感测来使用单声道麦克风提取精确的到达方向(DoA)信息。为了解决现有数据集缺乏微结构辅助语音记录的问题,我们通过使用LibriSpeech数据集综合创建了一个名为OmniTalk的数据集。这种空间信息与OpenAI的Whisper模型的语言嵌入相融合,允许每种模态学习互补的上下文表示。融合嵌入与LLaMA-3.2 3B模型的输入空间对齐,并使用轻量级自适应技术LoRA进行微调,以优化设备上的处理。SING支持空间感知的自动语音识别(ASR),实现了25.72美元的平均误差-与现有工作中的88.52美元的中位误差相比有了很大的改善-单词错误率(WER)为5.3。SING还支持声音景观,例如,推断有多少人在说话及其方向,最多5人,DoA错误中位数为16$^\circ$。我们的系统在空间语音理解方面表现出卓越的性能,同时解决了能效、隐私和硬件限制的挑战,为增强现实、可访问性和沉浸式体验方面的高级应用铺平了道路。
摘要:Integrating spatial context into large language models (LLMs) has the potential to revolutionize human-computer interaction, particularly in wearable devices. In this work, we present a novel system architecture that incorporates spatial speech understanding into LLMs, enabling contextually aware and adaptive applications for wearable technologies. Our approach leverages microstructure-based spatial sensing to extract precise Direction of Arrival (DoA) information using a monaural microphone. To address the lack of existing dataset for microstructure-assisted speech recordings, we synthetically create a dataset called OmniTalk by using the LibriSpeech dataset. This spatial information is fused with linguistic embeddings from OpenAI's Whisper model, allowing each modality to learn complementary contextual representations. The fused embeddings are aligned with the input space of LLaMA-3.2 3B model and fine-tuned with lightweight adaptation technique LoRA to optimize for on-device processing. SING supports spatially-aware automatic speech recognition (ASR), achieving a mean error of $25.72^\circ$-a substantial improvement compared to the 88.52$^\circ$ median error in existing work-with a word error rate (WER) of 5.3. SING also supports soundscaping, for example, inference how many people were talking and their directions, with up to 5 people and a median DoA error of 16$^\circ$. Our system demonstrates superior performance in spatial speech understanding while addressing the challenges of power efficiency, privacy, and hardware constraints, paving the way for advanced applications in augmented reality, accessibility, and immersive experiences.
机器翻译由腾讯交互翻译提供,仅供参考
