微信公众号:arXiv_Daily
cs.SD语音
链接:https://arxiv.org/abs/2504.07053
备注:Preprint. Work in progress
摘要:大型语言模型(LLM)在基于文本的自然语言处理任务中表现出色,但仍然受到其对文本输入和输出的依赖的限制。为了实现更自然的人类-LLM交互,最近的进展集中在导出一个口语模型(SLM),它不仅可以听,而且可以生成语音。为了实现这一点,一个很有前途的方向是进行语音-文本联合建模。然而,由于模态不匹配,最近的SLM仍然落后于文本LLM。一个显著的不匹配可能是语音和文本标记之间的序列长度。为了解决这个问题,我们引入了文本对齐语音标记和嵌入(TASTE),这是一种通过在标记化阶段将语音标记与相应的文本转录对齐来直接解决模态差距的方法。我们提出了一种方法,可以实现这一目标,通过特殊的聚合机制和语音重建的训练目标。我们进行了广泛的实验,并表明,TASTE可以保留基本的语言信息,同时大大减少令牌序列的长度。此外,通过利用TASTE,我们可以使用参数高效的微调技术(如低秩自适应(LoRA))将基于文本的LLM调整为有效的SLM。在SALMON和StoryCloze等基准任务上的实验结果表明,基于TASTE的SLM与以前的全微调方法具有相似的性能。据我们所知,TASTE是第一个端到端的方法,它利用重建目标来自动学习适合口语建模的文本对齐语音标记和嵌入。我们的演示、代码和模型可在https://github.com/mtkresearch/TASTE-SpokenLM上公开获取。
摘要:Large Language Models (LLMs) excel in text-based natural language processing tasks but remain constrained by their reliance on textual inputs and outputs. To enable more natural human-LLM interaction, recent progress have focused on deriving a spoken language model (SLM) that can not only listen but also generate speech. To achieve this, a promising direction is to conduct speech-text joint modeling. However, recent SLM still lag behind text LLM due to the modality mismatch. One significant mismatch can be the sequence lengths between speech and text tokens. To address this, we introduce Text-Aligned Speech Tokenization and Embedding (TASTE), a method that directly addresses the modality gap by aligning speech token with the corresponding text transcription during the tokenization stage. We propose a method that can achieve this through the special aggregation mechanism and with speech reconstruction as the training objective. We conduct extensive experiments and show that TASTE can preserve essential paralinguistic information while dramatically reducing the token sequence length. Furthermore, by leveraging TASTE, we can adapt text-based LLMs into effective SLMs with parameter-efficient fine-tuning techniques such as Low-Rank Adaptation (LoRA). Experimental results on benchmark tasks, including SALMON and StoryCloze, demonstrate that TASTE-based SLMs perform similarly to previous full-finetuning methods. To our knowledge, TASTE is the first end-to-end approach that utilizes a reconstruction objective to automatically learn a text-aligned speech tokenization and embedding suitable for spoken language modeling. Our demo, code, and models are publicly available at https://github.com/mtkresearch/TASTE-SpokenLM.
【2】 Controllable Automatic Foley Artist
标题: 可控自动弗利艺术家链接:https://arxiv.org/abs/2504.06778
摘要:Foley是视频制作中的关键元素,指的是将音频信号添加到无声视频中,同时确保语义和时间对齐的过程。近年来,个性化内容创建的兴起和自动视频到音频模型的进步增加了对过程中更大用户控制的需求。一种可能的方法是结合文本来指导音频生成。虽然有现有方法的支持,但在确保模式之间的兼容性方面仍然存在挑战,特别是当文本引入额外信息或与从视觉效果自然推断的声音相矛盾时。在这项工作中,我们介绍了CAFA(可控自动Foley艺术家)的视频和文本到音频模型,生成语义和时间对齐的音频为给定的视频,文本输入的指导。CAFA是建立在一个文本到音频模型,并通过模态适配器机制集成视频信息。通过合并文本,用户可以细化语义细节并引入创造性的变化,引导音频合成超出预期的视频上下文线索。实验表明,除了其优越的质量方面的语义对齐和视听同步,所提出的方法,使高文本可控性所证明的主观和客观的评价。
摘要:Foley is a key element in video production, refers to the process of adding an audio signal to a silent video while ensuring semantic and temporal alignment. In recent years, the rise of personalized content creation and advancements in automatic video-to-audio models have increased the demand for greater user control in the process. One possible approach is to incorporate text to guide audio generation. While supported by existing methods, challenges remain in ensuring compatibility between modalities, particularly when the text introduces additional information or contradicts the sounds naturally inferred from the visuals. In this work, we introduce CAFA (Controllable Automatic Foley Artist) a video-and-text-to-audio model that generates semantically and temporally aligned audio for a given video, guided by text input. CAFA is built upon a text-to-audio model and integrates video information through a modality adapter mechanism. By incorporating text, users can refine semantic details and introduce creative variations, guiding the audio synthesis beyond the expected video contextual cues. Experiments show that besides its superior quality in terms of semantic alignment and audio-visual synchronization the proposed method enable high textual controllability as demonstrated in subjective and objective evaluations.
【3】 Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory Perception
标题: 检测所有类型Deepfake音频:子波提示调整以增强听觉感知链接:https://arxiv.org/abs/2504.06753
摘要:音频生成技术的快速发展升级了恶意deepfake音频在语音、声音、歌声和音乐中的风险,威胁到多媒体的安全和信任。虽然现有的对策(CM)在单一类型的音频深度伪造检测(ADD)中表现良好,但它们的性能在跨类型场景中会下降。本文致力于研究各种类型的ADD任务。我们是第一个全面建立全类型ADD基准来评估当前CM的公司,包括跨语音,声音,歌声和音乐的交叉类型深度伪造检测。然后,我们介绍了提示调优自监督学习(PT-SSL)训练范例,它通过学习ADD的专用提示标记来优化SSL前端,所需的可训练参数比微调(FT)少458倍。考虑到不同音频类型的听觉感知,我们提出了小波提示调谐(WPT)-SSL方法,以从频域捕获类型不变的听觉deepfake信息,而不需要额外的训练参数,从而在所有类型的ADD任务中增强FT的性能。为了实现通用CM,我们利用所有类型的deepfake音频进行联合训练。实验结果表明,WPT-XLSR-AASIST取得了最好的性能,在所有评估集的平均EER为3.58%。代码可在线获取。
摘要:The rapid advancement of audio generation technologies has escalated the risks of malicious deepfake audio across speech, sound, singing voice, and music, threatening multimedia security and trust. While existing countermeasures (CMs) perform well in single-type audio deepfake detection (ADD), their performance declines in cross-type scenarios. This paper is dedicated to studying the alltype ADD task. We are the first to comprehensively establish an all-type ADD benchmark to evaluate current CMs, incorporating cross-type deepfake detection across speech, sound, singing voice, and music. Then, we introduce the prompt tuning self-supervised learning (PT-SSL) training paradigm, which optimizes SSL frontend by learning specialized prompt tokens for ADD, requiring 458x fewer trainable parameters than fine-tuning (FT). Considering the auditory perception of different audio types,we propose the wavelet prompt tuning (WPT)-SSL method to capture type-invariant auditory deepfake information from the frequency domain without requiring additional training parameters, thereby enhancing performance over FT in the all-type ADD task. To achieve an universally CM, we utilize all types of deepfake audio for co-training. Experimental results demonstrate that WPT-XLSR-AASIST achieved the best performance, with an average EER of 3.58% across all evaluation sets. The code is available online.
【4】 A Streamable Neural Audio Codec with Residual Scalar-Vector Quantization for Real-Time Communication
标题: 一种用于实时通信的基于残差标量-矢量量化的流式神经音频编解码器链接:https://arxiv.org/abs/2504.06561
备注:Accepted by IEEE Signal Processing Letters
摘要:本文提出了StreamCodec,一个可流式神经音频编解码器设计的实时通信。StreamCodec采用完全因果、对称的编解码器结构,并在改进的离散余弦变换(MDCT)域中操作,旨在实现低延迟推理和实时高效生成。为了提高码本利用率和补偿结构因果关系造成的音频质量损失,StreamCodec引入了一种新的残差标量矢量量化器(RSVQ)。RSVQ以残差的方式依次连接标量量化器和改进的矢量量化器,分别构建粗略的音频轮廓和细化声学细节。实验结果证实,所提出的StreamCodec实现了与高级非流式神经音频编解码器相当的解码音频质量。具体而言,在16 kHz LibriTTS数据集上,StreamCodec在1.5 kbps下的ViSQOL得分为4.30。它具有仅20 ms的固定延迟,在CPU上实现了近20倍的实时生成速度,轻量级模型大小仅为7M参数,非常适合实时通信应用。
摘要:This paper proposes StreamCodec, a streamable neural audio codec designed for real-time communication. StreamCodec adopts a fully causal, symmetric encoder-decoder structure and operates in the modified discrete cosine transform (MDCT) domain, aiming for low-latency inference and real-time efficient generation. To improve codebook utilization efficiency and compensate for the audio quality loss caused by structural causality, StreamCodec introduces a novel residual scalar-vector quantizer (RSVQ). The RSVQ sequentially connects scalar quantizers and improved vector quantizers in a residual manner, constructing coarse audio contours and refining acoustic details, respectively. Experimental results confirm that the proposed StreamCodec achieves decoded audio quality comparable to advanced non-streamable neural audio codecs. Specifically, on the 16 kHz LibriTTS dataset, StreamCodec attains a ViSQOL score of 4.30 at 1.5 kbps. It has a fixed latency of only 20 ms and achieves a generation speed nearly 20 times real-time on a CPU, with a lightweight model size of just 7M parameters, making it highly suitable for real-time communication applications.
【5】 A Cascaded Architecture for Extractive Summarization of Multimedia Content via Audio-to-Text Alignment
标题: 通过音频与文本对齐提取多媒体内容摘要的级联架构链接:https://arxiv.org/abs/2504.06275
摘要:本研究提出了一种通过音频到文本对齐进行多媒体内容提取摘要的级联架构。该框架解决了从YouTube视频等多媒体源中提取关键见解的挑战。它使用Microsoft Azure Speech与高级提取摘要模型(包括Whisper,Pegasus和Facebook BART XSum)集成了音频到文本转换。该系统使用Pytube、Pydub和SpeechRecognition等工具进行内容检索、音频提取和转录。通过命名实体识别和语义角色标注来增强语言分析。使用ROUGE和F1分数进行的评估表明,级联架构优于传统的摘要方法,尽管存在转录错误等挑战。未来的改进可能包括模型微调和实时处理。这项研究有助于多媒体摘要,提高信息检索,可访问性和用户体验。
摘要:This study presents a cascaded architecture for extractive summarization of multimedia content via audio-to-text alignment. The proposed framework addresses the challenge of extracting key insights from multimedia sources like YouTube videos. It integrates audio-to-text conversion using Microsoft Azure Speech with advanced extractive summarization models, including Whisper, Pegasus, and Facebook BART XSum. The system employs tools such as Pytube, Pydub, and SpeechRecognition for content retrieval, audio extraction, and transcription. Linguistic analysis is enhanced through named entity recognition and semantic role labeling. Evaluation using ROUGE and F1 scores demonstrates that the cascaded architecture outperforms conventional summarization methods, despite challenges like transcription errors. Future improvements may include model fine-tuning and real-time processing. This study contributes to multimedia summarization by improving information retrieval, accessibility, and user experience.
【6】 RNN-Transducer-based Losses for Speech Recognition on Noisy Targets
标题: 基于RNN变换器的噪声目标语音识别损失链接:https://arxiv.org/abs/2504.06963
备注:Final Project Report, Bachelor's Degree in Computer Science, University of London, March 2024
摘要:在嘈杂的转录本上训练语音识别系统是工业管道中的一个重大挑战,因为数据集非常庞大,很难确保每个实例的准确转录。在这项工作中,我们引入了新的损失函数来减轻RNN-换能器模型中转录错误的影响。我们的星形传感器损失通过在损失网格中加入“跳帧”转换来解决删除错误,与使用准确转录训练的模型相比,恢复了90%以上的系统性能。传感器丢失使用“跳过标记”转换来解决插入错误,恢复超过60%的质量。最后,目标鲁棒传感器损耗融合了这些方法,提供了针对任意误差的鲁棒性能。实验结果表明,与良好转录的数据相比,目标鲁棒换能器损失通过恢复超过70%的质量,显着提高了RNN-T在噪声数据上的性能。
摘要:Training speech recognition systems on noisy transcripts is a significant challenge in industrial pipelines, where datasets are enormous and ensuring accurate transcription for every instance is difficult. In this work, we introduce novel loss functions to mitigate the impact of transcription errors in RNN-Transducer models. Our Star-Transducer loss addresses deletion errors by incorporating "skip frame" transitions in the loss lattice, restoring over 90% of the system's performance compared to models trained with accurate transcripts. The Bypass-Transducer loss uses "skip token" transitions to tackle insertion errors, recovering more than 60% of the quality. Finally, the Target-Robust Transducer loss merges these approaches, offering robust performance against arbitrary errors. Experimental results demonstrate that the Target-Robust Transducer loss significantly improves RNN-T performance on noisy data by restoring over 70% of the quality compared to well-transcribed data.
标题: 基于RNN变换器的噪声目标语音识别损失
链接:https://arxiv.org/abs/2504.06963
备注:Final Project Report, Bachelor's Degree in Computer Science, University of London, March 2024
摘要:在嘈杂的转录本上训练语音识别系统是工业管道中的一个重大挑战,因为数据集非常庞大,很难确保每个实例的准确转录。在这项工作中,我们引入了新的损失函数来减轻RNN-换能器模型中转录错误的影响。我们的星形传感器损失通过在损失网格中加入“跳帧”转换来解决删除错误,与使用准确转录训练的模型相比,恢复了90%以上的系统性能。传感器丢失使用“跳过标记”转换来解决插入错误,恢复超过60%的质量。最后,目标鲁棒传感器损耗融合了这些方法,提供了针对任意误差的鲁棒性能。实验结果表明,与良好转录的数据相比,目标鲁棒换能器损失通过恢复超过70%的质量,显着提高了RNN-T在噪声数据上的性能。
摘要:Training speech recognition systems on noisy transcripts is a significant challenge in industrial pipelines, where datasets are enormous and ensuring accurate transcription for every instance is difficult. In this work, we introduce novel loss functions to mitigate the impact of transcription errors in RNN-Transducer models. Our Star-Transducer loss addresses deletion errors by incorporating "skip frame" transitions in the loss lattice, restoring over 90% of the system's performance compared to models trained with accurate transcripts. The Bypass-Transducer loss uses "skip token" transitions to tackle insertion errors, recovering more than 60% of the quality. Finally, the Target-Robust Transducer loss merges these approaches, offering robust performance against arbitrary errors. Experimental results demonstrate that the Target-Robust Transducer loss significantly improves RNN-T performance on noisy data by restoring over 70% of the quality compared to well-transcribed data.
【2】 TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
标题: TASTE:口语建模的文本对齐语音标记和嵌入链接:https://arxiv.org/abs/2504.07053
备注:Preprint. Work in progress
摘要:大型语言模型(LLM)在基于文本的自然语言处理任务中表现出色,但仍然受到其对文本输入和输出的依赖的限制。为了实现更自然的人类-LLM交互,最近的进展集中在导出一个口语模型(SLM),它不仅可以听,而且可以生成语音。为了实现这一点,一个很有前途的方向是进行语音-文本联合建模。然而,由于模态不匹配,最近的SLM仍然落后于文本LLM。一个显著的不匹配可能是语音和文本标记之间的序列长度。为了解决这个问题,我们引入了文本对齐语音标记和嵌入(TASTE),这是一种通过在标记化阶段将语音标记与相应的文本转录对齐来直接解决模态差距的方法。我们提出了一种方法,可以实现这一目标,通过特殊的聚合机制和语音重建的训练目标。我们进行了广泛的实验,并表明,TASTE可以保留基本的语言信息,同时大大减少令牌序列的长度。此外,通过利用TASTE,我们可以使用参数高效的微调技术(如低秩自适应(LoRA))将基于文本的LLM调整为有效的SLM。在SALMON和StoryCloze等基准任务上的实验结果表明,基于TASTE的SLM与以前的全微调方法具有相似的性能。据我们所知,TASTE是第一个端到端的方法,它利用重建目标来自动学习适合口语建模的文本对齐语音标记和嵌入。我们的演示、代码和模型可在https://github.com/mtkresearch/TASTE-SpokenLM上公开获取。
摘要:Large Language Models (LLMs) excel in text-based natural language processing tasks but remain constrained by their reliance on textual inputs and outputs. To enable more natural human-LLM interaction, recent progress have focused on deriving a spoken language model (SLM) that can not only listen but also generate speech. To achieve this, a promising direction is to conduct speech-text joint modeling. However, recent SLM still lag behind text LLM due to the modality mismatch. One significant mismatch can be the sequence lengths between speech and text tokens. To address this, we introduce Text-Aligned Speech Tokenization and Embedding (TASTE), a method that directly addresses the modality gap by aligning speech token with the corresponding text transcription during the tokenization stage. We propose a method that can achieve this through the special aggregation mechanism and with speech reconstruction as the training objective. We conduct extensive experiments and show that TASTE can preserve essential paralinguistic information while dramatically reducing the token sequence length. Furthermore, by leveraging TASTE, we can adapt text-based LLMs into effective SLMs with parameter-efficient fine-tuning techniques such as Low-Rank Adaptation (LoRA). Experimental results on benchmark tasks, including SALMON and StoryCloze, demonstrate that TASTE-based SLMs perform similarly to previous full-finetuning methods. To our knowledge, TASTE is the first end-to-end approach that utilizes a reconstruction objective to automatically learn a text-aligned speech tokenization and embedding suitable for spoken language modeling. Our demo, code, and models are publicly available at https://github.com/mtkresearch/TASTE-SpokenLM.
【3】 Controllable Automatic Foley Artist
标题: 可控自动弗利艺术家链接:https://arxiv.org/abs/2504.06778
摘要:Foley是视频制作中的关键元素,指的是将音频信号添加到无声视频中,同时确保语义和时间对齐的过程。近年来,个性化内容创建的兴起和自动视频到音频模型的进步增加了对过程中更大用户控制的需求。一种可能的方法是结合文本来指导音频生成。虽然有现有方法的支持,但在确保模式之间的兼容性方面仍然存在挑战,特别是当文本引入额外信息或与从视觉效果自然推断的声音相矛盾时。在这项工作中,我们介绍了CAFA(可控自动Foley艺术家)的视频和文本到音频模型,生成语义和时间对齐的音频为给定的视频,文本输入的指导。CAFA是建立在一个文本到音频模型,并通过模态适配器机制集成视频信息。通过合并文本,用户可以细化语义细节并引入创造性的变化,引导音频合成超出预期的视频上下文线索。实验表明,除了其优越的质量方面的语义对齐和视听同步,所提出的方法,使高文本可控性所证明的主观和客观的评价。
摘要:Foley is a key element in video production, refers to the process of adding an audio signal to a silent video while ensuring semantic and temporal alignment. In recent years, the rise of personalized content creation and advancements in automatic video-to-audio models have increased the demand for greater user control in the process. One possible approach is to incorporate text to guide audio generation. While supported by existing methods, challenges remain in ensuring compatibility between modalities, particularly when the text introduces additional information or contradicts the sounds naturally inferred from the visuals. In this work, we introduce CAFA (Controllable Automatic Foley Artist) a video-and-text-to-audio model that generates semantically and temporally aligned audio for a given video, guided by text input. CAFA is built upon a text-to-audio model and integrates video information through a modality adapter mechanism. By incorporating text, users can refine semantic details and introduce creative variations, guiding the audio synthesis beyond the expected video contextual cues. Experiments show that besides its superior quality in terms of semantic alignment and audio-visual synchronization the proposed method enable high textual controllability as demonstrated in subjective and objective evaluations.
【4】 A Cascaded Architecture for Extractive Summarization of Multimedia Content via Audio-to-Text Alignment
标题: 通过音频与文本对齐提取多媒体内容摘要的级联架构链接:https://arxiv.org/abs/2504.06275
摘要:本研究提出一个级联架构,透过音频到文字对齐的多媒体内容摘要提取。该框架解决了从YouTube视频等多媒体源中提取关键见解的挑战。它使用Microsoft Azure Speech与高级提取摘要模型(包括Whisper,Pegasus和Facebook BART XSum)集成了音频到文本转换。该系统使用Pytube、Pydub和SpeechRecognition等工具进行内容检索、音频提取和转录。通过命名实体识别和语义角色标记来增强语言分析。使用ROUGE和F1分数进行的评估表明,级联架构优于传统的摘要方法,尽管存在转录错误等挑战。未来的改进可能包括模型微调和实时处理。这项研究有助于多媒体摘要,提高信息检索,可访问性和用户体验。
摘要:This study presents a cascaded architecture for extractive summarization of multimedia content via audio-to-text alignment. The proposed framework addresses the challenge of extracting key insights from multimedia sources like YouTube videos. It integrates audio-to-text conversion using Microsoft Azure Speech with advanced extractive summarization models, including Whisper, Pegasus, and Facebook BART XSum. The system employs tools such as Pytube, Pydub, and SpeechRecognition for content retrieval, audio extraction, and transcription. Linguistic analysis is enhanced through named entity recognition and semantic role labeling. Evaluation using ROUGE and F1 scores demonstrates that the cascaded architecture outperforms conventional summarization methods, despite challenges like transcription errors. Future improvements may include model fine-tuning and real-time processing. This study contributes to multimedia summarization by improving information retrieval, accessibility, and user experience.
机器翻译由腾讯交互翻译提供,仅供参考
