今日论文合集:cs.SD语音4篇,eess.AS音频处理4篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音
【1】 DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided  Diffusion Model
标题:DualSec:通过双频谱图引导扩散模型生成文本到空间音频
链接:https://arxiv.org/abs/2502.18952
作者:Lei Zhao,  Sizhou Chen,  Linfeng Feng,  Xiao-Lei Zhang,  Xuelong Li
摘要:文本到音频(TTA),从文本描述生成音频信号,近年来受到了极大的关注。然而,最近的作品只专注于文本到单声道音频。如我们所知,空间音频提供比单声道音频更身临其境的听觉体验,例如在虚拟现实中。为了解决这个问题,我们提出了一个名为DualSpec的文本到空间音频(TTSA)生成框架。具体来说,它首先训练变分自编码器(VAE),用于从声音事件音频中提取潜在的声学表示。然后,给定描述声音事件和事件方向的文本,所提出的方法使用预训练的大型语言模型的编码器将文本转换为文本特征。最后,它从潜在的声学表示和文本特征训练扩散模型用于空间音频生成。在推理阶段,仅需要文本描述来生成空间音频。特别地,为了同时提高空间声事件的合成质量和方位精度,我们提出使用两种声学特征。一种是有利于提高合成质量的梅尔谱图,另一种是有利于提高方位精度的短时傅里叶变换谱图。我们提供了一个构建带有文本提示的空间音频数据集的管道,用于VAE和扩散模型的训练。我们还引入了新的空间感知评估指标来量化所生成的空间音频记录的方位角误差。实验结果表明,该方法可以生成具有较高方向性和事件一致性的空间音频。
摘要:Text-to-audio (TTA), which generates audio signals from textual descriptions,has received huge attention in recent years. However, recent works focused ontext to monaural audio only. As we know, spatial audio provides more immersiveauditory experience than monaural audio, e.g. in virtual reality. To addressthis issue, we propose a text-to-spatial-audio (TTSA) generation frameworknamed DualSpec.Specifically, it first trains variational autoencoders (VAEs)for extracting the latent acoustic representations from sound event audio.Then, given text that describes sound events and event directions, the proposedmethod uses the encoder of a pretrained large language model to transform thetext into text features. Finally, it trains a diffusion model from the latentacoustic representations and text features for the spatial audio generation. Inthe inference stage, only the text description is needed to generate spatialaudio. Particularly, to improve the synthesis quality and azimuth accuracy ofthe spatial sound events simultaneously, we propose to use two kinds ofacoustic features. One is the Mel spectrograms which is good for improving thesynthesis quality, and the other is the short-time Fourier transformspectrograms which is good at improving the azimuth accuracy. We provide apipeline of constructing spatial audio dataset with text prompts, for thetraining of the VAEs and diffusion model. We also introduce new spatial-awareevaluation metrics to quantify the azimuth errors of the generated spatialaudio recordings. Experimental results demonstrate that the proposed method cangenerate spatial audio with high directional and event consistency.

【2】 CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English  Code-Switching Dialogues for Speech Recognition
标题:CS-Dialogue:用于语音识别的自发普通英语代码切换对话的104小时数据集
链接:https://arxiv.org/abs/2502.18913
作者:Jiaming Zhou,  Yujie Guo,  Shiwan Zhao,  Haoqin Sun,  Hui Wang,  Jiabei He,  Aobo Kong,  Shiyao Wang,  Xi Yang,  Yequan Wang,  Yonghua Lin,  Yong Qin
摘要:语码转换(CS)是在一个会话中两种或多种语言之间的转换,它给自动语音识别(ASR)系统带来了巨大的挑战。现有的汉英语码转换数据集通常受到规模、自发性和缺乏完整的transmittance对话记录的限制,阻碍了为真实世界的会话场景开发鲁棒的ASR模型。本文介绍了CS-Dialogue,一个新的大规模的汉语-英语码转换语音数据集,包括104小时的自发对话,从200个扬声器。与以前的数据集不同,CS-Dialogue提供了完整的对话录音,捕捉了连续语音中自然的语码转换模式。我们描述了数据收集和注释过程,提供了数据集的详细统计数据,并使用最先进的模型建立了基准ASR性能。我们使用Transformer、Conformer和Branchformer的实验证明了代码切换ASR的挑战,并表明现有的预训练模型(如Whisper)仍有改进的空间。CS-Dialogue数据集将免费提供用于所有学术目的。
摘要:Code-switching (CS), the alternation between two or more languages within asingle conversation, presents significant challenges for automatic speechrecognition (ASR) systems. Existing Mandarin-English code-switching datasetsoften suffer from limitations in size, spontaneity, and the lack of full-lengthdialogue recordings with transcriptions, hindering the development of robustASR models for real-world conversational scenarios. This paper introducesCS-Dialogue, a novel large-scale Mandarin-English code-switching speech datasetcomprising 104 hours of spontaneous conversations from 200 speakers. Unlikeprevious datasets, CS-Dialogue provides full-length dialogue recordings withcomplete transcriptions, capturing naturalistic code-switching patterns incontinuous speech. We describe the data collection and annotation processes,present detailed statistics of the dataset, and establish benchmark ASRperformance using state-of-the-art models. Our experiments, using Transformer,Conformer, and Branchformer, demonstrate the challenges of code-switching ASR,and show that existing pre-trained models such as Whisper still have the spaceto improve. The CS-Dialogue dataset will be made freely available for allacademic purposes.

【3】 Clip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Huality  Text-to-Speech Method based on Contextual Semantic Understanding
标题:Clip-TTC:对比文本内容和Mel频谱图,一种基于上下文语义理解的高层次文本到语音方法
链接:https://arxiv.org/abs/2502.18889
作者:Tianyun Liu
摘要:传统的文本到语音(TTS)方法主要集中在建立音素和梅尔声谱图之间的映射。然而,在音素编码阶段,往往缺乏真正的梅尔语谱图辅助信息,这导致编码过程缺乏真正的语义理解。与此同时,传统的TTS系统通常很难平衡模型的推理速度与合成语音的质量。生成高质量合成语音的方法往往具有较慢的推理速度,而更快的推理方法通常会牺牲语音质量。在本文中,我提出了Clip-TTS,一种基于剪辑架构的TTS方法。该方法在文本编码阶段利用Clip框架建立文本内容与真实的mel-声谱图之间的联系,使文本编码器能够直接学习全局上下文的真实语义,从而保证合成语音的质量。在模型架构方面,采用了Transformer的基本结构,使得Clip-TTS能够实现快速的推理速度。实验结果表明,在LJSpeech和Baker数据集上,Clip-TTS生成的语音达到了最先进的MOS分数,并且在多情感数据集上也表现出色。音频样本可在https://ltydd1314.github.io/上获得。
摘要:Traditional text-to-speech (TTS) methods primarily focus on establishing amapping between phonemes and mel-spectrograms. However, during the phonemeencoding stage, there is often a lack of real mel-spectrogram auxiliaryinformation, which results in the encoding process lacking true semanticunderstanding. At the same time, traditional TTS systems often struggle tobalance the inference speed of the model with the quality of the synthesizedspeech. Methods that generate high-quality synthesized speech tend to haveslower inference speeds, while faster inference methods often sacrifice speechquality. In this paper, I propose Clip-TTS, a TTS method based on the Cliparchitecture. This method uses the Clip framework to establish a connectionbetween text content and real mel-spectrograms during the text encoding stage,enabling the text encoder to directly learn the true semantics of the globalcontext, thereby ensuring the quality of the synthesized speech. In terms ofmodel architecture, I adopt the basic structure of Transformer, which allowsClip-TTS to achieve fast inference speeds. Experimental results show that onthe LJSpeech and Baker datasets, the speech generated by Clip-TTS achievesstate-of-the-art MOS scores, and it also performs excellently on multi-emotiondatasets.Audio samples are available at: https://ltydd1314.github.io/.

【4】 Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot  Speech Synthesis
标题:用于零激发语音合成的稀疏对齐增强潜伏扩散Transformer
链接:https://arxiv.org/abs/2502.18924
作者:Ziyue Jiang,  Yi Ren,  Ruiqi Li,  Shengpeng Ji,  Zhenhui Ye,  Chen Zhang,  Bai Jionghao,  Xiaoda Yang,  Jialong Zuo,  Yu Zhang,  Rui Liu,  Xiang Yin,  Zhou Zhao
摘要:虽然最近的zero-shot文本到语音(TTS)模型显著地改善了语音质量和表达能力,但是主流系统仍然遭受与语音-文本对齐建模相关的问题:1)没有显式语音-文本对齐建模的模型表现出较低的鲁棒性,特别是对于实际应用中的硬句子; 2)预定义的基于语音对齐的模型遭受强制对齐的自然度约束。本文介绍了\textit{S-DiT},一个TTS系统,具有创新的稀疏对齐算法,指导潜在的扩散Transformer(DiT)。具体来说,我们为S-DiT提供稀疏对齐边界,以降低对齐学习的难度,而不限制搜索空间,从而实现高自然度。此外,我们采用了多条件分类器的指导策略,口音强度调整,并采用分段整流技术,以加快生成过程。实验表明,S-DiT实现了最先进的zero-shot TTS语音质量,并支持高度灵活的口音强度控制。值得注意的是,我们的系统可以生成高质量的一分钟的语音,只有8个采样步骤。音频样本可在https://sditdemo.github.io/sditdemo/上获得。
摘要:While recent zero-shot text-to-speech (TTS) models have significantlyimproved speech quality and expressiveness, mainstream systems still sufferfrom issues related to speech-text alignment modeling: 1) models withoutexplicit speech-text alignment modeling exhibit less robustness, especially forhard sentences in practical applications; 2) predefined alignment-based modelssuffer from naturalness constraints of forced alignments. This paper introduces\textit{S-DiT}, a TTS system featuring an innovative sparse alignment algorithmthat guides the latent diffusion transformer (DiT). Specifically, we providesparse alignment boundaries to S-DiT to reduce the difficulty of alignmentlearning without limiting the search space, thereby achieving high naturalness.Moreover, we employ a multi-condition classifier-free guidance strategy foraccent intensity adjustment and adopt the piecewise rectified flow technique toaccelerate the generation process. Experiments demonstrate that S-DiT achievesstate-of-the-art zero-shot TTS speech quality and supports highly flexiblecontrol over accent intensity. Notably, our system can generate high-qualityone-minute speech with only 8 sampling steps. Audio samples are available athttps://sditdemo.github.io/sditdemo/.

eess.AS音频处理

【1】 Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot  Speech Synthesis
标题:用于零激发语音合成的稀疏对齐增强潜伏扩散Transformer
链接:https://arxiv.org/abs/2502.18924
作者:Ziyue Jiang,  Yi Ren,  Ruiqi Li,  Shengpeng Ji,  Zhenhui Ye,  Chen Zhang,  Bai Jionghao,  Xiaoda Yang,  Jialong Zuo,  Yu Zhang,  Rui Liu,  Xiang Yin,  Zhou Zhao
摘要:虽然最近的zero-shot文本到语音(TTS)模型显著地改善了语音质量和表达能力,但是主流系统仍然遭受与语音-文本对齐建模相关的问题:1)没有显式语音-文本对齐建模的模型表现出较低的鲁棒性,特别是对于实际应用中的硬句子; 2)预定义的基于语音对齐的模型遭受强制对齐的自然度约束。本文介绍了\textit{S-DiT},一个TTS系统,具有创新的稀疏对齐算法,指导潜在的扩散Transformer(DiT)。具体来说,我们为S-DiT提供稀疏对齐边界,以降低对齐学习的难度,而不限制搜索空间,从而实现高自然度。此外,我们采用了多条件分类器的指导策略,口音强度调整,并采用分段整流技术,以加快生成过程。实验表明,S-DiT实现了最先进的zero-shot TTS语音质量,并支持高度灵活的口音强度控制。值得注意的是,我们的系统可以生成高质量的一分钟的语音,只有8个采样步骤。音频样本可在https://sditdemo.github.io/sditdemo/上获得。
摘要:While recent zero-shot text-to-speech (TTS) models have significantlyimproved speech quality and expressiveness, mainstream systems still sufferfrom issues related to speech-text alignment modeling: 1) models withoutexplicit speech-text alignment modeling exhibit less robustness, especially forhard sentences in practical applications; 2) predefined alignment-based modelssuffer from naturalness constraints of forced alignments. This paper introduces\textit{S-DiT}, a TTS system featuring an innovative sparse alignment algorithmthat guides the latent diffusion transformer (DiT). Specifically, we providesparse alignment boundaries to S-DiT to reduce the difficulty of alignmentlearning without limiting the search space, thereby achieving high naturalness.Moreover, we employ a multi-condition classifier-free guidance strategy foraccent intensity adjustment and adopt the piecewise rectified flow technique toaccelerate the generation process. Experiments demonstrate that S-DiT achievesstate-of-the-art zero-shot TTS speech quality and supports highly flexiblecontrol over accent intensity. Notably, our system can generate high-qualityone-minute speech with only 8 sampling steps. Audio samples are available athttps://sditdemo.github.io/sditdemo/.

【2】 DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided  Diffusion Model
标题:DualSec:通过双频谱图引导扩散模型生成文本到空间音频
链接:https://arxiv.org/abs/2502.18952
作者:Lei Zhao,  Sizhou Chen,  Linfeng Feng,  Xiao-Lei Zhang,  Xuelong Li
摘要:文本到音频(TTA),从文本描述生成音频信号,近年来受到了极大的关注。然而,最近的作品只专注于文本到单声道音频。如我们所知,空间音频提供比单声道音频更身临其境的听觉体验,例如在虚拟现实中。为了解决这个问题,我们提出了一个名为DualSpec的文本到空间音频(TTSA)生成框架。具体来说,它首先训练变分自编码器(VAE),用于从声音事件音频中提取潜在的声学表示。然后,给定描述声音事件和事件方向的文本,所提出的方法使用预训练的大型语言模型的编码器将文本转换为文本特征。最后,它从潜在的声学表示和文本特征训练扩散模型用于空间音频生成。在推理阶段,仅需要文本描述来生成空间音频。特别地,为了同时提高空间声事件的合成质量和方位精度,我们提出使用两种声学特征。一种是有利于提高合成质量的梅尔谱图,另一种是有利于提高方位精度的短时傅里叶变换谱图。我们提供了一个构建带有文本提示的空间音频数据集的管道,用于VAE和扩散模型的训练。我们还引入了新的空间感知评估指标来量化所生成的空间音频记录的方位角误差。实验结果表明,该方法可以生成具有较高方向性和事件一致性的空间音频。
摘要:Text-to-audio (TTA), which generates audio signals from textual descriptions,has received huge attention in recent years. However, recent works focused ontext to monaural audio only. As we know, spatial audio provides more immersiveauditory experience than monaural audio, e.g. in virtual reality. To addressthis issue, we propose a text-to-spatial-audio (TTSA) generation frameworknamed DualSpec.Specifically, it first trains variational autoencoders (VAEs)for extracting the latent acoustic representations from sound event audio.Then, given text that describes sound events and event directions, the proposedmethod uses the encoder of a pretrained large language model to transform thetext into text features. Finally, it trains a diffusion model from the latentacoustic representations and text features for the spatial audio generation. Inthe inference stage, only the text description is needed to generate spatialaudio. Particularly, to improve the synthesis quality and azimuth accuracy ofthe spatial sound events simultaneously, we propose to use two kinds ofacoustic features. One is the Mel spectrograms which is good for improving thesynthesis quality, and the other is the short-time Fourier transformspectrograms which is good at improving the azimuth accuracy. We provide apipeline of constructing spatial audio dataset with text prompts, for thetraining of the VAEs and diffusion model. We also introduce new spatial-awareevaluation metrics to quantify the azimuth errors of the generated spatialaudio recordings. Experimental results demonstrate that the proposed method cangenerate spatial audio with high directional and event consistency.

【3】 CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English  Code-Switching Dialogues for Speech Recognition
标题:CS-Dialogue:用于语音识别的自发普通英语代码切换对话的104小时数据集
链接:https://arxiv.org/abs/2502.18913
作者:Jiaming Zhou,  Yujie Guo,  Shiwan Zhao,  Haoqin Sun,  Hui Wang,  Jiabei He,  Aobo Kong,  Shiyao Wang,  Xi Yang,  Yequan Wang,  Yonghua Lin,  Yong Qin
摘要:语码转换(CS)是在一个会话中两种或多种语言之间的转换,它给自动语音识别(ASR)系统带来了巨大的挑战。现有的汉英语码转换数据集通常受到规模、自发性和缺乏完整的transmittance对话记录的限制,阻碍了为真实世界的会话场景开发鲁棒的ASR模型。本文介绍了CS-Dialogue,一个新的大规模的汉语-英语码转换语音数据集,包括104小时的自发对话,从200个扬声器。与以前的数据集不同,CS-Dialogue提供了完整的对话录音,捕捉了连续语音中自然的语码转换模式。我们描述了数据收集和注释过程,提供了数据集的详细统计数据,并使用最先进的模型建立了基准ASR性能。我们使用Transformer、Conformer和Branchformer的实验证明了代码切换ASR的挑战,并表明现有的预训练模型(如Whisper)仍有改进的空间。CS-Dialogue数据集将免费提供用于所有学术目的。
摘要:Code-switching (CS), the alternation between two or more languages within asingle conversation, presents significant challenges for automatic speechrecognition (ASR) systems. Existing Mandarin-English code-switching datasetsoften suffer from limitations in size, spontaneity, and the lack of full-lengthdialogue recordings with transcriptions, hindering the development of robustASR models for real-world conversational scenarios. This paper introducesCS-Dialogue, a novel large-scale Mandarin-English code-switching speech datasetcomprising 104 hours of spontaneous conversations from 200 speakers. Unlikeprevious datasets, CS-Dialogue provides full-length dialogue recordings withcomplete transcriptions, capturing naturalistic code-switching patterns incontinuous speech. We describe the data collection and annotation processes,present detailed statistics of the dataset, and establish benchmark ASRperformance using state-of-the-art models. Our experiments, using Transformer,Conformer, and Branchformer, demonstrate the challenges of code-switching ASR,and show that existing pre-trained models such as Whisper still have the spaceto improve. The CS-Dialogue dataset will be made freely available for allacademic purposes.

【4】 Clip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Huality  Text-to-Speech Method based on Contextual Semantic Understanding
标题:Clip-TTC:对比文本内容和Mel频谱图,一种基于上下文语义理解的高层次文本到语音方法
链接:https://arxiv.org/abs/2502.18889
作者:Tianyun Liu
摘要:传统的文本到语音(TTS)的方法主要集中在建立音素和梅尔声谱图之间的映射。然而,在音素编码阶段,往往缺乏真正的梅尔语谱图辅助信息,这导致编码过程缺乏真正的语义理解。与此同时,传统的TTS系统往往难以平衡模型的推理速度与合成语音的质量。生成高质量合成语音的方法往往具有较慢的推理速度,而更快的推理方法通常会牺牲语音质量。在本文中,我提出了Clip-TTS,一种基于剪辑架构的TTS方法。该方法在文本编码阶段利用Clip框架建立文本内容与真实的mel-声谱图之间的联系,使文本编码器能够直接学习全局上下文的真实语义,从而保证合成语音的质量。在模型架构方面,采用了Transformer的基本结构,使得Clip-TTS能够实现快速的推理速度。实验结果表明,在LJSpeech和Baker数据集上,Clip-TTS生成的语音达到了最先进的MOS分数,并且在多情感数据集上也表现出色。音频样本可在https://ltydd1314.github.io/上获得。
摘要:Traditional text-to-speech (TTS) methods primarily focus on establishing amapping between phonemes and mel-spectrograms. However, during the phonemeencoding stage, there is often a lack of real mel-spectrogram auxiliaryinformation, which results in the encoding process lacking true semanticunderstanding. At the same time, traditional TTS systems often struggle tobalance the inference speed of the model with the quality of the synthesizedspeech. Methods that generate high-quality synthesized speech tend to haveslower inference speeds, while faster inference methods often sacrifice speechquality. In this paper, I propose Clip-TTS, a TTS method based on the Cliparchitecture. This method uses the Clip framework to establish a connectionbetween text content and real mel-spectrograms during the text encoding stage,enabling the text encoder to directly learn the true semantics of the globalcontext, thereby ensuring the quality of the synthesized speech. In terms ofmodel architecture, I adopt the basic structure of Transformer, which allowsClip-TTS to achieve fast inference speeds. Experimental results show that onthe LJSpeech and Baker datasets, the speech generated by Clip-TTS achievesstate-of-the-art MOS scores, and it also performs excellently on multi-emotiondatasets.Audio samples are available at: https://ltydd1314.github.io/.

机器翻译由腾讯交互翻译提供,仅供参考