今日论文合集:cs.SD语音14篇,eess.AS音频处理17篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling
标题: VidMuse:一个具有长短期建模的简单视频到音乐生成框架
作者:Zeyue Tian,Zhaoyang Liu,Ruibin Yuan,Jiahao Pan,Xiaoqiang Huang,Qifeng Liu,Xu Tan,Qifeng Chen,Wei Xue,Yike Guo
备注:The code and datasets will be available at this https URL
链接:点击下载PDF文件
摘要:在这项工作中,我们系统地研究音乐生成条件下的视频。首先,我们提出了一个大规模的数据集,包括190K的视频音乐对,包括各种类型,如电影预告片,广告和纪录片。此外,我们提出了VidMuse,一个简单的框架,用于生成与视频输入对齐的音乐。VidMuse通过制作高保真音乐而脱颖而出,这些音乐在声学和语义上都与视频保持一致。通过整合本地和全局视觉线索,VidMuse能够创建音乐连贯的音轨,通过长短期建模始终匹配视频内容。通过大量的实验,VidMuse在音频质量,多样性和视听对齐方面优于现有模型。代码和数据集将在https: github.com ZeyueT VidMuse 上提供。摘要:In this work, we systematically study music generation conditioned solely on the video. First, we present a large-scale dataset comprising 190K video-music pairs, including various genres such as movie trailers, advertisements, and documentaries. Furthermore, we propose VidMuse, a simple framework for generating music aligned with video inputs. VidMuse stands out by producing high-fidelity music that is both acoustically and semantically aligned with the video. By incorporating local and global visual cues, VidMuse enables the creation of musically coherent audio tracks that consistently match the video content through Long-Short-Term modeling. Through extensive experiments, VidMuse outperforms existing models in terms of audio quality, diversity, and audio-visual alignment. The code and datasets will be available at https: github.com ZeyueT VidMuse .

【2】 STraDa: A Singer Traits Dataset
标题: STraDa:歌手特质数据集
作者:Yuexuan Kong,Viet-Anh Tran,Romain Hennequin
链接:点击下载PDF文件
摘要:有数量有限的大型公共数据集,包含可下载的音乐音频文件和丰富的主唱元数据。为了提供这样一个数据集,以利于歌唱声音的研究,我们创建了歌手特征数据集(STraDa),其中包括两个子集:自动strada和注释strada。automatic-strada包含超过5000个独特主唱的25000首曲目,这些曲目跨越了众多流派和语言,其中包括交叉验证的主唱元数据以及其他曲目元数据。注释-strada由200个轨道组成,在2个性别,5种语言和4个年龄组方面保持平衡。由于其元数据的丰富性和可下载的音频文件,为了显示其用于模型训练和偏差分析的用途,我们对歌手性别分类(SSC)进行了基准测试并进行了偏差分析。摘要:There is a limited amount of large-scale public datasets that contain downloadable music audio files and rich lead singer metadata. To provide such a dataset to benefit research in singing voices, we created Singer Traits Dataset (STraDa) with two subsets: automatic-strada and annotated-strada. The automatic-strada contains twenty-five thousand tracks across numerous genres and languages of more than five thousand unique lead singers, which includes cross-validated lead singer metadata as well as other track metadata. The annotated-strada consists of two hundred tracks that are balanced in terms of 2 genders, 5 languages, and 4 age groups. To show its use for model training and bias analysis thanks to its metadata's richness and downloadable audio files, we benchmarked singer sex classification (SSC) and conducted bias analysis.

【3】 Spontaneous Speech-Based Suicide Risk Detection Using Whisper and Large Language Models
标题: 使用Whisper和Large语言模型的基于自发言语的自杀风险检测
作者:Ziyun Cui,Chang Lei,Wen Wu,Yinan Duan,Diyang Qu,Ji Wu,Runsen Chen,Chao Zhang
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:早期发现自杀风险很重要,因为它使干预能够防止潜在的自杀企图。本文研究了基于青少年自发言语的自杀风险自动检测方法,并收集了1000多名10 ~ 18岁青少年15小时的自杀言语数据集进行实验。为了利用嵌入在自发语音中的各种声学和语言特征,Whisper语音模型和文本大语言模型(LLM)都用于自杀风险检测。使用全参数微调和参数有效微调方法来调整自杀风险检测的预训练模型,并评估多种音频-文本融合方法以结合Whisper和LLM的表示。所提出的系统实现了0.807的检测精度和F1分数为0.846的测试集与119个主题,表明有前途的潜在真正的自杀风险检测应用。摘要:The early detection of suicide risk is important since it enables the intervention to prevent potential suicide attempts. This paper studies the automatic detection of suicide risk based on spontaneous speech from adolescents, and collects a Mandarin dataset with 15 hours of suicide speech from more than a thousand adolescents aged from ten to eighteen for our experiments. To leverage the diverse acoustic and linguistic features embedded in spontaneous speech, both the Whisper speech model and textual large language models (LLMs) are used for suicide risk detection. Both all-parameter finetuning and parameter-efficient finetuning approaches are used to adapt the pre-trained models for suicide risk detection, and multiple audio-text fusion approaches are evaluated to combine the representations of Whisper and the LLM. The proposed system achieves a detection accuracy of 0.807 and an F1-score of 0.846 on the test set with 119 subjects, indicating promising potential for real suicide risk detection applications.

【4】 BLSP-Emo: Towards Empathetic Large Speech-Language Models
标题: BLSP-Emo:迈向同理心的大型演讲语言模型
作者:Chen Wang,Minpeng Liao,Zhongqiang Huang,Junhong Wu,Chengqing Zong,Jiajun Zhang
链接:点击下载PDF文件
摘要:最近发布的GPT-4 o展示了端到端多模态模型的潜力,不仅在低延迟方面,而且在理解和生成具有丰富情感的表达性语音的能力方面。虽然开放研究社区尚不清楚细节,但它可能涉及大量的策划数据和计算,这两者都不容易访问。在本文中,我们提出了BLSP-Emo(带情感支持的Bootstrapped语音预训练),这是一种开发端到端语音语言模型的新方法,该模型能够理解语音中的语义和情感,并生成移情反应。BLSP-Emo通过两个阶段的过程利用现有的语音识别(ASR)和语音情感识别(SER)数据集。第一阶段的重点是语义对齐,最近的工作是使用ASR数据预训练语音语言模型。第二阶段执行情绪对齐与预训练的语音语言模型上的情绪感知的连续任务从SER数据构建。我们的实验表明,BLSP-Emo模型在理解语音和提供移情反应方面表现出色,无论是在注意力跟随任务还是对话中。摘要:The recent release of GPT-4o showcased the potential of end-to-end multimodal models, not just in terms of low latency but also in their ability to understand and generate expressive speech with rich emotions. While the details are unknown to the open research community, it likely involves significant amounts of curated data and compute, neither of which is readily accessible. In this paper, we present BLSP-Emo (Bootstrapped Language-Speech Pretraining with Emotion support), a novel approach to developing an end-to-end speech-language model capable of understanding both semantics and emotions in speech and generate empathetic responses. BLSP-Emo utilizes existing speech recognition (ASR) and speech emotion recognition (SER) datasets through a two-stage process. The first stage focuses on semantic alignment, following recent work on pretraining speech-language models using ASR data. The second stage performs emotion alignment with the pretrained speech-language model on an emotion-aware continuation task constructed from SER data. Our experiments demonstrate that the BLSP-Emo model excels in comprehending speech and delivering empathetic responses, both in instruction-following tasks and conversations.

【5】 SilentCipher: Deep Audio Watermarking
标题: SilentCipher:深度音频水印
作者:Mayank Kumar Singh,Naoya Takahashi,Weihsiang Liao,Yuki Mitsufuji
链接:点击下载PDF文件
摘要:在音频水印领域,如何在增强水印信息容量和鲁棒性的同时,对不可感知的信息进行编码是一个具有挑战性的问题。尽管基于深度学习的方法的最新进展增强了消息容量和传统方法的鲁棒性,但编码消息引入了可听见的伪像,限制了它们在专业环境中的使用。在这项研究中,我们介绍了三个关键的创新。首先,我们的工作是第一个基于深度学习的模型,结合基于心理声学模型的阈值来实现不可感知的水印。其次,我们引入了伪可微压缩层,增强了水印算法的鲁棒性。最后,我们介绍了一种方法,以消除感知损失的需要,使我们能够实现SOTA在鲁棒性以及不可感知水印。我们的贡献导致我们SilentCipher,一个模型,使用户能够在44.1kHz采样的音频信号中编码消息。摘要:In the realm of audio watermarking, it is challenging to simultaneously encode imperceptible messages while enhancing the message capacity and robustness. Although recent advancements in deep learning-based methods bolster the message capacity and robustness over traditional methods, the encoded messages introduce audible artefacts that restricts their usage in professional settings. In this study, we introduce three key innovations. Firstly, our work is the first deep learning-based model to integrate psychoacoustic model based thresholding to achieve imperceptible watermarks. Secondly, we introduce psuedo-differentiable compression layers, enhancing the robustness of our watermarking algorithm. Lastly, we introduce a method to eliminate the need for perceptual losses, enabling us to achieve SOTA in both robustness as well as imperceptible watermarking. Our contributions lead us to SilentCipher, a model enabling users to encode messages within audio signals sampled at 44.1kHz.

【6】 Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
标题: 利用kNN-ctc和门控单语数据存储库改进Zero-Shot中英代码转换ASB
作者:Jiaming Zhou,Shiwan Zhao,Hui Wang,Tian-Hao Zhang,Haoqin Sun,Xuechen Wang,Yong Qin
链接:点击下载PDF文件
摘要:kNN-CTC模型已被证明是有效的单语自动语音识别(ASR)。然而,将其直接应用于多语言场景(如代码转换)带来了挑战。虽然存在性能改进的潜力,但利用单个双语词典的kNN-CTC模型可能会无意中从替代语言引入不期望的噪声。为了解决这个问题,我们提出了一种新的kNN-CTC为基础的代码切换ASR(CS-ASR)框架,采用双单语数据存储和门控的选择机制,以减少噪声干扰。我们的方法选择合适的解码器来解码每一帧,确保将特定于语言的信息注入ASR过程。我们将此框架应用于尖端的基于CTC的模型,开发了先进的CS-ASR系统。大量的实验表明,我们的门控搜索机制在提高zero-shot汉英CS-ASR的性能显着的效果。摘要:The kNN-CTC model has proven to be effective for monolingual automatic speech recognition (ASR). However, its direct application to multilingual scenarios like code-switching, presents challenges. Although there is potential for performance improvement, a kNN-CTC model utilizing a single bilingual datastore can inadvertently introduce undesirable noise from the alternative language. To address this, we propose a novel kNN-CTC-based code-switching ASR (CS-ASR) framework that employs dual monolingual datastores and a gated datastore selection mechanism to reduce noise interference. Our method selects the appropriate datastore for decoding each frame, ensuring the injection of language-specific information into the ASR process. We apply this framework to cutting-edge CTC-based models, developing an advanced CS-ASR system. Extensive experiments demonstrate the remarkable effectiveness of our gated datastore mechanism in enhancing the performance of zero-shot Chinese-English CS-ASR.

【7】 Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
标题: 基于预算的文本到语音合成中具有上下文感知对比语音-音频预训练的检索增强生成
作者:Jinlong Xue,Yayue Deng,Yingming Gao,Ya Li
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:最近的基于语音合成的文本到语音(TTS)模型可以克隆一个看不见的扬声器使用只有一个简短的语音提示。他们利用强大的上下文能力来模仿语音提示,包括说话者风格,韵律和情感。因此,语音提示的选择极大地影响了生成的语音,类似于大型语言模型(LLM)中提示的重要性。然而,目前基于语音识别的TTS模型手动或简单地随机选择语音提示。因此,在本文中,我们将LLM的检索增强生成(RAG)适应于基于提示的TTS。与传统的RAG方法不同,我们在检索过程中还考虑了上下文信息,并提出了上下文感知对比语言音频预训练(CA-CLAP)模型来提取上下文感知、风格相关的特征。客观和主观评估表明,我们提出的RAG方法优于基线,并且我们的CA-CLAP比纯文本检索方法取得了更好的结果。摘要:Recent prompt-based text-to-speech (TTS) models can clone an unseen speaker using only a short speech prompt. They leverage a strong in-context ability to mimic the speech prompts, including speaker style, prosody, and emotion. Therefore, the selection of a speech prompt greatly influences the generated speech, akin to the importance of a prompt in large language models (LLMs). However, current prompt-based TTS models choose the speech prompt manually or simply at random. Hence, in this paper, we adapt retrieval augmented generation (RAG) from LLMs to prompt-based TTS. Unlike traditional RAG methods, we additionally consider contextual information during the retrieval process and present a Context-Aware Contrastive Language-Audio Pre-training (CA-CLAP) model to extract context-aware, style-related features. The objective and subjective evaluations demonstrate that our proposed RAG method outperforms baselines, and our CA-CLAP achieves better results than text-only retrieval methods.

【8】 Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
标题: 利用多模式上下文和大语言模型改进基于音频编解码器的Zero-Shot文本到语音合成
作者:Jinlong Xue,Yayue Deng,Yicheng Han,Yingming Gao,Ya Li
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:大语言模型(LLM)的最新进展和音频编解码器的发展极大地推动了zero-shot TTS。他们可以合成个性化的语音,只有一个3秒的语音看不见的发言者作为声学提示。但是,它们只支持简短的语音提示,无法利用有声读物和会话TTS场景中所需的较长上下文信息。在本文中,我们介绍了一种新的基于音频编解码器的TTS模型,以适应上下文功能的多重增强。受Qformer成功的启发,我们提出了一个多模态上下文增强的Qformer(MMCE-Qformer),利用额外的多模态上下文信息。此外,我们采用预训练的LLM来利用其理解能力来预测语义令牌,并使用SoundStorm来生成声学令牌,从而提高音频质量和扬声器相似性。广泛的客观和主观评价表明,我们提出的方法优于基线在各种上下文TTS的情况下。摘要:Recent advances in large language models (LLMs) and development of audio codecs greatly propel the zero-shot TTS. They can synthesize personalized speech with only a 3-second speech of an unseen speaker as acoustic prompt. However, they only support short speech prompts and cannot leverage longer context information, as required in audiobook and conversational TTS scenarios. In this paper, we introduce a novel audio codec-based TTS model to adapt context features with multiple enhancements. Inspired by the success of Qformer, we propose a multi-modal context-enhanced Qformer (MMCE-Qformer) to utilize additional multi-modal context information. Besides, we adapt a pretrained LLM to leverage its understanding ability to predict semantic tokens, and use a SoundStorm to generate acoustic tokens thereby enhancing audio quality and speaker similarity. The extensive objective and subjective evaluations show that our proposed method outperforms baselines across various context TTS scenarios.

【9】 Harder or Different? Understanding Generalization of Audio Deepfake Detection
标题: 更难还是不同?了解音频Deepfake检测的一般化
作者:Nicolas M. Müller,Nicholas Evans,Hemlata Tak,Philip Sperl,Konstantin Böttinger
Journal-ref:Interspeech 2024
链接:点击下载PDF文件
摘要:最近的研究强调了语音深度伪造检测中的一个关键问题:在一组深度伪造上训练的模型在其他方面表现不佳。问题出现了:这是由于文本到语音(TTS)模型的质量不断提高,即,新的DeepFakes只是“更难”检测吗?或者,这是因为使用一种模型生成的deepfake与使用另一种模型生成的deepfake根本不同吗?我们通过将域内和域外测试数据之间的性能差距分解为“硬度”和“差异”组件来回答这个问题。使用ASVspoof数据库进行的实验表明,硬度分量几乎可以忽略不计,性能差距主要归因于差异分量。这对现实世界的deepfake检测有直接的影响,强调仅仅增加模型容量,目前占主导地位的研究趋势,可能无法有效地解决泛化挑战。摘要:Recent research has highlighted a key issue in speech deepfake detection: models trained on one set of deepfakes perform poorly on others. The question arises: is this due to the continuously improving quality of Text-to-Speech (TTS) models, i.e., are newer DeepFakes just 'harder' to detect? Or, is it because deepfakes generated with one model are fundamentally different to those generated using another model? We answer this question by decomposing the performance gap between in-domain and out-of-domain test data into 'hardness' and 'difference' components. Experiments performed using ASVspoof databases indicate that the hardness component is practically negligible, with the performance gap being attributed primarily to the difference component. This has direct implications for real-world deepfake detection, highlighting that merely increasing model capacity, the currently-dominant research trend, may not effectively address the generalization challenge.

【10】 Speech-based Clinical Depression Screening: An Empirical Study
标题: 基于言语的临床抑郁症筛查:一项实证研究
作者:Yangbin Chen,Chenyang Xu,Chunfeng Liang,Yanbao Tao,Chuan Shi
备注:5 pages, 3 figures
链接:点击下载PDF文件
摘要:这项研究调查了语音信号在各种交互场景中用于基于AI的抑郁症筛查的效用,包括精神病学访谈,聊天机器人对话和文本阅读。参与者包括北京大学第六医院门诊招募的抑郁症患者和来自社区的对照组成员,均由精神科医生按照标准化诊断方案进行诊断。我们从每个参与者的分段录音中提取声学和深度语音特征。使用神经网络或支持向量机进行分类,聚合剪辑结果确定最终评估。我们对交互场景、语音处理技术和特征类型的分析证实了语音是抑郁症筛查的关键标志。具体而言,人机交互匹配临床访谈功效,超过阅读任务。片段持续时间和数量显著影响模型性能,深度语音特征显著优于传统声学特征。摘要:This study investigates the utility of speech signals for AI-based depression screening across varied interaction scenarios, including psychiatric interviews, chatbot conversations, and text readings. Participants includes depressed patients recruited from the outpatient clinics of Peking University Sixth Hospital and control group members from the community, all diagnosed by psychiatrists following standardized diagnostic protocols. We extracted acoustic and deep speech features from each participant's segmented recordings. Classifications were made using neural networks or SVMs, with aggregated clip outcomes determining final assessments. Our analysis across interaction scenarios, speech processing techniques, and feature types confirms speech as a crucial marker for depression screening. Specifically, human-computer interaction matches clinical interview efficacy, surpassing reading tasks. Segment duration and quantity significantly affect model performance, with deep speech features substantially outperforming traditional acoustic features.

【11】 Beyond Performance Plateaus: A Comprehensive Study on Scalability in Speech Enhancement
标题: 超越性能平台:语音增强可扩展性的综合研究
作者:Wangyou Zhang,Kohei Saijo,Jee-weon Jung,Chenda Li,Shinji Watanabe,Yanmin Qian
备注:5 pages, 3 figures, 4 tables, Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:基于深度学习的语音增强(SE)模型在过去十年中取得了令人印象深刻的性能。许多先进的体系结构被设计为提供最先进的性能;然而,它们的可扩展性潜力仍然未被揭示。与此同时,大多数研究都集中在具有有限多样性的小规模数据集上,导致性能提高的平台。在本文中,我们的目标是通过探索SE模型在架构,模型大小,计算预算和数据集大小方面的可扩展性来解决上述问题。我们的调查涉及几个流行的SE架构和语音数据从不同的领域。实验揭示了SE和其他任务(如语音识别)的缩放效应之间的相似性和区别。这些发现进一步提供了对未充分探索的SE方向的见解,例如,更大规模的多领域语料库和高效可扩展的架构。摘要:Deep learning-based speech enhancement (SE) models have achieved impressive performance in the past decade. Numerous advanced architectures have been designed to deliver state-of-the-art performance; however, their scalability potential remains unrevealed. Meanwhile, the majority of research focuses on small-sized datasets with restricted diversity, leading to a plateau in performance improvement. In this paper, we aim to provide new insights for addressing the above issues by exploring the scalability of SE models in terms of architectures, model sizes, compute budgets, and dataset sizes. Our investigation involves several popular SE architectures and speech data from different domains. Experiments reveal both similarities and distinctions between the scaling effects in SE and other tasks such as speech recognition. These findings further provide insights into the under-explored SE directions, e.g., larger-scale multi-domain corpora and efficiently scalable architectures.

【12】 Sound Event Bounding Boxes
标题: 声音活动边界盒
作者:Janek Ebbers,Francois G. Germain,Gordon Wichern,Jonathan Le Roux
备注:Accepted for publication at Interspeech 2024
链接:点击下载PDF文件
摘要:声音事件检测是识别声音并确定其在音频片段中的范围(起始 偏移时间)的任务。现有系统通常在短时间帧内预测声音存在置信度。然后,阈值处理产生二进制帧级存在决策,通过合并连续的正帧来确定各个事件的程度。在本文中,我们表明,帧级阈值降低预测的事件程度耦合它与系统的声音存在的信心。我们建议通过引入SEBB来解耦事件程度和置信度的预测,SEBB将每个声音事件预测格式化为类类型,程度和总体置信度的元组。我们还提出了一种基于变化检测的算法,将传统的帧级输出转换为SEBB。我们发现该算法显著提高了DCASE 2023 Challenge系统的性能,将现有技术从.644提升到.686 PSDS 1。摘要:Sound event detection is the task of recognizing sounds and determining their extent (onset offset times) within an audio clip. Existing systems commonly predict sound presence confidence in short time frames. Then, thresholding produces binary frame-level presence decisions, with the extent of individual events determined by merging consecutive positive frames. In this paper, we show that frame-level thresholding degrades the prediction of the event extent by coupling it with the system's sound presence confidence. We propose to decouple the prediction of event extent and confidence by introducing SEBBs, which format each sound event prediction as a tuple of a class type, extent, and overall confidence. We also propose a change-detection-based algorithm to convert legacy frame-level outputs into SEBBs. We find the algorithm significantly improves the performance of DCASE 2023 Challenge systems, boosting the state of the art from .644 to .686 PSDS1.

【13】 UrBAN: Urban Beehive Acoustics and PheNotyping Dataset
标题: UrBAN:城市蜂巢声学和PheNotyping数据集
作者:Mahsa Abdollahi,Yi Zhu,Heitor R. Guimarães,Nico Coallier,Ségolène Maucourt,Pierre Giovenazzo,Tiago H. Falk
链接:点击下载PDF文件
摘要:在本文中,我们提出了一个多模态数据集,从一个蜜蜂殖民地在蒙特利尔,魁北克,加拿大,跨越2021年至2022年。这个养蜂场由10个蜂箱组成,麦克风记录了超过2000小时的高质量原始音频,还有传感器捕获温度和湿度。定期的蜂巢检查包括监测蜂群蜜蜂数量的变化,评估蜂王相关的状况,并记录蜂巢的整体健康状况。此外,还记录了瓦螨侵染率和冬季死亡率评估等健康指标,为影响蜂巢健康状况和恢复力的因素提供了有价值的见解。在这项研究中,我们首先概述了数据收集过程,传感器数据描述和数据集结构。此外,我们通过从原始音频中提取各种特征来展示该数据集的实际应用,以蜜蜂的帧数作为代理来预测群体数量。摘要:In this paper, we present a multimodal dataset obtained from a honey bee colony in Montr 'eal, Quebec, Canada, spanning the years of 2021 to 2022. This apiary comprised 10 beehives, with microphones recording more than 2000 hours of high quality raw audio, and also sensors capturing temperature, and humidity. Periodic hive inspections involved monitoring colony honey bee population changes, assessing queen-related conditions, and documenting overall hive health. Additionally, health metrics, such as Varroa mite infestation rates and winter mortality assessments were recorded, offering valuable insights into factors affecting hive health status and resilience. In this study, we first outline the data collection process, sensor data description, and dataset structure. Furthermore, we demonstrate a practical application of this dataset by extracting various features from the raw audio to predict colony population using the number of frames of bees as a proxy.

【14】 Style Mixture of Experts for Expressive Text-To-Speech Synthesis
标题: 表达性文本到语音合成的专家风格混合
作者:Ahad Jawaid,Shreeram Suresh Chandra,Junchen Lu,Berrak Sisman
链接:点击下载PDF文件
摘要:文体转换技术的最新进展提高了合成语音的表达能力。尽管有这些进步,从不同的和看不见的参考语音编码风格信息仍然具有挑战性。本文介绍了StyleMoE,一种方法,划分的嵌入空间,建模的风格编码器,到易处理的风格专家处理的子集。所提出的方法取代了TTS系统中的风格编码器与混合专家(MoE)层。通过利用选通网络将参考语音路由到不同的风格专家,每个专家在优化期间专门研究风格空间的各个方面。我们的实验客观和主观地证明了我们提出的方法在增加不同的和看不见的风格的风格空间的覆盖率的有效性。这种方法可以提高现有的国家的最先进的风格转移TTS模型的性能,标志着MoE的风格转移TTS研究,以我们的知识。摘要:Recent advances in style transfer text-to-speech (TTS) have improved the expressiveness of synthesized speech. Despite these advancements, encoding stylistic information from diverse and unseen reference speech remains challenging. This paper introduces StyleMoE, an approach that divides the embedding space, modeled by the style encoder, into tractable subsets handled by style experts. The proposed method replaces the style encoder in a TTS system with a Mixture of Experts (MoE) layer. By utilizing a gating network to route reference speeches to different style experts, each expert specializes in aspects of the style space during optimization. Our experiments objectively and subjectively demonstrate the effectiveness of our proposed method in increasing the coverage of the style space for diverse and unseen styles. This approach can enhance the performance of existing state-of-the-art style transfer TTS models, marking the first study of MoE in style transfer TTS to our knowledge.


eess.AS音频处理
【1】 Total-Duration-Aware Duration Modeling for Text-to-Speech Systems
标题: 文本转语音系统的总持续时间感知持续时间建模
作者:Sefik Emre Eskimez,Xiaofei Wang,Manthan Thakker,Chung-Hsien Tsai,Canrun Li,Zhen Xiao,Hemin Yang,Zirun Zhu,Min Tang,Jinyu Li,Sheng Zhao,Naoyuki Kanda
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:通过调整语音速率来精确控制生成语音的总持续时间对于各种文本到语音(TTS)应用至关重要。然而,调整语音速率对语音质量的影响,如可懂度和扬声器特性,一直未被充分探索。在这项工作中,我们提出了一种新的总持续时间感知(TDA)持续时间模型的TTS,其中音素持续时间预测不仅从文本输入,但也从一个额外的输入的总目标持续时间。我们还提出了一个基于MaskGIT的持续时间模型,提高了预测音素持续时间的多样性和质量。我们的研究结果表明,建议的TDA持续时间模型实现更好的可懂度和说话人相似性的各种语音速率配置相比,基线模型。我们还表明,所提出的MaskGIT为基础的模型可以生成音素持续时间具有更高的质量和多样性相比,其回归或流匹配的同行。摘要:Accurate control of the total duration of generated speech by adjusting the speech rate is crucial for various text-to-speech (TTS) applications. However, the impact of adjusting the speech rate on speech quality, such as intelligibility and speaker characteristics, has been underexplored. In this work, we propose a novel total-duration-aware (TDA) duration model for TTS, where phoneme durations are predicted not only from the text input but also from an additional input of the total target duration. We also propose a MaskGIT-based duration model that enhances the diversity and quality of the predicted phoneme durations. Our results demonstrate that the proposed TDA duration models achieve better intelligibility and speaker similarity for various speech rate configurations compared to the baseline models. We also show that the proposed MaskGIT-based model can generate phoneme durations with higher quality and diversity compared to its regression or flow-matching counterparts.

【2】 Beyond Performance Plateaus: A Comprehensive Study on Scalability in Speech Enhancement
标题: 超越性能平台:语音增强可扩展性的综合研究
作者:Wangyou Zhang,Kohei Saijo,Jee-weon Jung,Chenda Li,Shinji Watanabe,Yanmin Qian
备注:5 pages, 3 figures, 4 tables, Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:基于深度学习的语音增强(SE)模型在过去十年中取得了令人印象深刻的性能。许多先进的体系结构被设计为提供最先进的性能;然而,它们的可扩展性潜力仍然未被揭示。与此同时,大多数研究都集中在具有有限多样性的小规模数据集上,导致性能提高的平台。在本文中,我们的目标是通过探索SE模型在架构,模型大小,计算预算和数据集大小方面的可扩展性来解决上述问题。我们的调查涉及几个流行的SE架构和语音数据从不同的领域。实验揭示了SE和其他任务(如语音识别)的缩放效应之间的相似性和区别。这些发现进一步提供了对未充分探索的SE方向的见解,例如,更大规模的多领域语料库和高效可扩展的架构。摘要:Deep learning-based speech enhancement (SE) models have achieved impressive performance in the past decade. Numerous advanced architectures have been designed to deliver state-of-the-art performance; however, their scalability potential remains unrevealed. Meanwhile, the majority of research focuses on small-sized datasets with restricted diversity, leading to a plateau in performance improvement. In this paper, we aim to provide new insights for addressing the above issues by exploring the scalability of SE models in terms of architectures, model sizes, compute budgets, and dataset sizes. Our investigation involves several popular SE architectures and speech data from different domains. Experiments reveal both similarities and distinctions between the scaling effects in SE and other tasks such as speech recognition. These findings further provide insights into the under-explored SE directions, e.g., larger-scale multi-domain corpora and efficiently scalable architectures.

【3】 Sound Event Bounding Boxes
标题: 声音活动边界盒
作者:Janek Ebbers,Francois G. Germain,Gordon Wichern,Jonathan Le Roux
备注:Accepted for publication at Interspeech 2024
链接:点击下载PDF文件
摘要:声音事件检测是识别声音并确定其在音频片段中的范围(起始 偏移时间)的任务。现有系统通常在短时间帧内预测声音存在置信度。然后,阈值处理产生二进制帧级存在决策,通过合并连续的正帧来确定各个事件的程度。在本文中,我们表明,帧级阈值降低预测的事件程度耦合它与系统的声音存在的信心。我们建议通过引入SEBB来解耦事件程度和置信度的预测,SEBB将每个声音事件预测格式化为类类型,程度和总体置信度的元组。我们还提出了一种基于变化检测的算法,将传统的帧级输出转换为SEBB。我们发现该算法显著提高了DCASE 2023 Challenge系统的性能,将现有技术从.644提升到.686 PSDS 1。摘要:Sound event detection is the task of recognizing sounds and determining their extent (onset offset times) within an audio clip. Existing systems commonly predict sound presence confidence in short time frames. Then, thresholding produces binary frame-level presence decisions, with the extent of individual events determined by merging consecutive positive frames. In this paper, we show that frame-level thresholding degrades the prediction of the event extent by coupling it with the system's sound presence confidence. We propose to decouple the prediction of event extent and confidence by introducing SEBBs, which format each sound event prediction as a tuple of a class type, extent, and overall confidence. We also propose a change-detection-based algorithm to convert legacy frame-level outputs into SEBBs. We find the algorithm significantly improves the performance of DCASE 2023 Challenge systems, boosting the state of the art from .644 to .686 PSDS1.

【4】 Helsinki Speech Challenge 2024
标题: 2024年赫尔辛基演讲挑战赛
作者:Martin Ludvigsen,Elli Karvonen,Markus Juvonen,Samuli Siltanen
链接:点击下载PDF文件
摘要:2024年赫尔辛基演讲挑战赛(HSC 2024)邀请研究人员对语音录音进行增强和去卷积。我们记录了一个数据集,该数据集挑战参与者将语音增强和逆问题技术应用于记录的语音数据。该数据集包括人工智能生成的干净语音和相应录音的配对样本,这些样本具有不同程度的损坏,包括频率衰减和混响。挑战的重点是开发创新的去卷积方法,以准确地恢复原始音频。这些方法的有效性将使用语音识别模型进行定量评估,为评估现实世界场景中的增强功能提供相关指标。摘要:The Helsinki Speech Challenge 2024 (HSC2024) invites researchers to enhance and deconvolve speech audio recordings. We recorded a dataset that challenges participants to apply speech enhancement and inverse problems techniques to recorded speech data. This dataset includes paired samples of AI-generated clean speech and corresponding recordings, which feature varying levels of corruption, including frequency attenuation and reverberation. The challenge focuses on developing innovative deconvolution methods to accurately recover the original audio. The effectiveness of these methods will be quantitatively assessed using a speech recognition model, providing a relevant metric for evaluating enhancements in real-world scenarios.

【5】 PLDNet: PLD-Guided Lightweight Deep Network Boosted by Efficient Attention for Handheld Dual-Microphone Speech Enhancement
标题: PLDNet:通过手持式双麦克风语音增强的高效关注来推动WD引导的轻量级深度网络
作者:Nan Zhou,Youhai Jiang,Jialin Tan,Chongmin Qi
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:手机上的低复杂度语音增强在5G时代至关重要。因此,针对手持手机通信场景,基于功率级差(PLD)算法和轻量级U-Net,提出了PLD引导的轻量级深度网络(PLDNet),这是一种融合信号处理算法和轻量级注意力增强U-Net的极轻量级双麦克风语音增强方法。对于指导信息,我们采用PLD算法对双麦克风频谱进行预处理,并将输出馈送到后续的深度神经网络中,该网络利用我们提出的门控卷积增强频率注意力(GCAFA)模块的轻量级U-Net来提取所需的干净语音。实验结果表明,我们提出的方法实现了竞争力的性能与最近的顶级性能模型,同时减少了超过90%的计算成本,突出了低复杂度的语音增强在手机上的潜力。摘要:Low-complexity speech enhancement on mobile phones is crucial in the era of 5G. Thus, focusing on handheld mobile phone communication scenario, based on power level difference (PLD) algorithm and lightweight U-Net, we propose PLD-guided lightweight deep network (PLDNet), an extremely lightweight dual-microphone speech enhancement method that integrates the guidance of signal processing algorithm and lightweight attention-augmented U-Net. For the guidance information, we employ PLD algorithm to pre-process dual-microphone spectrum, and feed the output into subsequent deep neural network, which utilizes a lightweight U-Net with our proposed gated convolution augmented frequency attention (GCAFA) module to extract desired clean speech. Experimental results demonstrate that our proposed method achieves competitive performance with recent top-performing models while reducing computational cost by over 90%, highlighting the potential for low-complexity speech enhancement on mobile phones.

【6】 UrBAN: Urban Beehive Acoustics and PheNotyping Dataset
标题: UrBAN:城市蜂巢声学和PheNotyping数据集
作者:Mahsa Abdollahi,Yi Zhu,Heitor R. Guimarães,Nico Coallier,Ségolène Maucourt,Pierre Giovenazzo,Tiago H. Falk
链接:点击下载PDF文件
摘要:在本文中,我们提出了一个多模态数据集,从一个蜜蜂殖民地在蒙特利尔,魁北克,加拿大,跨越2021年至2022年。这个养蜂场由10个蜂箱组成,麦克风记录了超过2000小时的高质量原始音频,还有传感器捕获温度和湿度。定期的蜂巢检查包括监测蜂群蜜蜂数量的变化,评估蜂王相关的状况,并记录蜂巢的整体健康状况。此外,还记录了瓦螨侵染率和冬季死亡率评估等健康指标,为影响蜂巢健康状况和恢复力的因素提供了有价值的见解。在这项研究中,我们首先概述了数据收集过程,传感器数据描述和数据集结构。此外,我们通过从原始音频中提取各种特征来展示该数据集的实际应用,以蜜蜂的帧数作为代理来预测群体数量。摘要:In this paper, we present a multimodal dataset obtained from a honey bee colony in Montr 'eal, Quebec, Canada, spanning the years of 2021 to 2022. This apiary comprised 10 beehives, with microphones recording more than 2000 hours of high quality raw audio, and also sensors capturing temperature, and humidity. Periodic hive inspections involved monitoring colony honey bee population changes, assessing queen-related conditions, and documenting overall hive health. Additionally, health metrics, such as Varroa mite infestation rates and winter mortality assessments were recorded, offering valuable insights into factors affecting hive health status and resilience. In this study, we first outline the data collection process, sensor data description, and dataset structure. Furthermore, we demonstrate a practical application of this dataset by extracting various features from the raw audio to predict colony population using the number of frames of bees as a proxy.

【7】 Style Mixture of Experts for Expressive Text-To-Speech Synthesis
标题: 表达性文本到语音合成的专家风格混合
作者:Ahad Jawaid,Shreeram Suresh Chandra,Junchen Lu,Berrak Sisman
链接:点击下载PDF文件
摘要:文体转换技术的最新进展提高了合成语音的表达能力。尽管有这些进步,从不同的和看不见的参考语音编码风格信息仍然具有挑战性。本文介绍了StyleMoE,一种方法,划分的嵌入空间,建模的风格编码器,到易处理的风格专家处理的子集。所提出的方法取代了TTS系统中的风格编码器与混合专家(MoE)层。通过利用选通网络将参考语音路由到不同的风格专家,每个专家在优化期间专门研究风格空间的各个方面。我们的实验客观和主观地证明了我们提出的方法在增加不同的和看不见的风格的风格空间的覆盖率的有效性。这种方法可以提高现有的国家的最先进的风格转移TTS模型的性能,标志着MoE的风格转移TTS研究,以我们的知识。摘要:Recent advances in style transfer text-to-speech (TTS) have improved the expressiveness of synthesized speech. Despite these advancements, encoding stylistic information from diverse and unseen reference speech remains challenging. This paper introduces StyleMoE, an approach that divides the embedding space, modeled by the style encoder, into tractable subsets handled by style experts. The proposed method replaces the style encoder in a TTS system with a Mixture of Experts (MoE) layer. By utilizing a gating network to route reference speeches to different style experts, each expert specializes in aspects of the style space during optimization. Our experiments objectively and subjectively demonstrate the effectiveness of our proposed method in increasing the coverage of the style space for diverse and unseen styles. This approach can enhance the performance of existing state-of-the-art style transfer TTS models, marking the first study of MoE in style transfer TTS to our knowledge.

【8】 NeuRO: An Application for Code-Switched Autism Detection in Children
标题: NeuRO:儿童代码切换自闭症检测的应用
作者:Mohd Mujtaba Akhtar,Girish,Orchid Chetia Phukan,Muskaan Singh
备注:Accepted to INTERSPEECH 24 Show & Tell Demonstrations
链接:点击下载PDF文件
摘要:语码转换是一种常见的交际现象,指的是人们在同一会话中交替使用两种或两种以上的语言或语言风格。自闭症谱系障碍(ASD)是一种发育障碍,在社会互动,沟通和重复行为方面构成挑战。在具有代码转换场景的个体中检测ASD提出了独特的挑战。在本文中,我们通过构建一个应用程序NeuRO来解决这个问题,该应用程序旨在检测代码转换对话中自闭症的潜在迹象,促进对ASD患者的早期干预和支持。摘要:Code-switching is a common communication phenomenon where individuals alternate between two or more languages or linguistic styles within a single conversation. Autism Spectrum Disorder (ASD) is a developmental disorder posing challenges in social interaction, communication, and repetitive behaviors. Detecting ASD in individuals with code-switch scenario presents unique challenges. In this paper, we address this problem by building an application NeuRO which aims to detect potential signs of autism in code-switched conversations, facilitating early intervention and support for individuals with ASD.

【9】 STraDa: A Singer Traits Dataset
标题: STraDa:歌手特质数据集
作者:Yuexuan Kong,Viet-Anh Tran,Romain Hennequin
链接:点击下载PDF文件
摘要:有数量有限的大型公共数据集,包含可下载的音乐音频文件和丰富的主唱元数据。为了提供这样一个数据集,以利于歌唱声音的研究,我们创建了歌手特征数据集(STraDa),其中包括两个子集:自动strada和注释strada。automatic-strada包含超过5000个独特主唱的25000首曲目,这些曲目跨越了众多流派和语言,其中包括交叉验证的主唱元数据以及其他曲目元数据。注释-strada由200个轨道组成,在2个性别,5种语言和4个年龄组方面保持平衡。由于其元数据的丰富性和可下载的音频文件,为了显示其用于模型训练和偏差分析的用途,我们对歌手性别分类(SSC)进行了基准测试并进行了偏差分析。摘要:There is a limited amount of large-scale public datasets that contain downloadable music audio files and rich lead singer metadata. To provide such a dataset to benefit research in singing voices, we created Singer Traits Dataset (STraDa) with two subsets: automatic-strada and annotated-strada. The automatic-strada contains twenty-five thousand tracks across numerous genres and languages of more than five thousand unique lead singers, which includes cross-validated lead singer metadata as well as other track metadata. The annotated-strada consists of two hundred tracks that are balanced in terms of 2 genders, 5 languages, and 4 age groups. To show its use for model training and bias analysis thanks to its metadata's richness and downloadable audio files, we benchmarked singer sex classification (SSC) and conducted bias analysis.

【10】 Spontaneous Speech-Based Suicide Risk Detection Using Whisper and Large Language Models
标题: 使用Whisper和Large语言模型的基于自发言语的自杀风险检测
作者:Ziyun Cui,Chang Lei,Wen Wu,Yinan Duan,Diyang Qu,Ji Wu,Runsen Chen,Chao Zhang
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:早期发现自杀风险很重要,因为它使干预能够防止潜在的自杀企图。本文研究了基于青少年自发言语的自杀风险自动检测方法,并收集了1000多名10 ~ 18岁青少年15小时的自杀言语数据集进行实验。为了利用嵌入在自发语音中的各种声学和语言特征,Whisper语音模型和文本大语言模型(LLM)都用于自杀风险检测。使用全参数微调和参数有效微调方法来调整自杀风险检测的预训练模型,并评估多种音频-文本融合方法以结合Whisper和LLM的表示。所提出的系统实现了0.807的检测精度和F1分数为0.846的测试集与119个主题,表明有前途的潜在真正的自杀风险检测应用。摘要:The early detection of suicide risk is important since it enables the intervention to prevent potential suicide attempts. This paper studies the automatic detection of suicide risk based on spontaneous speech from adolescents, and collects a Mandarin dataset with 15 hours of suicide speech from more than a thousand adolescents aged from ten to eighteen for our experiments. To leverage the diverse acoustic and linguistic features embedded in spontaneous speech, both the Whisper speech model and textual large language models (LLMs) are used for suicide risk detection. Both all-parameter finetuning and parameter-efficient finetuning approaches are used to adapt the pre-trained models for suicide risk detection, and multiple audio-text fusion approaches are evaluated to combine the representations of Whisper and the LLM. The proposed system achieves a detection accuracy of 0.807 and an F1-score of 0.846 on the test set with 119 subjects, indicating promising potential for real suicide risk detection applications.

【11】 BLSP-Emo: Towards Empathetic Large Speech-Language Models
标题: BLSP-Emo:迈向同理心的大型演讲语言模型
作者:Chen Wang,Minpeng Liao,Zhongqiang Huang,Junhong Wu,Chengqing Zong,Jiajun Zhang
链接:点击下载PDF文件
摘要:最近发布的GPT-4 o展示了端到端多模态模型的潜力,不仅在低延迟方面,而且在理解和生成具有丰富情感的表达性语音的能力方面。虽然开放研究社区尚不清楚细节,但它可能涉及大量的策划数据和计算,这两者都不容易访问。在本文中,我们提出了BLSP-Emo(带情感支持的Bootstrapped语音预训练),这是一种开发端到端语音语言模型的新方法,该模型能够理解语音中的语义和情感,并生成移情反应。BLSP-Emo通过两个阶段的过程利用现有的语音识别(ASR)和语音情感识别(SER)数据集。第一阶段的重点是语义对齐,最近的工作是使用ASR数据预训练语音语言模型。第二阶段执行情绪对齐与预训练的语音语言模型上的情绪感知的连续任务从SER数据构建。我们的实验表明,BLSP-Emo模型在理解语音和提供移情反应方面表现出色,无论是在注意力跟随任务还是对话中。摘要:The recent release of GPT-4o showcased the potential of end-to-end multimodal models, not just in terms of low latency but also in their ability to understand and generate expressive speech with rich emotions. While the details are unknown to the open research community, it likely involves significant amounts of curated data and compute, neither of which is readily accessible. In this paper, we present BLSP-Emo (Bootstrapped Language-Speech Pretraining with Emotion support), a novel approach to developing an end-to-end speech-language model capable of understanding both semantics and emotions in speech and generate empathetic responses. BLSP-Emo utilizes existing speech recognition (ASR) and speech emotion recognition (SER) datasets through a two-stage process. The first stage focuses on semantic alignment, following recent work on pretraining speech-language models using ASR data. The second stage performs emotion alignment with the pretrained speech-language model on an emotion-aware continuation task constructed from SER data. Our experiments demonstrate that the BLSP-Emo model excels in comprehending speech and delivering empathetic responses, both in instruction-following tasks and conversations.

【12】 SilentCipher: Deep Audio Watermarking
标题: SilentCipher:深度音频水印
作者:Mayank Kumar Singh,Naoya Takahashi,Weihsiang Liao,Yuki Mitsufuji
链接:点击下载PDF文件
摘要:在音频水印领域,如何在增强水印信息容量和鲁棒性的同时,对不可感知的信息进行编码是一个具有挑战性的问题。尽管基于深度学习的方法的最新进展增强了消息容量和传统方法的鲁棒性,但编码消息引入了可听见的伪像,限制了它们在专业环境中的使用。在这项研究中,我们介绍了三个关键的创新。首先,我们的工作是第一个基于深度学习的模型,结合基于心理声学模型的阈值来实现不可感知的水印。其次,我们引入了伪可微压缩层,增强了水印算法的鲁棒性。最后,我们介绍了一种方法,以消除感知损失的需要,使我们能够实现SOTA在鲁棒性以及不可感知水印。我们的贡献导致我们SilentCipher,一个模型,使用户能够在44.1kHz采样的音频信号中编码消息。摘要:In the realm of audio watermarking, it is challenging to simultaneously encode imperceptible messages while enhancing the message capacity and robustness. Although recent advancements in deep learning-based methods bolster the message capacity and robustness over traditional methods, the encoded messages introduce audible artefacts that restricts their usage in professional settings. In this study, we introduce three key innovations. Firstly, our work is the first deep learning-based model to integrate psychoacoustic model based thresholding to achieve imperceptible watermarks. Secondly, we introduce psuedo-differentiable compression layers, enhancing the robustness of our watermarking algorithm. Lastly, we introduce a method to eliminate the need for perceptual losses, enabling us to achieve SOTA in both robustness as well as imperceptible watermarking. Our contributions lead us to SilentCipher, a model enabling users to encode messages within audio signals sampled at 44.1kHz.

【13】 Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
标题: 利用kNN-ctc和门控单语数据存储库改进Zero-Shot中英代码转换ASB
作者:Jiaming Zhou,Shiwan Zhao,Hui Wang,Tian-Hao Zhang,Haoqin Sun,Xuechen Wang,Yong Qin
链接:点击下载PDF文件
摘要:kNN-CTC模型已被证明是有效的单语自动语音识别(ASR)。然而,将其直接应用于多语言场景(如代码转换)带来了挑战。虽然存在性能改进的潜力,但利用单个双语词典的kNN-CTC模型可能会无意中从替代语言引入不期望的噪声。为了解决这个问题,我们提出了一种新的kNN-CTC为基础的代码切换ASR(CS-ASR)框架,采用双单语数据存储和门控的选择机制,以减少噪声干扰。我们的方法选择合适的解码器来解码每一帧,确保将特定于语言的信息注入ASR过程。我们将此框架应用于尖端的基于CTC的模型,开发了先进的CS-ASR系统。大量的实验表明,我们的门控搜索机制在提高zero-shot汉英CS-ASR的性能显着的效果。摘要:The kNN-CTC model has proven to be effective for monolingual automatic speech recognition (ASR). However, its direct application to multilingual scenarios like code-switching, presents challenges. Although there is potential for performance improvement, a kNN-CTC model utilizing a single bilingual datastore can inadvertently introduce undesirable noise from the alternative language. To address this, we propose a novel kNN-CTC-based code-switching ASR (CS-ASR) framework that employs dual monolingual datastores and a gated datastore selection mechanism to reduce noise interference. Our method selects the appropriate datastore for decoding each frame, ensuring the injection of language-specific information into the ASR process. We apply this framework to cutting-edge CTC-based models, developing an advanced CS-ASR system. Extensive experiments demonstrate the remarkable effectiveness of our gated datastore mechanism in enhancing the performance of zero-shot Chinese-English CS-ASR.

【14】 Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
标题: 基于预算的文本到语音合成中具有上下文感知对比语音-音频预训练的检索增强生成
作者:Jinlong Xue,Yayue Deng,Yingming Gao,Ya Li
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:最近的基于语音合成的文本到语音(TTS)模型可以克隆一个看不见的扬声器使用只有一个简短的语音提示。他们利用强大的上下文能力来模仿语音提示,包括说话者风格,韵律和情感。因此,语音提示的选择极大地影响了生成的语音,类似于大型语言模型(LLM)中提示的重要性。然而,目前基于语音识别的TTS模型手动或简单地随机选择语音提示。因此,在本文中,我们适应检索增广生成(RAG)从LLM到基于TTS。与传统的RAG方法不同,我们在检索过程中还考虑了上下文信息,并提出了一个上下文感知的对比存储音频预训练(CA-CLAP)模型来提取上下文感知的,风格相关的功能。客观和主观评价表明,我们提出的RAG方法优于基线,我们的CA-CLAP取得了更好的结果比纯文本检索方法。摘要:Recent prompt-based text-to-speech (TTS) models can clone an unseen speaker using only a short speech prompt. They leverage a strong in-context ability to mimic the speech prompts, including speaker style, prosody, and emotion. Therefore, the selection of a speech prompt greatly influences the generated speech, akin to the importance of a prompt in large language models (LLMs). However, current prompt-based TTS models choose the speech prompt manually or simply at random. Hence, in this paper, we adapt retrieval augmented generation (RAG) from LLMs to prompt-based TTS. Unlike traditional RAG methods, we additionally consider contextual information during the retrieval process and present a Context-Aware Contrastive Language-Audio Pre-training (CA-CLAP) model to extract context-aware, style-related features. The objective and subjective evaluations demonstrate that our proposed RAG method outperforms baselines, and our CA-CLAP achieves better results than text-only retrieval methods.

【15】 Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
标题: 利用多模式上下文和大语言模型改进基于音频编解码器的Zero-Shot文本到语音合成
作者:Jinlong Xue,Yayue Deng,Yicheng Han,Yingming Gao,Ya Li
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:大语言模型(LLM)的最新进展和音频编解码器的发展极大地推动了zero-shot TTS。他们可以合成个性化的语音,只有一个3秒的语音看不见的发言者作为声学提示。但是,它们只支持简短的语音提示,无法利用有声读物和会话TTS场景中所需的较长上下文信息。在本文中,我们介绍了一种新的基于音频编解码器的TTS模型,以适应上下文功能的多重增强。受到Qformer成功的启发,我们提出了一种多模态上下文增强Qformer(MMCE-Qformer)来利用额外的多模态上下文信息。此外,我们调整预训练的LLM以利用其理解能力来预测语义令牌,并使用SoundStorm生成声学令牌,从而提高音频质量和扬声器相似性。广泛的客观和主观评估表明,我们提出的方法在各种上下文TTS场景中的性能优于基线。摘要:Recent advances in large language models (LLMs) and development of audio codecs greatly propel the zero-shot TTS. They can synthesize personalized speech with only a 3-second speech of an unseen speaker as acoustic prompt. However, they only support short speech prompts and cannot leverage longer context information, as required in audiobook and conversational TTS scenarios. In this paper, we introduce a novel audio codec-based TTS model to adapt context features with multiple enhancements. Inspired by the success of Qformer, we propose a multi-modal context-enhanced Qformer (MMCE-Qformer) to utilize additional multi-modal context information. Besides, we adapt a pretrained LLM to leverage its understanding ability to predict semantic tokens, and use a SoundStorm to generate acoustic tokens thereby enhancing audio quality and speaker similarity. The extensive objective and subjective evaluations show that our proposed method outperforms baselines across various context TTS scenarios.

【16】 Harder or Different? Understanding Generalization of Audio Deepfake Detection
标题: 更难还是不同?了解音频Deepfake检测的一般化
作者:Nicolas M. Müller,Nicholas Evans,Hemlata Tak,Philip Sperl,Konstantin Böttinger
Journal-ref:Interspeech 2024
链接:点击下载PDF文件
摘要:最近的研究强调了语音深度伪造检测中的一个关键问题:在一组深度伪造上训练的模型在其他方面表现不佳。问题出现了:这是由于文本到语音(TTS)模型的质量不断提高,即,新的DeepFakes只是“更难”检测吗?或者,这是因为使用一种模型生成的deepfake与使用另一种模型生成的deepfake根本不同吗?我们通过将域内和域外测试数据之间的性能差距分解为“硬度”和“差异”组件来回答这个问题。使用ASVspoof数据库进行的实验表明,硬度分量几乎可以忽略不计,性能差距主要归因于差异分量。这对现实世界的deepfake检测有直接的影响,强调仅仅增加模型容量,目前占主导地位的研究趋势,可能无法有效地解决泛化挑战。摘要:Recent research has highlighted a key issue in speech deepfake detection: models trained on one set of deepfakes perform poorly on others. The question arises: is this due to the continuously improving quality of Text-to-Speech (TTS) models, i.e., are newer DeepFakes just 'harder' to detect? Or, is it because deepfakes generated with one model are fundamentally different to those generated using another model? We answer this question by decomposing the performance gap between in-domain and out-of-domain test data into 'hardness' and 'difference' components. Experiments performed using ASVspoof databases indicate that the hardness component is practically negligible, with the performance gap being attributed primarily to the difference component. This has direct implications for real-world deepfake detection, highlighting that merely increasing model capacity, the currently-dominant research trend, may not effectively address the generalization challenge.

【17】 Speech-based Clinical Depression Screening: An Empirical Study
标题: 基于言语的临床抑郁症筛查:一项实证研究
作者:Yangbin Chen,Chenyang Xu,Chunfeng Liang,Yanbao Tao,Chuan Shi
备注:5 pages, 3 figures
链接:点击下载PDF文件
摘要:这项研究调查了语音信号在各种交互场景中用于基于AI的抑郁症筛查的效用,包括精神病学访谈,聊天机器人对话和文本阅读。参与者包括北京大学第六医院门诊招募的抑郁症患者和来自社区的对照组成员,均由精神科医生按照标准化诊断方案进行诊断。我们从每个参与者的分段录音中提取声学和深度语音特征。使用神经网络或支持向量机进行分类,聚合剪辑结果确定最终评估。我们对交互场景、语音处理技术和特征类型的分析证实了语音是抑郁症筛查的关键标志。具体而言,人机交互匹配临床访谈功效,超过阅读任务。片段持续时间和数量显著影响模型性能,深度语音特征显著优于传统声学特征。摘要:This study investigates the utility of speech signals for AI-based depression screening across varied interaction scenarios, including psychiatric interviews, chatbot conversations, and text readings. Participants includes depressed patients recruited from the outpatient clinics of Peking University Sixth Hospital and control group members from the community, all diagnosed by psychiatrists following standardized diagnostic protocols. We extracted acoustic and deep speech features from each participant's segmented recordings. Classifications were made using neural networks or SVMs, with aggregated clip outcomes determining final assessments. Our analysis across interaction scenarios, speech processing techniques, and feature types confirms speech as a crucial marker for depression screening. Specifically, human-computer interaction matches clinical interview efficacy, surpassing reading tasks. Segment duration and quantity significantly affect model performance, with deep speech features substantially outperforming traditional acoustic features.


机器翻译,仅供参考