微信公众号:arXiv_Daily
cs.SD语音
标题: 量化一对一语音转换中的源扬声器泄漏
链接:https://arxiv.org/abs/2504.15822
备注:Accepted at IEEE 23rd International Conference of the Biometrics Special Interest Group (BIOSIG 2024)
摘要:使用一个多口音语料库的并行话语与商业语音设备的使用,我们提出了一个案例研究表明,它是可以量化的一对一的语音转换的情况下,一个源扬声器的身份的置信度。在使用HiFi-GAN声码器进行语音转换之后,我们比较了一系列扬声器特征的信息泄漏;假设“最坏情况”的白盒场景,我们量化了我们执行推断的信心,并缩小了可能的源扬声器的范围,加强了合成语音提供商必须确保其扬声器数据隐私的监管义务和道德责任。
摘要:Using a multi-accented corpus of parallel utterances for use with commercial speech devices, we present a case study to show that it is possible to quantify a degree of confidence about a source speaker's identity in the case of one-to-one voice conversion. Following voice conversion using a HiFi-GAN vocoder, we compare information leakage for a range speaker characteristics; assuming a "worst-case" white-box scenario, we quantify our confidence to perform inference and narrow the pool of likely source speakers, reinforcing the regulatory obligation and moral duty that providers of synthetic voices have to ensure the privacy of their speakers' data.
【2】 SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
标题: Sim 2S-LLM:解锁语音LLM的同时推理以实现语音到语音翻译链接:https://arxiv.org/abs/2504.15509
摘要:同步语音翻译(SST)与流式语音输入并行输出翻译,从而平衡翻译质量和延迟。虽然大型语言模型(LLM)已经扩展到处理语音模态,但由于语音被预先作为整个生成过程的提示,因此流仍然具有挑战性。为了解锁LLM流媒体功能,本文提出了SimulS 2S-LLM,离线训练语音LLM,并采用测试时间策略来指导同时推理。SimulS 2S-LLM通过提取边界感知语音提示来消除训练和推理之间的不匹配,从而使其与文本输入数据更好地匹配。SimulS 2S-LLM通过预测离散输出语音令牌,然后使用预先训练的声码器合成输出语音,实现同步语音到语音翻译(Simul-S2 ST)。设计了一种增量波束搜索算法,在不增加时延的情况下扩展了语音标记预测的搜索空间。在CVSS语音数据上的实验表明,与使用相同训练数据的现有方法相比,SimulS 2S-LLM提供了更好的翻译质量-延迟权衡,例如在类似延迟的情况下将ASR-BLEU分数提高了3分。
摘要:Simultaneous speech translation (SST) outputs translations in parallel with streaming speech input, balancing translation quality and latency. While large language models (LLMs) have been extended to handle the speech modality, streaming remains challenging as speech is prepended as a prompt for the entire generation process. To unlock LLM streaming capability, this paper proposes SimulS2S-LLM, which trains speech LLMs offline and employs a test-time policy to guide simultaneous inference. SimulS2S-LLM alleviates the mismatch between training and inference by extracting boundary-aware speech prompts that allows it to be better matched with text input data. SimulS2S-LLM achieves simultaneous speech-to-speech translation (Simul-S2ST) by predicting discrete output speech tokens and then synthesising output speech using a pre-trained vocoder. An incremental beam search is designed to expand the search space of speech token prediction without increasing latency. Experiments on the CVSS speech data show that SimulS2S-LLM offers a better translation quality-latency trade-off than existing methods that use the same training data, such as improving ASR-BLEU scores by 3 points at similar latency.
【3】 Exploring the User Experience of AI-Assisted Sound Searching Systems for Creative Workflows
标题: 探索创意工作流程的人工智能辅助声音搜索系统的用户体验链接:https://arxiv.org/abs/2504.15575
摘要:有效地定位正确的声音效果是音频制作的一个重要但具有挑战性的主题。大多数当前的声音搜索系统依赖于由人类创建的预先注释的音频标签,这可能是耗时的,并且容易出现不准确,限制了音频制作的效率。在对比语言音频预训练(CLAP)模型的最新进展之后,我们探索了一种基于CLAP的声音搜索系统(CLAP-UI),该系统不依赖于人类注释。为了评估CLAP-UI的有效性,我们与广泛使用的音效搜索平台BBC Sound Effect Library进行了对比实验。我们的研究通过基于专业声音搜索工作流程的生态有效任务评估用户性能,认知负荷和满意度。我们的研究结果表明,CLAP-UI表现出显着提高生产力和减少挫折,同时保持可比的认知需求。我们还定性分析了参与者的反馈,这为未来人工智能辅助声音搜索系统的设计提供了有价值的观点。
摘要:Locating the right sound effect efficiently is an important yet challenging topic for audio production. Most current sound-searching systems rely on pre-annotated audio labels created by humans, which can be time-consuming to produce and prone to inaccuracies, limiting the efficiency of audio production. Following the recent advancement of contrastive language-audio pre-training (CLAP) models, we explore an alternative CLAP-based sound-searching system (CLAP-UI) that does not rely on human annotations. To evaluate the effectiveness of CLAP-UI, we conducted comparative experiments with a widely used sound effect searching platform, the BBC Sound Effect Library. Our study evaluates user performance, cognitive load, and satisfaction through ecologically valid tasks based on professional sound-searching workflows. Our result shows that CLAP-UI demonstrated significantly enhanced productivity and reduced frustration while maintaining comparable cognitive demands. We also qualitatively analyzed the participants' feedback, which offered valuable perspectives on the design of future AI-assisted sound search systems.
标题: FASEL:利用证据深度学习进行不确定性感知的虚假音频检测
链接:https://arxiv.org/abs/2504.15663
备注:Accepted at ICASSP 2025
摘要:最近,虚假音频检测得到了极大的关注,因为语音合成和语音转换的进步增加了自动说话人验证(ASV)系统对欺骗攻击的脆弱性。这项任务的一个关键挑战是推广模型来检测看不见的分发外(OOD)攻击。虽然现有的方法已经显示出有希望的结果,但由于使用softmax进行分类,它们固有地存在过度自信问题,当遇到不可预测的欺骗尝试时,会产生不可靠的预测。为了解决这个问题,我们提出了一个新的框架,称为假音频检测与证据学习(FADEL)。通过使用Dirichlet分布对类概率进行建模,FADEL将模型的不确定性纳入其预测中,从而在OOD场景中实现更强大的性能。在ASVspoof2019逻辑访问(LA)和ASVspoof2021 LA数据集上的实验结果表明,该方法显著提高了基线模型的性能。此外,我们通过分析不同欺骗算法的平均不确定性和等错误率(EER)之间的强相关性,证明了不确定性估计的有效性。
摘要:Recently, fake audio detection has gained significant attention, as advancements in speech synthesis and voice conversion have increased the vulnerability of automatic speaker verification (ASV) systems to spoofing attacks. A key challenge in this task is generalizing models to detect unseen, out-of-distribution (OOD) attacks. Although existing approaches have shown promising results, they inherently suffer from overconfidence issues due to the usage of softmax for classification, which can produce unreliable predictions when encountering unpredictable spoofing attempts. To deal with this limitation, we propose a novel framework called fake audio detection with evidential learning (FADEL). By modeling class probabilities with a Dirichlet distribution, FADEL incorporates model uncertainty into its predictions, thereby leading to more robust performance in OOD scenarios. Experimental results on the ASVspoof2019 Logical Access (LA) and ASVspoof2021 LA datasets indicate that the proposed method significantly improves the performance of baseline models. Furthermore, we demonstrate the validity of uncertainty estimation by analyzing a strong correlation between average uncertainty and equal error rate (EER) across different spoofing algorithms.
【2】 Exploring the User Experience of AI-Assisted Sound Searching Systems for Creative Workflows
标题: 探索创意工作流程的人工智能辅助声音搜索系统的用户体验链接:https://arxiv.org/abs/2504.15575
摘要:有效地定位正确的声音效果是音频制作的一个重要但具有挑战性的主题。大多数当前的声音搜索系统依赖于由人类创建的预先注释的音频标签,这可能是耗时的,并且容易出现不准确,限制了音频制作的效率。在对比语言音频预训练(CLAP)模型的最新进展之后,我们探索了一种基于CLAP的声音搜索系统(CLAP-UI),该系统不依赖于人类注释。为了评估CLAP-UI的有效性,我们与广泛使用的音效搜索平台BBC Sound Effect Library进行了对比实验。我们的研究通过基于专业声音搜索工作流程的生态有效任务评估用户性能,认知负荷和满意度。我们的研究结果表明,CLAP-UI表现出显着提高生产力和减少挫折,同时保持可比的认知需求。我们还定性分析了参与者的反馈,这为未来人工智能辅助声音搜索系统的设计提供了有价值的观点。
摘要:Locating the right sound effect efficiently is an important yet challenging topic for audio production. Most current sound-searching systems rely on pre-annotated audio labels created by humans, which can be time-consuming to produce and prone to inaccuracies, limiting the efficiency of audio production. Following the recent advancement of contrastive language-audio pre-training (CLAP) models, we explore an alternative CLAP-based sound-searching system (CLAP-UI) that does not rely on human annotations. To evaluate the effectiveness of CLAP-UI, we conducted comparative experiments with a widely used sound effect searching platform, the BBC Sound Effect Library. Our study evaluates user performance, cognitive load, and satisfaction through ecologically valid tasks based on professional sound-searching workflows. Our result shows that CLAP-UI demonstrated significantly enhanced productivity and reduced frustration while maintaining comparable cognitive demands. We also qualitatively analyzed the participants' feedback, which offered valuable perspectives on the design of future AI-assisted sound search systems.
【3】 Quantifying Source Speaker Leakage in One-to-One Voice Conversion
标题: 量化一对一语音转换中的源扬声器泄漏链接:https://arxiv.org/abs/2504.15822
备注:Accepted at IEEE 23rd International Conference of the Biometrics Special Interest Group (BIOSIG 2024)
摘要:使用一个多口音语料库的并行话语与商业语音设备的使用,我们提出了一个案例研究表明,它是可以量化的一对一的语音转换的情况下,一个源扬声器的身份的置信度。在使用HiFi-GAN声码器进行语音转换之后,我们比较了一系列扬声器特征的信息泄漏;假设“最坏情况”的白盒场景,我们量化了我们执行推断的信心,并缩小了可能的源扬声器的范围,加强了合成语音提供商必须确保其扬声器数据隐私的监管义务和道德责任。
摘要:Using a multi-accented corpus of parallel utterances for use with commercial speech devices, we present a case study to show that it is possible to quantify a degree of confidence about a source speaker's identity in the case of one-to-one voice conversion. Following voice conversion using a HiFi-GAN vocoder, we compare information leakage for a range speaker characteristics; assuming a "worst-case" white-box scenario, we quantify our confidence to perform inference and narrow the pool of likely source speakers, reinforcing the regulatory obligation and moral duty that providers of synthetic voices have to ensure the privacy of their speakers' data.
【4】 SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
标题: Sim 2S-LLM:解锁语音LLM的同时推理以实现语音到语音翻译链接:https://arxiv.org/abs/2504.15509
摘要:同步语音翻译(SST)与流式语音输入并行输出翻译,从而平衡翻译质量和延迟。虽然大型语言模型(LLM)已经扩展到处理语音模态,但由于语音被预先作为整个生成过程的提示,因此流仍然具有挑战性。为了解锁LLM流媒体功能,本文提出了SimulS 2S-LLM,离线训练语音LLM,并采用测试时间策略来指导同时推理。SimulS 2S-LLM通过提取边界感知语音提示来消除训练和推理之间的不匹配,从而使其与文本输入数据更好地匹配。SimulS 2S-LLM通过预测离散输出语音令牌,然后使用预先训练的声码器合成输出语音,实现同步语音到语音翻译(Simul-S2 ST)。设计了一种增量波束搜索算法,在不增加时延的情况下扩展了语音标记预测的搜索空间。对CVSS语音数据的实验表明,与使用相同训练数据的现有方法相比,SimulS 2 S-LLM提供了更好的翻译质量-延迟权衡,例如在相似延迟的情况下将ASR-BLEU评分提高3分。
摘要:Simultaneous speech translation (SST) outputs translations in parallel with streaming speech input, balancing translation quality and latency. While large language models (LLMs) have been extended to handle the speech modality, streaming remains challenging as speech is prepended as a prompt for the entire generation process. To unlock LLM streaming capability, this paper proposes SimulS2S-LLM, which trains speech LLMs offline and employs a test-time policy to guide simultaneous inference. SimulS2S-LLM alleviates the mismatch between training and inference by extracting boundary-aware speech prompts that allows it to be better matched with text input data. SimulS2S-LLM achieves simultaneous speech-to-speech translation (Simul-S2ST) by predicting discrete output speech tokens and then synthesising output speech using a pre-trained vocoder. An incremental beam search is designed to expand the search space of speech token prediction without increasing latency. Experiments on the CVSS speech data show that SimulS2S-LLM offers a better translation quality-latency trade-off than existing methods that use the same training data, such as improving ASR-BLEU scores by 3 points at similar latency.
机器翻译由腾讯交互翻译提供,仅供参考
