微信公众号:arXiv_Daily
cs.SD语音
标题: 基于生成对抗网络的语音转换:技术、挑战和最新进展
链接:https://arxiv.org/abs/2504.19197
备注:19 pages, 12 figures, 1 table
摘要:语音转换(VC)是语音合成中的一个重要研究领域,它能够在保持语言内容的同时,将说话人的声音特征转换为另一个说话人的声音特征。该技术具有广泛的应用,包括自动电影配音、语音到歌唱转换以及用于病理性言语康复的辅助设备。随着对高质量和自然声音合成语音的需求不断增加,研究人员开发了各种VC技术。其中,基于生成对抗网络(GAN)的方法因其强大的特征映射能力和产生高度逼真语音的潜力而引起了相当大的关注。尽管取得了显着的进步,但确保训练稳定性,保持语言一致性和实现感知自然性等挑战继续阻碍基于GAN的VC系统的进展。这篇系统性综述对语音转换领域进行了全面分析,重点介绍了关键技术、关键挑战以及GANs在该领域的变革性影响。该调查对现有方法进行了分类,检查了技术障碍,并批判性地评估了基于GAN的VC的最新发展。通过整合和综合分散在文献中的研究结果,这篇综述提供了一个结构化的理解不同方法的优势和局限性。这项调查的重要性在于它能够通过识别现有的差距,提出潜在的方向,并为构建更强大,更高效的风险投资系统提供见解来指导未来的研究。总的来说,这项工作是研究人员、开发人员和从业人员的重要资源,旨在推进语音转换技术的最新发展(SOTA)。
摘要:Voice conversion (VC) stands as a crucial research area in speech synthesis, enabling the transformation of a speaker's vocal characteristics to resemble another while preserving the linguistic content. This technology has broad applications, including automated movie dubbing, speech-to-singing conversion, and assistive devices for pathological speech rehabilitation. With the increasing demand for high-quality and natural-sounding synthetic voices, researchers have developed a wide range of VC techniques. Among these, generative adversarial network (GAN)-based approaches have drawn considerable attention for their powerful feature-mapping capabilities and potential to produce highly realistic speech. Despite notable advancements, challenges such as ensuring training stability, maintaining linguistic consistency, and achieving perceptual naturalness continue to hinder progress in GAN-based VC systems. This systematic review presents a comprehensive analysis of the voice conversion landscape, highlighting key techniques, key challenges, and the transformative impact of GANs in the field. The survey categorizes existing methods, examines technical obstacles, and critically evaluates recent developments in GAN-based VC. By consolidating and synthesizing research findings scattered across the literature, this review provides a structured understanding of the strengths and limitations of different approaches. The significance of this survey lies in its ability to guide future research by identifying existing gaps, proposing potential directions, and offering insights for building more robust and efficient VC systems. Overall, this work serves as an essential resource for researchers, developers, and practitioners aiming to advance the state-of-the-art (SOTA) in voice conversion technology.
【2】 Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget
标题: Muyan-TTC:一种针对播客场景优化的可训练文本到语音模型,预算为5万美元链接:https://arxiv.org/abs/2504.19146
摘要:文本到语音(TTS)模型的最新进展是由大型语言模型(LLM)的集成,增强语义理解和提高语音自然度所驱动的。然而,现有的基于LLM的TTS模型往往缺乏开源的训练代码和有效的推理加速框架,限制了它们的可访问性和适应性。此外,没有公开可用的TTS模型专门针对播客场景进行优化,这对语音交互应用程序有很高的需求。为了解决这些限制,我们引入Muyan-TTS,这是一个开源的可训练的TTS模型,专为播客应用程序设计,预算为50,000美元。我们的模型基于超过100,000小时的播客音频数据进行了预训练,实现了具有高质量语音生成的zero-shot TTS合成。此外,Muyan-TTS支持数十分钟的目标语音的扬声器自适应,使其高度可定制的个人声音。除了开源模型外,我们还提供了一个全面的数据收集和处理管道,一个完整的训练过程,以及一个优化的推理框架,可以加速基于LLM的TTS合成。我们的代码和型号可在https://github.com/MYZY-AI/Muyan-TTS上获得。
摘要:Recent advancements in text-to-speech (TTS) models have been driven by the integration of large language models (LLMs), enhancing semantic comprehension and improving speech naturalness. However, existing LLM-based TTS models often lack open-source training code and efficient inference acceleration frameworks, limiting their accessibility and adaptability. Additionally, there is no publicly available TTS model specifically optimized for podcast scenarios, which are in high demand for voice interaction applications. To address these limitations, we introduce Muyan-TTS, an open-source trainable TTS model designed for podcast applications within a $50,000 budget. Our model is pre-trained on over 100,000 hours of podcast audio data, enabling zero-shot TTS synthesis with high-quality voice generation. Furthermore, Muyan-TTS supports speaker adaptation with dozens of minutes of target speech, making it highly customizable for individual voices. In addition to open-sourcing the model, we provide a comprehensive data collection and processing pipeline, a full training procedure, and an optimized inference framework that accelerates LLM-based TTS synthesis. Our code and models are available at https://github.com/MYZY-AI/Muyan-TTS.
【3】 Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning
标题: 通过迁移学习改进预训练的YAMNet以增强语音命令检测链接:https://arxiv.org/abs/2504.19030
备注:None
摘要:这项工作解决了提高语音命令识别系统的准确性和效率的需要,这是改善各种智能应用程序中用户交互的关键组成部分。利用强大的预训练YAMNet模型和迁移学习,本研究开发了一种显著提高语音命令识别的方法。我们调整并训练YAMNet深度学习模型,以有效地从音频信号中检测和解释语音命令。使用广泛注释的语音命令数据集(speech_commands_v0.01),我们的方法演示了迁移学习的实际应用,以准确识别预定义的语音命令集。数据集经过精心扩充,并策略性地提取特征以提高模型性能。最终模型的识别准确率达到了95.28%,凸显了先进机器学习技术对语音命令识别的影响。这一成果标志着音频处理技术取得了实质性进展,并为该领域未来的研究建立了新的基准。
摘要:This work addresses the need for enhanced accuracy and efficiency in speech command recognition systems, a critical component for improving user interaction in various smart applications. Leveraging the robust pretrained YAMNet model and transfer learning, this study develops a method that significantly improves speech command recognition. We adapt and train a YAMNet deep learning model to effectively detect and interpret speech commands from audio signals. Using the extensively annotated Speech Commands dataset (speech_commands_v0.01), our approach demonstrates the practical application of transfer learning to accurately recognize a predefined set of speech commands. The dataset is meticulously augmented, and features are strategically extracted to boost model performance. As a result, the final model achieved a recognition accuracy of 95.28%, underscoring the impact of advanced machine learning techniques on speech command recognition. This achievement marks substantial progress in audio processing technologies and establishes a new benchmark for future research in the field.
【4】 Speaker Retrieval in the Wild: Challenges, Effectiveness and Robustness
标题: 野外说话人检索:挑战、有效性和稳健性链接:https://arxiv.org/abs/2504.18950
备注:13 pages, 10 figures, 10 tables, 76 references
摘要:公众可获得的或公司拥有的音频/视频档案越来越丰富,突出表明有效获取所需内容和从这些档案中检索信息的重要性日益增加。本文研究的挑战,解决方案,有效性和鲁棒性的扬声器检索系统开发的“在野外”,其中涉及解决两个主要的挑战:从有限的元数据中提取任务相关的标签系统的开发和评估,以及不受约束的声学条件中遇到的档案,从安静的工作室,以不利的嘈杂环境。虽然我们专注于公开的BBC倒带档案(跨越1948年至1979年),我们的框架解决了更广泛的问题,扬声器检索广泛的和可能老化的档案,没有控制的内容和声学条件。通常情况下,这些档案提供了一个简短的和一般的文件描述,大多不足以为特定的应用程序,如扬声器检索,和这样的大规模的档案手动注释是不可行的。我们探讨了系统开发的各个方面(例如,发言人日记,嵌入提取,查询选择),并分析挑战,可能的解决方案,以及它们的功能。为了评估性能,我们在干净的设置和对各种失真模拟现实世界的应用进行系统的实验。我们的研究结果表明,所开发的说话人检索系统的有效性和鲁棒性,建立的多功能性和可扩展性的建议框架的广泛的应用范围以外的BBC倒带语料库。
摘要:There is a growing abundance of publicly available or company-owned audio/video archives, highlighting the increasing importance of efficient access to desired content and information retrieval from these archives. This paper investigates the challenges, solutions, effectiveness, and robustness of speaker retrieval systems developed "in the wild" which involves addressing two primary challenges: extraction of task-relevant labels from limited metadata for system development and evaluation, as well as the unconstrained acoustic conditions encountered in the archive, ranging from quiet studios to adverse noisy environments. While we focus on the publicly-available BBC Rewind archive (spanning 1948 to 1979), our framework addresses the broader issue of speaker retrieval on extensive and possibly aged archives with no control over the content and acoustic conditions. Typically, these archives offer a brief and general file description, mostly inadequate for specific applications like speaker retrieval, and manual annotation of such large-scale archives is unfeasible. We explore various aspects of system development (e.g., speaker diarisation, embedding extraction, query selection) and analyse the challenges, possible solutions, and their functionality. To evaluate the performance, we conduct systematic experiments in both clean setup and against various distortions simulating real-world applications. Our findings demonstrate the effectiveness and robustness of the developed speaker retrieval systems, establishing the versatility and scalability of the proposed framework for a wide range of applications beyond the BBC Rewind corpus.
【5】 A Survey on Multimodal Music Emotion Recognition
标题: 多模式音乐情感识别研究综述链接:https://arxiv.org/abs/2504.18799
摘要:None
摘要:Multimodal music emotion recognition (MMER) is an emerging discipline in music information retrieval that has experienced a surge in interest in recent years. This survey provides a comprehensive overview of the current state-of-the-art in MMER. Discussing the different approaches and techniques used in this field, the paper introduces a four-stage MMER framework, including multimodal data selection, feature extraction, feature processing, and final emotion prediction. The survey further reveals significant advancements in deep learning methods and the increasing importance of feature fusion techniques. Despite these advancements, challenges such as the need for large annotated datasets, datasets with more modalities, and real-time processing capabilities remain. This paper also contributes to the field by identifying critical gaps in current research and suggesting potential directions for future research. The gaps underscore the importance of developing robust, scalable, a interpretable models for MMER, with implications for applications in music recommendation systems, therapeutic tools, and entertainment.
【6】 Spatial Speech Translation: Translating Across Space With Binaural Hearables
标题: 空间语音翻译:使用双耳可听语言跨空间翻译链接:https://arxiv.org/abs/2504.18715
备注:Accepted by CHI2025
摘要:想象一下,在一个拥挤的空间里,人们说着不同的语言,而听觉设备可以将听觉空间转换成你的母语,同时保留所有说话者的空间线索。我们介绍了空间语音翻译,一个新的概念,翻译扬声器在佩戴者的环境,同时保持方向和独特的语音特征的双耳输出中的每个扬声器的听觉。为了实现这一目标,我们解决了盲源分离、定位、实时表达翻译和双耳渲染等多项技术挑战,以保留翻译音频中的扬声器方向,同时在Apple M2芯片上实现实时推理。我们使用原型双耳耳机进行的概念验证评估表明,与现有模型不同,现有模型在干扰的情况下会失败,尽管环境中其他扬声器的干扰很强,但在语言之间进行翻译时,我们的BLEU得分高达22.01。用户研究进一步证实了该系统的有效性,在空间渲染翻译的语音在以前看不见的现实世界的混响环境。退一步说,这项工作标志着将空间感知融入语音翻译的第一步。
摘要:Imagine being in a crowded space where people speak a different language and having hearables that transform the auditory space into your native language, while preserving the spatial cues for all speakers. We introduce spatial speech translation, a novel concept for hearables that translate speakers in the wearer's environment, while maintaining the direction and unique voice characteristics of each speaker in the binaural output. To achieve this, we tackle several technical challenges spanning blind source separation, localization, real-time expressive translation, and binaural rendering to preserve the speaker directions in the translated audio, while achieving real-time inference on the Apple M2 silicon. Our proof-of-concept evaluation with a prototype binaural headset shows that, unlike existing models, which fail in the presence of interference, we achieve a BLEU score of up to 22.01 when translating between languages, despite strong interference from other speakers in the environment. User studies further confirm the system's effectiveness in spatially rendering the translated speech in previously unseen real-world reverberant environments. Taking a step back, this work marks the first step towards integrating spatial perception into speech translation.
【7】 Unsupervised outlier detection to improve bird audio dataset labels
标题: 无监督离群值检测以改进鸟类音频数据集标签链接:https://arxiv.org/abs/2504.18650
备注:27 pages, 9 figures
摘要:Xeno-Canto鸟类音频库对于那些对世界各地鸟类的发声和其他声音感兴趣的人来说是一个非常宝贵的资源。对于试图提高分类模型的鸟类物种识别准确性的机器学习研究人员来说,情况尤其如此。然而,从这个众包存储库中找到的记录中提取标记数据集的任务面临着几个挑战。对于机器学习从业者来说,一个特别重要的挑战是,一个鸟类物种标签被应用于每个音频记录,但经常还捕获其他声音,包括其他鸟类、其他动物声音、人为和其他环境声音。这些非目标鸟类的声音可能会导致数据集标签差异,称为标签噪声。在这项工作中,我们提出了一个清洗过程,包括音频预处理,然后进行降维和无监督离群值检测(UOD),以减少来自Xeno-Canto录音的数据集中的标签噪声。我们研究了三种神经网络降维技术:两种卷积自动编码器和变分深度嵌入(VaDE(Jiang,2017))。虽然这两种方法在检测大多数鸟类物种数据集的离群值方面都显示出一定程度的有效性,但我们发现,从一个物种到下一个物种,方法的性能存在显着差异。我们相信,这项调查的结果表明,我们的清洗过程的应用程序可以有意义地减少来自Xeno-Canto音频库的鸟类物种数据集的标签噪声,但结果因物种而异。
摘要:The Xeno-Canto bird audio repository is an invaluable resource for those interested in vocalizations and other sounds made by birds around the world. This is particularly the case for machine learning researchers attempting to improve on the bird species recognition accuracy of classification models. However, the task of extracting labeled datasets from the recordings found in this crowd-sourced repository faces several challenges. One challenge of particular significance to machine learning practitioners is that one bird species label is applied to each audio recording, but frequently other sounds are also captured including other bird species, other animal sounds, anthropogenic and other ambient sounds. These non-target bird species sounds can result in dataset labeling discrepancies referred to as label noise. In this work we present a cleaning process consisting of audio preprocessing followed by dimensionality reduction and unsupervised outlier detection (UOD) to reduce the label noise in a dataset derived from Xeno-Canto recordings. We investigate three neural network dimensionality reduction techniques: two flavors of convolutional autoencoders and variational deep embedding (VaDE (Jiang, 2017)). While both methods show some degree of effectiveness at detecting outliers for most bird species datasets, we found significant variation in the performance of the methods from one species to the next. We believe that the results of this investigation demonstrate that the application of our cleaning process can meaningfully reduce the label noise of bird species datasets derived from Xeno-Canto audio repository but results vary across species.
【8】 Speaker Diarization for Low-Resource Languages Through Wav2vec Fine-Tuning
标题: 通过Wav2vec微调实现低资源语言的扬声器拨号链接:https://arxiv.org/abs/2504.18582
摘要:说话人日志化是语音处理中的一项基本任务,它涉及到按说话人划分音频流。虽然最先进的模型在高资源语言中具有先进的性能,但由于有限的注释数据,多种方言和频繁的代码切换,库尔德语等低资源语言构成了独特的挑战。在这项研究中,我们通过在专用的库尔德语语料库上训练Wav2Vec 2.0自监督学习模型来解决这些问题。通过利用迁移学习,我们调整了从其他语言中学习到的多语言表示,以捕捉库尔德语的语音和声学特征。相对于基线方法,我们的方法将日志化错误率降低了7.2%,并将聚类纯度提高了13%。这些研究结果表明,对现有模型的增强可以显着提高资源不足的语言的日记化性能。我们的工作对开发库尔德语媒体的转录服务以及多语言呼叫中心,电话会议和视频会议系统中的扬声器分割具有实际意义。研究结果为在其他未充分研究的语言中建立有效的日记系统奠定了基础,有助于提高语音技术的公平性。
摘要:Speaker diarization is a fundamental task in speech processing that involves dividing an audio stream by speaker. Although state-of-the-art models have advanced performance in high-resource languages, low-resource languages such as Kurdish pose unique challenges due to limited annotated data, multiple dialects and frequent code-switching. In this study, we address these issues by training the Wav2Vec 2.0 self-supervised learning model on a dedicated Kurdish corpus. By leveraging transfer learning, we adapted multilingual representations learned from other languages to capture the phonetic and acoustic characteristics of Kurdish speech. Relative to a baseline method, our approach reduced the diarization error rate by seven point two percent and improved cluster purity by thirteen percent. These findings demonstrate that enhancements to existing models can significantly improve diarization performance for under-resourced languages. Our work has practical implications for developing transcription services for Kurdish-language media and for speaker segmentation in multilingual call centers, teleconferencing and video-conferencing systems. The results establish a foundation for building effective diarization systems in other understudied languages, contributing to greater equity in speech technology.
【9】 A Comparative Study on Positional Encoding for Time-frequency Domain Dual-path Transformer-based Source Separation Models
标题: 基于时频域双路径变换器的源分离模型位置编码比较研究链接:https://arxiv.org/abs/2504.19605
备注:5 pages, 3 tables, 2 figures
摘要:在这项研究中,我们调查的位置编码(PE)的源分离性能和推广能力的长序列(长度外推)在基于变换的时间-频率(TF)域双路径模型的影响。TF域双路径模型中的长度外推能力是一个关键因素,因为它不仅影响它们在长持续时间输入上的性能,而且影响它们对具有不可见采样率的信号的可推广性。虽然PE是众所周知的显着影响长度外推,已经有有限的研究,探讨选择PE TF域双路径模型从这个角度来看。为了解决这一差距,我们比较了各种PE方法使用最近的国家的最先进的模型,TF-机车,作为基础架构。我们的分析产生了以下关键发现:(i)当处理与训练期间看到的序列长度相同或更短的序列时,具有PE的模型实现了更好的性能。(ii)然而,没有PE的模型表现出优越的长度外推。当模型包含卷积层时,这种趋势尤其明显。
摘要:In this study, we investigate the impact of positional encoding (PE) on source separation performance and the generalization ability to long sequences (length extrapolation) in Transformer-based time-frequency (TF) domain dual-path models. The length extrapolation capability in TF-domain dual-path models is a crucial factor, as it affects not only their performance on long-duration inputs but also their generalizability to signals with unseen sampling rates. While PE is known to significantly impact length extrapolation, there has been limited research that explores the choice of PEs for TF-domain dual-path models from this perspective. To address this gap, we compare various PE methods using a recent state-of-the-art model, TF-Locoformer, as the base architecture. Our analysis yields the following key findings: (i) When handling sequences that are the same length as or shorter than those seen during training, models with PEs achieve better performance. (ii) However, models without PE exhibit superior length extrapolation. This trend is particularly pronounced when the model contains convolutional layers.
【10】 Versatile Framework for Song Generation with Prompt-based Control
标题: 一种基于控制的多功能歌曲生成框架链接:https://arxiv.org/abs/2504.19062
摘要:歌曲生成专注于根据各种提示生成可控的高质量歌曲。然而,现有的方法难以用基于语音的控制和适当的对齐来生成人声和人声。此外,它们不足以支持各种任务。为了解决这些挑战,我们引入了VersBand,一个多任务歌曲生成框架,用于合成高质量的歌曲,并使用基于文本的控制对齐。VersBand包括这些主要模型:1)VocalBand,一个解耦的模型,利用流匹配方法生成演唱风格,音高和梅尔声谱图,允许快速,高质量的声乐生成与风格控制。2)AccompBand是一种基于流的Transformer模型,它结合了Band-MOE,选择合适的专家来增强质量、对齐和控制。该模型允许生成与人声对齐的可控、高质量的声音。3)两个生成模型,歌词的LyricBand和旋律的MelodyBand,有助于全面的多任务歌曲生成系统,允许基于多个提示进行广泛的控制。实验结果表明,VersBand在使用客观和主观指标的多个歌曲生成任务中的性能优于基线模型。音频样本可在https://VersBand.github.io上获得。
摘要:Song generation focuses on producing controllable high-quality songs based on various prompts. However, existing methods struggle to generate vocals and accompaniments with prompt-based control and proper alignment. Additionally, they fall short in supporting various tasks. To address these challenges, we introduce VersBand, a multi-task song generation framework for synthesizing high-quality, aligned songs with prompt-based control. VersBand comprises these primary models: 1) VocalBand, a decoupled model, leverages the flow-matching method for generating singing styles, pitches, and mel-spectrograms, allowing fast, high-quality vocal generation with style control. 2) AccompBand, a flow-based transformer model, incorporates the Band-MOE, selecting suitable experts for enhanced quality, alignment, and control. This model allows for generating controllable, high-quality accompaniments aligned with vocals. 3) Two generation models, LyricBand for lyrics and MelodyBand for melodies, contribute to the comprehensive multi-task song generation system, allowing for extensive control based on multiple prompts. Experimental results demonstrate that VersBand performs better over baseline models across multiple song generation tasks using objective and subjective metrics. Audio samples are available at https://VersBand.github.io.
【11】 Enhancing Cochlear Implant Signal Coding with Scaled Dot-Product Attention
标题: 基于尺度点积注意的相干植入信号编码增强链接:https://arxiv.org/abs/2504.19046
备注:None
摘要:耳蜗植入体(CI)通过直接用电信号刺激听神经,在恢复重度至极重度感音神经性听力损失患者的听力方面发挥着至关重要的作用。虽然传统的编码策略,如先进的组合编码器(ACE),已被证明是有效的,他们受到其适应性和精度。本文研究了使用深度学习(DL)技术来生成CI的心电图,将我们的模型作为一种先进的替代方案。我们比较了我们的模型与ACE策略的性能,通过使用短时客观可懂度(STOI)度量评估重建的音频信号的可懂度。结果表明,我们的模型达到了0.6031的STOI得分,接近ACE策略的0.6126得分,并提供了潜在的优势,在灵活性和适应性。这项研究强调了将人工智能(AI)纳入CI技术的好处,例如增强个性化和效率。
摘要:Cochlear implants (CIs) play a vital role in restoring hearing for individuals with severe to profound sensorineural hearing loss by directly stimulating the auditory nerve with electrical signals. While traditional coding strategies, such as the advanced combination encoder (ACE), have proven effective, they are constrained by their adaptability and precision. This paper investigates the use of deep learning (DL) techniques to generate electrodograms for CIs, presenting our model as an advanced alternative. We compared the performance of our model with the ACE strategy by evaluating the intelligibility of reconstructed audio signals using the short-time objective intelligibility (STOI) metric. The results indicate that our model achieves a STOI score of 0.6031, closely approximating the 0.6126 score of the ACE strategy, and offers potential advantages in flexibility and adaptability. This study underscores the benefits of incorporating artificial intelligent (AI) into CI technology, such as enhanced personalization and efficiency.
【12】 Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation
标题: 学习鲁棒视听语音表示的多任务损坏预测链接:https://arxiv.org/abs/2504.18539
备注:22 pages, 6 figures, 14 tables
摘要:视听语音识别(AVSR)结合了听觉和视觉模态,以提高识别精度,特别是在嘈杂的环境中,只有音频语音系统是不够的。虽然以前的研究主要是解决音频中断,很少有研究涉及视觉腐败,例如,嘴唇咬合或模糊的视频,这也是有害的。为了解决这一现实世界的挑战,我们提出了CAV2vec,这是一种新的自监督语音表示学习框架,专门用于处理视听联合腐败。CAV2vec采用了一种带有损坏预测任务的自蒸馏方法,其中学生模型学习预测教师模型生成的干净目标,并带有损坏的输入帧。具体来说,我们建议一种单峰多任务学习,它通过预测带有损坏视频的干净音频目标和带有损坏音频的干净视频目标来提取跨模态知识并对齐损坏的模态。该策略减轻了由损坏的模态引起的表示空间中的分散,从而导致更可靠和鲁棒的视听融合。我们在强大的AVSR基准测试上的实验表明,损坏的表示学习方法显着提高了在涉及各种类型的腐败的广义环境中的识别精度。
摘要:Audio-visual speech recognition (AVSR) incorporates auditory and visual modalities to improve recognition accuracy, particularly in noisy environments where audio-only speech systems are insufficient. While previous research has largely addressed audio disruptions, few studies have dealt with visual corruptions, e.g., lip occlusions or blurred videos, which are also detrimental. To address this real-world challenge, we propose CAV2vec, a novel self-supervised speech representation learning framework particularly designed to handle audio-visual joint corruption. CAV2vec employs a self-distillation approach with a corrupted prediction task, where the student model learns to predict clean targets, generated by the teacher model, with corrupted input frames. Specifically, we suggest a unimodal multi-task learning, which distills cross-modal knowledge and aligns the corrupted modalities, by predicting clean audio targets with corrupted videos, and clean video targets with corrupted audios. This strategy mitigates the dispersion in the representation space caused by corrupted modalities, leading to more reliable and robust audio-visual fusion. Our experiments on robust AVSR benchmarks demonstrate that the corrupted representation learning method significantly enhances recognition accuracy across generalized environments involving various types of corruption.
标题: 基于时频域双路径变换器的源分离模型位置编码比较研究
链接:https://arxiv.org/abs/2504.19605
备注:5 pages, 3 tables, 2 figures
摘要:在这项研究中,我们调查的位置编码(PE)的源分离性能和推广能力的长序列(长度外推)在基于变换的时间-频率(TF)域双路径模型的影响。TF域双路径模型中的长度外推能力是一个关键因素,因为它不仅影响它们在长持续时间输入上的性能,而且影响它们对具有不可见采样率的信号的可推广性。虽然PE是众所周知的显着影响长度外推,已经有有限的研究,探讨选择PE TF域双路径模型从这个角度来看。为了解决这一差距,我们比较了各种PE方法使用最近的国家的最先进的模型,TF-机车,作为基础架构。我们的分析产生了以下关键发现:(i)当处理与训练期间看到的序列长度相同或更短的序列时,具有PE的模型实现了更好的性能。(ii)然而,没有PE的模型表现出优越的长度外推。当模型包含卷积层时,这种趋势尤其明显。
摘要:In this study, we investigate the impact of positional encoding (PE) on source separation performance and the generalization ability to long sequences (length extrapolation) in Transformer-based time-frequency (TF) domain dual-path models. The length extrapolation capability in TF-domain dual-path models is a crucial factor, as it affects not only their performance on long-duration inputs but also their generalizability to signals with unseen sampling rates. While PE is known to significantly impact length extrapolation, there has been limited research that explores the choice of PEs for TF-domain dual-path models from this perspective. To address this gap, we compare various PE methods using a recent state-of-the-art model, TF-Locoformer, as the base architecture. Our analysis yields the following key findings: (i) When handling sequences that are the same length as or shorter than those seen during training, models with PEs achieve better performance. (ii) However, models without PE exhibit superior length extrapolation. This trend is particularly pronounced when the model contains convolutional layers.
【2】 Versatile Framework for Song Generation with Prompt-based Control
标题: 一种基于控制的多功能歌曲生成框架链接:https://arxiv.org/abs/2504.19062
摘要:歌曲生成专注于根据各种提示生成可控的高质量歌曲。然而,现有的方法难以用基于语音的控制和适当的对齐来生成人声和人声。此外,它们在支持各种任务方面也有不足。为了解决这些挑战,我们引入了VersBand,一个多任务歌曲生成框架,用于合成高质量的歌曲,并使用基于文本的控制对齐。VersBand包括这些主要模型:1)VocalBand,一个解耦的模型,利用流匹配方法生成演唱风格,音高和梅尔声谱图,允许快速,高质量的声乐生成与风格控制。2)AccompBand是一种基于流的Transformer模型,它结合了Band-MOE,选择合适的专家来增强质量、对齐和控制。该模型允许生成与人声对齐的可控、高质量的声音。3)两个生成模型,歌词的LyricBand和旋律的MelodyBand,有助于全面的多任务歌曲生成系统,允许基于多个提示进行广泛的控制。实验结果表明,VersBand在使用客观和主观指标的多个歌曲生成任务中的性能优于基线模型。音频样本可在https://VersBand.github.io上获得。
摘要:Song generation focuses on producing controllable high-quality songs based on various prompts. However, existing methods struggle to generate vocals and accompaniments with prompt-based control and proper alignment. Additionally, they fall short in supporting various tasks. To address these challenges, we introduce VersBand, a multi-task song generation framework for synthesizing high-quality, aligned songs with prompt-based control. VersBand comprises these primary models: 1) VocalBand, a decoupled model, leverages the flow-matching method for generating singing styles, pitches, and mel-spectrograms, allowing fast, high-quality vocal generation with style control. 2) AccompBand, a flow-based transformer model, incorporates the Band-MOE, selecting suitable experts for enhanced quality, alignment, and control. This model allows for generating controllable, high-quality accompaniments aligned with vocals. 3) Two generation models, LyricBand for lyrics and MelodyBand for melodies, contribute to the comprehensive multi-task song generation system, allowing for extensive control based on multiple prompts. Experimental results demonstrate that VersBand performs better over baseline models across multiple song generation tasks using objective and subjective metrics. Audio samples are available at https://VersBand.github.io.
【3】 Enhancing Cochlear Implant Signal Coding with Scaled Dot-Product Attention
标题: 基于尺度点积注意的相干植入信号编码增强链接:https://arxiv.org/abs/2504.19046
备注:None
摘要:耳蜗植入体(CI)通过直接用电信号刺激听神经,在恢复重度至极重度感音神经性听力损失患者的听力方面发挥着至关重要的作用。虽然传统的编码策略,如先进的组合编码器(ACE),已被证明是有效的,他们受到其适应性和精度。本文研究了使用深度学习(DL)技术来生成CI的心电图,将我们的模型作为一种先进的替代方案。我们比较了我们的模型与ACE策略的性能,通过使用短时客观可懂度(STOI)度量评估重建的音频信号的可懂度。结果表明,我们的模型达到了0.6031的STOI得分,接近ACE策略的0.6126得分,并提供了潜在的优势,在灵活性和适应性。这项研究强调了将人工智能(AI)纳入CI技术的好处,例如增强个性化和效率。
摘要:Cochlear implants (CIs) play a vital role in restoring hearing for individuals with severe to profound sensorineural hearing loss by directly stimulating the auditory nerve with electrical signals. While traditional coding strategies, such as the advanced combination encoder (ACE), have proven effective, they are constrained by their adaptability and precision. This paper investigates the use of deep learning (DL) techniques to generate electrodograms for CIs, presenting our model as an advanced alternative. We compared the performance of our model with the ACE strategy by evaluating the intelligibility of reconstructed audio signals using the short-time objective intelligibility (STOI) metric. The results indicate that our model achieves a STOI score of 0.6031, closely approximating the 0.6126 score of the ACE strategy, and offers potential advantages in flexibility and adaptability. This study underscores the benefits of incorporating artificial intelligent (AI) into CI technology, such as enhanced personalization and efficiency.
【4】 Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation
标题: 学习鲁棒视听语音表示的多任务损坏预测链接:https://arxiv.org/abs/2504.18539
备注:22 pages, 6 figures, 14 tables
摘要:视听语音识别(AVSR)结合了听觉和视觉模态,以提高识别精度,特别是在嘈杂的环境中,只有音频语音系统是不够的。虽然以前的研究主要是解决音频中断,很少有研究涉及视觉腐败,例如,嘴唇咬合或模糊的视频,这也是有害的。为了解决这一现实世界的挑战,我们提出了CAV2vec,这是一种新的自监督语音表示学习框架,专门用于处理视听联合腐败。CAV2vec采用了一种带有损坏预测任务的自蒸馏方法,其中学生模型学习预测教师模型生成的干净目标,并带有损坏的输入帧。具体来说,我们建议一种单峰多任务学习,它通过预测带有损坏视频的干净音频目标和带有损坏音频的干净视频目标来提取跨模态知识并对齐损坏的模态。该策略减轻了由损坏的模态引起的表示空间中的分散,从而导致更可靠和鲁棒的视听融合。我们在强大的AVSR基准测试上的实验表明,损坏的表示学习方法显着提高了在涉及各种类型的腐败的广义环境中的识别精度。
摘要:Audio-visual speech recognition (AVSR) incorporates auditory and visual modalities to improve recognition accuracy, particularly in noisy environments where audio-only speech systems are insufficient. While previous research has largely addressed audio disruptions, few studies have dealt with visual corruptions, e.g., lip occlusions or blurred videos, which are also detrimental. To address this real-world challenge, we propose CAV2vec, a novel self-supervised speech representation learning framework particularly designed to handle audio-visual joint corruption. CAV2vec employs a self-distillation approach with a corrupted prediction task, where the student model learns to predict clean targets, generated by the teacher model, with corrupted input frames. Specifically, we suggest a unimodal multi-task learning, which distills cross-modal knowledge and aligns the corrupted modalities, by predicting clean audio targets with corrupted videos, and clean video targets with corrupted audios. This strategy mitigates the dispersion in the representation space caused by corrupted modalities, leading to more reliable and robust audio-visual fusion. Our experiments on robust AVSR benchmarks demonstrate that the corrupted representation learning method significantly enhances recognition accuracy across generalized environments involving various types of corruption.
【5】 Generative Adversarial Network based Voice Conversion: Techniques, Challenges, and Recent Advancements
标题: 基于生成对抗网络的语音转换:技术、挑战和最新进展链接:https://arxiv.org/abs/2504.19197
备注:19 pages, 12 figures, 1 table
摘要:语音转换(VC)是语音合成中的一个重要研究领域,它能够在保持语言内容的同时,将说话人的声音特征转换为另一个说话人的声音特征。这项技术有着广泛的应用,包括自动电影配音,语音到歌唱转换,以及用于病理性语言康复的辅助设备。随着对高质量和自然声音合成语音的需求不断增加,研究人员开发了各种VC技术。其中,基于生成对抗网络(GAN)的方法因其强大的特征映射能力和产生高度逼真语音的潜力而引起了相当大的关注。尽管取得了显着的进步,但确保训练稳定性、保持语言一致性和实现感知自然性等挑战仍然阻碍着基于GAN的VC系统的进展。这篇系统性综述对语音转换领域进行了全面分析,重点介绍了关键技术、关键挑战以及GANs在该领域的变革性影响。该调查对现有方法进行了分类,检查了技术障碍,并批判性地评估了基于GAN的VC的最新发展。通过整合和综合分散在文献中的研究结果,这篇综述提供了一个结构化的理解不同方法的优势和局限性。这项调查的重要性在于它能够通过识别现有的差距,提出潜在的方向,并为构建更强大,更高效的风险投资系统提供见解来指导未来的研究。总的来说,这项工作是研究人员、开发人员和从业人员的重要资源,旨在推进语音转换技术的最新发展(SOTA)。
摘要:Voice conversion (VC) stands as a crucial research area in speech synthesis, enabling the transformation of a speaker's vocal characteristics to resemble another while preserving the linguistic content. This technology has broad applications, including automated movie dubbing, speech-to-singing conversion, and assistive devices for pathological speech rehabilitation. With the increasing demand for high-quality and natural-sounding synthetic voices, researchers have developed a wide range of VC techniques. Among these, generative adversarial network (GAN)-based approaches have drawn considerable attention for their powerful feature-mapping capabilities and potential to produce highly realistic speech. Despite notable advancements, challenges such as ensuring training stability, maintaining linguistic consistency, and achieving perceptual naturalness continue to hinder progress in GAN-based VC systems. This systematic review presents a comprehensive analysis of the voice conversion landscape, highlighting key techniques, key challenges, and the transformative impact of GANs in the field. The survey categorizes existing methods, examines technical obstacles, and critically evaluates recent developments in GAN-based VC. By consolidating and synthesizing research findings scattered across the literature, this review provides a structured understanding of the strengths and limitations of different approaches. The significance of this survey lies in its ability to guide future research by identifying existing gaps, proposing potential directions, and offering insights for building more robust and efficient VC systems. Overall, this work serves as an essential resource for researchers, developers, and practitioners aiming to advance the state-of-the-art (SOTA) in voice conversion technology.
【6】 Muyan-TTS: A Trainable Text-to-Speech Model Optimized for Podcast Scenarios with a $50K Budget
标题: Muyan-TTC:一种针对播客场景优化的可训练文本到语音模型,预算为5万美元链接:https://arxiv.org/abs/2504.19146
摘要:文本到语音(TTS)模型的最新进展是由大型语言模型(LLM)的集成,增强语义理解和提高语音自然度所驱动的。然而,现有的基于LLM的TTS模型往往缺乏开源的训练代码和有效的推理加速框架,限制了它们的可访问性和适应性。此外,没有公开可用的TTS模型专门针对播客场景进行优化,这对语音交互应用程序有很高的需求。为了解决这些限制,我们引入Muyan-TTS,这是一个开源的可训练的TTS模型,专为播客应用程序设计,预算为50,000美元。我们的模型基于超过100,000小时的播客音频数据进行了预训练,实现了具有高质量语音生成的zero-shot TTS合成。此外,Muyan-TTS支持数十分钟的目标语音的扬声器自适应,使其高度可定制的个人声音。除了开源模型外,我们还提供了一个全面的数据收集和处理管道,一个完整的训练过程,以及一个优化的推理框架,可以加速基于LLM的TTS合成。我们的代码和型号可在https://github.com/MYZY-AI/Muyan-TTS上获得。
摘要:Recent advancements in text-to-speech (TTS) models have been driven by the integration of large language models (LLMs), enhancing semantic comprehension and improving speech naturalness. However, existing LLM-based TTS models often lack open-source training code and efficient inference acceleration frameworks, limiting their accessibility and adaptability. Additionally, there is no publicly available TTS model specifically optimized for podcast scenarios, which are in high demand for voice interaction applications. To address these limitations, we introduce Muyan-TTS, an open-source trainable TTS model designed for podcast applications within a $50,000 budget. Our model is pre-trained on over 100,000 hours of podcast audio data, enabling zero-shot TTS synthesis with high-quality voice generation. Furthermore, Muyan-TTS supports speaker adaptation with dozens of minutes of target speech, making it highly customizable for individual voices. In addition to open-sourcing the model, we provide a comprehensive data collection and processing pipeline, a full training procedure, and an optimized inference framework that accelerates LLM-based TTS synthesis. Our code and models are available at https://github.com/MYZY-AI/Muyan-TTS.
【7】 Improving Pretrained YAMNet for Enhanced Speech Command Detection via Transfer Learning
标题: 通过迁移学习改进预训练的YAMNet以增强语音命令检测链接:https://arxiv.org/abs/2504.19030
备注:None
摘要:这项工作解决了提高语音命令识别系统的准确性和效率的需要,这是改善各种智能应用程序中用户交互的关键组成部分。利用强大的预训练YAMNet模型和迁移学习,本研究开发了一种显著提高语音命令识别的方法。我们调整并训练YAMNet深度学习模型,以有效地从音频信号中检测和解释语音命令。使用广泛注释的语音命令数据集(speech_commands_v0.01),我们的方法演示了迁移学习的实际应用,以准确识别预定义的语音命令集。数据集经过精心扩充,并策略性地提取特征以提高模型性能。最终模型的识别准确率达到了95.28%,凸显了先进机器学习技术对语音命令识别的影响。这一成就标志着音频处理技术的重大进展,并为该领域的未来研究建立了新的基准。
摘要:This work addresses the need for enhanced accuracy and efficiency in speech command recognition systems, a critical component for improving user interaction in various smart applications. Leveraging the robust pretrained YAMNet model and transfer learning, this study develops a method that significantly improves speech command recognition. We adapt and train a YAMNet deep learning model to effectively detect and interpret speech commands from audio signals. Using the extensively annotated Speech Commands dataset (speech_commands_v0.01), our approach demonstrates the practical application of transfer learning to accurately recognize a predefined set of speech commands. The dataset is meticulously augmented, and features are strategically extracted to boost model performance. As a result, the final model achieved a recognition accuracy of 95.28%, underscoring the impact of advanced machine learning techniques on speech command recognition. This achievement marks substantial progress in audio processing technologies and establishes a new benchmark for future research in the field.
【8】 Speaker Retrieval in the Wild: Challenges, Effectiveness and Robustness
标题: 野外说话人检索:挑战、有效性和稳健性链接:https://arxiv.org/abs/2504.18950
备注:13 pages, 10 figures, 10 tables, 76 references
摘要:None
摘要:There is a growing abundance of publicly available or company-owned audio/video archives, highlighting the increasing importance of efficient access to desired content and information retrieval from these archives. This paper investigates the challenges, solutions, effectiveness, and robustness of speaker retrieval systems developed "in the wild" which involves addressing two primary challenges: extraction of task-relevant labels from limited metadata for system development and evaluation, as well as the unconstrained acoustic conditions encountered in the archive, ranging from quiet studios to adverse noisy environments. While we focus on the publicly-available BBC Rewind archive (spanning 1948 to 1979), our framework addresses the broader issue of speaker retrieval on extensive and possibly aged archives with no control over the content and acoustic conditions. Typically, these archives offer a brief and general file description, mostly inadequate for specific applications like speaker retrieval, and manual annotation of such large-scale archives is unfeasible. We explore various aspects of system development (e.g., speaker diarisation, embedding extraction, query selection) and analyse the challenges, possible solutions, and their functionality. To evaluate the performance, we conduct systematic experiments in both clean setup and against various distortions simulating real-world applications. Our findings demonstrate the effectiveness and robustness of the developed speaker retrieval systems, establishing the versatility and scalability of the proposed framework for a wide range of applications beyond the BBC Rewind corpus.
【9】 A Survey on Multimodal Music Emotion Recognition
标题: 多模式音乐情感识别研究综述链接:https://arxiv.org/abs/2504.18799
摘要:多模态音乐情感识别(MMER)是音乐信息检索领域的一门新兴学科,近年来受到了广泛的关注。该调查提供了MMER当前最新技术水平的全面概述。讨论了在这一领域中使用的不同方法和技术,本文介绍了一个四阶段的MMER框架,包括多模态数据选择,特征提取,特征处理和最终的情感预测。该调查进一步揭示了深度学习方法的重大进步以及特征融合技术的重要性日益增加。尽管取得了这些进步,但仍然存在挑战,例如需要大型注释数据集,具有更多模态的数据集以及实时处理能力。本文还通过确定当前研究中的关键差距并为未来研究提出潜在方向来为该领域做出贡献。这些差距强调了开发强大的,可扩展的,可解释的MMER模型的重要性,并对音乐推荐系统,治疗工具和娱乐应用产生影响。
摘要:Multimodal music emotion recognition (MMER) is an emerging discipline in music information retrieval that has experienced a surge in interest in recent years. This survey provides a comprehensive overview of the current state-of-the-art in MMER. Discussing the different approaches and techniques used in this field, the paper introduces a four-stage MMER framework, including multimodal data selection, feature extraction, feature processing, and final emotion prediction. The survey further reveals significant advancements in deep learning methods and the increasing importance of feature fusion techniques. Despite these advancements, challenges such as the need for large annotated datasets, datasets with more modalities, and real-time processing capabilities remain. This paper also contributes to the field by identifying critical gaps in current research and suggesting potential directions for future research. The gaps underscore the importance of developing robust, scalable, a interpretable models for MMER, with implications for applications in music recommendation systems, therapeutic tools, and entertainment.
【10】 Spatial Speech Translation: Translating Across Space With Binaural Hearables
标题: 空间语音翻译:使用双耳可听语言跨空间翻译链接:https://arxiv.org/abs/2504.18715
备注:Accepted by CHI2025
摘要:想象一下,在一个拥挤的空间里,人们说着不同的语言,而听觉设备可以将听觉空间转换成你的母语,同时保留所有说话者的空间线索。我们介绍了空间语音翻译,一个新的概念,翻译扬声器在佩戴者的环境,同时保持方向和独特的语音特征的双耳输出中的每个扬声器的听觉。为了实现这一目标,我们解决了盲源分离、定位、实时表达翻译和双耳渲染等多项技术挑战,以保留翻译音频中的扬声器方向,同时在Apple M2芯片上实现实时推理。我们使用原型双耳耳机进行的概念验证评估表明,与现有模型不同,现有模型在干扰的情况下会失败,尽管环境中其他扬声器的干扰很强,但在语言之间进行翻译时,我们的BLEU得分高达22.01。用户研究进一步证实了该系统的有效性,在空间渲染翻译的语音在以前看不见的现实世界的混响环境。退一步说,这项工作标志着将空间感知融入语音翻译的第一步。
摘要:Imagine being in a crowded space where people speak a different language and having hearables that transform the auditory space into your native language, while preserving the spatial cues for all speakers. We introduce spatial speech translation, a novel concept for hearables that translate speakers in the wearer's environment, while maintaining the direction and unique voice characteristics of each speaker in the binaural output. To achieve this, we tackle several technical challenges spanning blind source separation, localization, real-time expressive translation, and binaural rendering to preserve the speaker directions in the translated audio, while achieving real-time inference on the Apple M2 silicon. Our proof-of-concept evaluation with a prototype binaural headset shows that, unlike existing models, which fail in the presence of interference, we achieve a BLEU score of up to 22.01 when translating between languages, despite strong interference from other speakers in the environment. User studies further confirm the system's effectiveness in spatially rendering the translated speech in previously unseen real-world reverberant environments. Taking a step back, this work marks the first step towards integrating spatial perception into speech translation.
【11】 Unsupervised outlier detection to improve bird audio dataset labels
标题: 无监督离群值检测以改进鸟类音频数据集标签链接:https://arxiv.org/abs/2504.18650
备注:27 pages, 9 figures
摘要:None
摘要:The Xeno-Canto bird audio repository is an invaluable resource for those interested in vocalizations and other sounds made by birds around the world. This is particularly the case for machine learning researchers attempting to improve on the bird species recognition accuracy of classification models. However, the task of extracting labeled datasets from the recordings found in this crowd-sourced repository faces several challenges. One challenge of particular significance to machine learning practitioners is that one bird species label is applied to each audio recording, but frequently other sounds are also captured including other bird species, other animal sounds, anthropogenic and other ambient sounds. These non-target bird species sounds can result in dataset labeling discrepancies referred to as label noise. In this work we present a cleaning process consisting of audio preprocessing followed by dimensionality reduction and unsupervised outlier detection (UOD) to reduce the label noise in a dataset derived from Xeno-Canto recordings. We investigate three neural network dimensionality reduction techniques: two flavors of convolutional autoencoders and variational deep embedding (VaDE (Jiang, 2017)). While both methods show some degree of effectiveness at detecting outliers for most bird species datasets, we found significant variation in the performance of the methods from one species to the next. We believe that the results of this investigation demonstrate that the application of our cleaning process can meaningfully reduce the label noise of bird species datasets derived from Xeno-Canto audio repository but results vary across species.
【12】 Speaker Diarization for Low-Resource Languages Through Wav2vec Fine-Tuning
标题: 通过Wav2vec微调实现低资源语言的扬声器拨号链接:https://arxiv.org/abs/2504.18582
摘要:说话人日志化是语音处理中的一项基本任务,它涉及到按说话人划分音频流。虽然最先进的模型在高资源语言中具有先进的性能,但由于有限的注释数据,多种方言和频繁的代码切换,库尔德语等低资源语言构成了独特的挑战。在这项研究中,我们通过在专用的库尔德语语料库上训练Wav2Vec 2.0自监督学习模型来解决这些问题。通过利用迁移学习,我们调整了从其他语言中学习到的多语言表示,以捕捉库尔德语的语音和声学特征。相对于基线方法,我们的方法将日志化错误率降低了7.2%,并将聚类纯度提高了13%。这些发现表明,对现有模型的增强可以显着提高资源不足语言的日记化性能。我们的工作对开发库尔德语媒体的转录服务以及多语言呼叫中心,电话会议和视频会议系统中的扬声器分割具有实际意义。研究结果为在其他未充分研究的语言中建立有效的日记系统奠定了基础,有助于提高语音技术的公平性。
摘要:Speaker diarization is a fundamental task in speech processing that involves dividing an audio stream by speaker. Although state-of-the-art models have advanced performance in high-resource languages, low-resource languages such as Kurdish pose unique challenges due to limited annotated data, multiple dialects and frequent code-switching. In this study, we address these issues by training the Wav2Vec 2.0 self-supervised learning model on a dedicated Kurdish corpus. By leveraging transfer learning, we adapted multilingual representations learned from other languages to capture the phonetic and acoustic characteristics of Kurdish speech. Relative to a baseline method, our approach reduced the diarization error rate by seven point two percent and improved cluster purity by thirteen percent. These findings demonstrate that enhancements to existing models can significantly improve diarization performance for under-resourced languages. Our work has practical implications for developing transcription services for Kurdish-language media and for speaker segmentation in multilingual call centers, teleconferencing and video-conferencing systems. The results establish a foundation for building effective diarization systems in other understudied languages, contributing to greater equity in speech technology.
机器翻译由腾讯交互翻译提供,仅供参考
