微信公众号:arXiv_Daily
1.cs.SD语音:
【1】 Empowering Global Voices: A Data-Efficient, Phoneme-Tone Adaptive Approach to High-Fidelity Speech Synthesis
标题: 增强全球之声:一种数据高效、音素音调自适应的高保真语音合成方法
链接:https://arxiv.org/abs/2504.07858
摘要:文语转换(TTS)技术在广泛使用的语言方面取得了令人印象深刻的成果,但许多资源不足的语言仍然面临着数据有限和语言复杂性的挑战。在本文中,我们提出了一种新的方法,将数据优化框架与先进的声学模型相结合,为低资源场景构建高质量的TTS系统。我们证明了我们的方法的有效性,泰国作为一个说明性的情况下,复杂的语音规则和稀疏的资源得到有效解决。我们的方法能够实现zero-shot语音克隆,并提高了从金融到医疗保健、教育和法律等各种客户端应用程序的性能。广泛的评估-主观和客观-证实我们的模型符合最先进的标准,为数据有限的环境中的TTS生产提供了可扩展的解决方案,对更广泛的行业采用和多语言可访问性具有重要意义。
摘要:Text-to-speech (TTS) technology has achieved impressive results for widelyspoken languages, yet many under-resourced languages remain challenged bylimited data and linguistic complexities. In this paper, we present a novelmethodology that integrates a data-optimized framework with an advancedacoustic model to build high-quality TTS systems for low-resource scenarios. Wedemonstrate the effectiveness of our approach using Thai as an illustrativecase, where intricate phonetic rules and sparse resources are effectivelyaddressed. Our method enables zero-shot voice cloning and improved performanceacross diverse client applications, ranging from finance to healthcare,education, and law. Extensive evaluations - both subjective and objective -confirm that our model meets state-of-the-art standards, offering a scalablesolution for TTS production in data-limited settings, with significantimplications for broader industry adoption and multilingual accessibility.
【2】 SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow
标题: SlimSpeech:轻量级、高效的文本到语音,具有纤薄的纠正流
链接:https://arxiv.org/abs/2504.07776
摘要:近年来,基于流匹配的语音合成技术在减少推理步骤的同时,显著地提高了合成语音的质量。在本文中,我们介绍了SlimSpeech,一个轻量级的,高效的语音合成系统的基础上整流。我们已经建立在现有的语音合成方法,利用整流模型,修改其结构,以减少参数,并作为一个教师模型。通过改进回流操作,我们直接从较大的模型中导出具有更直的采样轨迹的较小模型,同时利用蒸馏技术进一步提高模型性能。实验结果表明,我们提出的方法,与显着减少的模型参数,通过一步采样,实现了较大的模型相当的性能。
摘要:Recently, flow matching based speech synthesis has significantly enhanced thequality of synthesized speech while reducing the number of inference steps. Inthis paper, we introduce SlimSpeech, a lightweight and efficient speechsynthesis system based on rectified flow. We have built upon the existingspeech synthesis method utilizing the rectified flow model, modifying itsstructure to reduce parameters and serve as a teacher model. By refining thereflow operation, we directly derive a smaller model with a more straightsampling trajectory from the larger model, while utilizing distillationtechniques to further enhance the model performance. Experimental resultsdemonstrate that our proposed method, with significantly reduced modelparameters, achieves comparable performance to larger models through one-stepsampling.
【3】 Towards Generalizability to Tone and Content Variations in the Transcription of Amplifier Rendered Electric Guitar Audio
标题: 放大器渲染电吉他音频转录中音调和内容变化的通用性
链接:https://arxiv.org/abs/2504.07406
摘要:转录电吉他录音是具有挑战性的,因为缺乏不同的数据集和复杂的音调相关的变化,由放大器,橱柜和效果踏板。为了解决这些问题,我们引入了EGDB-PG,这是一种新颖的数据集,旨在捕获各种不同的调音台机柜配置中的各种音调相关特性。此外,我们提出了音调通知Transformer(TIT),一个基于transformer的转录模型增强了音调嵌入机制,利用学习表示,以提高模型的适应性音调相关的细微差别。实验表明,在EGDB-PG上训练的TIT在不同放大器类型上的表现优于现有基线,数据集的多样性和音调嵌入技术推动了转录准确性的提高。通过详细的基准测试和消融研究,我们评估了音调增强,内容增强,音频标准化和音调嵌入对转录性能的影响。这项工作通过克服数据集多样性和音调建模的限制来推进电吉他转录,为未来的研究提供了坚实的基础。
摘要:Transcribing electric guitar recordings is challenging due to the scarcity ofdiverse datasets and the complex tone-related variations introduced byamplifiers, cabinets, and effect pedals. To address these issues, we introduceEGDB-PG, a novel dataset designed to capture a wide range of tone-relatedcharacteristics across various amplifier-cabinet configurations. In addition,we propose the Tone-informed Transformer (TIT), a Transformer-basedtranscription model enhanced with a tone embedding mechanism that leverageslearned representations to improve the model's adaptability to tone-relatednuances. Experiments demonstrate that TIT, trained on EGDB-PG, outperformsexisting baselines across diverse amplifier types, with transcription accuracyimprovements driven by the dataset's diversity and the tone embeddingtechnique. Through detailed benchmarking and ablation studies, we evaluate theimpact of tone augmentation, content augmentation, audio normalization, andtone embedding on transcription performance. This work advances electric guitartranscription by overcoming limitations in dataset diversity and tone modeling,providing a robust foundation for future research.
【4】 Quantum-Inspired Genetic Algorithm for Robust Source Separation in Smart City Acoustics
标题: 量子启发的遗传算法用于智能城市声学中的鲁棒源分离
链接:https://arxiv.org/abs/2504.07345
备注:6 pages, 2 figures, IEEE International Conference on Communications (ICC 2025)
摘要:城市声音的不和谐对依赖精确声学场景分析的智能城市应用提出了重大挑战。有效分析这些复杂的音景,通常以重叠的声源,不同的声学事件和不可预测的噪声水平为特征,需要精确的源分离。当只有有限的训练数据可用时,这个任务变得更加复杂。本文介绍了一种新的量子启发遗传算法(p-QIGA)用于源分离,从量子信息理论中汲取灵感,以增强智能城市中的声学场景分析。通过利用量子叠加进行有效的解空间探索和纠缠来处理相关源,p-QIGA即使在有限的数据下也能实现稳健的分离。这些量子启发的概念被集成到遗传算法框架中,以优化源分离参数。我们的方法的有效性在两个数据集上得到了证明:TAU Urban Acoustic Scenes 2020 Mobile数据集,代表典型的城市音景,以及Silent Cities数据集,捕捉COVID-19大流行期间更安静的城市环境。实验结果表明,p-QIGA实现了与最先进的方法相当的精度,同时对噪声和有限的训练数据表现出卓越的弹性,在噪声环境中实现高达8.2 dB的信号失真比(SDR),并且仅用10%的训练数据就比基线方法高出2 dB。这项研究强调了p-QIGA在智能城市中推进声学信号处理的潜力,特别是在噪声污染监测和声学监控方面。
摘要:The cacophony of urban sounds presents a significant challenge for smart cityapplications that rely on accurate acoustic scene analysis. Effectivelyanalyzing these complex soundscapes, often characterized by overlapping soundsources, diverse acoustic events, and unpredictable noise levels, requiresprecise source separation. This task becomes more complicated when only limitedtraining data is available. This paper introduces a novel Quantum-InspiredGenetic Algorithm (p-QIGA) for source separation, drawing inspiration fromquantum information theory to enhance acoustic scene analysis in smart cities.By leveraging quantum superposition for efficient solution space explorationand entanglement to handle correlated sources, p-QIGA achieves robustseparation even with limited data. These quantum-inspired concepts areintegrated into a genetic algorithm framework to optimize source separationparameters. The effectiveness of our approach is demonstrated on two datasets:the TAU Urban Acoustic Scenes 2020 Mobile dataset, representing typical urbansoundscapes, and the Silent Cities dataset, capturing quieter urbanenvironments during the COVID-19 pandemic. Experimental results show that thep-QIGA achieves accuracy comparable to state-of-the-art methods whileexhibiting superior resilience to noise and limited training data, achieving upto 8.2 dB signal-to-distortion ratio (SDR) in noisy environments andoutperforming baseline methods by up to 2 dB with only 10% of the trainingdata. This research highlights the potential of p-QIGA to advance acousticsignal processing in smart cities, particularly for noise pollution monitoringand acoustic surveillance.
【5】 Artificial intelligence in creating, representing or expressing an immersive soundscape
标题: 人工智能创建、表示或表达沉浸式音景
链接:https://arxiv.org/abs/2504.07153
备注:Internoise 2024: 53rd International Congress and Exposition on Noise Control Engineering, The International Institute of Noise Control Engineering (I-INCE); Soci{\'e}t{\'e} Fran{\c c}aise d'Acoustique (SFA), Aug 2024, Nantes, France, Aug 2024, Nantes, France
摘要:在当今技术驱动的世界中,人工智能和虚拟现实已经取得了重大进展。这些发展推动研究探索它们在声景领域的交叉点。这些技术不仅提出了关于它们将如何彻底改变我们设计和创建音景的方式的问题,而且还对它们对人类感知,理解和听觉环境表达的影响进行了重要的调查。本文旨在回顾和讨论人工智能在这一领域的最新应用。它探讨了如何利用人工智能来创建虚拟现实沉浸式音景,利用其识别各种形式数据中复杂模式的能力。这包括在不同的模态(如文本、声音和动画)之间进行翻译,以及预测和生成跨这些领域的数据。它解决了围绕人工智能预测,检测和理解声景数据的能力的问题,最终旨在弥合声音和其他形式的人类可读数据之间的差距。1.
摘要:In today's tech-driven world, significant advancements in artificialintelligence and virtual reality have emerged. These developments driveresearch into exploring their intersection in the realm of soundscape. Not onlydo these technologies raise questions about how they will revolutionize the waywe design and create soundscapes, but they also draw significant inquiries intotheir impact on human perception, understanding, and expression of auditoryenvironments. This paper aims to review and discuss the latest applications ofartificial intelligence in this domain. It explores how artificial intelligencecan be utilized to create a virtual reality immersive soundscape, exploitingits ability to recognize complex patterns in various forms of data. Thisincludes translating between different modalities such as text, sounds, andanimations as well as predicting and generating data across these domains. Itaddresses questions surrounding artificial intelligence's capacity to predict,detect, and comprehend soundscape data, ultimately aiming to bridge the gapbetween sound and other forms of human-readable data. 1.
2.eess.AS音频处理:
【1】 Categorical Unsupervised Variational Acoustic Clustering
标题: 类别无监督变分声学聚集
链接:https://arxiv.org/abs/2504.07652
摘要:我们提出了一种分类的方法,在时间-频率域的音频数据的无监督变分声学聚类。分类分布的考虑强制执行更清晰的聚类,即使数据点在时间和频率上强烈重叠,这是大多数城市声学场景数据集的情况。为此,我们使用Gumbel-Softmax分布作为分类分布的软近似,允许通过反向传播进行训练。在此设置中,softmax温度作为优化群集性能的主要机制。结果表明,该模型可以获得令人印象深刻的聚类性能的所有考虑的数据集,即使当数据点强烈重叠的时间和频率。
摘要:We propose a categorical approach for unsupervised variational acousticclustering of audio data in the time-frequency domain. The consideration of acategorical distribution enforces sharper clustering even when data pointsstrongly overlap in time and frequency, which is the case for most datasets ofurban acoustic scenes. To this end, we use a Gumbel-Softmax distribution as asoft approximation to the categorical distribution, allowing for training viabackpropagation. In this settings, the softmax temperature serves as the mainmechanism to tune clustering performance. The results show that the proposedmodel can obtain impressive clustering performance for all considered datasets,even when data points strongly overlap in time and frequency.
【2】 Towards Generalizability to Tone and Content Variations in the Transcription of Amplifier Rendered Electric Guitar Audio
标题: 放大器渲染电吉他音频转录中音调和内容变化的通用性
链接:https://arxiv.org/abs/2504.07406
摘要:转录电吉他录音是具有挑战性的,因为缺乏不同的数据集和复杂的音调相关的变化,由放大器,橱柜和效果踏板。为了解决这些问题,我们引入了EGDB-PG,这是一种新颖的数据集,旨在捕获各种不同的调音台机柜配置中的各种音调相关特性。此外,我们提出了音调通知Transformer(TIT),一个基于transformer的转录模型增强了音调嵌入机制,利用学习表示,以提高模型的适应性音调相关的细微差别。实验表明,在EGDB-PG上训练的TIT在不同放大器类型上的表现优于现有基线,数据集的多样性和音调嵌入技术推动了转录准确性的提高。通过详细的基准测试和消融研究,我们评估了音调增强,内容增强,音频标准化和音调嵌入对转录性能的影响。这项工作通过克服数据集多样性和音调建模的限制来推进电吉他转录,为未来的研究提供了坚实的基础。
摘要:Transcribing electric guitar recordings is challenging due to the scarcity ofdiverse datasets and the complex tone-related variations introduced byamplifiers, cabinets, and effect pedals. To address these issues, we introduceEGDB-PG, a novel dataset designed to capture a wide range of tone-relatedcharacteristics across various amplifier-cabinet configurations. In addition,we propose the Tone-informed Transformer (TIT), a Transformer-basedtranscription model enhanced with a tone embedding mechanism that leverageslearned representations to improve the model's adaptability to tone-relatednuances. Experiments demonstrate that TIT, trained on EGDB-PG, outperformsexisting baselines across diverse amplifier types, with transcription accuracyimprovements driven by the dataset's diversity and the tone embeddingtechnique. Through detailed benchmarking and ablation studies, we evaluate theimpact of tone augmentation, content augmentation, audio normalization, andtone embedding on transcription performance. This work advances electric guitartranscription by overcoming limitations in dataset diversity and tone modeling,providing a robust foundation for future research.
【3】 Quantum-Inspired Genetic Algorithm for Robust Source Separation in Smart City Acoustics
标题: 量子启发的遗传算法用于智能城市声学中的鲁棒源分离
链接:https://arxiv.org/abs/2504.07345
备注:6 pages, 2 figures, IEEE International Conference on Communications (ICC 2025)
摘要:None
摘要:The cacophony of urban sounds presents a significant challenge for smart cityapplications that rely on accurate acoustic scene analysis. Effectivelyanalyzing these complex soundscapes, often characterized by overlapping soundsources, diverse acoustic events, and unpredictable noise levels, requiresprecise source separation. This task becomes more complicated when only limitedtraining data is available. This paper introduces a novel Quantum-InspiredGenetic Algorithm (p-QIGA) for source separation, drawing inspiration fromquantum information theory to enhance acoustic scene analysis in smart cities.By leveraging quantum superposition for efficient solution space explorationand entanglement to handle correlated sources, p-QIGA achieves robustseparation even with limited data. These quantum-inspired concepts areintegrated into a genetic algorithm framework to optimize source separationparameters. The effectiveness of our approach is demonstrated on two datasets:the TAU Urban Acoustic Scenes 2020 Mobile dataset, representing typical urbansoundscapes, and the Silent Cities dataset, capturing quieter urbanenvironments during the COVID-19 pandemic. Experimental results show that thep-QIGA achieves accuracy comparable to state-of-the-art methods whileexhibiting superior resilience to noise and limited training data, achieving upto 8.2 dB signal-to-distortion ratio (SDR) in noisy environments andoutperforming baseline methods by up to 2 dB with only 10% of the trainingdata. This research highlights the potential of p-QIGA to advance acousticsignal processing in smart cities, particularly for noise pollution monitoringand acoustic surveillance.
【4】 Visual-Aware Speech Recognition for Noisy Scenarios
标题: 针对噪音场景的视觉感知语音识别
链接:https://arxiv.org/abs/2504.07229
摘要:人类有能力利用视觉线索,如嘴唇运动和视觉场景,以增强听觉感知,特别是在嘈杂的环境中。然而,目前的自动语音识别(ASR)或视听语音识别(AVSR)模型往往在嘈杂的场景中挣扎。为了解决这个问题,我们提出了一个模型,通过将噪声源与视觉线索相关联来改善转录。与依赖于嘴唇运动,需要扬声器的可见性的作品,我们利用更广泛的视觉信息的环境。这使我们的模型能够自然地从噪声中过滤语音并改善转录,就像人类在嘈杂的场景中所做的那样。我们的方法重新使用预先训练的语音和视觉编码器,将它们与多头注意力联系起来。这种方法使得语音的转录和视频输入中的噪声标签的预测成为可能。我们引入了一个可扩展的管道来开发视听数据集,其中视觉线索与音频中的噪声相关。我们显示出显着的改进,在嘈杂的情况下,现有的仅音频模型。结果还强调,视觉线索在提高转录准确性方面发挥着至关重要的作用。
摘要:Humans have the ability to utilize visual cues, such as lip movements andvisual scenes, to enhance auditory perception, particularly in noisyenvironments. However, current Automatic Speech Recognition (ASR) orAudio-Visual Speech Recognition (AVSR) models often struggle in noisyscenarios. To solve this task, we propose a model that improves transcriptionby correlating noise sources to visual cues. Unlike works that rely on lipmotion and require the speaker's visibility, we exploit broader visualinformation from the environment. This allows our model to naturally filterspeech from noise and improve transcription, much like humans do in noisyscenarios. Our method re-purposes pretrained speech and visual encoders,linking them with multi-headed attention. This approach enables thetranscription of speech and the prediction of noise labels in video inputs. Weintroduce a scalable pipeline to develop audio-visual datasets, where visualcues correlate to noise in the audio. We show significant improvements overexisting audio-only models in noisy scenarios. Results also highlight thatvisual cues play a vital role in improved transcription accuracy.
机器翻译由腾讯交互翻译提供,仅供参考
