微信公众号:arXiv_Daily
cs.SD语音
标题:ASR模型在低资源语言上的适应性--Whisper和Wav 2 Vec-BERT在孟加拉语上的比较研究
链接:https://arxiv.org/abs/2507.01931
摘要:近年来,在大型多语言文本和语音数据集上训练的神经模型在支持低资源语言方面表现出了巨大的潜力。这项研究调查了两种最先进的自动语音识别(ASR)模型,OpenAI的Whisper(Small & Large-V2)和Facebook的Wav 2 Vec-BERT在孟加拉语(一种低资源语言)上的性能。我们使用两个公开的数据集进行了实验:Mozilla Common Voice-17和OpenSLR来评估模型性能。通过系统的微调和超参数优化,包括学习率,epoch和模型检查点选择,我们比较了基于单词错误率(WER),字符错误率(CER),训练时间和计算效率的模型。Wav 2 Vec-BERT模型在所有关键评估指标上都优于Whisper,表现出卓越的性能,同时需要更少的计算资源,并为在低资源语言环境中开发强大的语音识别系统提供了宝贵的见解。
摘要:In recent years, neural models trained on large multilingual text and speech datasets have shown great potential for supporting low-resource languages. This study investigates the performances of two state-of-the-art Automatic Speech Recognition (ASR) models, OpenAI's Whisper (Small & Large-V2) and Facebook's Wav2Vec-BERT on Bangla, a low-resource language. We have conducted experiments using two publicly available datasets: Mozilla Common Voice-17 and OpenSLR to evaluate model performances. Through systematic fine-tuning and hyperparameter optimization, including learning rate, epochs, and model checkpoint selection, we have compared the models based on Word Error Rate (WER), Character Error Rate (CER), Training Time, and Computational Efficiency. The Wav2Vec-BERT model outperformed Whisper across all key evaluation metrics, demonstrated superior performance while requiring fewer computational resources, and offered valuable insights to develop robust speech recognition systems in low-resource linguistic settings.
【2】A Dataset for Automatic Assessment of TTS Quality in Spanish
链接:https://arxiv.org/abs/2507.01805
备注:5 pages, 2 figures. Accepted at Interspeech 2025
摘要:这项工作解决了文本到语音(TTS)系统在西班牙语的自动评估的数据库的开发,旨在提高自然度预测模型的准确性。该数据集由来自52个不同TTS系统和人类声音的4,326个音频样本组成,据我们所知,这是西班牙语中的第一个。为了给音频贴上标签,根据ITU-T Rec.P.807标准设计了一个主观测试,由92名参与者完成。此外,通过训练自动自然度预测系统,验证了所收集数据集的实用性。我们探索了两种方法:微调最初为英语训练的现有模型,并在冻结的自监督语音模型上训练小型下游网络。我们的模型在五点MOS尺度上实现了0.8的平均绝对误差。进一步的分析表明,开发的数据集的质量和多样性,其潜力,以推进TTS研究在西班牙语。
摘要:This work addresses the development of a database for the automatic assessment of text-to-speech (TTS) systems in Spanish, aiming to improve the accuracy of naturalness prediction models. The dataset consists of 4,326 audio samples from 52 different TTS systems and human voices and is, up to our knowledge, the first of its kind in Spanish. To label the audios, a subjective test was designed based on the ITU-T Rec. P.807 standard and completed by 92 participants. Furthermore, the utility of the collected dataset was validated by training automatic naturalness prediction systems. We explored two approaches: fine-tuning an existing model originally trained for English, and training small downstream networks on top of frozen self-supervised speech models. Our models achieve a mean absolute error of 0.8 on a five-point MOS scale. Further analysis demonstrates the quality and diversity of the developed dataset, and its potential to advance TTS research in Spanish.
【3】Exploring Classical Piano Performance Generation with Expressive Music Variational AutoEncoder
链接:https://arxiv.org/abs/2507.01582
备注:Accepted by IEEE SMC 2025
摘要:古典音乐的创造力不仅来自于作曲家的音乐作品,也来自于表演者的表演,他们用富有表现力的细微差别来解释静态的符号。本文探讨了如何从零开始创作古典钢琴作品的挑战,旨在模拟作曲家和钢琴家在创作过程中的双重角色。我们介绍了表达复合词(ECP)表示,有效地捕捉古典表演的韵律结构和表达的细微差别。在此基础上,我们提出了表达性音乐变分自动编码器(XMVAE),一个具有两个分支的模型:矢量量化变分自动编码器(VQ-VAE)分支,它生成与分数相关的内容,代表作曲家,以及一个香草VAE分支,它生成表达性细节,履行钢琴家的角色。这些分支使用类似的Seq 2Seq架构进行联合训练,利用多尺度编码器捕获节拍级上下文信息,并利用正交Transformer解码器进行有效的复合令牌解码。客观和主观评估都表明,与最先进的模型相比,XMVAE产生的古典表演具有卓越的音乐质量。此外,在额外的乐谱数据集上预训练Composer分支有助于显著的性能提升。
摘要:The creativity of classical music arises not only from composers who craft the musical sheets but also from performers who interpret the static notations with expressive nuances. This paper addresses the challenge of generating classical piano performances from scratch, aiming to emulate the dual roles of composer and pianist in the creative process. We introduce the Expressive Compound Word (ECP) representation, which effectively captures both the metrical structure and expressive nuances of classical performances. Building on this, we propose the Expressive Music Variational AutoEncoder (XMVAE), a model featuring two branches: a Vector Quantized Variational AutoEncoder (VQ-VAE) branch that generates score-related content, representing the Composer, and a vanilla VAE branch that produces expressive details, fulfilling the role of Pianist. These branches are jointly trained with similar Seq2Seq architectures, leveraging a multiscale encoder to capture beat-level contextual information and an orthogonal Transformer decoder for efficient compound tokens decoding. Both objective and subjective evaluations demonstrate that XMVAE generates classical performances with superior musical quality compared to state-of-the-art models. Furthermore, pretraining the Composer branch on extra musical score datasets contribute to a significant performance gain.
【4】Real-Time Emergency Vehicle Siren Detection with Efficient CNNs on Embedded Hardware
链接:https://arxiv.org/abs/2507.01563
备注:10 pages, 10 figures, submitted to this https URL. arXiv admin note: text overlap with arXiv:2506.23437
摘要:我们提出了一个全栈的紧急车辆(EV)警笛检测系统设计的实时部署在嵌入式硬件上。所提出的方法是基于E2 PANNs,一个微调卷积神经网络从EPANNs派生,并优化二进制声音事件检测在城市声学条件下。一个关键的贡献是创建策划和语义结构化的数据集- AudioSet-EV,AudioSet-EV增强和Unified-EV -使用自定义AudioSet-Tools框架开发,以克服标准AudioSet注释的低可靠性。该系统部署在配备高保真DAC+麦克风板的Raspberry Pi 5上,实现了具有自适应帧大小、概率平滑和决策状态机的多线程推理引擎,以控制误报激活。远程WebSocket接口提供实时监控并促进现场演示功能。使用跨多个配置的逐帧和基于事件的指标来评估性能。实验结果表明,该系统实现了低延迟检测,在真实音频条件下具有更好的鲁棒性。这项工作证明了部署与IoS兼容的SED解决方案的可行性,这些解决方案可以形成分布式声学监控网络,通过低成本边缘设备上的WebSocket连接实现跨智能城市基础设施的协作紧急车辆跟踪。
摘要:We present a full-stack emergency vehicle (EV) siren detection system designed for real-time deployment on embedded hardware. The proposed approach is based on E2PANNs, a fine-tuned convolutional neural network derived from EPANNs, and optimized for binary sound event detection under urban acoustic conditions. A key contribution is the creation of curated and semantically structured datasets - AudioSet-EV, AudioSet-EV Augmented, and Unified-EV - developed using a custom AudioSet-Tools framework to overcome the low reliability of standard AudioSet annotations. The system is deployed on a Raspberry Pi 5 equipped with a high-fidelity DAC+microphone board, implementing a multithreaded inference engine with adaptive frame sizing, probability smoothing, and a decision-state machine to control false positive activations. A remote WebSocket interface provides real-time monitoring and facilitates live demonstration capabilities. Performance is evaluated using both framewise and event-based metrics across multiple configurations. Results show the system achieves low-latency detection with improved robustness under realistic audio conditions. This work demonstrates the feasibility of deploying IoS-compatible SED solutions that can form distributed acoustic monitoring networks, enabling collaborative emergency vehicle tracking across smart city infrastructures through WebSocket connectivity on low-cost edge devices.
【5】User-guided Generative Source Separation
链接:https://arxiv.org/abs/2507.01339
摘要:音乐源分离(MSS)的目的是从它们的混合中提取单个乐器源。虽然大多数现有的方法集中在广泛采用的四个干分离设置(人声,低音,鼓和其他乐器),这种方法缺乏现实世界的应用所需的灵活性。为了解决这个问题,我们提出了GuideSep,一个基于扩散的MSS模型,能够超越四个茎设置的仪器不可知分离。GuideSep以多个输入为条件:波形模仿条件,可以通过哼唱或播放目标旋律轻松提供,以及梅尔频谱图域掩码,为分离提供额外的指导。与以前的方法,依赖于固定的类标签或声音查询,我们的空调计划,再加上生成的方法,提供了更大的灵活性和适用性。此外,我们使用相同的模型架构设计了掩模预测基线,以系统地比较预测和生成方法。我们的客观和主观评估表明,GuideSep实现了高质量的分离,同时实现了更通用的仪器提取,突出了用户参与MSS基于扩散的生成过程的潜力。我们的代码和演示页面可以在https://yutongwen.github.io/GuideSep/上找到
摘要:Music source separation (MSS) aims to extract individual instrument sources from their mixture. While most existing methods focus on the widely adopted four-stem separation setup (vocals, bass, drums, and other instruments), this approach lacks the flexibility needed for real-world applications. To address this, we propose GuideSep, a diffusion-based MSS model capable of instrument-agnostic separation beyond the four-stem setup. GuideSep is conditioned on multiple inputs: a waveform mimicry condition, which can be easily provided by humming or playing the target melody, and mel-spectrogram domain masks, which offer additional guidance for separation. Unlike prior approaches that relied on fixed class labels or sound queries, our conditioning scheme, coupled with the generative approach, provides greater flexibility and applicability. Additionally, we design a mask-prediction baseline using the same model architecture to systematically compare predictive and generative approaches. Our objective and subjective evaluations demonstrate that GuideSep achieves high-quality separation while enabling more versatile instrument extraction, highlighting the potential of user participation in the diffusion-based generative process for MSS. Our code and demo page are available at https://yutongwen.github.io/GuideSep/
【6】A Review on Sound Source Localization in Robotics: Focusing on Deep Learning Methods
链接:https://arxiv.org/abs/2507.01143
备注:35 pages
摘要:声源定位(SSL)为听觉感知增加了空间维度,允许系统精确定位语音,机械噪声,警告音或其他声学事件的来源,这些功能有助于机器人导航,人机对话和状态监控。虽然现有的调查提供了有价值的历史背景,但它们通常针对一般的音频应用,并没有充分考虑机器人的限制或深度学习的最新进展。本文通过提供以机器人为重点的综合来解决这些差距,强调深度学习方法的最新进展。我们首先回顾经典的方法,如到达时间差(TDOA),波束形成,转向响应功率(SRP),和子空间分析。随后,我们深入研究了现代机器学习(ML)和深度学习(DL)方法,讨论了传统的ML和神经网络(NN),卷积神经网络(CNN),卷积递归神经网络(CRNN)和新兴的基于注意力的架构。数据和训练策略是基于DL的SSL的两个基石进行了探索。研究进一步分类的机器人类型和应用领域,以方便研究人员在确定相关的工作,为他们的具体情况。最后,我们强调了SSL工程目前面临的挑战一般,关于环境的鲁棒性,声源的多样性,在机器人技术的具体实施限制,以及数据和学习策略,在DL为基础的SSL。此外,我们勾画了有希望的方向,为下一代机器人提供一个可操作的路线图,以实现强大的,适应性强,高效的,可解释的基于DL的SSL。
摘要:Sound source localization (SSL) adds a spatial dimension to auditory perception, allowing a system to pinpoint the origin of speech, machinery noise, warning tones, or other acoustic events, capabilities that facilitate robot navigation, human-machine dialogue, and condition monitoring. While existing surveys provide valuable historical context, they typically address general audio applications and do not fully account for robotic constraints or the latest advancements in deep learning. This review addresses these gaps by offering a robotics-focused synthesis, emphasizing recent progress in deep learning methodologies. We start by reviewing classical methods such as Time Difference of Arrival (TDOA), beamforming, Steered-Response Power (SRP), and subspace analysis. Subsequently, we delve into modern machine learning (ML) and deep learning (DL) approaches, discussing traditional ML and neural networks (NNs), convolutional neural networks (CNNs), convolutional recurrent neural networks (CRNNs), and emerging attention-based architectures. The data and training strategy that are the two cornerstones of DL-based SSL are explored. Studies are further categorized by robot types and application domains to facilitate researchers in identifying relevant work for their specific contexts. Finally, we highlight the current challenges in SSL works in general, regarding environmental robustness, sound source multiplicity, and specific implementation constraints in robotics, as well as data and learning strategies in DL-based SSL. Also, we sketch promising directions to offer an actionable roadmap toward robust, adaptable, efficient, and explainable DL-based SSL for next-generation robots.
【7】Low-Complexity Neural Wind Noise Reduction for Audio Recordings
链接:https://arxiv.org/abs/2507.01821
摘要:风噪声会显著降低室外音频录制的质量,但在资源受限的设备上仍然难以实时抑制。在这项工作中,我们提出了一个低复杂度的单通道深度神经网络,它利用了风噪声的频谱特性。实验结果表明,我们的方法实现的性能与最先进的低复杂度ULCNet模型。该模型仅具有249K参数和大约73 MHz的计算能力,适用于嵌入式和移动音频应用。
摘要:Wind noise significantly degrades the quality of outdoor audio recordings, yet remains difficult to suppress in real-time on resource-constrained devices. In this work, we propose a low-complexity single-channel deep neural network that leverages the spectral characteristics of wind noise. Experimental results show that our method achieves performance comparable to the state-of-the-art low-complexity ULCNet model. The proposed model, with only 249K parameters and roughly 73 MHz of computational power, is suitable for embedded and mobile audio applications.
【8】Generalizable Detection of Audio Deepfakes
链接:https://arxiv.org/abs/2507.01750
备注:8 pages, 3 figures
摘要:在本文中,我们提出了旨在增强音频deepfake检测模型的泛化能力的综合研究。我们研究了各种预训练骨干的性能,包括Wav2Vec2,WavLM和Whisper,在不同的数据集上,包括来自ASVspoof挑战和其他来源的数据集。我们的实验集中在不同的数据增强策略和损失函数对模型性能的影响。我们的研究结果表明,音频deepfake检测模型的泛化能力得到了大幅增强,超过了ASVspoof 5挑战赛中排名第一的单个系统的性能。这项研究为优化音频模型以实现更强大的deepfake检测提供了有价值的见解,并促进了这一关键领域的未来研究。
摘要:In this paper, we present our comprehensive study aimed at enhancing the generalization capabilities of audio deepfake detection models. We investigate the performance of various pre-trained backbones, including Wav2Vec2, WavLM, and Whisper, across a diverse set of datasets, including those from the ASVspoof challenges and additional sources. Our experiments focus on the effects of different data augmentation strategies and loss functions on model performance. The results of our research demonstrate substantial enhancements in the generalization capabilities of audio deepfake detection models, surpassing the performance of the top-ranked single system in the ASVspoof 5 Challenge. This study contributes valuable insights into the optimization of audio models for more robust deepfake detection and facilitates future research in this critical area.
【9】QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model
链接:https://arxiv.org/abs/2507.01611
备注:This manuscript is currently under review for publication in the IEEE Transactions on Audio, Speech, and Language Processing. This work has been submitted to the IEEE for possible publication
摘要:声码器,将语音信号编码成声学特征并允许从它们重构语音信号,已经被研究了几十年。最近,深度学习的兴起特别推动了神经声码器的发展,以生成高质量的语音信号。另一方面,现有的端到端神经声码器具有黑盒性质,其使语音产生机制和语音的内在结构变得盲目,导致分别建模源激励和谐振特性的模糊性以及丧失灵活合成或修改高质量语音的能力。此外,它们的顺序波形生成通常需要复杂的网络,导致大量的时间消耗。在这项工作中,受准谐波模型(QHM),代表语音作为稀疏分量的启发,我们结合神经网络和QHM合成过程,提出了一种新的神经声码器的框架。因此,语音信号可以被编码成自回归移动平均(ARMA)函数以模拟谐振特性,从而产生在任何频率下准谐波的幅度和相位的准确估计。随后,语音可以重新合成和任意修改的音调移动和时间拉伸与高质量,而时间消耗和网络规模减少。实验表明,该方法利用QHM,ARMA模型和神经网络的优势,导致我们的方法优于其他方法的生成速度,合成质量和修改的灵活性。
摘要:Vocoders, encoding speech signals into acoustic features and allowing for speech signal reconstruction from them, have been studied for decades. Recently, the rise of deep learning has particularly driven the development of neural vocoders to generate high-quality speech signals. On the other hand, the existing end-to-end neural vocoders suffer from a black-box nature that blinds the speech production mechanism and the intrinsic structure of speech, resulting in the ambiguity of separately modeling source excitation and resonance characteristics and the loss of flexibly synthesizing or modifying speech with high quality. Moreover, their sequence-wise waveform generation usually requires complicated networks, leading to substantial time consumption. In this work, inspired by the quasi-harmonic model (QHM) that represents speech as sparse components, we combine the neural network and QHM synthesis process to propose a novel framework for the neural vocoder. Accordingly, speech signals can be encoded into autoregressive moving average (ARMA) functions to model the resonance characteristics, yielding accurate estimates of the amplitudes and phases of quasi-harmonics at any frequency. Subsequently, the speech can be resynthesized and arbitrarily modified in terms of pitch shifting and time stretching with high quality, whereas the time consumption and network size decrease. The experiments indicate that the proposed method leverages the strengths of QHM, the ARMA model, and neural networks, leading to the outperformance of our methods over other methods in terms of generation speed, synthesis quality, and modification flexibility.
【10】Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
链接:https://arxiv.org/abs/2507.01356
备注:Accepted at Interspeech 2025
摘要:感知到的声音可爱性在各种社交互动中起着至关重要的作用,例如合作伙伴选择和广告。一个提供针对目标受众的参考可爱语音样本的系统将使用户能够调整他们的说话风格和语音质量,促进更顺畅的沟通。为此,我们提出了一种语音转换方法,控制输入语音的可爱性,同时保留说话人的身份和语言内容。为了提高训练数据的可扩展性,我们在现有的语音喜好度数据集上训练了一个喜好度预测器,并利用它来自动注释具有喜好度评级的大型语音合成语料库。训练。实验评估揭示了预测器的输出和人类提供的可爱性评级之间的显着相关性。主观和客观的评价进一步表明,该方法有效地控制语音的可爱性,同时保持说话人身份和语言内容。
摘要:Perceived voice likability plays a crucial role in various social interactions, such as partner selection and advertising. A system that provides reference likable voice samples tailored to target audiences would enable users to adjust their speaking style and voice quality, facilitating smoother communication. To this end, we propose a voice conversion method that controls the likability of input speech while preserving both speaker identity and linguistic content. To improve training data scalability, we train a likability predictor on an existing voice likability dataset and employ it to automatically annotate a large speech synthesis corpus with likability ratings. Experimental evaluations reveal a significant correlation between the predictor's outputs and human-provided likability ratings. Subjective and objective evaluations further demonstrate that the proposed approach effectively controls voice likability while preserving both speaker identity and linguistic content.
【11】IdolSongsJp Corpus: A Multi-Singer Song Corpus in the Style of Japanese Idol Groups
链接:https://arxiv.org/abs/2507.01349
备注:Accepted at ISMIR 2025
摘要:日本偶像团体,由被称为“偶像”的表演者组成,是日本流行文化不可或缺的一部分。他们经常出现在现场音乐会和电视节目中,用他们的歌舞娱乐观众。与其他日本流行歌曲类似,偶像组合的音乐风格广泛,有各种类型的和弦进行和乐器安排。这些曲目通常具有许多乐器,并采用复杂的母带技术,导致高信号响度。此外,大多数歌曲包括一个歌曲分工(utawari)结构,其中成员之间的交替唱独唱和表演在一起。因此,这些歌曲非常适合用于基准测试各种音乐信息处理技术,例如歌手日记化、音乐源分离和在具有挑战性的条件下的自动和弦估计。针对这些特点,我们委托专业作曲家创作了15首日本偶像团体风格的歌曲,构建了一个名为IdolSongsJp的歌曲语料库。该语料库不仅包括掌握的音轨,而且还包括用于音乐源分离的茎、干声乐音轨和和弦注释。本文提供了一个详细的描述语料库,通过与现实世界中的偶像团体歌曲的比较,展示了它的多样性,并提出了它的应用评估几个音乐信息处理技术。
摘要:Japanese idol groups, comprising performers known as "idols," are an indispensable part of Japanese pop culture. They frequently appear in live concerts and television programs, entertaining audiences with their singing and dancing. Similar to other J-pop songs, idol group music covers a wide range of styles, with various types of chord progressions and instrumental arrangements. These tracks often feature numerous instruments and employ complex mastering techniques, resulting in high signal loudness. Additionally, most songs include a song division (utawari) structure, in which members alternate between singing solos and performing together. Hence, these songs are well-suited for benchmarking various music information processing techniques such as singer diarization, music source separation, and automatic chord estimation under challenging conditions. Focusing on these characteristics, we constructed a song corpus titled IdolSongsJp by commissioning professional composers to create 15 tracks in the style of Japanese idol groups. This corpus includes not only mastered audio tracks but also stems for music source separation, dry vocal tracks, and chord annotations. This paper provides a detailed description of the corpus, demonstrates its diversity through comparisons with real-world idol group songs, and presents its application in evaluating several music information processing techniques.
【12】SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
链接:https://arxiv.org/abs/2507.01348
备注:10 pages, includes references, 4 figures, 4 tables
摘要:外国口音转换在语音处理中一直是一个具有挑战性的课题。基于大语言模型(LLM)在文本到语音(TTS)任务中的显着成功,本研究探讨了基于LLM的FAC技术的适应性,我们称之为SpeechAccentLLM。在这个框架的核心,我们引入SpeechCodeVAE,第一个模型集成连接主义时间分类(CTC)直接到码本离散语音内容标记化。这种新的架构生成令牌具有独特的“局部性”的属性,通过实验证明内容忠实性,时间一致性和结构可恢复性之间的最佳权衡验证。然后,为了解决FAC模块的数据稀缺问题,我们采用了一种多任务学习策略,联合训练FAC和TTS模块。除了减轻数据限制外,与独立的FAC训练相比,这种方法还可以加速收敛并获得更高的语音质量。此外,利用我们的离散语音表示的显着特性,我们引入SpeechRestorer,旨在完善LLM生成的输出后处理架构。该模块有效地减轻了LLM推理管道中普遍存在的随机错误,同时增强了韵律连续性,消融实验验证了这一点。
摘要:Foreign accent conversion (FAC) in speech processing remains a challenging task. Building on the remarkable success of large language models (LLMs) in Text-to-Speech (TTS) tasks, this study investigates the adaptation of LLM-based techniques for FAC, which we term SpeechAccentLLM. At the core of this framework, we introduce SpeechCodeVAE, the first model to integrate connectionist temporal classification (CTC) directly into codebook discretization for speech content tokenization. This novel architecture generates tokens with a unique "locality" property, as validated by experiments demonstrating optimal trade-offs among content faithfulness, temporal coherence, and structural recoverability. Then, to address data scarcity for the FAC module, we adopted a multitask learning strategy that jointly trains the FAC and TTS modules. Beyond mitigating data limitations, this approach yielded accelerated convergence and superior speech quality compared to standalone FAC training. Moreover, leveraging the salient properties of our discrete speech representations, we introduce SpeechRestorer, a postprocessing architecture designed to refine LLM-generated outputs. This module effectively mitigates stochastic errors prevalent in LLM inference pipelines while enhancing prosodic continuity, as validated by ablation experiments.
【13】Hello Afrika: Speech Commands in Kinyarwanda
链接:https://arxiv.org/abs/2507.01024
备注:Data Science Africa, 2024
摘要:语音或语音命令是一种语言的更广泛口语语料库的子集,对于日常生活中使用的设备(特别是残疾人)中的大型AI系统的非接触式控制和激活至关重要。目前,非洲语言缺乏语音命令模型。Hello Afrika项目旨在解决这一问题,其第一次迭代的重点是Kinyarwanda语言,因为该国对开发语音识别技术表现出兴趣,最终成为Mozilla Common Voice上最大的数据集之一。该模型是建立在一个自定义语音命令语料库,由一般指令,数字和唤醒词。最终模型部署在多个设备(PC,移动电话和边缘设备)上,并使用适当的指标评估性能。
摘要:Voice or Speech Commands are a subset of the broader Spoken Word Corpus of a language which are essential for non-contact control of and activation of larger AI systems in devices used in everyday life especially for persons with disabilities. Currently, there is a dearth of speech command models for African languages. The Hello Afrika project aims to address this issue and its first iteration is focused on the Kinyarwanda language since the country has shown interest in developing speech recognition technologies culminating in one of the largest datasets on Mozilla Common Voice. The model was built off a custom speech command corpus made up of general directives, numbers, and a wake word. The final model was deployed on multiple devices (PC, Mobile Phone and Edge Devices) and the performance was assessed using suitable metrics.
【14】Workflow-Based Evaluation of Music Generation Systems
链接:https://arxiv.org/abs/2507.01022
备注:54 pages, 3 figures, 6 tables, 5 appendices
摘要:本研究提出了一个探索性的评估音乐生成系统(MGS)在当代音乐制作工作流程,通过检查八个开源系统。该评估框架通过专门设计的标准将技术见解与实际实验相结合,以调查音乐制作的迭代,非线性性质中系统的实用和创造性启示。本研究采用单一评估者的方法作为初步阶段,采用混合方法,利用定性方法形成假设,随后通过定量指标进行评估。所选系统代表了符号和基于音频的音乐生成方法的建筑多样性,涵盖了作曲,安排和声音设计任务。调查解决了当前MGS在音乐制作中的局限性,工作流程集成的挑战和机遇,以及在保持艺术真实性的同时作为协作工具的发展潜力。研究结果表明,这些系统的功能主要是作为补充工具,加强而不是取代人类的专业知识。它们在保持主题和结构连贯性方面表现出局限性,强调人类创造力在要求情感深度和复杂决策的任务中不可或缺的作用。本研究提供了一个结构化的评估框架,认为音乐创作的迭代性质。它确定了后续全面评估所需的方法改进,并确定了人工智能集成的可行领域,作为创意工作流程中的协作工具。该研究提供了基于实践的见解,以指导该领域的未来发展。
摘要:This study presents an exploratory evaluation of Music Generation Systems (MGS) within contemporary music production workflows by examining eight open-source systems. The evaluation framework combines technical insights with practical experimentation through criteria specifically designed to investigate the practical and creative affordances of the systems within the iterative, non-linear nature of music production. Employing a single-evaluator methodology as a preliminary phase, this research adopts a mixed approach utilizing qualitative methods to form hypotheses subsequently assessed through quantitative metrics. The selected systems represent architectural diversity across both symbolic and audio-based music generation approaches, spanning composition, arrangement, and sound design tasks. The investigation addresses limitations of current MGS in music production, challenges and opportunities for workflow integration, and development potential as collaborative tools while maintaining artistic authenticity. Findings reveal these systems function primarily as complementary tools enhancing rather than replacing human expertise. They exhibit limitations in maintaining thematic and structural coherence that emphasize the indispensable role of human creativity in tasks demanding emotional depth and complex decision-making. This study contributes a structured evaluation framework that considers the iterative nature of music creation. It identifies methodological refinements necessary for subsequent comprehensive evaluations and determines viable areas for AI integration as collaborative tools in creative workflows. The research provides empirically-grounded insights to guide future development in the field.
【15】Scalable Offline ASR for Command-Style Dictation in Courtrooms
链接:https://arxiv.org/abs/2507.01021
备注:Accepted to Interspeech 2025 Show & Tell
摘要:我们提出了一个开源的命令式听写框架,解决了资源密集型在线系统和高延迟批处理之间的差距。我们的方法使用语音活动检测(VAD)来分割音频,并使用Whisper模型并行转录这些片段,从而实现音频之间的高效多路复用。与SuperWhisper等专有系统不同,该框架还兼容大多数ASR架构,包括广泛使用的基于CTC的模型。我们的多路复用技术最大限度地提高了现实环境中的计算利用率,印度约15%的法庭都采用了这种技术。对实时数据的评估显示,与顺序批处理相比,随着用户并发性的增加,延迟持续减少。现场演示将展示我们的开源实现,并允许与会者与之实时互动。
摘要:We propose an open-source framework for Command-style dictation that addresses the gap between resource-intensive Online systems and high-latency Batch processing. Our approach uses Voice Activity Detection (VAD) to segment audio and transcribes these segments in parallel using Whisper models, enabling efficient multiplexing across audios. Unlike proprietary systems like SuperWhisper, this framework is also compatible with most ASR architectures, including widely used CTC-based models. Our multiplexing technique maximizes compute utilization in real-world settings, as demonstrated by its deployment in around 15% of India's courtrooms. Evaluations on live data show consistent latency reduction as user concurrency increases, compared to sequential batch processing. The live demonstration will showcase our open-sourced implementation and allow attendees to interact with it in real-time.
【1】Perceptual Ratings Predict Speech Inversion Articulatory Kinematics in Childhood Speech Sound Disorders
链接:https://arxiv.org/abs/2507.01888
备注:This manuscript is in submission for publication. It has not yet been peer reviewed
摘要:目的:本研究评估是否发音运动学,推断发音语音反转神经网络,与知觉评级的/r/和/s/在言语声音障碍的儿童的讲话。 方法:对118名2.25-45岁的儿童和3名成人的5,961个发音进行发音语音学声道变量推断。使用新颖的5点感知评定量表和训练方案对感知评定进行标准化。两个研究问题检查了推断声道变量的发音模式是否与所研究的电话的感知错误类别一致(例如,舌尖在齿状/s/中比在正确的/s/中更靠前)。第三个研究问题,如果梯度感知等级量表分数预测发音接近正确的产品。 结果:从线性混合模型估计的边际平均值支持17个18 /r/假设,涉及舌尖和舌体收缩。对于/s/,从第二线性混合模型估计的边际均值支持15个假设中的7个,特别是与舌尖相关的假设。第三个线性混合模型显示,感知等级量表分数显着预测发音接近的erados手机,以正确的生产。 结论:推断声道变量区分类别和发音错误的幅度/r/,并在较小程度上/s/,与感知判断。这些发现支持言语倒置声道变量和感知评定量表在量化与目标声音的发音接近度方面的临床可解释性,特别是对于/r/。
摘要:Purpose: This study evaluated whether articulatory kinematics, inferred by Articulatory Phonology speech inversion neural networks, aligned with perceptual ratings of /r/ and /s/ in the speech of children with speech sound disorders. Methods: Articulatory Phonology vocal tract variables were inferred for 5,961 utterances from 118 children and 3 adults, aged 2.25-45 years. Perceptual ratings were standardized using the novel 5-point PERCEPT Rating Scale and training protocol. Two research questions examined if the articulatory patterns of inferred vocal tract variables aligned with the perceptual error category for the phones investigated (e.g., tongue tip is more anterior in dentalized /s/ productions than in correct /s/). A third research question examined if gradient PERCEPT Rating Scale scores predicted articulatory proximity to correct productions. Results: Estimated marginal means from linear mixed models supported 17 of 18 /r/ hypotheses, involving tongue tip and tongue body constrictions. For /s/, estimated marginal means from a second linear mixed model supported 7 of 15 hypotheses, particularly those related to the tongue tip. A third linear mixed model revealed that PERCEPT Rating Scale scores significantly predicted articulatory proximity of errored phones to correct productions. Conclusion: Inferred vocal tract variables differentiated category and magnitude of articulatory errors for /r/, and to a lesser extent for /s/, aligning with perceptual judgments. These findings support the clinical interpretability of speech inversion vocal tract variables and the PERCEPT Rating Scale in quantifying articulatory proximity to the target sound, particularly for /r/.
【2】Low-Complexity Neural Wind Noise Reduction for Audio Recordings
链接:https://arxiv.org/abs/2507.01821
摘要:风噪声会显著降低室外音频录制的质量,但在资源受限的设备上仍然难以实时抑制。在这项工作中,我们提出了一个低复杂度的单通道深度神经网络,它利用了风噪声的频谱特性。实验结果表明,我们的方法实现的性能与最先进的低复杂度ULCNet模型。该模型仅具有249K参数和大约73 MHz的计算能力,适用于嵌入式和移动音频应用。
摘要:Wind noise significantly degrades the quality of outdoor audio recordings, yet remains difficult to suppress in real-time on resource-constrained devices. In this work, we propose a low-complexity single-channel deep neural network that leverages the spectral characteristics of wind noise. Experimental results show that our method achieves performance comparable to the state-of-the-art low-complexity ULCNet model. The proposed model, with only 249K parameters and roughly 73 MHz of computational power, is suitable for embedded and mobile audio applications.
【3】First Steps Towards Voice Anonymization for Code-Switching Speech
链接:https://arxiv.org/abs/2507.01765
备注:accepted at Interspeech 2025
摘要:语音匿名化的目标是修改音频,以隐藏其说话者的真实身份。对该任务的研究通常限于相同的英语阅读语音数据集,因此当前方法对其他类型语音数据的有效性仍然未知。在本文中,我们提出了第一次调查的语音匿名化的多语种现象的代码转换语音。我们准备了两个语料库,这项任务,并提出了适应多语言匿名化模型,使其适用于代码切换语音。通过在数据集上测试这种方法和两种独立于语言的方法的匿名化性能,我们发现只有多语言系统在隐私和实用性保护方面表现良好。此外,我们观察到的挑战,因为它的自发性和有限的代码切换支持的多语言语音识别模型在此数据上进行效用评估。
摘要:The goal of voice anonymization is to modify an audio such that the true identity of its speaker is hidden. Research on this task is typically limited to the same English read speech datasets, thus the efficacy of current methods for other types of speech data remains unknown. In this paper, we present the first investigation of voice anonymization for the multilingual phenomenon of code-switching speech. We prepare two corpora for this task and propose adaptations to a multilingual anonymization model to make it applicable for code-switching speech. By testing the anonymization performance of this and two language-independent methods on the datasets, we find that only the multilingual system performs well in terms of privacy and utility preservation. Furthermore, we observe challenges in performing utility evaluations on this data because of its spontaneous character and the limited code-switching support by the multilingual speech recognition model.
【4】Generalizable Detection of Audio Deepfakes
链接:https://arxiv.org/abs/2507.01750
备注:8 pages, 3 figures
摘要:在本文中,我们提出了旨在增强音频deepfake检测模型的泛化能力的综合研究。我们研究了各种预训练骨干的性能,包括Wav2Vec2,WavLM和Whisper,在不同的数据集上,包括来自ASVspoof挑战和其他来源的数据集。我们的实验集中在不同的数据增强策略和损失函数对模型性能的影响。我们的研究结果表明,音频deepfake检测模型的泛化能力得到了大幅增强,超过了ASVspoof 5挑战赛中排名第一的单个系统的性能。这项研究为优化音频模型以实现更强大的deepfake检测提供了有价值的见解,并促进了这一关键领域的未来研究。
摘要:In this paper, we present our comprehensive study aimed at enhancing the generalization capabilities of audio deepfake detection models. We investigate the performance of various pre-trained backbones, including Wav2Vec2, WavLM, and Whisper, across a diverse set of datasets, including those from the ASVspoof challenges and additional sources. Our experiments focus on the effects of different data augmentation strategies and loss functions on model performance. The results of our research demonstrate substantial enhancements in the generalization capabilities of audio deepfake detection models, surpassing the performance of the top-ranked single system in the ASVspoof 5 Challenge. This study contributes valuable insights into the optimization of audio models for more robust deepfake detection and facilitates future research in this critical area.
【5】QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model
链接:https://arxiv.org/abs/2507.01611
备注:This manuscript is currently under review for publication in the IEEE Transactions on Audio, Speech, and Language Processing. This work has been submitted to the IEEE for possible publication
摘要:声码器,将语音信号编码成声学特征并允许从它们重构语音信号,已经被研究了几十年。最近,深度学习的兴起特别推动了神经声码器的发展,以生成高质量的语音信号。另一方面,现有的端到端神经声码器具有黑盒性质,其使语音产生机制和语音的内在结构变得盲目,导致分别建模源激励和谐振特性的模糊性以及丧失灵活合成或修改高质量语音的能力。此外,它们的顺序波形生成通常需要复杂的网络,导致大量的时间消耗。在这项工作中,受准谐波模型(QHM),代表语音作为稀疏分量的启发,我们结合神经网络和QHM合成过程,提出了一种新的神经声码器的框架。因此,语音信号可以被编码成自回归移动平均(ARMA)函数以模拟谐振特性,从而产生在任何频率下准谐波的幅度和相位的准确估计。随后,语音可以重新合成和任意修改的音调移动和时间拉伸与高质量,而时间消耗和网络规模减少。实验表明,该方法利用QHM,ARMA模型和神经网络的优势,导致我们的方法优于其他方法的生成速度,合成质量和修改的灵活性。
摘要:Vocoders, encoding speech signals into acoustic features and allowing for speech signal reconstruction from them, have been studied for decades. Recently, the rise of deep learning has particularly driven the development of neural vocoders to generate high-quality speech signals. On the other hand, the existing end-to-end neural vocoders suffer from a black-box nature that blinds the speech production mechanism and the intrinsic structure of speech, resulting in the ambiguity of separately modeling source excitation and resonance characteristics and the loss of flexibly synthesizing or modifying speech with high quality. Moreover, their sequence-wise waveform generation usually requires complicated networks, leading to substantial time consumption. In this work, inspired by the quasi-harmonic model (QHM) that represents speech as sparse components, we combine the neural network and QHM synthesis process to propose a novel framework for the neural vocoder. Accordingly, speech signals can be encoded into autoregressive moving average (ARMA) functions to model the resonance characteristics, yielding accurate estimates of the amplitudes and phases of quasi-harmonics at any frequency. Subsequently, the speech can be resynthesized and arbitrarily modified in terms of pitch shifting and time stretching with high quality, whereas the time consumption and network size decrease. The experiments indicate that the proposed method leverages the strengths of QHM, the ARMA model, and neural networks, leading to the outperformance of our methods over other methods in terms of generation speed, synthesis quality, and modification flexibility.
【6】Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
链接:https://arxiv.org/abs/2507.01356
备注:Accepted at Interspeech 2025
摘要:感知到的声音可爱性在各种社交互动中起着至关重要的作用,例如合作伙伴选择和广告。一个提供针对目标受众的参考可爱语音样本的系统将使用户能够调整他们的说话风格和语音质量,促进更顺畅的沟通。为此,我们提出了一种语音转换方法,控制输入语音的可爱性,同时保留说话人的身份和语言内容。为了提高训练数据的可扩展性,我们在现有的语音喜好度数据集上训练了一个喜好度预测器,并利用它来自动注释具有喜好度评级的大型语音合成语料库。训练。实验评估揭示了预测器的输出和人类提供的可爱性评级之间的显着相关性。主观和客观的评价进一步表明,该方法有效地控制语音的可爱性,同时保持说话人身份和语言内容。
摘要:Perceived voice likability plays a crucial role in various social interactions, such as partner selection and advertising. A system that provides reference likable voice samples tailored to target audiences would enable users to adjust their speaking style and voice quality, facilitating smoother communication. To this end, we propose a voice conversion method that controls the likability of input speech while preserving both speaker identity and linguistic content. To improve training data scalability, we train a likability predictor on an existing voice likability dataset and employ it to automatically annotate a large speech synthesis corpus with likability ratings. Experimental evaluations reveal a significant correlation between the predictor's outputs and human-provided likability ratings. Subjective and objective evaluations further demonstrate that the proposed approach effectively controls voice likability while preserving both speaker identity and linguistic content.
【7】IdolSongsJp Corpus: A Multi-Singer Song Corpus in the Style of Japanese Idol Groups
链接:https://arxiv.org/abs/2507.01349
备注:Accepted at ISMIR 2025
摘要:日本偶像团体,由被称为“偶像”的表演者组成,是日本流行文化不可或缺的一部分。他们经常出现在现场音乐会和电视节目中,用他们的歌舞娱乐观众。与其他日本流行歌曲类似,偶像组合的音乐风格广泛,有各种类型的和弦进行和乐器安排。这些曲目通常具有许多乐器,并采用复杂的母带技术,导致高信号响度。此外,大多数歌曲包括一个歌曲分工(utawari)结构,其中成员之间的交替唱独唱和表演在一起。因此,这些歌曲非常适合用于基准测试各种音乐信息处理技术,例如歌手日记化、音乐源分离和在具有挑战性的条件下的自动和弦估计。针对这些特点,我们委托专业作曲家创作了15首日本偶像团体风格的歌曲,构建了一个名为IdolSongsJp的歌曲语料库。该语料库不仅包括掌握的音轨,而且还包括用于音乐源分离的茎、干声乐音轨和和弦注释。本文提供了一个详细的描述语料库,通过与现实世界中的偶像团体歌曲的比较,展示了它的多样性,并提出了它的应用评估几个音乐信息处理技术。
摘要:Japanese idol groups, comprising performers known as "idols," are an indispensable part of Japanese pop culture. They frequently appear in live concerts and television programs, entertaining audiences with their singing and dancing. Similar to other J-pop songs, idol group music covers a wide range of styles, with various types of chord progressions and instrumental arrangements. These tracks often feature numerous instruments and employ complex mastering techniques, resulting in high signal loudness. Additionally, most songs include a song division (utawari) structure, in which members alternate between singing solos and performing together. Hence, these songs are well-suited for benchmarking various music information processing techniques such as singer diarization, music source separation, and automatic chord estimation under challenging conditions. Focusing on these characteristics, we constructed a song corpus titled IdolSongsJp by commissioning professional composers to create 15 tracks in the style of Japanese idol groups. This corpus includes not only mastered audio tracks but also stems for music source separation, dry vocal tracks, and chord annotations. This paper provides a detailed description of the corpus, demonstrates its diversity through comparisons with real-world idol group songs, and presents its application in evaluating several music information processing techniques.
【8】SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
链接:https://arxiv.org/abs/2507.01348
备注:10 pages, includes references, 4 figures, 4 tables
摘要:外国口音转换在语音处理中一直是一个具有挑战性的课题。基于大语言模型(LLM)在文本到语音(TTS)任务中的显着成功,本研究探讨了基于LLM的FAC技术的适应性,我们称之为SpeechAccentLLM。在这个框架的核心,我们引入SpeechCodeVAE,第一个模型集成连接主义时间分类(CTC)直接到码本离散语音内容标记化。这种新的架构生成令牌具有独特的“局部性”的属性,通过实验证明内容忠实性,时间一致性和结构可恢复性之间的最佳权衡验证。然后,为了解决FAC模块的数据稀缺问题,我们采用了一种多任务学习策略,联合训练FAC和TTS模块。除了减轻数据限制外,与独立的FAC训练相比,这种方法还可以加速收敛并获得更高的语音质量。此外,利用我们的离散语音表示的显着特性,我们引入SpeechRestorer,旨在完善LLM生成的输出后处理架构。该模块有效地减轻了LLM推理管道中普遍存在的随机错误,同时增强了韵律连续性,消融实验验证了这一点。
摘要:Foreign accent conversion (FAC) in speech processing remains a challenging task. Building on the remarkable success of large language models (LLMs) in Text-to-Speech (TTS) tasks, this study investigates the adaptation of LLM-based techniques for FAC, which we term SpeechAccentLLM. At the core of this framework, we introduce SpeechCodeVAE, the first model to integrate connectionist temporal classification (CTC) directly into codebook discretization for speech content tokenization. This novel architecture generates tokens with a unique "locality" property, as validated by experiments demonstrating optimal trade-offs among content faithfulness, temporal coherence, and structural recoverability. Then, to address data scarcity for the FAC module, we adopted a multitask learning strategy that jointly trains the FAC and TTS modules. Beyond mitigating data limitations, this approach yielded accelerated convergence and superior speech quality compared to standalone FAC training. Moreover, leveraging the salient properties of our discrete speech representations, we introduce SpeechRestorer, a postprocessing architecture designed to refine LLM-generated outputs. This module effectively mitigates stochastic errors prevalent in LLM inference pipelines while enhancing prosodic continuity, as validated by ablation experiments.
【9】Classical Guitar Duet Separation using GuitarDuets -- a Dataset of Real and Synthesized Guitar Recordings
链接:https://arxiv.org/abs/2507.01172
备注:In Proceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR 2024), San Francisco, USA, November 2024. The dataset is available at: https://zenodo.org/records/12802440
摘要:音乐源分离(MSS)的最新进展集中在多音色的情况下,与现有的架构为不同乐器的分离定制,从而忽略了分离乐器具有相似的音色特性的挑战。为了解决这个问题,我们的工作重点是单音色MSS,特别是在古典吉他二重奏的背景下。为此,我们引入了GuitarDuets数据集,该数据集总共包含大约三个小时的真实和合成古典吉他二重奏录音,以及合成二重奏的音符级注释。我们进行了广泛的跨数据集的评估,适应Demucs,一个国家的最先进的MSS架构,单音色源分离。此外,我们开发了一个联合置换不变的转录和分离框架,利用笔记事件预测作为辅助信息。我们的研究结果表明,利用真实的和合成的子集的GuitarDuets导致独立记录的测试集相比,利用单独的一个子集,提高分离性能。我们还发现,虽然地面实况音符标签的可用性极大地帮助了分离网络的性能,但预测的音符估计仅导致边际改进。最后,我们讨论了常用的度量,如SDR和SI-SDR,在单音色MSS的上下文中的行为。
摘要:Recent advancements in music source separation (MSS) have focused in the multi-timbral case, with existing architectures tailored for the separation of distinct instruments, overlooking thus the challenge of separating instruments with similar timbral characteristics. Addressing this gap, our work focuses on monotimbral MSS, specifically within the context of classical guitar duets. To this end, we introduce the GuitarDuets dataset, featuring a combined total of approximately three hours of real and synthesized classical guitar duet recordings, as well as note-level annotations of the synthesized duets. We perform an extensive cross-dataset evaluation by adapting Demucs, a state-of-the-art MSS architecture, to monotimbral source separation. Furthermore, we develop a joint permutation-invariant transcription and separation framework, to exploit note event predictions as auxiliary information. Our results indicate that utilizing both the real and synthesized subsets of GuitarDuets leads to improved separation performance in an independently recorded test set compared to utilizing solely one subset. We also find that while the availability of ground-truth note labels greatly helps the performance of the separation network, the predicted note estimates result only in marginal improvement. Finally, we discuss the behavior of commonly utilized metrics, such as SDR and SI-SDR, in the context of monotimbral MSS.
【10】Hello Afrika: Speech Commands in Kinyarwanda
链接:https://arxiv.org/abs/2507.01024
备注:Data Science Africa, 2024
摘要:语音或语音命令是一种语言的更广泛口语语料库的子集,对于日常生活中使用的设备(特别是残疾人)中的大型AI系统的非接触式控制和激活至关重要。目前,非洲语言缺乏语音命令模型。Hello Afrika项目旨在解决这一问题,其第一次迭代的重点是Kinyarwanda语言,因为该国对开发语音识别技术表现出兴趣,最终成为Mozilla Common Voice上最大的数据集之一。该模型是建立在一个自定义语音命令语料库,由一般指令,数字和唤醒词。最终模型部署在多个设备(PC,移动电话和边缘设备)上,并使用适当的指标评估性能。
摘要:Voice or Speech Commands are a subset of the broader Spoken Word Corpus of a language which are essential for non-contact control of and activation of larger AI systems in devices used in everyday life especially for persons with disabilities. Currently, there is a dearth of speech command models for African languages. The Hello Afrika project aims to address this issue and its first iteration is focused on the Kinyarwanda language since the country has shown interest in developing speech recognition technologies culminating in one of the largest datasets on Mozilla Common Voice. The model was built off a custom speech command corpus made up of general directives, numbers, and a wake word. The final model was deployed on multiple devices (PC, Mobile Phone and Edge Devices) and the performance was assessed using suitable metrics.
【11】Workflow-Based Evaluation of Music Generation Systems
链接:https://arxiv.org/abs/2507.01022
备注:54 pages, 3 figures, 6 tables, 5 appendices
摘要:本研究提出了一个探索性的评估音乐生成系统(MGS)在当代音乐制作工作流程,通过检查八个开源系统。该评估框架通过专门设计的标准将技术见解与实际实验相结合,以调查音乐制作的迭代,非线性性质中系统的实用和创造性启示。本研究采用单一评估者的方法作为初步阶段,采用混合方法,利用定性方法形成假设,随后通过定量指标进行评估。所选系统代表了符号和基于音频的音乐生成方法的建筑多样性,涵盖了作曲,安排和声音设计任务。调查解决了当前MGS在音乐制作中的局限性,工作流程集成的挑战和机遇,以及在保持艺术真实性的同时作为协作工具的发展潜力。研究结果表明,这些系统的功能主要是作为补充工具,加强而不是取代人类的专业知识。它们在保持主题和结构连贯性方面表现出局限性,强调人类创造力在要求情感深度和复杂决策的任务中不可或缺的作用。本研究提供了一个结构化的评估框架,认为音乐创作的迭代性质。它确定了后续全面评估所需的方法改进,并确定了人工智能集成的可行领域,作为创意工作流程中的协作工具。该研究提供了基于实践的见解,以指导该领域的未来发展。
摘要:This study presents an exploratory evaluation of Music Generation Systems (MGS) within contemporary music production workflows by examining eight open-source systems. The evaluation framework combines technical insights with practical experimentation through criteria specifically designed to investigate the practical and creative affordances of the systems within the iterative, non-linear nature of music production. Employing a single-evaluator methodology as a preliminary phase, this research adopts a mixed approach utilizing qualitative methods to form hypotheses subsequently assessed through quantitative metrics. The selected systems represent architectural diversity across both symbolic and audio-based music generation approaches, spanning composition, arrangement, and sound design tasks. The investigation addresses limitations of current MGS in music production, challenges and opportunities for workflow integration, and development potential as collaborative tools while maintaining artistic authenticity. Findings reveal these systems function primarily as complementary tools enhancing rather than replacing human expertise. They exhibit limitations in maintaining thematic and structural coherence that emphasize the indispensable role of human creativity in tasks demanding emotional depth and complex decision-making. This study contributes a structured evaluation framework that considers the iterative nature of music creation. It identifies methodological refinements necessary for subsequent comprehensive evaluations and determines viable areas for AI integration as collaborative tools in creative workflows. The research provides empirically-grounded insights to guide future development in the field.
【12】Scalable Offline ASR for Command-Style Dictation in Courtrooms
链接:https://arxiv.org/abs/2507.01021
备注:Accepted to Interspeech 2025 Show & Tell
摘要:我们提出了一个开源的命令式听写框架,解决了资源密集型在线系统和高延迟批处理之间的差距。我们的方法使用语音活动检测(VAD)来分割音频,并使用Whisper模型并行转录这些片段,从而实现音频之间的高效多路复用。与SuperWhisper等专有系统不同,该框架还兼容大多数ASR架构,包括广泛使用的基于CTC的模型。我们的多路复用技术最大限度地提高了现实环境中的计算利用率,印度约15%的法庭都采用了这种技术。对实时数据的评估显示,与顺序批处理相比,随着用户并发性的增加,延迟持续减少。现场演示将展示我们的开源实现,并允许与会者与之实时互动。
摘要:We propose an open-source framework for Command-style dictation that addresses the gap between resource-intensive Online systems and high-latency Batch processing. Our approach uses Voice Activity Detection (VAD) to segment audio and transcribes these segments in parallel using Whisper models, enabling efficient multiplexing across audios. Unlike proprietary systems like SuperWhisper, this framework is also compatible with most ASR architectures, including widely used CTC-based models. Our multiplexing technique maximizes compute utilization in real-world settings, as demonstrated by its deployment in around 15% of India's courtrooms. Evaluations on live data show consistent latency reduction as user concurrency increases, compared to sequential batch processing. The live demonstration will showcase our open-sourced implementation and allow attendees to interact with it in real-time.
【13】Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla
链接:https://arxiv.org/abs/2507.01931
摘要:近年来,在大型多语言文本和语音数据集上训练的神经模型在支持低资源语言方面表现出了巨大的潜力。这项研究调查了两种最先进的自动语音识别(ASR)模型,OpenAI的Whisper(Small & Large-V2)和Facebook的Wav 2 Vec-BERT在孟加拉语(一种低资源语言)上的性能。我们使用两个公开的数据集进行了实验:Mozilla Common Voice-17和OpenSLR来评估模型性能。通过系统的微调和超参数优化,包括学习率,epoch和模型检查点选择,我们比较了基于单词错误率(WER),字符错误率(CER),训练时间和计算效率的模型。Wav 2 Vec-BERT模型在所有关键评估指标上都优于Whisper,表现出卓越的性能,同时需要更少的计算资源,并为在低资源语言环境中开发强大的语音识别系统提供了宝贵的见解。
摘要:In recent years, neural models trained on large multilingual text and speech datasets have shown great potential for supporting low-resource languages. This study investigates the performances of two state-of-the-art Automatic Speech Recognition (ASR) models, OpenAI's Whisper (Small & Large-V2) and Facebook's Wav2Vec-BERT on Bangla, a low-resource language. We have conducted experiments using two publicly available datasets: Mozilla Common Voice-17 and OpenSLR to evaluate model performances. Through systematic fine-tuning and hyperparameter optimization, including learning rate, epochs, and model checkpoint selection, we have compared the models based on Word Error Rate (WER), Character Error Rate (CER), Training Time, and Computational Efficiency. The Wav2Vec-BERT model outperformed Whisper across all key evaluation metrics, demonstrated superior performance while requiring fewer computational resources, and offered valuable insights to develop robust speech recognition systems in low-resource linguistic settings.
【14】A Dataset for Automatic Assessment of TTS Quality in Spanish
链接:https://arxiv.org/abs/2507.01805
备注:5 pages, 2 figures. Accepted at Interspeech 2025
摘要:这项工作解决了文本到语音(TTS)系统在西班牙语的自动评估的数据库的开发,旨在提高自然度预测模型的准确性。该数据集由来自52个不同TTS系统和人类声音的4,326个音频样本组成,据我们所知,这是西班牙语中的第一个。为了给音频贴上标签,根据ITU-T Rec.P.807标准设计了一个主观测试,由92名参与者完成。此外,通过训练自动自然度预测系统,验证了所收集数据集的实用性。我们探索了两种方法:微调最初为英语训练的现有模型,并在冻结的自监督语音模型上训练小型下游网络。我们的模型在五点MOS尺度上实现了0.8的平均绝对误差。进一步的分析表明,开发的数据集的质量和多样性,其潜力,以推进TTS研究在西班牙语。
摘要:This work addresses the development of a database for the automatic assessment of text-to-speech (TTS) systems in Spanish, aiming to improve the accuracy of naturalness prediction models. The dataset consists of 4,326 audio samples from 52 different TTS systems and human voices and is, up to our knowledge, the first of its kind in Spanish. To label the audios, a subjective test was designed based on the ITU-T Rec. P.807 standard and completed by 92 participants. Furthermore, the utility of the collected dataset was validated by training automatic naturalness prediction systems. We explored two approaches: fine-tuning an existing model originally trained for English, and training small downstream networks on top of frozen self-supervised speech models. Our models achieve a mean absolute error of 0.8 on a five-point MOS scale. Further analysis demonstrates the quality and diversity of the developed dataset, and its potential to advance TTS research in Spanish.
【15】Exploring Classical Piano Performance Generation with Expressive Music Variational AutoEncoder
链接:https://arxiv.org/abs/2507.01582
备注:Accepted by IEEE SMC 2025
摘要:古典音乐的创造力不仅来自于作曲家的音乐作品,也来自于表演者的表演,他们用富有表现力的细微差别来解释静态的符号。本文探讨了如何从零开始创作古典钢琴作品的挑战,旨在模拟作曲家和钢琴家在创作过程中的双重角色。我们介绍了表达复合词(ECP)表示,有效地捕捉古典表演的韵律结构和表达的细微差别。在此基础上,我们提出了表达性音乐变分自动编码器(XMVAE),一个具有两个分支的模型:矢量量化变分自动编码器(VQ-VAE)分支,它生成与分数相关的内容,代表作曲家,以及一个香草VAE分支,它生成表达性细节,履行钢琴家的角色。这些分支使用类似的Seq 2Seq架构进行联合训练,利用多尺度编码器捕获节拍级上下文信息,并利用正交Transformer解码器进行有效的复合令牌解码。客观和主观的评估表明,XMVAE产生的经典表演与卓越的音乐质量相比,国家的最先进的模型。此外,在额外的乐谱数据集上预训练Composer分支有助于显著的性能提升。
摘要:The creativity of classical music arises not only from composers who craft the musical sheets but also from performers who interpret the static notations with expressive nuances. This paper addresses the challenge of generating classical piano performances from scratch, aiming to emulate the dual roles of composer and pianist in the creative process. We introduce the Expressive Compound Word (ECP) representation, which effectively captures both the metrical structure and expressive nuances of classical performances. Building on this, we propose the Expressive Music Variational AutoEncoder (XMVAE), a model featuring two branches: a Vector Quantized Variational AutoEncoder (VQ-VAE) branch that generates score-related content, representing the Composer, and a vanilla VAE branch that produces expressive details, fulfilling the role of Pianist. These branches are jointly trained with similar Seq2Seq architectures, leveraging a multiscale encoder to capture beat-level contextual information and an orthogonal Transformer decoder for efficient compound tokens decoding. Both objective and subjective evaluations demonstrate that XMVAE generates classical performances with superior musical quality compared to state-of-the-art models. Furthermore, pretraining the Composer branch on extra musical score datasets contribute to a significant performance gain.
【16】Real-Time Emergency Vehicle Siren Detection with Efficient CNNs on Embedded Hardware
链接:https://arxiv.org/abs/2507.01563
备注:10 pages, 10 figures, submitted to this https URL. arXiv admin note: text overlap with arXiv:2506.23437
摘要:我们提出了一个全栈的紧急车辆(EV)警笛检测系统设计的实时部署在嵌入式硬件上。所提出的方法是基于E2 PANNs,一个微调卷积神经网络从EPANNs派生,并优化二进制声音事件检测在城市声学条件下。一个关键的贡献是创建策划和语义结构化的数据集- AudioSet-EV,AudioSet-EV增强和Unified-EV -使用自定义AudioSet-Tools框架开发,以克服标准AudioSet注释的低可靠性。该系统部署在配备高保真DAC+麦克风板的Raspberry Pi 5上,实现了具有自适应帧大小、概率平滑和决策状态机的多线程推理引擎,以控制误报激活。远程WebSocket接口提供实时监控并促进现场演示功能。使用跨多个配置的逐帧和基于事件的指标来评估性能。实验结果表明,该系统实现了低延迟检测,在真实音频条件下具有更好的鲁棒性。这项工作证明了部署与IoS兼容的SED解决方案的可行性,这些解决方案可以形成分布式声学监控网络,通过低成本边缘设备上的WebSocket连接实现跨智能城市基础设施的协作紧急车辆跟踪。
摘要:We present a full-stack emergency vehicle (EV) siren detection system designed for real-time deployment on embedded hardware. The proposed approach is based on E2PANNs, a fine-tuned convolutional neural network derived from EPANNs, and optimized for binary sound event detection under urban acoustic conditions. A key contribution is the creation of curated and semantically structured datasets - AudioSet-EV, AudioSet-EV Augmented, and Unified-EV - developed using a custom AudioSet-Tools framework to overcome the low reliability of standard AudioSet annotations. The system is deployed on a Raspberry Pi 5 equipped with a high-fidelity DAC+microphone board, implementing a multithreaded inference engine with adaptive frame sizing, probability smoothing, and a decision-state machine to control false positive activations. A remote WebSocket interface provides real-time monitoring and facilitates live demonstration capabilities. Performance is evaluated using both framewise and event-based metrics across multiple configurations. Results show the system achieves low-latency detection with improved robustness under realistic audio conditions. This work demonstrates the feasibility of deploying IoS-compatible SED solutions that can form distributed acoustic monitoring networks, enabling collaborative emergency vehicle tracking across smart city infrastructures through WebSocket connectivity on low-cost edge devices.
【17】User-guided Generative Source Separation
链接:https://arxiv.org/abs/2507.01339
摘要:音乐源分离(MSS)的目的是从它们的混合中提取单个乐器源。虽然大多数现有的方法集中在广泛采用的四个干分离设置(人声,低音,鼓和其他乐器),这种方法缺乏现实世界的应用所需的灵活性。为了解决这个问题,我们提出了GuideSep,一个基于扩散的MSS模型,能够超越四个茎设置的仪器不可知分离。GuideSep以多个输入为条件:波形模仿条件,可以通过哼唱或播放目标旋律轻松提供,以及梅尔频谱图域掩码,为分离提供额外的指导。与以前的方法,依赖于固定的类标签或声音查询,我们的空调计划,再加上生成的方法,提供了更大的灵活性和适用性。此外,我们使用相同的模型架构设计了掩模预测基线,以系统地比较预测和生成方法。我们的客观和主观评估表明,GuideSep实现了高质量的分离,同时实现了更通用的仪器提取,突出了用户参与MSS基于扩散的生成过程的潜力。我们的代码和演示页面可以在https://yutongwen.github.io/GuideSep/上找到
摘要:Music source separation (MSS) aims to extract individual instrument sources from their mixture. While most existing methods focus on the widely adopted four-stem separation setup (vocals, bass, drums, and other instruments), this approach lacks the flexibility needed for real-world applications. To address this, we propose GuideSep, a diffusion-based MSS model capable of instrument-agnostic separation beyond the four-stem setup. GuideSep is conditioned on multiple inputs: a waveform mimicry condition, which can be easily provided by humming or playing the target melody, and mel-spectrogram domain masks, which offer additional guidance for separation. Unlike prior approaches that relied on fixed class labels or sound queries, our conditioning scheme, coupled with the generative approach, provides greater flexibility and applicability. Additionally, we design a mask-prediction baseline using the same model architecture to systematically compare predictive and generative approaches. Our objective and subjective evaluations demonstrate that GuideSep achieves high-quality separation while enabling more versatile instrument extraction, highlighting the potential of user participation in the diffusion-based generative process for MSS. Our code and demo page are available at https://yutongwen.github.io/GuideSep/
【18】A Review on Sound Source Localization in Robotics: Focusing on Deep Learning Methods
链接:https://arxiv.org/abs/2507.01143
备注:35 pages
摘要:声源定位(SSL)为听觉感知增加了空间维度,允许系统精确定位语音,机械噪声,警告音或其他声学事件的来源,这些功能有助于机器人导航,人机对话和状态监控。虽然现有的调查提供了有价值的历史背景,但它们通常针对一般的音频应用,并没有充分考虑机器人的限制或深度学习的最新进展。本文通过提供以机器人为重点的综合来解决这些差距,强调深度学习方法的最新进展。我们首先回顾经典的方法,如到达时间差(TDOA),波束形成,转向响应功率(SRP),和子空间分析。随后,我们深入研究了现代机器学习(ML)和深度学习(DL)方法,讨论了传统的ML和神经网络(NN),卷积神经网络(CNN),卷积递归神经网络(CRNN)和新兴的基于注意力的架构。数据和训练策略是基于DL的SSL的两个基石进行了探索。研究进一步分类的机器人类型和应用领域,以方便研究人员在确定相关的工作,为他们的具体情况。最后,我们强调了SSL工作当前面临的总体挑战,涉及环境鲁棒性、声源多样性和机器人技术中的特定实施限制,以及基于DL的SSL中的数据和学习策略。此外,我们勾画了有希望的方向,为下一代机器人提供一个可操作的路线图,以实现强大的,适应性强,高效的,可解释的基于DL的SSL。
摘要:Sound source localization (SSL) adds a spatial dimension to auditory perception, allowing a system to pinpoint the origin of speech, machinery noise, warning tones, or other acoustic events, capabilities that facilitate robot navigation, human-machine dialogue, and condition monitoring. While existing surveys provide valuable historical context, they typically address general audio applications and do not fully account for robotic constraints or the latest advancements in deep learning. This review addresses these gaps by offering a robotics-focused synthesis, emphasizing recent progress in deep learning methodologies. We start by reviewing classical methods such as Time Difference of Arrival (TDOA), beamforming, Steered-Response Power (SRP), and subspace analysis. Subsequently, we delve into modern machine learning (ML) and deep learning (DL) approaches, discussing traditional ML and neural networks (NNs), convolutional neural networks (CNNs), convolutional recurrent neural networks (CRNNs), and emerging attention-based architectures. The data and training strategy that are the two cornerstones of DL-based SSL are explored. Studies are further categorized by robot types and application domains to facilitate researchers in identifying relevant work for their specific contexts. Finally, we highlight the current challenges in SSL works in general, regarding environmental robustness, sound source multiplicity, and specific implementation constraints in robotics, as well as data and learning strategies in DL-based SSL. Also, we sketch promising directions to offer an actionable roadmap toward robust, adaptable, efficient, and explainable DL-based SSL for next-generation robots.
机器翻译由腾讯交互翻译提供,仅供参考
