微信公众号:arXiv_Daily
cs.SD语音
标题: Smart:通过音频领域美学奖励调整象征性音乐生成系统
链接:https://arxiv.org/abs/2504.16839
摘要:最近的工作提出了训练机器学习模型来预测音乐音频的美学评级。我们的工作探讨了这种模型是否可以用于通过强化学习微调符号音乐生成系统,以及这对系统输出有什么影响。为了测试这一点,我们使用组相对策略优化来微调钢琴模型,其中音频渲染输出的Meta Audiobox Aesthetics评级作为奖励。我们发现,这种优化产生的输出的多个低级别的功能上的影响,并提高了平均主观评分在初步的听力研究与14 $$参与者。我们还发现,过度优化大大降低了模型输出的多样性。
摘要:Recent work has proposed training machine learning models to predict aesthetic ratings for music audio. Our work explores whether such models can be used to finetune a symbolic music generation system with reinforcement learning, and what effect this has on the system outputs. To test this, we use group relative policy optimization to finetune a piano MIDI model with Meta Audiobox Aesthetics ratings of audio-rendered outputs as the reward. We find that this optimization has effects on multiple low-level features of the generated outputs, and improves the average subjective ratings in a preliminary listening study with $14$ participants. We also find that over-optimization dramatically reduces diversity of model outputs.
【2】 Insect-Computer Hybrid Speaker: Speaker using Chirp of the Cicada Controlled by Electrical Muscle Stimulation
标题: 昆虫-计算机混合扬声器:使用由肌肉电刺激控制的蝉鸣的扬声器链接:https://arxiv.org/abs/2504.16459
备注:6 pages, 3 figures
摘要:我们提出了“昆虫-计算机混合扬声器”,这使我们能够使音乐从计算机和昆虫的组合。许多研究已经提出了控制昆虫和获得反馈的方法和界面。然而,关于利用昆虫与第三方互动的研究较少。在本文中,我们提出了一种方法,其中蝉被用作扬声器触发使用肌肉电刺激(EMS)。我们探索并研究了蝉啁啾的合适控制波形、合适的电压范围以及蝉能啁啾的最大音调。
摘要:We propose "Insect-Computer Hybrid Speaker", which enables us to make musics made from combinations of computer and insects. Lots of studies have proposed methods and interfaces for controlling insects and obtaining feedback. However, there have been less research on the use of insects for interaction with third parties. In this paper, we propose a method in which cicadas are used as speakers triggered by using Electrical Muscle Stimulation (EMS). We explored and investigated the suitable waveform of chirp to be controlled, the appropriate voltage range, and the maximum pitch at which cicadas can chirp.
【3】 An Automated Pipeline for Few-Shot Bird Call Classification: A Case Study with the Tooth-Billed Pigeon
标题: 用于少量鸟叫声分类的自动管道:牙齿鸟叫声的案例研究链接:https://arxiv.org/abs/2504.16276
备注:16 pages, 5 figures, 4 tables
摘要:本文提出了一种自动化的一次性鸟鸣分类管道,设计用于稀有物种缺乏大型公开可用的分类器,如BirdNET和Perch。虽然这些模型在检测具有丰富训练数据的常见鸟类方面表现出色,但它们缺乏对只有1-3个已知记录的物种的选择-这对于保护主义者监测濒危鸟类的最后剩余个体来说是一个关键限制。为了解决这个问题,我们利用大型鸟类分类网络的嵌入空间,并使用余弦相似性开发分类器,结合滤波和去噪预处理技术,以最少的训练数据优化检测。我们使用聚类度量评估各种嵌入空间,并在Xeno-Canto录音的模拟场景和对极度濒危的齿嘴鸽(Didunculus strigirostris)的真实测试中验证我们的方法,该鸽子没有现有的分类器,只有三个确认的录音。最终的模型在检测齿嘴鸽叫声时达到了1.0的召回率和0.95的准确率,使其在该领域具有实用性。这个开源系统为寻求检测和监测濒临灭绝的稀有物种的保护主义者提供了实用的工具。
摘要:This paper presents an automated one-shot bird call classification pipeline designed for rare species absent from large publicly available classifiers like BirdNET and Perch. While these models excel at detecting common birds with abundant training data, they lack options for species with only 1-3 known recordings-a critical limitation for conservationists monitoring the last remaining individuals of endangered birds. To address this, we leverage the embedding space of large bird classification networks and develop a classifier using cosine similarity, combined with filtering and denoising preprocessing techniques, to optimize detection with minimal training data. We evaluate various embedding spaces using clustering metrics and validate our approach in both a simulated scenario with Xeno-Canto recordings and a real-world test on the critically endangered tooth-billed pigeon (Didunculus strigirostris), which has no existing classifiers and only three confirmed recordings. The final model achieved 1.0 recall and 0.95 accuracy in detecting tooth-billed pigeon calls, making it practical for use in the field. This open-source system provides a practical tool for conservationists seeking to detect and monitor rare species on the brink of extinction.
【4】 Using Phonemes in cascaded S2S translation pipeline
标题: 在级联S2 S翻译管道中使用音素链接:https://arxiv.org/abs/2504.16234
备注:Accepted at Swiss NLP Conference 2025
摘要:本文探讨了在传统的多语言同步语音到语音翻译管道中使用音素作为文本表示的想法,而不是传统的依赖于基于文本的语言表示。为了研究这一点,我们在WMT17数据集上以两种格式训练了一个开源的序列到序列模型:一种使用标准文本表示,另一种使用音素表示。这两种方法的性能使用BLEU度量进行评估。我们的研究结果表明,音素的方法提供了相当的质量,但提供了几个优势,包括较低的资源要求或更适合低资源的语言。
摘要:This paper explores the idea of using phonemes as a textual representation within a conventional multilingual simultaneous speech-to-speech translation pipeline, as opposed to the traditional reliance on text-based language representations. To investigate this, we trained an open-source sequence-to-sequence model on the WMT17 dataset in two formats: one using standard textual representation and the other employing phonemic representation. The performance of both approaches was assessed using the BLEU metric. Our findings shows that the phonemic approach provides comparable quality but offers several advantages, including lower resource requirements or better suitability for low-resource languages.
【5】 TinyML for Speech Recognition
标题: 用于语音识别的TinyML链接:https://arxiv.org/abs/2504.16213
摘要:我们训练并部署了一个量化的1D卷积神经网络模型,以在资源高度受限的物联网边缘设备上进行语音识别。这在各种物联网(IoT)应用中可能很有用,例如智能家居以及老年人和残疾人的环境辅助生活,仅举几个例子。在本文中,我们首先创建了一个新的数据集,其中包含一个多小时的音频数据,使我们的研究能够进行,并将对该领域的未来研究有用。其次,我们利用Edge Impulse提供的技术来增强模型的性能,并在我们的数据集上实现高达97%的高准确度。为了验证,我们使用Arduino Nano 33 BLE Sense微控制器板实现了我们的原型。该微控制器板专为物联网和人工智能应用而设计,使其成为我们目标用例场景的理想选择。虽然大多数现有的研究集中在有限的关键字集,我们的模型可以处理23个不同的关键字,使复杂的命令。
摘要:We train and deploy a quantized 1D convolutional neural network model to conduct speech recognition on a highly resource-constrained IoT edge device. This can be useful in various Internet of Things (IoT) applications, such as smart homes and ambient assisted living for the elderly and people with disabilities, just to name a few examples. In this paper, we first create a new dataset with over one hour of audio data that enables our research and will be useful to future studies in this field. Second, we utilize the technologies provided by Edge Impulse to enhance our model's performance and achieve a high Accuracy of up to 97% on our dataset. For the validation, we implement our prototype using the Arduino Nano 33 BLE Sense microcontroller board. This microcontroller board is specifically designed for IoT and AI applications, making it an ideal choice for our target use case scenarios. While most existing research focuses on a limited set of keywords, our model can process 23 different keywords, enabling complex commands.
【6】 SoCov: Semi-Orthogonal Parametric Pooling of Covariance Matrix for Speaker Recognition
标题: SoCov:用于说话人识别的协方差矩阵半垂直参数池链接:https://arxiv.org/abs/2504.16441
备注:This paper has been accepted by IEEE ICASSP2025
摘要:在传统的深度说话人嵌入框架中,池化层随着时间的推移聚合所有帧级特征,并计算它们的均值和标准差统计数据,作为后续段级层的输入。这种统计池策略从可变长度的语音段产生固定长度的表示。然而,该方法平等地对待不同的帧级特征,并丢弃协方差信息。本文提出了协方差矩阵的半正交参数池化(SoCov)方法。SoCov池化从自关注帧级特征计算协方差矩阵,并使用半正交参数向量化将其压缩成向量,然后将其与加权标准偏差向量级联以形成到片段级层的输入。基于SoCov的深度嵌入被称为“sc-vector”。建议的sc-向量进行比较,几个不同的基线SRE 21开发和评估集。sc-向量系统显著优于常规x-向量系统,SRE 21 Eval的EER相对降低15.5%。当使用自我关注的深度特征时,SoCov有助于将SRE 21 Eval上的EER相对于传统的“平均值+标准差”统计减少约30.9%。
摘要:In conventional deep speaker embedding frameworks, the pooling layer aggregates all frame-level features over time and computes their mean and standard deviation statistics as inputs to subsequent segment-level layers. Such statistics pooling strategy produces fixed-length representations from variable-length speech segments. However, this method treats different frame-level features equally and discards covariance information. In this paper, we propose the Semi-orthogonal parameter pooling of Covariance matrix (SoCov) method. The SoCov pooling computes the covariance matrix from the self-attentive frame-level features and compresses it into a vector using the semi-orthogonal parametric vectorization, which is then concatenated with the weighted standard deviation vector to form inputs to the segment-level layers. Deep embedding based on SoCov is called ``sc-vector''. The proposed sc-vector is compared to several different baselines on the SRE21 development and evaluation sets. The sc-vector system significantly outperforms the conventional x-vector system, with a relative reduction in EER of 15.5% on SRE21Eval. When using self-attentive deep feature, SoCov helps to reduce EER on SRE21Eval by about 30.9% relatively to the conventional ``mean + standard deviation'' statistics.
【7】 Deep, data-driven modeling of room acoustics: literature review and research perspectives
标题: 房间声学的深度、数据驱动建模:文献回顾和研究观点链接:https://arxiv.org/abs/2504.16289
摘要:None
摘要:Our everyday auditory experience is shaped by the acoustics of the indoor environments in which we live. Room acoustics modeling is aimed at establishing mathematical representations of acoustic wave propagation in such environments. These representations are relevant to a variety of problems ranging from echo-aided auditory indoor navigation to restoring speech understanding in cocktail party scenarios. Many disciplines in science and engineering have recently witnessed a paradigm shift powered by deep learning (DL), and room acoustics research is no exception. The majority of deep, data-driven room acoustics models are inspired by DL-based speech and image processing, and hence lack the intrinsic space-time structure of acoustic wave propagation. More recently, DL-based models for room acoustics that include either geometric or wave-based information have delivered promising results, primarily for the problem of sound field reconstruction. In this review paper, we will provide an extensive and structured literature review on deep, data-driven modeling in room acoustics. Moreover, we position these models in a framework that allows for a conceptual comparison with traditional physical and data-driven models. Finally, we identify strengths and shortcomings of deep, data-driven room acoustics models and outline the main challenges for further research.
【1】 SoCov: Semi-Orthogonal Parametric Pooling of Covariance Matrix for Speaker Recognition
标题: SoCov:用于说话人识别的协方差矩阵半垂直参数池链接:https://arxiv.org/abs/2504.16441
备注:This paper has been accepted by IEEE ICASSP2025
摘要:在传统的深度说话人嵌入框架中,池化层随着时间的推移聚合所有帧级特征,并计算它们的均值和标准差统计数据,作为后续段级层的输入。这种统计池策略从可变长度的语音段产生固定长度的表示。然而,该方法平等地对待不同的帧级特征,并丢弃协方差信息。本文提出了协方差矩阵的半正交参数池化(SoCov)方法。SoCov池化从自关注帧级特征计算协方差矩阵,并使用半正交参数向量化将其压缩成向量,然后将其与加权标准偏差向量级联以形成到片段级层的输入。基于SoCov的深度嵌入被称为“sc-vector”。建议的sc-向量进行比较,几个不同的基线SRE 21开发和评估集。sc-向量系统显著优于常规x-向量系统,SRE 21 Eval的EER相对降低15.5%。当使用自我关注的深度特征时,SoCov有助于将SRE 21 Eval上的EER相对于传统的“平均值+标准差”统计减少约30.9%。
摘要:In conventional deep speaker embedding frameworks, the pooling layer aggregates all frame-level features over time and computes their mean and standard deviation statistics as inputs to subsequent segment-level layers. Such statistics pooling strategy produces fixed-length representations from variable-length speech segments. However, this method treats different frame-level features equally and discards covariance information. In this paper, we propose the Semi-orthogonal parameter pooling of Covariance matrix (SoCov) method. The SoCov pooling computes the covariance matrix from the self-attentive frame-level features and compresses it into a vector using the semi-orthogonal parametric vectorization, which is then concatenated with the weighted standard deviation vector to form inputs to the segment-level layers. Deep embedding based on SoCov is called ``sc-vector''. The proposed sc-vector is compared to several different baselines on the SRE21 development and evaluation sets. The sc-vector system significantly outperforms the conventional x-vector system, with a relative reduction in EER of 15.5% on SRE21Eval. When using self-attentive deep feature, SoCov helps to reduce EER on SRE21Eval by about 30.9% relatively to the conventional ``mean + standard deviation'' statistics.
【2】 Deep, data-driven modeling of room acoustics: literature review and research perspectives
标题: 房间声学的深度、数据驱动建模:文献回顾和研究观点链接:https://arxiv.org/abs/2504.16289
摘要:我们日常的听觉体验是由我们生活的室内环境的声学塑造的。室内声学建模的目的是建立在这样的环境中的声波传播的数学表示。这些表示是相关的各种问题,从回声辅助听觉室内导航恢复语音理解鸡尾酒会的情况。科学和工程领域的许多学科最近都见证了由深度学习(DL)驱动的范式转变,室内声学研究也不例外。大多数深层数据驱动的房间声学模型受到基于DL的语音和图像处理的启发,因此缺乏声波传播的固有时空结构。最近,DL为基础的模型,包括几何或波为基础的信息,室内声学提供了有前途的结果,主要是声场重建的问题。在这篇综述论文中,我们将提供一个广泛的和结构化的文献综述深入,数据驱动的室内声学建模。此外,我们将这些模型定位在一个框架中,该框架允许与传统的物理和数据驱动模型进行概念比较。最后,我们确定了深度,数据驱动的房间声学模型的优点和缺点,并概述了进一步研究的主要挑战。
摘要:Our everyday auditory experience is shaped by the acoustics of the indoor environments in which we live. Room acoustics modeling is aimed at establishing mathematical representations of acoustic wave propagation in such environments. These representations are relevant to a variety of problems ranging from echo-aided auditory indoor navigation to restoring speech understanding in cocktail party scenarios. Many disciplines in science and engineering have recently witnessed a paradigm shift powered by deep learning (DL), and room acoustics research is no exception. The majority of deep, data-driven room acoustics models are inspired by DL-based speech and image processing, and hence lack the intrinsic space-time structure of acoustic wave propagation. More recently, DL-based models for room acoustics that include either geometric or wave-based information have delivered promising results, primarily for the problem of sound field reconstruction. In this review paper, we will provide an extensive and structured literature review on deep, data-driven modeling in room acoustics. Moreover, we position these models in a framework that allows for a conceptual comparison with traditional physical and data-driven models. Finally, we identify strengths and shortcomings of deep, data-driven room acoustics models and outline the main challenges for further research.
【3】 Perceptual Audio Coding: A 40-Year Historical Perspective
标题: 感知音频编码:40年的历史视角链接:https://arxiv.org/abs/2504.16223
备注:None
摘要:在音频和声学信号处理的历史上,感知音频编码无疑因其在几乎所有数字媒体设备(例如计算机、平板电脑、手机、机顶盒和数字收音机)中的普遍部署而成为一个光明的成功故事。从技术的角度来看,感知音频编码已经经历了巨大的发展,从第一个非常基本的感知驱动的编码器(包括流行的mp3格式)到今天的成熟的集成编码/渲染系统。本文通过精确定位感知音频编码演变中的关键发展步骤,提供了这一研究历程的历史概述。最后,它提供了关于这一领域未来方向的想法。
摘要:In the history of audio and acoustic signal processing, perceptual audio coding has certainly excelled as a bright success story by its ubiquitous deployment in virtually all digital media devices, such as computers, tablets, mobile phones, set-top-boxes, and digital radios. From a technology perspective, perceptual audio coding has undergone tremendous development from the first very basic perceptually driven coders (including the popular mp3 format) to today's full-blown integrated coding/rendering systems. This paper provides a historical overview of this research journey by pinpointing the pivotal development steps in the evolution of perceptual audio coding. Finally, it provides thoughts about future directions in this area.
【4】 Using Phonemes in cascaded S2S translation pipeline
标题: 在级联S2 S翻译管道中使用音素链接:https://arxiv.org/abs/2504.16234
备注:Accepted at Swiss NLP Conference 2025
摘要:本文探讨了在传统的多语言同步语音到语音翻译管道中使用音素作为文本表示的想法,而不是传统的依赖于基于文本的语言表示。为了研究这一点,我们在WMT17数据集上以两种格式训练了一个开源的序列到序列模型:一种使用标准文本表示,另一种使用音素表示。这两种方法的性能使用BLEU度量进行评估。我们的研究结果表明,音素的方法提供了相当的质量,但提供了几个优势,包括较低的资源要求或更适合低资源的语言。
摘要:This paper explores the idea of using phonemes as a textual representation within a conventional multilingual simultaneous speech-to-speech translation pipeline, as opposed to the traditional reliance on text-based language representations. To investigate this, we trained an open-source sequence-to-sequence model on the WMT17 dataset in two formats: one using standard textual representation and the other employing phonemic representation. The performance of both approaches was assessed using the BLEU metric. Our findings shows that the phonemic approach provides comparable quality but offers several advantages, including lower resource requirements or better suitability for low-resource languages.
【5】 TinyML for Speech Recognition
标题: 用于语音识别的TinyML链接:https://arxiv.org/abs/2504.16213
摘要:我们训练并部署了一个量化的1D卷积神经网络模型,以在资源高度受限的物联网边缘设备上进行语音识别。这在各种物联网(IoT)应用中可能很有用,例如智能家居以及老年人和残疾人的环境辅助生活,仅举几个例子。在本文中,我们首先创建了一个新的数据集,其中包含一个多小时的音频数据,使我们的研究能够进行,并将对该领域的未来研究有用。其次,我们利用Edge Impulse提供的技术来增强模型的性能,并在我们的数据集上实现高达97%的高准确度。为了验证,我们使用Arduino Nano 33 BLE Sense微控制器板实现了我们的原型。该微控制器板专为物联网和人工智能应用而设计,使其成为我们目标用例场景的理想选择。虽然大多数现有的研究集中在有限的关键字集,我们的模型可以处理23个不同的关键字,使复杂的命令。
摘要:We train and deploy a quantized 1D convolutional neural network model to conduct speech recognition on a highly resource-constrained IoT edge device. This can be useful in various Internet of Things (IoT) applications, such as smart homes and ambient assisted living for the elderly and people with disabilities, just to name a few examples. In this paper, we first create a new dataset with over one hour of audio data that enables our research and will be useful to future studies in this field. Second, we utilize the technologies provided by Edge Impulse to enhance our model's performance and achieve a high Accuracy of up to 97% on our dataset. For the validation, we implement our prototype using the Arduino Nano 33 BLE Sense microcontroller board. This microcontroller board is specifically designed for IoT and AI applications, making it an ideal choice for our target use case scenarios. While most existing research focuses on a limited set of keywords, our model can process 23 different keywords, enabling complex commands.
机器翻译由腾讯交互翻译提供,仅供参考
