今日论文合集:cs.SD语音13篇,eess.AS音频处理15篇。本文经arXiv每日学术速递授权转载
【1】Deep Space Separable Distillation for Lightweight Acoustic Scene Classification
摘要:声场景分类在现实世界中具有重要意义。最近,基于深度学习的方法已被广泛用于声学场景分类。然而,这些方法目前不够轻量级,以及他们的性能不令人满意。为了解决这些问题,我们提出了一种深空可分离蒸馏网络。首先,网络对log-mel谱图进行高低频分解,在保持模型性能的同时显著降低计算复杂度。其次,我们专门为ASC设计了三个轻量级算子,包括可分离卷积(SC),正交可分离卷积(OSC)和可分离部分卷积(SPC)。这些运营商表现出高效的特征提取能力,在声学场景分类任务。实验结果表明,与目前流行的深度学习方法相比,该方法实现了9.8%的性能增益,同时具有更小的参数数量和计算复杂度。摘要:Acoustic scene classification (ASC) is highly important in the real world. Recently, deep learning-based methods have been widely employed for acoustic scene classification. However, these methods are currently not lightweight enough as well as their performance is not satisfactory. To solve these problems, we propose a deep space separable distillation network. Firstly, the network performs high-low frequency decomposition on the log-mel spectrogram, significantly reducing computational complexity while maintaining model performance. Secondly, we specially design three lightweight operators for ASC, including Separable Convolution (SC), Orthonormal Separable Convolution (OSC), and Separable Partial Convolution (SPC). These operators exhibit highly efficient feature extraction capabilities in acoustic scene classification tasks. The experimental results demonstrate that the proposed method achieves a performance gain of 9.8% compared to the currently popular deep learning methods, while also having smaller parameter count and computational complexity.【2】 Whispy: Adapting STT Whisper Models to Real-Time Environments标题:Whispy:将STT Whisper模型适应实时环境作者:Antonio Bevilacqua,Paolo Saviano,Alessandro Amirante,Simon Pietro Romano摘要:大型通用Transformer模型最近已成为语音分析领域的中流砥柱。特别是,Whisper在语音识别、翻译、语言识别和语音活动检测等相关任务中取得了最先进的成果。然而,Whisper模型并不是设计用于实时条件,这种限制使它们不适合大量的实际应用。在本文中,我们介绍了Whispy,一个旨在为Whisper预训练模型带来实时功能的系统。由于许多架构优化,Whispy能够消耗现场音频流并生成高水平,连贯的语音传输,同时仍然保持低计算成本。我们评估了我们的系统在一个大型的公开可用的语音数据集存储库的性能,调查如何转录机制引入Whispy的Whisper输出的影响。实验结果表明Whispy在鲁棒性,快速性和准确性方面表现出色。摘要:Large general-purpose transformer models have recently become the mainstay in the realm of speech analysis. In particular, Whisper achieves state-of-the-art results in relevant tasks such as speech recognition, translation, language identification, and voice activity detection. However, Whisper models are not designed to be used in real-time conditions, and this limitation makes them unsuitable for a vast plethora of practical applications. In this paper, we introduce Whispy, a system intended to bring live capabilities to the Whisper pretrained models. As a result of a number of architectural optimisations, Whispy is able to consume live audio streams and generate high level, coherent voice transcriptions, while still maintaining a low computational cost. We evaluate the performance of our system on a large repository of publicly available speech datasets, investigating how the transcription mechanism introduced by Whispy impacts on the Whisper output. Experimental results show how Whispy excels in robustness, promptness, and accuracy.
【3】 Fully Reversing the Shoebox Image Source Method: From Impulse Responses to Room Parameters标题:完全颠倒鞋盒图像源方法:从脉冲响应到房间参数作者:Tom Sprunck,Antoine Deleforge,Yannick Privat,Cédric Foy摘要:我们提出了一种算法,完全逆转鞋盒图像源方法(ISM),一个流行的和广泛使用的房间脉冲响应(RIR)模拟器介绍了艾伦和伯克利在1979年的长方体房间。更确切地说,给定由鞋盒ISM针对已知几何形状的麦克风阵列生成的离散多通道RIR,该算法可靠地恢复18个输入参数。这些是3D源位置、房间的3个维度、6个自由度的房间平移和定向以及6个房间边界中的每一个的吸收系数。该方法建立在最近提出的无网格图像源定位技术结合新的程序,房间轴恢复和一阶反射识别。大量的模拟实验表明,在2 × 2 × 2 ~ 10 × 10 × 5米的房间内,对于32单元、8.4 cm宽的球形麦克风阵列和16 kHz的采样率,使用完全随机化的输入参数,实现了所有参数的接近精确的恢复。估计误差衰减到零时,增加阵列的大小和采样率。该方法也被证明是强烈优于一个已知的基线,并证明其能力外推RIR在新的位置。至关重要的是,该方法严格限于使用香草鞋盒ISM模拟的低通离散RIR。尽管如此,据我们所知,它代表了第一个算法证明,这个困难的逆问题是在原则上完全可解的各种配置。摘要:We present an algorithm that fully reverses the shoebox image source method (ISM), a popular and widely used room impulse response (RIR) simulator for cuboid rooms introduced by Allen and Berkley in 1979. More precisely, given a discrete multichannel RIR generated by the shoebox ISM for a microphone array of known geometry, the algorithm reliably recovers the 18 input parameters. These are the 3D source position, the 3 dimensions of the room, the 6-degrees-of-freedom room translation and orientation, and an absorption coefficient for each of the 6 room boundaries. The approach builds on a recently proposed gridless image source localization technique combined with new procedures for room axes recovery and first-order-reflection identification. Extensive simulated experiments reveal that near-exact recovery of all parameters is achieved for a 32-element, 8.4-cm-wide spherical microphone array and a sampling rate of 16~kHz using fully randomized input parameters within rooms of size 2X2X2 to 10X10X5 meters. Estimation errors decay towards zero when increasing the array size and sampling rate. The method is also shown to strongly outperform a known baseline, and its ability to extrapolate RIRs at new positions is demonstrated. Crucially, the approach is strictly limited to low-passed discrete RIRs simulated using the vanilla shoebox ISM. Nonetheless, it represents to our knowledge the first algorithmic demonstration that this difficult inverse problem is in-principle fully solvable over a wide range of configurations.
【4】 Enhancing Aeroacoustic Wind Tunnel Studies through Massive Channel Upscaling with MEMS Microphones标题:通过使用微机电麦克风进行大规模通道升级来增强气动声学风洞研究作者:Daniel Ernst,Armin Goudarzi,Reinhard Geisler,Florian Philipp,Thomas Ahlefeldt,Carsten Spehr备注:30th AIAA/CEAS Aeroacoustics Conference摘要:本文介绍了一种大口径6m × 3m的7200 MEMS传声器阵列.该阵列的设计,使具有优化的点扩散函数的子阵列可以用于波束形成,从而使风洞设施中的源方向性的研究。整个阵列由800个模块化麦克风面板组成,每个面板由四个独特的PCB板设计组成。这种模块化架构允许对任意数量的面板进行时间同步测量,从而测量孔径大小和传感器总数。面板可以无间隙地安装,使得阵列的麦克风方向图避免点扩散函数中的高旁瓣。该阵列的能力是在DNW—NWB的开放式风洞中在1:9.5机身半模型上进行评估的。总的源发射的量化和方向性的评估与波束形成。额外的远场麦克风来验证结果。摘要:This paper presents a large 6~m x 3~m aperture 7200 MEMS microphone array. The array is designed so that sub-arrays with optimized point spread functions can be used for beamforming and thus, enable the research of source directivity in wind tunnel facilities. The total array consists of modular 800 microphone panels, each consisting of four unique PCB board designs. This modular architecture allows for the time-synchronized measurement of an arbitrary number of panels and thus, aperture size and total number of sensors. The panels can be installed without a gap so that the array's microphone pattern avoids high sidelobes in the point spread function. The array's capabilities are evaluated on a 1:9.5 airframe half model in an open wind tunnel at DNW-NWB. The total source emission is quantified and the directivity is evaluated with beamforming. Additional far-field microphones are employed to validate the results.
【5】 POPDG: Popular 3D Dance Generation with PopDanceSet标题:POPDG:PopDanceSet的流行3D舞蹈一代作者:Zhenye Luo,Min Ren,Xuecai Hu,Yongzhen Huang,Li Yao摘要:生成既逼真又与音乐保持一致的舞蹈仍然是跨模态领域的一项具有挑战性的任务。本文介绍了PopDanceSet,这是第一个为年轻观众的喜好量身定制的数据集,可以生成以美学为导向的舞蹈。它在音乐流派多样性和舞蹈动作的复杂性和深度方面超过了AIST++数据集。此外,在iDDPM框架内提出的POPDG模型增强了舞蹈的多样性,并通过空间增强算法加强了人体关节之间的空间物理连接,确保增加的多样性不会影响生成质量。一个流线型的对齐模块也被设计来改善舞蹈和音乐之间的时间对齐。大量的实验表明,POPDG在两个数据集上实现了SOTA结果。此外,本文还扩展了当前的评估指标。数据集和代码可在https://github.com/Luke-Luo1/POPDG上获得。摘要:Generating dances that are both lifelike and well-aligned with music continues to be a challenging task in the cross-modal domain. This paper introduces PopDanceSet, the first dataset tailored to the preferences of young audiences, enabling the generation of aesthetically oriented dances. And it surpasses the AIST++ dataset in music genre diversity and the intricacy and depth of dance movements. Moreover, the proposed POPDG model within the iDDPM framework enhances dance diversity and, through the Space Augmentation Algorithm, strengthens spatial physical connections between human body joints, ensuring that increased diversity does not compromise generation quality. A streamlined Alignment Module is also designed to improve the temporal alignment between dance and music. Extensive experiments show that POPDG achieves SOTA results on two datasets. Furthermore, the paper also expands on current evaluation metrics. The dataset and code are available at https://github.com/Luke-Luo1/POPDG.
【6】 Transhuman Ansambl - Voice Beyond Language标题:Transshuman Ansambl -超越语言的声音作者:Lucija Ivsic,Jon McCormack,Vince Dziekan摘要:在本文中,我们提出了设计和开发的transshuman Ansambl,一种新型的交互式唱歌的声音界面,它的感觉,它的环境和响应的声音输入与发声使用人类的声音。ansambl是为现场表演而设计的,它是一个独立的声音装置,由16个定制的虚拟歌手组成,排成一个圆圈。在现场表演时,虚拟歌手会倾听人类表演者的声音,并通过阅读音高、语调和音量提示来回应他们的歌声。在独立的声音安装模式下,歌手使用超声波距离传感器来感知观众的存在。作为第一作者的实践为基础的博士和艺术实践的一部分,作为一个现场表演者,这项工作采用唱歌的声音,以探索语音交互在HCI超越语言,和现场表演的创新方式。技术是如何支持通过声音产生的亲密效果的?用响应式虚拟歌手包围观众的行为是否挑战了表演者-听众的传统角色?为了回答这些问题,我们借鉴了第一作者的系统经验,以及语音研究的跨学科领域,认为语音是独立于语言的声音媒介,能够在身体之间建立相互联系。摘要:In this paper we present the design and development of the Transhuman Ansambl, a novel interactive singing-voice interface which senses its environment and responds to vocal input with vocalisations using human voice. Designed for live performance with a human performer and as a standalone sound installation, the ansambl consists of sixteen bespoke virtual singers arranged in a circle. When performing live, the virtual singers listen to the human performer and respond to their singing by reading pitch, intonation and volume cues. In a standalone sound installation mode, singers use ultrasonic distance sensors to sense audience presence. Developed as part of the 1st author's practice-based PhD and artistic practice as a live performer, this work employs the singing-voice to explore voice interactions in HCI beyond language, and innovative ways of live performing. How is technology supporting the effect of intimacy produced through voice? Does the act of surrounding the audience with responsive virtual singers challenge the traditional roles of performer-listener? To answer these questions, we draw upon the 1st author's experience with the system, and the interdisciplinary field of voice studies that consider the voice as the sound medium independent of language, capable of enacting a reciprocal connection between bodies.【7】 Determined Multichannel Blind Source Separation with Clustered Source Model作者:Jianyu Wang,Shanzheng Guan摘要:独立低秩矩阵分析(ILRMA)方法是一种重要的多通道盲音频源分离技术。它利用非负矩阵分解(NMF)和非负典型多元分解(NCPD)来模拟源参数。虽然它有效地捕捉低秩结构的来源,NMF模型忽略了通道间的依赖性。另一方面,NCPD保留了固有的结构,但缺乏可解释的潜在因素,使其具有挑战性,将先验信息作为约束。为了解决这些限制,我们引入了一个集群源模型的基础上非负块项分解(NBTD)。该模型将块定义为向量(聚类)和矩阵(用于光谱结构建模)的外积,提供可解释的潜在向量。此外,它能够直接集成的正交约束,以确保源图像之间的独立性。实验结果表明,我们提出的方法优于ILRMA及其扩展在消声条件下,并超过原来的ILRMA在模拟混响环境。摘要:The independent low-rank matrix analysis (ILRMA) method stands out as a prominent technique for multichannel blind audio source separation. It leverages nonnegative matrix factorization (NMF) and nonnegative canonical polyadic decomposition (NCPD) to model source parameters. While it effectively captures the low-rank structure of sources, the NMF model overlooks inter-channel dependencies. On the other hand, NCPD preserves intrinsic structure but lacks interpretable latent factors, making it challenging to incorporate prior information as constraints. To address these limitations, we introduce a clustered source model based on nonnegative block-term decomposition (NBTD). This model defines blocks as outer products of vectors (clusters) and matrices (for spectral structure modeling), offering interpretable latent vectors. Moreover, it enables straightforward integration of orthogonality constraints to ensure independence among source images. Experimental results demonstrate that our proposed method outperforms ILRMA and its extensions in anechoic conditions and surpasses the original ILRMA in simulated reverberant environments.
【8】 RepAugment: Input-Agnostic Representation-Level Augmentation for Respiratory Sound Classification作者:June-Woo Kim,Miika Toikkanen,Sangmin Bae,Minseok Kim,Ho-Young Jung摘要:人工智能的最新进展使其作为医疗助理的部署民主化。虽然来自大规模视觉和音频数据集的预训练模型已被证明适用于这一任务,但令人惊讶的是,没有研究探索过预训练的语音模型,这些语音模型作为人类起源的声音,直观上与肺部声音更相似。本文探讨了预训练的语音模型对呼吸声分类的有效性。我们发现,语音和肺音样本之间存在一个表征差距,为了弥合这一差距,数据增强是必不可少的。然而,用于音频和语音的最广泛使用的增强技术SpecAugment需要二维频谱图格式,并且不能应用于在语音波形上预训练的模型。为了解决这个问题,我们提出了RepAugment,这是一种与输入无关的表示级增强技术,其性能优于SpecAugment,但也适用于使用波形预训练模型的呼吸声分类。实验结果表明,我们的方法优于SpecAugment,在少数疾病类别的准确率方面有了实质性的提高,达到了7.14%。摘要:Recent advancements in AI have democratized its deployment as a healthcare assistant. While pretrained models from large-scale visual and audio datasets have demonstrably generalized to this task, surprisingly, no studies have explored pretrained speech models, which, as human-originated sounds, intuitively would share closer resemblance to lung sounds. This paper explores the efficacy of pretrained speech models for respiratory sound classification. We find that there is a characterization gap between speech and lung sound samples, and to bridge this gap, data augmentation is essential. However, the most widely used augmentation technique for audio and speech, SpecAugment, requires 2-dimensional spectrogram format and cannot be applied to models pretrained on speech waveforms. To address this, we propose RepAugment, an input-agnostic representation-level augmentation technique that outperforms SpecAugment, but is also suitable for respiratory sound classification with waveform pretrained models. Experimental results show that our approach outperforms the SpecAugment, demonstrating a substantial improvement in the accuracy of minority disease classes, reaching up to 7.14%.【9】 Steered Response Power for Sound Source Localization: A Tutorial Review作者:Eric Grinstein,Elisa Tengan,Bilgesu Çakmak,Thomas Dietzen,Leonardo Nunes,Toon van Waterschoot,Mike Brookes,Patrick A. Naylor摘要:在过去的三十年中,转向响应功率(SRP)方法已被广泛用于声源定位(SSL)的任务,由于其令人满意的定位性能在中度混响和噪声的情况下。许多工作已经分析和扩展了原来的SRP方法,以减少其计算成本,使其能够定位多个源,或提高其在恶劣环境中的性能。在这项工作中,我们回顾了200多篇关于SRP方法及其变体的论文,重点是SRP-PHAT方法。我们还提出了eXtensible-SRP,或X-SRP,一个广义和模块化的版本的SRP算法,它允许审查的扩展来实现。我们提供了一个Python实现的算法,其中包括从文献中选择的扩展。摘要:In the last three decades, the Steered Response Power (SRP) method has been widely used for the task of Sound Source Localization (SSL), due to its satisfactory localization performance on moderately reverberant and noisy scenarios. Many works have analyzed and extended the original SRP method to reduce its computational cost, to allow it to locate multiple sources, or to improve its performance in adverse environments. In this work, we review over 200 papers on the SRP method and its variants, with emphasis on the SRP-PHAT method. We also present eXtensible-SRP, or X-SRP, a generalized and modularized version of the SRP algorithm which allows the reviewed extensions to be implemented. We provide a Python implementation of the algorithm which includes selected extensions from the literature.
【10】 Sim2Real Transfer for Audio-Visual Navigation with Frequency-Adaptive Acoustic Field Prediction标题:具有频率自适应声学场预测的视听导航Sim 2 Real传输作者:Changan Chen,Jordi Ramos,Anshul Tomar,Kristen Grauman摘要:Sim2real传输最近受到越来越多的关注,因为它在端到端仿真中成功地学习了机器人任务。虽然在转移基于视觉的导航策略方面取得了很大进展,但现有的用于视听导航的sim2real策略凭经验执行数据增强,而不测量声学间隙。声音与光的不同之处在于它跨越更宽的频率,因此需要sim2real的不同解决方案。我们提出了第一个治疗的sim2real视听导航分解成声场预测(AFP)和航点导航。我们首先在SoundSpaces模拟器中验证我们的设计选择,并在Continuous AudioGoal导航基准测试中显示改进。然后,我们收集真实世界的数据来测量模拟和真实世界之间的频谱差异,通过训练AFP模型,只需要一个特定的频率子带作为输入。我们进一步提出了一种频率自适应策略,该策略基于测量的频谱差和接收音频的能量分布智能地选择最佳频带进行预测,从而提高了对真实数据的性能。最后,我们建立了一个真正的机器人平台,并表明转移的政策可以成功地导航到发声对象。这项工作展示了构建智能代理的潜力,这些智能代理可以完全从模拟中看到,听到和采取行动,并将它们转移到现实世界中。摘要:Sim2real transfer has received increasing attention lately due to the success of learning robotic tasks in simulation end-to-end. While there has been a lot of progress in transferring vision-based navigation policies, the existing sim2real strategy for audio-visual navigation performs data augmentation empirically without measuring the acoustic gap. The sound differs from light in that it spans across much wider frequencies and thus requires a different solution for sim2real. We propose the first treatment of sim2real for audio-visual navigation by disentangling it into acoustic field prediction (AFP) and waypoint navigation. We first validate our design choice in the SoundSpaces simulator and show improvement on the Continuous AudioGoal navigation benchmark. We then collect real-world data to measure the spectral difference between the simulation and the real world by training AFP models that only take a specific frequency subband as input. We further propose a frequency-adaptive strategy that intelligently selects the best frequency band for prediction based on both the measured spectral difference and the energy distribution of the received audio, which improves the performance on the real data. Lastly, we build a real robot platform and show that the transferred policy can successfully navigate to sounding objects. This work demonstrates the potential of building intelligent agents that can see, hear, and act entirely from simulation, and transferring them to the real world.【11】 Mozart's Touch: A Lightweight Multi-modal Music Generation Framework Based on Pre-Trained Large Models标题:莫扎特的触摸:基于预先训练的大型模型的轻量级多模式音乐生成框架作者:Tianze Xu,Jiajun Li,Xuesong Chen,Yinrui Yao,Shuchang Liu备注:7 pages, 2 figures, submitted to ACM MM 2024摘要:近年来,人工智能生成的内容(AIGC)取得了快速发展,促进了各行各业音乐、图像和其他艺术表达形式的生成。然而,目前对通用的多模态音乐生成模型的研究还很缺乏。为了填补这一空白,我们提出了一个多模态的音乐生成框架莫扎特的触摸。它可以生成与跨模态输入(如图像,视频和文本)对齐的音乐。莫扎特的触摸是由三个主要组件:多模态字幕模块,大语言模型(LLM)理解和桥接模块,和音乐生成模块。与传统方法不同,Mozart's Touch不需要训练或微调预先训练的模型,通过清晰,可解释的提示提供效率和透明度。我们还引入了“LLM桥”方法来解决不同模态描述性文本之间的异构表示问题。我们对所提出的模型进行了一系列客观和主观的评价,结果表明,我们的模型超越了当前最先进的模型的性能。我们的代码和示例可在https://github.com/WangTooNaive/MozartsTouch上获得摘要:In recent years, AI-Generated Content (AIGC) has witnessed rapid advancements, facilitating the generation of music, images, and other forms of artistic expression across various industries. However, researches on general multi-modal music generation model remain scarce. To fill this gap, we propose a multi-modal music generation framework Mozart's Touch. It could generate aligned music with the cross-modality inputs, such as images, videos and text. Mozart's Touch is composed of three main components: Multi-modal Captioning Module, Large Language Model (LLM) Understanding & Bridging Module, and Music Generation Module. Unlike traditional approaches, Mozart's Touch requires no training or fine-tuning pre-trained models, offering efficiency and transparency through clear, interpretable prompts. We also introduce "LLM-Bridge" method to resolve the heterogeneous representation problems between descriptive texts of different modalities. We conduct a series of objective and subjective evaluations on the proposed model, and results indicate that our model surpasses the performance of current state-of-the-art models. Our codes and examples is availble at: https://github.com/WangTooNaive/MozartsTouch
【12】 Quranic Audio Dataset: Crowdsourced and Labeled Recitation from Non-Arabic Speakers标题:《古兰经》音频数据集:非阿拉伯语使用者的众包和标签背诵作者:Raghad Salameh,Mohamad Al Mdfaa,Nursultan Askarbekuly,Manuel Mazzara摘要:本文讨论了学习背诵古兰经为非阿拉伯语的挑战。我们探索了众包一个仔细注释的古兰经数据集的可能性,在此基础上可以构建人工智能模型来简化学习过程。特别是,我们使用基于志愿者的众包类型,并实现众包API来收集音频资产。我们将API集成到一个名为NamazApp的现有移动应用程序中,以收集音频背诵。我们开发了一个名为“古兰经之声”的众包平台,用于注释收集的音频资产。因此,我们从超过11个非阿拉伯国家的1287名参与者中收集了大约7000个古兰经背诵,我们从六个类别的数据集中注释了1166个背诵。我们已经实现了0.77的人群准确度,注释者之间的评分者间一致性为0.63,算法分配的标签与专家判断之间的一致性为0.89。摘要:This paper addresses the challenge of learning to recite the Quran for non-Arabic speakers. We explore the possibility of crowdsourcing a carefully annotated Quranic dataset, on top of which AI models can be built to simplify the learning process. In particular, we use the volunteer-based crowdsourcing genre and implement a crowdsourcing API to gather audio assets. We integrated the API into an existing mobile application called NamazApp to collect audio recitations. We developed a crowdsourcing platform called Quran Voice for annotating the gathered audio assets. As a result, we have collected around 7000 Quranic recitations from a pool of 1287 participants across more than 11 non-Arabic countries, and we have annotated 1166 recitations from the dataset in six categories. We have achieved a crowd accuracy of 0.77, an inter-rater agreement of 0.63 between the annotators, and 0.89 between the labels assigned by the algorithm and the expert judgments.【13】 Speech Technology Services for Oral History Research作者:Christoph Draxler,Henk van den Heuvel,Arjan van Hessen,Pavel Ircing,Jan Lehečka备注:5 pages plus references, 3 figures摘要:口述历史是关于历史事件的目击者和评论者的口述资料。语音技术是一个重要的工具来处理这样的录音,以获得转录和进一步增强结构的口头帐户在这方面的贡献,我们解决了转录门户网站和网络服务与语音处理在BAS,语音解决方案在LINDAT开发,如何做到这一点自己与耳语,剩余的挑战,以及未来的发展。摘要:Oral history is about oral sources of witnesses and commentors on historical events. Speech technology is an important instrument to process such recordings in order to obtain transcription and further enhancements to structure the oral account In this contribution we address the transcription portal and the webservices associated with speech processing at BAS, speech solutions developed at LINDAT, how to do it yourself with Whisper, remaining challenges, and future developments.【1】 Automatic Assessment of Dysarthria Using Audio-visual Vowel Graph Attention Network作者:Xiaokang Liu,Xiaoxia Du,Juan Liu,Rongfeng Su,Manwa Lawrence Ng,Yumei Zhang,Yudong Yang,Shaofeng Zhao,Lan Wang,Nan Yan备注:10 pages, 6 figures, 7 tables摘要:Automatic assessment of dysarthria remains a highly challenging task due to high variability in acoustic signals and the limited data. Currently, research on the automatic assessment of dysarthria primarily focuses on two approaches: one that utilizes expert features combined with machine learning, and the other that employs data-driven deep learning methods to extract representations. Research has demonstrated that expert features are effective in representing pathological characteristics, while deep learning methods excel at uncovering latent features. Therefore, integrating the advantages of expert features and deep learning to construct a neural network architecture based on expert knowledge may be beneficial for interpretability and assessment performance. In this context, the present paper proposes a vowel graph attention network based on audio-visual information, which effectively integrates the strengths of expert knowledges and deep learning. Firstly, various features were combined as inputs, including knowledge based acoustical features and deep learning based pre-trained representations. Secondly, the graph network structure based on vowel space theory was designed, allowing for a deep exploration of spatial correlations among vowels. Finally, visual information was incorporated into the model to further enhance its robustness and generalizability. The method exhibited superior performance in regression experiments targeting Frenchay scores compared to existing approaches.【2】 MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition标题:MMGER:利用LLM进行多模式和多粒度生成式错误纠正,用于联合口音和语音识别作者:Bingshen Mu,Yangze Li,Qijie Shao,Kun Wei,Xucheng Wan,Naijun Zheng,Huan Zhou,Lei Xie摘要:尽管自动语音识别(ASR)取得了显着的进步,但当面临不利条件时,性能往往会下降。生成纠错(GER)利用大型语言模型(LLM)的卓越文本理解能力,在ASR纠错中提供令人印象深刻的性能,其中N—best假设为转录预测提供了有价值的信息。然而,GER遇到的挑战,如固定的N—best假设,声学信息的利用不足,和有限的特异性多口音的情况下。在本文中,我们探讨了GER在多口音场景中的应用。口音代表了标准发音规范的偏差,同时ASR和口音识别(AR)的多任务学习框架有效地解决了多口音场景,使其成为一个突出的解决方案。在这项工作中,我们提出了一个统一的ASR—AR GER模型,命名为MMGER,利用多模态校正和多粒度校正。采用多任务ASR—AR学习来提供动态1—最佳假设和口音嵌入。多模态校正通过将语音的声学特征与对应的字符级1最佳假设序列强制对齐来实现细粒度的帧级校正。多粒度校正通过在细粒度多模态校正之上结合常规1—最佳假设来补充全局语言信息,以实现粗粒度话语级校正。MMGER有效地缓解了GER的局限性,并为多口音场景定制了基于LLM的ASR纠错。在多口音普通话KeSpeech数据集上进行的实验证明了MMGER的有效性,与完善的标准基线相比,AR准确率相对提高了26.72%,ASR字符错误率相对降低了27.55%。摘要:Despite notable advancements in automatic speech recognition (ASR), performance tends to degrade when faced with adverse conditions. Generative error correction (GER) leverages the exceptional text comprehension capabilities of large language models (LLM), delivering impressive performance in ASR error correction, where N-best hypotheses provide valuable information for transcription prediction. However, GER encounters challenges such as fixed N-best hypotheses, insufficient utilization of acoustic information, and limited specificity to multi-accent scenarios. In this paper, we explore the application of GER in multi-accent scenarios. Accents represent deviations from standard pronunciation norms, and the multi-task learning framework for simultaneous ASR and accent recognition (AR) has effectively addressed the multi-accent scenarios, making it a prominent solution. In this work, we propose a unified ASR-AR GER model, named MMGER, leveraging multi-modal correction, and multi-granularity correction. Multi-task ASR-AR learning is employed to provide dynamic 1-best hypotheses and accent embeddings. Multi-modal correction accomplishes fine-grained frame-level correction by force-aligning the acoustic features of speech with the corresponding character-level 1-best hypothesis sequence. Multi-granularity correction supplements the global linguistic information by incorporating regular 1-best hypotheses atop fine-grained multi-modal correction to achieve coarse-grained utterance-level correction. MMGER effectively mitigates the limitations of GER and tailors LLM-based ASR error correction for the multi-accent scenarios. Experiments conducted on the multi-accent Mandarin KeSpeech dataset demonstrate the efficacy of MMGER, achieving a 26.72% relative improvement in AR accuracy and a 27.55% relative reduction in ASR character error rate, compared to a well-established standard baseline.
【3】 Deep Space Separable Distillation for Lightweight Acoustic Scene Classification摘要:声场景分类在现实世界中具有重要意义。最近,基于深度学习的方法已被广泛用于声学场景分类。然而,这些方法目前不够轻量级,以及他们的性能不令人满意。为了解决这些问题,我们提出了一种深空可分离蒸馏网络。首先,网络对log—mel谱图进行高低频分解,在保持模型性能的同时显著降低计算复杂度。其次,我们专门为ASC设计了三个轻量级算子,包括可分离卷积(SC),正交可分离卷积(OSC)和可分离部分卷积(SPC)。这些运营商表现出高效的特征提取能力,在声学场景分类任务。实验结果表明,与目前流行的深度学习方法相比,该方法实现了9.8%的性能增益,同时具有更小的参数数量和计算复杂度。摘要:Acoustic scene classification (ASC) is highly important in the real world. Recently, deep learning-based methods have been widely employed for acoustic scene classification. However, these methods are currently not lightweight enough as well as their performance is not satisfactory. To solve these problems, we propose a deep space separable distillation network. Firstly, the network performs high-low frequency decomposition on the log-mel spectrogram, significantly reducing computational complexity while maintaining model performance. Secondly, we specially design three lightweight operators for ASC, including Separable Convolution (SC), Orthonormal Separable Convolution (OSC), and Separable Partial Convolution (SPC). These operators exhibit highly efficient feature extraction capabilities in acoustic scene classification tasks. The experimental results demonstrate that the proposed method achieves a performance gain of 9.8% compared to the currently popular deep learning methods, while also having smaller parameter count and computational complexity.
【4】 Whispy: Adapting STT Whisper Models to Real-Time Environments标题:Whispy:将STT Whisper模型适应实时环境作者:Antonio Bevilacqua,Paolo Saviano,Alessandro Amirante,Simon Pietro Romano摘要:大型通用Transformer模型最近已成为语音分析领域的中流砥柱。特别是,Whisper在语音识别、翻译、语言识别和语音活动检测等相关任务中取得了最先进的成果。然而,Whisper模型并不是设计用于实时条件,这种限制使它们不适合大量的实际应用。在本文中,我们介绍了Whispy,一个旨在为Whisper预训练模型带来实时功能的系统。由于许多架构优化,Whispy能够消耗现场音频流并生成高水平,连贯的语音传输,同时仍然保持低计算成本。我们评估了我们的系统在一个大型的公开可用的语音数据集存储库的性能,调查如何转录机制引入Whispy的Whisper输出的影响。实验结果表明Whispy在鲁棒性,快速性和准确性方面表现出色。摘要:Large general-purpose transformer models have recently become the mainstay in the realm of speech analysis. In particular, Whisper achieves state-of-the-art results in relevant tasks such as speech recognition, translation, language identification, and voice activity detection. However, Whisper models are not designed to be used in real-time conditions, and this limitation makes them unsuitable for a vast plethora of practical applications. In this paper, we introduce Whispy, a system intended to bring live capabilities to the Whisper pretrained models. As a result of a number of architectural optimisations, Whispy is able to consume live audio streams and generate high level, coherent voice transcriptions, while still maintaining a low computational cost. We evaluate the performance of our system on a large repository of publicly available speech datasets, investigating how the transcription mechanism introduced by Whispy impacts on the Whisper output. Experimental results show how Whispy excels in robustness, promptness, and accuracy.【5】 Fully Reversing the Shoebox Image Source Method: From Impulse Responses to Room Parameters标题:完全颠倒鞋盒图像源方法:从脉冲响应到房间参数作者:Tom Sprunck,Antoine Deleforge,Yannick Privat,Cédric Foy摘要:我们提出了一种算法,完全逆转鞋盒图像源方法(ISM),一个流行的和广泛使用的房间脉冲响应(RIR)模拟器介绍了艾伦和伯克利在1979年的长方体房间。更确切地说,给定由鞋盒ISM针对已知几何形状的麦克风阵列生成的离散多通道RIR,该算法可靠地恢复18个输入参数。这些是3D源位置、房间的3个维度、6个自由度的房间平移和定向以及6个房间边界中的每一个的吸收系数。该方法建立在最近提出的无网格图像源定位技术结合新的程序,房间轴恢复和一阶反射识别。大量的模拟实验表明,在2 × 2 × 2 ~ 10 × 10 × 5米的房间内,对于32单元、8.4 cm宽的球形麦克风阵列和16 kHz的采样率,使用完全随机化的输入参数,实现了所有参数的接近精确的恢复。估计误差衰减到零时,增加阵列的大小和采样率。该方法也被证明是强烈优于一个已知的基线,并证明其能力外推RIR在新的位置。至关重要的是,该方法严格限于使用香草鞋盒ISM模拟的低通离散RIR。尽管如此,据我们所知,它代表了第一个算法证明,这个困难的逆问题是在原则上完全可解的各种配置。摘要:We present an algorithm that fully reverses the shoebox image source method (ISM), a popular and widely used room impulse response (RIR) simulator for cuboid rooms introduced by Allen and Berkley in 1979. More precisely, given a discrete multichannel RIR generated by the shoebox ISM for a microphone array of known geometry, the algorithm reliably recovers the 18 input parameters. These are the 3D source position, the 3 dimensions of the room, the 6-degrees-of-freedom room translation and orientation, and an absorption coefficient for each of the 6 room boundaries. The approach builds on a recently proposed gridless image source localization technique combined with new procedures for room axes recovery and first-order-reflection identification. Extensive simulated experiments reveal that near-exact recovery of all parameters is achieved for a 32-element, 8.4-cm-wide spherical microphone array and a sampling rate of 16~kHz using fully randomized input parameters within rooms of size 2X2X2 to 10X10X5 meters. Estimation errors decay towards zero when increasing the array size and sampling rate. The method is also shown to strongly outperform a known baseline, and its ability to extrapolate RIRs at new positions is demonstrated. Crucially, the approach is strictly limited to low-passed discrete RIRs simulated using the vanilla shoebox ISM. Nonetheless, it represents to our knowledge the first algorithmic demonstration that this difficult inverse problem is in-principle fully solvable over a wide range of configurations.【6】 Enhancing Aeroacoustic Wind Tunnel Studies through Massive Channel Upscaling with MEMS Microphones标题:通过使用微机电麦克风进行大规模通道升级来增强气动声学风洞研究作者:Daniel Ernst,Armin Goudarzi,Reinhard Geisler,Florian Philipp,Thomas Ahlefeldt,Carsten Spehr备注:30th AIAA/CEAS Aeroacoustics Conference摘要:本文介绍了一种大口径6 m × 3 m的7200 MEMS传声器阵列.该阵列的设计,使具有优化的点扩散函数的子阵列可以用于波束形成,从而使风洞设施中的源方向性的研究。整个阵列由800个模块化麦克风面板组成,每个面板由四个独特的PCB板设计组成。这种模块化架构允许对任意数量的面板进行时间同步测量,从而测量孔径大小和传感器总数。面板可以无间隙地安装,使得阵列的麦克风方向图避免点扩散函数中的高旁瓣。该阵列的能力是在DNW-NWB的开放式风洞中在1:9.5机身半模型上进行评估的。总的源发射的量化和方向性的评估与波束形成。额外的远场麦克风来验证结果。摘要:This paper presents a large 6~m x 3~m aperture 7200 MEMS microphone array. The array is designed so that sub-arrays with optimized point spread functions can be used for beamforming and thus, enable the research of source directivity in wind tunnel facilities. The total array consists of modular 800 microphone panels, each consisting of four unique PCB board designs. This modular architecture allows for the time-synchronized measurement of an arbitrary number of panels and thus, aperture size and total number of sensors. The panels can be installed without a gap so that the array's microphone pattern avoids high sidelobes in the point spread function. The array's capabilities are evaluated on a 1:9.5 airframe half model in an open wind tunnel at DNW-NWB. The total source emission is quantified and the directivity is evaluated with beamforming. Additional far-field microphones are employed to validate the results.
【7】 POPDG: Popular 3D Dance Generation with PopDanceSet标题:POPDG:PopDanceSet的流行3D舞蹈一代作者:Zhenye Luo,Min Ren,Xuecai Hu,Yongzhen Huang,Li Yao摘要:生成既逼真又与音乐保持一致的舞蹈仍然是跨模态领域的一项具有挑战性的任务。本文介绍了PopDanceSet,这是第一个为年轻观众的喜好量身定制的数据集,可以生成以美学为导向的舞蹈。它在音乐流派多样性和舞蹈动作的复杂性和深度方面超过了AIST++数据集。此外,在iDDPM框架内提出的POPDG模型增强了舞蹈的多样性,并通过空间增强算法加强了人体关节之间的空间物理连接,确保增加的多样性不会影响生成质量。一个流线型的对齐模块也被设计来改善舞蹈和音乐之间的时间对齐。大量的实验表明,POPDG在两个数据集上实现了SOTA结果。此外,本文还扩展了当前的评估指标。数据集和代码可在https://github.com/Luke-Luo1/POPDG上获得。摘要:Generating dances that are both lifelike and well-aligned with music continues to be a challenging task in the cross-modal domain. This paper introduces PopDanceSet, the first dataset tailored to the preferences of young audiences, enabling the generation of aesthetically oriented dances. And it surpasses the AIST++ dataset in music genre diversity and the intricacy and depth of dance movements. Moreover, the proposed POPDG model within the iDDPM framework enhances dance diversity and, through the Space Augmentation Algorithm, strengthens spatial physical connections between human body joints, ensuring that increased diversity does not compromise generation quality. A streamlined Alignment Module is also designed to improve the temporal alignment between dance and music. Extensive experiments show that POPDG achieves SOTA results on two datasets. Furthermore, the paper also expands on current evaluation metrics. The dataset and code are available at https://github.com/Luke-Luo1/POPDG.
【8】 Transhuman Ansambl - Voice Beyond Language标题:Transshuman Ansambl -超越语言的声音作者:Lucija Ivsic,Jon McCormack,Vince Dziekan摘要:在本文中,我们提出了设计和开发的transshuman Ansambl,一种新型的交互式唱歌的声音界面,它的感觉,它的环境和响应的声音输入与发声使用人类的声音。ansambl是为现场表演而设计的,它是一个独立的声音装置,由16个定制的虚拟歌手组成,排成一个圆圈。在现场表演时,虚拟歌手会倾听人类表演者的声音,并通过阅读音高、语调和音量提示来回应他们的歌声。在独立的声音安装模式下,歌手使用超声波距离传感器来感知观众的存在。作为第一作者的实践为基础的博士和艺术实践的一部分,作为一个现场表演者,这项工作采用唱歌的声音,以探索语音交互在HCI超越语言,和现场表演的创新方式。技术是如何支持通过声音产生的亲密效果的?用响应式虚拟歌手包围观众的行为是否挑战了表演者-听众的传统角色?为了回答这些问题,我们借鉴了第一作者的系统经验,以及语音研究的跨学科领域,认为语音是独立于语言的声音媒介,能够在身体之间建立相互联系。摘要:In this paper we present the design and development of the Transhuman Ansambl, a novel interactive singing-voice interface which senses its environment and responds to vocal input with vocalisations using human voice. Designed for live performance with a human performer and as a standalone sound installation, the ansambl consists of sixteen bespoke virtual singers arranged in a circle. When performing live, the virtual singers listen to the human performer and respond to their singing by reading pitch, intonation and volume cues. In a standalone sound installation mode, singers use ultrasonic distance sensors to sense audience presence. Developed as part of the 1st author's practice-based PhD and artistic practice as a live performer, this work employs the singing-voice to explore voice interactions in HCI beyond language, and innovative ways of live performing. How is technology supporting the effect of intimacy produced through voice? Does the act of surrounding the audience with responsive virtual singers challenge the traditional roles of performer-listener? To answer these questions, we draw upon the 1st author's experience with the system, and the interdisciplinary field of voice studies that consider the voice as the sound medium independent of language, capable of enacting a reciprocal connection between bodies.
【9】 Determined Multichannel Blind Source Separation with Clustered Source Model作者:Jianyu Wang,Shanzheng Guan摘要:独立低秩矩阵分析(ILRMA)方法是一种重要的多通道盲音频源分离技术。它利用非负矩阵分解(NMF)和非负典型多元分解(NCPD)来模拟源参数。虽然它有效地捕捉低秩结构的来源,NMF模型忽略了通道间的依赖性。另一方面,NCPD保留了固有的结构,但缺乏可解释的潜在因素,使其具有挑战性,将先验信息作为约束。为了解决这些限制,我们引入了一个集群源模型的基础上非负块项分解(NBTD)。该模型将块定义为向量(聚类)和矩阵(用于光谱结构建模)的外积,提供可解释的潜在向量。此外,它能够直接集成的正交约束,以确保源图像之间的独立性。实验结果表明,我们提出的方法优于ILRMA及其扩展在消声条件下,并超过原来的ILRMA在模拟混响环境。摘要:The independent low-rank matrix analysis (ILRMA) method stands out as a prominent technique for multichannel blind audio source separation. It leverages nonnegative matrix factorization (NMF) and nonnegative canonical polyadic decomposition (NCPD) to model source parameters. While it effectively captures the low-rank structure of sources, the NMF model overlooks inter-channel dependencies. On the other hand, NCPD preserves intrinsic structure but lacks interpretable latent factors, making it challenging to incorporate prior information as constraints. To address these limitations, we introduce a clustered source model based on nonnegative block-term decomposition (NBTD). This model defines blocks as outer products of vectors (clusters) and matrices (for spectral structure modeling), offering interpretable latent vectors. Moreover, it enables straightforward integration of orthogonality constraints to ensure independence among source images. Experimental results demonstrate that our proposed method outperforms ILRMA and its extensions in anechoic conditions and surpasses the original ILRMA in simulated reverberant environments.【10】 RepAugment: Input-Agnostic Representation-Level Augmentation for Respiratory Sound Classification作者:June-Woo Kim,Miika Toikkanen,Sangmin Bae,Minseok Kim,Ho-Young Jung摘要:人工智能的最新进展使其作为医疗助理的部署民主化。虽然来自大规模视觉和音频数据集的预训练模型已被证明适用于这一任务,但令人惊讶的是,没有研究探索过预训练的语音模型,这些语音模型作为人类起源的声音,直观上与肺部声音更相似。本文探讨了预训练的语音模型对呼吸声分类的有效性。我们发现,语音和肺音样本之间存在一个表征差距,为了弥合这一差距,数据增强是必不可少的。然而,用于音频和语音的最广泛使用的增强技术SpecAugment需要二维频谱图格式,并且不能应用于在语音波形上预训练的模型。为了解决这个问题,我们提出了RepAugment,这是一种与输入无关的表示级增强技术,其性能优于SpecAugment,但也适用于使用波形预训练模型的呼吸声分类。实验结果表明,我们的方法优于SpecAugment,在少数疾病类别的准确率方面有了实质性的提高,达到了7.14%。摘要:Recent advancements in AI have democratized its deployment as a healthcare assistant. While pretrained models from large-scale visual and audio datasets have demonstrably generalized to this task, surprisingly, no studies have explored pretrained speech models, which, as human-originated sounds, intuitively would share closer resemblance to lung sounds. This paper explores the efficacy of pretrained speech models for respiratory sound classification. We find that there is a characterization gap between speech and lung sound samples, and to bridge this gap, data augmentation is essential. However, the most widely used augmentation technique for audio and speech, SpecAugment, requires 2-dimensional spectrogram format and cannot be applied to models pretrained on speech waveforms. To address this, we propose RepAugment, an input-agnostic representation-level augmentation technique that outperforms SpecAugment, but is also suitable for respiratory sound classification with waveform pretrained models. Experimental results show that our approach outperforms the SpecAugment, demonstrating a substantial improvement in the accuracy of minority disease classes, reaching up to 7.14%.
【11】 Steered Response Power for Sound Source Localization: A Tutorial Review作者:Eric Grinstein,Elisa Tengan,Bilgesu Çakmak,Thomas Dietzen,Leonardo Nunes,Toon van Waterschoot,Mike Brookes,Patrick A. Naylor摘要:在过去的三十年中,转向响应功率(SRP)方法已被广泛用于声源定位(SSL)的任务,由于其令人满意的定位性能在中度混响和噪声的情况下。许多工作已经分析和扩展了原来的SRP方法,以减少其计算成本,使其能够定位多个源,或提高其在恶劣环境中的性能。在这项工作中,我们回顾了200多篇关于SRP方法及其变体的论文,重点是SRP-PHAT方法。我们还提出了eXtensible-SRP,或X-SRP,一个广义和模块化的版本的SRP算法,它允许审查的扩展来实现。我们提供了一个Python实现的算法,其中包括从文献中选择的扩展。摘要:In the last three decades, the Steered Response Power (SRP) method has been widely used for the task of Sound Source Localization (SSL), due to its satisfactory localization performance on moderately reverberant and noisy scenarios. Many works have analyzed and extended the original SRP method to reduce its computational cost, to allow it to locate multiple sources, or to improve its performance in adverse environments. In this work, we review over 200 papers on the SRP method and its variants, with emphasis on the SRP-PHAT method. We also present eXtensible-SRP, or X-SRP, a generalized and modularized version of the SRP algorithm which allows the reviewed extensions to be implemented. We provide a Python implementation of the algorithm which includes selected extensions from the literature.【12】 Sim2Real Transfer for Audio-Visual Navigation with Frequency-Adaptive Acoustic Field Prediction标题:具有频率自适应声学场预测的视听导航Sim 2 Real传输作者:Changan Chen,Jordi Ramos,Anshul Tomar,Kristen Grauman摘要:Sim2real传输最近受到越来越多的关注,因为它在端到端仿真中成功地学习了机器人任务。虽然在转移基于视觉的导航策略方面取得了很大进展,但现有的用于视听导航的sim2real策略根据经验执行数据增强,而不测量声学间隙。声音与光的不同之处在于它跨越更宽的频率,因此需要sim2real的不同解决方案。我们提出了第一个治疗的sim2real视听导航分解成声场预测(AFP)和航点导航。我们首先在SoundSpaces模拟器中验证我们的设计选择,并在Continuous AudioGoal导航基准测试中显示改进。然后,我们收集真实世界的数据来测量模拟和真实世界之间的频谱差异,通过训练AFP模型,只需要一个特定的频率子带作为输入。我们进一步提出了一种频率自适应策略,该策略基于测量的频谱差和接收音频的能量分布智能地选择最佳频带进行预测,从而提高了对真实数据的性能。最后,我们建立了一个真正的机器人平台,并表明转移的政策可以成功地导航到发声对象。这项工作展示了构建智能代理的潜力,这些智能代理可以完全从模拟中看到,听到和采取行动,并将它们转移到现实世界中。摘要:Sim2real transfer has received increasing attention lately due to the success of learning robotic tasks in simulation end-to-end. While there has been a lot of progress in transferring vision-based navigation policies, the existing sim2real strategy for audio-visual navigation performs data augmentation empirically without measuring the acoustic gap. The sound differs from light in that it spans across much wider frequencies and thus requires a different solution for sim2real. We propose the first treatment of sim2real for audio-visual navigation by disentangling it into acoustic field prediction (AFP) and waypoint navigation. We first validate our design choice in the SoundSpaces simulator and show improvement on the Continuous AudioGoal navigation benchmark. We then collect real-world data to measure the spectral difference between the simulation and the real world by training AFP models that only take a specific frequency subband as input. We further propose a frequency-adaptive strategy that intelligently selects the best frequency band for prediction based on both the measured spectral difference and the energy distribution of the received audio, which improves the performance on the real data. Lastly, we build a real robot platform and show that the transferred policy can successfully navigate to sounding objects. This work demonstrates the potential of building intelligent agents that can see, hear, and act entirely from simulation, and transferring them to the real world.【13】 Mozart's Touch: A Lightweight Multi-modal Music Generation Framework Based on Pre-Trained Large Models标题:莫扎特的触摸:基于预先训练的大型模型的轻量级多模式音乐生成框架作者:Tianze Xu,Jiajun Li,Xuesong Chen,Yinrui Yao,Shuchang Liu备注:7 pages, 2 figures, submitted to ACM MM 2024摘要:近年来,人工智能生成的内容(AIGC)取得了快速发展,促进了各行各业音乐、图像和其他艺术表达形式的生成。然而,目前对通用的多模态音乐生成模型的研究还很缺乏。为了填补这一空白,我们提出了一个多模态的音乐生成框架莫扎特的触摸。它可以生成与跨模态输入(如图像,视频和文本)对齐的音乐。莫扎特的触摸是由三个主要组件:多模态字幕模块,大语言模型(LLM)理解和桥接模块,和音乐生成模块。与传统方法不同,Mozart's Touch不需要训练或微调预先训练的模型,通过清晰,可解释的提示提供效率和透明度。我们还引入了“LLM桥”方法来解决不同模态描述性文本之间的异构表示问题。我们对所提出的模型进行了一系列客观和主观的评价,结果表明,我们的模型超越了当前最先进的模型的性能。我们的代码和示例可在https://github.com/WangTooNaive/MozartsTouch上获得摘要:In recent years, AI-Generated Content (AIGC) has witnessed rapid advancements, facilitating the generation of music, images, and other forms of artistic expression across various industries. However, researches on general multi-modal music generation model remain scarce. To fill this gap, we propose a multi-modal music generation framework Mozart's Touch. It could generate aligned music with the cross-modality inputs, such as images, videos and text. Mozart's Touch is composed of three main components: Multi-modal Captioning Module, Large Language Model (LLM) Understanding & Bridging Module, and Music Generation Module. Unlike traditional approaches, Mozart's Touch requires no training or fine-tuning pre-trained models, offering efficiency and transparency through clear, interpretable prompts. We also introduce "LLM-Bridge" method to resolve the heterogeneous representation problems between descriptive texts of different modalities. We conduct a series of objective and subjective evaluations on the proposed model, and results indicate that our model surpasses the performance of current state-of-the-art models. Our codes and examples is availble at: https://github.com/WangTooNaive/MozartsTouch
【14】 Quranic Audio Dataset: Crowdsourced and Labeled Recitation from Non-Arabic Speakers标题:《古兰经》音频数据集:非阿拉伯语使用者的众包和标签背诵作者:Raghad Salameh,Mohamad Al Mdfaa,Nursultan Askarbekuly,Manuel Mazzara摘要:本文讨论了学习背诵古兰经为非阿拉伯语的挑战。我们探索了众包一个仔细注释的古兰经数据集的可能性,在此基础上可以构建人工智能模型来简化学习过程。特别是,我们使用基于志愿者的众包类型,并实现众包API来收集音频资产。我们将API集成到一个名为NamazApp的现有移动应用程序中,以收集音频背诵。我们开发了一个名为“古兰经之声”的众包平台,用于注释收集的音频资产。因此,我们从超过11个非阿拉伯国家的1287名参与者中收集了大约7000个古兰经背诵,我们从六个类别的数据集中注释了1166个背诵。我们已经实现了0.77的人群准确度,注释者之间的评分者间一致性为0.63,算法分配的标签与专家判断之间的一致性为0.89。摘要:This paper addresses the challenge of learning to recite the Quran for non-Arabic speakers. We explore the possibility of crowdsourcing a carefully annotated Quranic dataset, on top of which AI models can be built to simplify the learning process. In particular, we use the volunteer-based crowdsourcing genre and implement a crowdsourcing API to gather audio assets. We integrated the API into an existing mobile application called NamazApp to collect audio recitations. We developed a crowdsourcing platform called Quran Voice for annotating the gathered audio assets. As a result, we have collected around 7000 Quranic recitations from a pool of 1287 participants across more than 11 non-Arabic countries, and we have annotated 1166 recitations from the dataset in six categories. We have achieved a crowd accuracy of 0.77, an inter-rater agreement of 0.63 between the annotators, and 0.89 between the labels assigned by the algorithm and the expert judgments.
【15】 Speech Technology Services for Oral History Research作者:Christoph Draxler,Henk van den Heuvel,Arjan van Hessen,Pavel Ircing,Jan Lehečka备注:5 pages plus references, 3 figures摘要:口述历史是关于历史事件的目击者和评论者的口述资料。语音技术是一个重要的工具来处理这样的录音,以获得转录和进一步增强结构的口头帐户在这方面的贡献,我们解决了转录门户网站和网络服务与语音处理在BAS,语音解决方案在LINDAT开发,如何做到这一点自己与耳语,剩余的挑战,以及未来的发展。摘要:Oral history is about oral sources of witnesses and commentors on historical events. Speech technology is an important instrument to process such recordings in order to obtain transcription and further enhancements to structure the oral account In this contribution we address the transcription portal and the webservices associated with speech processing at BAS, speech solutions developed at LINDAT, how to do it yourself with Whisper, remaining challenges, and future developments.