今日论文合集:cs.SD语音6篇,eess.AS音频处理8篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】Quality Over Quantity? LLM-Based Curation for a Data-Efficient  Audio-Video Foundation Model

标题: 质量重于数量?基于LLM的数据高效音频视频基础模型的治疗
链接:https://arxiv.org/abs/2503.09205
作者: Ali Vosoughi,  Dimitra Emmanouilidou,  Hannes Gamper
备注:5 pages, 5 figures, 3 tables
摘要:整合音频和视觉数据来训练多模态基础模型仍然具有挑战性。我们提出了音视频矢量对齐(AVVA),它通过基于大语言模型(LLM)的数据策展管道将视听(AV)场景内容对齐到不仅仅是时间同步。具体而言,AVVA在双编码器对比学习框架内使用Whisper(基于语音的音频基础模型)对音频进行评分并选择高质量的训练片段,并使用DINOv 2对视频进行评分。对AudioCaps,VALOR和VGGSound的评估表明,这种方法可以用更少的数据实现显着的准确性增益。例如,与ImageBind相比,AVVA在VGGSound上的音频到视频检索的top-1准确率提高了7.6%,尽管仅对192小时的仔细过滤数据进行了训练(与5800小时以上)。此外,一项消融研究强调,用数据量换取数据质量可以提高性能,AudioCaps、VALOR和VGGSound的前三名准确率分别比未策划的基线提高了47.8、48.4和58.0个百分点。虽然这些结果强调了AVVA的数据效率,但我们还讨论了LLM驱动的策展的开销,以及如何在更大的域中扩展或近似。总的来说,AVVA提供了一个可行的路径,以更强大的,无文本的视听学习,提高检索准确性。
摘要:Integrating audio and visual data for training multimodal foundational modelsremains challenging. We present Audio-Video Vector Alignment (AVVA), whichaligns audiovisual (AV) scene content beyond mere temporal synchronization viaa Large Language Model (LLM)-based data curation pipeline. Specifically, AVVAscores and selects high-quality training clips using Whisper (speech-basedaudio foundation model) for audio and DINOv2 for video within a dual-encodercontrastive learning framework. Evaluations on AudioCaps, VALOR, and VGGSounddemonstrate that this approach can achieve significant accuracy gains withsubstantially less curated data. For instance, AVVA yields a 7.6% improvementin top-1 accuracy for audio-to-video retrieval on VGGSound compared toImageBind, despite training on only 192 hours of carefully filtered data (vs.5800+ hours). Moreover, an ablation study highlights that trading data quantityfor data quality improves performance, yielding respective top-3 accuracyincreases of 47.8, 48.4, and 58.0 percentage points on AudioCaps, VALOR, andVGGSound over uncurated baselines. While these results underscore AVVA's dataefficiency, we also discuss the overhead of LLM-driven curation and how it maybe scaled or approximated in larger domains. Overall, AVVA provides a viablepath toward more robust, text-free audiovisual learning with improved retrievalaccuracy.

【2】 Zero to 16383 Through the Wire: Transmitting High- Resolution MIDI with  WebSockets and the Browser
标题: 通过电线从零到16383:使用WebSockets和浏览器传输高分辨率收件箱
链接:https://arxiv.org/abs/2503.09055
作者: Daniel McKemie
摘要:本文概述了如何利用Web API和Web技术将JavaScript中的数字数据转换为最高有效字节和最低有效字节组合,将数据作为双并发CC消息进行分段,使用WebSockets将其发送到多个端点,并将浏览器连接到其他音乐软件。这种方法允许用户通过14位的以太网消息传递来控制他们自己的本地应用程序,甚至是位于远程源上的应用程序。由于该技术利用了WebSocket,因此它不依赖于本地网络进行连接,并为世界上任何地方的远程软件控制和协作提供了可能性。虽然从Web控制音乐软件的选项并不缺乏,但Web API允许更简化的最终用户体验,因为它无缝地链接到核心OS API功能。本文将分享一个通过浏览器传输高分辨率数据并将其转换为控制电压数据以供模块化合成器使用的用例。
摘要:This paper outlines how to leverage the Web MIDI API and web technologies toconvert numerical data in JavaScript to Most Significant Byte and LeastSignificant Byte combos, stage the data as dual concurrent CC messages, useWebSockets to send it to multiple endpoints, and wire the browser to othermusic software. This method allows users to control their own nativeapplication via 14-bit MIDI messaging and even applications housed on a remotesource. Because the technology utilizes WebSockets, it is not reliant on localnetworks for connectivity and opens the possibilities of remote softwarecontrol and collaboration anywhere in the world. While no shortage of optionsexists for controlling music software from the web, the Web MIDI API allows fora more streamlined end user experience as it seamlessly links to core OS MIDIfunctionality. The paper will share a use case of transmitting high-resolutionMIDI through the browser and translating it to control voltage data for usewith a modular synthesizer.

【3】 Control Surfaces: Using the Commodore 64 and Analog Synthesizer to  Expand Musical Boundaries
标题: 控制面:使用Commodore 64和模拟合成器扩展音乐边界
链接:https://arxiv.org/abs/2503.09053
作者: Daniel McKemie
摘要:模拟-数字混合电子音乐系统曾经出于必要而存在,以便促进用于创作现场计算机音乐的灵活工作环境。随着计算能力随着更快的微处理器的发展而增加,对模拟声音产生的数字功能的需求减少,计算机变得更有能力处理这两项任务。鉴于这些系统的排他性和使用时间相对较短,几乎没有探讨过这种系统的可能性。Jose Vicente Asuar的工作最好地证明了这种系统的可访问性,但他从未得到任何机构的支持,以使他的机器得到广泛的关注。本文以他的方法为模型,使用一个CoreCore 64(或免费提供的操作系统仿真器)和模拟模块化硬件,旨在打造一个易于访问、负担得起、易于使用、具有教育意义和音乐丰富的系统。
摘要:Analog-digital hybrid electronic music systems once existed out of necessityin order to facilitate a flexible work environment for the creation of livecomputer music. As computational power increased with the development of fastermicroprocessors, the need for digital functionality with analog soundproduction decreased, with the computer becoming more capable of handling bothtasks. Given the exclusivity of these systems and the relatively short timethey were in use, the possibilities of such systems were hardly explored. Thework of Jos\'e Vicente Asuar best demonstrated a push for accessibility of suchsystems, but he never received the support of any institution in order to bringhis machine widespread attention. Modeled after his approach, using a Commodore64 (or freely available OS emulator) and analog modular hardware, this paperaims to fashion a system that is accessible, affordable, easy to use,educational, and musically rich in nature.

【4】 Learning Control of Neural Sound Effects Synthesis from Physically  Inspired Models
标题: 物理启发模型的神经声音效果合成的学习控制
链接:https://arxiv.org/abs/2503.08806
作者: Yisu Zong,  Joshua Reiss
备注:ICASSP 2025
摘要:音效模型设计通常使用具有完全控制能力的数字信号处理技术,但很难在有限的参数范围内实现真实感。近来,神经声音效果合成方法已经作为用于生成高质量和逼真的声音的有前途的方法出现,但是合成期望的声音的过程在控制方面造成困难。本文提出了一种由物理启发模型指导的实时神经合成模型,在继承物理启发模型的控制界面的同时,能够生成高质量的声音。我们展示了我们的模型在音质和控制方面的卓越性能。
摘要:Sound effects model design commonly uses digital signal processing techniqueswith full control ability, but it is difficult to achieve realism within alimited number of parameters. Recently, neural sound effects synthesis methodshave emerged as a promising approach for generating high-quality and realisticsounds, but the process of synthesizing the desired sound poses difficulties interms of control. This paper presents a real-time neural synthesis model guidedby a physically inspired model, enabling the generation of high-quality soundswhile inheriting the control interface of the physically inspired model. Weshowcase the superior performance of our model in terms of sound quality andcontrol.

【5】 Contextual Speech Extraction: Leveraging Textual History as an Implicit  Cue for Target Speech Extraction
标题: 上下文语音提取:利用文本历史作为目标语音提取的隐式线索
链接:https://arxiv.org/abs/2503.08798
作者: Minsu Kim,  Rodrigo Mira,  Honglie Chen,  Stavros Petridis,  Maja Pantic
备注:Accepted to ICASSP 2025
摘要:在本文中,我们研究了一种新的方法,目标语音提取(TSE),它完全依赖于文本上下文来提取目标语音。我们将此任务称为上下文语音提取(CSE)。与传统的TSE方法,依赖于预先录制的注册话语,视频的目标扬声器的脸,空间信息,或其他明确的线索,以确定目标流,我们提出的方法只需要几个回合的以前的对话(或独白)的历史。这种方法在移动消息环境中自然是可行的,在移动消息环境中,语音记录通常在文本对话之前,可以隐式地利用文本对话。我们提出了三个CSE模型,并分析了它们在三个数据集上的性能。通过我们的实验,我们证明,即使当模型完全依赖于对话历史,它可以达到90%以上的准确率,在识别正确的目标流,只有两个以前的对话轮。此外,我们表明,通过在训练过程中利用文本上下文和注册话语作为线索,我们进一步提高了模型的灵活性和有效性,使我们能够在推理过程中使用任何线索,或将两者结合起来以提高性能。示例和代码可在https://miraodasilva.github.io/cse-project-page上获得。
摘要:In this paper, we investigate a novel approach for Target Speech Extraction(TSE), which relies solely on textual context to extract the target speech. Werefer to this task as Contextual Speech Extraction (CSE). Unlike traditionalTSE methods that rely on pre-recorded enrollment utterances, video of thetarget speaker's face, spatial information, or other explicit cues to identifythe target stream, our proposed method requires only a few turns of previousdialogue (or monologue) history. This approach is naturally feasible in mobilemessaging environments where voice recordings are typically preceded by textualdialogue that can be leveraged implicitly. We present three CSE models andanalyze their performances on three datasets. Through our experiments, wedemonstrate that even when the model relies purely on dialogue history, it canachieve over 90 % accuracy in identifying the correct target stream with onlytwo previous dialogue turns. Furthermore, we show that by leveraging bothtextual context and enrollment utterances as cues during training, we furtherenhance our model's flexibility and effectiveness, allowing us to use eithercue during inference, or combine both for improved performance. Samples andcode available on https://miraodasilva.github.io/cse-project-page .

【6】 Performance Modeling for Correlation-based Neural Decoding of Auditory  Attention to Speech
标题: 基于相关性的语音听觉注意力神经解码的性能建模
链接:https://arxiv.org/abs/2503.09349
作者: Simon Geirnaert,  Jonas Vanthornhout,  Tom Francart,  Alexander Bertrand
摘要:基于相关性的听觉注意力解码(AAD)算法利用神经跟踪机制来确定竞争语音源之间的收听者注意力,例如,脑电图信号解码的神经反应和编码的语音刺激的不同的扬声器之间的相关系数,然后作为AAD决策变量。时间分辨率(用于计算这些相关性的决策窗口长度)和AAD准确性之间存在关键的权衡。这种权衡的典型特征在于评估跨多个窗口长度的AAD精度,从而得出性能曲线。我们提出了一种新的方法来模拟这种权衡曲线,使用标记的相关性,只有一个单一的决策窗口长度。我们的方法在应用Fisher变换后对(未)参与的相关性进行正态分布建模,从而在不同的窗口长度上实现准确的AAD准确度预测。我们在两个不同的AAD实现上验证了该方法:线性解码器和非线性VLAAI深度神经网络,在不同的数据集上进行评估。结果显示,建模误差始终较低,约为2%,94%的真实准确度在估计的95%置信区间内。所提出的方法能够实现高效的性能曲线建模,而无需大量的多窗口长度评估,从而促进实际应用,例如,在神经操纵的听力设备中的性能跟踪,以随着时间的推移连续地适应系统参数。
摘要:Correlation-based auditory attention decoding (AAD) algorithms exploit neuraltracking mechanisms to determine listener attention among competing speechsources via, e.g., electroencephalography signals. The correlation coefficientsbetween the decoded neural responses and encoded speech stimuli of thedifferent speakers then serve as AAD decision variables. A critical trade-offexists between the temporal resolution (the decision window length used tocompute these correlations) and the AAD accuracy. This trade-off is typicallycharacterized by evaluating AAD accuracy across multiple window lengths,leading to the performance curve. We propose a novel method to model thistrade-off curve using labeled correlations from only a single decision windowlength. Our approach models the (un)attended correlations with a normaldistribution after applying the Fisher transformation, enabling accurate AADaccuracy prediction across different window lengths. We validate the method ontwo distinct AAD implementations: a linear decoder and the non-linear VLAAIdeep neural network, evaluated on separate datasets. Results show consistentlylow modeling errors of approximately 2 percent points, with 94% of trueaccuracies falling within estimated 95%-confidence intervals. The proposedmethod enables efficient performance curve modeling without extensivemulti-window length evaluation, facilitating practical applications in, e.g.,performance tracking in neuro-steered hearing devices to continuously adapt thesystem parameters over time.

eess.AS音频处理
【1】 Multiple Speaker Separation from Noisy Sources in Reverberant Rooms  using Relative Transfer Matrix
标题: 使用相对传递矩阵分离回响室中的多说话人与噪音源
链接:https://arxiv.org/abs/2503.09412
作者: Wageesha N. Manamperi,  Thushara D. Abhayapala
备注:5 pages, 4 figures, submitted to 33th European Signal Processing Conference (EUSIPCO) 2025
摘要:在强混响和多背景噪声源的环境中,同时激活的多个说话人的分离是一项困难的任务。本文使用的相对传递矩阵(ReTM),一个房间的相对传递函数的推广,提出了一种简单而新颖的方法分离并发扬声器使用嘈杂的多通道麦克风录音。所提出的方法(i)允许多个语音和背景噪声源,(ii)包括混响,(iii)不需要语音和噪声源的位置的知识,也不需要麦克风位置和它们的相对几何形状,以及(iv)使用相对小的记录段进行训练。我们说明了语音源分离能力,提高了清晰度使用的模拟研究,包括四个扬声器在混响室中存在三个噪声源。我们还显示了该方法在一个实际的实验中,在一个真实的房间的适用性。
摘要:Separation of simultaneously active multiple speakers is a difficult task inenvironments with strong reverberation and many background noise sources. Thispaper uses the relative transfer matrix (ReTM), a generalization of therelative transfer function of a room, to propose a simple yet novel approachfor separating concurrent speakers using noisy multichannel microphonerecordings. The proposed method (i) allows multiple speech and background noisesources, (ii) includes reverberation, (iii) does not need the knowledge of thelocations of speech and noise sources nor microphone locations and theirrelative geometry, and (iv) uses relatively small segment of recordings fortraining. We illustrate the speech source separation capability with improvedintelligibility using a simulation study consisting of four speakers in thepresence of three noise sources in a reverberant room. We also show theapplicability of the method in a practical experiment in a real room.

【2】 Performance Modeling for Correlation-based Neural Decoding of Auditory  Attention to Speech
标题: 基于相关性的语音听觉注意力神经解码的性能建模
链接:https://arxiv.org/abs/2503.09349
作者: Simon Geirnaert,  Jonas Vanthornhout,  Tom Francart,  Alexander Bertrand
摘要:基于相关性的听觉注意力解码(AAD)算法利用神经跟踪机制来确定竞争语音源之间的收听者注意力,例如,脑电图信号解码的神经反应和编码的语音刺激的不同的扬声器之间的相关系数,然后作为AAD决策变量。时间分辨率(用于计算这些相关性的决策窗口长度)和AAD精度之间存在关键的权衡。这种权衡的典型特征在于评估跨多个窗口长度的AAD精度,从而得出性能曲线。我们提出了一种新的方法来模拟这种权衡曲线,使用标记的相关性,只有一个单一的决策窗口长度。我们的方法在应用Fisher变换后对(未)参与的相关性进行正态分布建模,从而在不同的窗口长度上实现准确的AAD准确度预测。我们在两个不同的AAD实现上验证了该方法:线性解码器和非线性VLAAI深度神经网络,在不同的数据集上进行评估。结果显示,建模误差始终较低,约为2%,94%的真实准确度在估计的95%置信区间内。所提出的方法能够实现高效的性能曲线建模,无需大量的多窗口长度评估,从而促进了实际应用,例如,在神经操纵的听力设备中的性能跟踪,以随着时间的推移连续地适应系统参数。
摘要:Correlation-based auditory attention decoding (AAD) algorithms exploit neuraltracking mechanisms to determine listener attention among competing speechsources via, e.g., electroencephalography signals. The correlation coefficientsbetween the decoded neural responses and encoded speech stimuli of thedifferent speakers then serve as AAD decision variables. A critical trade-offexists between the temporal resolution (the decision window length used tocompute these correlations) and the AAD accuracy. This trade-off is typicallycharacterized by evaluating AAD accuracy across multiple window lengths,leading to the performance curve. We propose a novel method to model thistrade-off curve using labeled correlations from only a single decision windowlength. Our approach models the (un)attended correlations with a normaldistribution after applying the Fisher transformation, enabling accurate AADaccuracy prediction across different window lengths. We validate the method ontwo distinct AAD implementations: a linear decoder and the non-linear VLAAIdeep neural network, evaluated on separate datasets. Results show consistentlylow modeling errors of approximately 2 percent points, with 94% of trueaccuracies falling within estimated 95%-confidence intervals. The proposedmethod enables efficient performance curve modeling without extensivemulti-window length evaluation, facilitating practical applications in, e.g.,performance tracking in neuro-steered hearing devices to continuously adapt thesystem parameters over time.

【3】 An Exhaustive Evaluation of TTS- and VC-based Data Augmentation for ASR
标题: 基于TTC和VC的ASB数据增强的详尽评估
链接:https://arxiv.org/abs/2503.08954
作者: Sewade Ogun,  Vincent Colotte,  Emmanuel Vincent
摘要:近年来,利用文本到语音(TTS)或语音转换(VC)生成的合成数据来增强自动语音识别(ASR)系统的训练数据已经变得流行。一些作品已经证明了使用这种增强方法的ASR性能的改善。然而,由于合成语音的多样性较低,天真地将合成数据和真实数据结合起来通常不会产生最佳结果。在这项工作中,我们利用最近提出的基于流的TTS/VC模型,允许更大的语音多样性,并评估各自的影响,增强各种语音属性的单词错误率(WER)实现了几个ASR模型。音高增强和基于VC的扬声器增强被发现是无效的,在我们的设置。与仅在真实数据上进行训练相比,联合增强所有其他属性将Conformer-Transducer模型的WER相对于Common Voice降低了11%,相对于LibriSpeech降低了高达35%。
摘要:Augmenting the training data of automatic speech recognition (ASR) systemswith synthetic data generated by text-to-speech (TTS) or voice conversion (VC)has gained popularity in recent years. Several works have demonstratedimprovements in ASR performance using this augmentation approach. However,because of the lower diversity of synthetic speech, naively combining syntheticand real data often does not yield the best results. In this work, we leveragerecently proposed flow-based TTS/VC models allowing greater speech diversity,and assess the respective impact of augmenting various speech attributes on theword error rate (WER) achieved by several ASR models. Pitch augmentation andVC-based speaker augmentation are found to be ineffective in our setup. Jointlyaugmenting all other attributes reduces the WER of a Conformer-Transducer modelby 11\% relative on Common Voice and by up to 35\% relative on LibriSpeechcompared to training on real data only.

【4】 Quality Over Quantity? LLM-Based Curation for a Data-Efficient  Audio-Video Foundation Model
标题: 质量重于数量?基于LLM的数据高效音频视频基础模型的治疗
链接:https://arxiv.org/abs/2503.09205
作者: Ali Vosoughi,  Dimitra Emmanouilidou,  Hannes Gamper
备注:5 pages, 5 figures, 3 tables
摘要:整合音频和视觉数据来训练多模态基础模型仍然具有挑战性。我们提出了音视频矢量对齐(AVVA),它通过基于大语言模型(LLM)的数据策展管道将视听(AV)场景内容对齐到不仅仅是时间同步。具体而言,AVVA在双编码器对比学习框架内使用Whisper(基于语音的音频基础模型)对音频进行评分并选择高质量的训练片段,并使用DINOv 2对视频进行评分。对AudioCaps,VALOR和VGGSound的评估表明,这种方法可以用更少的数据实现显着的准确性增益。例如,与ImageBind相比,AVVA在VGGSound上的音频到视频检索的top-1准确率提高了7.6%,尽管仅对192小时的仔细过滤数据进行了训练(与5800小时以上)。此外,一项消融研究强调,用数据量换取数据质量可以提高性能,AudioCaps、VALOR和VGGSound的前三名准确率分别比未策划的基线提高了47.8、48.4和58.0个百分点。虽然这些结果强调了AVVA的数据效率,但我们还讨论了LLM驱动的策展的开销,以及如何在更大的域中扩展或近似。总的来说,AVVA提供了一个可行的路径,以更强大的,无文本的视听学习,提高检索准确性。
摘要:Integrating audio and visual data for training multimodal foundational modelsremains challenging. We present Audio-Video Vector Alignment (AVVA), whichaligns audiovisual (AV) scene content beyond mere temporal synchronization viaa Large Language Model (LLM)-based data curation pipeline. Specifically, AVVAscores and selects high-quality training clips using Whisper (speech-basedaudio foundation model) for audio and DINOv2 for video within a dual-encodercontrastive learning framework. Evaluations on AudioCaps, VALOR, and VGGSounddemonstrate that this approach can achieve significant accuracy gains withsubstantially less curated data. For instance, AVVA yields a 7.6% improvementin top-1 accuracy for audio-to-video retrieval on VGGSound compared toImageBind, despite training on only 192 hours of carefully filtered data (vs.5800+ hours). Moreover, an ablation study highlights that trading data quantityfor data quality improves performance, yielding respective top-3 accuracyincreases of 47.8, 48.4, and 58.0 percentage points on AudioCaps, VALOR, andVGGSound over uncurated baselines. While these results underscore AVVA's dataefficiency, we also discuss the overhead of LLM-driven curation and how it maybe scaled or approximated in larger domains. Overall, AVVA provides a viablepath toward more robust, text-free audiovisual learning with improved retrievalaccuracy.

【5】 Zero to 16383 Through the Wire: Transmitting High- Resolution MIDI with  WebSockets and the Browser
标题: 通过电线从零到16383:使用WebSockets和浏览器传输高分辨率收件箱
链接:https://arxiv.org/abs/2503.09055
作者: Daniel McKemie
摘要:本文概述了如何利用Web API和Web技术将JavaScript中的数字数据转换为最高有效字节和最低有效字节组合,将数据作为双并发CC消息进行分段,使用WebSockets将其发送到多个端点,并将浏览器连接到其他音乐软件。这种方法允许用户通过14位的以太网消息传递来控制他们自己的本地应用程序,甚至是位于远程源上的应用程序。由于该技术利用了WebSocket,因此它不依赖于本地网络进行连接,并为世界上任何地方的远程软件控制和协作提供了可能性。虽然从Web控制音乐软件的选项并不缺乏,但Web API允许更简化的最终用户体验,因为它无缝地链接到核心OS API功能。本文将分享一个通过浏览器传输高分辨率数据并将其转换为控制电压数据以供模块化合成器使用的用例。
摘要:This paper outlines how to leverage the Web MIDI API and web technologies toconvert numerical data in JavaScript to Most Significant Byte and LeastSignificant Byte combos, stage the data as dual concurrent CC messages, useWebSockets to send it to multiple endpoints, and wire the browser to othermusic software. This method allows users to control their own nativeapplication via 14-bit MIDI messaging and even applications housed on a remotesource. Because the technology utilizes WebSockets, it is not reliant on localnetworks for connectivity and opens the possibilities of remote softwarecontrol and collaboration anywhere in the world. While no shortage of optionsexists for controlling music software from the web, the Web MIDI API allows fora more streamlined end user experience as it seamlessly links to core OS MIDIfunctionality. The paper will share a use case of transmitting high-resolutionMIDI through the browser and translating it to control voltage data for usewith a modular synthesizer.

【6】 Control Surfaces: Using the Commodore 64 and Analog Synthesizer to  Expand Musical Boundaries
标题: 控制面:使用Commodore 64和模拟合成器扩展音乐边界
链接:https://arxiv.org/abs/2503.09053
作者: Daniel McKemie
摘要:模拟-数字混合电子音乐系统曾经出于必要而存在,以便促进用于创作现场计算机音乐的灵活工作环境。随着计算能力随着更快的微处理器的发展而增加,对模拟声音产生的数字功能的需求减少,计算机变得更有能力处理这两项任务。鉴于这些系统的排他性和使用时间相对较短,几乎没有探讨过这种系统的可能性。Jose Vicente Asuar的工作最好地证明了这种系统的可访问性,但他从未得到任何机构的支持,以使他的机器得到广泛的关注。本文以他的方法为模型,使用一个CoreCore 64(或免费提供的操作系统仿真器)和模拟模块化硬件,旨在打造一个易于访问、负担得起、易于使用、具有教育意义和音乐丰富的系统。
摘要:Analog-digital hybrid electronic music systems once existed out of necessityin order to facilitate a flexible work environment for the creation of livecomputer music. As computational power increased with the development of fastermicroprocessors, the need for digital functionality with analog soundproduction decreased, with the computer becoming more capable of handling bothtasks. Given the exclusivity of these systems and the relatively short timethey were in use, the possibilities of such systems were hardly explored. Thework of Jos\'e Vicente Asuar best demonstrated a push for accessibility of suchsystems, but he never received the support of any institution in order to bringhis machine widespread attention. Modeled after his approach, using a Commodore64 (or freely available OS emulator) and analog modular hardware, this paperaims to fashion a system that is accessible, affordable, easy to use,educational, and musically rich in nature.

【7】 Learning Control of Neural Sound Effects Synthesis from Physically  Inspired Models
标题: 物理启发模型的神经声音效果合成的学习控制
链接:https://arxiv.org/abs/2503.08806
作者: Yisu Zong,  Joshua Reiss
备注:ICASSP 2025
摘要:音效模型设计通常使用具有完全控制能力的数字信号处理技术,但很难在有限的参数范围内实现真实感。近来,神经声音效果合成方法已经作为用于生成高质量和逼真的声音的有前途的方法出现,但是合成期望的声音的过程在控制方面造成困难。本文提出了一种由物理启发模型指导的实时神经合成模型,在继承物理启发模型的控制界面的同时,能够生成高质量的声音。我们展示了我们的模型在音质和控制方面的卓越性能。
摘要:Sound effects model design commonly uses digital signal processing techniqueswith full control ability, but it is difficult to achieve realism within alimited number of parameters. Recently, neural sound effects synthesis methodshave emerged as a promising approach for generating high-quality and realisticsounds, but the process of synthesizing the desired sound poses difficulties interms of control. This paper presents a real-time neural synthesis model guidedby a physically inspired model, enabling the generation of high-quality soundswhile inheriting the control interface of the physically inspired model. Weshowcase the superior performance of our model in terms of sound quality andcontrol.

【8】 Contextual Speech Extraction: Leveraging Textual History as an Implicit  Cue for Target Speech Extraction
标题: 上下文语音提取:利用文本历史作为目标语音提取的隐式线索
链接:https://arxiv.org/abs/2503.08798
作者: Minsu Kim,  Rodrigo Mira,  Honglie Chen,  Stavros Petridis,  Maja Pantic
备注:Accepted to ICASSP 2025
摘要:在本文中,我们研究了一种新的方法,目标语音提取(TSE),它完全依赖于文本上下文来提取目标语音。我们将此任务称为上下文语音提取(CSE)。与传统的TSE方法,依赖于预先录制的注册话语,视频的目标扬声器的脸,空间信息,或其他明确的线索,以确定目标流,我们提出的方法只需要几个回合的以前的对话(或独白)的历史。这种方法在移动消息环境中自然是可行的,在移动消息环境中,语音记录通常在文本对话之前,可以隐式地利用文本对话。我们提出了三个CSE模型,并分析了它们在三个数据集上的性能。通过我们的实验,我们证明,即使当模型完全依赖于对话历史,它可以达到90%以上的准确率,在识别正确的目标流,只有两个以前的对话轮。此外,我们表明,通过在训练过程中利用文本上下文和注册话语作为线索,我们进一步提高了模型的灵活性和有效性,使我们能够在推理过程中使用任何线索,或将两者结合起来以提高性能。示例和代码可在https://miraodasilva.github.io/cse-project-page上获得。
摘要:In this paper, we investigate a novel approach for Target Speech Extraction(TSE), which relies solely on textual context to extract the target speech. Werefer to this task as Contextual Speech Extraction (CSE). Unlike traditionalTSE methods that rely on pre-recorded enrollment utterances, video of thetarget speaker's face, spatial information, or other explicit cues to identifythe target stream, our proposed method requires only a few turns of previousdialogue (or monologue) history. This approach is naturally feasible in mobilemessaging environments where voice recordings are typically preceded by textualdialogue that can be leveraged implicitly. We present three CSE models andanalyze their performances on three datasets. Through our experiments, wedemonstrate that even when the model relies purely on dialogue history, it canachieve over 90 % accuracy in identifying the correct target stream with onlytwo previous dialogue turns. Furthermore, we show that by leveraging bothtextual context and enrollment utterances as cues during training, we furtherenhance our model's flexibility and effectiveness, allowing us to use eithercue during inference, or combine both for improved performance. Samples andcode available on https://miraodasilva.github.io/cse-project-page .

机器翻译由腾讯交互翻译提供,仅供参考