今日论文合集:cs.SD语音8篇,eess.AS音频处理10篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation  via Fine-grained AI Feedback
标题: T2 A反馈:通过细粒度人工智能反馈提高文本到音频生成的基本能力
链接:https://arxiv.org/abs/2505.10561
作者: Zehan Wang,  Ke Lei,  Chen Zhu,  Jiawei Huang,  Sashuai Zhou,  Luping Liu,  Xize Cheng,  Shengpeng Ji,  Zhenhui Ye,  Tao Jin,  Zhou Zhao 
备注:ACL 2025
摘要:文本到音频(T2 A)生成在从语言提示生成各种音频输出方面取得了显著进展。然而,当前最先进的T2 A模型在生成复杂的多事件音频时仍然难以满足人类对跟随性和声学质量的偏好。为了提高模型在这些高级应用中的性能,我们建议通过AI反馈学习来增强模型的基本功能。首先,我们引入了细粒度的AI音频评分管道:1)验证文本提示中的每个事件是否存在于音频中(事件发生分数),2)检测事件序列与语言描述的偏差(事件序列分数),以及3)评估生成的音频的整体声学和谐波质量(声学和谐波质量)。我们评估了这三个自动评分管道,发现它们与人类偏好的相关性明显优于其他评估指标。这突出了它们作为反馈信号和评估指标的价值。利用我们强大的评分管道,我们构建了一个大型音频偏好数据集T2 A-FeedBack,其中包含41 k提示和249 k音频,每个都附有详细的分数。此外,我们还介绍了T2 A-EpicBench,这是一个专注于长字幕,多事件和讲故事场景的基准测试,旨在评估T2 A模型的高级功能。最后,我们展示了T2 A-FeedBack如何增强当前最先进的音频模型。通过简单的偏好调整,音频生成模型在简单(AudioCaps测试集)和复杂(T2 A-EpicBench)场景中都表现出显着的改进。
摘要:Text-to-audio (T2A) generation has achieved remarkable progress in generating a variety of audio outputs from language prompts. However, current state-of-the-art T2A models still struggle to satisfy human preferences for prompt-following and acoustic quality when generating complex multi-event audio. To improve the performance of the model in these high-level applications, we propose to enhance the basic capabilities of the model with AI feedback learning. First, we introduce fine-grained AI audio scoring pipelines to: 1) verify whether each event in the text prompt is present in the audio (Event Occurrence Score), 2) detect deviations in event sequences from the language description (Event Sequence Score), and 3) assess the overall acoustic and harmonic quality of the generated audio (Acoustic&Harmonic Quality). We evaluate these three automatic scoring pipelines and find that they correlate significantly better with human preferences than other evaluation metrics. This highlights their value as both feedback signals and evaluation metrics. Utilizing our robust scoring pipelines, we construct a large audio preference dataset, T2A-FeedBack, which contains 41k prompts and 249k audios, each accompanied by detailed scores. Moreover, we introduce T2A-EpicBench, a benchmark that focuses on long captions, multi-events, and story-telling scenarios, aiming to evaluate the advanced capabilities of T2A models. Finally, we demonstrate how T2A-FeedBack can enhance current state-of-the-art audio model. With simple preference tuning, the audio generation model exhibits significant improvements in both simple (AudioCaps test set) and complex (T2A-EpicBench) scenarios.


【2】 Learning Nonlinear Dynamics in Physical Modelling Synthesis using Neural  Ordinary Differential Equations

标题: 使用神经常微方程学习物理模型合成中的非线性动力学
链接:https://arxiv.org/abs/2505.10511
作者: Victor Zheleznov,  Stefan Bilbao,  Alec Wright,  Simon King 
备注:Accepted for publication in Proceedings of the 28th International Conference on Digital Audio Effects (DAFx25), Ancona, Italy, September 2025
摘要:模态综合方法是一种长期存在的分布式音乐系统建模方法。在某些情况下,扩展是可能的,以处理几何非线性。一个这样的情况是弦的高振幅振动,其中几何非线性效应导致感知上重要的效应,包括音高滑动和亮度对打击振幅的依赖性。模态分解导致耦合的非线性常微分方程组。应用机器学习方法(特别是神经常微分方程)的最新工作已被用于从数据自动建模集总动态系统,如电子电路。在这项工作中,我们研究如何模态分解可以与神经常微分方程建模分布式音乐系统相结合。该模型利用系统的线性振动模式的解析解,并采用神经网络来解释非线性动态行为。系统的物理参数在训练之后保持容易访问,而不需要网络架构中的参数编码器。作为概念的初始证明,我们生成了非线性横向弦的合成数据,并表明该模型可以被训练来再现系统的非线性动力学。声音的例子。
摘要:Modal synthesis methods are a long-standing approach for modelling distributed musical systems. In some cases extensions are possible in order to handle geometric nonlinearities. One such case is the high-amplitude vibration of a string, where geometric nonlinear effects lead to perceptually important effects including pitch glides and a dependence of brightness on striking amplitude. A modal decomposition leads to a coupled nonlinear system of ordinary differential equations. Recent work in applied machine learning approaches (in particular neural ordinary differential equations) has been used to model lumped dynamic systems such as electronic circuits automatically from data. In this work, we examine how modal decomposition can be combined with neural ordinary differential equations for modelling distributed musical systems. The proposed model leverages the analytical solution for linear vibration of system's modes and employs a neural network to account for nonlinear dynamic behaviour. Physical parameters of a system remain easily accessible after the training without the need for a parameter encoder in the network architecture. As an initial proof of concept, we generate synthetic data for a nonlinear transverse string and show that the model can be trained to reproduce the nonlinear dynamics of the system. Sound examples are presented.


【3】 ListenNet: A Lightweight Spatio-Temporal Enhancement Nested Network for  Auditory Attention Detection

标题: ListenNet:用于听觉注意力检测的轻量级时空增强嵌套网络
链接:https://arxiv.org/abs/2505.10348
作者: Cunhang Fan,  Xiaoke Yang,  Hongyu Zhang,  Ying Chen,  Lu Li,  Jian Zhou,  Zhao Lv 
摘要:听觉注意检测(AAD)的目的是在多说话人环境中从脑信号(如脑电图(EEG)信号)中识别被关注说话人的方向。然而,现有的基于EEG的AAD方法忽略了EEG信号的时空依赖性,限制了它们的解码和泛化能力。为了解决这些问题,本文提出了一个轻量级的时空增强嵌套网络(ListenNet)的AAD。ListenNet有三个关键组件:时空相关编码器(STDE),多尺度时间增强(MSTE)和交叉嵌套注意力(CNA)。STDE重建跨通道的连续时间窗口之间的依赖关系,提高了动态模式提取的鲁棒性。MSTE在多个尺度上捕获时间特征,以表示细粒度和长范围的时间模式。此外,CNA通过新颖的动态注意机制更有效地集成了层次特征,以捕获深度时空相关性。在三个公共数据集上的实验结果表明,ListenNet在受试者相关和具有挑战性的受试者无关设置中优于最先进的方法,同时将可训练参数计数减少了约7倍。代码可从以下网址获得:https://github.com/fchest/ListenNet。
摘要:Auditory attention detection (AAD) aims to identify the direction of the attended speaker in multi-speaker environments from brain signals, such as Electroencephalography (EEG) signals. However, existing EEG-based AAD methods overlook the spatio-temporal dependencies of EEG signals, limiting their decoding and generalization abilities. To address these issues, this paper proposes a Lightweight Spatio-Temporal Enhancement Nested Network (ListenNet) for AAD. The ListenNet has three key components: Spatio-temporal Dependency Encoder (STDE), Multi-scale Temporal Enhancement (MSTE), and Cross-Nested Attention (CNA). The STDE reconstructs dependencies between consecutive time windows across channels, improving the robustness of dynamic pattern extraction. The MSTE captures temporal features at multiple scales to represent both fine-grained and long-range temporal patterns. In addition, the CNA integrates hierarchical features more effectively through novel dynamic attention mechanisms to capture deep spatio-temporal correlations. Experimental results on three public datasets demonstrate the superiority of ListenNet over state-of-the-art methods in both subject-dependent and challenging subject-independent settings, while reducing the trainable parameter count by approximately 7 times. Code is available at:https://github.com/fchest/ListenNet.


【4】 LAV: Audio-Driven Dynamic Visual Generation with Neural Compression and  StyleGAN2

标题: LAV:采用神经压缩和StyleGAN 2的音频驱动动态视觉生成
链接:https://arxiv.org/abs/2505.10101
作者: Jongmin Jung,  Dasaem Jeong 
备注:Paper accepted at ISEA 2025, The 30th International Symposium on Electronic/Emerging Art, Seoul, Republic of Korea, 23 - 29 May 2025
摘要:本文介绍了LAV(Latent Audio-Visual),这是一个集成了EnCodec的神经音频压缩与StyleGAN 2的生成功能的系统,可以生成由预先录制的音频驱动的视觉动态输出。与之前依赖显式特征映射的作品不同,LAV使用EnCodec嵌入作为潜在表示,通过随机初始化的线性映射直接转换为StyleGAN 2的风格潜在空间。这种方法在转换中保留了语义的丰富性,从而实现了细微差别和语义连贯的视听翻译。该框架展示了使用预先训练的音频压缩模型进行艺术和计算应用的潜力。
摘要:This paper introduces LAV (Latent Audio-Visual), a system that integrates EnCodec's neural audio compression with StyleGAN2's generative capabilities to produce visually dynamic outputs driven by pre-recorded audio. Unlike previous works that rely on explicit feature mappings, LAV uses EnCodec embeddings as latent representations, directly transformed into StyleGAN2's style latent space via randomly initialized linear mapping. This approach preserves semantic richness in the transformation, enabling nuanced and semantically coherent audio-visual translations. The framework demonstrates the potential of using pretrained audio compression models for artistic and computational applications.


【5】 Theoretical Model of Acoustic Power Transfer Through Solids

标题: 固体声功率传递的理论模型
链接:https://arxiv.org/abs/2505.09784
作者: Ippokratis Kochliaridis,  Michail E. Kiziroglou 
备注:8th International Workoshop on Microsystems, International Hellenic University
摘要:声功率传输是一项相对较新的技术。它是一种现代类型的无线接口,其中使用机械波通过介质传输数据信号和电源电压。这种系统最简单的应用是测量音频扬声器的频率响应。它由可变信号发生器、驱动声源的测量放大器和扬声器驱动器组成。接收器包含一个带有电平记录器的麦克风电路。声功率传输可以有许多应用,例如:耳蜗植入物,声纳系统和无线充电。但是,这是一项新技术,因此需要进一步研究。
摘要:Acoustic Power Transfer is a relatively new technology. It is a modern type of a wireless interface, where data signals and supply voltages are transmitted, with the use of mechanical waves, through a medium. The simplest application of such systems is the measurement of frequency response for audio speakers. It consists of a variable signal generator, a measuring amplifier which drives an acoustic source and the loudspeaker driver. The receiver contains a microphone circuit with a level recorder. Acoustic Power Transfer could have many applications, such as: Cochlear Implants, Sonar Systems and Wireless Charging. However, it is a new technology, thus it needs further investigation.


【6】 Introducing voice timbre attribute detection

标题: 引入语音音色属性检测
链接:https://arxiv.org/abs/2505.09661
作者: Jinghao He,  Zhengyan Sheng,  Liping Chen,  Kong Aik Lee,  Zhen-Hua Ling 
摘要:本文着重于解释语音信号所传达的音色,并介绍了一个被称为语音音色属性检测(vessel)的任务。在这项任务中,语音音色解释与一组感官属性描述其人类的感知。处理一对语音话语,并且在指定的音色描述符中比较它们的强度。此外,提出了一个框架,这是建立在从语音话语提取的说话人嵌入。该研究是在VCTK-RVA数据集上进行的。对ECAPA-TDNN和FACodec说话人编码器的实验研究表明:1)ECAPA-TDNN说话人编码器在可见场景中的能力更强,其中测试说话人被包括在训练集中; 2)FACodec说话人编码器在不可见场景中更优越,其中测试说话人不是训练的一部分,这表明增强的泛化能力。VCTK-RVA数据集和开源代码可在网站https://github.com/vTAD2025-Challenge/vTAD上获得。
摘要:This paper focuses on explaining the timbre conveyed by speech signals and introduces a task termed voice timbre attribute detection (vTAD). In this task, voice timbre is explained with a set of sensory attributes describing its human perception. A pair of speech utterances is processed, and their intensity is compared in a designated timbre descriptor. Moreover, a framework is proposed, which is built upon the speaker embeddings extracted from the speech utterances. The investigation is conducted on the VCTK-RVA dataset. Experimental examinations on the ECAPA-TDNN and FACodec speaker encoders demonstrated that: 1) the ECAPA-TDNN speaker encoder was more capable in the seen scenario, where the testing speakers were included in the training set; 2) the FACodec speaker encoder was superior in the unseen scenario, where the testing speakers were not part of the training, indicating enhanced generalization capability. The VCTK-RVA dataset and open-source code are available on the website https://github.com/vTAD2025-Challenge/vTAD.


【7】 Detecting Musical Deepfakes

标题: 检测音乐Deepfake
链接:https://arxiv.org/abs/2505.09633
作者: Nick Sunday 
备注:Submitted as part of coursework at UT Austin. Accompanying code available at: this https URL
摘要:文本到音乐(TTM)平台的激增使音乐创作民主化,使用户能够毫不费力地生成高质量的作品。然而,这一创新也给音乐家和更广泛的音乐产业带来了新的挑战。这项研究通过将音频分类为deepfake或人类来研究使用FakeMusicCaps数据集检测AI生成的歌曲。为了模拟真实世界的对抗条件,对数据集应用了节奏拉伸和音高变换。从修改后的音频生成Mel频谱图,然后用于训练和评估卷积神经网络。除了展示技术成果外,这项工作还探讨了TTM平台的伦理和社会影响,认为精心设计的检测系统对于保护艺术家和释放音乐中生成AI的积极潜力至关重要。
摘要:The proliferation of Text-to-Music (TTM) platforms has democratized music creation, enabling users to effortlessly generate high-quality compositions. However, this innovation also presents new challenges to musicians and the broader music industry. This study investigates the detection of AI-generated songs using the FakeMusicCaps dataset by classifying audio as either deepfake or human. To simulate real-world adversarial conditions, tempo stretching and pitch shifting were applied to the dataset. Mel spectrograms were generated from the modified audio, then used to train and evaluate a convolutional neural network. In addition to presenting technical results, this work explores the ethical and societal implications of TTM platforms, arguing that carefully designed detection systems are essential to both protecting artists and unlocking the positive potential of generative AI in music.


【8】 SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for  Attacking Anonymized Speech

标题: SpecWav-Attack:利用频谱图大小调整和Wav 2 Vec 2.0攻击匿名语音
链接:https://arxiv.org/abs/2505.09616
作者: Yuqi Li,  Yuanzhong Zheng,  Zhongtian Guo,  Yaoxuan Wang,  Jianjun Yin,  Haojun Fei 
备注:2 pages,3 figures,1 chart
摘要:本文提出了SpecWav-Attack,一种用于检测匿名语音中说话人的对抗模型。它利用Wav 2 Vec 2进行特征提取,并结合了频谱图分析和增量训练以提高性能。在librispeech-dev和librispeech-test上进行评估后,SpecWav-Attack的性能优于传统攻击,揭示了匿名语音系统中的漏洞,并强调了以ICASSP 2025攻击者挑战为基准进行更强大防御的必要性。
摘要:This paper presents SpecWav-Attack, an adversarial model for detecting speakers in anonymized speech. It leverages Wav2Vec2 for feature extraction and incorporates spectrogram resizing and incremental training for improved performance. Evaluated on librispeech-dev and librispeech-test, SpecWav-Attack outperforms conventional attacks, revealing vulnerabilities in anonymized speech systems and emphasizing the need for stronger defenses, benchmarked against the ICASSP 2025 Attacker Challenge.


eess.AS音频处理


【1】 Quantized Approximate Signal Processing (QASP): Towards Homomorphic  Encryption for audio

标题: 量化近似信号处理(QISP):迈向音频的同质加密
链接:https://arxiv.org/abs/2505.10500
作者: Tu Duyen Nguyen,  Adrien Lesage,  Clotilde Cantini,  Rachid Riad 
备注:34 pages, 5 figures
摘要:音频和语音数据越来越多地用于机器学习应用,如语音识别、说话人识别和心理健康监测。然而,音频收听设备被动收集这些数据引起了严重的隐私问题。全同态加密(FHE)通过对加密数据进行计算并保护用户隐私,提供了一种很有前途的解决方案。尽管FHE具有潜力,但将其应用于音频处理的先前尝试面临挑战,特别是在安全计算时间频率表示方面,这是许多音频任务中的关键步骤。   在这里,我们通过引入一个完全安全的管道来解决这个差距,该管道使用FHE和量化神经网络操作来计算四个基本的时频表示:短时傅立叶变换(STFT),Mel滤波器组,Mel频率倒谱系数(MFCC)和伽马滤波器。我们的方法还支持音频描述符和卷积神经网络(CNN)分类器的私有计算。此外,我们提出了近似STFT算法,减轻计算和比特使用的统计和机器学习分析。   我们在VocalSet和OxVoc数据集上进行了实验,展示了我们方法的完全私有计算。我们在音频标记的私人统计分析中使用STFT近似以及使用CNN进行声乐练习分类时表现出显着的性能改善。我们的研究结果表明,我们的近似大大降低了错误率相比,传统的STFT实现FHE。我们还展示了基于原始音频的完全私有分类,用于性别和发声练习分类。最后,我们提供了一个实用的启发式参数选择,使量化的近似信号处理的研究人员和从业人员,旨在保护敏感的音频数据。
摘要:Audio and speech data are increasingly used in machine learning applications such as speech recognition, speaker identification, and mental health monitoring. However, the passive collection of this data by audio listening devices raises significant privacy concerns. Fully homomorphic encryption (FHE) offers a promising solution by enabling computations on encrypted data and preserving user privacy. Despite its potential, prior attempts to apply FHE to audio processing have faced challenges, particularly in securely computing time frequency representations, a critical step in many audio tasks.   Here, we addressed this gap by introducing a fully secure pipeline that computes, with FHE and quantized neural network operations, four fundamental time-frequency representations: Short-Time Fourier Transform (STFT), Mel filterbanks, Mel-frequency cepstral coefficients (MFCCs), and gammatone filters. Our methods also support the private computation of audio descriptors and convolutional neural network (CNN) classifiers. Besides, we proposed approximate STFT algorithms that lighten computation and bit use for statistical and machine learning analyses.   We ran experiments on the VocalSet and OxVoc datasets demonstrating the fully private computation of our approach. We showed significant performance improvements with STFT approximation in private statistical analysis of audio markers, and for vocal exercise classification with CNNs. Our results reveal that our approximations substantially reduce error rates compared to conventional STFT implementations in FHE. We also demonstrated a fully private classification based on the raw audio for gender and vocal exercise classification. Finally, we provided a practical heuristic for parameter selection, making quantized approximate signal processing accessible to researchers and practitioners aiming to protect sensitive audio data.


【2】 Spatially Selective Active Noise Control for Open-fitting Hearables with  Acausal Optimization

标题: 具有并行优化的开放式助听器的空间选择性主动噪音控制
链接:https://arxiv.org/abs/2505.10372
作者: Tong Xiao,  Simon Doclo 
备注:Forum Acusticum/Euronoise 2025
摘要:有源噪声控制的最新进展使得具有空间选择性的可听设备的开发成为可能,其主动抑制不期望的噪声,同时保留来自特定方向的期望声音。在这项工作中,我们提出了一种改进的方法,空间选择性有源噪声控制,将非因果相对脉冲响应的优化过程中,导致显着改善的因果设计的性能。我们评估系统通过模拟使用一对开放式拟合听觉与空间本地化的语音和噪声源在消声环境中。性能评估方面的语音失真,降噪,并在不同的延迟和程度的非因果性的信噪比改善。结果表明,所提出的因果优化始终优于因果的方法在所有的指标和情况下,因果滤波器更有效地表征所需的源的响应。
摘要:Recent advances in active noise control have enabled the development of hearables with spatial selectivity, which actively suppress undesired noise while preserving desired sound from specific directions. In this work, we propose an improved approach to spatially selective active noise control that incorporates acausal relative impulse responses into the optimization process, resulting in significantly improved performance over the causal design. We evaluate the system through simulations using a pair of open-fitting hearables with spatially localized speech and noise sources in an anechoic environment. Performance is evaluated in terms of speech distortion, noise reduction, and signal-to-noise ratio improvement across different delays and degrees of acausality. Results show that the proposed acausal optimization consistently outperforms the causal approach across all metrics and scenarios, as acausal filters more effectively characterize the response of the desired source.


【3】 Who Said What WSW 2.0? Enhanced Automated Analysis of Preschool  Classroom Speech

标题: 谁说的WSW 2.0?增强的学前课堂语音自动分析
链接:https://arxiv.org/abs/2505.09972
作者: Anchen Sun,  Tiantian Feng,  Gabriela Gutierrez,  Juan J Londono,  Anfeng Xu,  Batya Elbaum,  Shrikanth Narayanan,  Lynn K Perry,  Daniel S Messinger 
备注:8 pages, 2 figures, 5 tables
摘要:本文介绍了一个自动化框架WSW 2.0,用于分析学前班教室中的语音交互,通过整合基于wav2vec2的说话人分类和Whisper(大v2和大v3)语音转录来提高准确性和可扩展性。总共235分钟的音频记录(160分钟来自12名儿童,75分钟来自5名教师)用于将系统输出与专家人工注释进行比较。WSW 2.0在说话人分类(儿童与教师)方面实现了.845的加权F1得分,.846的准确性和.672的纠错Kappa。转录质量是中等偏高,教师的单词错误率为0.119,儿童为0.238。WSW 2.0表现出相对较高的绝对协议组内相关性(ICC)与专家transmittance的一系列课堂语言功能。这些包括教师和儿童的平均话语长度,词汇多样性,提问,以及对问题和其他话语的反应,这些都显示出.64和.98之间的绝对一致性。为了建立可扩展性,我们将该框架应用于跨越两年的广泛数据集和超过1,592小时的课堂录音,展示了该框架在广泛的现实应用中的鲁棒性。这些发现突出了深度学习和自然语言处理技术通过提供学前班课堂语音关键特征的准确测量来彻底改变教育研究的潜力,最终指导更有效的干预策略并支持幼儿语言发展。
摘要:This paper introduces an automated framework WSW2.0 for analyzing vocal interactions in preschool classrooms, enhancing both accuracy and scalability through the integration of wav2vec2-based speaker classification and Whisper (large-v2 and large-v3) speech transcription. A total of 235 minutes of audio recordings (160 minutes from 12 children and 75 minutes from 5 teachers), were used to compare system outputs to expert human annotations. WSW2.0 achieves a weighted F1 score of .845, accuracy of .846, and an error-corrected kappa of .672 for speaker classification (child vs. teacher). Transcription quality is moderate to high with word error rates of .119 for teachers and .238 for children. WSW2.0 exhibits relatively high absolute agreement intraclass correlations (ICC) with expert transcriptions for a range of classroom language features. These include teacher and child mean utterance length, lexical diversity, question asking, and responses to questions and other utterances, which show absolute agreement intraclass correlations between .64 and .98. To establish scalability, we apply the framework to an extensive dataset spanning two years and over 1,592 hours of classroom audio recordings, demonstrating the framework's robustness for broad real-world applications. These findings highlight the potential of deep learning and natural language processing techniques to revolutionize educational research by providing accurate measures of key features of preschool classroom speech, ultimately guiding more effective intervention strategies and supporting early childhood language development.


【4】 T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation  via Fine-grained AI Feedback

标题: T2 A反馈:通过细粒度人工智能反馈提高文本到音频生成的基本能力
链接:https://arxiv.org/abs/2505.10561
作者: Zehan Wang,  Ke Lei,  Chen Zhu,  Jiawei Huang,  Sashuai Zhou,  Luping Liu,  Xize Cheng,  Shengpeng Ji,  Zhenhui Ye,  Tao Jin,  Zhou Zhao 
备注:ACL 2025
摘要:文本到音频(T2 A)生成在从语言提示生成各种音频输出方面取得了显著进展。然而,当前最先进的T2 A模型在生成复杂的多事件音频时仍然难以满足人类对跟随性和声学质量的偏好。为了提高模型在这些高级应用中的性能,我们建议通过AI反馈学习来增强模型的基本功能。首先,我们引入了细粒度的AI音频评分管道:1)验证文本提示中的每个事件是否存在于音频中(事件发生分数),2)检测事件序列与语言描述的偏差(事件序列分数),以及3)评估生成的音频的整体声学和谐波质量(声学和谐波质量)。我们评估了这三个自动评分管道,发现它们与人类偏好的相关性明显优于其他评估指标。这突出了它们作为反馈信号和评估指标的价值。利用我们强大的评分管道,我们构建了一个大型音频偏好数据集T2 A-FeedBack,其中包含41 k提示和249 k音频,每个都附有详细的分数。此外,我们还介绍了T2 A-EpicBench,这是一个专注于长字幕,多事件和讲故事场景的基准测试,旨在评估T2 A模型的高级功能。最后,我们展示了T2 A-FeedBack如何增强当前最先进的音频模型。通过简单的偏好调整,音频生成模型在简单(AudioCaps测试集)和复杂(T2 A-EpicBench)场景中都表现出显着的改进。
摘要:Text-to-audio (T2A) generation has achieved remarkable progress in generating a variety of audio outputs from language prompts. However, current state-of-the-art T2A models still struggle to satisfy human preferences for prompt-following and acoustic quality when generating complex multi-event audio. To improve the performance of the model in these high-level applications, we propose to enhance the basic capabilities of the model with AI feedback learning. First, we introduce fine-grained AI audio scoring pipelines to: 1) verify whether each event in the text prompt is present in the audio (Event Occurrence Score), 2) detect deviations in event sequences from the language description (Event Sequence Score), and 3) assess the overall acoustic and harmonic quality of the generated audio (Acoustic&Harmonic Quality). We evaluate these three automatic scoring pipelines and find that they correlate significantly better with human preferences than other evaluation metrics. This highlights their value as both feedback signals and evaluation metrics. Utilizing our robust scoring pipelines, we construct a large audio preference dataset, T2A-FeedBack, which contains 41k prompts and 249k audios, each accompanied by detailed scores. Moreover, we introduce T2A-EpicBench, a benchmark that focuses on long captions, multi-events, and story-telling scenarios, aiming to evaluate the advanced capabilities of T2A models. Finally, we demonstrate how T2A-FeedBack can enhance current state-of-the-art audio model. With simple preference tuning, the audio generation model exhibits significant improvements in both simple (AudioCaps test set) and complex (T2A-EpicBench) scenarios.


【5】 Learning Nonlinear Dynamics in Physical Modelling Synthesis using Neural  Ordinary Differential Equations

标题: 使用神经常微方程学习物理模型合成中的非线性动力学
链接:https://arxiv.org/abs/2505.10511
作者: Victor Zheleznov,  Stefan Bilbao,  Alec Wright,  Simon King 
备注:Accepted for publication in Proceedings of the 28th International Conference on Digital Audio Effects (DAFx25), Ancona, Italy, September 2025
摘要:模态综合方法是一种长期存在的分布式音乐系统建模方法。在某些情况下,扩展是可能的,以处理几何非线性。一个这样的情况是弦的高振幅振动,其中几何非线性效应导致感知上重要的效应,包括音高滑动和亮度对打击振幅的依赖性。模态分解导致耦合的非线性常微分方程组。应用机器学习方法(特别是神经常微分方程)的最新工作已被用于从数据自动建模集总动态系统,如电子电路。在这项工作中,我们研究如何模态分解可以与神经常微分方程建模分布式音乐系统相结合。该模型利用系统的线性振动模式的解析解,并采用神经网络来解释非线性动态行为。系统的物理参数在训练之后保持容易访问,而不需要网络架构中的参数编码器。作为概念的初始证明,我们生成了非线性横向弦的合成数据,并表明该模型可以被训练来再现系统的非线性动力学。声音的例子。
摘要:Modal synthesis methods are a long-standing approach for modelling distributed musical systems. In some cases extensions are possible in order to handle geometric nonlinearities. One such case is the high-amplitude vibration of a string, where geometric nonlinear effects lead to perceptually important effects including pitch glides and a dependence of brightness on striking amplitude. A modal decomposition leads to a coupled nonlinear system of ordinary differential equations. Recent work in applied machine learning approaches (in particular neural ordinary differential equations) has been used to model lumped dynamic systems such as electronic circuits automatically from data. In this work, we examine how modal decomposition can be combined with neural ordinary differential equations for modelling distributed musical systems. The proposed model leverages the analytical solution for linear vibration of system's modes and employs a neural network to account for nonlinear dynamic behaviour. Physical parameters of a system remain easily accessible after the training without the need for a parameter encoder in the network architecture. As an initial proof of concept, we generate synthetic data for a nonlinear transverse string and show that the model can be trained to reproduce the nonlinear dynamics of the system. Sound examples are presented.


【6】 ListenNet: A Lightweight Spatio-Temporal Enhancement Nested Network for  Auditory Attention Detection

标题: ListenNet:用于听觉注意力检测的轻量级时空增强嵌套网络
链接:https://arxiv.org/abs/2505.10348
作者: Cunhang Fan,  Xiaoke Yang,  Hongyu Zhang,  Ying Chen,  Lu Li,  Jian Zhou,  Zhao Lv 
摘要:听觉注意检测(AAD)的目的是在多说话人环境中从脑信号(如脑电图(EEG)信号)中识别被关注说话人的方向。然而,现有的基于EEG的AAD方法忽略了EEG信号的时空依赖性,限制了它们的解码和泛化能力。为了解决这些问题,本文提出了一个轻量级的时空增强嵌套网络(ListenNet)的AAD。ListenNet有三个关键组件:时空相关编码器(STDE),多尺度时间增强(MSTE)和交叉嵌套注意力(CNA)。STDE重建跨通道的连续时间窗口之间的依赖关系,提高了动态模式提取的鲁棒性。MSTE在多个尺度上捕获时间特征,以表示细粒度和长范围的时间模式。此外,CNA通过新颖的动态注意机制更有效地集成了层次特征,以捕获深度时空相关性。在三个公共数据集上的实验结果表明,ListenNet在受试者相关和具有挑战性的受试者无关设置中优于最先进的方法,同时将可训练参数计数减少了约7倍。代码可从以下网址获得:https://github.com/fchest/ListenNet。
摘要:Auditory attention detection (AAD) aims to identify the direction of the attended speaker in multi-speaker environments from brain signals, such as Electroencephalography (EEG) signals. However, existing EEG-based AAD methods overlook the spatio-temporal dependencies of EEG signals, limiting their decoding and generalization abilities. To address these issues, this paper proposes a Lightweight Spatio-Temporal Enhancement Nested Network (ListenNet) for AAD. The ListenNet has three key components: Spatio-temporal Dependency Encoder (STDE), Multi-scale Temporal Enhancement (MSTE), and Cross-Nested Attention (CNA). The STDE reconstructs dependencies between consecutive time windows across channels, improving the robustness of dynamic pattern extraction. The MSTE captures temporal features at multiple scales to represent both fine-grained and long-range temporal patterns. In addition, the CNA integrates hierarchical features more effectively through novel dynamic attention mechanisms to capture deep spatio-temporal correlations. Experimental results on three public datasets demonstrate the superiority of ListenNet over state-of-the-art methods in both subject-dependent and challenging subject-independent settings, while reducing the trainable parameter count by approximately 7 times. Code is available at:https://github.com/fchest/ListenNet.


【7】 LAV: Audio-Driven Dynamic Visual Generation with Neural Compression and  StyleGAN2

标题: LAV:采用神经压缩和StyleGAN 2的音频驱动动态视觉生成
链接:https://arxiv.org/abs/2505.10101
作者: Jongmin Jung,  Dasaem Jeong 
备注:Paper accepted at ISEA 2025, The 30th International Symposium on Electronic/Emerging Art, Seoul, Republic of Korea, 23 - 29 May 2025
摘要:本文介绍了LAV(Latent Audio-Visual),这是一个集成了EnCodec的神经音频压缩与StyleGAN 2的生成功能的系统,可以生成由预先录制的音频驱动的视觉动态输出。与之前依赖显式特征映射的作品不同,LAV使用EnCodec嵌入作为潜在表示,通过随机初始化的线性映射直接转换为StyleGAN 2的风格潜在空间。这种方法在转换中保留了语义的丰富性,从而实现了细微差别和语义连贯的视听翻译。该框架展示了使用预先训练的音频压缩模型进行艺术和计算应用的潜力。
摘要:This paper introduces LAV (Latent Audio-Visual), a system that integrates EnCodec's neural audio compression with StyleGAN2's generative capabilities to produce visually dynamic outputs driven by pre-recorded audio. Unlike previous works that rely on explicit feature mappings, LAV uses EnCodec embeddings as latent representations, directly transformed into StyleGAN2's style latent space via randomly initialized linear mapping. This approach preserves semantic richness in the transformation, enabling nuanced and semantically coherent audio-visual translations. The framework demonstrates the potential of using pretrained audio compression models for artistic and computational applications.


【8】 Theoretical Model of Acoustic Power Transfer Through Solids

标题: 固体声功率传递的理论模型
链接:https://arxiv.org/abs/2505.09784
作者: Ippokratis Kochliaridis,  Michail E. Kiziroglou 
备注:8th International Workoshop on Microsystems, International Hellenic University
摘要:声功率传输是一项相对较新的技术。它是一种现代类型的无线接口,其中使用机械波通过介质传输数据信号和电源电压。这种系统最简单的应用是测量音频扬声器的频率响应。它由可变信号发生器、驱动声源的测量放大器和扬声器驱动器组成。接收器包含一个带有电平记录器的麦克风电路。声功率传输可以有许多应用,例如:耳蜗植入物,声纳系统和无线充电。但是,这是一项新技术,因此需要进一步研究。
摘要:Acoustic Power Transfer is a relatively new technology. It is a modern type of a wireless interface, where data signals and supply voltages are transmitted, with the use of mechanical waves, through a medium. The simplest application of such systems is the measurement of frequency response for audio speakers. It consists of a variable signal generator, a measuring amplifier which drives an acoustic source and the loudspeaker driver. The receiver contains a microphone circuit with a level recorder. Acoustic Power Transfer could have many applications, such as: Cochlear Implants, Sonar Systems and Wireless Charging. However, it is a new technology, thus it needs further investigation.


【9】 Introducing voice timbre attribute detection

标题: 引入语音音色属性检测
链接:https://arxiv.org/abs/2505.09661
作者: Jinghao He,  Zhengyan Sheng,  Liping Chen,  Kong Aik Lee,  Zhen-Hua Ling 
摘要:本文着重于解释语音信号所传达的音色,并介绍了一个被称为语音音色属性检测(vessel)的任务。在这项任务中,语音音色解释与一组感官属性描述其人类的感知。处理一对语音话语,并且在指定的音色描述符中比较它们的强度。此外,提出了一个框架,这是建立在从语音话语提取的说话人嵌入。该研究是在VCTK-RVA数据集上进行的。对ECAPA-TDNN和FACodec说话人编码器的实验研究表明:1)ECAPA-TDNN说话人编码器在可见场景中的能力更强,其中测试说话人被包括在训练集中; 2)FACodec说话人编码器在不可见场景中更优越,其中测试说话人不是训练的一部分,这表明增强的泛化能力。VCTK-RVA数据集和开源代码可在网站https://github.com/vTAD2025-Challenge/vTAD上获得。
摘要:This paper focuses on explaining the timbre conveyed by speech signals and introduces a task termed voice timbre attribute detection (vTAD). In this task, voice timbre is explained with a set of sensory attributes describing its human perception. A pair of speech utterances is processed, and their intensity is compared in a designated timbre descriptor. Moreover, a framework is proposed, which is built upon the speaker embeddings extracted from the speech utterances. The investigation is conducted on the VCTK-RVA dataset. Experimental examinations on the ECAPA-TDNN and FACodec speaker encoders demonstrated that: 1) the ECAPA-TDNN speaker encoder was more capable in the seen scenario, where the testing speakers were included in the training set; 2) the FACodec speaker encoder was superior in the unseen scenario, where the testing speakers were not part of the training, indicating enhanced generalization capability. The VCTK-RVA dataset and open-source code are available on the website https://github.com/vTAD2025-Challenge/vTAD.


【10】 SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for  Attacking Anonymized Speech

标题: SpecWav-Attack:利用频谱图大小调整和Wav 2 Vec 2.0攻击匿名语音
链接:https://arxiv.org/abs/2505.09616
作者: Yuqi Li,  Yuanzhong Zheng,  Zhongtian Guo,  Yaoxuan Wang,  Jianjun Yin,  Haojun Fei 
备注:2 pages,3 figures,1 chart
摘要:本文提出了SpecWav-Attack,一种用于检测匿名语音中说话人的对抗模型。它利用Wav 2 Vec 2进行特征提取,并结合了频谱图分析和增量训练以提高性能。在librispeech-dev和librispeech-test上进行评估后,SpecWav-Attack的性能优于传统攻击,揭示了匿名语音系统中的漏洞,并强调了以ICASSP 2025攻击者挑战为基准进行更强大防御的必要性。
摘要:This paper presents SpecWav-Attack, an adversarial model for detecting speakers in anonymized speech. It leverages Wav2Vec2 for feature extraction and incorporates spectrogram resizing and incremental training for improved performance. Evaluated on librispeech-dev and librispeech-test, SpecWav-Attack outperforms conventional attacks, revealing vulnerabilities in anonymized speech systems and emphasizing the need for stronger defenses, benchmarked against the ICASSP 2025 Attacker Challenge.


机器翻译由腾讯交互翻译提供,仅供参考