今日论文合集:cs.SD语音9篇,eess.AS音频处理7篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 Dopamine Audiobook: A Training-free MLLM Agent for Emotional and  Human-like Audiobook Generation
标题: 多巴胺有声读物:情感和类人有声读物生成的免训练MLLM代理
链接:https://arxiv.org/abs/2504.11002
作者: Yan Rong,  Shan Yang,  Guangzhi Lei,  Li Liu 
摘要:有声书生成,创造生动和情感丰富的音频作品,面临着传达复杂的情感,实现人性化的品质,并与人类的喜好调整评价的挑战。现有的文本到语音(TTS)方法通常限于特定的场景,与情感过渡作斗争,并且缺乏自动的人类对齐的评估基准,而是依赖于不对齐的自动化指标或昂贵的人类评估。为了解决这些问题,我们提出了Dopamine Audiobook,这是一个新的统一的无训练系统,利用多模态大型语言模型(MLLM)作为AI代理,用于情感和类人有声读物的生成和评估。具体来说,我们首先设计了一个基于流的情感增强框架,将复杂的情感语音合成分解为可控的子任务。然后,我们提出了一个自适应模型选择模块,动态地选择最合适的TTS方法,从一组现有的国家的最先进的(SOTA)TTS方法为不同的场景。我们进一步提高情感表达能力,通过在词和话语水平上的双语言增强和韵律检索。对于评估,我们提出了一种新的基于GPT的评估框架,包括自我批评,观点采择和心理MagicEmo提示,以确保人类对齐和自我对齐的评估。实验表明,我们的方法生成的长语音具有优越的情感表达SOTA TTS模型在各种指标。重要的是,我们的评估框架展示了更好的与人类偏好的一致性和跨音频任务的可转移性。包含音频样本的项目网站可以在https://dopamine-audiobook.github.io上找到。
摘要:Audiobook generation, which creates vivid and emotion-rich audio works, faces challenges in conveying complex emotions, achieving human-like qualities, and aligning evaluations with human preferences. Existing text-to-speech (TTS) methods are often limited to specific scenarios, struggle with emotional transitions, and lack automatic human-aligned evaluation benchmarks, instead relying on either misaligned automated metrics or costly human assessments. To address these issues, we propose Dopamine Audiobook, a new unified training-free system leveraging a multimodal large language model (MLLM) as an AI agent for emotional and human-like audiobook generation and evaluation. Specifically, we first design a flow-based emotion-enhanced framework that decomposes complex emotional speech synthesis into controllable sub-tasks. Then, we propose an adaptive model selection module that dynamically selects the most suitable TTS methods from a set of existing state-of-the-art (SOTA) TTS methods for diverse scenarios. We further enhance emotional expressiveness through paralinguistic augmentation and prosody retrieval at word and utterance levels. For evaluation, we propose a novel GPT-based evaluation framework incorporating self-critique, perspective-taking, and psychological MagicEmo prompts to ensure human-aligned and self-aligned assessments. Experiments show that our method generates long speech with superior emotional expression to SOTA TTS models in various metrics. Importantly, our evaluation framework demonstrates better alignment with human preferences and transferability across audio tasks. Project website with audio samples can be found at https://dopamine-audiobook.github.io.


【2】 Real-Time Word-Level Temporal Segmentation in Streaming Speech  Recognition

标题: 流语音识别中的实时词级时态分割
链接:https://arxiv.org/abs/2504.10849
作者: Naoto Nishida,  Hirotaka Hiraki,  Jun Rekimoto,  Yoshio Ishiguro 
备注:3 pages, 1 figures
摘要:富文本字幕对于帮助聋人和听力障碍(DHH)人群、第二语言学习者以及自闭症谱系障碍(ASD)患者进行交流至关重要。它们还在将语音转换为文本时保留细微差别,增强演示文稿脚本和对话或语音日志的真实感。然而,目前的实时字幕系统缺乏改变文本属性(例如,大写、大小和字体),从而阻碍了以语音的音调或语调表达的说话者意图的准确传达。例如,“你应该这样做”往往被认为是指示“你”作为句子的焦点,而“你应该这样做”往往是“这”作为焦点。本文提出了一种在单词级实时改变文本装饰的解决方案。作为原型,我们开发了一个应用程序,可以根据每个口语单词的响度调整单词大小。来自用户的反馈表明,该系统有助于传达说话者的意图,提供更吸引人和更容易获得的字幕体验。
摘要:Rich-text captions are essential to help communication for Deaf and hard-of-hearing (DHH) people, second-language learners, and those with autism spectrum disorder (ASD). They also preserve nuances when converting speech to text, enhancing the realism of presentation scripts and conversation or speech logs. However, current real-time captioning systems lack the capability to alter text attributes (ex. capitalization, sizes, and fonts) at the word level, hindering the accurate conveyance of speaker intent that is expressed in the tones or intonations of the speech. For example, ''YOU should do this'' tends to be considered as indicating ''You'' as the focus of the sentence, whereas ''You should do THIS'' tends to be ''This'' as the focus. This paper proposes a solution that changes the text decorations at the word level in real time. As a prototype, we developed an application that adjusts word size based on the loudness of each spoken word. Feedback from users implies that this system helped to convey the speaker's intent, offering a more engaging and accessible captioning experience.


【3】 SteerMusic: Enhanced Musical Consistency for Zero-shot Text-Guided and  Personalized Music Editing

标题: SteerMusic:增强的音乐一致性,实现Zero-Shot文本引导和个性化音乐编辑
链接:https://arxiv.org/abs/2504.10826
作者: Xinlei Niu,  Kin Wai Cheuk,  Jing Zhang,  Naoki Murata,  Chieh-Hsin Lai,  Michele Mancusi,  Woosung Choi,  Giorgio Fabbro,  Wei-Hsiang Liao,  Charles Patrick Martin,  Yuki Mitsufuji 
摘要:音乐编辑是音乐制作中的重要环节,在游戏开发和电影制作中有着广泛的应用。大多数现有的zero-shot文本引导方法都依赖于预训练的扩散模型,涉及向前-向后扩散过程进行编辑。然而,这些方法通常难以保持音乐内容的一致性。此外,单独的文本指令通常无法准确地描述所需的音乐。在本文中,我们提出了两种音乐编辑方法,提高了原始和编辑的音乐之间的一致性,利用分数蒸馏。第一种方法,SteerMusic,是一种粗粒度的zero-shot编辑方法,使用增量去噪得分。第二种方法是SteerMusic+,通过操纵表示用户定义的音乐风格的概念令牌来实现细粒度的个性化音乐编辑。SteerMusic+允许将音乐编辑成任何用户定义的音乐风格,这些音乐风格无法单独通过文本指令实现。实验结果表明,我们的方法优于现有的方法在保持音乐内容的一致性和编辑保真度。用户研究进一步验证了我们的方法实现了卓越的音乐编辑质量。音频示例可在https://steermusic.pages.dev/上获得。
摘要:Music editing is an important step in music production, which has broad applications, including game development and film production. Most existing zero-shot text-guided methods rely on pretrained diffusion models by involving forward-backward diffusion processes for editing. However, these methods often struggle to maintain the music content consistency. Additionally, text instructions alone usually fail to accurately describe the desired music. In this paper, we propose two music editing methods that enhance the consistency between the original and edited music by leveraging score distillation. The first method, SteerMusic, is a coarse-grained zero-shot editing approach using delta denoising score. The second method, SteerMusic+, enables fine-grained personalized music editing by manipulating a concept token that represents a user-defined musical style. SteerMusic+ allows for the editing of music into any user-defined musical styles that cannot be achieved by the text instructions alone. Experimental results show that our methods outperform existing approaches in preserving both music content consistency and editing fidelity. User studies further validate that our methods achieve superior music editing quality. Audio examples are available on https://steermusic.pages.dev/.


【4】 Progressive Rock Music Classification

标题: 进步摇滚音乐分类
链接:https://arxiv.org/abs/2504.10821
作者: Arpan Nagar,  Joseph Bensabat,  Jokent Gaza,  Moinak Dey 
备注:20 pages
摘要:本研究旨在探讨前卫摇滚乐的分类,这是一种有别于其他音乐风格的音乐类型,其特点是复杂的作曲和多样的乐器。为了解决这个音乐信息检索(MIR)任务,我们使用Librosa库从歌曲片段中提取了全面的音频特征,包括频谱图,梅尔频率倒谱系数(MFCC),色度图和节拍位置。采用赢家通吃的投票策略将片段级预测聚合到最终的歌曲分类中。我们对各种机器学习技术进行了比较分析。探索了包括Bagging(Random Forest,ExtraTrees,Bagging Classifier)和Boosting(XGBoost,Gradient Boosting)在内的Entrim方法,利用主成分分析(PCA)进行降维,以管理高维特征集的计算约束。此外,还研究了深度学习方法,包括开发定制的1D卷积神经网络(1D CNN)架构(名为“Zuck”和“Satya”),其具有特定的层配置、归一化和激活函数。此外,我们微调了最先进的音频频谱图Transformer(AST)模型,利用其基于注意力的音频分类机制。对验证集和测试集的性能评估显示,不同模型的有效性不同,像Extra Trees这样的集成方法的测试准确率高达76.38%。这项研究为渐进式摇滚流派分类的细致入微的任务提供了不同机器学习范式的应用和相对性能的见解。
摘要:This study investigates the classification of progressive rock music, a genre characterized by complex compositions and diverse instrumentation, distinct from other musical styles. Addressing this Music Information Retrieval (MIR) task, we extracted comprehensive audio features, including spectrograms, Mel-Frequency Cepstral Coefficients (MFCCs), chromagrams, and beat positions from song snippets using the Librosa library. A winner-take-all voting strategy was employed to aggregate snippet-level predictions into final song classifications. We conducted a comparative analysis of various machine learning techniques. Ensemble methods, encompassing Bagging (Random Forest, ExtraTrees, Bagging Classifier) and Boosting (XGBoost, Gradient Boosting), were explored, utilizing Principal Component Analysis (PCA) for dimensionality reduction to manage computational constraints with high-dimensional feature sets. Additionally, deep learning approaches were investigated, including the development of custom 1D Convolutional Neural Network (1D CNN) architectures (named "Zuck" and "Satya") featuring specific layer configurations, normalization, and activation functions. Furthermore, we fine-tuned a state-of-the-art Audio Spectrogram Transformer (AST) model, leveraging its attention-based mechanisms for audio classification. Performance evaluation on validation and test sets revealed varying effectiveness across models, with ensemble methods like Extra Trees achieving test accuracies up to 76.38%. This research provides insights into the application and relative performance of diverse machine learning paradigms for the nuanced task of progressive rock genre classification.


【5】 Generalized Audio Deepfake Detection Using Frame-level Latent  Information Entropy

标题: 使用帧级潜在信息量的广义音频深度伪造检测
链接:https://arxiv.org/abs/2504.10819
作者: Botao Zhao,  Zuheng Kang,  Yayun He,  Xiaoyang Qu,  Junqing Peng,  Jing Xiao,  Jianzong Wang 
备注:Accpeted by IEEE International Conference on Multimedia & Expo 2025 (ICME 2025)
摘要:由于文本到语音(TTS)和语音转换(VC)技术的快速发展,泛化能力,即鲁棒模型在看不见的数据上有效执行的能力,对于音频deepfake检测至关重要。区分真实样本和欺骗样本的一种有前途的方法在于识别内在差异以增强模型的泛化能力。从信息论的角度来看,我们假设的信息内容是一个内在的差异:真正的样本代表了一个密集的,信息丰富的采样的真实世界,而欺骗样本通常是来自低维,信息量较少的表示。为了实现这一点,我们引入了帧级潜在信息熵检测器(f-InfoED),这是一个框架,它从帧级的潜在表示中提取独特的信息熵,以识别音频deepfake。此外,我们还介绍了AdaLAM,它使用可训练适配器扩展了大型预训练音频模型,以增强特征提取。为了便于全面评估,音频deepfake取证2024(ADFF 2024)数据集是通过最新的TTS和VC方法构建的。大量的实验表明,我们提出的方法达到了最先进的性能,并表现出显着的泛化能力。进一步的分析研究证实了AdaLAM在提取区别性音频特征方面的有效性,以及f-InfoED在利用潜在熵信息进行更广泛的deepfake检测方面的有效性。
摘要:Generalizability, the capacity of a robust model to perform effectively on unseen data, is crucial for audio deepfake detection due to the rapid evolution of text-to-speech (TTS) and voice conversion (VC) technologies. A promising approach to differentiate between bonafide and spoof samples lies in identifying intrinsic disparities to enhance model generalizability. From an information-theoretic perspective, we hypothesize the information content is one of the intrinsic differences: bonafide sample represents a dense, information-rich sampling of the real world, whereas spoof sample is typically derived from lower-dimensional, less informative representations. To implement this, we introduce frame-level latent information entropy detector(f-InfoED), a framework that extracts distinctive information entropy from latent representations at the frame level to identify audio deepfakes. Furthermore, we present AdaLAM, which extends large pre-trained audio models with trainable adapters for enhanced feature extraction. To facilitate comprehensive evaluation, the audio deepfake forensics 2024 (ADFF 2024) dataset was built by the latest TTS and VC methods. Extensive experiments demonstrate that our proposed approach achieves state-of-the-art performance and exhibits remarkable generalization capabilities. Further analytical studies confirms the efficacy of AdaLAM in extracting discriminative audio features and f-InfoED in leveraging latent entropy information for more generalized deepfake detection.


【6】 SonicSieve: Bringing Directional Speech Extraction to Smartphones Using  Acoustic Microstructures

标题: SonicSieve:使用声学微结构将定向语音提取引入智能手机
链接:https://arxiv.org/abs/2504.10793
作者: Kuang Yuan,  Yifeng Wang,  Xiyuxing Zhang,  Chengyi Shen,  Swarun Kumar,  Justin Chan 
摘要:想象一下,将智能手机放在一家嘈杂的餐厅的桌子上,清晰地捕捉到坐在你周围的朋友的声音,或者在一个回荡的礼堂里清晰地记录下演讲者的声音。我们介绍SonicSieve,第一个智能定向语音提取系统的智能手机使用生物启发的声学微结构。我们的无源设计将方向提示嵌入到传入的语音中,而无需任何额外的电子设备。它连接到低成本有线耳机的内嵌麦克风,可以连接到智能手机。我们提出了一个端到端的神经网络,在移动设备上实时处理原始音频混合。我们的研究结果表明,SonicSieve实现了5.0 dB的信号质量改善时,集中在30{\deg}角区域。此外,我们的系统的性能的基础上只有两个麦克风超过传统的5麦克风阵列。
摘要:Imagine placing your smartphone on a table in a noisy restaurant and clearly capturing the voices of friends seated around you, or recording a lecturer's voice with clarity in a reverberant auditorium. We introduce SonicSieve, the first intelligent directional speech extraction system for smartphones using a bio-inspired acoustic microstructure. Our passive design embeds directional cues onto incoming speech without any additional electronics. It attaches to the in-line mic of low-cost wired earphones which can be attached to smartphones. We present an end-to-end neural network that processes the raw audio mixtures in real-time on mobile devices. Our results show that SonicSieve achieves a signal quality improvement of 5.0 dB when focusing on a 30{\deg} angular region. Additionally, the performance of our system based on only two microphones exceeds that of conventional 5-microphone arrays.


【7】 Deep Audio Watermarks are Shallow: Limitations of Post-Hoc Watermarking  Techniques for Speech

标题: 深度音频水印很浅:语音事后水印技术的局限性
链接:https://arxiv.org/abs/2504.10782
作者: Patrick O'Reilly,  Zeyu Jin,  Jiaqi Su,  Bryan Pardo 
备注:ICLR 2025 Workshop on GenAI Watermarking
摘要:在音频模态中,最先进的水印方法利用深度神经网络来允许在生成的音频中嵌入人类无法感知的签名。理想的情况是嵌入签名,当水印音频通过压缩,过滤或其他变换改变时,可以高精度地检测到这些签名。现有的音频水印技术以事后方式运行,在生成后操纵音频记录的“低级”特征(例如,通过添加低幅度水印信号)。我们表明,这种事后制定现有的音频水印容易受到基于变换的删除攻击。专注于语音音频,我们(1)统一和扩展现有的音频变换对水印检测能力的影响的评估,和(2)表明,国家的最先进的事后音频水印可以被删除,没有知识的水印方案和最小的音频质量下降。
摘要:In the audio modality, state-of-the-art watermarking methods leverage deep neural networks to allow the embedding of human-imperceptible signatures in generated audio. The ideal is to embed signatures that can be detected with high accuracy when the watermarked audio is altered via compression, filtering, or other transformations. Existing audio watermarking techniques operate in a post-hoc manner, manipulating "low-level" features of audio recordings after generation (e.g. through the addition of a low-magnitude watermark signal). We show that this post-hoc formulation makes existing audio watermarks vulnerable to transformation-based removal attacks. Focusing on speech audio, we (1) unify and extend existing evaluations of the effect of audio transformations on watermark detectability, and (2) demonstrate that state-of-the-art post-hoc audio watermarks can be removed with no knowledge of the watermarking scheme and minimal degradation in audio quality.


【8】 Hearing Anywhere in Any Environment

标题: 在任何环境中的任何地方都能听到
链接:https://arxiv.org/abs/2504.10746
作者: Xiulong Liu,  Anurag Kumar,  Paul Calamia,  Sebastia V. Amengual,  Calvin Murdock,  Ishwarya Ananthabhotla,  Philip Robinson,  Eli Shlizerman,  Vamsi Krishna Ithapu,  Ruohan Gao 
备注:CVPR 2025
摘要:在混合现实应用中,空间环境中逼真的声学体验与实现真正沉浸感的视觉体验一样重要。尽管用于房间脉冲响应(RIR)估计的神经方法最近取得了进展,但大多数现有方法仅限于训练它们的单一环境,缺乏推广到具有不同几何形状和表面材料的新房间的能力。我们的目标是开发一个统一的模型,能够重建任何环境的空间声学经验,最少的额外测量。为此,我们提出了xRIR,跨房间RIR预测的框架。我们的可推广的方法的核心在于结合的几何特征提取器,它捕获从全景深度图像的空间上下文,与RIR编码器,提取详细的声学特征,从只有几个参考RIR样本。为了评估我们的方法,我们引入了ACOUSTICROOMS,这是一个新的数据集,具有来自260个房间的300,000多个RIR的高保真模拟。实验表明,我们的方法大大优于一系列的基线。此外,我们通过在四个真实世界环境中评估我们的模型,成功地执行了模拟到真实的转换,证明了我们方法的通用性和数据集的真实性。
摘要:In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances in neural approaches for Room Impulse Response (RIR) estimation, most existing methods are limited to the single environment on which they are trained, lacking the ability to generalize to new rooms with different geometries and surface materials. We aim to develop a unified model capable of reconstructing the spatial acoustic experience of any environment with minimum additional measurements. To this end, we present xRIR, a framework for cross-room RIR prediction. The core of our generalizable approach lies in combining a geometric feature extractor, which captures spatial context from panorama depth images, with a RIR encoder that extracts detailed acoustic features from only a few reference RIR samples. To evaluate our method, we introduce ACOUSTICROOMS, a new dataset featuring high-fidelity simulation of over 300,000 RIRs from 260 rooms. Experiments show that our method strongly outperforms a series of baselines. Furthermore, we successfully perform sim-to-real transfer by evaluating our model on four real-world environments, demonstrating the generalizability of our approach and the realism of our dataset.


【9】 Focal Loss based Residual Convolutional Neural Network for Speech  Emotion Recognition

标题: 基于焦失的剩余卷积神经网络用于语音情感识别
链接:https://arxiv.org/abs/1906.05682
作者: Suraj Tripathi,  Abhay Kumar,  Abhiram Ramesh,  Chirag Singh,  Promod Yenigalla 
备注:Accepted in CICLing 2019
摘要:本文提出了一种基于语音特征的残差卷积神经网络(ResNet),并在聚焦损失下进行训练,以识别语音中的情感。语音特征,如频谱图和梅尔频率倒谱系数(MFCC)已经显示出比纯文本更好地表征情感的能力。进一步的焦点损失,首先用于一阶段对象检测器,已经显示出将训练过程更多地集中在硬示例上的能力,并降低分配给分类良好的示例的损失的权重,从而防止模型被容易分类的示例淹没。
摘要:This paper proposes a Residual Convolutional Neural Network (ResNet) based on speech features and trained under Focal Loss to recognize emotion in speech. Speech features such as Spectrogram and Mel-frequency Cepstral Coefficients (MFCCs) have shown the ability to characterize emotion better than just plain text. Further Focal Loss, first used in One-Stage Object Detectors, has shown the ability to focus the training process more towards hard-examples and down-weight the loss assigned to well-classified examples, thus preventing the model from being overwhelmed by easily classifiable examples.


eess.AS音频处理


【1】 Respiratory Inhaler Sound Event Classification Using Self-Supervised  Learning

标题: 使用自我监督学习的呼吸吸入器声音事件分类
链接:https://arxiv.org/abs/2504.11246
作者: Davoud Shariat Panah,  Alessandro N Franciosi,  Cormac McCarthy,  Andrew Hines 
备注:Accepted at the IEEE EMBC 2025 Conference
摘要:哮喘是一种慢性呼吸系统疾病,影响着全世界数百万人。虽然这种情况可以通过手持吸入器给予控制药物来管理,但临床研究表明,正确的吸入器使用技术的依从性较低。因此,许多患者可能无法获得药物的全部益处。最近研究了吸入器声音的自动分类以评估药物依从性。然而,现有的分类模型通常是使用来自特定吸入器类型的数据进行训练的,并且它们对来自不同吸入器的声音进行概括的能力仍未得到探索。在这项研究中,我们采用了wav2vec 2.0自监督学习模型,通过对吸入器声音进行预训练和微调,对吸入器声音进行分类。该模型在使用干粉吸入器和智能手表设备收集的数据集上显示出98%的平衡准确度。结果还表明,对来自目标吸入器的最小数据的该模型进行微调是使通用吸入器声音分类模型适应不同吸入器设备和音频捕获硬件的有前途的方法。这是该领域的第一项研究,证明了智能手表作为辅助技术的潜力,用于使用机器学习模型个性化监测吸入器依从性。
摘要:Asthma is a chronic respiratory condition that affects millions of people worldwide. While this condition can be managed by administering controller medications through handheld inhalers, clinical studies have shown low adherence to the correct inhaler usage technique. Consequently, many patients may not receive the full benefit of their medication. Automated classification of inhaler sounds has recently been studied to assess medication adherence. However, the existing classification models were typically trained using data from specific inhaler types, and their ability to generalize to sounds from different inhalers remains unexplored. In this study, we adapted the wav2vec 2.0 self-supervised learning model for inhaler sound classification by pre-training and fine-tuning this model on inhaler sounds. The proposed model shows a balanced accuracy of 98% on a dataset collected using a dry powder inhaler and smartwatch device. The results also demonstrate that re-finetuning this model on minimal data from a target inhaler is a promising approach to adapting a generic inhaler sound classification model to a different inhaler device and audio capture hardware. This is the first study in the field to demonstrate the potential of smartwatches as assistive technologies for the personalized monitoring of inhaler adherence using machine learning models.


【2】 Dopamine Audiobook: A Training-free MLLM Agent for Emotional and  Human-like Audiobook Generation

标题: 多巴胺有声读物:情感和类人有声读物生成的免训练MLLM代理
链接:https://arxiv.org/abs/2504.11002
作者: Yan Rong,  Shan Yang,  Guangzhi Lei,  Li Liu 
摘要:有声书生成,创造生动和情感丰富的音频作品,面临着传达复杂的情感,实现人性化的品质,并与人类的喜好调整评价的挑战。现有的文本到语音(TTS)方法通常限于特定的场景,与情感过渡作斗争,并且缺乏自动的人类对齐的评估基准,而是依赖于不对齐的自动化指标或昂贵的人类评估。为了解决这些问题,我们提出了Dopamine Audiobook,这是一个新的统一的无训练系统,利用多模态大型语言模型(MLLM)作为AI代理,用于情感和类人有声读物的生成和评估。具体来说,我们首先设计了一个基于流的情感增强框架,将复杂的情感语音合成分解为可控的子任务。然后,我们提出了一个自适应模型选择模块,动态地选择最合适的TTS方法,从一组现有的国家的最先进的(SOTA)TTS方法为不同的场景。我们进一步提高情感表达能力,通过在词和话语水平上的语言增强和韵律检索。对于评估,我们提出了一种新的基于GPT的评估框架,包括自我批评,观点采择和心理MagicEmo提示,以确保人类对齐和自我对齐的评估。实验表明,我们的方法生成的长语音具有优越的情感表达SOTA TTS模型在各种指标。重要的是,我们的评估框架展示了更好的与人类偏好的一致性和跨音频任务的可转移性。包含音频样本的项目网站可以在https://dopamine-audiobook.github.io上找到。
摘要:Audiobook generation, which creates vivid and emotion-rich audio works, faces challenges in conveying complex emotions, achieving human-like qualities, and aligning evaluations with human preferences. Existing text-to-speech (TTS) methods are often limited to specific scenarios, struggle with emotional transitions, and lack automatic human-aligned evaluation benchmarks, instead relying on either misaligned automated metrics or costly human assessments. To address these issues, we propose Dopamine Audiobook, a new unified training-free system leveraging a multimodal large language model (MLLM) as an AI agent for emotional and human-like audiobook generation and evaluation. Specifically, we first design a flow-based emotion-enhanced framework that decomposes complex emotional speech synthesis into controllable sub-tasks. Then, we propose an adaptive model selection module that dynamically selects the most suitable TTS methods from a set of existing state-of-the-art (SOTA) TTS methods for diverse scenarios. We further enhance emotional expressiveness through paralinguistic augmentation and prosody retrieval at word and utterance levels. For evaluation, we propose a novel GPT-based evaluation framework incorporating self-critique, perspective-taking, and psychological MagicEmo prompts to ensure human-aligned and self-aligned assessments. Experiments show that our method generates long speech with superior emotional expression to SOTA TTS models in various metrics. Importantly, our evaluation framework demonstrates better alignment with human preferences and transferability across audio tasks. Project website with audio samples can be found at https://dopamine-audiobook.github.io.


【3】 Progressive Rock Music Classification

标题: 进步摇滚音乐分类
链接:https://arxiv.org/abs/2504.10821
作者: Arpan Nagar,  Joseph Bensabat,  Jokent Gaza,  Moinak Dey 
备注:20 pages
摘要:None
摘要:This study investigates the classification of progressive rock music, a genre characterized by complex compositions and diverse instrumentation, distinct from other musical styles. Addressing this Music Information Retrieval (MIR) task, we extracted comprehensive audio features, including spectrograms, Mel-Frequency Cepstral Coefficients (MFCCs), chromagrams, and beat positions from song snippets using the Librosa library. A winner-take-all voting strategy was employed to aggregate snippet-level predictions into final song classifications. We conducted a comparative analysis of various machine learning techniques. Ensemble methods, encompassing Bagging (Random Forest, ExtraTrees, Bagging Classifier) and Boosting (XGBoost, Gradient Boosting), were explored, utilizing Principal Component Analysis (PCA) for dimensionality reduction to manage computational constraints with high-dimensional feature sets. Additionally, deep learning approaches were investigated, including the development of custom 1D Convolutional Neural Network (1D CNN) architectures (named "Zuck" and "Satya") featuring specific layer configurations, normalization, and activation functions. Furthermore, we fine-tuned a state-of-the-art Audio Spectrogram Transformer (AST) model, leveraging its attention-based mechanisms for audio classification. Performance evaluation on validation and test sets revealed varying effectiveness across models, with ensemble methods like Extra Trees achieving test accuracies up to 76.38%. This research provides insights into the application and relative performance of diverse machine learning paradigms for the nuanced task of progressive rock genre classification.


【4】 Generalized Audio Deepfake Detection Using Frame-level Latent  Information Entropy

标题: 使用帧级潜在信息量的广义音频深度伪造检测
链接:https://arxiv.org/abs/2504.10819
作者: Botao Zhao,  Zuheng Kang,  Yayun He,  Xiaoyang Qu,  Junqing Peng,  Jing Xiao,  Jianzong Wang 
备注:Accpeted by IEEE International Conference on Multimedia & Expo 2025 (ICME 2025)
摘要:由于文本到语音(TTS)和语音转换(VC)技术的快速发展,泛化能力,即鲁棒模型在看不见的数据上有效执行的能力,对于音频deepfake检测至关重要。区分真实样本和欺骗样本的一种有前途的方法在于识别内在差异以增强模型的泛化能力。从信息论的角度来看,我们假设的信息内容是一个内在的差异:真正的样本代表了一个密集的,信息丰富的采样的真实世界,而欺骗样本通常是来自低维,信息量较少的表示。为了实现这一点,我们引入了帧级潜在信息熵检测器(f-InfoED),这是一个框架,它从帧级的潜在表示中提取独特的信息熵,以识别音频deepfake。此外,我们还介绍了AdaLAM,它使用可训练适配器扩展了大型预训练音频模型,以增强特征提取。为了便于全面评估,音频deepfake取证2024(ADFF 2024)数据集是通过最新的TTS和VC方法构建的。大量的实验表明,我们提出的方法达到了最先进的性能,并表现出显着的泛化能力。进一步的分析研究证实了AdaLAM在提取区别性音频特征方面的有效性,以及f-InfoED在利用潜在熵信息进行更广泛的deepfake检测方面的有效性。
摘要:Generalizability, the capacity of a robust model to perform effectively on unseen data, is crucial for audio deepfake detection due to the rapid evolution of text-to-speech (TTS) and voice conversion (VC) technologies. A promising approach to differentiate between bonafide and spoof samples lies in identifying intrinsic disparities to enhance model generalizability. From an information-theoretic perspective, we hypothesize the information content is one of the intrinsic differences: bonafide sample represents a dense, information-rich sampling of the real world, whereas spoof sample is typically derived from lower-dimensional, less informative representations. To implement this, we introduce frame-level latent information entropy detector(f-InfoED), a framework that extracts distinctive information entropy from latent representations at the frame level to identify audio deepfakes. Furthermore, we present AdaLAM, which extends large pre-trained audio models with trainable adapters for enhanced feature extraction. To facilitate comprehensive evaluation, the audio deepfake forensics 2024 (ADFF 2024) dataset was built by the latest TTS and VC methods. Extensive experiments demonstrate that our proposed approach achieves state-of-the-art performance and exhibits remarkable generalization capabilities. Further analytical studies confirms the efficacy of AdaLAM in extracting discriminative audio features and f-InfoED in leveraging latent entropy information for more generalized deepfake detection.


【5】 SonicSieve: Bringing Directional Speech Extraction to Smartphones Using  Acoustic Microstructures

标题: SonicSieve:使用声学微结构将定向语音提取引入智能手机
链接:https://arxiv.org/abs/2504.10793
作者: Kuang Yuan,  Yifeng Wang,  Xiyuxing Zhang,  Chengyi Shen,  Swarun Kumar,  Justin Chan 
摘要:想象一下,将智能手机放在一家嘈杂的餐厅的桌子上,清晰地捕捉到坐在你周围的朋友的声音,或者在一个回荡的礼堂里清晰地记录下演讲者的声音。我们介绍SonicSieve,第一个智能定向语音提取系统的智能手机使用生物启发的声学微结构。我们的无源设计将方向提示嵌入到传入的语音中,而无需任何额外的电子设备。它连接到低成本有线耳机的内嵌麦克风,可以连接到智能手机。我们提出了一个端到端的神经网络,在移动设备上实时处理原始音频混合。我们的研究结果表明,SonicSieve实现了5.0 dB的信号质量改善时,集中在30{\deg}角区域。此外,我们的系统的性能的基础上只有两个麦克风超过传统的5麦克风阵列。
摘要:Imagine placing your smartphone on a table in a noisy restaurant and clearly capturing the voices of friends seated around you, or recording a lecturer's voice with clarity in a reverberant auditorium. We introduce SonicSieve, the first intelligent directional speech extraction system for smartphones using a bio-inspired acoustic microstructure. Our passive design embeds directional cues onto incoming speech without any additional electronics. It attaches to the in-line mic of low-cost wired earphones which can be attached to smartphones. We present an end-to-end neural network that processes the raw audio mixtures in real-time on mobile devices. Our results show that SonicSieve achieves a signal quality improvement of 5.0 dB when focusing on a 30{\deg} angular region. Additionally, the performance of our system based on only two microphones exceeds that of conventional 5-microphone arrays.


【6】 Deep Audio Watermarks are Shallow: Limitations of Post-Hoc Watermarking  Techniques for Speech

标题: 深度音频水印很浅:语音事后水印技术的局限性
链接:https://arxiv.org/abs/2504.10782
作者: Patrick O'Reilly,  Zeyu Jin,  Jiaqi Su,  Bryan Pardo 
备注:ICLR 2025 Workshop on GenAI Watermarking
摘要:在音频模态中,最先进的水印方法利用深度神经网络来允许在生成的音频中嵌入人类无法感知的签名。理想的情况是嵌入签名,当水印音频通过压缩,过滤或其他变换改变时,可以高精度地检测到这些签名。现有的音频水印技术以事后方式运行,在生成后操纵音频记录的“低级”特征(例如,通过添加低幅度水印信号)。我们表明,这种事后制定现有的音频水印容易受到基于变换的删除攻击。专注于语音音频,我们(1)统一和扩展现有的音频变换对水印检测能力的影响的评估,和(2)表明,国家的最先进的事后音频水印可以被删除,没有知识的水印方案和最小的音频质量下降。
摘要:In the audio modality, state-of-the-art watermarking methods leverage deep neural networks to allow the embedding of human-imperceptible signatures in generated audio. The ideal is to embed signatures that can be detected with high accuracy when the watermarked audio is altered via compression, filtering, or other transformations. Existing audio watermarking techniques operate in a post-hoc manner, manipulating "low-level" features of audio recordings after generation (e.g. through the addition of a low-magnitude watermark signal). We show that this post-hoc formulation makes existing audio watermarks vulnerable to transformation-based removal attacks. Focusing on speech audio, we (1) unify and extend existing evaluations of the effect of audio transformations on watermark detectability, and (2) demonstrate that state-of-the-art post-hoc audio watermarks can be removed with no knowledge of the watermarking scheme and minimal degradation in audio quality.


【7】 Will AI shape the way we speak? The emerging sociolinguistic influence  of synthetic voices

标题: 人工智能会改变我们说话的方式吗?合成声音新兴的社会语言学影响
链接:https://arxiv.org/abs/2504.10650
作者: Éva Székely,  Jūra Miniota,  Míša (Michaela) Hejná 
备注:5 pages, 0 figures, International Workshop on Spoken Dialogue Systems Technology (IWSDS) 2025
摘要:在语音和语言技术发展的推动下,会话语音界面的日益普及,提出了关于它们对人类交流影响的重要问题。虽然书面交流可以通过词汇和风格的选择来表明身份,但基于语音的互动本质上会放大社会索引元素-如口音,语调和演讲风格-更突出地传达社会身份和群体归属。有证据表明,即使是像电视这样的被动媒体也可能影响观众的语言模式。与被动媒体不同,对话式人工智能是互动的,创造了一种更具沉浸感和互惠性的动态,具有更大的潜力来影响个人在日常互动中的说话方式。这种增强的影响可以预期产生的现象,如声学韵律夹带和语言的调节,这自然发生在互动过程中,使用户能够适应他们的语音模式,以响应系统。虽然这一现象仍在出现,但其潜在的社会影响可能会为组织、运动和品牌提供一个微妙而强大的途径,来塑造和控制公众的看法和社会认同。我们认为,人工智能生成的语音的社会索引影响值得关注,并应成为跨学科研究的重点,利用新的和现有的方法和技术,以更好地了解其影响。
摘要:The growing prevalence of conversational voice interfaces, powered by developments in both speech and language technologies, raises important questions about their influence on human communication. While written communication can signal identity through lexical and stylistic choices, voice-based interactions inherently amplify socioindexical elements - such as accent, intonation, and speech style - which more prominently convey social identity and group affiliation. There is evidence that even passive media such as television is likely to influence the audience's linguistic patterns. Unlike passive media, conversational AI is interactive, creating a more immersive and reciprocal dynamic that holds a greater potential to impact how individuals speak in everyday interactions. Such heightened influence can be expected to arise from phenomena such as acoustic-prosodic entrainment and linguistic accommodation, which occur naturally during interaction and enable users to adapt their speech patterns in response to the system. While this phenomenon is still emerging, its potential societal impact could provide organisations, movements, and brands with a subtle yet powerful avenue for shaping and controlling public perception and social identity. We argue that the socioindexical influence of AI-generated speech warrants attention and should become a focus of interdisciplinary research, leveraging new and existing methodologies and technologies to better understand its implications.


机器翻译由腾讯交互翻译提供,仅供参考