今日论文合集:cs.SD语音11篇,eess.AS音频处理5篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Enhancing Quranic Learning: A Multimodal Deep Learning Approach for Arabic Phoneme Recognition
标题:增强古兰经学习:阿拉伯语音素识别的多模式深度学习方法
链接:https://arxiv.org/abs/2511.17477

作者:Ayhan Kucukmanisa,Derya Gelmez,Sukru Selim Calik,Zeynep Hilal Kilimci
备注:11 pages, 2 figures, 3 tables
摘要:多模态深度学习的最新进展大大增强了语音分析和发音评估系统的能力。准确的发音检测仍然是阿拉伯语的一个关键挑战,特别是在古兰经背诵的背景下,细微的语音差异可能会改变含义。针对这一挑战,本研究提出了一种基于变换器的多模态框架,结合声学和文本表示,以实现更高的精度和鲁棒性的阿拉伯语音素发音错误检测。该框架将UniSpeech衍生的声学嵌入与从Whisper transmitting中提取的基于BERT的文本嵌入相集成,创建了一个统一的表示,可以捕获语音细节和语言上下文。为了确定最有效的整合策略,早期,中期和后期融合方法的实施和评估两个数据集包含29个阿拉伯语音素,包括8个hafiz声音,由11个母语。从公开的YouTube录音中收集的其他语音样本被纳入,以提高数据的多样性和通用性。使用标准评估指标评估模型性能:准确度,精确度,召回率和F1分数,允许对融合策略进行详细比较。实验结果表明,UniSpeech-BERT多模态配置提供了强大的结果,基于融合的Transformer架构是有效的音素级发音错误检测。该研究有助于开发智能的,独立于说话者的,多模式的计算机辅助语言学习(CALL)系统,为技术支持的古兰经发音训练和更广泛的基于语音的教育应用提供了实际的一步。
摘要:Recent advances in multimodal deep learning have greatly enhanced the capability of systems for speech analysis and pronunciation assessment. Accurate pronunciation detection remains a key challenge in Arabic, particularly in the context of Quranic recitation, where subtle phonetic differences can alter meaning. Addressing this challenge, the present study proposes a transformer-based multimodal framework for Arabic phoneme mispronunciation detection that combines acoustic and textual representations to achieve higher precision and robustness. The framework integrates UniSpeech-derived acoustic embeddings with BERT-based textual embeddings extracted from Whisper transcriptions, creating a unified representation that captures both phonetic detail and linguistic context. To determine the most effective integration strategy, early, intermediate, and late fusion methods were implemented and evaluated on two datasets containing 29 Arabic phonemes, including eight hafiz sounds, articulated by 11 native speakers. Additional speech samples collected from publicly available YouTube recordings were incorporated to enhance data diversity and generalization. Model performance was assessed using standard evaluation metrics: accuracy, precision, recall, and F1-score, allowing a detailed comparison of the fusion strategies. Experimental findings show that the UniSpeech-BERT multimodal configuration provides strong results and that fusion-based transformer architectures are effective for phoneme-level mispronunciation detection. The study contributes to the development of intelligent, speaker-independent, and multimodal Computer-Aided Language Learning (CALL) systems, offering a practical step toward technology-supported Quranic pronunciation training and broader speech-based educational applications.


【2】Semantic and Semiotic Interplays in Text-to-Audio AI: Exploring Cognitive Dynamics and Musical Interactions
标题:文本到音频人工智能中的语义和符号相互作用:探索认知动力学和音乐相互作用
链接:https://arxiv.org/abs/2511.17429

作者:Guilherme Coelho
摘要:本文研究了人工智能(AI)中新兴的文本到音频范式,研究了其对音乐创作,解释和认知的变革性影响。我探讨了复杂的语义和符号的相互作用时发生的描述性自然语言提示翻译成细微差别的声音对象在文本到音频的方式。从结构主义和后结构主义的角度,以及图式动力学和元认知的认知理论,本文探讨了这些人工智能系统如何重新配置音乐的意义过程和导航既定的认知框架。该研究分析了人工智能介导的音乐中发挥作用的一些认知动力学,包括图式同化和适应,元认知反射和建设性感知的过程。本文认为,文本到音频的人工智能模型作为音乐意义的准对象,同时稳定和不稳定的传统形式,同时培养新的聆听模式和审美反射。本研究以Udio为主要案例研究,探讨这些模型如何导航语言提示和声音输出之间的阈空间。这一过程不仅产生了新颖的音乐表现形式,而且促使听众进行批判性的和“结构意识的倾听”。“,鼓励更深入地了解音乐的结构,符号学的细微差别,以及塑造我们音乐认知的社会文化背景。本文最后反思了文本到音频AI模型作为认知工具和准对象的潜力,促进了音乐互动的重大转变,并邀请用户对音乐的认知和文化基础有更细致入微的理解。
摘要:This paper investigates the emerging text-to-audio paradigm in artificial intelligence (AI), examining its transformative implications for musical creation, interpretation, and cognition. I explore the complex semantic and semiotic interplays that occur when descriptive natural language prompts are translated into nuanced sound objects across the text-to-audio modality. Drawing from structuralist and post-structuralist perspectives, as well as cognitive theories of schema dynamics and metacognition, the paper explores how these AI systems reconfigure musical signification processes and navigate established cognitive frameworks. The research analyzes some of the cognitive dynamics at play in AI-mediated musicking, including processes of schema assimilation and accommodation, metacognitive reflection, and constructive perception. The paper argues that text-to-audio AI models function as quasi-objects of musical signification, simultaneously stabilizing and destabilizing conventional forms while fostering new modes of listening and aesthetic reflexivity.Using Udio as a primary case study, this study explores how these models navigate the liminal spaces between linguistic prompts and sonic outputs. This process not only generates novel musical expressions but also prompts listeners to engage in forms of critical and "structurally-aware listening.", encouraging a deeper understanding of music's structures, semiotic nuances, and the socio-cultural contexts that shape our musical cognition. The paper concludes by reflecting on the potential of text-to-audio AI models to serve as epistemic tools and quasi-objects, facilitating a significant shift in musical interactions and inviting users to develop a more nuanced comprehension of the cognitive and cultural foundations of music.


【3】AI in Music and Sound: Pedagogical Reflections, Post-Structuralist Approaches and Creative Outcomes in Seminar Practice
标题:音乐和声音中的人工智能:研讨会实践中的教学反思、后结构主义方法和创造性成果
链接:https://arxiv.org/abs/2511.17425

作者:Guilherme Coelho
摘要:本文介绍了一个教学和概念帐户的过程中,AI在音乐和声音:模式,工具和创造性的应用程序,提供内的音乐信息学和媒体艺术模块的硕士在音频通信。该课程让学生参与一系列AI模式,如符号合成,语音合成,音色转移,神经音频合成和文本到音频系统,将理论反思与基于实践的实验相结合。它的核心教学举措是一个成对的练习设计:每一种形式首先通过其预期的启示,然后通过一个故意重新定义或“误用”的练习,表面代表性的限制和替代行为。在媒介理论和后结构主义探究的框架下,我们将人工智能视为一个跨模态的机器人-一个跨文本,符号,音色和音频领域翻译和干扰音乐符号的系统。来自学生工作和反思的证据表明,技术流利性,媒介意识和批判性素养的增长,以及实验方法和过程导向的听力的培养。本文概述了课程架构,评估设计和代表性项目,并提炼出一套人工智能音乐教学法的设计模式(例如,文本到音频中的无条件相互作用和语义不稳定;音色转移中的潜在空间唯物主义)。它的结论是教学建议,将创造性实践与媒介意识和人工智能技术的文化认知分析相结合,使学生能够参与人工智能如何被理解,开发和部署在创意社区。
摘要:This paper presents a pedagogical and conceptual account of the course AI in Music and Sound: Modalities, Tools and Creative Applications, offered within the Music Informatics and Media Art module of an M.Sc. in Audio Communication. The course engaged students with a range of AI modalities such as symbolic composition, voice synthesis, timbre transfer, neural audio synthesis, and text-to-audio systems, combining theoretical reflection with practice-based experimentation. Its central pedagogical move is a paired-études design: each modality is approached first through its intended affordances and then through a deliberately reframed or "misused" exercise that surfaces representational limits and alternative behaviours. Framed by medium theory and post-structuralist inquiry, we treated AI as a transmodal conduit-a system that translates and perturbs musical signs across textual, symbolic, timbral and audio domains. Evidence from student work and reflection indicates growth in technical fluency, medium awareness, and critical literacy, alongside the cultivation of experimental method and process-oriented listening. The paper outlines the course architecture, assessment design, and representative projects, and distils a set of design patterns for AI-music pedagogy (eg., prompt-conditioned interplays and semantic destabilisation in text-to-audio; latent space materialism in timbre transfer). It concludes with pedagogical recommendations that integrate creative practice with medium awareness and with cultural-epistemic analysis of AI technologies, preparing students to participate in how AI is understood, developed, and deployed with creative communities.


【4】The Artist is Present: Traces of Artists Resigind and Spawning in Text-to-Audio AI
标题:艺术家在场:艺术家在文本转音频人工智能中辞职和繁衍的痕迹
链接:https://arxiv.org/abs/2511.17404

作者:Guilherme Coelho
摘要:文本到音频(TTA)系统正在迅速改变音乐创作和发行,像Udio和Suno这样的平台每天生成数千首曲目,并融入主流音乐平台和生态系统。这些系统在大量未公开的数据集上训练,从根本上重塑了音乐的制作、复制和消费方式。本文提出的经验证据表明,艺术家的条件区域可以系统地通过基于元标签的提示设计微定位,有效地使产卵的艺术家一样的内容,通过战略提示工程。通过系统地探索基于元标签的提示工程技术,这项研究揭示了用户如何访问特定艺术家的独特声音签名,证明它们包含在训练数据集中。使用从公共音乐分类中提取的描述符星座,本文展示了与Bon Iver,Philip Glass,Panda Bear和William Basinski等艺术家的可复制接近。结果表明,稳定的文本音频对应与艺术家特定的训练信号一致,使精确遍历风格的微位置,而无需明确命名的艺术家。这种召唤艺术家特定输出的能力表明,艺术家的创造性作品是这些系统产生新内容的基础材料,通常没有明确的同意或归属。从概念上讲,这项工作澄清了文本描述符如何在高维表示空间中充当导航线索;从方法上讲,它提供了一个可复制的协议,用于审计风格归纳。这些发现为治理--归属、同意和披露标准--和创造性实践提出了直接的质疑,在创造性实践中,诱导的风格接近使所有权、复制、模仿、创造性机构和算法创造的道德之间的界限复杂化。
摘要:Text-to-audio (TTA) systems are rapidly transforming music creation and distribution, with platforms like Udio and Suno generating thousands of tracks daily and integrating into mainstream music platforms and ecosystems. These systems, trained on vast and largely undisclosed datasets, are fundamentally reshaping how music is produced, reproduced and consumed. This paper presents empirical evidence that artist-conditioned regions can be systematically microlocated through metatag-based prompt design, effectively enabling the spawning of artist-like content through strategic prompt engineering. Through systematic exploration of metatag-based prompt engineering techniques this research reveals how users can access the distinctive sonic signatures of specific artists, evidencing their inclusion in training datasets. Using descriptor constellations drawn from public music taxonomies, the paper demonstrates reproducible proximity to artists such as Bon Iver, Philip Glass, Panda Bear and William Basinski. The results indicate stable text-audio correspondences consistent with artist-specific training signals, enabling precise traversal of stylistic microlocations without explicitly naming artists. This capacity to summon artist-specific outputs shows that artists' creative works fuction as foundational material from which these systems generate new content, often without explicit consent or attribuition. Conceptually, the work clarifies how textual descriptors act as navigational cues in high-dimensional representation spaces; methodologically, it provides a replicable protocol for auditing stylistic inducibility. The findings raise immediate queestions for governance-attribution, consent and disclosure standards-and for creative practice, where induced stylistic proximity complicates boundaries between ownership, reproduction, imitation, creative agency and the ethics of algorithmic creation.


【5】Is Phase Really Needed for Weakly-Supervised Dereverberation ?
标题:监督薄弱的消除回响真的需要阶段吗?
链接:https://arxiv.org/abs/2511.17346

作者:Marius Rodrigues,Louis Bahrman,Roland Badeau,Gaël Richard
摘要:在用于语音去混响的无监督或弱监督方法中,目标干净(干)信号在训练期间被认为是未知的。在这种情况下,评估在何种程度上可以从混响(湿)语音的唯一知识检索信息变得至关重要。这项工作调查的混响(湿)阶段在时间-频率域中的作用。基于统计波场理论,我们表明,后期混响扰动相位分量与白色,均匀分布的噪声,除了在低频率。因此,湿相携带有限的有用信息,并且对于弱监督去混响不是必需的。为了验证这一发现,我们训练去混响模型在最近的弱监管框架下,并证明,性能可以显着提高,从损失函数中排除混响阶段。
摘要:In unsupervised or weakly-supervised approaches for speech dereverberation, the target clean (dry) signals are considered to be unknown during training. In that context, evaluating to what extent information can be retrieved from the sole knowledge of reverberant (wet) speech becomes critical. This work investigates the role of the reverberant (wet) phase in the time-frequency domain. Based on Statistical Wave Field Theory, we show that late reverberation perturbs phase components with white, uniformly distributed noise, except at low frequencies. Consequently, the wet phase carries limited useful information and is not essential for weakly supervised dereverberation. To validate this finding, we train dereverberation models under a recent weak supervision framework and demonstrate that performance can be significantly improved by excluding the reverberant phase from the loss function.


【6】A new kid on the block: Distributional semantics predicts the word-specific tone signatures of monosyllabic words in conversational Taiwan Mandarin
标题:一个新人:分布语义预测台湾普通话对话中单音节词的词特定语气特征
链接:https://arxiv.org/abs/2511.17337

作者:Xiaoyun Jin,Mirjam Ernestus,R. Harald Baayen
备注:arXiv admin note: text overlap with arXiv:2409.07891
摘要:本文以语料库为基础,研究了汉语单音节词的音高轮廓是如何在自然会话中实现的,重点考察了词义的影响。我们使用广义加性模型将给定的观察到的音高轮廓分解成一组与不同的控制变量和语义预测因子相关联的音高轮廓。即使当变量,如词的持续时间,性别,扬声器的身份,音调上下文,元音的高度,和话语的位置进行控制,字的效果仍然是一个强有力的预测音调的实现。我们提出的证据表明,这种效果的字是一个语义的影响:词义被证明是一个更好的预测比字,和异形同音异义词被证明有不同的音高轮廓。语义重要性的最有力证据是,单个单词标记的音高轮廓可以从它们的上下文嵌入中预测,其准确性大大超过了排列基线。对于语音学来说,分布语义学是一个新生事物。虽然我们的研究结果挑战了普通话声调的标准理论,但它们符合区分性词汇模型的理论框架。
摘要:We present a corpus-based investigation of how the pitch contours of monosyllabic words are realized in spontaneous conversational Mandarin, focusing on the effects of words' meanings. We used the generalized additive model to decompose a given observed pitch contour into a set of component pitch contours that are tied to different control variables and semantic predictors. Even when variables such as word duration, gender, speaker identity, tonal context, vowel height, and utterance position are controlled for, the effect of word remains a strong predictor of tonal realization. We present evidence that this effect of word is a semantic effect: word sense is shown to be a better predictor than word, and heterographic homophones are shown to have different pitch contours. The strongest evidence for the importance of semantics is that the pitch contours of individual word tokens can be predicted from their contextualized embeddings with an accuracy that substantially exceeds a permutation baseline. For phonetics, distributional semantics is a new kid on the block. Although our findings challenge standard theories of Mandarin tone, they fit well within the theoretical framework of the Discriminative Lexicon Model.


【7】Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM
标题:基于长上下文Q-Former和多模态LLM的机器人确认生成和动作规划
链接:https://arxiv.org/abs/2511.17335

作者:Chiori Hori,Yoshiki Masuyama,Siddarth Jain,Radu Corcodel,Devesh Jha,Diego Romeres,Jonathan Le Roux
备注:Accepted to ASRU 2025
摘要:为了实现共同目标,人机协作需要机器人理解人类的行为以及与周围环境的交互。本文主要研究基于人机对话的人机交互,人机对话依赖于机器人的动作确认和动作步骤的生成,使用多模态场景理解。国家的最先进的方法使用多模态Transformers生成机器人动作步骤与机器人动作确认对齐,从一个单一的剪辑显示由多个微步骤组成的任务。虽然在整个视频中,对长视野任务的动作相互依赖,但目前的方法主要集中在剪辑级处理上,而没有利用长上下文信息。本文提出了一种长上下文Q-former,在完整的视频中结合左右上下文依赖。此外,本文提出了一种文本条件化方法,将文本嵌入直接馈送到LLM解码器中,以减轻Q-former对文本中信息的高度抽象。YouCook 2语料库的实验表明,确认生成的准确性是动作规划性能的主要因素。此外,我们证明了长上下文Q-前提高了确认和行动规划,通过整合VideoLLaMA 3。
摘要:Human-robot collaboration towards a shared goal requires robots to understand human action and interaction with the surrounding environment. This paper focuses on human-robot interaction (HRI) based on human-robot dialogue that relies on the robot action confirmation and action step generation using multimodal scene understanding. The state-of-the-art approach uses multimodal transformers to generate robot action steps aligned with robot action confirmation from a single clip showing a task composed of multiple micro steps. Although actions towards a long-horizon task depend on each other throughout an entire video, the current approaches mainly focus on clip-level processing and do not leverage long-context information. This paper proposes a long-context Q-former incorporating left and right context dependency in full videos. Furthermore, this paper proposes a text-conditioning approach to feed text embeddings directly into the LLM decoder to mitigate the high abstraction of the information in text by Q-former. Experiments with the YouCook2 corpus show that the accuracy of confirmation generation is a major factor in the performance of action planning. Furthermore, we demonstrate that the long-context Q-former improves the confirmation and action planning by integrating VideoLLaMA3.


【8】MusicAIR: A Multimodal AI Music Generation Framework Powered by an Algorithm-Driven Core
标题:MusicAir:由边界驱动核心提供支持的多模式人工智能音乐生成框架
链接:https://arxiv.org/abs/2511.17323

作者:Callie C. Liao,Duoduo Liao,Ellie L. Zhang
备注:Accepted by IEEE Big Data 2025
摘要:生成式人工智能的最新进展使音乐生成成为一个突出的研究焦点。然而,许多基于神经的模型依赖于大型数据集,这引发了对版权侵权和高性能成本的担忧。相比之下,我们提出了MusicAIR,这是一个创新的多模态AI音乐生成框架,由一个新颖的算法驱动的符号音乐核心提供支持,有效地降低了版权侵权风险。音乐核心算法将关键的歌词和节奏信息连接起来,自动导出音乐特征,仅从歌词中创建完整,连贯的旋律乐谱。MusicAIR框架促进了从歌词、文本和图像生成音乐。生成的乐谱遵循音乐理论、抒情结构和节奏惯例的既定原则。我们开发了Generate AI Music(GenAIM),这是一个使用MusicAIR进行歌词到歌曲,文本到音乐和图像到音乐生成的Web工具。在我们的实验中,我们使用标准音乐指标和创新分析来评估系统生成的AI生成的乐谱,并将这些作品与原创作品进行比较。该系统实现了85%的平均关键置信度,超过了人类作曲家的79%,并与既定的音乐理论标准密切相关,证明了其生成多样化,类人作品的能力。作为一个辅助工具,GenAIM可以作为一个可靠的音乐作曲助手和可能的教育作曲导师,同时降低所有有抱负的音乐家的进入门槛,这是创新的,并为音乐生成的AI做出了重大贡献。
摘要:Recent advances in generative AI have made music generation a prominent research focus. However, many neural-based models rely on large datasets, raising concerns about copyright infringement and high-performance costs. In contrast, we propose MusicAIR, an innovative multimodal AI music generation framework powered by a novel algorithm-driven symbolic music core, effectively mitigating copyright infringement risks. The music core algorithms connect critical lyrical and rhythmic information to automatically derive musical features, creating a complete, coherent melodic score solely from the lyrics. The MusicAIR framework facilitates music generation from lyrics, text, and images. The generated score adheres to established principles of music theory, lyrical structure, and rhythmic conventions. We developed Generate AI Music (GenAIM), a web tool using MusicAIR for lyric-to-song, text-to-music, and image-to-music generation. In our experiments, we evaluated AI-generated music scores produced by the system using both standard music metrics and innovative analysis that compares these compositions with original works. The system achieves an average key confidence of 85%, outperforming human composers at 79%, and aligns closely with established music theory standards, demonstrating its ability to generate diverse, human-like compositions. As a co-pilot tool, GenAIM can serve as a reliable music composition assistant and a possible educational composition tutor while simultaneously lowering the entry barrier for all aspiring musicians, which is innovative and significantly contributes to AI for music generation.


【9】Investigating self-supervised representations for audio-visual deepfake detection
标题:调查用于视听深度伪造检测的自我监督表示
链接:https://arxiv.org/abs/2511.17181

作者:Dragos-Alexandru Boldisor,Stefan Smeu,Dan Oneata,Elisabeta Oneata
摘要:自我监督表示在许多视觉和语音任务中表现出色,但它们在视听深度伪造检测方面的潜力仍有待探索。与以前的工作,使用这些功能在隔离或埋在复杂的架构,我们系统地评估它们跨模态(音频,视频,多模态)和域(嘴唇运动,通用的视觉内容)。我们评估三个关键方面:检测有效性,编码信息的可解释性和跨模态互补性。我们发现,大多数自监督特征都能捕捉到与deepfake相关的信息,而且这些信息是互补的。此外,模型主要关注语义上有意义的区域,而不是虚假的工件。然而,没有一个能可靠地在数据集之间推广。这种泛化失败可能源于数据集特征,而不是特征本身锁定在表面模式上。这些结果揭示了自监督表示用于深度伪造检测的前景和基本挑战:虽然它们学习有意义的模式,但实现强大的跨域性能仍然是难以捉摸的。
摘要:Self-supervised representations excel at many vision and speech tasks, but their potential for audio-visual deepfake detection remains underexplored. Unlike prior work that uses these features in isolation or buried within complex architectures, we systematically evaluate them across modalities (audio, video, multimodal) and domains (lip movements, generic visual content). We assess three key dimensions: detection effectiveness, interpretability of encoded information, and cross-modal complementarity. We find that most self-supervised features capture deepfake-relevant information, and that this information is complementary. Moreover, models primarily attend to semantically meaningful regions rather than spurious artifacts. Yet none generalize reliably across datasets. This generalization failure likely stems from dataset characteristics, not from the features themselves latching onto superficial patterns. These results expose both the promise and fundamental challenges of self-supervised representations for deepfake detection: while they learn meaningful patterns, achieving robust cross-domain performance remains elusive.


【10】Device-Guided Music Transfer
标题:设备引导音乐传输
链接:https://arxiv.org/abs/2511.17136

作者:Manh Pham Hung,Changshuo Hu,Ting Dang,Dong Ma
摘要:设备引导的音乐传输可为缺少设备的用户调整在看不见的设备上的播放。现有方法主要集中于修改音色、节奏、和声或乐器以模仿流派或艺术家,而忽略了回放设备的各种硬件属性(即,发言人)。因此,我们提出DeMT,它使用视觉语言模型来提取设备嵌入,将扬声器的频率响应曲线处理为线图。然后,这些嵌入通过特征线性调制来调节混合Transformer。DeMT在自我收集的数据集上进行了微调,为看不见的设备实现了有效的扬声器风格传输和强大的Few-Shot适应,支持设备风格增强和质量增强等应用。
摘要:Device-guided music transfer adapts playback across unseen devices for users who lack them. Existing methods mainly focus on modifying the timbre, rhythm, harmony, or instrumentation to mimic genres or artists, overlooking the diverse hardware properties of the playback device (i.e., speaker). Therefore, we propose DeMT, which processes a speaker's frequency response curve as a line graph using a vision-language model to extract device embeddings. These embeddings then condition a hybrid transformer via feature-wise linear modulation. Fine-tuned on a self-collected dataset, DeMT enables effective speaker-style transfer and robust few-shot adaptation for unseen devices, supporting applications like device-style augmentation and quality enhancement.


【11】Better audio representations are more brain-like: linking model-brain alignment with performance in downstream auditory tasks
标题:更好的音频表示更像大脑:将模型-大脑对齐与下游听觉任务的表现联系起来
链接:https://arxiv.org/abs/2511.16849

作者:Leonardo Pepino,Pablo Riera,Juan Kamienkowski,Luciana Ferrer
摘要:人工神经网络(ANN)是越来越强大的大脑计算模型,但目前尚不清楚提高其任务性能是否也会使其内部表征更类似于大脑信号。为了在听觉领域解决这个问题,我们量化了来自两个独立的fMRI数据集的36种不同音频模型的内部表示和大脑活动之间的对齐。使用体素和分量回归以及表示相似性分析(RSA),我们发现,最近的自监督音频模型在各种下游任务中具有强大的性能,比旧的和更专业的模型更好地预测了听觉皮层活动。为了评估音频表示的质量,我们在HEAREval基准的6个听觉任务中评估了这些模型,包括音乐,语音和环境声音。这揭示了模型的整体任务表现与其与大脑表征的一致性之间的强正皮尔逊相关性(r>0.7$)。最后,我们分析了EnCodecMAE预训练过程中音频和大脑表征之间相似性的演变。我们发现,大脑的相似性逐渐增加,并在预训练过程中早期出现,尽管模型没有为此目标进行明确优化。这表明,类脑表征可能是学习从自然主义音频数据中重建缺失信息的一个新兴副产品。
摘要:Artificial neural networks (ANNs) are increasingly powerful models of brain computation, yet it remains unclear whether improving their task performance also makes their internal representations more similar to brain signals. To address this question in the auditory domain, we quantified the alignment between the internal representations of 36 different audio models and brain activity from two independent fMRI datasets. Using voxel-wise and component-wise regression, and representation similarity analysis (RSA), we found that recent self-supervised audio models with strong performance in diverse downstream tasks are better predictors of auditory cortex activity than older and more specialized models. To assess the quality of the audio representations, we evaluated these models in 6 auditory tasks from the HEAREval benchmark, spanning music, speech, and environmental sounds. This revealed strong positive Pearson correlations ($r>0.7$) between a model's overall task performance and its alignment with brain representations. Finally, we analyzed the evolution of the similarity between audio and brain representations during the pretraining of EnCodecMAE. We discovered that brain similarity increases progressively and emerges early during pretraining, despite the model not being explicitly optimized for this objective. This suggests that brain-like representations can be an emergent byproduct of learning to reconstruct missing information from naturalistic audio data.


eess.AS音频处理


【1】Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
标题:重新审视学习通用音频表示的音频语言预训练
链接:https://arxiv.org/abs/2511.16757

作者:Wei-Cheng Tseng,Xuanru Zhou,Mingyue Huo,Yiwen Shao,Hao Zhang,Dong Yu
备注:Work in progress
摘要:音频语言预训练为通用音频理解提供了希望,但与视觉对应物相比,仍然未得到充分探索。虽然像CLIP这样的视觉语言模型被广泛采用,但现有的音频语言模型主要擅长检索任务,而作为通用编码器的采用有限。我们确定了三个主要障碍:有限的大规模音频文本语料库,字幕的多样性不足,缺乏系统的探索和评估。为此,我们引入CaptionStew,这是一个10.7M的字幕数据集,它聚合了跨多个领域和字幕风格的各种开源音频文本语料库。使用此资源,我们进行了第一次全面评估,比较了语音,音乐和环境声音任务中音频表示学习的对比和字幕目标。我们的研究结果表明,音频语言预训练产生有竞争力的,可转移的表示。通过系统的数据缩放实验,我们揭示了互补的客观优势:对比学习在较小的尺度上实现了卓越的数据效率,而字幕在涉及语言的音频理解任务上表现出更好的可扩展性。我们还发现,常见的监督初始化实践提供了规模收益递减,挑战当前的方法。这些发现建立了音频语言预训练作为一个可行的途径,以通用的音频表示,指导未来的研究。为了加快进度,我们发布了数据准备配方、训练协议和预训练模型,为通用音频理解铺平了道路。
摘要:Audio-language pretraining holds promise for general-purpose audio understanding, yet remains underexplored compared to its vision counterpart. While vision-language models like CLIP serve as widely adopted foundations, existing audio-language models primarily excel at retrieval tasks with limited adoption as general-purpose encoders. We identify three key barriers: limited large-scale audio-text corpora, insufficient caption diversity, and lack of systematic exploration and evaluation. To this end, we introduce CaptionStew, a 10.7M caption dataset aggregating diverse open-source audio-text corpora across multiple domains and captioning styles. Using this resource, we conduct the first comprehensive evaluation comparing contrastive and captioning objectives for audio representation learning across speech, music, and environmental sound tasks. Our results demonstrate that audio-language pretraining yields competitive, transferable representations. Through systematic data-scaling experiments, we reveal complementary objective strengths: contrastive learning achieves superior data efficiency at smaller scales, while captioning demonstrates better scalability on language-involved audio understanding tasks. We also find that common supervised initialization practices provide diminishing returns at scale, challenging current approaches. These findings establish audio-language pretraining as a viable pathway toward general-purpose audio representations, guiding future research. To accelerate progress, we release data preparation recipes, training protocols, and pretrained models, paving the way toward universal audio understanding.


【2】Semantic and Semiotic Interplays in Text-to-Audio AI: Exploring Cognitive Dynamics and Musical Interactions
标题:文本到音频人工智能中的语义和符号相互作用:探索认知动力学和音乐相互作用
链接:https://arxiv.org/abs/2511.17429

作者:Guilherme Coelho
摘要:本文研究了人工智能(AI)中新兴的文本到音频范式,研究了其对音乐创作,解释和认知的变革性影响。我探讨了复杂的语义和符号的相互作用时发生的描述性自然语言提示翻译成细微差别的声音对象在文本到音频的方式。从结构主义和后结构主义的角度,以及图式动力学和元认知的认知理论,本文探讨了这些人工智能系统如何重新配置音乐的意义过程和导航既定的认知框架。该研究分析了人工智能介导的音乐中发挥作用的一些认知动力学,包括图式同化和适应,元认知反射和建设性感知的过程。本文认为,文本到音频的人工智能模型作为音乐意义的准对象,同时稳定和不稳定的传统形式,同时培养新的聆听模式和审美反射。本研究以Udio为主要案例研究,探讨这些模型如何导航语言提示和声音输出之间的阈空间。这一过程不仅产生了新颖的音乐表现形式,而且促使听众进行批判性的和“结构意识的倾听”。“,鼓励更深入地了解音乐的结构,符号学的细微差别,以及塑造我们音乐认知的社会文化背景。本文最后反思了文本到音频AI模型作为认知工具和准对象的潜力,促进了音乐互动的重大转变,并邀请用户对音乐的认知和文化基础有更细致入微的理解。
摘要:This paper investigates the emerging text-to-audio paradigm in artificial intelligence (AI), examining its transformative implications for musical creation, interpretation, and cognition. I explore the complex semantic and semiotic interplays that occur when descriptive natural language prompts are translated into nuanced sound objects across the text-to-audio modality. Drawing from structuralist and post-structuralist perspectives, as well as cognitive theories of schema dynamics and metacognition, the paper explores how these AI systems reconfigure musical signification processes and navigate established cognitive frameworks. The research analyzes some of the cognitive dynamics at play in AI-mediated musicking, including processes of schema assimilation and accommodation, metacognitive reflection, and constructive perception. The paper argues that text-to-audio AI models function as quasi-objects of musical signification, simultaneously stabilizing and destabilizing conventional forms while fostering new modes of listening and aesthetic reflexivity.Using Udio as a primary case study, this study explores how these models navigate the liminal spaces between linguistic prompts and sonic outputs. This process not only generates novel musical expressions but also prompts listeners to engage in forms of critical and "structurally-aware listening.", encouraging a deeper understanding of music's structures, semiotic nuances, and the socio-cultural contexts that shape our musical cognition. The paper concludes by reflecting on the potential of text-to-audio AI models to serve as epistemic tools and quasi-objects, facilitating a significant shift in musical interactions and inviting users to develop a more nuanced comprehension of the cognitive and cultural foundations of music.


【3】AI in Music and Sound: Pedagogical Reflections, Post-Structuralist Approaches and Creative Outcomes in Seminar Practice
标题:音乐和声音中的人工智能:研讨会实践中的教学反思、后结构主义方法和创造性成果
链接:https://arxiv.org/abs/2511.17425

作者:Guilherme Coelho
摘要:本文介绍了一个教学和概念帐户的过程中,AI在音乐和声音:模式,工具和创造性的应用程序,提供内的音乐信息学和媒体艺术模块的硕士在音频通信。该课程让学生参与一系列AI模式,如符号合成,语音合成,音色转移,神经音频合成和文本到音频系统,将理论反思与基于实践的实验相结合。它的核心教学举措是一个成对的练习设计:每一种形式首先通过其预期的启示,然后通过一个故意重新定义或"误用"的练习,表面代表性的限制和替代行为。在媒介理论和后结构主义探究的框架下,我们将人工智能视为一个跨模态的机器人-一个跨文本,符号,音色和音频领域翻译和干扰音乐符号的系统。来自学生工作和反思的证据表明,技术流利性,媒介意识和批判性素养的增长,以及实验方法和过程导向的听力的培养。本文概述了课程架构,评估设计和代表性项目,并提炼出一套人工智能音乐教学法的设计模式(例如,文本到音频中的无条件相互作用和语义不稳定;音色转移中的潜在空间唯物主义)。它的结论是教学建议,将创造性实践与媒介意识和人工智能技术的文化认知分析相结合,使学生能够参与人工智能如何被理解,开发和部署在创意社区。
摘要:This paper presents a pedagogical and conceptual account of the course AI in Music and Sound: Modalities, Tools and Creative Applications, offered within the Music Informatics and Media Art module of an M.Sc. in Audio Communication. The course engaged students with a range of AI modalities such as symbolic composition, voice synthesis, timbre transfer, neural audio synthesis, and text-to-audio systems, combining theoretical reflection with practice-based experimentation. Its central pedagogical move is a paired-études design: each modality is approached first through its intended affordances and then through a deliberately reframed or "misused" exercise that surfaces representational limits and alternative behaviours. Framed by medium theory and post-structuralist inquiry, we treated AI as a transmodal conduit-a system that translates and perturbs musical signs across textual, symbolic, timbral and audio domains. Evidence from student work and reflection indicates growth in technical fluency, medium awareness, and critical literacy, alongside the cultivation of experimental method and process-oriented listening. The paper outlines the course architecture, assessment design, and representative projects, and distils a set of design patterns for AI-music pedagogy (eg., prompt-conditioned interplays and semantic destabilisation in text-to-audio; latent space materialism in timbre transfer). It concludes with pedagogical recommendations that integrate creative practice with medium awareness and with cultural-epistemic analysis of AI technologies, preparing students to participate in how AI is understood, developed, and deployed with creative communities.


【4】The Artist is Present: Traces of Artists Resigind and Spawning in Text-to-Audio AI
标题:艺术家在场:艺术家在文本转音频人工智能中辞职和繁衍的痕迹
链接:https://arxiv.org/abs/2511.17404

作者:Guilherme Coelho
摘要:文本到音频(TTA)系统正在迅速改变音乐创作和发行,像Udio和Suno这样的平台每天生成数千首曲目,并融入主流音乐平台和生态系统。这些系统在大量未公开的数据集上训练,从根本上重塑了音乐的制作、复制和消费方式。本文提出的经验证据表明,艺术家的条件区域可以系统地通过基于元标签的提示设计微定位,有效地使产卵的艺术家一样的内容,通过战略提示工程。通过系统地探索基于元标签的提示工程技术,这项研究揭示了用户如何访问特定艺术家的独特声音签名,证明它们包含在训练数据集中。使用从公共音乐分类中提取的描述符星座,本文展示了与Bon Iver,Philip Glass,Panda Bear和William Basinski等艺术家的可复制接近。结果表明,稳定的文本音频对应与艺术家特定的训练信号一致,使精确遍历风格的微位置,而无需明确命名的艺术家。这种召唤艺术家特定输出的能力表明,艺术家的创造性作品是这些系统产生新内容的基础材料,通常没有明确的同意或归属。从概念上讲,这项工作澄清了文本描述符如何在高维表示空间中充当导航线索;从方法上讲,它提供了一个可复制的协议,用于审计风格归纳。这些发现为治理--归属、同意和披露标准--和创造性实践提出了直接的质疑,在创造性实践中,诱导的风格接近使所有权、复制、模仿、创造性机构和算法创造的道德之间的界限复杂化。
摘要:Text-to-audio (TTA) systems are rapidly transforming music creation and distribution, with platforms like Udio and Suno generating thousands of tracks daily and integrating into mainstream music platforms and ecosystems. These systems, trained on vast and largely undisclosed datasets, are fundamentally reshaping how music is produced, reproduced and consumed. This paper presents empirical evidence that artist-conditioned regions can be systematically microlocated through metatag-based prompt design, effectively enabling the spawning of artist-like content through strategic prompt engineering. Through systematic exploration of metatag-based prompt engineering techniques this research reveals how users can access the distinctive sonic signatures of specific artists, evidencing their inclusion in training datasets. Using descriptor constellations drawn from public music taxonomies, the paper demonstrates reproducible proximity to artists such as Bon Iver, Philip Glass, Panda Bear and William Basinski. The results indicate stable text-audio correspondences consistent with artist-specific training signals, enabling precise traversal of stylistic microlocations without explicitly naming artists. This capacity to summon artist-specific outputs shows that artists' creative works fuction as foundational material from which these systems generate new content, often without explicit consent or attribuition. Conceptually, the work clarifies how textual descriptors act as navigational cues in high-dimensional representation spaces; methodologically, it provides a replicable protocol for auditing stylistic inducibility. The findings raise immediate queestions for governance-attribution, consent and disclosure standards-and for creative practice, where induced stylistic proximity complicates boundaries between ownership, reproduction, imitation, creative agency and the ethics of algorithmic creation.


【5】Robot Confirmation Generation and Action Planning Using Long-context Q-Former Integrated with Multimodal LLM
标题:基于长上下文Q-Former和多模态LLM的机器人确认生成和动作规划
链接:https://arxiv.org/abs/2511.17335

作者:Chiori Hori,Yoshiki Masuyama,Siddarth Jain,Radu Corcodel,Devesh Jha,Diego Romeres,Jonathan Le Roux
备注:Accepted to ASRU 2025
摘要:为了实现共同目标,人机协作需要机器人理解人类的行为以及与周围环境的交互。本文主要研究基于人机对话的人机交互,人机对话依赖于机器人的动作确认和动作步骤的生成,使用多模态场景理解。国家的最先进的方法使用多模态Transformers生成机器人动作步骤与机器人动作确认对齐,从一个单一的剪辑显示由多个微步骤组成的任务。虽然在整个视频中,对长视野任务的动作相互依赖,但目前的方法主要集中在剪辑级处理上,而没有利用长上下文信息。本文提出了一种长上下文Q-former,在完整的视频中结合左右上下文依赖。此外,本文提出了一种文本条件化方法,将文本嵌入直接馈送到LLM解码器中,以减轻Q-former对文本中信息的高度抽象。YouCook 2语料库的实验表明,确认生成的准确性是动作规划性能的主要因素。此外,我们证明了长上下文Q-前提高了确认和行动规划,通过整合VideoLLaMA 3。
摘要:Human-robot collaboration towards a shared goal requires robots to understand human action and interaction with the surrounding environment. This paper focuses on human-robot interaction (HRI) based on human-robot dialogue that relies on the robot action confirmation and action step generation using multimodal scene understanding. The state-of-the-art approach uses multimodal transformers to generate robot action steps aligned with robot action confirmation from a single clip showing a task composed of multiple micro steps. Although actions towards a long-horizon task depend on each other throughout an entire video, the current approaches mainly focus on clip-level processing and do not leverage long-context information. This paper proposes a long-context Q-former incorporating left and right context dependency in full videos. Furthermore, this paper proposes a text-conditioning approach to feed text embeddings directly into the LLM decoder to mitigate the high abstraction of the information in text by Q-former. Experiments with the YouCook2 corpus show that the accuracy of confirmation generation is a major factor in the performance of action planning. Furthermore, we demonstrate that the long-context Q-former improves the confirmation and action planning by integrating VideoLLaMA3.


机器翻译由腾讯交互翻译提供,仅供参考