今日论文合集:cs.SD语音10篇,eess.AS音频处理3篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Linear Complexity Self-Supervised Learning for Music Understanding with Random Quantizer
标题:使用随机量化器进行音乐理解的线性复杂性自我监督学习
链接:https://arxiv.org/abs/2601.09603

作者:Petros Vavaroutsos,Theodoros Palamas,Pantelis Vikatos
备注:accepted by ACM/SIGAPP Symposium on Applied Computing (SAC 2026)
摘要:近年来,基础模型由于其出色的性能而变得非常受欢迎,主要是在首次引入它们的自然语言(NLP)任务中。这些模型通常由数亿甚至数十亿个参数组成,使其在训练和生产系统中成为资源密集型,从而导致成本增加。本文重点研究了在应用于音乐信息检索(MIR)任务时减少基金会的模型大小。我们的研究结合了Branchformer架构和SummaryMixing,它们首先应用于语音识别,以及随机量化过程。为了促进可重复性,我们对公开可用的数据集进行了预训练,并辅之以与文献中报道的其他私人数据集规模相当的专有数据集。我们通过使用由各种下游MIR任务组成的框架来确保强大的评估。我们的研究结果表明,我们的架构实现了竞争力的性能相比,其他国家的最先进的模型,使用多头自注意力,同时减少模型的大小从8.5%到12.3%。
摘要:In recent years, foundation models have become very popular due to their exceptional performance, mainly in natural language (NLP) tasks where they were first introduced. These models usually consist of hundreds of millions, or even billions, of parameters, making them resource-intensive during training and in production systems, leading to increased costs. This paper focuses on the reduction of a foundation's model size when applied to music information retrieval (MIR) tasks. Our research combines the Branchformer architecture with SummaryMixing, which were first applied in speech recognition, along with a random quantization process. To facilitate reproducibility, we conduct pre-training on publicly available datasets, complemented by a proprietary dataset comparable in scale to other private datasets reported in the literature. We ensure robust evaluation by using a framework consisting of a variety of downstream MIR tasks. Our results show that our architecture achieves competitive performance when compared with other state-of-the-art models that use multi-head self-attention, while reducing the model size from 8.5% up to 12.3%.


【2】Towards Realistic Synthetic Data for Automatic Drum Transcription
标题:迈向自动鼓抄写的真实合成数据
链接:https://arxiv.org/abs/2601.09520

作者:Pierfrancesco Melucci,Paolo Merialdo,Taketo Akama
摘要:深度学习模型定义了自动鼓转录(ADT)的最新技术,但它们的性能取决于大规模的配对音频数据集,而这些数据集是稀缺的。使用合成数据的现有解决方案通常会引入显著的域差距,因为它们通常依赖于缺乏声学多样性的低保真度SoundFont库。虽然高质量的一次性样本提供了一个更好的选择,但它们没有适合训练的标准化大规模格式。本文介绍了一种新的ADT范式,可以避免对配对音频-MIDI训练数据的需求。我们的主要贡献是一个半监督的方法来自动策划一个大的和不同的语料库的一杆鼓样本从未标记的音频源。然后,我们使用这个语料库来合成一个高质量的数据集,从单独的数据文件,我们用它来训练一个序列到序列的转录模型。我们在ENST和MDB测试集上评估了我们的模型,在那里它获得了最先进的结果,显著优于完全监督方法和以前的合成数据方法。复制我们实验的代码可在https://github.com/pier-maker92/ADT_STR上公开获得。
摘要:Deep learning models define the state-of-the-art in Automatic Drum Transcription (ADT), yet their performance is contingent upon large-scale, paired audio-MIDI datasets, which are scarce. Existing workarounds that use synthetic data often introduce a significant domain gap, as they typically rely on low-fidelity SoundFont libraries that lack acoustic diversity. While high-quality one-shot samples offer a better alternative, they are not available in a standardized, large-scale format suitable for training. This paper introduces a new paradigm for ADT that circumvents the need for paired audio-MIDI training data. Our primary contribution is a semi-supervised method to automatically curate a large and diverse corpus of one-shot drum samples from unlabeled audio sources. We then use this corpus to synthesize a high-quality dataset from MIDI files alone, which we use to train a sequence-to-sequence transcription model. We evaluate our model on the ENST and MDB test sets, where it achieves new state-of-the-art results, significantly outperforming both fully supervised methods and previous synthetic-data approaches. The code for reproducing our experiments is publicly available at https://github.com/pier-maker92/ADT_STR


【3】Analysis of the Maximum Prediction Gain of Short-Term Prediction on Sustained Speech
标题:持续语音短期预测的最大预测收益分析
链接:https://arxiv.org/abs/2601.09461

作者:Reemt Hinrichs,Muhamad Fadli Damara,Stephan Preihs,Jörn Ostermann
备注:Rejected at Eurasip for practical irrelevancy. Submitted here for reference. Originally accepted at DCC 2020 (Poster) but withdrawn due to page count limit
摘要:信号预测被广泛地用于,例如,经济预测、回声消除和数据压缩,特别是语音和音乐的预测编码。预测编码算法通过信号预测来降低数据传输或存储所需的比特率。预测增益是预测器质量的应用信号编码中的经典度量,因为它将均方预测误差与预测编码器的信号-量化噪声联系起来。为了评估预测器模型,需要关于独立于预测器模型的最大可实现预测增益的知识。在本文中,应用Nadaraya-Watson核回归(NWKR)和信息论上界来分析新记录的持续语音/音素数据集的预测增益上界。结果发现,对于清音语音的线性预测器总是达到最大预测增益在最多0.3分贝。在有声语音上,发现最佳单抽头预测器是线性的,但是从两个抽头开始,发现最大可实现的预测增益比线性预测器的预测增益高大约2dB到6dB。观察到说话者/受试者之间的显着差异。   创建的数据集以及代码可以根据要求获得用于研究目的。
摘要:Signal prediction is widely used in, e.g., economic forecasting, echo cancellation and in data compression, particularly in predictive coding of speech and music. Predictive coding algorithms reduce the bit-rate required for data transmission or storage by signal prediction. The prediction gain is a classic measure in applied signal coding of the quality of a predictor, as it links the mean-squared prediction error to the signal-to-quantization-noise of predictive coders. To evaluate predictor models, knowledge about the maximum achievable prediction gain independent of a predictor model is desirable. In this manuscript, Nadaraya-Watson kernel-regression (NWKR) and an information theoretic upper bound are applied to analyze the upper bound of the prediction gain on a newly recorded dataset of sustained speech/phonemes. It was found that for unvoiced speech a linear predictor always achieves the maximum prediction gain within at most 0.3 dB. On voiced speech, the optimum one-tap predictor was found to be linear but starting with two taps, the maximum achievable prediction gain was found to be about 2 dB to 6 dB above the prediction gain of the linear predictor. Significant differences between speakers/subjects were observed.   The created dataset as well as the code can be obtained for research purpose upon request.


【4】Population-Aligned Audio Reproduction With LLM-Based Equalizers
标题:使用基于LLM的均衡器实现人口一致的音频复制
链接:https://arxiv.org/abs/2601.09448

作者:Ioannis Stylianou,Jon Francombe,Pablo Martinez-Nuevo,Sven Ewan Shepstone,Zheng-Hua Tan
备注:12 pages, 13 figures, 2 tables, IEEE JSTSP journal submission under first revision
摘要:传统的音频均衡是静态过程,其需要手动和繁琐的调整以适应改变的收听上下文(例如,情绪、位置或社会环境)。在本文中,我们介绍了一种基于大语言模型(LLM)的替代方案,将自然语言文本提示映射到均衡设置。这使得对话式方法能够用于声音系统控制。通过利用从受控听力实验中收集的数据,我们的模型利用上下文学习和参数有效的微调技术,以可靠地与人群偏好的均衡设置保持一致。我们的评估方法,利用分布指标,捕捉用户的不同偏好,显示统计上显着的改进,分布对齐随机抽样和静态预设基线。这些结果表明,LLM可以作为“人工均衡器”,有助于开发更容易访问的,上下文感知的和专家级的音频调谐方法。
摘要:Conventional audio equalization is a static process that requires manual and cumbersome adjustments to adapt to changing listening contexts (e.g., mood, location, or social setting). In this paper, we introduce a Large Language Model (LLM)-based alternative that maps natural language text prompts to equalization settings. This enables a conversational approach to sound system control. By utilizing data collected from a controlled listening experiment, our models exploit in-context learning and parameter-efficient fine-tuning techniques to reliably align with population-preferred equalization settings. Our evaluation methods, which leverage distributional metrics that capture users' varied preferences, show statistically significant improvements in distributional alignment over random sampling and static preset baselines. These results indicate that LLMs could function as "artificial equalizers," contributing to the development of more accessible, context-aware, and expert-level audio tuning methods.


【5】Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
标题:Speech-Hands:一种全感知的语音识别和音频推理的自反射语音分析方法
链接:https://arxiv.org/abs/2601.09413

作者:Zhen Wan,Chao-Han Huck Yang,Jinchuan Tian,Hanrong Ye,Ankita Pasad,Szu-wei Fu,Arushi Goel,Ryo Hachiuma,Shizhe Diao,Kunal Dhawan,Sreyan Ghosh,Yusuke Hirota,Zhehuai Chen,Rafael Valle,Ehsan Hosseini Asl,Chenhui Chu,Shinji Watanabe,Yu-Chiang Frank Wang,Boris Ginsburg
备注:Preprint. The version was submitted in October 2025
摘要:我们引入了一个语音代理框架,学习一个关键的全方位理解技能:知道什么时候信任自己,什么时候咨询外部音频感知。我们的工作是由一个关键但违反直觉的发现所推动的:在语音识别和外部声音理解任务上天真地微调全模型通常会降低性能,因为模型很容易被嘈杂的假设误导。为了解决这个问题,我们的框架,Speech-Hands,将这个问题重新定义为一个明确的自我反思决定。这个可学习的反射原语被证明可以有效地防止模型被有缺陷的外部候选者破坏。我们表明,这种代理的行动机制自然从语音识别到复杂的,多选择的音频推理。在OpenASR排行榜上,Speech-Hands在七个基准测试中的表现始终优于强基线12.1%。该模型在音频问答决策上也达到了77.37%的准确率和高F1,在不同的音频问答数据集上表现出了强大的泛化能力和可靠性。通过统一感知和决策,我们的工作为实现更可靠和更有弹性的音频智能提供了一条切实可行的道路。
摘要:We introduce a voice-agentic framework that learns one critical omni-understanding skill: knowing when to trust itself versus when to consult external audio perception. Our work is motivated by a crucial yet counterintuitive finding: naively fine-tuning an omni-model on both speech recognition and external sound understanding tasks often degrades performance, as the model can be easily misled by noisy hypotheses. To address this, our framework, Speech-Hands, recasts the problem as an explicit self-reflection decision. This learnable reflection primitive proves effective in preventing the model from being derailed by flawed external candidates. We show that this agentic action mechanism generalizes naturally from speech recognition to complex, multiple-choice audio reasoning. Across the OpenASR leaderboard, Speech-Hands consistently outperforms strong baselines by 12.1% WER on seven benchmarks. The model also achieves 77.37% accuracy and high F1 on audio QA decisions, showing robust generalization and reliability across diverse audio question answering datasets. By unifying perception and decision-making, our work offers a practical path toward more reliable and resilient audio intelligence.


【6】SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing
标题:SLAM-LLM:模块化、开源多模式大型语言模型框架和语音、语言、音频和音乐处理的最佳实践
链接:https://arxiv.org/abs/2601.09385

作者:Ziyang Ma,Guanrou Yang,Wenxi Chen,Zhifu Gao,Yexing Du,Xiquan Li,Zhisheng Zheng,Haina Zhu,Jianheng Zhuo,Zheshu Song,Ruiyang Xu,Tiranrui Wang,Yifan Yang,Yanqiao Zhu,Zhikang Niu,Liumeng Xue,Yinghao Ma,Ruibin Yuan,Shiliang Zhang,Kai Yu,Eng Siong Chng,Xie Chen
备注:Published in IEEE Journal of Selected Topics in Signal Processing (JSTSP)
摘要:最近开源多模态大型语言模型(MLLM)框架(如LLaVA)的激增为人工智能开发人员和研究人员提供了一个方便的开端。然而,大多数MLLM框架将视觉作为主要的输入模态,并且对语音、音频和音乐的模态提供有限的深入支持。这种情况阻碍了音频语言模型的发展,并迫使研究人员花费大量精力编写代码和调整超参数。我们介绍了SLAM-LLM,这是一个开源的深度学习框架,旨在训练定制的MLLM,专注于语音,语言,音频和音乐处理。SLAM-LLM提供不同编码器、投影仪、LLM和参数高效微调插件的模块化配置。SLAM-LLM还包括用于主流任务的详细训练和推理配方,以及基于LLM的自动语音识别(ASR),自动音频字幕(AAC)和音乐字幕(MC)等高性能检查点。其中一些配方已经达到或接近最先进的性能,一些相关技术也被学术论文所接受。我们希望SLAM-LLM将加速迭代,开发,数据工程和研究人员的模型培训。我们致力于通过这个开源框架不断推进基于音频的MLLM,并呼吁社区为基于LLM的语音,音频和音乐处理做出贡献。
摘要:The recent surge in open-source Multimodal Large Language Models (MLLM) frameworks, such as LLaVA, provides a convenient kickoff for artificial intelligence developers and researchers. However, most of the MLLM frameworks take vision as the main input modality, and provide limited in-depth support for the modality of speech, audio, and music. This situation hinders the development of audio-language models, and forces researchers to spend a lot of effort on code writing and hyperparameter tuning. We present SLAM-LLM, an open-source deep learning framework designed to train customized MLLMs, focused on speech, language, audio, and music processing. SLAM-LLM provides a modular configuration of different encoders, projectors, LLMs, and parameter-efficient fine-tuning plugins. SLAM-LLM also includes detailed training and inference recipes for mainstream tasks, along with high-performance checkpoints like LLM-based Automatic Speech Recognition (ASR), Automated Audio Captioning (AAC), and Music Captioning (MC). Some of these recipes have already reached or are nearing state-of-the-art performance, and some relevant techniques have also been accepted by academic papers. We hope SLAM-LLM will accelerate iteration, development, data engineering, and model training for researchers. We are committed to continually pushing forward audio-based MLLMs through this open-source framework, and call on the community to contribute to the LLM-based speech, audio and music processing.


【7】Research on Piano Timbre Transformation System Based on Diffusion Model
标题:基于扩散模型的钢琴音色转换系统研究
链接:https://arxiv.org/abs/2601.09333

作者:Chun-Chieh Hsu,Tsai-Ling Hsu,Chen-Chen Yeh,Shao-Chien Lu,Cheng-Han Wu,Bing-Ze Liu,Timothy K. Shih,Yu-Cheng Lin
摘要:本文提出了一种基于Diffusion结构的音色转换模型,用于将各种乐器演奏的音乐精确地转换为钢琴版本。该模型采用音高编码器和响度编码器来提取音乐的音高和响度特征,这些特征作为Dif-fusion模型解码器的条件输入,生成高质量的钢琴音色。实例分析结果表明,该模型在音高精度和音色相似性方面表现优异,在不同音乐风格(古典、爵士、流行)和长度(从短片段到完整片段)之间保持稳定的转换。特别是,该模型在处理快速变化的音符和复杂的音乐结构时仍能保持较高的音质和精度,表现出良好的泛化能力。此外,该模型具有实时音乐转换的潜力,适用于现场表演和数字音乐创作工具。未来的研究将集中在增强响度动态的处理,并纳入额外的音乐功能(如音色变化和节奏的复杂性),以提高模型的适应性和表现力。我们计划探索该模型在其他音色转换任务中的应用潜力,例如将人声转换为乐器声音或与数字钢琴集成,进一步扩大基于扩散的音色转换模型在音乐生成领域的应用范围。
摘要:We propose a timbre conversion model based on the Diffusion architecture de-signed to precisely translate music played by various instruments into piano ver-sions. The model employs a Pitch Encoder and Loudness Encoder to extract pitch and loudness features of the music, which serve as conditional inputs to the Dif-fusion Model's decoder, generating high-quality piano timbres. Case analysis re-sults show that the model performs excellently in terms of pitch accuracy and timbral similarity, maintaining stable conversion across different musical styles (classical, jazz, pop) and lengths (from short clips to full pieces). Particularly, the model maintains high sound quality and accuracy even when dealing with rapidly changing notes and complex musical structures, demonstrating good generaliza-tion capability. Additionally, the model has the potential for real-time musical conversion and is suitable for live performances and digital music creation tools. Future research will focus on enhancing the handling of loudness dynamics and incorporating additional musical features (such as timbral variations and rhythmic complexity) to improve the model's adaptability and expressiveness. We plan to explore the model's application potential in other timbre conversion tasks, such as converting vocals to instrumental sounds or integration with MIDI digital pianos, further expanding the application scope of the Diffusion-based timbre conversion model in the field of music generation.


【8】DSA-Tokenizer: Disentangled Semantic-Acoustic Tokenization via Flow Matching-based Hierarchical Fusion
标题:DSA-令牌化器:通过基于流匹配的分层融合来解开语义-声学令牌化
链接:https://arxiv.org/abs/2601.09239

作者:Hanlin Zhang,Daxin Tan,Dehua Tao,Xiao Chen,Haochen Tan,Yunhe Li,Yuchen Cao,Jianping Wang,Linqi Song
摘要:语音标记器是离散语音大语言模型(Speech LLM)的基石。现有的分词器要么优先进行语义编码,要么将语义内容与声学风格不可分割地融合在一起,要么实现了不完全的语义-声学分离。为了实现更好的解纠缠,我们提出了DSA-Tokenizer,它通过不同的优化约束明确地将语音解纠缠成离散的语义和声学令牌。具体而言,语义令牌由ASR监督以捕获语言内容,而声学令牌专注于梅尔频谱图恢复以编码风格。为了消除两个序列之间严格的长度限制,我们引入了一个分层的流匹配解码器,进一步提高了语音生成质量,并且采用了一个联合重构-重组训练策略来加强这种分离。DSA标记器通过强大的解纠缠实现高保真重建和灵活重组,促进语音LLM中的可控生成。我们的分析强调了解开标记作为未来语音建模的关键范式。音频样本可在https://anonymous.4open.science/w/DSA_Tokenizer_demo/上获得。该代码和模型将在论文被接受后公开提供。
摘要:Speech tokenizers serve as the cornerstone of discrete Speech Large Language Models (Speech LLMs). Existing tokenizers either prioritize semantic encoding, fuse semantic content with acoustic style inseparably, or achieve incomplete semantic-acoustic disentanglement. To achieve better disentanglement, we propose DSA-Tokenizer, which explicitly disentangles speech into discrete semantic and acoustic tokens via distinct optimization constraints. Specifically, semantic tokens are supervised by ASR to capture linguistic content, while acoustic tokens focus on mel-spectrograms restoration to encode style. To eliminate rigid length constraints between the two sequences, we introduce a hierarchical Flow-Matching decoder that further improve speech generation quality.Furthermore, We employ a joint reconstruction-recombination training strategy to enforce this separation. DSA-Tokenizer enables high fidelity reconstruction and flexible recombination through robust disentanglement, facilitating controllable generation in speech LLMs. Our analysis highlights disentangled tokenization as a pivotal paradigm for future speech modeling. Audio samples are avaialble at https://anonymous.4open.science/w/DSA_Tokenizer_demo/. The code and model will be made publicly available after the paper has been accepted.


【9】Echoes of Ideology: Toward an Audio Analysis Pipeline to Unveil Character Traits in Historical Nazi Propaganda Films
标题:意识形态的回声:走向音频分析管道,揭示历史纳粹宣传片中的人物特征
链接:https://arxiv.org/abs/2601.08879

作者:Nicolas Ruth,Manuel Burghardt
摘要:本研究旨在探讨使用计算音频分析,以检查纳粹宣传电影中的意识形态叙事。它采用了三个步骤,说话人日记,音频转录和心理语言学分析,揭示了人物的意识形态模式。尽管目前的问题与扬声器日记,该方法提供了洞察性格特征和宣传叙事,建议可扩展的应用程序。
摘要:This study investigates the use of computational audio analysis to examine ideological narratives in Nazi propaganda films. Employing a three-step pipeline, speaker diarization, audio transcription and psycholinguistic analysis, it reveals ideological patterns in characters. Despite current issues with speaker diarization, the methodology provides insights into character traits and propaganda narratives, suggesting scalable applications.


【10】Semantic visually-guided acoustic highlighting with large vision-language models
标题:使用大型视觉语言模型的语义视觉引导声学突出显示
链接:https://arxiv.org/abs/2601.08871

作者:Junhua Huang,Chao Huang,Chenliang Xu
摘要:平衡对话,音乐和音效与伴随的视频对于沉浸式讲故事至关重要,但目前的音频混合工作流程仍然主要是手动和劳动密集型的。虽然最近的进步已经引入了视觉引导的声学突出任务,它隐含地使用多模态指导来重新平衡音频源,但仍不清楚哪些视觉方面作为条件反射信号最有效。我们通过系统研究深度视频理解是否可以改善音频混音来解决这一差距。使用文本描述作为视觉分析的代理,我们提示大型视觉语言模型提取六种类型的视觉语义方面,包括对象和角色外观,情感,相机焦点,音调,场景背景和推断的声音相关线索。通过大量的实验,相机焦点,色调和场景背景一致地产生感知混合质量的最大改进超过最先进的基线。我们的研究结果(i)确定哪些视觉语义线索最强烈地支持连贯和视觉对齐的音频混音,以及(ii)使用来自大型视觉语言模型的轻量级指导,概述了自动化电影级声音设计的实用路径。
摘要:Balancing dialogue, music, and sound effects with accompanying video is crucial for immersive storytelling, yet current audio mixing workflows remain largely manual and labor-intensive. While recent advancements have introduced the visually guided acoustic highlighting task, which implicitly rebalances audio sources using multimodal guidance, it remains unclear which visual aspects are most effective as conditioning signals.We address this gap through a systematic study of whether deep video understanding improves audio remixing. Using textual descriptions as a proxy for visual analysis, we prompt large vision-language models to extract six types of visual-semantic aspects, including object and character appearance, emotion, camera focus, tone, scene background, and inferred sound-related cues. Through extensive experiments, camera focus, tone, and scene background consistently yield the largest improvements in perceptual mix quality over state-of-the-art baselines. Our findings (i) identify which visual-semantic cues most strongly support coherent and visually aligned audio remixing, and (ii) outline a practical path toward automating cinema-grade sound design using lightweight guidance derived from large vision-language models.


eess.AS音频处理


【1】Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
标题:Speech-Hands:一种全感知的语音识别和音频推理的自反射语音分析方法
链接:https://arxiv.org/abs/2601.09413

作者:Zhen Wan,Chao-Han Huck Yang,Jinchuan Tian,Hanrong Ye,Ankita Pasad,Szu-wei Fu,Arushi Goel,Ryo Hachiuma,Shizhe Diao,Kunal Dhawan,Sreyan Ghosh,Yusuke Hirota,Zhehuai Chen,Rafael Valle,Ehsan Hosseini Asl,Chenhui Chu,Shinji Watanabe,Yu-Chiang Frank Wang,Boris Ginsburg
备注:Preprint. The version was submitted in October 2025
摘要:我们引入了一个语音代理框架,学习一个关键的全方位理解技能:知道什么时候信任自己,什么时候咨询外部音频感知。我们的工作是由一个关键但违反直觉的发现所推动的:在语音识别和外部声音理解任务上天真地微调全模型通常会降低性能,因为模型很容易被嘈杂的假设误导。为了解决这个问题,我们的框架,Speech-Hands,将这个问题重新定义为一个明确的自我反思决定。这个可学习的反射原语被证明可以有效地防止模型被有缺陷的外部候选者破坏。我们表明,这种代理的行动机制自然从语音识别到复杂的,多选择的音频推理。在OpenASR排行榜上,Speech-Hands在七个基准测试中的表现始终优于强基线12.1%。该模型在音频问答决策上也达到了77.37%的准确率和高F1,在不同的音频问答数据集上表现出了强大的泛化能力和可靠性。通过统一感知和决策,我们的工作为实现更可靠和更有弹性的音频智能提供了一条切实可行的道路。
摘要:We introduce a voice-agentic framework that learns one critical omni-understanding skill: knowing when to trust itself versus when to consult external audio perception. Our work is motivated by a crucial yet counterintuitive finding: naively fine-tuning an omni-model on both speech recognition and external sound understanding tasks often degrades performance, as the model can be easily misled by noisy hypotheses. To address this, our framework, Speech-Hands, recasts the problem as an explicit self-reflection decision. This learnable reflection primitive proves effective in preventing the model from being derailed by flawed external candidates. We show that this agentic action mechanism generalizes naturally from speech recognition to complex, multiple-choice audio reasoning. Across the OpenASR leaderboard, Speech-Hands consistently outperforms strong baselines by 12.1% WER on seven benchmarks. The model also achieves 77.37% accuracy and high F1 on audio QA decisions, showing robust generalization and reliability across diverse audio question answering datasets. By unifying perception and decision-making, our work offers a practical path toward more reliable and resilient audio intelligence.


【2】Echoes of Ideology: Toward an Audio Analysis Pipeline to Unveil Character Traits in Historical Nazi Propaganda Films
标题:意识形态的回声:走向音频分析管道,揭示历史纳粹宣传片中的人物特征
链接:https://arxiv.org/abs/2601.08879

作者:Nicolas Ruth,Manuel Burghardt
摘要:本研究旨在探讨使用计算音频分析,以检查纳粹宣传电影中的意识形态叙事。它采用了三个步骤,说话人日记,音频转录和心理语言学分析,揭示了人物的意识形态模式。尽管目前的问题与扬声器日记,该方法提供了洞察性格特征和宣传叙事,建议可扩展的应用程序。
摘要:This study investigates the use of computational audio analysis to examine ideological narratives in Nazi propaganda films. Employing a three-step pipeline, speaker diarization, audio transcription and psycholinguistic analysis, it reveals ideological patterns in characters. Despite current issues with speaker diarization, the methodology provides insights into character traits and propaganda narratives, suggesting scalable applications.


【3】Semantic visually-guided acoustic highlighting with large vision-language models
标题:使用大型视觉语言模型的语义视觉引导声学突出显示
链接:https://arxiv.org/abs/2601.08871

作者:Junhua Huang,Chao Huang,Chenliang Xu
摘要:平衡对话,音乐和音效与伴随的视频对于沉浸式讲故事至关重要,但目前的音频混合工作流程仍然主要是手动和劳动密集型的。虽然最近的进步已经引入了视觉引导的声学突出任务,它隐含地使用多模态指导来重新平衡音频源,但仍不清楚哪些视觉方面作为条件反射信号最有效。我们通过系统研究深度视频理解是否可以改善音频混音来解决这一差距。使用文本描述作为视觉分析的代理,我们提示大型视觉语言模型提取六种类型的视觉语义方面,包括对象和角色外观,情感,相机焦点,音调,场景背景和推断的声音相关线索。通过大量的实验,相机焦点,色调和场景背景一致地产生感知混合质量的最大改进超过最先进的基线。我们的研究结果(i)确定哪些视觉语义线索最强烈地支持连贯和视觉对齐的音频混音,以及(ii)使用来自大型视觉语言模型的轻量级指导,概述了自动化电影级声音设计的实用路径。
摘要:Balancing dialogue, music, and sound effects with accompanying video is crucial for immersive storytelling, yet current audio mixing workflows remain largely manual and labor-intensive. While recent advancements have introduced the visually guided acoustic highlighting task, which implicitly rebalances audio sources using multimodal guidance, it remains unclear which visual aspects are most effective as conditioning signals.We address this gap through a systematic study of whether deep video understanding improves audio remixing. Using textual descriptions as a proxy for visual analysis, we prompt large vision-language models to extract six types of visual-semantic aspects, including object and character appearance, emotion, camera focus, tone, scene background, and inferred sound-related cues. Through extensive experiments, camera focus, tone, and scene background consistently yield the largest improvements in perceptual mix quality over state-of-the-art baselines. Our findings (i) identify which visual-semantic cues most strongly support coherent and visually aligned audio remixing, and (ii) outline a practical path toward automating cinema-grade sound design using lightweight guidance derived from large vision-language models.


机器翻译由腾讯交互翻译提供,仅供参考