今日论文合集:cs.SD语音11篇,eess.AS音频处理10篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】TASLA: Text-Aligned Speech Tokens with Multiple Layer-Aggregation
标题:TASLA:具有多层聚合的文本对齐语音令牌
链接:https://arxiv.org/abs/2510.14934

作者:Ming-Hao Hsu, Liang-Hsuan Tseng, Hung-yi Lee, Zhizheng Wu
摘要:我们提出了具有多层聚合的文本对齐语音令牌(TASLA),这是一个文本对齐语音令牌化框架,旨在解决在低帧速率和文本对齐制度下,单源语音令牌在重建过程中可能丢失声学细节的问题。另一方面,本文进一步解释了不同的编码器层如何协作以捕获用于标记化的全面声学特征。以前的工作,TASTE,提出了文本对齐的语音标记化框架,这是一个LM友好的架构,但努力捕捉声学细节。我们通过两个组件来解决这种权衡:多层动态注意力(MLDA),它让每个文本位置自适应地混合来自冻结语音编码器的浅/深特征,以及有限标量量化(FSQ),一种简单的逐维离散化,具有平滑优化。在大约2.62 Hz(tokens/s)时,TASLA持续改善韵律,并在域内(LibriSpeech)和OOD(EXPRESSO,Voxceleb)集上实现了与TASTE相比具有竞争力的质量。我们进一步证明了动态层混合与光谱通量相关,并解释了为什么MLDA在低帧速率下保留韵律与极端特征压缩。
摘要:We propose Text-Aligned Speech Tokens with Multiple Layer-Aggregation (TASLA), which is a text-aligned speech tokenization framework that aims to address the problem that under a low-frame-rate and text-aligned regime, single-source speech tokens may lose acoustic details during reconstruction. On the other hand, this paper further explains how different encoder layers collaborate to capture comprehensive acoustic features for tokenization. Previous work, TASTE, proposed the text-aligned speech tokenization framework, which is a LM-friendly architecture, but struggles to capture acoustic details. We address this trade-off with two components: Multi-Layer Dynamic Attention (MLDA), which lets each text position adaptively mix shallow/deep features from a frozen speech encoder, and Finite Scalar Quantization (FSQ), a simple per-dimension discretization with smooth optimization. At about 2.62 Hz (tokens/s), TASLA consistently improves prosody and achieves competitive quality over TASTE on in-domain (LibriSpeech) and OOD (EXPRESSO, Voxceleb) sets. We further demonstrate that dynamic layer mixing is correlated with spectral flux and explains why MLDA preserves prosody under a low frame rate with extreme feature compression.


【2】If You Hold Me Without Hurting Me: Pathways to Designing Game Audio for Healthy Escapism and Player Well-being
标题:如果你抱着我而不伤害我:为健康逃避主义和玩家福祉设计游戏音频的途径
链接:https://arxiv.org/abs/2510.14691

作者:Caio Nunes, Bosco Borges, Georgia Cruz, Ticianne Darin
备注:5 pages. Presented and discussed at the CHI PLAY 2025 Workshop Exploring Future Directions for Healthy Escapism and Self-Regulation in Games, Pittsburgh, USA, October 13, 2025
摘要:游戏中的逃避主义可以支持恢复或导致有害的回避。自我调节,被理解为将自主权与积极成果相结合,是这一区别的关键。我们认为,音频,往往被忽视,在监管中发挥着核心作用。它可以调节唤醒,标记过渡,并提供关闭,但它对幸福的贡献仍然没有得到充分的探索。本文指出了限制人们对音频潜力认识的方法和可访问性差距,并概述了解决这些问题的方法。我们的目标是鼓励研究人员和开发人员更有意识地将音频整合到更健康的逃避现实游戏的设计和研究中。
摘要:Escapism in games can support recovery or lead to harmful avoidance. Self-regulation, understood as combining autonomy with positive outcomes, is key to this distinction. We argue that audio, often overlooked, plays a central role in regulation. It can modulate arousal, mark transitions, and provide closure, yet its contribution to well-being remains underexplored. This paper identifies methodological and accessibility gaps that limit recognition of audio's potential and outlines ways to address them. We aim to encourage researchers and developers to integrate audio more deliberately into the design and study of healthier escapist play.


【3】SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
标题:LLM作为评委演讲:走向通用且可解释的演讲质量评估
链接:https://arxiv.org/abs/2510.14664

作者:Hui Wang, Jinghua Zhao, Yifan Yang, Shujie Liu, Junyang Chen, Yanzhe Zhang, Shiwan Zhao, Jinyu Li, Jiaming Zhou, Haoqin Sun, Yan Lu, Yong Qin
摘要:生成语音技术正在迅速发展,但评估合成语音的感知质量仍然是一个核心挑战。现有的方法通常依赖于标量分数或二元决策,缺乏跨任务和语言的可解释性和泛化性。我们提出了SpeechLLM-as-Judges,这是一种使大型语言模型(LLM)能够进行结构化和基于简化的语音质量评估的新范式。为了支持这一方向,我们引入了SpeechEval,这是一个包含32,207个多语言语音片段和128,754个注释的大规模数据集,涵盖四个任务:质量评估,成对比较,改进建议和deepfake检测。基于此资源,我们开发了SQ-LLM,这是一种语音质量感知的LLM,经过思维链推理和奖励优化训练,以提高能力。实验结果表明,SQ-LLM在任务和语言上都有很强的性能,揭示了这种范式在推进语音质量评估方面的潜力。相关资源将开放。
摘要:Generative speech technologies are progressing rapidly, but evaluating the perceptual quality of synthetic speech remains a core challenge. Existing methods typically rely on scalar scores or binary decisions, which lack interpretability and generalization across tasks and languages. We present SpeechLLM-as-Judges, a new paradigm for enabling large language models (LLMs) to conduct structured and explanation-based speech quality evaluation. To support this direction, we introduce SpeechEval, a large-scale dataset containing 32,207 multilingual speech clips and 128,754 annotations spanning four tasks: quality assessment, pairwise comparison, improvement suggestion, and deepfake detection. Based on this resource, we develop SQ-LLM, a speech-quality-aware LLM trained with chain-of-thought reasoning and reward optimization to improve capability. Experimental results show that SQ-LLM delivers strong performance across tasks and languages, revealing the potential of this paradigm for advancing speech quality evaluation. Relevant resources will be open-sourced.


【4】AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
标题:AudioEval:文本到音频生成的自动双视角和多维评估
链接:https://arxiv.org/abs/2510.14570

作者:Hui Wang, Jinghua Zhao, Cheng Liu, Yuhang Jia, Haoqin Sun, Jiaming Zhou, Yong Qin
摘要:文本到音频(TTA)正在迅速发展,在虚拟现实,可访问性和创意媒体方面具有广泛的潜力。然而,评估TTA质量仍然很困难:人类的评级是昂贵的和有限的,而现有的客观指标只捕捉感知质量的部分方面。为了解决这一差距,我们引入了AudioEval,这是第一个大规模的TTA评估数据集,包含来自24个系统的4,200个音频样本,在五个感知维度上有126,000个评级,由专家和非专家注释。基于此资源,我们提出了Qwen-DisQA,一个多模态评分模型,联合处理文本提示和生成的音频来预测类人的质量评级。实验结果表明,该方法能有效地提供可靠的、可扩展的评价结果.该数据集将公开提供,以加速未来的研究。
摘要:Text-to-audio (TTA) is rapidly advancing, with broad potential in virtual reality, accessibility, and creative media. However, evaluating TTA quality remains difficult: human ratings are costly and limited, while existing objective metrics capture only partial aspects of perceptual quality. To address this gap, we introduce AudioEval, the first large-scale TTA evaluation dataset, containing 4,200 audio samples from 24 systems with 126,000 ratings across five perceptual dimensions, annotated by both experts and non-experts. Based on this resource, we propose Qwen-DisQA, a multimodal scoring model that jointly processes text prompts and generated audio to predict human-like quality ratings. Experiments show its effectiveness in providing reliable and scalable evaluation. The dataset will be made publicly available to accelerate future research.


【5】Big Data Approaches to Bovine Bioacoustics: A FAIR-Compliant Dataset and Scalable ML Framework for Precision Livestock Welfare
标题:牛生物声学的大数据方法:符合FAIR的数据集和可扩展ML框架,用于精确牲畜福利
链接:https://arxiv.org/abs/2510.14443

作者:Mayuri Kate, Suresh Neethirajan
备注:40 pages, 14 figures, 9 Tables
摘要:物联网传感、边缘计算和机器学习的融合正在改变精准畜牧业。然而,由于计算复杂性和生态有效性的挑战,生物声学数据流仍然没有得到充分利用。我们提出了迄今为止最全面的牛发声数据集之一,其中569个策划的片段涵盖了48个行为类别,使用多个麦克风阵列在三个商业奶牛场记录,并通过域信息增强扩展到2900个样本。这个符合FAIR标准的资源解决了大数据的主要挑战-容量(90小时的录音,65.6 GB),多样性(多场和多区域声学),速度(实时处理)和准确性(噪声鲁棒特征提取)。我们的分布式处理框架集成了使用iZotope RX的高级去噪,通过音频和视频对齐实现的多模式同步,以及使用Praat,librosa和openSMILE生成的24个声学描述符的标准化特征工程。初步的基准揭示了不同的类水平的声学模式,发情检测,遇险分类,和产妇的沟通。数据集的生态现实主义,反映了真实的谷仓声学,而不是控制设置,确保现场部署的准备。这项工作为以动物为中心的人工智能奠定了基础,生物声学数据可以在工业规模上进行连续和非侵入性的福利评估。通过发布标准化的管道和详细的元数据,我们促进了可重复的研究,将大数据分析,可持续农业和精确的牲畜管理联系起来。该框架支持联合国可持续发展目标9,展示了数据科学如何将传统农业转变为智能的、福利优化的系统,以满足全球粮食需求,同时维护合乎道德的动物护理。
摘要:The convergence of IoT sensing, edge computing, and machine learning is transforming precision livestock farming. Yet bioacoustic data streams remain underused because of computational complexity and ecological validity challenges. We present one of the most comprehensive bovine vocalization datasets to date, with 569 curated clips covering 48 behavioral classes, recorded across three commercial dairy farms using multiple microphone arrays and expanded to 2900 samples through domain informed augmentation. This FAIR compliant resource addresses major Big Data challenges - volume (90 hours of recordings, 65.6 GB), variety (multi farm and multi zone acoustics), velocity (real time processing), and veracity (noise robust feature extraction). Our distributed processing framework integrates advanced denoising using iZotope RX, multimodal synchronization through audio and video alignment, and standardized feature engineering with 24 acoustic descriptors generated from Praat, librosa, and openSMILE. Preliminary benchmarks reveal distinct class level acoustic patterns for estrus detection, distress classification, and maternal communication. The datasets ecological realism, reflecting authentic barn acoustics rather than controlled settings, ensures readiness for field deployment. This work establishes a foundation for animal centered AI, where bioacoustic data enable continuous and non invasive welfare assessment at industrial scale. By releasing standardized pipelines and detailed metadata, we promote reproducible research that connects Big Data analytics, sustainable agriculture, and precision livestock management. The framework supports UN SDG 9, showing how data science can turn traditional farming into intelligent, welfare optimized systems that meet global food needs while upholding ethical animal care.


【6】Revisit Modality Imbalance at the Decision Layer
标题:重新审视决策层的模式失衡
链接:https://arxiv.org/abs/2510.14411

作者:Xiaoyu Ma, Hao Chen
备注:Some Insights in Balanced Multimodal Learning
摘要:多模态学习集成了来自不同模态的信息以提高模型性能,但它经常遭受模态不平衡,在联合优化过程中,主导模态会掩盖较弱的模态。本文揭示了这种不平衡不仅发生在表示学习过程中,而且在决策层也表现得很明显。在视听数据集(CREMAD和Kinetic-Sounds)上的实验表明,即使经过广泛的预训练和平衡优化,模型仍然表现出对某些模态(如音频)的系统性偏见。进一步的分析表明,这种偏见源于特征空间和决策权重分布的内在差异,而不仅仅是优化动态。我们认为,在融合阶段聚合未校准的模态输出会导致有偏见的决策层加权,阻碍较弱的模态有效地作出贡献。为了解决这个问题,我们建议,未来的多模态系统应该更多地关注在决策层纳入自适应权重分配机制,使相对平衡,根据每个模态的能力。
摘要:Multimodal learning integrates information from different modalities to enhance model performance, yet it often suffers from modality imbalance, where dominant modalities overshadow weaker ones during joint optimization. This paper reveals that such an imbalance not only occurs during representation learning but also manifests significantly at the decision layer. Experiments on audio-visual datasets (CREMAD and Kinetic-Sounds) show that even after extensive pretraining and balanced optimization, models still exhibit systematic bias toward certain modalities, such as audio. Further analysis demonstrates that this bias originates from intrinsic disparities in feature-space and decision-weight distributions rather than from optimization dynamics alone. We argue that aggregating uncalibrated modality outputs at the fusion stage leads to biased decision-layer weighting, hindering weaker modalities from contributing effectively. To address this, we propose that future multimodal systems should focus more on incorporate adaptive weight allocation mechanisms at the decision layer, enabling relative balanced according to the capabilities of each modality.


【7】Beat Detection as Object Detection
标题:节拍检测作为对象检测
链接:https://arxiv.org/abs/2510.14391

作者:Jaehoon Ahn, Moon-Ryul Jung
备注:11 pages, 4 figures, 5 tables
摘要:最近的节拍和下拍跟踪模型(例如,RNN、TCN、Transformers)输出帧级激活。我们建议将此任务重新定义为对象检测,其中节拍和强拍被建模为时间“对象”。“将FCOS检测器从计算机视觉调整到1D音频,我们用WaveBeat的时间特征提取器取代了其原始主干,并添加了一个特征金字塔网络来捕获多尺度时间模式。该模型预测具有置信度分数的重叠搏动/下拍间隔,然后是非最大抑制(NMS)以选择最终预测。这个NMS步骤与传统跟踪器中的DBN类似,但更简单,启发性更低。在标准音乐数据集上进行评估,我们的方法取得了有竞争力的结果,表明对象检测技术可以有效地以最小的适应性来模拟音乐节拍。
摘要:Recent beat and downbeat tracking models (e.g., RNNs, TCNs, Transformers) output frame-level activations. We propose reframing this task as object detection, where beats and downbeats are modeled as temporal "objects." Adapting the FCOS detector from computer vision to 1D audio, we replace its original backbone with WaveBeat's temporal feature extractor and add a Feature Pyramid Network to capture multi-scale temporal patterns. The model predicts overlapping beat/downbeat intervals with confidence scores, followed by non-maximum suppression (NMS) to select final predictions. This NMS step serves a similar role to DBNs in traditional trackers, but is simpler and less heuristic. Evaluated on standard music datasets, our approach achieves competitive results, showing that object detection techniques can effectively model musical beats with minimal adaptation.


【8】Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?
标题:联合语音-音频嵌入会编码感知音色语义吗?
链接:https://arxiv.org/abs/2510.14249

作者:Qixin Deng, Bryan Pardo, Thrasyvoulos N Pappas
摘要:理解和建模语言和声音之间的关系对于音乐信息检索、文本引导音乐生成和音频字幕等应用至关重要。这些任务的核心是使用联合语言音频嵌入空间,将文本描述和听觉内容映射到共享的嵌入空间中。虽然多模态嵌入模型(如MS-CLAP,LAION-CLAP和MuQ-MuLan)在对齐语言和音频方面表现出了很强的性能,但它们与人类对音色的感知的对应关系,这是一个多方面的属性,包括亮度,粗糙度和温暖度等质量,仍然没有得到充分的探索。在本文中,我们评估了上述三个联合语言音频嵌入模型的能力,以捕捉感知维度的音色。我们的研究结果表明,LAION-CLAP始终提供了最可靠的对齐与人类感知的音色语义在乐器的声音和音频效果。
摘要:Understanding and modeling the relationship between language and sound is critical for applications such as music information retrieval,text-guided music generation, and audio captioning. Central to these tasks is the use of joint language-audio embedding spaces, which map textual descriptions and auditory content into a shared embedding space. While multimodal embedding models such as MS-CLAP, LAION-CLAP, and MuQ-MuLan have shown strong performance in aligning language and audio, their correspondence to human perception of timbre, a multifaceted attribute encompassing qualities such as brightness, roughness, and warmth, remains underexplored. In this paper, we evaluate the above three joint language-audio embedding models on their ability to capture perceptual dimensions of timbre. Our findings show that LAION-CLAP consistently provides the most reliable alignment with human-perceived timbre semantics across both instrumental sounds and audio effects.


【9】Sound Masking Strategies for Interference with Mosquito Hearing
标题:干扰蚊子听力的声音掩蔽策略
链接:https://arxiv.org/abs/2510.14921

作者:Justin Faber, Alexandros C Alampounti, Marcos Georgiades, Joerg T Albert, Dolores Bozovic
摘要:听觉掩蔽的使用长期以来一直是心理声学和工程目的的兴趣,以掩盖对人类或栖息地与我们重叠的物种具有破坏性的声音。在大多数情况下,我们力求尽量减少对野生动物交流的干扰。然而,在携带病原体的昆虫的情况下,我们可能希望将这些干扰最大化,作为控制种群的一种方式。在目前的工作中,我们探索候选人掩蔽策略的通用模型的主动听觉系统和蚊子听觉系统的模型。对于这两种模型,我们发现所有声功率都集中在一个或几个频率上的口罩性能最好。我们建议,快速频率调制的基础上的面具是最有效的最大中断的信息传输和最小的可懂度。我们希望这些结果将有助于指导避免或选择可能的声学信号,分别最大限度地提高或最小化通信。
摘要:The use of auditory masking has long been of interest in psychoacoustics and for engineering purposes, in order to cover sounds that are disruptive to humans or to species whose habitats overlap with ours. In most cases, we seek to minimize the disturbances to the communication of wildlife. However, in the case of pathogen-carrying insects, we may want to maximize these disturbances as a way to control populations. In the current work, we explore candidate masking strategies for a generic model of active auditory systems and a model of the mosquito auditory system. For both models, we find that masks with all acoustic power focused into just one or a few frequencies perform best. We propose that masks based on rapid frequency modulation are most effective for maximal disruption of information transfer and minimizing intelligibility. We hope that these results will serve to guide the avoidance or selection of possible acoustic signals for, respectively, maximizing or minimizing communication.


【10】Musical consonance: a review of theory and evidence on perception and preference of auditory roughness in humans and other animals
标题:音乐和谐:人类和其他动物对听觉粗糙度的感知和偏好的理论和证据回顾
链接:https://arxiv.org/abs/2510.14159

作者:John M. McBride
摘要:人类音乐中和谐音的起源长期以来一直存在争议,今天有三个主要假设:厌恶粗糙,偏好和谐,以及从文化接触中习得的偏好。虽然目前的证据不足以解开这些假设的贡献,我提出了几个原因,粗糙度是一个特别有前途的领域,为未来的研究。本文的目的是总结和批判性地评价粗糙度理论和模型,实验数据,突出值得进一步研究的领域。我确定了两个关键领域:由于粗糙度定义中的同义反复以及经验测量中缺乏独立性,结果的定义和解释存在根本性问题。尽管广泛的模型开发,有许多重复和模型有问题的数据质量和过度拟合。未来的理论发展应该以模型的简单性为目标,并对额外的假设、特征和参数进行系统的评估。模型评估的目标应该是最大限度地扩大预测的刺激范围。
摘要:The origins of consonance in human music has long been contested, and today there are three primary hypotheses: aversion to roughness, preference for harmonicity, and learned preferences from cultural exposure. While the evidence is currently insufficient to disentangle the contributions of these hypotheses, I propose several reasons why roughness is an especially promising area for future study. The aim of this review is to summarize and critically evaluate roughness theory and models, experimental data, to highlight areas that deserve further research. I identify 2 key areas: There are fundamental issues with the definition and interpretation of results due to tautology in the definition of roughness, and the lack of independence in empirical measurements. Despite extensive model development, there are many duplications and models have issues with data quality and overfitting. Future theory development should aim for model simplicity, and extra assumptions, features and parameters should be evaluated systematically. Model evaluation should aim to maximise the breadth of stimuli that are predicted.


【11】Switchboard-Affect: Emotion Perception Labels from Conversational Speech
标题:总机情感:来自对话演讲的情感感知标签
链接:https://arxiv.org/abs/2510.13906

作者:Amrit Romana, Jaya Narain, Tien Dung Tran, Andrea Davis, Jason Fong, Ramya Rasipuram, Vikramjit Mitra
备注:2025 13th International Conference on Affective Computing and Intelligent Interaction (ACII) this https URL
摘要:理解语音情感数据集管理和标记的细微差别对于评估语音情感识别(SER)模型在现实世界应用中的潜力至关重要。大多数训练和评估数据集包含表演或伪表演语音(例如,播客语音),其中情感表达可能被夸大或以其他方式被有意修改。此外,基于人群感知标记的数据集通常缺乏关于注释者的指导方针的透明度。这些因素使得很难理解模型的性能和确定需要改进的地方。为了解决这个问题,我们将Switchboard语料库确定为自然主义对话语音的一个有前途的来源,我们训练了一群人来标记分类情绪(愤怒,蔑视,厌恶,恐惧,悲伤,惊讶,幸福,温柔,平静和中性)和维度属性(激活,效价和优势)的数据集。我们将此标签集称为开关板影响(SWB-Affect)。在这项工作中,我们详细介绍了我们的方法,包括提供给注释者的定义和可能在他们的感知中发挥作用的词汇和语言学线索的分析。此外,我们评估了最先进的SER模型,我们发现在情绪类别中的可变性能,尤其是愤怒的泛化能力较差。这些发现强调了使用捕捉语音中自然情感变化的数据集进行评估的重要性。我们发布了SWB-Affect的标签,以便在该领域进行进一步分析。
摘要:Understanding the nuances of speech emotion dataset curation and labeling is essential for assessing speech emotion recognition (SER) model potential in real-world applications. Most training and evaluation datasets contain acted or pseudo-acted speech (e.g., podcast speech) in which emotion expressions may be exaggerated or otherwise intentionally modified. Furthermore, datasets labeled based on crowd perception often lack transparency regarding the guidelines given to annotators. These factors make it difficult to understand model performance and pinpoint necessary areas for improvement. To address this gap, we identified the Switchboard corpus as a promising source of naturalistic conversational speech, and we trained a crowd to label the dataset for categorical emotions (anger, contempt, disgust, fear, sadness, surprise, happiness, tenderness, calmness, and neutral) and dimensional attributes (activation, valence, and dominance). We refer to this label set as Switchboard-Affect (SWB-Affect). In this work, we present our approach in detail, including the definitions provided to annotators and an analysis of the lexical and paralinguistic cues that may have played a role in their perception. In addition, we evaluate state-of-the-art SER models, and we find variable performance across the emotion categories with especially poor generalization for anger. These findings underscore the importance of evaluation with datasets that capture natural affective variations in speech. We release the labels for SWB-Affect to enable further analysis in this domain.


eess.AS音频处理


【1】Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization
标题:用于多通道神经说话人拨号的空间感知自监督模型
链接:https://arxiv.org/abs/2510.14551

作者:Jiangyu Han, Ruoyu Wang, Yoshiki Masuyama, Marc Delcroix, Johan Rohdin, Jun Du, Lukas Burget
备注:Submitted to ICASSP 2026
摘要:WavLM等自监督模型在神经说话人日记化方面表现出了强大的性能。然而,这些模型通常在单通道记录上进行预训练,限制了它们在多通道场景中的有效性。基于这些模型构建的现有日记系统通常依赖DOVER-Lap来组合来自各个通道的输出。虽然有效,这种方法会产生大量的计算开销,并未能充分利用空间信息。在这项工作中,基于DiariZen,一个将基于WavLM的本地端到端神经日志化与说话人嵌入聚类相结合的管道,我们引入了一种轻量级的方法,通过将信道通信模块插入早期层来使预训练的WavLM具有空间感知能力。我们的方法是不可知的麦克风通道和阵列拓扑结构的数量,确保广泛的适用性。我们还提出了利用空间注意力权重融合多通道说话人嵌入。对五个公共数据集的评估显示,与单通道基线相比,性能和效率都有了持续的提高,与DOVER-Lap相比,性能和效率都有了显著提高。我们的源代码可在https://github.com/BUTSpeechFIT/DiariZen上公开获取。
摘要:Self-supervised models such as WavLM have demonstrated strong performance for neural speaker diarization. However, these models are typically pre-trained on single-channel recordings, limiting their effectiveness in multi-channel scenarios. Existing diarization systems built on these models often rely on DOVER-Lap to combine outputs from individual channels. Although effective, this approach incurs substantial computational overhead and fails to fully exploit spatial information. In this work, building on DiariZen, a pipeline that combines WavLM-based local endto-end neural diarization with speaker embedding clustering, we introduce a lightweight approach to make pre-trained WavLM spatially aware by inserting channel communication modules into the early layers. Our method is agnostic to both the number of microphone channels and array topologies, ensuring broad applicability. We further propose to fuse multi-channel speaker embeddings by leveraging spatial attention weights. Evaluations on five public datasets show consistent improvements over single-channel baselines and demonstrate superior performance and efficiency compared with DOVER-Lap. Our source code is publicly available at https://github.com/BUTSpeechFIT/DiariZen.


【2】Switchboard-Affect: Emotion Perception Labels from Conversational Speech
标题:总机情感:来自对话演讲的情感感知标签
链接:https://arxiv.org/abs/2510.13906

作者:Amrit Romana, Jaya Narain, Tien Dung Tran, Andrea Davis, Jason Fong, Ramya Rasipuram, Vikramjit Mitra
备注:2025 13th International Conference on Affective Computing and Intelligent Interaction (ACII) this https URL
摘要:理解语音情感数据集管理和标记的细微差别对于评估语音情感识别(SER)模型在现实世界应用中的潜力至关重要。大多数训练和评估数据集包含表演或伪表演语音(例如,播客语音),其中情感表达可能被夸大或以其他方式被有意修改。此外,基于人群感知标记的数据集通常缺乏关于注释者的指导方针的透明度。这些因素使得很难理解模型的性能和确定需要改进的地方。为了解决这个问题,我们将Switchboard语料库确定为自然主义对话语音的一个有前途的来源,我们训练了一群人来标记分类情绪(愤怒,蔑视,厌恶,恐惧,悲伤,惊讶,幸福,温柔,平静和中性)和维度属性(激活,效价和优势)的数据集。我们将此标签集称为开关板影响(SWB-Affect)。在这项工作中,我们详细介绍了我们的方法,包括提供给注释者的定义和可能在他们的感知中发挥作用的词汇和语言学线索的分析。此外,我们评估了最先进的SER模型,我们发现在情绪类别中的可变性能,尤其是愤怒的泛化能力较差。这些发现强调了使用捕捉语音中自然情感变化的数据集进行评估的重要性。我们发布了SWB-Affect的标签,以便在该领域进行进一步分析。
摘要:Understanding the nuances of speech emotion dataset curation and labeling is essential for assessing speech emotion recognition (SER) model potential in real-world applications. Most training and evaluation datasets contain acted or pseudo-acted speech (e.g., podcast speech) in which emotion expressions may be exaggerated or otherwise intentionally modified. Furthermore, datasets labeled based on crowd perception often lack transparency regarding the guidelines given to annotators. These factors make it difficult to understand model performance and pinpoint necessary areas for improvement. To address this gap, we identified the Switchboard corpus as a promising source of naturalistic conversational speech, and we trained a crowd to label the dataset for categorical emotions (anger, contempt, disgust, fear, sadness, surprise, happiness, tenderness, calmness, and neutral) and dimensional attributes (activation, valence, and dominance). We refer to this label set as Switchboard-Affect (SWB-Affect). In this work, we present our approach in detail, including the definitions provided to annotators and an analysis of the lexical and paralinguistic cues that may have played a role in their perception. In addition, we evaluate state-of-the-art SER models, and we find variable performance across the emotion categories with especially poor generalization for anger. These findings underscore the importance of evaluation with datasets that capture natural affective variations in speech. We release the labels for SWB-Affect to enable further analysis in this domain.


【3】TRI-DEP: A Trimodal Comparative Study for Depression Detection Using Speech, Text, and EEG
标题:TRI-DPP:使用语音、文本和脑电检测抑郁症的三峰比较研究
链接:https://arxiv.org/abs/2510.14922

作者:Annisaa Fitri Nurfidausi, Eleonora Mancini, Paolo Torroni
摘要:抑郁症是一种普遍的心理健康障碍,但其自动检测仍然具有挑战性。先前的工作已经探索了单峰和多峰方法,多峰系统通过利用互补信号显示出希望。然而,现有的研究范围有限,缺乏系统的比较功能,并遭受不一致的评价协议。我们通过系统地探索EEG以及语音和文本的特征表示和建模策略来解决这些差距。我们评估手工制作的功能与预训练的嵌入,评估不同的神经编码器的有效性,比较单峰,双峰和三峰配置,并分析融合策略,注意EEG的作用。应用一致的独立于受试者的分割,以确保稳健、可重现的基准。我们的研究结果表明,(i)EEG,语音和文本模态的组合增强了多模态检测,(ii)预训练的嵌入优于手工特征,(iii)精心设计的三模态模型实现了最先进的性能。我们的工作奠定了基础,为未来的研究多模态抑郁症检测。
摘要:Depression is a widespread mental health disorder, yet its automatic detection remains challenging. Prior work has explored unimodal and multimodal approaches, with multimodal systems showing promise by leveraging complementary signals. However, existing studies are limited in scope, lack systematic comparisons of features, and suffer from inconsistent evaluation protocols. We address these gaps by systematically exploring feature representations and modelling strategies across EEG, together with speech and text. We evaluate handcrafted features versus pre-trained embeddings, assess the effectiveness of different neural encoders, compare unimodal, bimodal, and trimodal configurations, and analyse fusion strategies with attention to the role of EEG. Consistent subject-independent splits are applied to ensure robust, reproducible benchmarking. Our results show that (i) the combination of EEG, speech and text modalities enhances multimodal detection, (ii) pretrained embeddings outperform handcrafted features, and (iii) carefully designed trimodal models achieve state-of-the-art performance. Our work lays the groundwork for future research in multimodal depression detection.


【4】SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
标题:LLM作为评委演讲:走向通用且可解释的演讲质量评估
链接:https://arxiv.org/abs/2510.14664

作者:Hui Wang, Jinghua Zhao, Yifan Yang, Shujie Liu, Junyang Chen, Yanzhe Zhang, Shiwan Zhao, Jinyu Li, Jiaming Zhou, Haoqin Sun, Yan Lu, Yong Qin
摘要:生成语音技术正在迅速发展,但评估合成语音的感知质量仍然是一个核心挑战。现有的方法通常依赖于标量分数或二元决策,缺乏跨任务和语言的可解释性和泛化性。我们提出了SpeechLLM-as-Judges,这是一种使大型语言模型(LLM)能够进行结构化和基于简化的语音质量评估的新范式。为了支持这一方向,我们引入了SpeechEval,这是一个包含32,207个多语言语音片段和128,754个注释的大规模数据集,涵盖四个任务:质量评估,成对比较,改进建议和deepfake检测。基于此资源,我们开发了SQ-LLM,这是一种语音质量感知的LLM,经过思维链推理和奖励优化训练,以提高能力。实验结果表明,SQ-LLM在任务和语言上都有很强的性能,揭示了这种范式在推进语音质量评估方面的潜力。相关资源将开放。
摘要:Generative speech technologies are progressing rapidly, but evaluating the perceptual quality of synthetic speech remains a core challenge. Existing methods typically rely on scalar scores or binary decisions, which lack interpretability and generalization across tasks and languages. We present SpeechLLM-as-Judges, a new paradigm for enabling large language models (LLMs) to conduct structured and explanation-based speech quality evaluation. To support this direction, we introduce SpeechEval, a large-scale dataset containing 32,207 multilingual speech clips and 128,754 annotations spanning four tasks: quality assessment, pairwise comparison, improvement suggestion, and deepfake detection. Based on this resource, we develop SQ-LLM, a speech-quality-aware LLM trained with chain-of-thought reasoning and reward optimization to improve capability. Experimental results show that SQ-LLM delivers strong performance across tasks and languages, revealing the potential of this paradigm for advancing speech quality evaluation. Relevant resources will be open-sourced.


【5】AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
标题:AudioEval:文本到音频生成的自动双视角和多维评估
链接:https://arxiv.org/abs/2510.14570

作者:Hui Wang, Jinghua Zhao, Cheng Liu, Yuhang Jia, Haoqin Sun, Jiaming Zhou, Yong Qin
摘要:文本到音频(TTA)正在迅速发展,在虚拟现实,可访问性和创意媒体方面具有广泛的潜力。然而,评估TTA质量仍然很困难:人类的评级是昂贵的和有限的,而现有的客观指标只捕捉感知质量的部分方面。为了解决这一差距,我们引入了AudioEval,这是第一个大规模的TTA评估数据集,包含来自24个系统的4,200个音频样本,在五个感知维度上有126,000个评级,由专家和非专家注释。基于此资源,我们提出了Qwen-DisQA,一个多模态评分模型,联合处理文本提示和生成的音频来预测类人的质量评级。实验结果表明,该方法能有效地提供可靠的、可扩展的评价结果.该数据集将公开提供,以加速未来的研究。
摘要:Text-to-audio (TTA) is rapidly advancing, with broad potential in virtual reality, accessibility, and creative media. However, evaluating TTA quality remains difficult: human ratings are costly and limited, while existing objective metrics capture only partial aspects of perceptual quality. To address this gap, we introduce AudioEval, the first large-scale TTA evaluation dataset, containing 4,200 audio samples from 24 systems with 126,000 ratings across five perceptual dimensions, annotated by both experts and non-experts. Based on this resource, we propose Qwen-DisQA, a multimodal scoring model that jointly processes text prompts and generated audio to predict human-like quality ratings. Experiments show its effectiveness in providing reliable and scalable evaluation. The dataset will be made publicly available to accelerate future research.


【6】Big Data Approaches to Bovine Bioacoustics: A FAIR-Compliant Dataset and Scalable ML Framework for Precision Livestock Welfare
标题:牛生物声学的大数据方法:符合FAIR的数据集和可扩展ML框架,用于精确牲畜福利
链接:https://arxiv.org/abs/2510.14443

作者:Mayuri Kate, Suresh Neethirajan
备注:40 pages, 14 figures, 9 Tables
摘要:物联网传感、边缘计算和机器学习的融合正在改变精准畜牧业。然而,由于计算复杂性和生态有效性的挑战,生物声学数据流仍然没有得到充分利用。我们提出了迄今为止最全面的牛发声数据集之一,其中569个策划的片段涵盖了48个行为类别,使用多个麦克风阵列在三个商业奶牛场记录,并通过域信息增强扩展到2900个样本。这个符合FAIR标准的资源解决了大数据的主要挑战-容量(90小时的录音,65.6 GB),多样性(多场和多区域声学),速度(实时处理)和准确性(噪声鲁棒特征提取)。我们的分布式处理框架集成了使用iZotope RX的高级去噪,通过音频和视频对齐实现的多模式同步,以及使用Praat,librosa和openSMILE生成的24个声学描述符的标准化特征工程。初步的基准揭示了不同的类水平的声学模式,发情检测,遇险分类,和产妇的沟通。数据集的生态现实主义反映了真实的谷仓声学而不是受控的设置,确保了现场部署的准备就绪。这项工作为以动物为中心的人工智能奠定了基础,生物声学数据可以在工业规模上进行连续和非侵入性的福利评估。通过发布标准化的管道和详细的元数据,我们促进了可重复的研究,将大数据分析,可持续农业和精确的牲畜管理联系起来。该框架支持联合国可持续发展目标9,展示了数据科学如何将传统农业转变为智能的、福利优化的系统,以满足全球粮食需求,同时维护合乎道德的动物护理。
摘要:The convergence of IoT sensing, edge computing, and machine learning is transforming precision livestock farming. Yet bioacoustic data streams remain underused because of computational complexity and ecological validity challenges. We present one of the most comprehensive bovine vocalization datasets to date, with 569 curated clips covering 48 behavioral classes, recorded across three commercial dairy farms using multiple microphone arrays and expanded to 2900 samples through domain informed augmentation. This FAIR compliant resource addresses major Big Data challenges - volume (90 hours of recordings, 65.6 GB), variety (multi farm and multi zone acoustics), velocity (real time processing), and veracity (noise robust feature extraction). Our distributed processing framework integrates advanced denoising using iZotope RX, multimodal synchronization through audio and video alignment, and standardized feature engineering with 24 acoustic descriptors generated from Praat, librosa, and openSMILE. Preliminary benchmarks reveal distinct class level acoustic patterns for estrus detection, distress classification, and maternal communication. The datasets ecological realism, reflecting authentic barn acoustics rather than controlled settings, ensures readiness for field deployment. This work establishes a foundation for animal centered AI, where bioacoustic data enable continuous and non invasive welfare assessment at industrial scale. By releasing standardized pipelines and detailed metadata, we promote reproducible research that connects Big Data analytics, sustainable agriculture, and precision livestock management. The framework supports UN SDG 9, showing how data science can turn traditional farming into intelligent, welfare optimized systems that meet global food needs while upholding ethical animal care.


【7】Revisit Modality Imbalance at the Decision Layer
标题:重新审视决策层的模式失衡
链接:https://arxiv.org/abs/2510.14411

作者:Xiaoyu Ma, Hao Chen
备注:Some Insights in Balanced Multimodal Learning
摘要:多模态学习集成了来自不同模态的信息以提高模型性能,但它经常遭受模态不平衡,在联合优化过程中,主导模态会掩盖较弱的模态。本文揭示了这种不平衡不仅发生在表示学习过程中,而且在决策层也表现得很明显。在视听数据集(CREMAD和Kinetic-Sounds)上的实验表明,即使经过广泛的预训练和平衡优化,模型仍然表现出对某些模态(如音频)的系统性偏见。进一步的分析表明,这种偏见源于特征空间和决策权重分布的内在差异,而不仅仅是优化动态。我们认为,在融合阶段聚合未校准的模态输出会导致有偏见的决策层加权,阻碍较弱的模态有效地作出贡献。为了解决这个问题,我们建议,未来的多模态系统应该更多地关注在决策层纳入自适应权重分配机制,使相对平衡,根据每个模态的能力。
摘要:Multimodal learning integrates information from different modalities to enhance model performance, yet it often suffers from modality imbalance, where dominant modalities overshadow weaker ones during joint optimization. This paper reveals that such an imbalance not only occurs during representation learning but also manifests significantly at the decision layer. Experiments on audio-visual datasets (CREMAD and Kinetic-Sounds) show that even after extensive pretraining and balanced optimization, models still exhibit systematic bias toward certain modalities, such as audio. Further analysis demonstrates that this bias originates from intrinsic disparities in feature-space and decision-weight distributions rather than from optimization dynamics alone. We argue that aggregating uncalibrated modality outputs at the fusion stage leads to biased decision-layer weighting, hindering weaker modalities from contributing effectively. To address this, we propose that future multimodal systems should focus more on incorporate adaptive weight allocation mechanisms at the decision layer, enabling relative balanced according to the capabilities of each modality.


【8】A Robust Classification Method using Hybrid Word Embedding for Early Diagnosis of Alzheimer's Disease
标题:使用混合词嵌入进行阿尔茨海默病早期诊断的稳健分类方法
链接:https://arxiv.org/abs/2510.14332

作者:Yangyang Li
备注:Peer-reviewed and published in Proceedings of the 2020 3rd   International Conference on Algorithms, Computing and Artificial Intelligence   (ACAI 2020). 7 pages, 5 figures
摘要:阿尔茨海默病(AD)的早期发现对AD患者非常有益,从而导致早期治疗,减轻症状并减轻医疗保健的经济负担。语言能力的改变是AD的主要表现之一,可用于AD的早期诊断。在本文中,我开发了一种强大的分类方法,使用混合词嵌入和微调超参数,以实现最先进的准确性,在早期检测AD。具体来说,我们创建一个混合词嵌入的基础上,从Doc 2 Vec和ELMo的词向量,以获得困惑分数的句子。分数识别句子是否流利,并捕获句子的语义上下文。我通过添加语言特征来分析句法和语义来丰富词嵌入。此外,我们将嵌入的特征向量输入逻辑回归,并在整个管道中微调超参数。通过调整机器学习流水线的超参数(例如,模型正则化参数,学习率和Doc 2 Vec的向量大小,以及ELMo的向量大小),我在区分早期AD和健康受试者方面实现了91%的分类准确率和97%的曲线下面积(AUC)。根据我的知识,我的模型具有91%的准确性和97%的AUC,优于现有最好的NLP模型,其准确性为88% [32]。我通过反复实验研究了模型的稳定性,发现即使训练数据被随机分割,模型也是稳定的(准确度的标准差= 0.0403; AUC的标准差= 0.0174)。这证实了我们提出的方法是准确和稳定的。该模型可作为AD的大规模筛查方法,也可作为医生检测AD的补充检查。
摘要:Early detection of Alzheimer's Disease (AD) is greatly beneficial to AD patients, leading to early treatments that lessen symptoms and alleviating financial burden of health care. As one of the leading signs of AD, language capability changes can be used for early diagnosis of AD. In this paper, I develop a robust classification method using hybrid word embedding and fine-tuned hyperparameters to achieve state-of-the-art accuracy in the early detection of AD. Specifically, we create a hybrid word embedding based on word vectors from Doc2Vec and ELMo to obtain perplexity scores of the sentences. The scores identify whether a sentence is fluent or not and capture semantic context of the sentences. I enrich the word embedding by adding linguistic features to analyze syntax and semantics. Further, we input an embedded feature vector into logistic regression and fine tune hyperparameters throughout the pipeline. By tuning hyperparameters of the machine learning pipeline (e.g., model regularization parameter, learning rate and vector size of Doc2Vec, and vector size of ELMo), I achieve 91% classification accuracy and an Area Under the Curve (AUC) of 97% in distinguishing early AD from healthy subjects. Based on my knowledge, my model with 91% accuracy and 97% AUC outperforms the best existing NLP model for AD diagnosis with an accuracy of 88% [32]. I study the model stability through repeated experiments and find that the model is stable even though the training data is split randomly (standard deviation of accuracy = 0.0403; standard deviation of AUC = 0.0174). This affirms our proposed method is accurate and stable. This model can be used as a large-scale screening method for AD, as well as a complementary examination for doctors to detect AD.


【9】Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?
标题:联合语音-音频嵌入会编码感知音色语义吗?
链接:https://arxiv.org/abs/2510.14249

作者:Qixin Deng, Bryan Pardo, Thrasyvoulos N Pappas
摘要:理解和建模语言和声音之间的关系对于音乐信息检索、文本引导音乐生成和音频字幕等应用至关重要。这些任务的核心是使用联合语言音频嵌入空间,将文本描述和听觉内容映射到共享的嵌入空间中。虽然多模态嵌入模型(如MS-CLAP,LAION-CLAP和MuQ-MuLan)在对齐语言和音频方面表现出了很强的性能,但它们与人类对音色的感知的对应关系,这是一个多方面的属性,包括亮度,粗糙度和温暖度等质量,仍然没有得到充分的探索。在本文中,我们评估了上述三个联合语言音频嵌入模型的能力,以捕捉感知维度的音色。我们的研究结果表明,LAION-CLAP始终提供了最可靠的对齐与人类感知的音色语义在乐器的声音和音频效果。
摘要:Understanding and modeling the relationship between language and sound is critical for applications such as music information retrieval,text-guided music generation, and audio captioning. Central to these tasks is the use of joint language-audio embedding spaces, which map textual descriptions and auditory content into a shared embedding space. While multimodal embedding models such as MS-CLAP, LAION-CLAP, and MuQ-MuLan have shown strong performance in aligning language and audio, their correspondence to human perception of timbre, a multifaceted attribute encompassing qualities such as brightness, roughness, and warmth, remains underexplored. In this paper, we evaluate the above three joint language-audio embedding models on their ability to capture perceptual dimensions of timbre. Our findings show that LAION-CLAP consistently provides the most reliable alignment with human-perceived timbre semantics across both instrumental sounds and audio effects.


【10】Musical consonance: a review of theory and evidence on perception and preference of auditory roughness in humans and other animals
标题:音乐和谐:人类和其他动物对听觉粗糙度的感知和偏好的理论和证据回顾
链接:https://arxiv.org/abs/2510.14159

作者:John M. McBride
摘要:人类音乐中和谐音的起源长期以来一直存在争议,今天有三个主要假设:厌恶粗糙,偏好和谐,以及从文化接触中习得的偏好。虽然目前的证据不足以解开这些假设的贡献,我提出了几个原因,粗糙度是一个特别有前途的领域,为未来的研究。本文的目的是总结和批判性地评价粗糙度理论和模型,实验数据,突出值得进一步研究的领域。我确定了两个关键领域:由于粗糙度定义中的同义反复以及经验测量中缺乏独立性,结果的定义和解释存在根本性问题。尽管广泛的模型开发,有许多重复和模型有问题的数据质量和过度拟合。未来的理论发展应该以模型的简单性为目标,并对额外的假设、特征和参数进行系统的评估。模型评估的目标应该是最大限度地扩大预测的刺激范围。
摘要:The origins of consonance in human music has long been contested, and today there are three primary hypotheses: aversion to roughness, preference for harmonicity, and learned preferences from cultural exposure. While the evidence is currently insufficient to disentangle the contributions of these hypotheses, I propose several reasons why roughness is an especially promising area for future study. The aim of this review is to summarize and critically evaluate roughness theory and models, experimental data, to highlight areas that deserve further research. I identify 2 key areas: There are fundamental issues with the definition and interpretation of results due to tautology in the definition of roughness, and the lack of independence in empirical measurements. Despite extensive model development, there are many duplications and models have issues with data quality and overfitting. Future theory development should aim for model simplicity, and extra assumptions, features and parameters should be evaluated systematically. Model evaluation should aim to maximise the breadth of stimuli that are predicted.


机器翻译由腾讯交互翻译提供,仅供参考