今日论文合集:cs.SD语音13篇,eess.AS音频处理15篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】SpeechIQ: Speech Intelligence Quotient Across Cognitive Levels in Voice Understanding Large Language Models
标题:SpeechIQ:语音理解大型语言模型中跨认知水平的言语智商
链接:https://arxiv.org/abs/2507.19361

作者: Zhen Wan, Chao-Han Huck Yang, Yahan Yu, Jinchuan Tian, Sheng Li, Ke Hu, Zhehuai Chen, Shinji Watanabe, Fei Cheng, Chenhui Chu, Sadao Kurohashi
备注:Our Speech-IQ leaderboard will be hosted at this http URL. ACL 2025 main
摘要:我们引入了基于语音的智商(SIQ)作为一种新形式的人类认知启发的评估管道,用于语音理解大型语言模型LLM Voice,旨在评估其语音理解能力。超越流行的语音理解指标,如单词错误率(WER),SIQ检查LLM语音跨越三个认知水平由布卢姆的分类:(1)记住(即,(2)理解(即,法学硕士解释的相似性);和(3)应用(即,模拟下游任务的QA精度)。我们证明,SIQ不仅量化了语音理解能力,而且还提供了级联方法之间的统一比较(例如,ASR LLM)和端到端模型,识别现有基准测试中的注释错误,并检测LLM语音中的幻觉。我们的框架代表了第一个同类智力考试,将认知原则与语音基准相结合,同时暴露了多模式培训中被忽视的挑战。
摘要:We introduce Speech-based Intelligence Quotient (SIQ) as a new form of human cognition-inspired evaluation pipeline for voice understanding large language models, LLM Voice, designed to assess their voice understanding ability. Moving beyond popular voice understanding metrics such as word error rate (WER), SIQ examines LLM Voice across three cognitive levels motivated by Bloom's Taxonomy: (1) Remembering (i.e., WER for verbatim accuracy); (2) Understanding (i.e., similarity of LLM's interpretations); and (3) Application (i.e., QA accuracy for simulating downstream tasks). We demonstrate that SIQ not only quantifies voice understanding abilities but also provides unified comparisons between cascaded methods (e.g., ASR LLM) and end-to-end models, identifies annotation errors in existing benchmarks, and detects hallucinations in LLM Voice. Our framework represents a first-of-its-kind intelligence examination that bridges cognitive principles with voice-oriented benchmarks, while exposing overlooked challenges in multi-modal training.


【2】The Eloquence team submission for task 1 of MLC-SLM challenge
标题:口才团队提交MLC-LAM挑战任务1
链接:https://arxiv.org/abs/2507.19308

作者:Lorenzo Concina, Jordi Luque, Alessio Brutti, Marco Matassoni, Yuchen Zhang
备注:Technical Report for MLC-SLM Challenge of Interspeech2025
摘要:在本文中,我们提出了我们的研究和实验进行的任务1的挑战和研讨会上的多语种会话语音语言模型(MLC-SLM),重点是通过语音语言模型架构的发展,推进多语种会话语音识别。鉴于现实世界的会话数据越来越多的相关性,建立强大的口语对话系统,我们探讨三种方法来多语言ASR。首先,我们对官方基线进行评估,以更好地了解其优势和局限性,通过使用不同的基础模型训练两个投影仪(线性和qformer)。其次,我们利用SLAM-ASR框架来训练一个定制的多语言线性投影仪。最后,我们研究了对比学习和扩展的会话上下文在增强识别鲁棒性方面的作用。
摘要:In this paper, we present our studies and experiments carried out for the task 1 of the Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM), which focuses on advancing multilingual conversational speech recognition through the development of speech language models architectures. Given the increasing relevance of real-world conversational data for building robust Spoken Dialogue Systems, we explore three approaches to multilingual ASR. First, we conduct an evaluation of the official baseline to better understand its strengths and limitations, by training two projectors (linear and qformer) with different foundation models. Second we leverage the SLAM-ASR framework to train a custom multilingual linear projector. Finally we investigate the role of contrastive learning and the extended conversational context in enhancing the robustness of recognition.


【3】Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation
标题:Face 2Poker同步:用于文本驱动会说话的面部生成的轻量级面部语音一致性
链接:https://arxiv.org/abs/2507.19225

作者:Fang Kang, Yin Cao, Haoyu Chen
摘要:最近在语音驱动的说话人脸生成方面的研究取得了令人鼓舞的成果,但它们对固定驱动语音的依赖限制了进一步的应用(例如,面部-语音不匹配)。因此,我们将任务扩展到一个更具挑战性的设置:给定一个面部图像和文本说话,生成说话的面部动画和相应的语音。因此,我们提出了一个新的框架,Face 2 VoiceSync,具有几个新的贡献:1)语音-面部对齐,确保生成的语音匹配面部外观; 2)多样性\&操纵,使生成的语音控制在非语言特征空间; 3)高效训练,使用轻量级VAE桥接视觉和音频大型预训练模型,与现有方法相比,可训练参数明显减少; 4)新的评估指标,公平地评估多样性和身份一致性。实验表明,Face 2 VoiceSync在单个40 GB GPU上实现了最先进的视觉和音频性能。
摘要:Recent studies in speech-driven talking face generation achieve promising results, but their reliance on fixed-driven speech limits further applications (e.g., face-voice mismatch). Thus, we extend the task to a more challenging setting: given a face image and text to speak, generating both talking face animation and its corresponding speeches. Accordingly, we propose a novel framework, Face2VoiceSync, with several novel contributions: 1) Voice-Face Alignment, ensuring generated voices match facial appearance; 2) Diversity \& Manipulation, enabling generated voice control over paralinguistic features space; 3) Efficient Training, using a lightweight VAE to bridge visual and audio large-pretrained models, with significantly fewer trainable parameters than existing methods; 4) New Evaluation Metric, fairly assessing the diversity and identity consistency. Experiments show Face2VoiceSync achieves both visual and audio state-of-the-art performances on a single 40GB GPU.


【4】Latent Granular Resynthesis using Neural Audio Codecs
标题:使用神经音频编解码器的潜在颗粒再合成
链接:https://arxiv.org/abs/2507.19202

作者:Nao Tokui, Tom Baker
备注:Accepted at ISMIR 2025 Late Breaking Demos
摘要:我们介绍了一种新的技术,创造性的音频再合成,通过改造的概念,颗粒合成在潜在的向量水平。我们的方法通过将源音频语料编码成潜在向量段来创建“粒度码本”,然后将目标音频信号的每个潜在颗粒与码本中最接近的对应物进行匹配。解码得到的混合序列以产生保留目标的时间结构同时采用源的音色特征的音频。这种技术不需要模型训练,适用于各种音频材料,并且通过编解码器在解码期间的隐式插值自然避免了传统级联合成的典型不连续性。我们在https://github.com/naotokui/latentgranular/上提供了补充材料,以及一个概念验证实现,允许用户在https://huggingface.co/spaces/naotokui/latentgranular上试验他们自己的声音。
摘要:We introduce a novel technique for creative audio resynthesis that operates by reworking the concept of granular synthesis at the latent vector level. Our approach creates a "granular codebook" by encoding a source audio corpus into latent vector segments, then matches each latent grain of a target audio signal to its closest counterpart in the codebook. The resulting hybrid sequence is decoded to produce audio that preserves the target's temporal structure while adopting the source's timbral characteristics. This technique requires no model training, works with diverse audio materials, and naturally avoids the discontinuities typical of traditional concatenative synthesis through the codec's implicit interpolation during decoding. We include supplementary material at https://github.com/naotokui/latentgranular/ , as well as a proof-of-concept implementation to allow users to experiment with their own sounds at https://huggingface.co/spaces/naotokui/latentgranular .


【5】From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models
标题:从连续到离散:通过分层语言模型的跨领域协作通用语音增强
链接:https://arxiv.org/abs/2507.19062

作者:Zhaoxi Mu, Rilin Chen, Andong Li, Meng Yu, Xinyu Yang, Dong Yu
备注:ACMMM 2025
摘要:本文介绍了OmniGSE,一种新的通用语音增强(GSE)框架,旨在减轻语音信号在现实世界中遇到的各种失真。这些失真包括背景噪声、混响、带宽限制、信号削波和网络数据包丢失。现有的方法通常专注于针对单一类型的失真进行优化,通常难以有效地处理复杂场景中同时存在的多个失真。OmniGSE通过两阶段架构整合了判别式和生成式方法的优势,从而弥合了这一差距,该架构支持跨领域协作优化。在第一阶段,使用轻量级通道分割NAC-RoFormer增强连续特征。在第二阶段中,生成离散令牌以通过语言模型重建高质量的语音。具体来说,我们设计了一个层次化的语言模型结构,包括一个RootLM和多个BranchLM。RootLM模型跨码本层的一般声学特征,而BranchLM明确地捕获不同码本级别之间的渐进关系。实验结果表明,OmniGSE在多个基准测试中超越了现有模型,特别是在涉及复合失真的场景中表现出色。这些发现强调了该框架在现实世界应用中的强大和通用的语音增强的潜力。
摘要:This paper introduces OmniGSE, a novel general speech enhancement (GSE) framework designed to mitigate the diverse distortions that speech signals encounter in real-world scenarios. These distortions include background noise, reverberation, bandwidth limitations, signal clipping, and network packet loss. Existing methods typically focus on optimizing for a single type of distortion, often struggling to effectively handle the simultaneous presence of multiple distortions in complex scenarios. OmniGSE bridges this gap by integrating the strengths of discriminative and generative approaches through a two-stage architecture that enables cross-domain collaborative optimization. In the first stage, continuous features are enhanced using a lightweight channel-split NAC-RoFormer. In the second stage, discrete tokens are generated to reconstruct high-quality speech through language models. Specifically, we designed a hierarchical language model structure consisting of a RootLM and multiple BranchLMs. The RootLM models general acoustic features across codebook layers, while the BranchLMs explicitly capture the progressive relationships between different codebook levels. Experimental results demonstrate that OmniGSE surpasses existing models across multiple benchmarks, particularly excelling in scenarios involving compound distortions. These findings underscore the framework's potential for robust and versatile speech enhancement in real-world applications.


【6】MLLM-based Speech Recognition: When and How is Multimodality Beneficial?
标题:基于MLLM的语音识别:多模式何时以及如何受益?
链接:https://arxiv.org/abs/2507.19037

作者:Yiwen Guan, Viet Anh Trinh, Vivek Voleti, Jacob Whitehill
摘要:多模态大型语言模型(MLLM)的最新进展为语音、文本、图像和其他模态的统一建模开辟了新的可能性。在我们先前工作的基础上,本文研究了多输入模态可以提高噪声环境中自动语音识别(ASR)准确性的条件和模型架构。通过对合成数据和真实数据的实验,我们发现(1)利用更多的模态通常会提高ASR的准确性,因为每种模态都提供了互补的信息,但这种提高取决于听觉噪声的数量。(2)同步模态(例如,嘴唇运动)在高噪声水平下更有用而非同步模态(例如,图像上下文)在中等噪声水平下最有帮助。(3)更高质量的视觉表示不断提高ASR的准确性,突出了开发更强大的视觉编码器的重要性。(4)曼巴表现出与Transformers相似的多模态优势趋势。(5)模态的输入顺序以及它们在损失函数中的权重可以显著影响准确性。这些发现提供了实用的见解,并有助于加深我们对具有挑战性的条件下多模态语音识别的理解。
摘要:Recent advances in multi-modal large language models (MLLMs) have opened new possibilities for unified modeling of speech, text, images, and other modalities. Building on our prior work, this paper examines the conditions and model architectures under which multiple input modalities can improve automatic speech recognition (ASR) accuracy in noisy environments. Through experiments on synthetic and real-world data, we find that (1) harnessing more modalities usually improves ASR accuracy, as each modality provides complementary information, but the improvement depends on the amount of auditory noise. (2) Synchronized modalities (e.g., lip movements) are more useful at high noise levels whereas unsynchronized modalities (e.g., image context) are most helpful at moderate noise levels. (3) Higher-quality visual representations consistently improve ASR accuracy, highlighting the importance of developing more powerful visual encoders. (4) Mamba exhibits similar trends regarding the benefits of multimodality as do Transformers. (5) The input order of modalities as well as their weights in the loss function can significantly impact accuracy. These findings both offer practical insights and help to deepen our understanding of multi-modal speech recognition under challenging conditions.


【7】HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling
标题:HH-Codec:用于口语建模的高压缩高保真离散神经编解码器
链接:https://arxiv.org/abs/2507.18897

作者:Rongkun Xue, Yazhe Niu, Shuai Hu, Zixin Yin, Yongqiang Yao, Jing Yang
摘要:离散语音标记化是语音编解码器中的基本组成部分。然而,在大规模语音到语音系统中,来自多个量化器的并行流的复杂性和高时间维编解码器的计算成本构成了重大挑战。在本文中,我们介绍了HH-Codec,这是一种神经编解码器,可以在24 kHz音频中以每秒24个令牌的速度实现极端压缩,同时依赖于单量化器推理。我们的方法涉及到一个精心设计的矢量量化空间口语建模,优化压缩效率,同时最大限度地减少信息丢失。在此基础上,我们提出了一种非对称的编码器-解码器架构(Audio-VQ-Mel-Audio),该架构利用双重监督和渐进式训练来增强重建稳定性和保真度。HH-Codec以0.3 kbps的超低带宽实现了最先进的语音重建性能。我们进一步评估其有效性的码本利用率和生成模型自适应,与广泛的消融验证每个模块的必要性。HH-Codec可在https://github.com/opendilab/HH-Codec上获得。
摘要:Discrete speech tokenization is a fundamental component in speech codecs. However, in large-scale speech-to-speech systems, the complexity of parallel streams from multiple quantizers and the computational cost of high-time-dimensional codecs pose significant challenges. In this paper, we introduce HH-Codec, a neural codec that achieves extreme compression at 24 tokens per second for 24 kHz audio while relying on single-quantizer inference. Our approach involves a carefully designed Vector Quantization space for Spoken Language Modeling, optimizing compression efficiency while minimizing information loss. Building on this, we propose an asymmetric encoder-decoder architecture (Audio-VQ-Mel-Audio) that leverages dual supervision and progressive training to enhance reconstruction stability and fidelity. HH-Codec achieves state-of-the-art performance in speech reconstruction with an ultra-low bandwidth of 0.3 kbps. We further evaluate its effectiveness in codebook utilization and generative model adaptation, with extensive ablations validating the necessity of each module. HH-Codec is available at https://github.com/opendilab/HH-Codec.


【8】CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation
标题:CatchPhrase:用于音频到图像生成的Exext-guided编码器自适应
链接:https://arxiv.org/abs/2507.18750

作者:Hyunwoo Oh, SeungJu Cha, Kwanyoung Lee, Si-Woo Kim, Dong-Jin Kim
摘要:我们提出了CatchPhrase,一种新的音频到图像生成框架,旨在减轻音频输入和生成的图像之间的语义错位。虽然多模态编码器的最新进展使跨模态生成取得了进展,但源于同形异义词和听觉错觉的模糊性继续阻碍准确对齐。为了解决这个问题,CatchPhrase通过利用大型语言模型(LLM)和音频字幕模型(ACM)从弱类标签中生成丰富的跨模态语义提示(EXPrompt Mining)。为了解决类级别和实例级别的不一致,我们应用多模态过滤和检索来为每个音频样本选择语义上最一致的提示(EXPrompt)。然后训练轻量级映射网络以使预训练的文本到图像生成模型适应音频输入。在多个音频分类数据集上进行的大量实验表明,CatchPhrase可以改善音频到图像的对齐,并通过减轻语义不对齐来不断提高生成质量。
摘要:We propose CatchPhrase, a novel audio-to-image generation framework designed to mitigate semantic misalignment between audio inputs and generated images. While recent advances in multi-modal encoders have enabled progress in cross-modal generation, ambiguity stemming from homographs and auditory illusions continues to hinder accurate alignment. To address this issue, CatchPhrase generates enriched cross-modal semantic prompts (EXPrompt Mining) from weak class labels by leveraging large language models (LLMs) and audio captioning models (ACMs). To address both class-level and instance-level misalignment, we apply multi-modal filtering and retrieval to select the most semantically aligned prompt for each audio sample (EXPrompt Selector). A lightweight mapping network is then trained to adapt pre-trained text-to-image generation models to audio input. Extensive experiments on multiple audio classification datasets demonstrate that CatchPhrase improves audio-to-image alignment and consistently enhances generation quality by mitigating semantic misalignment.


【9】KuiSCIMA v2.0: Improved Baselines, Calibration, and Cross-Notation Generalization for Historical Chinese Music Notations in Jiang Kui's Baishidaoren Gequ
标题:KuiSCIMA v2.0:姜奎《白石真人歌曲》中中国历史音乐符号的改进基线、校准和交叉符号概括
链接:https://arxiv.org/abs/2507.18741

作者:Tristan Repolusk, Eduardo Veas
备注:International Conference on Document Analysis and Recognition. This preprint has not undergone any post-submission improvements or corrections. The Version of Record of this contribution is published in "19th International Conference on Document Analysis and Recognition (ICDAR 2025), Wuhan, China, September 16-21, 2025, Proceedings", and is available online at the External DOI field below
摘要:由于类别不平衡和训练数据有限,中国历史乐谱(如苏子谱和l“ul“upu)的光学音乐识别(OMR)面临着独特的挑战。本文介绍了姜夔《白石道人歌曲》自1202年以来的重要研究成果。在这项工作中,我们开发和评估一个字符识别模型的稀缺不平衡数据。我们通过将suzipu的字符错误率(CER)从10.4%降低到7.1%来改进先前的基线,尽管使用了77个高度不平衡的类,并且l\“ul\“upu的CER达到了0.9%。我们的模型优于人类转录器,平均人类CER为15.9%,最佳情况CER为7.6%。我们采用温度缩放来实现良好校准的模型,其预期校准误差(ECE)低于0.0162。使用留一个版本的交叉验证方法,我们确保了在五个历史版本中的稳健性能。此外,我们扩展了KuiSCIMA数据集,以包括来自白石道人歌曲的所有109个片段,包括suzipu,l\“ul\“upu和jianzipu标记。我们的研究结果推进了中国历史音乐的数字化和可访问性,促进了OMR的文化多样性,并将其适用性扩展到代表性不足的音乐传统。
摘要:Optical Music Recognition (OMR) for historical Chinese musical notations, such as suzipu and l\"ul\"upu, presents unique challenges due to high class imbalance and limited training data. This paper introduces significant advancements in OMR for Jiang Kui's influential collection Baishidaoren Gequ from 1202. In this work, we develop and evaluate a character recognition model for scarce imbalanced data. We improve upon previous baselines by reducing the Character Error Rate (CER) from 10.4% to 7.1% for suzipu, despite working with 77 highly imbalanced classes, and achieve a remarkable CER of 0.9% for l\"ul\"upu. Our models outperform human transcribers, with an average human CER of 15.9% and a best-case CER of 7.6%. We employ temperature scaling to achieve a well-calibrated model with an Expected Calibration Error (ECE) below 0.0162. Using a leave-one-edition-out cross-validation approach, we ensure robust performance across five historical editions. Additionally, we extend the KuiSCIMA dataset to include all 109 pieces from Baishidaoren Gequ, encompassing suzipu, l\"ul\"upu, and jianzipu notations. Our findings advance the digitization and accessibility of historical Chinese music, promoting cultural diversity in OMR and expanding its applicability to underrepresented music traditions.


【10】SCORE-SET: A dataset of GuitarPro files for Music Phrase Generation and Sequence Learning
标题:SCORE-SET:用于音乐短语生成和序列学习的GuitarPro文件数据集
链接:https://arxiv.org/abs/2507.18723

作者:Vishakh Begari
备注:6 pages, 6 figures
摘要:提供了一个精心策划的Guitar Pro指法文件数据集(.gp5格式),专为涉及吉他音乐生成、序列建模和表演感知学习的任务而量身定制。该数据集来自MAESTRO和Giantstrom中的音符,这些音符已被改编为节奏吉他曲目。这些音轨经过进一步处理,以包括吉他演奏中典型的各种表情设置,例如压音、滑动、颤音和手掌静音,以更好地反映现实世界吉他演奏的细微差别。
摘要:A curated dataset of Guitar Pro tablature files (.gp5 format), tailored for tasks involving guitar music generation, sequence modeling, and performance-aware learning is provided. The dataset is derived from MIDI notes in MAESTRO and GiantMIDI which have been adapted into rhythm guitar tracks. These tracks are further processed to include a variety of expression settings typical of guitar performance, such as bends, slides, vibrato, and palm muting, to better reflect the nuances of real-world guitar playing.


【11】Binaural Target Speaker Extraction using HRTFs and a Complex-Valued Neural Network
标题:使用HRTTF和复值神经网络的双耳目标说话人提取
链接:https://arxiv.org/abs/2507.19369

作者:Yoav Ellinson, Sharon Gannot
摘要:在这项工作中,我们的目标是模仿人类的能力,有选择地参加一个单一的扬声器,即使在多个同时说话的存在。我们提出了一种新的方法,利用双耳目标说话人提取听者的头相关传递函数(HRTF),以隔离所需的扬声器。值得注意的是,我们的方法不依赖于说话者嵌入,使其与说话者无关,并能够在不同语言的多个语音数据集上实现强大的泛化。   我们采用了一个完全复值的神经网络,直接对混合音频信号的复值短时傅立叶变换(STFT)进行操作。这偏离了使用频谱图或将STFT的实部和虚部视为单独实值输入的传统方法。   我们首先评估的方法在消声,无噪声的情况下,它表现出优异的提取性能,同时有效地保留目标信号的双耳线索。然后,我们测试一个修改后的变种在温和的混响条件下。该版本在混响环境中保持鲁棒性,保持语音清晰度,保持源方向性,同时减少混响。
摘要:In this work, we aim to imitate the human ability to selectively attend to a single speaker, even in the presence of multiple simultaneous talkers. We propose a novel approach for binaural target speaker extraction that leverages the listener's Head-Related Transfer Function (HRTF) to isolate the desired speaker. Notably, our method does not rely on speaker embeddings, making it speaker-independent and enabling strong generalization across multiple speech datasets in different languages.   We employ a fully complex-valued neural network that operates directly on the complex-valued Short-Time Fourier Transform (STFT) of the mixed audio signals. This deviates from conventional approaches that use spectrograms or treat the real and imaginary components of the STFT as separate real-valued inputs.   We first evaluate the method in an anechoic, noise-free scenario, where it demonstrates excellent extraction performance while effectively preserving the binaural cues of the target signal. We then test a modified variant under mild reverberation conditions. This version remains robust in reverberant environments, maintaining speech clarity, preserving source directionality, and simultaneously reducing reverberation.


【12】Should Top-Down Clustering Affect Boundaries in Unsupervised Word Discovery?
标题:自顶向下聚类是否会影响无监督词发现中的边界?
链接:https://arxiv.org/abs/2507.19204

作者:Simon Malan, Benjamin van Niekerk, Herman Kamper
备注:5 figures, 5 tables
摘要:我们调查的问题分割成词样单位和聚类这些无标记的语音创建一个词典。以前的工作可以分为两个框架。自底向上方法首先确定边界,然后将固定的分割词聚类到词典中。相比之下,自上而下的方法结合了来自聚类词的信息来通知边界选择。然而,目前还不清楚自上而下的信息是否是必要的,以改善分割。为了探索这一点,我们看两个类似的方法,不同的是自上而下的集群通知边界选择。我们简单的自下而上的策略使用相邻的自监督特征之间的相异性来预测单词边界,然后将所得到的片段聚类以构建词典。我们的自顶向下系统是ES-KMeans动态规划方法的更新版本,该方法迭代地使用K-means来更新其边界。在五种语言的ZeroSpeech基准测试中,这两种方法都取得了相当的最先进的结果,自底向上的系统快了近五倍。通过详细的分析,我们表明ES-KMeans的自上而下的影响可能是有益的(取决于候选边界等因素),但在许多情况下,简单的自下而上的方法也同样有效。对于这两种方法,我们表明,聚类步骤是一个限制因素。因此,我们建议未来的工作重点放在改进的聚类技术和学习更多的歧视性词样表示。项目代码库:https://github.com/s-malan/prom-seg-clus。
摘要:We investigate the problem of segmenting unlabeled speech into word-like units and clustering these to create a lexicon. Prior work can be categorized into two frameworks. Bottom-up methods first determine boundaries and then cluster the fixed segmented words into a lexicon. In contrast, top-down methods incorporate information from the clustered words to inform boundary selection. However, it is unclear whether top-down information is necessary to improve segmentation. To explore this, we look at two similar approaches that differ in whether top-down clustering informs boundary selection. Our simple bottom-up strategy predicts word boundaries using the dissimilarity between adjacent self-supervised features, then clusters the resulting segments to construct a lexicon. Our top-down system is an updated version of the ES-KMeans dynamic programming method that iteratively uses K-means to update its boundaries. On the five-language ZeroSpeech benchmarks, both approaches achieve comparable state-of-the-art results, with the bottom-up system being nearly five times faster. Through detailed analyses, we show that the top-down influence of ES-KMeans can be beneficial (depending on factors like the candidate boundaries), but in many cases the simple bottom-up method performs just as well. For both methods, we show that the clustering step is a limiting factor. Therefore, we recommend that future work focus on improved clustering techniques and learning more discriminative word-like representations. Project code repository: https://github.com/s-malan/prom-seg-clus.


【13】Assessment of Personality Dimensions Across Situations Using Conversational Speech
标题:使用对话言语评估跨情境的性格维度
链接:https://arxiv.org/abs/2507.19137

作者:Alice Zhang, Skanda Muralidhar, Daniel Gatica-Perez, Mathew Magimai-Doss
摘要:先前的研究表明,用户更喜欢与自己个性一致的辅助技术。这引发了人们对自动人格感知(APP)的兴趣,APP旨在预测个人的感知人格特质。先前的APP研究将人格视为静态特征,独立于上下文。然而,正如心理学研究所显示的那样,感知的个性会因环境和情况而异。在这项研究中,我们调查了两种工作情况下(一个中立的采访和紧张的客户互动)的参与者的会话言语和感知的个性之间的关系。我们的主要发现是:1)感知的个性在交互中显著不同,2)响度、声音水平和频谱通量特征指示中性交互中感知的外向性、宜人性、开放性,而神经质在压力情境中与这些特征相关,3)手工制作的声学特征和非语言特征在感知的个性的推断中优于说话者嵌入,4)压力性互动更能预测神经质,这与现有的心理学研究一致。
摘要:Prior research indicates that users prefer assistive technologies whose personalities align with their own. This has sparked interest in automatic personality perception (APP), which aims to predict an individual's perceived personality traits. Previous studies in APP have treated personalities as static traits, independent of context. However, perceived personalities can vary by context and situation as shown in psychological research. In this study, we investigate the relationship between conversational speech and perceived personality for participants engaged in two work situations (a neutral interview and a stressful client interaction). Our key findings are: 1) perceived personalities differ significantly across interactions, 2) loudness, sound level, and spectral flux features are indicative of perceived extraversion, agreeableness, conscientiousness, and openness in neutral interactions, while neuroticism correlates with these features in stressful contexts, 3) handcrafted acoustic features and non-verbal features outperform speaker embeddings in inference of perceived personality, and 4) stressful interactions are more predictive of neuroticism, aligning with existing psychological research.


eess.AS音频处理


【1】Binaural Target Speaker Extraction using HRTFs and a Complex-Valued Neural Network
标题:使用HRTTF和复值神经网络的双耳目标说话人提取
链接:https://arxiv.org/abs/2507.19369

作者:Yoav Ellinson, Sharon Gannot
摘要:在这项工作中,我们的目标是模仿人类的能力,有选择地参加一个单一的扬声器,即使在多个同时说话的存在。我们提出了一种新的方法,利用双耳目标说话人提取听者的头相关传递函数(HRTF),以隔离所需的扬声器。值得注意的是,我们的方法不依赖于说话者嵌入,使其与说话者无关,并能够在不同语言的多个语音数据集上实现强大的泛化。   我们采用了一个完全复值的神经网络,直接对混合音频信号的复值短时傅立叶变换(STFT)进行操作。这偏离了使用频谱图或将STFT的实部和虚部视为单独实值输入的传统方法。   我们首先评估的方法在消声,无噪声的情况下,它表现出优异的提取性能,同时有效地保留目标信号的双耳线索。然后,我们测试一个修改后的变种在温和的混响条件下。该版本在混响环境中保持鲁棒性,保持语音清晰度,保持源方向性,同时减少混响。
摘要:In this work, we aim to imitate the human ability to selectively attend to a single speaker, even in the presence of multiple simultaneous talkers. We propose a novel approach for binaural target speaker extraction that leverages the listener's Head-Related Transfer Function (HRTF) to isolate the desired speaker. Notably, our method does not rely on speaker embeddings, making it speaker-independent and enabling strong generalization across multiple speech datasets in different languages.   We employ a fully complex-valued neural network that operates directly on the complex-valued Short-Time Fourier Transform (STFT) of the mixed audio signals. This deviates from conventional approaches that use spectrograms or treat the real and imaginary components of the STFT as separate real-valued inputs.   We first evaluate the method in an anechoic, noise-free scenario, where it demonstrates excellent extraction performance while effectively preserving the binaural cues of the target signal. We then test a modified variant under mild reverberation conditions. This version remains robust in reverberant environments, maintaining speech clarity, preserving source directionality, and simultaneously reducing reverberation.


【2】Comparison of Knowledge Distillation Methods for Low-complexity Multi-microphone Speech Enhancement using the FT-JNF Architecture
标题:使用FT-JNF架构的低复杂度多麦克风语音增强的知识提取方法比较
链接:https://arxiv.org/abs/2507.19208

作者:Robert Metzger, Mattes Ohlenbusch, Christian Rollwage, Simon Doclo
备注:Accepted at the ITG Conference on Speech Communication 2025 in Berlin
摘要:近年来,使用深度神经网络(DNN)的多麦克风语音增强取得了显着进展。然而,许多提出的基于DNN的语音增强算法无法在具有有限硬件资源的设备上实现。仅通过减少参数的数量来降低这种系统的复杂性通常会导致更差的性能。知识蒸馏(KD)是一种在保持性能的同时减小DNN模型大小的有前途的方法。在本文中,我们考虑最近提出的频率-时间联合非线性滤波器(FT-JNF)架构,并研究了几种KD方法,以训练较小的(学生)模型从一个大的预训练(教师)模型。五KD方法进行评估,使用直接输出匹配,自相似的中间层,融合多层损失。在使用具有五个麦克风的紧凑阵列的模拟数据集上的实验结果表明,与没有KD的训练相比,三种KD方法大大提高了学生模型的性能。仅具有教师模型参数的25%的学生模型在0 dB SNR下实现了相当的PESQ分数。此外,模型大小可以减少高达96%,而PESQ评分仅略有下降。
摘要:Multi-microphone speech enhancement using deep neural networks (DNNs) has significantly progressed in recent years. However, many proposed DNN-based speech enhancement algorithms cannot be implemented on devices with limited hardware resources. Only lowering the complexity of such systems by reducing the number of parameters often results in worse performance. Knowledge Distillation (KD) is a promising approach for reducing DNN model size while preserving performance. In this paper, we consider the recently proposed Frequency-Time Joint Non-linear Filter (FT-JNF) architecture and investigate several KD methods to train smaller (student) models from a large pre-trained (teacher) model. Five KD methods are evaluated using direct output matching, the self-similarity of intermediate layers, and fused multi-layer losses. Experimental results on a simulated dataset using a compact array with five microphones show that three KD methods substantially improve the performance of student models compared to training without KD. A student model with only 25% of the teacher model's parameters achieves comparable PESQ scores at 0 dB SNR. Furthermore, a reduction of up to 96% in model size can be achieved with only a minimal decrease in PESQ scores.


【3】Should Top-Down Clustering Affect Boundaries in Unsupervised Word Discovery?
标题:自顶向下聚类是否会影响无监督词发现中的边界?
链接:https://arxiv.org/abs/2507.19204

作者:Simon Malan, Benjamin van Niekerk, Herman Kamper
备注:5 figures, 5 tables
摘要:我们调查的问题分割成词样单位和聚类这些无标记的语音创建一个词典。以前的工作可以分为两个框架。自底向上方法首先确定边界,然后将固定的分割词聚类到词典中。相比之下,自上而下的方法结合了来自聚类词的信息来通知边界选择。然而,目前还不清楚自上而下的信息是否是必要的,以改善分割。为了探索这一点,我们看两个类似的方法,不同的是自上而下的集群通知边界选择。我们简单的自下而上的策略使用相邻的自监督特征之间的相异性来预测单词边界,然后将所得到的片段聚类以构建词典。我们的自顶向下系统是ES-KMeans动态规划方法的更新版本,该方法迭代地使用K-means来更新其边界。在五种语言的ZeroSpeech基准测试中,这两种方法都取得了相当的最先进的结果,自底向上的系统快了近五倍。通过详细的分析,我们表明ES-KMeans的自上而下的影响可能是有益的(取决于候选边界等因素),但在许多情况下,简单的自下而上的方法也同样有效。对于这两种方法,我们表明,聚类步骤是一个限制因素。因此,我们建议未来的工作重点放在改进的聚类技术和学习更多的歧视性词样表示。项目代码库:https://github.com/s-malan/prom-seg-clus。
摘要:We investigate the problem of segmenting unlabeled speech into word-like units and clustering these to create a lexicon. Prior work can be categorized into two frameworks. Bottom-up methods first determine boundaries and then cluster the fixed segmented words into a lexicon. In contrast, top-down methods incorporate information from the clustered words to inform boundary selection. However, it is unclear whether top-down information is necessary to improve segmentation. To explore this, we look at two similar approaches that differ in whether top-down clustering informs boundary selection. Our simple bottom-up strategy predicts word boundaries using the dissimilarity between adjacent self-supervised features, then clusters the resulting segments to construct a lexicon. Our top-down system is an updated version of the ES-KMeans dynamic programming method that iteratively uses K-means to update its boundaries. On the five-language ZeroSpeech benchmarks, both approaches achieve comparable state-of-the-art results, with the bottom-up system being nearly five times faster. Through detailed analyses, we show that the top-down influence of ES-KMeans can be beneficial (depending on factors like the candidate boundaries), but in many cases the simple bottom-up method performs just as well. For both methods, we show that the clustering step is a limiting factor. Therefore, we recommend that future work focus on improved clustering techniques and learning more discriminative word-like representations. Project code repository: https://github.com/s-malan/prom-seg-clus.


【4】Assessment of Personality Dimensions Across Situations Using Conversational Speech
标题:使用对话言语评估跨情境的性格维度
链接:https://arxiv.org/abs/2507.19137

作者:Alice Zhang, Skanda Muralidhar, Daniel Gatica-Perez, Mathew Magimai-Doss
摘要:先前的研究表明,用户更喜欢与自己个性一致的辅助技术。这引发了人们对自动人格感知(APP)的兴趣,APP旨在预测个人的感知人格特质。先前的APP研究将人格视为静态特征,独立于上下文。然而,正如心理学研究所显示的那样,感知的个性会因环境和情况而异。在这项研究中,我们调查了两种工作情况下(一个中立的采访和紧张的客户互动)的参与者的会话言语和感知的个性之间的关系。我们的主要发现是:1)感知的个性在交互中显著不同,2)响度、声音水平和频谱通量特征指示中性交互中感知的外向性、宜人性、开放性,而神经质在压力情境中与这些特征相关,3)手工制作的声学特征和非语言特征在感知的个性的推断中优于说话者嵌入,4)压力性互动更能预测神经质,这与现有的心理学研究一致。
摘要:Prior research indicates that users prefer assistive technologies whose personalities align with their own. This has sparked interest in automatic personality perception (APP), which aims to predict an individual's perceived personality traits. Previous studies in APP have treated personalities as static traits, independent of context. However, perceived personalities can vary by context and situation as shown in psychological research. In this study, we investigate the relationship between conversational speech and perceived personality for participants engaged in two work situations (a neutral interview and a stressful client interaction). Our key findings are: 1) perceived personalities differ significantly across interactions, 2) loudness, sound level, and spectral flux features are indicative of perceived extraversion, agreeableness, conscientiousness, and openness in neutral interactions, while neuroticism correlates with these features in stressful contexts, 3) handcrafted acoustic features and non-verbal features outperform speaker embeddings in inference of perceived personality, and 4) stressful interactions are more predictive of neuroticism, aligning with existing psychological research.


【5】FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems
标题:FD-Bench:一种为全双工语音对话系统设计的全双工基准测试流水线
链接:https://arxiv.org/abs/2507.19040

作者:Yizhou Peng, Yi-Wen Chao, Dianwen Ng, Yukun Ma, Chongjia Ni, Bin Ma, Eng Siong Chng
备注:Accepted to Interspeech 2025. 5 pages
摘要:与依赖话轮转换的传统SDS相比,全双工口语对话系统(FDSDS)允许实时用户中断和反向通道,从而实现更自然的人机交互。然而,现有的基准缺乏针对FD场景的度量,例如,在用户中断期间评估模型性能。在本文中,我们提出了一个全面的FD基准测试管道,利用LLM,TTS和ASR来解决这一差距。它评估了FDSDS处理用户中断、管理延迟以及在具有挑战性的场景中保持鲁棒性的能力。我们将我们的基准应用于三个开源FDSDS(Moshi,Freeze-omni和VITA-1.5),使用超过40小时的生成语音,293次模拟对话和1,200次中断。结果表明,所有模型都继续面临挑战,例如在频繁中断和噪声条件下无法响应用户中断。将发布演示、数据和代码。
摘要:Full-duplex spoken dialogue systems (FDSDS) enable more natural human-machine interactions by allowing real-time user interruptions and backchanneling, compared to traditional SDS that rely on turn-taking. However, existing benchmarks lack metrics for FD scenes, e.g., evaluating model performance during user interruptions. In this paper, we present a comprehensive FD benchmarking pipeline utilizing LLMs, TTS, and ASR to address this gap. It assesses FDSDS's ability to handle user interruptions, manage delays, and maintain robustness in challenging scenarios with diverse novel metrics. We applied our benchmark to three open-source FDSDS (Moshi, Freeze-omni, and VITA-1.5) using over 40 hours of generated speech, with 293 simulated conversations and 1,200 interruptions. The results show that all models continue to face challenges, such as failing to respond to user interruptions, under frequent disruptions and noisy conditions. Demonstrations, data, and code will be released.


【6】SpeechIQ: Speech Intelligence Quotient Across Cognitive Levels in Voice Understanding Large Language Models
标题:SpeechIQ:语音理解大型语言模型中跨认知水平的言语智商
链接:https://arxiv.org/abs/2507.19361

作者: Zhen Wan, Chao-Han Huck Yang, Yahan Yu, Jinchuan Tian, Sheng Li, Ke Hu, Zhehuai Chen, Shinji Watanabe, Fei Cheng, Chenhui Chu, Sadao Kurohashi
备注:Our Speech-IQ leaderboard will be hosted at this http URL. ACL 2025 main
摘要:我们引入了基于语音的智商(SIQ)作为一种新形式的人类认知启发的评估管道,用于语音理解大型语言模型LLM Voice,旨在评估其语音理解能力。超越流行的语音理解指标,如单词错误率(WER),SIQ检查LLM语音跨越三个认知水平由布卢姆的分类:(1)记住(即,(2)理解(即,法学硕士解释的相似性);和(3)应用(即,模拟下游任务的QA准确性)。我们证明,SIQ不仅量化了语音理解能力,而且还提供了级联方法之间的统一比较(例如,ASR LLM)和端到端模型,识别现有基准测试中的注释错误,并检测LLM语音中的幻觉。我们的框架代表了第一个同类智力考试,将认知原则与语音基准相结合,同时暴露了多模式培训中被忽视的挑战。
摘要:We introduce Speech-based Intelligence Quotient (SIQ) as a new form of human cognition-inspired evaluation pipeline for voice understanding large language models, LLM Voice, designed to assess their voice understanding ability. Moving beyond popular voice understanding metrics such as word error rate (WER), SIQ examines LLM Voice across three cognitive levels motivated by Bloom's Taxonomy: (1) Remembering (i.e., WER for verbatim accuracy); (2) Understanding (i.e., similarity of LLM's interpretations); and (3) Application (i.e., QA accuracy for simulating downstream tasks). We demonstrate that SIQ not only quantifies voice understanding abilities but also provides unified comparisons between cascaded methods (e.g., ASR LLM) and end-to-end models, identifies annotation errors in existing benchmarks, and detects hallucinations in LLM Voice. Our framework represents a first-of-its-kind intelligence examination that bridges cognitive principles with voice-oriented benchmarks, while exposing overlooked challenges in multi-modal training.


【7】The Eloquence team submission for task 1 of MLC-SLM challenge
标题:口才团队提交MLC-LAM挑战任务1
链接:https://arxiv.org/abs/2507.19308

作者:Lorenzo Concina, Jordi Luque, Alessio Brutti, Marco Matassoni, Yuchen Zhang
备注:Technical Report for MLC-SLM Challenge of Interspeech2025
摘要:在本文中,我们提出了我们的研究和实验进行的任务1的挑战和研讨会上的多语种会话语音语言模型(MLC-SLM),重点是通过语音语言模型架构的发展,推进多语种会话语音识别。鉴于现实世界的会话数据越来越多的相关性,建立强大的口语对话系统,我们探讨三种方法来多语言ASR。首先,我们对官方基线进行评估,以更好地了解其优势和局限性,通过使用不同的基础模型训练两个投影仪(线性和qformer)。其次,我们利用SLAM-ASR框架来训练一个定制的多语言线性投影仪。最后,我们研究了对比学习和扩展的会话上下文在增强识别鲁棒性方面的作用。
摘要:In this paper, we present our studies and experiments carried out for the task 1 of the Challenge and Workshop on Multilingual Conversational Speech Language Model (MLC-SLM), which focuses on advancing multilingual conversational speech recognition through the development of speech language models architectures. Given the increasing relevance of real-world conversational data for building robust Spoken Dialogue Systems, we explore three approaches to multilingual ASR. First, we conduct an evaluation of the official baseline to better understand its strengths and limitations, by training two projectors (linear and qformer) with different foundation models. Second we leverage the SLAM-ASR framework to train a custom multilingual linear projector. Finally we investigate the role of contrastive learning and the extended conversational context in enhancing the robustness of recognition.


【8】Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation
标题:Face 2Poker同步:用于文本驱动会说话的面部生成的轻量级面部语音一致性
链接:https://arxiv.org/abs/2507.19225

作者:Fang Kang, Yin Cao, Haoyu Chen
摘要:最近在语音驱动的说话人脸生成中的研究取得了有希望的结果,但是它们对固定驱动语音的依赖限制了进一步的应用(例如,面部-语音不匹配)。因此,我们将任务扩展到一个更具挑战性的设置:给定一个面部图像和文本说话,生成说话的面部动画和相应的语音。因此,我们提出了一个新的框架,Face 2 VoiceSync,具有几个新的贡献:1)语音-面部对齐,确保生成的语音匹配面部外观; 2)多样性\&操纵,使生成的语音控制在非语言特征空间; 3)高效训练,使用轻量级VAE桥接视觉和音频大型预训练模型,与现有方法相比,可训练参数明显减少; 4)新的评价指标,公平地评估多样性和同一性的一致性。实验表明,Face 2 VoiceSync在单个40 GB GPU上实现了最先进的视觉和音频性能。
摘要:Recent studies in speech-driven talking face generation achieve promising results, but their reliance on fixed-driven speech limits further applications (e.g., face-voice mismatch). Thus, we extend the task to a more challenging setting: given a face image and text to speak, generating both talking face animation and its corresponding speeches. Accordingly, we propose a novel framework, Face2VoiceSync, with several novel contributions: 1) Voice-Face Alignment, ensuring generated voices match facial appearance; 2) Diversity \& Manipulation, enabling generated voice control over paralinguistic features space; 3) Efficient Training, using a lightweight VAE to bridge visual and audio large-pretrained models, with significantly fewer trainable parameters than existing methods; 4) New Evaluation Metric, fairly assessing the diversity and identity consistency. Experiments show Face2VoiceSync achieves both visual and audio state-of-the-art performances on a single 40GB GPU.


【9】Latent Granular Resynthesis using Neural Audio Codecs
标题:使用神经音频编解码器的潜在颗粒再合成
链接:https://arxiv.org/abs/2507.19202

作者:Nao Tokui, Tom Baker
备注:Accepted at ISMIR 2025 Late Breaking Demos
摘要:我们介绍了一种新的技术,创造性的音频再合成,通过改造的概念,颗粒合成在潜在的向量水平。我们的方法通过将源音频语料编码成潜在向量段来创建“粒度码本”,然后将目标音频信号的每个潜在颗粒与码本中最接近的对应物进行匹配。将所得混合序列解码以产生保留目标时间结构同时采用源音色特征的音频。这种技术不需要模型训练,适用于各种音频材料,并且通过编解码器在解码期间的隐式插值自然避免了传统级联合成的典型不连续性。我们在https://github.com/naotokui/latentgranular/上提供了补充材料,以及一个概念验证实现,允许用户在https://huggingface.co/spaces/naotokui/latentgranular上试验他们自己的声音。
摘要:We introduce a novel technique for creative audio resynthesis that operates by reworking the concept of granular synthesis at the latent vector level. Our approach creates a "granular codebook" by encoding a source audio corpus into latent vector segments, then matches each latent grain of a target audio signal to its closest counterpart in the codebook. The resulting hybrid sequence is decoded to produce audio that preserves the target's temporal structure while adopting the source's timbral characteristics. This technique requires no model training, works with diverse audio materials, and naturally avoids the discontinuities typical of traditional concatenative synthesis through the codec's implicit interpolation during decoding. We include supplementary material at https://github.com/naotokui/latentgranular/ , as well as a proof-of-concept implementation to allow users to experiment with their own sounds at https://huggingface.co/spaces/naotokui/latentgranular .


【10】From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models
标题:从连续到离散:通过分层语言模型的跨领域协作通用语音增强
链接:https://arxiv.org/abs/2507.19062

作者:Zhaoxi Mu, Rilin Chen, Andong Li, Meng Yu, Xinyu Yang, Dong Yu
备注:ACMMM 2025
摘要:本文介绍了OmniGSE,一种新的通用语音增强(GSE)框架,旨在减轻语音信号在现实世界中遇到的各种失真。这些失真包括背景噪声、混响、带宽限制、信号削波和网络数据包丢失。现有的方法通常专注于针对单一类型的失真进行优化,通常难以有效地处理复杂场景中同时存在的多个失真。OmniGSE通过两阶段架构整合了判别式和生成式方法的优势,从而弥合了这一差距,该架构支持跨领域协作优化。在第一阶段,使用轻量级通道分割NAC-RoFormer增强连续特征。在第二阶段中,生成离散令牌以通过语言模型重建高质量的语音。具体来说,我们设计了一个层次化的语言模型结构,包括一个RootLM和多个BranchLM。RootLM模型跨码本层的一般声学特征,而BranchLM明确地捕获不同码本级别之间的渐进关系。实验结果表明,OmniGSE在多个基准测试中超越了现有模型,特别是在涉及复合失真的场景中表现出色。这些发现强调了该框架在现实世界应用中的强大和通用的语音增强的潜力。
摘要:This paper introduces OmniGSE, a novel general speech enhancement (GSE) framework designed to mitigate the diverse distortions that speech signals encounter in real-world scenarios. These distortions include background noise, reverberation, bandwidth limitations, signal clipping, and network packet loss. Existing methods typically focus on optimizing for a single type of distortion, often struggling to effectively handle the simultaneous presence of multiple distortions in complex scenarios. OmniGSE bridges this gap by integrating the strengths of discriminative and generative approaches through a two-stage architecture that enables cross-domain collaborative optimization. In the first stage, continuous features are enhanced using a lightweight channel-split NAC-RoFormer. In the second stage, discrete tokens are generated to reconstruct high-quality speech through language models. Specifically, we designed a hierarchical language model structure consisting of a RootLM and multiple BranchLMs. The RootLM models general acoustic features across codebook layers, while the BranchLMs explicitly capture the progressive relationships between different codebook levels. Experimental results demonstrate that OmniGSE surpasses existing models across multiple benchmarks, particularly excelling in scenarios involving compound distortions. These findings underscore the framework's potential for robust and versatile speech enhancement in real-world applications.


【11】MLLM-based Speech Recognition: When and How is Multimodality Beneficial?
标题:基于MLLM的语音识别:多模式何时以及如何受益?
链接:https://arxiv.org/abs/2507.19037

作者:Yiwen Guan, Viet Anh Trinh, Vivek Voleti, Jacob Whitehill
摘要:多模态大型语言模型(MLLM)的最新进展为语音、文本、图像和其他模态的统一建模开辟了新的可能性。在我们先前工作的基础上,本文研究了多输入模态可以提高噪声环境中自动语音识别(ASR)准确性的条件和模型架构。通过对合成数据和真实数据的实验,我们发现(1)利用更多的模态通常会提高ASR的准确性,因为每种模态都提供了互补的信息,但这种提高取决于听觉噪声的数量。(2)同步模态(例如,嘴唇运动)在高噪声水平下更有用而非同步模态(例如,图像上下文)在中等噪声水平下最有帮助。(3)更高质量的视觉表示不断提高ASR的准确性,突出了开发更强大的视觉编码器的重要性。(4)曼巴表现出与Transformers相似的多模态优势趋势。(5)模态的输入顺序以及它们在损失函数中的权重可以显著影响准确性。这些发现提供了实用的见解,并有助于加深我们对具有挑战性的条件下多模态语音识别的理解。
摘要:Recent advances in multi-modal large language models (MLLMs) have opened new possibilities for unified modeling of speech, text, images, and other modalities. Building on our prior work, this paper examines the conditions and model architectures under which multiple input modalities can improve automatic speech recognition (ASR) accuracy in noisy environments. Through experiments on synthetic and real-world data, we find that (1) harnessing more modalities usually improves ASR accuracy, as each modality provides complementary information, but the improvement depends on the amount of auditory noise. (2) Synchronized modalities (e.g., lip movements) are more useful at high noise levels whereas unsynchronized modalities (e.g., image context) are most helpful at moderate noise levels. (3) Higher-quality visual representations consistently improve ASR accuracy, highlighting the importance of developing more powerful visual encoders. (4) Mamba exhibits similar trends regarding the benefits of multimodality as do Transformers. (5) The input order of modalities as well as their weights in the loss function can significantly impact accuracy. These findings both offer practical insights and help to deepen our understanding of multi-modal speech recognition under challenging conditions.


【12】HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling
标题:HH-Codec:用于口语建模的高压缩高保真离散神经编解码器
链接:https://arxiv.org/abs/2507.18897

作者:Rongkun Xue, Yazhe Niu, Shuai Hu, Zixin Yin, Yongqiang Yao, Jing Yang
摘要:离散语音标记化是语音编解码器中的基本组成部分。然而,在大规模语音到语音系统中,来自多个量化器的并行流的复杂性和高时间维编解码器的计算成本构成了重大挑战。在本文中,我们介绍了HH-Codec,这是一种神经编解码器,可以在24 kHz音频中以每秒24个令牌的速度实现极端压缩,同时依赖于单量化器推理。我们的方法涉及到一个精心设计的矢量量化空间口语建模,优化压缩效率,同时最大限度地减少信息丢失。在此基础上,我们提出了一种非对称的编码器-解码器架构(Audio-VQ-Mel-Audio),该架构利用双重监督和渐进式训练来增强重建稳定性和保真度。HH-Codec以0.3 kbps的超低带宽实现了最先进的语音重建性能。我们进一步评估其有效性的码本利用率和生成模型自适应,与广泛的消融验证每个模块的必要性。HH-Codec可在https://github.com/opendilab/HH-Codec上获得。
摘要:Discrete speech tokenization is a fundamental component in speech codecs. However, in large-scale speech-to-speech systems, the complexity of parallel streams from multiple quantizers and the computational cost of high-time-dimensional codecs pose significant challenges. In this paper, we introduce HH-Codec, a neural codec that achieves extreme compression at 24 tokens per second for 24 kHz audio while relying on single-quantizer inference. Our approach involves a carefully designed Vector Quantization space for Spoken Language Modeling, optimizing compression efficiency while minimizing information loss. Building on this, we propose an asymmetric encoder-decoder architecture (Audio-VQ-Mel-Audio) that leverages dual supervision and progressive training to enhance reconstruction stability and fidelity. HH-Codec achieves state-of-the-art performance in speech reconstruction with an ultra-low bandwidth of 0.3 kbps. We further evaluate its effectiveness in codebook utilization and generative model adaptation, with extensive ablations validating the necessity of each module. HH-Codec is available at https://github.com/opendilab/HH-Codec.


【13】CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation
标题:CatchPhrase:用于音频到图像生成的Exext-guided编码器自适应
链接:https://arxiv.org/abs/2507.18750

作者:Hyunwoo Oh, SeungJu Cha, Kwanyoung Lee, Si-Woo Kim, Dong-Jin Kim
摘要:我们提出了CatchPhrase,一种新的音频到图像生成框架,旨在减轻音频输入和生成的图像之间的语义错位。虽然多模态编码器的最新进展使跨模态生成取得了进展,但源于同形异义词和听觉错觉的模糊性继续阻碍准确对齐。为了解决这个问题,CatchPhrase通过利用大型语言模型(LLM)和音频字幕模型(ACM)从弱类标签中生成丰富的跨模态语义提示(EXPrompt Mining)。为了解决类级别和实例级别的不一致,我们应用多模态过滤和检索来为每个音频样本选择语义上最一致的提示(EXPrompt)。然后训练轻量级映射网络以使预训练的文本到图像生成模型适应音频输入。在多个音频分类数据集上进行的大量实验表明,CatchPhrase可以改善音频到图像的对齐,并通过减轻语义不对齐来不断提高生成质量。
摘要:We propose CatchPhrase, a novel audio-to-image generation framework designed to mitigate semantic misalignment between audio inputs and generated images. While recent advances in multi-modal encoders have enabled progress in cross-modal generation, ambiguity stemming from homographs and auditory illusions continues to hinder accurate alignment. To address this issue, CatchPhrase generates enriched cross-modal semantic prompts (EXPrompt Mining) from weak class labels by leveraging large language models (LLMs) and audio captioning models (ACMs). To address both class-level and instance-level misalignment, we apply multi-modal filtering and retrieval to select the most semantically aligned prompt for each audio sample (EXPrompt Selector). A lightweight mapping network is then trained to adapt pre-trained text-to-image generation models to audio input. Extensive experiments on multiple audio classification datasets demonstrate that CatchPhrase improves audio-to-image alignment and consistently enhances generation quality by mitigating semantic misalignment.


【14】KuiSCIMA v2.0: Improved Baselines, Calibration, and Cross-Notation Generalization for Historical Chinese Music Notations in Jiang Kui's Baishidaoren Gequ
标题:KuiSCIMA v2.0:姜奎《白石真人歌曲》中中国历史音乐符号的改进基线、校准和交叉符号概括
链接:https://arxiv.org/abs/2507.18741

作者:Tristan Repolusk, Eduardo Veas
备注:International Conference on Document Analysis and Recognition. This preprint has not undergone any post-submission improvements or corrections. The Version of Record of this contribution is published in "19th International Conference on Document Analysis and Recognition (ICDAR 2025), Wuhan, China, September 16-21, 2025, Proceedings", and is available online at the External DOI field below
摘要:由于类别不平衡和训练数据有限,中国历史乐谱(如苏子谱和l“ul“upu)的光学音乐识别(OMR)面临着独特的挑战。本文介绍了姜夔《白石道人歌曲》自1202年以来的重要研究成果。在这项工作中,我们开发和评估一个字符识别模型的稀缺不平衡数据。我们通过将suzipu的字符错误率(CER)从10.4%降低到7.1%来改进先前的基线,尽管使用了77个高度不平衡的类,并且l\“ul\“upu的CER达到了0.9%。我们的模型优于人类转录器,平均人类CER为15.9%,最佳情况CER为7.6%。我们采用温度缩放来实现良好校准的模型,其预期校准误差(ECE)低于0.0162。使用留一个版本的交叉验证方法,我们确保了在五个历史版本中的稳健性能。此外,我们扩展了KuiSCIMA数据集,以包括来自白石道人歌曲的所有109个片段,包括suzipu,l\“ul\“upu和jianzipu标记。我们的研究结果推进了中国历史音乐的数字化和可访问性,促进了OMR的文化多样性,并将其适用性扩展到代表性不足的音乐传统。
摘要:Optical Music Recognition (OMR) for historical Chinese musical notations, such as suzipu and l\"ul\"upu, presents unique challenges due to high class imbalance and limited training data. This paper introduces significant advancements in OMR for Jiang Kui's influential collection Baishidaoren Gequ from 1202. In this work, we develop and evaluate a character recognition model for scarce imbalanced data. We improve upon previous baselines by reducing the Character Error Rate (CER) from 10.4% to 7.1% for suzipu, despite working with 77 highly imbalanced classes, and achieve a remarkable CER of 0.9% for l\"ul\"upu. Our models outperform human transcribers, with an average human CER of 15.9% and a best-case CER of 7.6%. We employ temperature scaling to achieve a well-calibrated model with an Expected Calibration Error (ECE) below 0.0162. Using a leave-one-edition-out cross-validation approach, we ensure robust performance across five historical editions. Additionally, we extend the KuiSCIMA dataset to include all 109 pieces from Baishidaoren Gequ, encompassing suzipu, l\"ul\"upu, and jianzipu notations. Our findings advance the digitization and accessibility of historical Chinese music, promoting cultural diversity in OMR and expanding its applicability to underrepresented music traditions.


【15】SCORE-SET: A dataset of GuitarPro files for Music Phrase Generation and Sequence Learning
标题:SCORE-SET:用于音乐短语生成和序列学习的GuitarPro文件数据集
链接:https://arxiv.org/abs/2507.18723

作者:Vishakh Begari
备注:6 pages, 6 figures
摘要:Guitar Pro指法文件(.gp5格式)的策划数据集,为涉及吉他音乐生成,序列建模和性能感知学习的任务量身定制。该数据集来自MAESTRO和Giantstrom中的音符,这些音符已被改编为节奏吉他曲目。这些音轨经过进一步处理,以包括吉他演奏中典型的各种表情设置,例如压音、滑动、颤音和手掌静音,以更好地反映现实世界吉他演奏的细微差别。
摘要:A curated dataset of Guitar Pro tablature files (.gp5 format), tailored for tasks involving guitar music generation, sequence modeling, and performance-aware learning is provided. The dataset is derived from MIDI notes in MAESTRO and GiantMIDI which have been adapted into rhythm guitar tracks. These tracks are further processed to include a variety of expression settings typical of guitar performance, such as bends, slides, vibrato, and palm muting, to better reflect the nuances of real-world guitar playing.


机器翻译由腾讯交互翻译提供,仅供参考