微信公众号:arXiv_Daily
cs.SD语音
【1】PROFASR-BENCH: A Benchmark for Context-Conditioned ASR in High-Stakes Professional Speech
标题:PROFASR-BENCH:高风险专业演讲中上下文条件的ASB基准
链接:https://arxiv.org/abs/2512.23686
备注:Benchmark dataset and evaluation suite. Data and code available at: https://huggingface.co/datasets/prdeepakbabu/ProfASR-Bench https://github.com/prdeepakbabu/ProfASR-Bench
摘要:专业环境中的自动语音识别(ASR)面临着现有基准未充分考虑的挑战:密集的领域术语,正式的寄存器变化以及对关键实体错误的近零容忍度。我们介绍ProfASR-Bench,这是一个专业的评估套件,适用于金融,医学,法律和技术等高风险应用。每个示例将自然语言提示(域提示和/或说话者简档)与实体丰富的目标话语配对,从而实现上下文条件识别的受控测量。语料库支持传统的ASR指标以及实体感知分数和按口音和性别的切片报告。在匹配的无上下文、配置文件、域+配置文件、Oracle和对抗性条件下,使用代表性家族Whisper(编码器-解码器ASR)和Qwen-Omni(音频语言模型),我们发现了一个一致的模式:轻量级文本上下文对平均单词错误率(WER)几乎没有影响,即使使用Oracle提示,对抗性提示也不会可靠地降低性能。我们称之为上下文利用率差距(CUG):目前的系统名义上是可验证的,但未充分利用现成的边信息。ProfASR-Bench提供了一个标准化的上下文阶梯,具有置信区间的实体和切片感知报告,以及一个可重复的测试平台,用于比较模型家族之间的融合策略。 数据集:https://huggingface.co/datasets/prdeepakbabu/ProfASR-Bench 代码:https://github.com/prdeepakbabu/ProfASR-Bench
摘要:Automatic Speech Recognition (ASR) in professional settings faces challenges that existing benchmarks underplay: dense domain terminology, formal register variation, and near-zero tolerance for critical entity errors. We present ProfASR-Bench, a professional-talk evaluation suite for high-stakes applications across finance, medicine, legal, and technology. Each example pairs a natural-language prompt (domain cue and/or speaker profile) with an entity-rich target utterance, enabling controlled measurement of context-conditioned recognition. The corpus supports conventional ASR metrics alongside entity-aware scores and slice-wise reporting by accent and gender. Using representative families Whisper (encoder-decoder ASR) and Qwen-Omni (audio language models) under matched no-context, profile, domain+profile, oracle, and adversarial conditions, we find a consistent pattern: lightweight textual context produces little to no change in average word error rate (WER), even with oracle prompts, and adversarial prompts do not reliably degrade performance. We term this the context-utilization gap (CUG): current systems are nominally promptable yet underuse readily available side information. ProfASR-Bench provides a standardized context ladder, entity- and slice-aware reporting with confidence intervals, and a reproducible testbed for comparing fusion strategies across model families. Dataset: https://huggingface.co/datasets/prdeepakbabu/ProfASR-Bench Code: https://github.com/prdeepakbabu/ProfASR-Bench
【2】Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models
标题:风格错觉:调查多轮口语模型中的说话风格退化和缓解
链接:https://arxiv.org/abs/2512.23578
备注:Work in progress
摘要:在本文中,我们表明,当口语模型(SLMs)被指示在一个特定的说话风格在一个多回合的对话开始时,他们不能保持所需的说话风格后,几个回合的互动,我们称之为风格健忘症的SLMs。我们专注于语言的说话风格,包括情感,口音,音量和语速。我们评估了三个专有和两个开源的SLM,证明这些模型都不能保持一致的说话风格时,指示这样做。我们进一步表明,当SLM被要求回忆的风格指令在以后的回合,他们可以回忆的风格指令,但他们无法表达它在整个会话。我们还表明,明确要求模型回忆风格指令可以部分减轻风格健忘症。此外,我们研究了各种提示策略,并发现SLM的斗争,按照所需的风格时,指令被放置在系统消息,而不是用户消息,这与系统提示的预期功能相矛盾。
摘要:In this paper, we show that when spoken language models (SLMs) are instructed to speak in a specific speaking style at the beginning of a multi-turn conversation, they cannot maintain the required speaking styles after several turns of interaction; we refer to this as the style amnesia of SLMs. We focus on paralinguistic speaking styles, including emotion, accent, volume, and speaking speed. We evaluate three proprietary and two open-source SLMs, demonstrating that none of these models can maintain a consistent speaking style when instructed to do so. We further show that when SLMs are asked to recall the style instruction in later turns, they can recall the style instruction, but they fail to express it throughout the conversation. We also show that explicitly asking the model to recall the style instruction can partially mitigate style amnesia. In addition, we examine various prompting strategies and find that SLMs struggle to follow the required style when the instruction is placed in system messages rather than user messages, which contradicts the intended function of system prompts.
【3】Mobile-Efficient Speech Emotion Recognition Using DistilHuBERT: A Cross-Corpus Validation Study
标题:基于DistilHuBERT的移动高效语音情感识别:跨语料库验证研究
链接:https://arxiv.org/abs/2512.23435
备注:5 pages, 2 tables, 1 figure. Submitted to IEEE conference
摘要:语音情感识别(SER)在移动应用中具有巨大的潜力,但部署仍然受到最先进的Transformer架构的计算需求的限制。本文提出了一种基于DistilHuBERT的移动高效SER系统,DistilHuBERT是一种经过蒸馏和8位量化的Transformer,与满量程Wav 2 Vec 2.0模型相比,其参数减少了92%,同时保持了具有竞争力的精度。我们对IEMOCAP数据集进行了严格的5重Leave-One-Session-Out(LOSO)交叉验证,以确保说话人独立性,并在CREMA-D上进行了跨语料库训练以增强泛化能力。使用CREMA-D进行跨语料库训练,加权准确度提高了1.2%,宏F1分数提高了1.4%,交叉方差降低了32%,中性类表现出最大的好处,F1分数提高了5.4%。我们的方法实现了61.4%的未加权精度,量化模型占用空间仅为23 MB,约占全尺寸基线性能的91%。RAVDESS上的跨语料库评估表明,表演情绪的戏剧性质导致预测按唤醒水平而不是效价聚类:由于高能量表达中的声学饱和,幸福与愤怒被系统地混淆了。尽管这种戏剧性效应将整体RAVDESS准确率降低到43.29%,但该模型保持了强大的唤醒检测,其中97%的愤怒回忆和64%的悲伤回忆。这些发现建立了模型大小和准确性之间的帕累托最优权衡,使资源受限的移动设备上的实际影响识别。
摘要:Speech Emotion Recognition (SER) has significant potential for mobile applications, yet deployment remains constrained by the computational demands of state-of-the-art transformer architectures. This paper presents a mobile-efficient SER system based on DistilHuBERT, a distilled and 8-bit quantized transformer that achieves 92% parameter reduction compared to full-scale Wav2Vec 2.0 models while maintaining competitive accuracy. We conduct a rigorous 5-fold Leave-One-Session-Out (LOSO) cross-validation on the IEMOCAP dataset to ensure speaker independence, augmented with cross-corpus training on CREMA-D to enhance generalization. Cross-corpus training with CREMA-D yields a 1.2% improvement in Weighted Accuracy, a 1.4% gain in Macro F1-score, and a 32% reduction in cross-fold variance, with the Neutral class showing the most substantial benefit at 5.4% F1-score improvement. Our approach achieves an Unweighted Accuracy of 61.4% with a quantized model footprint of only 23 MB, representing approximately 91% of full-scale baseline performance. Cross-corpus evaluation on RAVDESS reveals that the theatrical nature of acted emotions causes predictions to cluster by arousal level rather than valence: happiness is systematically confused with anger due to acoustic saturation in high-energy expressions. Despite this theatricality effect reducing overall RAVDESS accuracy to 43.29%, the model maintains robust arousal detection with 97% recall for anger and 64% for sadness. These findings establish a Pareto-optimal tradeoff between model size and accuracy, enabling practical affect recognition on resource-constrained mobile devices.
【4】Chord Recognition with Deep Learning
标题:利用深度学习进行和弦识别
链接:https://arxiv.org/abs/2512.22621
摘要:自深度学习在该领域出现以来,自动和弦识别的进展一直很缓慢。为了理解为什么,我对现有方法进行实验,并测试生成模型最近发展所支持的假设。研究结果表明,和弦分类器在罕见的和弦上表现不佳,音高增强可以提高准确性。从生成模型中提取的特征没有帮助,合成数据为未来的工作提供了一个令人兴奋的途径。最后,我通过提高模型输出的可解释性与节拍检测,报告了一些最好的结果在该领域,并提供定性分析。虽然自动和弦识别还有很多工作要做,但我希望这篇论文能为其他人指明一条道路。
摘要:Progress in automatic chord recognition has been slow since the advent of deep learning in the field. To understand why, I conduct experiments on existing methods and test hypotheses enabled by recent developments in generative models. Findings show that chord classifiers perform poorly on rare chords and that pitch augmentation boosts accuracy. Features extracted from generative models do not help and synthetic data presents an exciting avenue for future work. I conclude by improving the interpretability of model outputs with beat detection, reporting some of the best results in the field and providing qualitative analysis. Much work remains to solve automatic chord recognition, but I hope this thesis will chart a path for others to try.
【5】AudioGAN: A Compact and Efficient Framework for Real-Time High-Fidelity Text-to-Audio Generation
标题:AudioGAN:一个紧凑高效的实时高保真文本到音频生成框架
链接:https://arxiv.org/abs/2512.22166
备注:10 pages, 6 figures, Accepted to AES AIMLA 2025
摘要:文本到音频(TTA)生成可以通过降低生产成本和提高工作效率而使媒体行业受益匪浅。然而,大多数当前的TTA模型(主要是基于扩散的)遭受缓慢的推理速度和高计算成本。在本文中,我们介绍了AudioGAN,这是第一个成功的基于生成对抗网络(GANs)的TTA框架,它可以在单次通过中生成音频,从而降低模型复杂度和推理时间。为了克服训练GANs的固有困难,我们整合了多个对比损失,并提出了创新的组件单双三(SDT)注意力和时频交叉注意力(TF-CA)。在AudioCaps数据集上进行的大量实验表明,AudioGAN实现了最先进的性能,同时使用的参数减少了90%,运行速度提高了20倍,在不到一秒的时间内合成音频。这些结果使AudioGAN成为实时TTA的实用而强大的解决方案。
摘要:Text-to-audio (TTA) generation can significantly benefit the media industry by reducing production costs and enhancing work efficiency. However, most current TTA models (primarily diffusion-based) suffer from slow inference speeds and high computational costs. In this paper, we introduce AudioGAN, the first successful Generative Adversarial Networks (GANs)-based TTA framework that generates audio in a single pass, thereby reducing model complexity and inference time. To overcome the inherent difficulties in training GANs, we integrate multiple ,contrastive losses and propose innovative components Single-Double-Triple (SDT) Attention and Time-Frequency Cross-Attention (TF-CA). Extensive experiments on the AudioCaps dataset demonstrate that AudioGAN achieves state-of-the-art performance while using 90% fewer parameters and running 20 times faster, synthesizing audio in under one second. These results establish AudioGAN as a practical and powerful solution for real-time TTA.
【6】Marco-ASR: A Principled and Metric-Driven Framework for Fine-Tuning Large-Scale ASR Models for Domain Adaptation
标题:Marco-ASB:一个原则性和度量驱动的框架,用于微调大规模ASB模型以实现领域自适应
链接:https://arxiv.org/abs/2512.22165
备注:Technical Report
摘要:自动语音识别(ASR)模型在一般情况下已经取得了显着的准确性,但它们的性能往往下降,在特定领域的应用,由于数据不匹配和语言的变化。对于现代基于大型语言模型(LLM)的ASR系统来说,这一挑战被放大了,其庞大的规模和复杂的训练动态使得有效的微调变得非常重要。为了解决这一差距,本文提出了一个原则性和度量驱动的微调框架,以适应传统的和基于LLM的ASR模型的专业领域。该框架强调基于性能指标的学习率优化,结合特定领域的数据转换和增强。我们在最先进的模型上对我们的框架进行了经验评估,包括Whisper,Whisper-Turbo和Qwen 2-Audio,跨多域,多语言和多长度数据集。我们的研究结果不仅验证了所提出的框架,而且还建立了实用的协议,以提高特定领域的ASR性能,同时防止过拟合。
摘要:Automatic Speech Recognition (ASR) models have achieved remarkable accuracy in general settings, yet their performance often degrades in domain-specific applications due to data mismatch and linguistic variability. This challenge is amplified for modern Large Language Model (LLM)-based ASR systems, whose massive scale and complex training dynamics make effective fine-tuning non-trivial. To address this gap, this paper proposes a principled and metric-driven fine-tuning framework for adapting both traditional and LLM-based ASR models to specialized domains. The framework emphasizes learning rate optimization based on performance metrics, combined with domain-specific data transformation and augmentation. We empirically evaluate our framework on state-of-the-art models, including Whisper, Whisper-Turbo, and Qwen2-Audio, across multi-domain, multilingual, and multi-length datasets. Our results not only validate the proposed framework but also establish practical protocols for improving domain-specific ASR performance while preventing overfitting.
【7】A Robust framework for sound event localization and detection on real recordings
标题:用于对真实录音进行声音事件定位和检测的稳健框架
链接:https://arxiv.org/abs/2512.22156
备注:Technical Report submitted to DCASE 2022 Challenge Task 3 (Winner of the Judge's Award)
摘要:本技术报告介绍了提交给DCASE2022挑战任务3的系统:声音事件定位和检测(SELD)。该任务旨在检测声音事件的发生,并指定它们的类别,进而估计它们的位置。我们的系统利用了一个基于ResNet的模型下,提出了一个强大的框架SELD。为了保证真实世界声音场景的通用性能,我们设计了增强技术的总体框架,一个来自真实世界声音场景和仿真的混合数据集的管道,以及测试时间增强。增强技术和外部声源的利用使得能够训练不同的样本,并且通过保持批次中的真实记录样本的数量来保持足够训练真实世界上下文的机会。此外,我们设计了一个测试时间增强和基于聚类的模型集成方法来聚合可信的预测。实验结果表明,该框架下的模型优于基线方法,并在现实世界的录音中取得了竞争力的表现。
摘要:This technical report describes the systems submitted to the DCASE2022 challenge task 3: sound event localization and detection (SELD). The task aims to detect occurrences of sound events and specify their class, furthermore estimate their position. Our system utilizes a ResNet-based model under a proposed robust framework for SELD. To guarantee the generalized performance on the real-world sound scenes, we design the total framework with augmentation techniques, a pipeline of mixing datasets from real-world sound scenes and emulations, and test time augmentation. Augmentation techniques and exploitation of external sound sources enable training diverse samples and keeping the opportunity to train the real-world context enough by maintaining the number of the real recording samples in the batch. In addition, we design a test time augmentation and a clustering-based model ensemble method to aggregate confident predictions. Experimental results show that the model under a proposed framework outperforms the baseline methods and achieves competitive performance in real-world sound recordings.
【8】Rethinking Leveraging Pre-Trained Multi-Layer Representations for Speaker Verification
标题:重新思考利用预先训练的多层表示进行说话人验证
链接:https://arxiv.org/abs/2512.22148
备注:Accepted to Interspeech 2025
摘要:最近的说话人确认研究通过利用来自预训练的Transformer模型的逐层输出取得了显着的成功。然而,很少有人探索在静态加权平均之外聚合这些多层次特征的进步。我们提出了层注意池(Layer Attentive Pooling),这是一种新的策略,用于从预先训练的语音模型中聚合层间表示,以进行说话人验证。它从多个角度动态地评估每一层的重要性,并采用最大池化而不是平均化。此外,我们提出了一个轻量级的后端扬声器模型,包括语音识别和注意统计时间池(ASTP),从预训练的模型输出中提取扬声器嵌入。在VoxCeleb基准测试上的实验表明,我们的紧凑架构实现了最先进的性能,同时大大减少了训练时间。我们进一步分析了语音识别器的设计及其动态加权机制,以捕捉说话人特征。
摘要:Recent speaker verification studies have achieved notable success by leveraging layer-wise output from pre-trained Transformer models. However, few have explored the advancements in aggregating these multi-level features beyond the static weighted average. We present Layer Attentive Pooling (LAP), a novel strategy for aggregating inter-layer representations from pre-trained speech models for speaker verification. LAP assesses the significance of each layer from multiple perspectives time-dynamically, and employs max pooling instead of averaging. Additionally, we propose a lightweight backend speaker model comprising LAP and Attentive Statistical Temporal Pooling (ASTP) to extract speaker embeddings from pre-trained model output. Experiments on the VoxCeleb benchmark reveal that our compact architecture achieves state-of-the-art performance while greatly reducing the training time. We further analyzed LAP design and its dynamic weighting mechanism for capturing speaker characteristics.
【9】Geometry-Aware Optimization for Respiratory Sound Classification: Enhancing Sensitivity with SAM-Optimized Audio Spectrogram Transformers
标题:呼吸声分类的几何感知优化:使用ASM优化的音频频谱图变换器提高灵敏度
链接:https://arxiv.org/abs/2512.22564
备注:10 pages, 3 figures,2 tables
摘要:呼吸音分类受到ICBHI 2017等基准数据集的有限大小,高噪声水平和严重类别不平衡的阻碍。虽然基于Transformer的模型提供了强大的特征提取功能,但它们容易过拟合,并且在对此类受限的医疗数据进行训练时,往往会收敛到损失情况中的最小值。为了解决这个问题,我们介绍了一个框架,增强音频频谱图Transformer(AST)使用清晰度感知最小化(SAM)。我们的方法不是仅仅最小化训练损失,而是优化损失表面的几何形状,引导模型朝向更平坦的最小值,更好地推广到看不见的患者。我们还实现了一个加权抽样策略,以有效地处理类不平衡。我们的方法在ICBHI 2017数据集上获得了68.10%的最新得分,优于现有的CNN和混合基线。更重要的是,它达到了68.31%的灵敏度,这对于可靠的临床筛查来说是一个至关重要的进步。使用t-SNE和注意力图的进一步分析证实,该模型学习了鲁棒的、有区别的特征,而不是记忆背景噪声。
摘要:Respiratory sound classification is hindered by the limited size, high noise levels, and severe class imbalance of benchmark datasets like ICBHI 2017. While Transformer-based models offer powerful feature extraction capabilities, they are prone to overfitting and often converge to sharp minima in the loss landscape when trained on such constrained medical data. To address this, we introduce a framework that enhances the Audio Spectrogram Transformer (AST) using Sharpness-Aware Minimization (SAM). Instead of merely minimizing the training loss, our approach optimizes the geometry of the loss surface, guiding the model toward flatter minima that generalize better to unseen patients. We also implement a weighted sampling strategy to handle class imbalance effectively. Our method achieves a state-of-the-art score of 68.10% on the ICBHI 2017 dataset, outperforming existing CNN and hybrid baselines. More importantly, it reaches a sensitivity of 68.31%, a crucial improvement for reliable clinical screening. Further analysis using t-SNE and attention maps confirms that the model learns robust, discriminative features rather than memorizing background noise.
【10】EEG-to-Voice Decoding of Spoken and Imagined speech Using Non-Invasive EEG
标题:使用无创脑电对口语和想象语音进行脑电到语音解码
链接:https://arxiv.org/abs/2512.22146
备注:20 pages, 7 figures, 4 tables
摘要:从神经信号中恢复语音通信是脑机接口研究的中心目标,但由于有限的空间分辨率,对噪声的敏感性以及想象语音中缺乏时间对齐的声学目标,基于EEG的语音重建仍然具有挑战性。在这项研究中,我们提出了一个EEG语音范式,直接重建语音从非侵入性的EEG信号,而无需动态时间规整(DTW)或明确的时间对齐。所提出的管道生成梅尔频谱从EEG在一个开环的方式使用特定于主题的发生器,然后由预训练的声码器和自动语音识别(ASR)模块合成语音波形和解码文本。分别为口语和想象语音训练生成器,并通过对口语进行预训练和对想象语音进行自适应来应用基于迁移学习的域自适应。可选地应用基于最小语言模型的校正模块来校正有限的ASR错误,同时保留语义结构。该框架进行了评估,在2秒和4秒的语音条件下,使用声学水平的指标(PCC,RMSE,MCD)和语言水平的指标(CER,WER)。稳定的声学重建和可比的语言准确性,观察到口头讲话和想象的讲话。虽然声学相似性下降较长的话语,文本级的解码性能在很大程度上得到保留,和字位置分析显示,解码错误朝着后面的句子部分略有增加。基于语言模型的校正一致地减少CER和WER,而不引入语义失真。这些结果表明,直接,开环EEG语音重建口语和想象的语音没有明确的时间对齐的可行性。
摘要:Restoring speech communication from neural signals is a central goal of brain-computer interface research, yet EEG-based speech reconstruction remains challenging due to limited spatial resolution, susceptibility to noise, and the absence of temporally aligned acoustic targets in imagined speech. In this study, we propose an EEG-to-Voice paradigm that directly reconstructs speech from non-invasive EEG signals without dynamic time warping (DTW) or explicit temporal alignment. The proposed pipeline generates mel-spectrograms from EEG in an open-loop manner using a subject-specific generator, followed by pretrained vocoder and automatic speech recognition (ASR) modules to synthesize speech waveforms and decode text. Separate generators were trained for spoken speech and imagined speech, and transfer learning-based domain adaptation was applied by pretraining on spoken speech and adapting to imagined speech. A minimal language model-based correction module was optionally applied to correct limited ASR errors while preserving semantic structure. The framework was evaluated under 2 s and 4 s speech conditions using acoustic-level metrics (PCC, RMSE, MCD) and linguistic-level metrics (CER, WER). Stable acoustic reconstruction and comparable linguistic accuracy were observed for both spoken speech and imagined speech. While acoustic similarity decreased for longer utterances, text-level decoding performance was largely preserved, and word-position analysis revealed a mild increase in decoding errors toward later parts of sentences. The language model-based correction consistently reduced CER and WER without introducing semantic distortion. These results demonstrate the feasibility of direct, open-loop EEG-to-Voice reconstruction for spoken speech and imagined speech without explicit temporal alignment.
【1】Single Channel Blind Dereverberation of Speech Signals
标题:语音信号的单通道盲去回响
链接:https://arxiv.org/abs/2512.23322
摘要:语音信号的去混响是语音处理中最重要的问题之一。在目前的工作中,目标是理解和实现去混响技术,旨在增强混响语音信号的幅度谱图,以消除混响的影响。提出了一种从混响语音谱图中提取纯净语音谱图的方法。这是通过非负矩阵因子反卷积(NMFD)实现的。此外,这种方法扩展使用语音幅度谱图的NMF表示。为了利用时间依赖性,将基于卷积NMF的表示和帧堆叠模型合并到用于语音的NMFD框架中。本文还提出了一种新的混响去噪方法,即将NMFD应用于混响幅度谱图的激活矩阵。最后,基于两个关键的客观指标- PESQ和倒谱失真,使用TIMIT数据库中的句子录音和Reverb 2014挑战赛中记录的房间脉冲响应,对所列技术的性能进行了比较分析。虽然我们能够定性地验证文献中关于这些技术的声明,但无法匹配确切的结果。新的方法,因为它是建议,提供了改进的定量指标,但不一致
摘要:Dereverberation of recorded speech signals is one of the most pertinent problems in speech processing. In the present work, the objective is to understand and implement dereverberation techniques that aim at enhancing the magnitude spectrogram of reverberant speech signals to remove the reverberant effects introduced. An approach to estimate a clean speech spectrogram from the reverberant speech spectrogram is proposed. This is achieved through non-negative matrix factor deconvolution(NMFD). Further, this approach is extended using the NMF representation for speech magnitude spectrograms. To exploit temporal dependencies, a convolutive NMF-based representation and a frame-stacked model are incorporated into the NMFD framework for speech. A novel approach for dereverberation by applying NMFD to the activation matrix of the reverberated magnitude spectrogram is also proposed. Finally, a comparative analysis of the performance of the listed techniques, using sentence recordings from the TIMIT database and recorded room impulse responses from the Reverb 2014 challenge, is presented based on two key objective measures - PESQ and Cepstral Distortion.\\ Although we were qualitatively able to verify the claims made in literature regarding these techniques, exact results could not be matched. The novel approach, as it is suggested, provides improvement in quantitative metrics, but is not consistent
【2】Flow2GAN: Hybrid Flow Matching and GAN with Multi-Resolution Network for Few-step High-Fidelity Audio Generation
标题:Flow 2GAN:具有多分辨率网络的混合流匹配和GAN,用于少步高保真音频生成
链接:https://arxiv.org/abs/2512.23278
摘要:现有的主要音频生成方法包括生成对抗网络(GANs)和基于扩散的方法,如流匹配。GANs在训练过程中收敛缓慢,可能会出现模式崩溃,而扩散方法需要多步推理,这会带来相当大的计算开销。在这项工作中,我们引入了Flow 2 GAN,这是一个两阶段的框架,它将用于学习生成能力的Flow Matching训练与用于高效的几步推理的GAN微调相结合。具体来说,鉴于音频的独特属性,我们首先通过以下方式改进音频建模的流匹配:1)将目标重新定义为端点估计,避免涉及空区域时的速度估计困难; 2)应用基于频谱能量的损失缩放来强调感知上突出的安静区域。在这些流匹配适应的基础上,我们证明了轻量级GAN微调的进一步阶段使我们能够获得产生高质量音频的一步生成器。此外,我们开发了一个多分支网络架构,处理傅立叶系数在不同的时间-频率分辨率,这提高了建模能力相比,以前的单分辨率设计。实验结果表明,我们的Flow 2GAN可以从Mel频谱图或离散音频令牌中生成高保真音频,比现有的最先进的基于GAN和基于流匹配的方法实现更好的质量效率权衡。在线演示示例可在https://flow2gan.github.io上获得,源代码可在https://github.com/k2-fsa/Flow2GAN上发布。
摘要:Existing dominant methods for audio generation include Generative Adversarial Networks (GANs) and diffusion-based methods like Flow Matching. GANs suffer from slow convergence and potential mode collapse during training, while diffusion methods require multi-step inference that introduces considerable computational overhead. In this work, we introduce Flow2GAN, a two-stage framework that combines Flow Matching training for learning generative capabilities with GAN fine-tuning for efficient few-step inference. Specifically, given audio's unique properties, we first improve Flow Matching for audio modeling through: 1) reformulating the objective as endpoint estimation, avoiding velocity estimation difficulties when involving empty regions; 2) applying spectral energy-based loss scaling to emphasize perceptually salient quieter regions. Building on these Flow Matching adaptations, we demonstrate that a further stage of lightweight GAN fine-tuning enables us to obtain one-step generator that produces high-quality audio. In addition, we develop a multi-branch network architecture that processes Fourier coefficients at different time-frequency resolutions, which improves the modeling capabilities compared to prior single-resolution designs. Experimental results indicate that our Flow2GAN delivers high-fidelity audio generation from Mel-spectrograms or discrete audio tokens, achieving better quality-efficiency trade-offs than existing state-of-the-art GAN-based and Flow Matching-based methods. Online demo samples are available at https://flow2gan.github.io, and the source code is released at https://github.com/k2-fsa/Flow2GAN.
【3】Spatial Interpolation of Room Impulse Responses based on Deeper Physics-Informed Neural Networks with Residual Connections
标题:基于具有剩余连接的更深物理信息神经网络的房间脉冲响应空间内插
链接:https://arxiv.org/abs/2512.22915
备注:This work has been submitted to the IEEE for possible publication
摘要:在线性时不变假设下,房间脉冲响应(RIR)表征了声音在房间中从扬声器到麦克风的传播。从有限数量的测量点估计RIR对于声音传播分析和可视化至关重要。最近引入了物理信息神经网络(PINN),通过将管理物理定律嵌入深度学习模型来精确估计RIR;然而,网络深度的作用尚未得到系统研究。在这项研究中,我们开发了一个更深层次的PINN架构与剩余连接,并分析了网络深度如何影响估计性能。我们进一步比较了激活函数,包括双曲正切和正弦激活。我们的研究结果表明,与正弦激活的残留PINN实现了最高的精度内插和外推的RIR。此外,所提出的架构能够随着深度的增加进行稳定的训练,并在估计反射分量方面产生显着的改进。这些结果提供了实际的指导方针,设计深和稳定的PINN的声学逆问题。
摘要:The room impulse response (RIR) characterizes sound propagation in a room from a loudspeaker to a microphone under the linear time-invariant assumption. Estimating RIRs from a limited number of measurement points is crucial for sound propagation analysis and visualization. Physics-informed neural networks (PINNs) have recently been introduced for accurate RIR estimation by embedding governing physical laws into deep learning models; however, the role of network depth has not been systematically investigated. In this study, we developed a deeper PINN architecture with residual connections and analyzed how network depth affects estimation performance. We further compared activation functions, including tanh and sinusoidal activations. Our results indicate that the residual PINN with sinusoidal activations achieves the highest accuracy for both interpolation and extrapolation of RIRs. Moreover, the proposed architecture enables stable training as the depth increases and yields notable improvements in estimating reflection components. These results provide practical guidelines for designing deep and stable PINNs for acoustic-inverse problems.
【4】Geometry-Aware Optimization for Respiratory Sound Classification: Enhancing Sensitivity with SAM-Optimized Audio Spectrogram Transformers
标题:呼吸声分类的几何感知优化:使用ASM优化的音频频谱图变换器提高灵敏度
链接:https://arxiv.org/abs/2512.22564
备注:10 pages, 3 figures,2 tables
摘要:呼吸音分类受到ICBHI 2017等基准数据集的有限大小,高噪声水平和严重类别不平衡的阻碍。虽然基于Transformer的模型提供了强大的特征提取功能,但它们容易过拟合,并且在对此类受限的医疗数据进行训练时,往往会收敛到损失情况中的最小值。为了解决这个问题,我们介绍了一个框架,增强音频频谱图Transformer(AST)使用清晰度感知最小化(SAM)。我们的方法不是仅仅最小化训练损失,而是优化损失表面的几何形状,引导模型朝向更平坦的最小值,更好地推广到看不见的患者。我们还实现了一个加权抽样策略,以有效地处理类不平衡。我们的方法在ICBHI 2017数据集上获得了68.10%的最新得分,优于现有的CNN和混合基线。更重要的是,它达到了68.31%的灵敏度,这对于可靠的临床筛查来说是一个至关重要的进步。使用t-SNE和注意力图的进一步分析证实,该模型学习了鲁棒的、有区别的特征,而不是记忆背景噪声。
摘要:Respiratory sound classification is hindered by the limited size, high noise levels, and severe class imbalance of benchmark datasets like ICBHI 2017. While Transformer-based models offer powerful feature extraction capabilities, they are prone to overfitting and often converge to sharp minima in the loss landscape when trained on such constrained medical data. To address this, we introduce a framework that enhances the Audio Spectrogram Transformer (AST) using Sharpness-Aware Minimization (SAM). Instead of merely minimizing the training loss, our approach optimizes the geometry of the loss surface, guiding the model toward flatter minima that generalize better to unseen patients. We also implement a weighted sampling strategy to handle class imbalance effectively. Our method achieves a state-of-the-art score of 68.10% on the ICBHI 2017 dataset, outperforming existing CNN and hybrid baselines. More importantly, it reaches a sensitivity of 68.31%, a crucial improvement for reliable clinical screening. Further analysis using t-SNE and attention maps confirms that the model learns robust, discriminative features rather than memorizing background noise.
【5】AudioGAN: A Compact and Efficient Framework for Real-Time High-Fidelity Text-to-Audio Generation
标题:AudioGAN:一个紧凑高效的实时高保真文本到音频生成框架
链接:https://arxiv.org/abs/2512.22166
备注:10 pages, 6 figures, Accepted to AES AIMLA 2025
摘要:文本到音频(TTA)生成可以通过降低生产成本和提高工作效率而使媒体行业受益匪浅。然而,大多数当前的TTA模型(主要是基于扩散的)遭受缓慢的推理速度和高计算成本。在本文中,我们介绍了AudioGAN,这是第一个成功的基于生成对抗网络(GANs)的TTA框架,它可以在单次通过中生成音频,从而降低模型复杂度和推理时间。为了克服训练GANs的固有困难,我们整合了多个对比损失,并提出了创新的组件单双三(SDT)注意力和时频交叉注意力(TF-CA)。在AudioCaps数据集上进行的大量实验表明,AudioGAN实现了最先进的性能,同时使用的参数减少了90%,运行速度提高了20倍,在不到一秒的时间内合成音频。这些结果使AudioGAN成为实时TTA的实用而强大的解决方案。
摘要:Text-to-audio (TTA) generation can significantly benefit the media industry by reducing production costs and enhancing work efficiency. However, most current TTA models (primarily diffusion-based) suffer from slow inference speeds and high computational costs. In this paper, we introduce AudioGAN, the first successful Generative Adversarial Networks (GANs)-based TTA framework that generates audio in a single pass, thereby reducing model complexity and inference time. To overcome the inherent difficulties in training GANs, we integrate multiple ,contrastive losses and propose innovative components Single-Double-Triple (SDT) Attention and Time-Frequency Cross-Attention (TF-CA). Extensive experiments on the AudioCaps dataset demonstrate that AudioGAN achieves state-of-the-art performance while using 90% fewer parameters and running 20 times faster, synthesizing audio in under one second. These results establish AudioGAN as a practical and powerful solution for real-time TTA.
机器翻译由腾讯交互翻译提供,仅供参考
