今日论文合集:cs.SD语音10篇,eess.AS音频处理10篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 BowelRCNN: Region-based Convolutional Neural Network System for Bowel  Sound Auscultation
标题: BowelRCNN:用于肠道音听诊的基于区域的卷积神经网络系统
链接:https://arxiv.org/abs/2504.08659
作者: Igor Matynia,  Robert Nowak 
备注:10 pages, 3 figures
摘要:表示肠道活动检测的声音事件是具有识别胃肠道状况的潜力的诊断工具。本文介绍了BowelRCNN,这是一种新型的肠音检测系统,它使用音频记录,频谱图分析和基于区域的卷积神经网络(RCNN)架构。该系统在从19名患者收集的真实记录数据集上进行了训练和验证,包括60分钟的准备和注释的音频数据。BowelRCNN的分类准确率为96%,F1评分为71%。这项研究强调了使用CNN架构进行肠音听诊的可行性,实现了与递归卷积方法相当的结果。
摘要:Sound events representing intestinal activity detection is a diagnostic tool with potential to identify gastrointestinal conditions. This article introduces BowelRCNN, a novel bowel sound detection system that uses audio recording, spectrogram analysys and region-based convolutional neural network (RCNN) architecture. The system was trained and validated on a real recording dataset gathered from 19 patients, comprising 60 minutes of prepared and annotated audio data. BowelRCNN achieved a classification accuracy of 96% and an F1 score of 71%. This research highlights the feasibility of using CNN architectures for bowel sound auscultation, achieving results comparable to those of recurrent-convolutional methods.


【2】 On The Landscape of Spoken Language Models: A Comprehensive Survey

标题: 口语模型的格局:全面调查
链接:https://arxiv.org/abs/2504.08528
作者: Siddhant Arora,  Kai-Wei Chang,  Chung-Ming Chien,  Yifan Peng,  Haibin Wu,  Yossi Adi,  Emmanuel Dupoux,  Hung-Yi Lee,  Karen Livescu,  Shinji Watanabe 
摘要:口语处理领域正在经历从训练定制的、特定于任务的模型向使用和优化充当通用语音处理系统的口语模型(SLM)的转变。这一趋势类似于(文本)自然语言处理领域向通用语言模型的发展。SLM既包括语音的“纯”语言模型--标记化语音序列的分布模型--也包括将语音编码器与文本语言模型相结合的模型,通常包括口头和书面输入或输出。这一领域的工作多种多样,有一系列术语和评价背景。本文的目的是通过一个统一的文献调查领域的发展背景下,最近的工作,有助于提高对可持续土地管理的理解。我们的调查按模型架构、培训和评估选择对这一领域的工作进行了分类,并描述了未来工作的一些关键挑战和方向。
摘要:The field of spoken language processing is undergoing a shift from training custom-built, task-specific models toward using and optimizing spoken language models (SLMs) which act as universal speech processing systems. This trend is similar to the progression toward universal language models that has taken place in the field of (text) natural language processing. SLMs include both "pure" language models of speech -- models of the distribution of tokenized speech sequences -- and models that combine speech encoders with text language models, often including both spoken and written input or output. Work in this area is very diverse, with a range of terminology and evaluation settings. This paper aims to contribute an improved understanding of SLMs via a unifying literature survey of recent work in the context of the evolution of the field. Our survey categorizes the work in this area by model architecture, training, and evaluation choices, and describes some key challenges and directions for future work.


【3】 On the Design of Diffusion-based Neural Speech Codecs

标题: 基于扩散的神经语音编解码器的设计
链接:https://arxiv.org/abs/2504.08470
作者: Pietro Foti,  Andreas Brendel 
摘要:最近,作为生成模型训练的神经语音编解码器(NSC)在低比特率下显示出比传统编解码器更优越的性能。尽管大多数最先进的NSC都被训练为生成对抗网络(GANs),但扩散模型(DM)(一类最近的生成模型)代表了一种有前途的替代方案,因为它们在图像生成方面的性能相对于GANs。因此,DM已经成功地应用于各种其他音频生成应用中的音频和语音编码。然而,基于扩散的神经干细胞的设计尚未被系统地探索。我们解决这个问题,提供一个全面的分析扩散为基础的神经干细胞分为三个贡献。首先,我们提出了一个分类的基础上的条件和输出域的DM。这个简单的概念框架,使我们能够定义一个设计空间的扩散为基础的神经干细胞,并在文献中现有的方法分配一个类别。其次,我们系统地研究未开发的设计,通过创建和评估新的扩散为基础的神经干细胞的概念框架内。最后,我们通过客观指标和主观听力测试将所提出的模型与现有的GAN和DM基线进行比较。
摘要:Recently, neural speech codecs (NSCs) trained as generative models have shown superior performance compared to conventional codecs at low bitrates. Although most state-of-the-art NSCs are trained as Generative Adversarial Networks (GANs), Diffusion Models (DMs), a recent class of generative models, represent a promising alternative due to their superior performance in image generation relative to GANs. Consequently, DMs have been successfully applied for audio and speech coding among various other audio generation applications. However, the design of diffusion-based NSCs has not yet been explored in a systematic way. We address this by providing a comprehensive analysis of diffusion-based NSCs divided into three contributions. First, we propose a categorization based on the conditioning and output domains of the DM. This simple conceptual framework allows us to define a design space for diffusion-based NSCs and to assign a category to existing approaches in the literature. Second, we systematically investigate unexplored designs by creating and evaluating new diffusion-based NSCs within the conceptual framework. Finally, we compare the proposed models to existing GAN and DM baselines through objective metrics and subjective listening tests.


【4】 Passive Underwater Acoustic Signal Separation based on Feature  Decoupling Dual-path Network

标题: 基于特征解耦双径网络的被动水声信号分离
链接:https://arxiv.org/abs/2504.08371
作者: Yucheng Liu,  Longyu Jiang 
备注:10pages,4 figures
摘要:被动水声领域的信号分离在很大程度上依赖于深度学习技术来隔离船舶辐射噪声。然而,该领域常用的分离网络源于语音分离应用,可能事先没有充分考虑水声的独特方面,例如不同传播介质、信号频率和调制特性的影响。这一疏忽突出了需要量身定制的方法,考虑到水下声音传播的具体特点。本文提出了一种新的时域网络设计,采用双路径模型和特征解耦的方法来分离船舶辐射噪声。混合信号的特征被转换到一个空间中,在这个空间中它们表现出更大的独立性,每个维度的重要性被解耦。随后,在分离层中采用局部和全局注意机制的融合。与其他流行的网络模型相比,广泛的比较显示了该方法的有效性,其在ShipsEar和DeepShip数据集中的性能证明了这一点。
摘要:Signal separation in the passive underwater acoustic domain has heavily relied on deep learning techniques to isolate ship radiated noise. However, the separation networks commonly used in this domain stem from speech separation applications and may not fully consider the unique aspects of underwater acoustics beforehand, such as the influence of different propagation media, signal frequencies and modulation characteristics. This oversight highlights the need for tailored approaches that account for the specific characteristics of underwater sound propagation. This study introduces a novel temporal network designed to separate ship radiated noise by employing a dual-path model and a feature decoupling approach. The mixed signals' features are transformed into a space where they exhibit greater independence, with each dimension's significance decoupled. Subsequently, a fusion of local and global attention mechanisms is employed in the separation layer. Extensive comparisons showcase the effectiveness of this method when compared to other prevalent network models, as evidenced by its performance in the ShipsEar and DeepShip datasets.


【5】 Location-Oriented Sound Event Localization and Detection with Spatial  Mapping and Regression Localization

标题: 利用空间映射和回归定位的面向位置的声音事件定位和检测
链接:https://arxiv.org/abs/2504.08365
作者: Xueping Zhang,  Yaxiong Chen,  Ruilin Yao,  Yunfei Zi,  Shengwu Xiong 
摘要:声音事件定位与检测(SELD)将声音事件检测(SED)与相应的到达方向(DOA)相结合。目前,采用的面向事件的多声道方法由于声道数量的限制,影响了多声道环境下的通用性。为了提高在复音环境中的通用性,我们提出了空间映射和回归定位SELD(SMRL-SELD)。SMRL-SELD分割三维空间,将其映射到二维平面,并提出了一种新的回归定位损失,以帮助结果收敛到相应的事件的位置。SMRL-SELD是面向位置的,允许模型基于方向学习事件特征。因此,该方法使模型能够处理复调声音,而不管重叠事件的数量。我们在STARSS 23和STARSS 22数据集上进行了实验,我们提出的SMRL-SELD在整体评估和复调环境中优于现有的SELD方法。
摘要:Sound Event Localization and Detection (SELD) combines the Sound Event Detection (SED) with the corresponding Direction Of Arrival (DOA). Recently, adopted event oriented multi-track methods affect the generality in polyphonic environments due to the limitation of the number of tracks. To enhance the generality in polyphonic environments, we propose Spatial Mapping and Regression Localization for SELD (SMRL-SELD). SMRL-SELD segments the 3D spatial space, mapping it to a 2D plane, and a new regression localization loss is proposed to help the results converge toward the location of the corresponding event. SMRL-SELD is location-oriented, allowing the model to learn event features based on orientation. Thus, the method enables the model to process polyphonic sounds regardless of the number of overlapping events. We conducted experiments on STARSS23 and STARSS22 datasets and our proposed SMRL-SELD outperforms the existing SELD methods in overall evaluation and polyphony environments.


【6】 Generalized Multilingual Text-to-Speech Generation with Language-Aware  Style Adaptation

标题: 具有地理感知风格适应的广义多语言文本到语音生成
链接:https://arxiv.org/abs/2504.08274
作者: Haowei Lou,  Hye-young Paik,  Sheng Li,  Wen Hu,  Lina Yao 
摘要:文本到语音(TTS)模型可以通过将音素转换为波形来生成跨多种语言的自然的、类似人类的语音。然而,由于音素词汇的差异以及不同语言之间韵律和说话风格的差异,多语言TTS仍然具有挑战性。现有的方法要么为每种语言训练单独的模型,以增加计算资源为代价实现高性能,要么为多种语言使用统一的模型,努力捕捉细粒度的,语言特定的风格变化。在这项工作中,我们提出了LanStyleTTS,一个非自回归,语言感知的风格自适应TTS框架,该框架能够识别音素表示,并能够跨语言进行细粒度的音素级风格控制。这种设计支持一个统一的多语言TTS模型,能够产生准确和高质量的语音,而不需要训练语言特定的模型。我们评估LanStyleTTS集成它与几个国家的最先进的非自回归TTS架构。结果表明,在不同的模型骨干一致的性能改进。此外,我们调查了一系列的声学特征表示,包括梅尔频谱图和自动编码器派生的潜在功能。我们的实验表明,潜在的编码可以显着降低模型的大小和计算成本,同时保持高质量的语音生成。
摘要:Text-to-Speech (TTS) models can generate natural, human-like speech across multiple languages by transforming phonemes into waveforms. However, multilingual TTS remains challenging due to discrepancies in phoneme vocabularies and variations in prosody and speaking style across languages. Existing approaches either train separate models for each language, which achieve high performance at the cost of increased computational resources, or use a unified model for multiple languages that struggles to capture fine-grained, language-specific style variations. In this work, we propose LanStyleTTS, a non-autoregressive, language-aware style adaptive TTS framework that standardizes phoneme representations and enables fine-grained, phoneme-level style control across languages. This design supports a unified multilingual TTS model capable of producing accurate and high-quality speech without the need to train language-specific models. We evaluate LanStyleTTS by integrating it with several state-of-the-art non-autoregressive TTS architectures. Results show consistent performance improvements across different model backbones. Furthermore, we investigate a range of acoustic feature representations, including mel-spectrograms and autoencoder-derived latent features. Our experiments demonstrate that latent encodings can significantly reduce model size and computational cost while preserving high-quality speech generation.


【7】 From Speech to Summary: A Comprehensive Survey of Speech Summarization

标题: 从演讲到总结:演讲总结综述
链接:https://arxiv.org/abs/2504.08024
作者: Fabian Retkowski,  Maike Züfle,  Andreas Sudmann,  Dinah Pfau,  Jan Niehues,  Alexander Waibel 
摘要:语音摘要已成为有效管理和访问不断增长的语音和视听内容的重要工具。然而,尽管其重要性越来越大,语音摘要仍然没有明确定义,并与几个研究领域交叉,包括语音识别,文本摘要和会议摘要等特定应用。该调查不仅检查了现有的数据集和评估方法,这对于评估摘要方法的有效性至关重要,而且还综合了该领域的最新发展,突出了从传统系统到高级模型的转变,如微调级联架构和端到端解决方案。
摘要:Speech summarization has become an essential tool for efficiently managing and accessing the growing volume of spoken and audiovisual content. However, despite its increasing importance, speech summarization is still not clearly defined and intersects with several research areas, including speech recognition, text summarization, and specific applications like meeting summarization. This survey not only examines existing datasets and evaluation methodologies, which are crucial for assessing the effectiveness of summarization approaches but also synthesizes recent developments in the field, highlighting the shift from traditional systems to advanced models like fine-tuned cascaded architectures and end-to-end solutions.


【8】 Reverberation-based Features for Sound Event Localization and Detection  with Distance Estimation

标题: 基于回响的声音事件定位和检测特征以及距离估计
链接:https://arxiv.org/abs/2504.08644
作者: Davide Berghi,  Philip J. B. Jackson 
摘要:声音事件定位和检测(SELD)涉及随着时间的推移预测活动的声音事件类别,同时估计它们的位置。SELD中的定位子任务通常被视为一个波达方向估计问题,忽略了源距离。直到最近,SELD才通过结合距离估计扩展到3D,从而能够预测3D空间中的声音事件位置(3D SELD)。然而,现有的方法缺乏设计用于距离估计的输入特征。我们认为,混响编码有价值的信息,这项任务。本文介绍了两种基于混响的3D SELD的新特征格式:一种使用直接混响比(DRR),另一种利用信号自相关为模型提供早期反射的见解。对合成数据进行预训练可以提高相对距离误差(RDE)和整体SELD分数,基于自相关的特征可以将STARSS 23数据集上的RDE降低超过3个百分点。提取特征的代码可在github.com/dberghi/SELD-distance-features上获得。
摘要:Sound event localization and detection (SELD) involves predicting active sound event classes over time while estimating their positions. The localization subtask in SELD is usually treated as a direction of arrival estimation problem, ignoring source distance. Only recently, SELD was extended to 3D by incorporating distance estimation, enabling the prediction of sound event positions in 3D space (3D SELD). However, existing methods lack input features designed for distance estimation. We argue that reverberation encodes valuable information for this task. This paper introduces two novel feature formats for 3D SELD based on reverberation: one using direct-to-reverberant ratio (DRR) and another leveraging signal autocorrelation to provide the model with insights into early reflections. Pre-training on synthetic data improves relative distance error (RDE) and overall SELD score, with autocorrelation-based features reducing RDE by over 3 percentage points on the STARSS23 dataset. The code to extract the features is available at github.com/dberghi/SELD-distance-features.


【9】 TorchFX: A modern approach to Audio DSP with PyTorch and GPU  acceleration

标题: TorchFX:具有PyTorch和图形处理器加速的音频DSP现代方法
链接:https://arxiv.org/abs/2504.08624
作者: Matteo Spanio,  Antonio Rodà 
备注:Submitted to DAFx 2025
摘要:音频信号日益增长的复杂性和实时处理需求需要优化算法,以利用图形处理单元(GPU)的计算能力。现有的数字信号处理(DSP)库通常无法提供必要的效率和灵活性,特别是在集成人工智能(AI)模型方面。作为回应,我们引入了TorchFX:一个用于DSP的GPU加速Python库,专门设计用于促进复杂的音频信号处理。TorchFX构建在PyTorch框架之上,提供了一个面向对象的界面,可以模拟torchaudio的可用性,通过新颖的管道操作符增强功能,实现直观的过滤器链接。该库提供了一套全面的有限脉冲响应(FIR)和无限脉冲响应(IIR)滤波器,重点关注多通道音频文件,从而促进了DSP和基于AI的方法的集成。我们的基准测试结果表明,与SciPy等传统库相比,效率有了显著提高,特别是在多通道环境中。尽管目前在GPU兼容性方面存在限制,但正在进行的开发有望提供更广泛的支持和实时处理能力。TorchFX的目标是成为社区的有用工具,通过GPU加速为DSP的创新和进步做出贡献。TorchFX在GitHub上公开发布,网址为https://github.com/matteospanio/torchfx。
摘要:The burgeoning complexity and real-time processing demands of audio signals necessitate optimized algorithms that harness the computational prowess of Graphics Processing Units (GPUs). Existing Digital Signal Processing (DSP) libraries often fall short in delivering the requisite efficiency and flexibility, particularly in integrating Artificial Intelligence (AI) models. In response, we introduce TorchFX: a GPU-accelerated Python library for DSP, specifically engineered to facilitate sophisticated audio signal processing. Built atop the PyTorch framework, TorchFX offers an Object-Oriented interface that emulates the usability of torchaudio, enhancing functionality with a novel pipe operator for intuitive filter chaining. This library provides a comprehensive suite of Finite Impulse Response (FIR) and Infinite Impulse Response (IIR) filters, with a focus on multichannel audio files, thus facilitating the integration of DSP and AI-based approaches. Our benchmarking results demonstrate significant efficiency gains over traditional libraries like SciPy, particularly in multichannel contexts. Despite current limitations in GPU compatibility, ongoing developments promise broader support and real-time processing capabilities. TorchFX aims to become a useful tool for the community, contributing to innovation and progress in DSP with GPU acceleration. TorchFX is publicly available on GitHub at https://github.com/matteospanio/torchfx.


【10】 Mitigating Timbre Leakage with Universal Semantic Mapping Residual Block  for Voice Conversion

标题: 用通用语义映射剩余块缓解语音转换的音色泄漏
链接:https://arxiv.org/abs/2504.08524
作者: Na Li,  Chuke Wang,  Yu Gu,  Zhifeng Li 
摘要:语音转换(VC)通过保留内容将源语音转换为目标语音。然而,来自源说话者的音色信息固有地嵌入在内容表示中,导致显著的音色泄漏并且降低与目标说话者的相似性。为了解决这个问题,我们将残差块引入内容提取器。剩余块由两个加权分支组成:1)基于通用语义词典的内容特征重表达(CFR)模块,提供无音色的内容表示。2)跳过与原始内容层的连接,提供补充的细粒度信息。在CFR模块中,通用语义词典中的每个词典条目表示使用来自多个说话者的语音统计地计算的音素类,从而创建稳定的、与说话者无关的语义集。我们引入了一种CFR方法,通过使用相应的音素后验子作为权重将每个内容帧表示为字典条目的加权线性组合来获得无音色的内容表示。在各种VC框架上的大量实验表明,我们的方法有效地减轻了音色泄漏,并显着提高了与目标说话人的相似性。
摘要:Voice conversion (VC) transforms source speech into a target voice by preserving the content. However, timbre information from the source speaker is inherently embedded in the content representations, causing significant timbre leakage and reducing similarity to the target speaker. To address this, we introduce a residual block to a content extractor. The residual block consists of two weighted branches: 1) universal semantic dictionary based Content Feature Re-expression (CFR) module, supplying timbre-free content representation. 2) skip connection to the original content layer, providing complementary fine-grained information. In the CFR module, each dictionary entry in the universal semantic dictionary represents a phoneme class, computed statistically using speech from multiple speakers, creating a stable, speaker-independent semantic set. We introduce a CFR method to obtain timbre-free content representations by expressing each content frame as a weighted linear combination of dictionary entries using corresponding phoneme posteriors as weights. Extensive experiments across various VC frameworks demonstrate that our approach effectively mitigates timbre leakage and significantly improves similarity to the target speaker.


eess.AS音频处理


【1】 Reverberation-based Features for Sound Event Localization and Detection  with Distance Estimation
标题: 基于回响的声音事件定位和检测特征以及距离估计
链接:https://arxiv.org/abs/2504.08644
作者: Davide Berghi,  Philip J. B. Jackson 
摘要:声音事件定位和检测(SELD)涉及随着时间的推移预测活动的声音事件类别,同时估计它们的位置。SELD中的定位子任务通常被视为一个波达方向估计问题,忽略了源距离。直到最近,SELD才通过结合距离估计扩展到3D,从而能够预测3D空间中的声音事件位置(3D SELD)。然而,现有的方法缺乏设计用于距离估计的输入特征。我们认为,混响编码有价值的信息,这项任务。本文介绍了两种基于混响的3D SELD的新特征格式:一种使用直接混响比(DRR),另一种利用信号自相关为模型提供早期反射的见解。对合成数据进行预训练可以提高相对距离误差(RDE)和整体SELD分数,基于自相关的特征可以将STARSS 23数据集上的RDE降低超过3个百分点。提取特征的代码可在github.com/dberghi/SELD-distance-features上获得。
摘要:Sound event localization and detection (SELD) involves predicting active sound event classes over time while estimating their positions. The localization subtask in SELD is usually treated as a direction of arrival estimation problem, ignoring source distance. Only recently, SELD was extended to 3D by incorporating distance estimation, enabling the prediction of sound event positions in 3D space (3D SELD). However, existing methods lack input features designed for distance estimation. We argue that reverberation encodes valuable information for this task. This paper introduces two novel feature formats for 3D SELD based on reverberation: one using direct-to-reverberant ratio (DRR) and another leveraging signal autocorrelation to provide the model with insights into early reflections. Pre-training on synthetic data improves relative distance error (RDE) and overall SELD score, with autocorrelation-based features reducing RDE by over 3 percentage points on the STARSS23 dataset. The code to extract the features is available at github.com/dberghi/SELD-distance-features.


【2】 TorchFX: A modern approach to Audio DSP with PyTorch and GPU  acceleration

标题: TorchFX:具有PyTorch和图形处理器加速的音频DSP现代方法
链接:https://arxiv.org/abs/2504.08624
作者: Matteo Spanio,  Antonio Rodà 
备注:Submitted to DAFx 2025
摘要:音频信号日益增长的复杂性和实时处理需求需要优化算法,以利用图形处理单元(GPU)的计算能力。现有的数字信号处理(DSP)库通常无法提供必要的效率和灵活性,特别是在集成人工智能(AI)模型方面。作为回应,我们引入了TorchFX:一个用于DSP的GPU加速Python库,专门设计用于促进复杂的音频信号处理。TorchFX构建在PyTorch框架之上,提供了一个面向对象的界面,可以模拟torchaudio的可用性,通过新颖的管道操作符增强功能,实现直观的过滤器链接。该库提供了一套全面的有限脉冲响应(FIR)和无限脉冲响应(IIR)滤波器,重点关注多通道音频文件,从而促进了DSP和基于AI的方法的集成。我们的基准测试结果表明,与SciPy等传统库相比,效率有了显著提高,特别是在多通道环境中。尽管目前在GPU兼容性方面存在限制,但正在进行的开发有望提供更广泛的支持和实时处理能力。TorchFX的目标是成为社区的有用工具,通过GPU加速为DSP的创新和进步做出贡献。TorchFX在GitHub上公开发布,网址为https://github.com/matteospanio/torchfx。
摘要:The burgeoning complexity and real-time processing demands of audio signals necessitate optimized algorithms that harness the computational prowess of Graphics Processing Units (GPUs). Existing Digital Signal Processing (DSP) libraries often fall short in delivering the requisite efficiency and flexibility, particularly in integrating Artificial Intelligence (AI) models. In response, we introduce TorchFX: a GPU-accelerated Python library for DSP, specifically engineered to facilitate sophisticated audio signal processing. Built atop the PyTorch framework, TorchFX offers an Object-Oriented interface that emulates the usability of torchaudio, enhancing functionality with a novel pipe operator for intuitive filter chaining. This library provides a comprehensive suite of Finite Impulse Response (FIR) and Infinite Impulse Response (IIR) filters, with a focus on multichannel audio files, thus facilitating the integration of DSP and AI-based approaches. Our benchmarking results demonstrate significant efficiency gains over traditional libraries like SciPy, particularly in multichannel contexts. Despite current limitations in GPU compatibility, ongoing developments promise broader support and real-time processing capabilities. TorchFX aims to become a useful tool for the community, contributing to innovation and progress in DSP with GPU acceleration. TorchFX is publicly available on GitHub at https://github.com/matteospanio/torchfx.


【3】 Mitigating Timbre Leakage with Universal Semantic Mapping Residual Block  for Voice Conversion

标题: 用通用语义映射剩余块缓解语音转换的音色泄漏
链接:https://arxiv.org/abs/2504.08524
作者: Na Li,  Chuke Wang,  Yu Gu,  Zhifeng Li 
摘要:语音转换(VC)通过保留内容将源语音转换为目标语音。然而,来自源说话者的音色信息固有地嵌入在内容表示中,导致显著的音色泄漏并且降低与目标说话者的相似性。为了解决这个问题,我们将残差块引入内容提取器。剩余块由两个加权分支组成:1)基于通用语义词典的内容特征重表达(CFR)模块,提供无音色的内容表示。2)跳过与原始内容层的连接,提供补充的细粒度信息。在CFR模块中,通用语义词典中的每个词典条目表示使用来自多个说话者的语音统计地计算的音素类,从而创建稳定的、与说话者无关的语义集。我们引入了一种CFR方法,通过使用相应的音素后验子作为权重将每个内容帧表示为字典条目的加权线性组合来获得无音色的内容表示。在各种VC框架上的大量实验表明,我们的方法有效地减轻了音色泄漏,并显着提高了与目标说话人的相似性。
摘要:Voice conversion (VC) transforms source speech into a target voice by preserving the content. However, timbre information from the source speaker is inherently embedded in the content representations, causing significant timbre leakage and reducing similarity to the target speaker. To address this, we introduce a residual block to a content extractor. The residual block consists of two weighted branches: 1) universal semantic dictionary based Content Feature Re-expression (CFR) module, supplying timbre-free content representation. 2) skip connection to the original content layer, providing complementary fine-grained information. In the CFR module, each dictionary entry in the universal semantic dictionary represents a phoneme class, computed statistically using speech from multiple speakers, creating a stable, speaker-independent semantic set. We introduce a CFR method to obtain timbre-free content representations by expressing each content frame as a weighted linear combination of dictionary entries using corresponding phoneme posteriors as weights. Extensive experiments across various VC frameworks demonstrate that our approach effectively mitigates timbre leakage and significantly improves similarity to the target speaker.


【4】 BowelRCNN: Region-based Convolutional Neural Network System for Bowel  Sound Auscultation

标题: BowelRCNN:用于肠道音听诊的基于区域的卷积神经网络系统
链接:https://arxiv.org/abs/2504.08659
作者: Igor Matynia,  Robert Nowak 
备注:10 pages, 3 figures
摘要:表示肠道活动检测的声音事件是具有识别胃肠道状况的潜力的诊断工具。本文介绍了BowelRCNN,这是一种新型的肠音检测系统,它使用音频记录,频谱图分析和基于区域的卷积神经网络(RCNN)架构。该系统在从19名患者收集的真实记录数据集上进行了训练和验证,包括60分钟的准备和注释的音频数据。BowelRCNN的分类准确率为96%,F1评分为71%。这项研究强调了使用CNN架构进行肠音听诊的可行性,实现了与递归卷积方法相当的结果。
摘要:Sound events representing intestinal activity detection is a diagnostic tool with potential to identify gastrointestinal conditions. This article introduces BowelRCNN, a novel bowel sound detection system that uses audio recording, spectrogram analysys and region-based convolutional neural network (RCNN) architecture. The system was trained and validated on a real recording dataset gathered from 19 patients, comprising 60 minutes of prepared and annotated audio data. BowelRCNN achieved a classification accuracy of 96% and an F1 score of 71%. This research highlights the feasibility of using CNN architectures for bowel sound auscultation, achieving results comparable to those of recurrent-convolutional methods.


【5】 On The Landscape of Spoken Language Models: A Comprehensive Survey

标题: 口语模型的格局:全面调查
链接:https://arxiv.org/abs/2504.08528
作者: Siddhant Arora,  Kai-Wei Chang,  Chung-Ming Chien,  Yifan Peng,  Haibin Wu,  Yossi Adi,  Emmanuel Dupoux,  Hung-Yi Lee,  Karen Livescu,  Shinji Watanabe 
摘要:口语处理领域正在经历从训练定制的、特定于任务的模型向使用和优化充当通用语音处理系统的口语模型(SLM)的转变。这一趋势类似于(文本)自然语言处理领域向通用语言模型的发展。SLM既包括语音的“纯”语言模型--标记化语音序列的分布模型--也包括将语音编码器与文本语言模型相结合的模型,通常包括口头和书面输入或输出。这一领域的工作多种多样,有一系列术语和评价背景。本文的目的是通过一个统一的文献调查领域的发展背景下,最近的工作,有助于提高对可持续土地管理的理解。我们的调查按模型架构、培训和评估选择对这一领域的工作进行了分类,并描述了未来工作的一些关键挑战和方向。
摘要:The field of spoken language processing is undergoing a shift from training custom-built, task-specific models toward using and optimizing spoken language models (SLMs) which act as universal speech processing systems. This trend is similar to the progression toward universal language models that has taken place in the field of (text) natural language processing. SLMs include both "pure" language models of speech -- models of the distribution of tokenized speech sequences -- and models that combine speech encoders with text language models, often including both spoken and written input or output. Work in this area is very diverse, with a range of terminology and evaluation settings. This paper aims to contribute an improved understanding of SLMs via a unifying literature survey of recent work in the context of the evolution of the field. Our survey categorizes the work in this area by model architecture, training, and evaluation choices, and describes some key challenges and directions for future work.


【6】 On the Design of Diffusion-based Neural Speech Codecs

标题: 基于扩散的神经语音编解码器的设计
链接:https://arxiv.org/abs/2504.08470
作者: Pietro Foti,  Andreas Brendel 
摘要:最近,作为生成模型训练的神经语音编解码器(NSC)在低比特率下显示出比传统编解码器更优越的性能。虽然大多数最先进的NSC都是作为生成对抗网络(GANs)训练的,但扩散模型(DM)是最近一类生成模型,由于其在图像生成方面相对于GANs的优越性能,因此代表了一种有前途的替代方案。因此,DM已经成功地应用于各种其他音频生成应用中的音频和语音编码。然而,基于扩散的神经干细胞的设计尚未被系统地探索。我们解决这个问题,提供一个全面的分析扩散为基础的神经干细胞分为三个贡献。首先,我们提出了一个分类的基础上的条件和输出域的DM。这个简单的概念框架,使我们能够定义一个设计空间的扩散为基础的神经干细胞,并在文献中现有的方法分配一个类别。其次,我们系统地研究未开发的设计,通过创建和评估新的扩散为基础的神经干细胞的概念框架内。最后,我们通过客观指标和主观听力测试将所提出的模型与现有的GAN和DM基线进行比较。
摘要:Recently, neural speech codecs (NSCs) trained as generative models have shown superior performance compared to conventional codecs at low bitrates. Although most state-of-the-art NSCs are trained as Generative Adversarial Networks (GANs), Diffusion Models (DMs), a recent class of generative models, represent a promising alternative due to their superior performance in image generation relative to GANs. Consequently, DMs have been successfully applied for audio and speech coding among various other audio generation applications. However, the design of diffusion-based NSCs has not yet been explored in a systematic way. We address this by providing a comprehensive analysis of diffusion-based NSCs divided into three contributions. First, we propose a categorization based on the conditioning and output domains of the DM. This simple conceptual framework allows us to define a design space for diffusion-based NSCs and to assign a category to existing approaches in the literature. Second, we systematically investigate unexplored designs by creating and evaluating new diffusion-based NSCs within the conceptual framework. Finally, we compare the proposed models to existing GAN and DM baselines through objective metrics and subjective listening tests.


【7】 Passive Underwater Acoustic Signal Separation based on Feature  Decoupling Dual-path Network

标题: 基于特征解耦双径网络的被动水声信号分离
链接:https://arxiv.org/abs/2504.08371
作者: Yucheng Liu,  Longyu Jiang 
备注:10pages,4 figures
摘要:被动水声领域的信号分离在很大程度上依赖于深度学习技术来隔离船舶辐射噪声。然而,该领域常用的分离网络源于语音分离应用,可能事先没有充分考虑水声的独特方面,例如不同传播介质、信号频率和调制特性的影响。这一疏忽突出了需要量身定制的方法,考虑到水下声音传播的具体特点。本文提出了一种新的时域网络设计,采用双路径模型和特征解耦的方法来分离船舶辐射噪声。混合信号的特征被转换到一个空间中,在该空间中它们表现出更大的独立性,并且每个维度的重要性被解耦。随后,在分离层中采用局部和全局注意机制的融合。与其他流行的网络模型相比,广泛的比较显示了该方法的有效性,其在ShipsEar和DeepShip数据集中的性能证明了这一点。
摘要:Signal separation in the passive underwater acoustic domain has heavily relied on deep learning techniques to isolate ship radiated noise. However, the separation networks commonly used in this domain stem from speech separation applications and may not fully consider the unique aspects of underwater acoustics beforehand, such as the influence of different propagation media, signal frequencies and modulation characteristics. This oversight highlights the need for tailored approaches that account for the specific characteristics of underwater sound propagation. This study introduces a novel temporal network designed to separate ship radiated noise by employing a dual-path model and a feature decoupling approach. The mixed signals' features are transformed into a space where they exhibit greater independence, with each dimension's significance decoupled. Subsequently, a fusion of local and global attention mechanisms is employed in the separation layer. Extensive comparisons showcase the effectiveness of this method when compared to other prevalent network models, as evidenced by its performance in the ShipsEar and DeepShip datasets.


【8】 Location-Oriented Sound Event Localization and Detection with Spatial  Mapping and Regression Localization

标题: 利用空间映射和回归定位的面向位置的声音事件定位和检测
链接:https://arxiv.org/abs/2504.08365
作者: Xueping Zhang,  Yaxiong Chen,  Ruilin Yao,  Yunfei Zi,  Shengwu Xiong 
摘要:声音事件定位与检测(SELD)将声音事件检测(SED)与相应的到达方向(DOA)相结合。目前,采用的面向事件的多声道方法由于声道数量的限制,影响了多声道环境下的通用性。为了提高在复音环境中的通用性,我们提出了空间映射和回归定位SELD(SMRL-SELD)。SMRL-SELD分割三维空间,将其映射到二维平面,并提出了一种新的回归定位损失,以帮助结果收敛到相应的事件的位置。SMRL-SELD是面向位置的,允许模型基于方向学习事件特征。因此,该方法使模型能够处理复调声音,而不管重叠事件的数量。我们在STARSS 23和STARSS 22数据集上进行了实验,我们提出的SMRL-SELD在整体评估和复调环境中优于现有的SELD方法。
摘要:Sound Event Localization and Detection (SELD) combines the Sound Event Detection (SED) with the corresponding Direction Of Arrival (DOA). Recently, adopted event oriented multi-track methods affect the generality in polyphonic environments due to the limitation of the number of tracks. To enhance the generality in polyphonic environments, we propose Spatial Mapping and Regression Localization for SELD (SMRL-SELD). SMRL-SELD segments the 3D spatial space, mapping it to a 2D plane, and a new regression localization loss is proposed to help the results converge toward the location of the corresponding event. SMRL-SELD is location-oriented, allowing the model to learn event features based on orientation. Thus, the method enables the model to process polyphonic sounds regardless of the number of overlapping events. We conducted experiments on STARSS23 and STARSS22 datasets and our proposed SMRL-SELD outperforms the existing SELD methods in overall evaluation and polyphony environments.


【9】 Generalized Multilingual Text-to-Speech Generation with Language-Aware  Style Adaptation

标题: 具有地理感知风格适应的广义多语言文本到语音生成
链接:https://arxiv.org/abs/2504.08274
作者: Haowei Lou,  Hye-young Paik,  Sheng Li,  Wen Hu,  Lina Yao 
摘要:文本到语音(TTS)模型可以通过将音素转换为波形来跨多种语言生成自然的、类似人类的语音。然而,由于音素词汇的差异以及不同语言之间韵律和说话风格的差异,多语言TTS仍然具有挑战性。现有的方法要么为每种语言训练单独的模型,以增加计算资源为代价实现高性能,要么为多种语言使用统一的模型,努力捕捉细粒度的,语言特定的风格变化。在这项工作中,我们提出了LanStyleTTS,一个非自回归,语言感知的风格自适应TTS框架,该框架能够识别音素表示,并能够跨语言进行细粒度的音素级风格控制。这种设计支持一个统一的多语言TTS模型,能够产生准确和高质量的语音,而不需要训练语言特定的模型。我们评估LanStyleTTS集成它与几个国家的最先进的非自回归TTS架构。结果表明,在不同的模型骨干一致的性能改进。此外,我们调查了一系列的声学特征表示,包括梅尔频谱图和自动编码器派生的潜在功能。我们的实验表明,潜在的编码可以显着降低模型的大小和计算成本,同时保持高质量的语音生成。
摘要:Text-to-Speech (TTS) models can generate natural, human-like speech across multiple languages by transforming phonemes into waveforms. However, multilingual TTS remains challenging due to discrepancies in phoneme vocabularies and variations in prosody and speaking style across languages. Existing approaches either train separate models for each language, which achieve high performance at the cost of increased computational resources, or use a unified model for multiple languages that struggles to capture fine-grained, language-specific style variations. In this work, we propose LanStyleTTS, a non-autoregressive, language-aware style adaptive TTS framework that standardizes phoneme representations and enables fine-grained, phoneme-level style control across languages. This design supports a unified multilingual TTS model capable of producing accurate and high-quality speech without the need to train language-specific models. We evaluate LanStyleTTS by integrating it with several state-of-the-art non-autoregressive TTS architectures. Results show consistent performance improvements across different model backbones. Furthermore, we investigate a range of acoustic feature representations, including mel-spectrograms and autoencoder-derived latent features. Our experiments demonstrate that latent encodings can significantly reduce model size and computational cost while preserving high-quality speech generation.


【10】 From Speech to Summary: A Comprehensive Survey of Speech Summarization

标题: 从演讲到总结:演讲总结综述
链接:https://arxiv.org/abs/2504.08024
作者: Fabian Retkowski,  Maike Züfle,  Andreas Sudmann,  Dinah Pfau,  Jan Niehues,  Alexander Waibel 
摘要:语音摘要已成为有效管理和访问不断增长的语音和视听内容的重要工具。然而,尽管其重要性越来越大,语音摘要仍然没有明确定义,并与几个研究领域交叉,包括语音识别,文本摘要和会议摘要等特定应用。该调查不仅检查了现有的数据集和评估方法,这对于评估摘要方法的有效性至关重要,而且还综合了该领域的最新发展,突出了从传统系统到高级模型的转变,如微调级联架构和端到端解决方案。
摘要:Speech summarization has become an essential tool for efficiently managing and accessing the growing volume of spoken and audiovisual content. However, despite its increasing importance, speech summarization is still not clearly defined and intersects with several research areas, including speech recognition, text summarization, and specific applications like meeting summarization. This survey not only examines existing datasets and evaluation methodologies, which are crucial for assessing the effectiveness of summarization approaches but also synthesizes recent developments in the field, highlighting the shift from traditional systems to advanced models like fine-tuned cascaded architectures and end-to-end solutions.


机器翻译由腾讯交互翻译提供,仅供参考