今日论文合集:cs.SD语音5篇,eess.AS音频处理5篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech  Translation

标题: MultiMed-ST:大规模多对多语言医学语音翻译
链接:https://arxiv.org/abs/2504.03546
作者: Khai Le-Duc,  Tuyen Tran,  Bach Phan Tat,  Nguyen Kim Hai Bui,  Quan Dang,  Hung-Phong Tran,  Thanh-Thuy Nguyen,  Ly Nguyen,  Tuan-Minh Phan,  Thi Thu Phuong Tran,  Chris Ngo,  Nguyen X. Khanh,  Thanh Nguyen-Tang 
备注:Preprint, 122 pages
摘要:医疗领域的多语言语音翻译(ST)通过跨越语言障碍实现高效沟通,缓解专业劳动力短缺,并促进改善诊断和治疗,特别是在流行病期间,增强了患者护理。在这项工作中,我们提出了第一个系统的研究医学ST,据我们所知,通过发布MultiMed-ST,一个大规模的ST数据集的医疗领域,跨越所有翻译方向的五种语言:越南语,英语,德语,法语,繁体中文和简体中文,连同模型。我们的数据集拥有290,000个样本,是最大的医学机器翻译(MT)数据集,也是所有领域中最大的多对多语言ST。其次,我们提出了迄今为止ST研究中最广泛的分析研究,包括:经验基线,双语-多语比较研究,端到端与级联比较研究,特定任务与多任务序列到序列(seq 2seq)比较研究,代码转换分析,以及定量-定性错误分析。所有代码、数据和模型均可在线获取:https://github.com/leduckhai/MultiMed-ST。
摘要:Multilingual speech translation (ST) in the medical domain enhances patient care by enabling efficient communication across language barriers, alleviating specialized workforce shortages, and facilitating improved diagnosis and treatment, particularly during pandemics. In this work, we present the first systematic study on medical ST, to our best knowledge, by releasing MultiMed-ST, a large-scale ST dataset for the medical domain, spanning all translation directions in five languages: Vietnamese, English, German, French, Traditional Chinese and Simplified Chinese, together with the models. With 290,000 samples, our dataset is the largest medical machine translation (MT) dataset and the largest many-to-many multilingual ST among all domains. Secondly, we present the most extensive analysis study in ST research to date, including: empirical baselines, bilingual-multilingual comparative study, end-to-end vs. cascaded comparative study, task-specific vs. multi-task sequence-to-sequence (seq2seq) comparative study, code-switch analysis, and quantitative-qualitative error analysis. All code, data, and models are available online: https://github.com/leduckhai/MultiMed-ST.


【2】 An Efficient GPU-based Implementation for Noise Robust Sound Source  Localization
标题: 一种基于GOP的有效噪音鲁棒性光源定位实现
链接:https://arxiv.org/abs/2504.03373

作者: Zirui Lin,  Masayuki Takigahira,  Naoya Terakado,  Haris Gulzar,  Monikka Roslianna Busto,  Takeharu Eda,  Katsutoshi Itoyama,  Kazuhiro Nakadai,  Hideharu Amano 
备注:6 pages, 2 figures
摘要:机器人听觉,包括声源定位(SSL),声源分离(SSS)和自动语音识别(ASR),使机器人和智能设备能够获得类似于人类听觉的听觉能力。尽管它们具有广泛的适用性,但在SSL中处理来自麦克风阵列的多声道音频信号涉及计算密集型矩阵运算,这可能会阻碍在中央处理单元(CPU)上的有效部署,特别是在CPU资源有限的嵌入式系统中。本文介绍了一种基于GPU的SSL机器人试听的实现,利用广义奇异值分解为基础的多信号分类(GSVD-MUSIC),一种抗噪声的算法,在HARK平台,一个开源的软件套件。对于一个60通道的麦克风阵列,所提出的实施方案实现了显着的性能改善。在由NVIDIA GPU和ARM Cortex-A78 AE v8.2 64位CPU驱动的嵌入式设备Jetson AGX Orin上,我们观察到GSVD计算的加速比为4645.1倍,SSL模块的加速比为8.8倍,而在配置了NVIDIA A100 GPU和AMD EPYC 7352 CPU的服务器上,GSVD计算的加速比为2223.4倍,整个SSL模块的加速比为8.95倍,这使得实时处理对于大规模麦克风阵列是可行的,并且为潜在的后续机器学习或深度学习任务的实时处理提供充足的容量。
摘要:Robot audition, encompassing Sound Source Localization (SSL), Sound Source Separation (SSS), and Automatic Speech Recognition (ASR), enables robots and smart devices to acquire auditory capabilities similar to human hearing. Despite their wide applicability, processing multi-channel audio signals from microphone arrays in SSL involves computationally intensive matrix operations, which can hinder efficient deployment on Central Processing Units (CPUs), particularly in embedded systems with limited CPU resources. This paper introduces a GPU-based implementation of SSL for robot audition, utilizing the Generalized Singular Value Decomposition-based Multiple Signal Classification (GSVD-MUSIC), a noise-robust algorithm, within the HARK platform, an open-source software suite. For a 60-channel microphone array, the proposed implementation achieves significant performance improvements. On the Jetson AGX Orin, an embedded device powered by an NVIDIA GPU and ARM Cortex-A78AE v8.2 64-bit CPUs, we observe speedups of 4645.1x for GSVD calculations and 8.8x for the SSL module, while speedups of 2223.4x for GSVD calculation and 8.95x for the entire SSL module on a server configured with an NVIDIA A100 GPU and AMD EPYC 7352 CPUs, making real-time processing feasible for large-scale microphone arrays and providing ample capacity for real-time processing of potential subsequent machine learning or deep learning tasks.


【3】 RWKVTTS: Yet another TTS based on RWKV-7
标题: RWKVTTC:基于RWKN-7的又一款TTC
链接:https://arxiv.org/abs/2504.03289

作者: Lin yueyu,  Liu Xiao 
摘要:人机交互在直观和高效的界面上蓬勃发展,其中语音作为一种特别自然和可访问的方式脱颖而出。基于转换器的文本到语音(TTS)系统的最新进展,如Fish-Speech,CosyVoice和MegaTTS 3,在质量和真实性方面取得了显着的改进,推动了TTS领域的重大发展。在本文中,我们介绍了RWKV-7 \cite{peng 2025 rwav},一个尖端的基于RNN的体系结构,专为TTS应用。与传统的Transformer模型不同,RWKV-7利用递归神经网络的优势来实现更高的计算效率和可扩展性,同时保持高质量的输出。我们全面的基准测试表明,RWKV-7在多个关键指标(包括合成速度、语音自然度和资源效率)方面优于基于变换器的模型。此外,我们还探讨了它对不同语言背景和低资源环境的适应性,展示了它使TTS技术民主化的潜力。这些发现将RWKV-7定位为功能强大且具有创新性的替代方案,为现实世界应用中更易于访问和通用的语音合成解决方案铺平了道路。我们的代码和权重为https://github.com/yynil/RWKVTTS,https://huggingface.co/spaces/RWKV-Red-Team
摘要:Human-AI interaction thrives on intuitive and efficient interfaces, among which voice stands out as a particularly natural and accessible modality. Recent advancements in transformer-based text-to-speech (TTS) systems, such as Fish-Speech, CosyVoice, and MegaTTS 3, have delivered remarkable improvements in quality and realism, driving a significant evolution in the TTS domain. In this paper, we introduce RWKV-7 \cite{peng2025rwkv}, a cutting-edge RNN-based architecture tailored for TTS applications. Unlike traditional transformer models, RWKV-7 leverages the strengths of recurrent neural networks to achieve greater computational efficiency and scalability, while maintaining high-quality output. Our comprehensive benchmarks demonstrate that RWKV-7 outperforms transformer-based models across multiple key metrics, including synthesis speed, naturalness of speech, and resource efficiency. Furthermore, we explore its adaptability to diverse linguistic contexts and low-resource environments, showcasing its potential to democratize TTS technology. These findings position RWKV-7 as a powerful and innovative alternative, paving the way for more accessible and versatile voice synthesis solutions in real-world applications.Our code and weights are https://github.com/yynil/RWKVTTS, https://huggingface.co/spaces/RWKV-Red-Team


【4】 Generating Diverse Audio-Visual 360 Soundscapes for Sound Event  Localization and Detection
标题: 生成多样化的视听360声景以进行声音事件定位和检测
链接:https://arxiv.org/abs/2504.02988

作者: Adrian S. Roman,  Aiden Chang,  Gerardo Meza,  Iran R. Roman 
摘要:我们提出SELDVisualSynth,用于生成合成视频的视听声音事件定位和检测(SELD)的工具。我们的方法结合了真实世界的背景图像,以提高合成视听SELD数据的真实性,同时也确保视听空间对齐。该工具创建360合成视频,其中对象移动匹配合成SELD音频数据及其注释。实验结果表明,使用该数据训练的模型在多个指标上获得了性能提升,实现了卓越的定位召回率(56.4 LR)和有竞争力的定位错误率(21.9deg LE)。我们开源了我们的数据生成工具,供SELD研究社区的成员最大限度地使用。
摘要:We present SELDVisualSynth, a tool for generating synthetic videos for audio-visual sound event localization and detection (SELD). Our approach incorporates real-world background images to improve realism in synthetic audio-visual SELD data while also ensuring audio-visual spatial alignment. The tool creates 360 synthetic videos where objects move matching synthetic SELD audio data and its annotations. Experimental results demonstrate that a model trained with this data attains performance gains across multiple metrics, achieving superior localization recall (56.4 LR) and competitive localization error (21.9deg LE). We open-source our data generation tool for maximal use by members of the SELD research community.


【5】 Mind the Prompt: Prompting Strategies in Audio Generations for Improving  Sound Classification
标题: 注意提示:改进声音分类的音频生成策略
链接:https://arxiv.org/abs/2504.03329

作者: Francesca Ronchini,  Ho-Hsiang Wu,  Wei-Cheng Lin,  Fabio Antonacci 
备注:Accepted at Generative Data Augmentation for Real-World Signal Processing Applications Workshop
摘要:本文研究了使用文本到音频(TTA)模型生成真实数据集的有效提示策略的设计。我们还分析了不同的技术,有效地结合这些数据集,以提高其在声音分类任务的效用。通过使用两个TTA模型评估两个声音分类数据集,我们应用了一系列提示策略。我们的研究结果表明,特定任务的提示策略显着优于基本提示方法在数据生成。此外,事实证明,合并使用不同TTA模型生成的数据集比仅仅增加训练数据集大小更有效地增强分类性能。总的来说,我们的研究结果强调了这些方法作为使用合成数据的有效数据增强技术的优势。
摘要:This paper investigates the design of effective prompt strategies for generating realistic datasets using Text-To-Audio (TTA) models. We also analyze different techniques for efficiently combining these datasets to enhance their utility in sound classification tasks. By evaluating two sound classification datasets with two TTA models, we apply a range of prompt strategies. Our findings reveal that task-specific prompt strategies significantly outperform basic prompt approaches in data generation. Furthermore, merging datasets generated using different TTA models proves to enhance classification performance more effectively than merely increasing the training dataset size. Overall, our results underscore the advantages of these methods as effective data augmentation techniques using synthetic data.


eess.AS音频处理


【1】 Mind the Prompt: Prompting Strategies in Audio Generations for Improving  Sound Classification
标题: 注意提示:改进声音分类的音频生成策略
链接:https://arxiv.org/abs/2504.03329

作者: Francesca Ronchini,  Ho-Hsiang Wu,  Wei-Cheng Lin,  Fabio Antonacci 
备注:Accepted at Generative Data Augmentation for Real-World Signal Processing Applications Workshop
摘要:本文研究了使用文本到音频(TTA)模型生成真实数据集的有效提示策略的设计。我们还分析了不同的技术,有效地结合这些数据集,以提高其在声音分类任务的效用。通过使用两个TTA模型评估两个声音分类数据集,我们应用了一系列提示策略。我们的研究结果表明,特定任务的提示策略显着优于基本提示方法在数据生成。此外,合并使用不同TTA模型生成的数据集证明比仅仅增加训练数据集大小更有效地提高分类性能。总的来说,我们的研究结果强调了这些方法作为使用合成数据的有效数据增强技术的优势。
摘要:This paper investigates the design of effective prompt strategies for generating realistic datasets using Text-To-Audio (TTA) models. We also analyze different techniques for efficiently combining these datasets to enhance their utility in sound classification tasks. By evaluating two sound classification datasets with two TTA models, we apply a range of prompt strategies. Our findings reveal that task-specific prompt strategies significantly outperform basic prompt approaches in data generation. Furthermore, merging datasets generated using different TTA models proves to enhance classification performance more effectively than merely increasing the training dataset size. Overall, our results underscore the advantages of these methods as effective data augmentation techniques using synthetic data.


【2】 MultiMed-ST: Large-scale Many-to-many Multilingual Medical Speech  Translation
标题: MultiMed-ST:大规模多对多语言医学语音翻译
链接:https://arxiv.org/abs/2504.03546

作者: Khai Le-Duc,  Tuyen Tran,  Bach Phan Tat,  Nguyen Kim Hai Bui,  Quan Dang,  Hung-Phong Tran,  Thanh-Thuy Nguyen,  Ly Nguyen,  Tuan-Minh Phan,  Thi Thu Phuong Tran,  Chris Ngo,  Nguyen X. Khanh,  Thanh Nguyen-Tang 
备注:Preprint, 122 pages
摘要:医疗领域的多语言语音翻译(ST)通过跨越语言障碍实现高效沟通,缓解专业劳动力短缺,并促进改善诊断和治疗,特别是在流行病期间,增强了患者护理。在这项工作中,我们提出了第一个系统的研究医学ST,据我们所知,通过发布MultiMed-ST,一个大规模的ST数据集的医疗领域,跨越所有翻译方向的五种语言:越南语,英语,德语,法语,繁体中文和简体中文,连同模型。我们的数据集拥有290,000个样本,是最大的医学机器翻译(MT)数据集,也是所有领域中最大的多对多语言ST。其次,我们提出了迄今为止ST研究中最广泛的分析研究,包括:经验基线,双语-多语比较研究,端到端与级联比较研究,特定任务与多任务序列到序列(seq 2seq)比较研究,代码转换分析,以及定量-定性错误分析。所有代码、数据和模型均可在线获取:https://github.com/leduckhai/MultiMed-ST。
摘要:Multilingual speech translation (ST) in the medical domain enhances patient care by enabling efficient communication across language barriers, alleviating specialized workforce shortages, and facilitating improved diagnosis and treatment, particularly during pandemics. In this work, we present the first systematic study on medical ST, to our best knowledge, by releasing MultiMed-ST, a large-scale ST dataset for the medical domain, spanning all translation directions in five languages: Vietnamese, English, German, French, Traditional Chinese and Simplified Chinese, together with the models. With 290,000 samples, our dataset is the largest medical machine translation (MT) dataset and the largest many-to-many multilingual ST among all domains. Secondly, we present the most extensive analysis study in ST research to date, including: empirical baselines, bilingual-multilingual comparative study, end-to-end vs. cascaded comparative study, task-specific vs. multi-task sequence-to-sequence (seq2seq) comparative study, code-switch analysis, and quantitative-qualitative error analysis. All code, data, and models are available online: https://github.com/leduckhai/MultiMed-ST.


【3】 An Efficient GPU-based Implementation for Noise Robust Sound Source  Localization
标题: 一种基于GOP的有效噪音鲁棒性光源定位实现
链接:https://arxiv.org/abs/2504.03373

作者: Zirui Lin,  Masayuki Takigahira,  Naoya Terakado,  Haris Gulzar,  Monikka Roslianna Busto,  Takeharu Eda,  Katsutoshi Itoyama,  Kazuhiro Nakadai,  Hideharu Amano 
备注:6 pages, 2 figures
摘要:机器人听觉,包括声源定位(SSL),声源分离(SSS)和自动语音识别(ASR),使机器人和智能设备能够获得类似于人类听觉的听觉能力。尽管它们具有广泛的适用性,但在SSL中处理来自麦克风阵列的多声道音频信号涉及计算密集型矩阵运算,这可能会阻碍在中央处理单元(CPU)上的有效部署,特别是在CPU资源有限的嵌入式系统中。本文介绍了一种基于GPU的SSL机器人试听的实现,利用广义奇异值分解为基础的多信号分类(GSVD-MUSIC),一种抗噪声的算法,在HARK平台,一个开源的软件套件。对于一个60通道的麦克风阵列,所提出的实施方案实现了显着的性能改善。在由NVIDIA GPU和ARM Cortex-A78 AE v8.2 64位CPU驱动的嵌入式设备Jetson AGX Orin上,我们观察到GSVD计算的加速比为4645.1倍,SSL模块的加速比为8.8倍,而在配置了NVIDIA A100 GPU和AMD EPYC 7352 CPU的服务器上,GSVD计算的加速比为2223.4倍,整个SSL模块的加速比为8.95倍,这使得实时处理对于大规模麦克风阵列是可行的,并且为潜在的后续机器学习或深度学习任务的实时处理提供充足的容量。
摘要:Robot audition, encompassing Sound Source Localization (SSL), Sound Source Separation (SSS), and Automatic Speech Recognition (ASR), enables robots and smart devices to acquire auditory capabilities similar to human hearing. Despite their wide applicability, processing multi-channel audio signals from microphone arrays in SSL involves computationally intensive matrix operations, which can hinder efficient deployment on Central Processing Units (CPUs), particularly in embedded systems with limited CPU resources. This paper introduces a GPU-based implementation of SSL for robot audition, utilizing the Generalized Singular Value Decomposition-based Multiple Signal Classification (GSVD-MUSIC), a noise-robust algorithm, within the HARK platform, an open-source software suite. For a 60-channel microphone array, the proposed implementation achieves significant performance improvements. On the Jetson AGX Orin, an embedded device powered by an NVIDIA GPU and ARM Cortex-A78AE v8.2 64-bit CPUs, we observe speedups of 4645.1x for GSVD calculations and 8.8x for the SSL module, while speedups of 2223.4x for GSVD calculation and 8.95x for the entire SSL module on a server configured with an NVIDIA A100 GPU and AMD EPYC 7352 CPUs, making real-time processing feasible for large-scale microphone arrays and providing ample capacity for real-time processing of potential subsequent machine learning or deep learning tasks.


【4】 RWKVTTS: Yet another TTS based on RWKV-7
标题: RWKVTTC:基于RWKN-7的又一款TTC
链接:https://arxiv.org/abs/2504.03289

作者: Lin yueyu,  Liu Xiao 
摘要:人机交互在直观和高效的界面上蓬勃发展,其中语音作为一种特别自然和可访问的方式脱颖而出。基于转换器的文本到语音(TTS)系统的最新进展,如Fish-Speech,CosyVoice和MegaTTS 3,在质量和真实性方面取得了显着的改进,推动了TTS领域的重大发展。在本文中,我们介绍了RWKV-7 \cite{peng 2025 rwav},一个尖端的基于RNN的体系结构,专为TTS应用。与传统的Transformer模型不同,RWKV-7利用递归神经网络的优势来实现更高的计算效率和可扩展性,同时保持高质量的输出。我们全面的基准测试表明,RWKV-7在多个关键指标(包括合成速度、语音自然度和资源效率)方面优于基于变换器的模型。此外,我们还探讨了它对不同语言背景和低资源环境的适应性,展示了它使TTS技术民主化的潜力。这些发现将RWKV-7定位为功能强大且具有创新性的替代方案,为现实世界应用中更易于访问和通用的语音合成解决方案铺平了道路。我们的代码和权重为https://github.com/yynil/RWKVTTS,https://huggingface.co/spaces/RWKV-Red-Team
摘要:Human-AI interaction thrives on intuitive and efficient interfaces, among which voice stands out as a particularly natural and accessible modality. Recent advancements in transformer-based text-to-speech (TTS) systems, such as Fish-Speech, CosyVoice, and MegaTTS 3, have delivered remarkable improvements in quality and realism, driving a significant evolution in the TTS domain. In this paper, we introduce RWKV-7 \cite{peng2025rwkv}, a cutting-edge RNN-based architecture tailored for TTS applications. Unlike traditional transformer models, RWKV-7 leverages the strengths of recurrent neural networks to achieve greater computational efficiency and scalability, while maintaining high-quality output. Our comprehensive benchmarks demonstrate that RWKV-7 outperforms transformer-based models across multiple key metrics, including synthesis speed, naturalness of speech, and resource efficiency. Furthermore, we explore its adaptability to diverse linguistic contexts and low-resource environments, showcasing its potential to democratize TTS technology. These findings position RWKV-7 as a powerful and innovative alternative, paving the way for more accessible and versatile voice synthesis solutions in real-world applications.Our code and weights are https://github.com/yynil/RWKVTTS, https://huggingface.co/spaces/RWKV-Red-Team


【5】 Generating Diverse Audio-Visual 360 Soundscapes for Sound Event  Localization and Detection
标题: 生成多样化的视听360声景以进行声音事件定位和检测
链接:https://arxiv.org/abs/2504.02988

作者: Adrian S. Roman,  Aiden Chang,  Gerardo Meza,  Iran R. Roman 
摘要:我们提出SELDVisualSynth,用于生成合成视频的视听声音事件定位和检测(SELD)的工具。我们的方法结合了真实世界的背景图像,以提高合成视听SELD数据的真实性,同时也确保视听空间对齐。该工具创建360合成视频,其中对象移动匹配合成SELD音频数据及其注释。实验结果表明,使用该数据训练的模型在多个指标上获得了性能提升,实现了卓越的定位召回率(56.4 LR)和有竞争力的定位错误率(21.9deg LE)。我们开源了我们的数据生成工具,供SELD研究社区的成员最大限度地使用。
摘要:We present SELDVisualSynth, a tool for generating synthetic videos for audio-visual sound event localization and detection (SELD). Our approach incorporates real-world background images to improve realism in synthetic audio-visual SELD data while also ensuring audio-visual spatial alignment. The tool creates 360 synthetic videos where objects move matching synthetic SELD audio data and its annotations. Experimental results demonstrate that a model trained with this data attains performance gains across multiple metrics, achieving superior localization recall (56.4 LR) and competitive localization error (21.9deg LE). We open-source our data generation tool for maximal use by members of the SELD research community.


机器翻译由腾讯交互翻译提供,仅供参考