今日论文合集:cs.SD语音6篇,eess.AS音频处理7篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】 Seeing Soundscapes: Audio-Visual Generation and Separation from  Soundscapes Using Audio-Visual Separator
标题: 看到声景:视听生成和使用视听分离器从声景中分离
链接:https://arxiv.org/abs/2504.18283
作者: Minjae Kang,  Martim Brandão 
备注:Originally submitted to CVPR 2025 on 2024-11-15 with paper ID 15808
摘要:最近的视听生成模型在从音频生成图像方面取得了实质性的进展。然而,现有的方法集中于从单类音频生成图像,而不能从混合音频生成图像。为了解决这个问题,我们提出了一个视听生成和分离模型(AV-GAS),用于从音景(包含多个类的混合音频)生成图像。我们的贡献有三个方面:首先,我们提出了一个新的挑战,在视听生成任务,这是生成一个图像给定的多类音频输入,我们提出了一种方法,解决了这个任务,使用视听分离器。其次,我们引入了一个新的视听分离任务,它涉及到为混合音频输入中的每个类生成单独的图像。最后,我们提出了新的评估指标的视听生成任务:类表示分数(CRS)和修改后的R@K。我们的模型是在VGGSound数据集上训练和评估的。我们证明了我们的方法优于最先进的方法,在生成具有混合音频的合理图像时,CRS提高了7%,R@2* 提高了4%。
摘要:Recent audio-visual generative models have made substantial progress in generating images from audio. However, existing approaches focus on generating images from single-class audio and fail to generate images from mixed audio. To address this, we propose an Audio-Visual Generation and Separation model (AV-GAS) for generating images from soundscapes (mixed audio containing multiple classes). Our contribution is threefold: First, we propose a new challenge in the audio-visual generation task, which is to generate an image given a multi-class audio input, and we propose a method that solves this task using an audio-visual separator. Second, we introduce a new audio-visual separation task, which involves generating separate images for each class present in a mixed audio input. Lastly, we propose new evaluation metrics for the audio-visual generation task: Class Representation Score (CRS) and a modified R@K. Our model is trained and evaluated on the VGGSound dataset. We show that our method outperforms the state-of-the-art, achieving 7% higher CRS and 4% higher R@2* in generating plausible images with mixed audio.


【2】 Tracking Articulatory Dynamics in Speech with a Fixed-Weight BiLSTM-CNN  Architecture

标题: 使用固定权重BiLSTM-CNN架构跟踪语音中的关节动态
链接:https://arxiv.org/abs/2504.18099
作者: Leena G Pillai,  D. Muhammad Noorul Mubarak,  Elizabeth Sherly 
备注:10 pages with 8 figures. This paper presented in an international Conference
摘要:言语产生是一个复杂的连续过程,涉及到各种发音特征的协调。其中,舌头是一个高度灵活的主动发音器官,负责塑造气流,以产生有针对性的语音,是智力,清晰,独特的。本文提出了一种新的方法来预测舌头和嘴唇发音特征涉及一个给定的语音声学使用堆叠双向长短期记忆(BiLSTM)架构,结合一维卷积神经网络(CNN)的后处理与固定权重初始化。建议的网络是用两个数据集训练的,这两个数据集由同时记录的语音和电磁关节成像(EMA)数据集组成,每个数据集都引入了地理起源,语言特征,语音多样性和记录设备方面的变化。该模型的性能进行了评估,在说话人相关(SD),说话人无关(SI),语料库相关(CD)和跨语料库(CC)模式。实验结果表明,所提出的模型与固定的权重方法优于自适应权重初始化在相对最少的训练时期。这些研究结果有助于发展强大而有效的发音特征预测模型,为语音产生研究和应用的进步铺平了道路。
摘要:Speech production is a complex sequential process which involve the coordination of various articulatory features. Among them tongue being a highly versatile active articulator responsible for shaping airflow to produce targeted speech sounds that are intellectual, clear, and distinct. This paper presents a novel approach for predicting tongue and lip articulatory features involved in a given speech acoustics using a stacked Bidirectional Long Short-Term Memory (BiLSTM) architecture, combined with a one-dimensional Convolutional Neural Network (CNN) for post-processing with fixed weights initialization. The proposed network is trained with two datasets consisting of simultaneously recorded speech and Electromagnetic Articulography (EMA) datasets, each introducing variations in terms of geographical origin, linguistic characteristics, phonetic diversity, and recording equipment. The performance of the model is assessed in Speaker Dependent (SD), Speaker Independent (SI), corpus dependent (CD) and cross corpus (CC) modes. Experimental results indicate that the proposed model with fixed weights approach outperformed the adaptive weights initialization with in relatively minimal number of training epochs. These findings contribute to the development of robust and efficient models for articulatory feature prediction, paving the way for advancements in speech production research and applications.


【3】 STNet: Prediction of Underwater Sound Speed Profiles with An Advanced  Semi-Transformer Neural Network

标题: STNet:使用先进的半Transformer神经网络预测水下音速轮廓
链接:https://arxiv.org/abs/2504.17912
作者: Wei Huang,  Jiajun Lu,  Hao Zhang,  Tianhe Xu 
摘要:实时获取准确的水下声速剖面是跟踪水声信号传播轨迹的关键,在海洋通信定位中发挥着关键作用。SSP可以通过仪器直接测量或利用声场数据进行反演。虽然测量技术提供了良好的精度,但它们受到有限的空间覆盖范围的限制,并且需要大量的时间投资。基于声场数据实时测量的反演方法提高了业务效率,但由于其对海洋观测基础设施的严格要求,失去了SSP估计的准确性,并且受到有限的空间适用性的影响。为了实现准确的长期海洋SSP独立于实时水下数据测量的估计,我们提出了一个半变压器神经网络(STNet),专门设计用于模拟声速分布模式的时间序列预测的角度。建议的网络架构采用了优化的自我注意机制,以有效地捕捉历史声速时间序列数据中的长距离时间依赖关系,便于准确估计当前的SSP或预测未来的SSP。通过对Transformer框架的架构优化和时间编码机制的集成,STNet可以有效地提高计算效率。对比实验结果表明,STNet在预测精度方面优于最先进的模型,并保持良好的计算效率,表明其具有实现准确的长期全深度海洋SSP预测的潜力。
摘要:Real time acquisition of accurate underwater sound velocity profile (SSP) is crucial for tracking the propagation trajectory of underwater acoustic signals, making it play a key role in ocean communication positioning. SSPs can be directly measured by instruments or inverted leveraging sound field data. Although measurement techniques provide a good accuracy, they are constrained by limited spatial coverage and require substantial time investment. The inversion method based on real-time measurement of acoustic field data improves operational efficiency, but loses the accuracy of SSP estimation and suffers from limited spatial applicability due to its stringent requirements for ocean observation infrastructure. To achieve accurate long-term ocean SSP estimation independent of real-time underwater data measurements, we propose a Semi-Transformer neural network (STNet) specifically designed for simulating sound velocity distribution patterns from the perspective of time series prediction. The proposed network architecture incorporates an optimized self-attention mechanism to effectively capture long-range temporal dependencies within historical sound velocity time-series data, facilitating accurate estimation of current SSPs or prediction of future SSPs. Through architectural optimization of the Transformer framework and integration of a time encoding mechanism, STNet could effectively improve computational efficiency. Comparative experimental results reveal that STNet outperforms state-of-the-art models in predictive accuracy and maintain good computational efficiency, demonstrating its potential for enabling accurate long-term full-depth ocean SSP forecasting.


【4】 Kimi-Audio Technical Report

标题: Kimi音频技术报告
链接:https://arxiv.org/abs/2504.18425
作者: KimiTeam,  Ding Ding,  Zeqian Ju,  Yichong Leng,  Songxiang Liu,  Tong Liu,  Zeyu Shang,  Kai Shen,  Wei Song,  Xu Tan,  Heyi Tang,  Zhengtao Wang,  Chu Wei,  Yifei Xin,  Xinran Xu,  Jianwei Yu,  Yutao Zhang,  Xinyu Zhou,  Y. Charles,  Jun Chen,  Yanru Chen,  Yulun Du,  Weiran He,  Zhenxing Hu,  Guokun Lai,  Qingcheng Li,  Yangyang Liu,  Weidong Sun,  Jianzhou Wang,  Yuzhi Wang,  Yuefeng Wu,  Yuxin Wu,  Dongchao Yang,  Hao Yang,  Ying Yang,  Zhilin Yang,  Aoxiong Yin,  Ruibin Yuan,  Yutong Zhang,  Zaida Zhou 
摘要:我们介绍Kimi-Audio,一个开源音频基础模型,在音频理解,生成和对话方面表现出色。我们详细介绍了构建Kimi-Audio的实践,包括模型架构、数据策展、训练配方、推理部署和评估。具体来说,我们利用12.5Hz的音频tokenizer,设计了一种新的基于LLM的架构,连续的功能作为输入和离散的令牌作为输出,并开发了一个基于流匹配的分块流去tokenizer。我们策划了一个由超过1300万小时的音频数据组成的预训练数据集,涵盖了包括语音,声音和音乐在内的各种形式,并建立了一个管道来构建高质量和多样化的后训练数据。Kimi-Audio从预训练的LLM初始化,通过几个精心设计的任务对音频和文本数据进行持续的预训练,然后进行微调以支持各种与音频相关的任务。广泛的评估表明,Kimi-Audio在一系列音频基准测试中达到了最先进的性能,包括语音识别,音频理解,音频问答和语音对话。我们在https://github.com/MoonshotAI/Kimi-Audio上发布代码、模型检查点以及评估工具包。
摘要:We present Kimi-Audio, an open-source audio foundation model that excels in audio understanding, generation, and conversation. We detail the practices in building Kimi-Audio, including model architecture, data curation, training recipe, inference deployment, and evaluation. Specifically, we leverage a 12.5Hz audio tokenizer, design a novel LLM-based architecture with continuous features as input and discrete tokens as output, and develop a chunk-wise streaming detokenizer based on flow matching. We curate a pre-training dataset that consists of more than 13 million hours of audio data covering a wide range of modalities including speech, sound, and music, and build a pipeline to construct high-quality and diverse post-training data. Initialized from a pre-trained LLM, Kimi-Audio is continual pre-trained on both audio and text data with several carefully designed tasks, and then fine-tuned to support a diverse of audio-related tasks. Extensive evaluation shows that Kimi-Audio achieves state-of-the-art performance on a range of audio benchmarks including speech recognition, audio understanding, audio question answering, and speech conversation. We release the codes, model checkpoints, as well as the evaluation toolkits in https://github.com/MoonshotAI/Kimi-Audio.


【5】 DOSE : Drum One-Shot Extraction from Music Mixture

标题: 剂量:从音乐混合物中一次提取鼓
链接:https://arxiv.org/abs/2504.18157
作者: Suntae Hwang,  Seonghyeon Kang,  Kyungsu Kim,  Semin Ahn,  Kyogu Lee 
备注:Published in IEEE ICASSP 2025
摘要:鼓单次采样对于音乐制作至关重要,特别是在声音设计和电子音乐中。本文介绍了鼓单镜头提取,其中的目标是提取鼓单镜头的音乐混合物中存在的任务。为了促进这一点,我们提出了随机混合单次数据集(RMOD),包括大规模的,随机排列的音乐混合与相应的鼓单次样本配对。我们提出的模型Drum One-Shot Extractor(DOSE)利用神经音频编解码器语言模型进行端到端提取,绕过传统的源分离步骤。此外,我们引入了一种新的发病损失,旨在鼓励准确预测的初始瞬态鼓一杆,这是必不可少的捕捉音色特征。我们将这种方法与基于源分离的提取方法作为基线进行比较。使用Frechet音频距离(FAD)和多尺度频谱损失(MSS)评估的结果表明,DOSE,增强了起始损失,优于基线,从音乐混合物中提供更准确和更高质量的鼓单镜头。代码、模型检查点和音频示例可在https://github.com/HSUNEH/DOSE上获得
摘要:Drum one-shot samples are crucial for music production, particularly in sound design and electronic music. This paper introduces Drum One-Shot Extraction, a task in which the goal is to extract drum one-shots that are present in the music mixture. To facilitate this, we propose the Random Mixture One-shot Dataset (RMOD), comprising large-scale, randomly arranged music mixtures paired with corresponding drum one-shot samples. Our proposed model, Drum One- Shot Extractor (DOSE), leverages neural audio codec language models for end-to-end extraction, bypassing traditional source separation steps. Additionally, we introduce a novel onset loss, designed to encourage accurate prediction of the initial transient of drum one-shots, which is essential for capturing timbral characteristics. We compare this approach against a source separation-based extraction method as a baseline. The results, evaluated using Frechet Audio Distance (FAD) and Multi-Scale Spectral loss (MSS), demonstrate that DOSE, enhanced with onset loss, outperforms the baseline, providing more accurate and higher-quality drum one-shots from music mixtures. The code, model checkpoint, and audio examples are available at https://github.com/HSUNEH/DOSE


【6】 Assessing the Utility of Audio Foundation Models for Heart and  Respiratory Sound Analysis

标题: 评估音频基金会模型用于心脏和呼吸声分析的实用性
链接:https://arxiv.org/abs/2504.18004
作者: Daisuke Niizumi,  Daiki Takeuchi,  Masahiro Yasuda,  Binh Thien Nguyen,  Yasunori Ohishi,  Noboru Harada 
备注:4 pages, 1 figure, and 4 tables. Accepted by IEEE EMBC 2025
摘要:预训练的深度学习模型,称为基础模型,已成为机器学习领域(如自然语言处理和图像领域)的重要构建模块。这种趋势已经扩展到呼吸和心音模型,这些模型已经证明了作为现成的特征提取器的有效性。然而,它们的评价基准受到限制,导致与最新技术水平(SOTA)性能不相容,从而妨碍证明它们的有效性。本研究通过比较现有音频基础模型在四个呼吸和心脏声音任务中的性能与SOTA微调结果,研究了它们的实际有效性。实验表明,模型在两个有噪声数据的任务上表现不佳,但在其他有干净数据的任务上达到了SOTA性能。此外,通用音频模型优于呼吸声模型,突出了其更广泛的适用性。通过获得的见解和发布的代码,我们为未来开发和利用呼吸和心音基础模型的研究做出了贡献。
摘要:Pre-trained deep learning models, known as foundation models, have become essential building blocks in machine learning domains such as natural language processing and image domains. This trend has extended to respiratory and heart sound models, which have demonstrated effectiveness as off-the-shelf feature extractors. However, their evaluation benchmarking has been limited, resulting in incompatibility with state-of-the-art (SOTA) performance, thus hindering proof of their effectiveness. This study investigates the practical effectiveness of off-the-shelf audio foundation models by comparing their performance across four respiratory and heart sound tasks with SOTA fine-tuning results. Experiments show that models struggled on two tasks with noisy data but achieved SOTA performance on the other tasks with clean data. Moreover, general-purpose audio models outperformed a respiratory sound model, highlighting their broader applicability. With gained insights and the released code, we contribute to future research on developing and leveraging foundation models for respiratory and heart sounds.


eess.AS音频处理

【1】 Music Tempo Estimation on Solo Instrumental Performance
标题: 器乐独奏演奏的音乐节奏估算
链接:https://arxiv.org/abs/2504.18502
作者: Zhanhong He,  Roberto Togneri,  Xiangyu Zhang 
备注:4 pages, rejected paper by WASPAA2023
摘要:最近,自动音乐转录使得将音乐音频转换为准确的音频成为可能。然而,最终的乐谱缺乏音乐符号,如节奏,这阻碍了它转换成乐谱。在本文中,我们调查国家的最先进的速度估计技术,并评估其性能独奏器乐。这些包括时间卷积网络(TCN)和递归神经网络(RNN)模型,这些模型是在大量混合人声和器乐上进行预训练的,以及专门用独奏乐器表演训练的TCN模型。通过对鼓,吉他和古典钢琴数据集的评估,我们的TCN模型与新的训练方案实现了最佳性能。我们新训练的TCN模型将吉他速度估计的Acc1指标提高了38.6%,而预训练的TCN模型的Acc1为61.1%。虽然我们训练的TCN模型在估计古典钢琴速度方面的准确度是预训练的TCN模型的两倍,但其Acc1仅为50.9%。为了提高深度学习模型的性能,我们研究了它们与各种后处理方法的组合。当深度学习模型难以估计特定乐器的节奏时,这些后处理技术有效地提高了深度学习模型的性能。
摘要:Recently, automatic music transcription has made it possible to convert musical audio into accurate MIDI. However, the resulting MIDI lacks music notations such as tempo, which hinders its conversion into sheet music. In this paper, we investigate state-of-the-art tempo estimation techniques and evaluate their performance on solo instrumental music. These include temporal convolutional network (TCN) and recurrent neural network (RNN) models that are pretrained on massive of mixed vocals and instrumental music, as well as TCN models trained specifically with solo instrumental performances. Through evaluations on drum, guitar, and classical piano datasets, our TCN models with the new training scheme achieved the best performance. Our newly trained TCN model increases the Acc1 metric by 38.6% for guitar tempo estimation, compared to the pretrained TCN model with an Acc1 of 61.1%. Although our trained TCN model is twice as accurate as the pretrained TCN model in estimating classical piano tempo, its Acc1 is only 50.9%. To improve the performance of deep learning models, we investigate their combinations with various post-processing methods. These post-processing techniques effectively enhance the performance of deep learning models when they struggle to estimate the tempo of specific instruments.


【2】 Kimi-Audio Technical Report

标题: Kimi音频技术报告
链接:https://arxiv.org/abs/2504.18425
作者: KimiTeam,  Ding Ding,  Zeqian Ju,  Yichong Leng,  Songxiang Liu,  Tong Liu,  Zeyu Shang,  Kai Shen,  Wei Song,  Xu Tan,  Heyi Tang,  Zhengtao Wang,  Chu Wei,  Yifei Xin,  Xinran Xu,  Jianwei Yu,  Yutao Zhang,  Xinyu Zhou,  Y. Charles,  Jun Chen,  Yanru Chen,  Yulun Du,  Weiran He,  Zhenxing Hu,  Guokun Lai,  Qingcheng Li,  Yangyang Liu,  Weidong Sun,  Jianzhou Wang,  Yuzhi Wang,  Yuefeng Wu,  Yuxin Wu,  Dongchao Yang,  Hao Yang,  Ying Yang,  Zhilin Yang,  Aoxiong Yin,  Ruibin Yuan,  Yutong Zhang,  Zaida Zhou 
摘要:我们介绍Kimi-Audio,一个开源音频基础模型,在音频理解,生成和对话方面表现出色。我们详细介绍了构建Kimi-Audio的实践,包括模型架构、数据策展、训练配方、推理部署和评估。具体来说,我们利用12.5Hz的音频tokenizer,设计了一种新的基于LLM的架构,连续的功能作为输入和离散的令牌作为输出,并开发了一个基于流匹配的分块流去tokenizer。我们策划了一个由超过1300万小时的音频数据组成的预训练数据集,涵盖了包括语音,声音和音乐在内的各种形式,并建立了一个管道来构建高质量和多样化的后训练数据。Kimi-Audio从预训练的LLM初始化,通过几个精心设计的任务对音频和文本数据进行持续的预训练,然后进行微调以支持各种与音频相关的任务。广泛的评估表明,Kimi-Audio在一系列音频基准测试中达到了最先进的性能,包括语音识别,音频理解,音频问答和语音对话。我们在https://github.com/MoonshotAI/Kimi-Audio上发布代码、模型检查点以及评估工具包。
摘要:We present Kimi-Audio, an open-source audio foundation model that excels in audio understanding, generation, and conversation. We detail the practices in building Kimi-Audio, including model architecture, data curation, training recipe, inference deployment, and evaluation. Specifically, we leverage a 12.5Hz audio tokenizer, design a novel LLM-based architecture with continuous features as input and discrete tokens as output, and develop a chunk-wise streaming detokenizer based on flow matching. We curate a pre-training dataset that consists of more than 13 million hours of audio data covering a wide range of modalities including speech, sound, and music, and build a pipeline to construct high-quality and diverse post-training data. Initialized from a pre-trained LLM, Kimi-Audio is continual pre-trained on both audio and text data with several carefully designed tasks, and then fine-tuned to support a diverse of audio-related tasks. Extensive evaluation shows that Kimi-Audio achieves state-of-the-art performance on a range of audio benchmarks including speech recognition, audio understanding, audio question answering, and speech conversation. We release the codes, model checkpoints, as well as the evaluation toolkits in https://github.com/MoonshotAI/Kimi-Audio.


【3】 DOSE : Drum One-Shot Extraction from Music Mixture

标题: 剂量:从音乐混合物中一次提取鼓
链接:https://arxiv.org/abs/2504.18157
作者: Suntae Hwang,  Seonghyeon Kang,  Kyungsu Kim,  Semin Ahn,  Kyogu Lee 
备注:Published in IEEE ICASSP 2025
摘要:None
摘要:Drum one-shot samples are crucial for music production, particularly in sound design and electronic music. This paper introduces Drum One-Shot Extraction, a task in which the goal is to extract drum one-shots that are present in the music mixture. To facilitate this, we propose the Random Mixture One-shot Dataset (RMOD), comprising large-scale, randomly arranged music mixtures paired with corresponding drum one-shot samples. Our proposed model, Drum One- Shot Extractor (DOSE), leverages neural audio codec language models for end-to-end extraction, bypassing traditional source separation steps. Additionally, we introduce a novel onset loss, designed to encourage accurate prediction of the initial transient of drum one-shots, which is essential for capturing timbral characteristics. We compare this approach against a source separation-based extraction method as a baseline. The results, evaluated using Frechet Audio Distance (FAD) and Multi-Scale Spectral loss (MSS), demonstrate that DOSE, enhanced with onset loss, outperforms the baseline, providing more accurate and higher-quality drum one-shots from music mixtures. The code, model checkpoint, and audio examples are available at https://github.com/HSUNEH/DOSE


【4】 Assessing the Utility of Audio Foundation Models for Heart and  Respiratory Sound Analysis

标题: 评估音频基金会模型用于心脏和呼吸声分析的实用性
链接:https://arxiv.org/abs/2504.18004
作者: Daisuke Niizumi,  Daiki Takeuchi,  Masahiro Yasuda,  Binh Thien Nguyen,  Yasunori Ohishi,  Noboru Harada 
备注:4 pages, 1 figure, and 4 tables. Accepted by IEEE EMBC 2025
摘要:预训练的深度学习模型,称为基础模型,已成为机器学习领域(如自然语言处理和图像领域)的重要构建模块。这种趋势已经扩展到呼吸和心音模型,这些模型已经证明了作为现成的特征提取器的有效性。然而,它们的评价基准受到限制,导致与最新技术水平(SOTA)性能不相容,从而妨碍证明它们的有效性。本研究通过比较现有音频基础模型在四个呼吸和心脏声音任务中的性能与SOTA微调结果,研究了它们的实际有效性。实验表明,模型在两个有噪声数据的任务上表现不佳,但在其他有干净数据的任务上达到了SOTA性能。此外,通用音频模型优于呼吸声模型,突出了其更广泛的适用性。通过获得的见解和发布的代码,我们为未来开发和利用呼吸和心音基础模型的研究做出了贡献。
摘要:Pre-trained deep learning models, known as foundation models, have become essential building blocks in machine learning domains such as natural language processing and image domains. This trend has extended to respiratory and heart sound models, which have demonstrated effectiveness as off-the-shelf feature extractors. However, their evaluation benchmarking has been limited, resulting in incompatibility with state-of-the-art (SOTA) performance, thus hindering proof of their effectiveness. This study investigates the practical effectiveness of off-the-shelf audio foundation models by comparing their performance across four respiratory and heart sound tasks with SOTA fine-tuning results. Experiments show that models struggled on two tasks with noisy data but achieved SOTA performance on the other tasks with clean data. Moreover, general-purpose audio models outperformed a respiratory sound model, highlighting their broader applicability. With gained insights and the released code, we contribute to future research on developing and leveraging foundation models for respiratory and heart sounds.


【5】 Seeing Soundscapes: Audio-Visual Generation and Separation from  Soundscapes Using Audio-Visual Separator

标题: 看到声景:视听生成和使用视听分离器从声景中分离
链接:https://arxiv.org/abs/2504.18283
作者: Minjae Kang,  Martim Brandão 
备注:Originally submitted to CVPR 2025 on 2024-11-15 with paper ID 15808
摘要:最近的视听生成模型在从音频生成图像方面取得了实质性的进展。然而,现有的方法集中于从单类音频生成图像,而不能从混合音频生成图像。为了解决这个问题,我们提出了一个视听生成和分离模型(AV-GAS),用于从音景(包含多个类的混合音频)生成图像。我们的贡献有三个方面:首先,我们提出了一个新的挑战,在视听生成任务,这是生成一个图像给定的多类音频输入,我们提出了一种方法,解决了这个任务,使用视听分离器。其次,我们引入了一个新的视听分离任务,它涉及到为混合音频输入中的每个类生成单独的图像。最后,我们提出了新的评估指标的视听生成任务:类表示分数(CRS)和修改后的R@K。我们的模型是在VGGSound数据集上训练和评估的。我们证明了我们的方法优于最先进的方法,在生成具有混合音频的合理图像时,CRS提高了7%,R@2* 提高了4%。
摘要:Recent audio-visual generative models have made substantial progress in generating images from audio. However, existing approaches focus on generating images from single-class audio and fail to generate images from mixed audio. To address this, we propose an Audio-Visual Generation and Separation model (AV-GAS) for generating images from soundscapes (mixed audio containing multiple classes). Our contribution is threefold: First, we propose a new challenge in the audio-visual generation task, which is to generate an image given a multi-class audio input, and we propose a method that solves this task using an audio-visual separator. Second, we introduce a new audio-visual separation task, which involves generating separate images for each class present in a mixed audio input. Lastly, we propose new evaluation metrics for the audio-visual generation task: Class Representation Score (CRS) and a modified R@K. Our model is trained and evaluated on the VGGSound dataset. We show that our method outperforms the state-of-the-art, achieving 7% higher CRS and 4% higher R@2* in generating plausible images with mixed audio.


【6】 Tracking Articulatory Dynamics in Speech with a Fixed-Weight BiLSTM-CNN  Architecture

标题: 使用固定权重BiLSTM-CNN架构跟踪语音中的关节动态
链接:https://arxiv.org/abs/2504.18099
作者: Leena G Pillai,  D. Muhammad Noorul Mubarak,  Elizabeth Sherly 
备注:10 pages with 8 figures. This paper presented in an international Conference
摘要:言语产生是一个复杂的连续过程,涉及到各种发音特征的协调。其中,舌头是一个高度灵活的主动发音器官,负责塑造气流,以产生有针对性的语音,是智力,清晰,独特的。本文提出了一种新的方法来预测舌头和嘴唇发音特征涉及一个给定的语音声学使用堆叠双向长短期记忆(BiLSTM)架构,结合一维卷积神经网络(CNN)的后处理与固定权重初始化。建议的网络是用两个数据集训练的,这两个数据集由同时记录的语音和电磁关节成像(EMA)数据集组成,每个数据集都引入了地理起源,语言特征,语音多样性和记录设备方面的变化。该模型的性能进行了评估,在说话人相关(SD),说话人无关(SI),语料库相关(CD)和跨语料库(CC)模式。实验结果表明,所提出的模型与固定的权重方法优于自适应权重初始化在相对最少的训练时期。这些研究结果有助于发展强大而有效的发音特征预测模型,为语音产生研究和应用的进步铺平了道路。
摘要:Speech production is a complex sequential process which involve the coordination of various articulatory features. Among them tongue being a highly versatile active articulator responsible for shaping airflow to produce targeted speech sounds that are intellectual, clear, and distinct. This paper presents a novel approach for predicting tongue and lip articulatory features involved in a given speech acoustics using a stacked Bidirectional Long Short-Term Memory (BiLSTM) architecture, combined with a one-dimensional Convolutional Neural Network (CNN) for post-processing with fixed weights initialization. The proposed network is trained with two datasets consisting of simultaneously recorded speech and Electromagnetic Articulography (EMA) datasets, each introducing variations in terms of geographical origin, linguistic characteristics, phonetic diversity, and recording equipment. The performance of the model is assessed in Speaker Dependent (SD), Speaker Independent (SI), corpus dependent (CD) and cross corpus (CC) modes. Experimental results indicate that the proposed model with fixed weights approach outperformed the adaptive weights initialization with in relatively minimal number of training epochs. These findings contribute to the development of robust and efficient models for articulatory feature prediction, paving the way for advancements in speech production research and applications.


【7】 STNet: Prediction of Underwater Sound Speed Profiles with An Advanced  Semi-Transformer Neural Network

标题: STNet:使用先进的半Transformer神经网络预测水下音速轮廓
链接:https://arxiv.org/abs/2504.17912
作者: Wei Huang,  Jiajun Lu,  Hao Zhang,  Tianhe Xu 
摘要:实时获取准确的水下声速剖面是跟踪水声信号传播轨迹的关键,在海洋通信定位中发挥着关键作用。SSP可以通过仪器直接测量或利用声场数据进行反演。虽然测量技术提供了良好的精度,但它们受到有限的空间覆盖范围的限制,并且需要大量的时间投资。基于声场数据实时测量的反演方法提高了业务效率,但由于其对海洋观测基础设施的严格要求,失去了SSP估计的准确性,并且受到有限的空间适用性的影响。为了实现准确的长期海洋SSP独立于实时水下数据测量的估计,我们提出了一个半变压器神经网络(STNet),专门设计用于模拟声速分布模式的时间序列预测的角度。建议的网络架构采用了优化的自我注意机制,以有效地捕捉历史声速时间序列数据中的长距离时间依赖关系,便于准确估计当前的SSP或预测未来的SSP。通过对Transformer框架的架构优化和时间编码机制的集成,STNet可以有效地提高计算效率。对比实验结果表明,STNet在预测精度方面优于最先进的模型,并保持良好的计算效率,表明其具有实现准确的长期全深度海洋SSP预测的潜力。
摘要:Real time acquisition of accurate underwater sound velocity profile (SSP) is crucial for tracking the propagation trajectory of underwater acoustic signals, making it play a key role in ocean communication positioning. SSPs can be directly measured by instruments or inverted leveraging sound field data. Although measurement techniques provide a good accuracy, they are constrained by limited spatial coverage and require substantial time investment. The inversion method based on real-time measurement of acoustic field data improves operational efficiency, but loses the accuracy of SSP estimation and suffers from limited spatial applicability due to its stringent requirements for ocean observation infrastructure. To achieve accurate long-term ocean SSP estimation independent of real-time underwater data measurements, we propose a Semi-Transformer neural network (STNet) specifically designed for simulating sound velocity distribution patterns from the perspective of time series prediction. The proposed network architecture incorporates an optimized self-attention mechanism to effectively capture long-range temporal dependencies within historical sound velocity time-series data, facilitating accurate estimation of current SSPs or prediction of future SSPs. Through architectural optimization of the Transformer framework and integration of a time encoding mechanism, STNet could effectively improve computational efficiency. Comparative experimental results reveal that STNet outperforms state-of-the-art models in predictive accuracy and maintain good computational efficiency, demonstrating its potential for enabling accurate long-term full-depth ocean SSP forecasting.


机器翻译由腾讯交互翻译提供,仅供参考