本文经arXiv每日学术速递授权转载
cs.SD语音
【1】 Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction
标题: 利用LLM嵌入进行跨数据集标签对齐和零镜头音乐情感预测
作者: Renhang Liu, Abhinaba Roy, Dorien Herremans
链接:点击下载PDF文件
【2】 Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
标题: Dist-SAGee:使用扩散模型的端到端空间音频生成
作者: Saksham Singh Kushwaha, Jianbo Ma, Mark R. P. Thomas, Yapeng Tian, Avery Bruni
链接:点击下载PDF文件
【3】 Investigation of Speaker Representation for Target-Speaker Speech Processing
标题: 目标说话人语音处理中说话人表示的研究
作者: Takanori Ashihara, Takafumi Moriya, Shota Horiguchi, Junyi Peng, Tsubasa Ochiai, Marc Delcroix, Kohei Matsuura, Hiroshi Sato
备注:Accepted at IEEE SLT 2024
链接:点击下载PDF文件
【4】 Audio-based Kinship Verification Using Age Domain Conversion
标题: 使用年龄域转换的音频亲属关系验证
作者: Qiyang Sun, Alican Akman, Xin Jing, Manuel Milling, Björn W. Schuller
备注:4 pages, 2 figures, submitted to IEEE Signal Processing Letters
链接:点击下载PDF文件
【5】 Character-aware audio-visual subtitling in context
标题: 背景下的用户感知视听字幕
作者: Jaesung Huh, Andrew Zisserman
备注:ACCV 2024
链接:点击下载PDF文件
【6】 CleanUMamba: A Compact Mamba Network for Speech Denoising using Channel Pruning
标题: CleanUmamba:一个紧凑的Mamba网络,用于使用通道修剪进行语音降噪
作者: Sjoerd Groot, Qinyu Chen, Jan C. van Gemert, Chang Gao
备注:5 pages, 4 figures
链接:点击下载PDF文件
【7】 GraFPrint: A GNN-Based Approach for Audio Identification
标题: GraFPrint:一种基于GNN的音频识别方法
作者: Aditya Bhattacharjee, Shubhr Singh, Emmanouil Benetos
备注:Submitted to IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2025)
链接:点击下载PDF文件
【8】 Audio Captioning via Generative Pair-to-Pair Retrieval with Refined Knowledge Base
标题: 通过具有精细知识库的生成性成对检索的音频字幕
作者: Choi Changin, Lim Sungjun, Rhee Wonjong
链接:点击下载PDF文件
【9】 LLM Gesticulator: Leveraging Large Language Models for Scalable and Controllable Co-Speech Gesture Synthesis
标题: LLM Gesticulator:利用大型语言模型进行可扩展和可控的同声手势合成
作者: Haozhou Pang, Tianwei Ding, Lanshan He, Qi Gan
链接:点击下载PDF文件
【10】 The importance of spatial and spectral information in multiple speaker tracking
标题: 空间和频谱信息在多说话人跟踪中的重要性
作者: Hanan Beit-On, Vladimir Tourbabin, Boaz Rafaely
链接:点击下载PDF文件
【11】 Mini-Omni2: Towards Open-source GPT-4o Model with Vision, Speech and Duplex
标题: Mini-Omni 2:迈向具有视觉、语音和双重功能的开源GPT-4 o模型
作者: Zhifei Xie, Changqiao Wu
备注:13 pages, 6 figures
链接:点击下载PDF文件
【12】 DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection
标题: DARNet:具有时空结构的双重注意细化网络,用于听觉注意检测
作者: Sheng Yan, Cunhang fan, Hongyu Zhang, Xiaoke Yang, Jianhua Tao, Zhao Lv
链接:点击下载PDF文件
【13】 DMDSpeech: Distilled Diffusion Model Surpassing The Teacher in Zero-shot Speech Synthesis via Direct Metric Optimization
标题: DMZ Speech:通过直接度量优化在Zero-Shot语音合成中超越教师的蒸馏扩散模型
作者: Yingahao Aaron Li, Rithesh Kumar, Zeyu Jin
链接:点击下载PDF文件
【14】 Code Drift: Towards Idempotent Neural Audio Codecs
标题: 代码漂移:迈向等能神经音频编解码器
作者: Patrick O'Reilly, Prem Seetharaman, Jiaqi Su, Zeyu Jin, Bryan Pardo
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
标题: 空间和频谱信息在多说话人跟踪中的重要性
作者: Hanan Beit-On, Vladimir Tourbabin, Boaz Rafaely
链接:点击下载PDF文件
【2】 Mini-Omni2: Towards Open-source GPT-4o Model with Vision, Speech and Duplex
标题: Mini-Omni 2:迈向具有视觉、语音和双重功能的开源GPT-4 o模型
作者: Zhifei Xie, Changqiao Wu
备注:13 pages, 6 figures
链接:点击下载PDF文件
【3】 DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection
标题: DARNet:具有时空结构的双重注意细化网络,用于听觉注意检测
作者: Sheng Yan, Cunhang fan, Hongyu Zhang, Xiaoke Yang, Jianhua Tao, Zhao Lv
链接:点击下载PDF文件
【4】 DMDSpeech: Distilled Diffusion Model Surpassing The Teacher in Zero-shot Speech Synthesis via Direct Metric Optimization
标题: DMZ Speech:通过直接度量优化在Zero-Shot语音合成中超越教师的蒸馏扩散模型
作者: Yingahao Aaron Li, Rithesh Kumar, Zeyu Jin
链接:点击下载PDF文件
【5】 Code Drift: Towards Idempotent Neural Audio Codecs
标题: 代码漂移:迈向等能神经音频编解码器
作者: Patrick O'Reilly, Prem Seetharaman, Jiaqi Su, Zeyu Jin, Bryan Pardo
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【6】 Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction
标题: 利用LLM嵌入进行跨数据集标签对齐和零镜头音乐情感预测
作者: Renhang Liu, Abhinaba Roy, Dorien Herremans
链接:点击下载PDF文件
【7】 Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
标题: Dist-SAGee:使用扩散模型的端到端空间音频生成
作者: Saksham Singh Kushwaha, Jianbo Ma, Mark R. P. Thomas, Yapeng Tian, Avery Bruni
链接:点击下载PDF文件
【8】 Investigation of Speaker Representation for Target-Speaker Speech Processing
标题: 目标说话人语音处理中说话人表示的研究
作者: Takanori Ashihara, Takafumi Moriya, Shota Horiguchi, Junyi Peng, Tsubasa Ochiai, Marc Delcroix, Kohei Matsuura, Hiroshi Sato
备注:Accepted at IEEE SLT 2024
链接:点击下载PDF文件
【9】 Audio-based Kinship Verification Using Age Domain Conversion
标题: 使用年龄域转换的音频亲属关系验证
作者: Qiyang Sun, Alican Akman, Xin Jing, Manuel Milling, Björn W. Schuller
备注:4 pages, 2 figures, submitted to IEEE Signal Processing Letters
链接:点击下载PDF文件
【10】 Character-aware audio-visual subtitling in context
标题: 背景下的用户感知视听字幕
作者: Jaesung Huh, Andrew Zisserman
备注:ACCV 2024
链接:点击下载PDF文件
【11】 CleanUMamba: A Compact Mamba Network for Speech Denoising using Channel Pruning
标题: CleanUmamba:一个紧凑的Mamba网络,用于使用通道修剪进行语音降噪
作者: Sjoerd Groot, Qinyu Chen, Jan C. van Gemert, Chang Gao
备注:5 pages, 4 figures
链接:点击下载PDF文件
【12】 GraFPrint: A GNN-Based Approach for Audio Identification
标题: GraFPrint:一种基于GNN的音频识别方法
作者: Aditya Bhattacharjee, Shubhr Singh, Emmanouil Benetos
备注:Submitted to IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2025)
链接:点击下载PDF文件
【13】 Audio Captioning via Generative Pair-to-Pair Retrieval with Refined Knowledge Base
标题: 通过具有精细知识库的生成性成对检索的音频字幕
作者: Choi Changin, Lim Sungjun, Rhee Wonjong
链接:点击下载PDF文件
【14】 LLM Gesticulator: Leveraging Large Language Models for Scalable and Controllable Co-Speech Gesture Synthesis
标题: LLM Gesticulator:利用大型语言模型进行可扩展和可控的同声手势合成
作者: Haozhou Pang, Tianwei Ding, Lanshan He, Qi Gan
链接:点击下载PDF文件
标题: 利用LLM嵌入进行跨数据集标签对齐和零镜头音乐情感预测
作者: Renhang Liu, Abhinaba Roy, Dorien Herremans
链接:点击下载PDF文件
摘要:在这项工作中,我们提出了一种新的方法,音乐情感识别,利用大语言模型(LLM)嵌入标签对齐跨多个数据集和zero-shot预测新的类别。首先,我们计算情感标签的LLM嵌入,并在包含不相交标签的多个数据集上应用非参数聚类对相似标签进行分组。我们使用这些聚类中心将音乐特征(MERT)映射到LLM嵌入空间。为了进一步增强模型,我们引入了一个对齐正则化,使MERT嵌入从不同的集群分离。这进一步增强了模型更好地适应未知数据集的能力。我们通过对一个新的数据集进行zero-shot推理来证明我们的方法的有效性,展示了它在没有额外训练的情况下推广到看不见的标签的能力。摘要:In this work, we present a novel method for music emotion recognition that leverages Large Language Model (LLM) embeddings for label alignment across multiple datasets and zero-shot prediction on novel categories. First, we compute LLM embeddings for emotion labels and apply non-parametric clustering to group similar labels, across multiple datasets containing disjoint labels. We use these cluster centers to map music features (MERT) to the LLM embedding space. To further enhance the model, we introduce an alignment regularization that enables dissociation of MERT embeddings from different clusters. This further enhances the model's ability to better adaptation to unseen datasets. We demonstrate the effectiveness of our approach by performing zero-shot inference on a new dataset, showcasing its ability to generalize to unseen labels without additional training.
【2】 Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
标题: Dist-SAGee:使用扩散模型的端到端空间音频生成
作者: Saksham Singh Kushwaha, Jianbo Ma, Mark R. P. Thomas, Yapeng Tian, Avery Bruni
链接:点击下载PDF文件
摘要:空间音频是创造沉浸式体验的关键组成部分。传统的基于模拟的方法来生成空间音频依赖于专业知识,具有有限的可扩展性,并假设语义和空间信息之间的独立性。为了解决这些问题,我们探索端到端的空间音频生成。我们介绍并制定了一个新的任务,产生一阶高保真度立体声(FOA)给定的声音类别和声源的空间位置。我们提出了Diff-SAGe,一个端到端的,基于流的扩散变压器模型,这项任务。Diff-SAGe利用复杂的频谱图表示FOA,保留相位信息的准确空间线索至关重要。此外,多条件编码器将输入条件集成到统一的表示中,指导从噪声中生成FOA波形。通过对两个数据集的广泛评估,我们证明了我们的方法在客观和主观指标上始终优于传统的基于模拟的基线。摘要:Spatial audio is a crucial component in creating immersive experiences. Traditional simulation-based approaches to generate spatial audio rely on expertise, have limited scalability, and assume independence between semantic and spatial information. To address these issues, we explore end-to-end spatial audio generation. We introduce and formulate a new task of generating first-order Ambisonics (FOA) given a sound category and sound source spatial location. We propose Diff-SAGe, an end-to-end, flow-based diffusion-transformer model for this task. Diff-SAGe utilizes a complex spectrogram representation for FOA, preserving the phase information crucial for accurate spatial cues. Additionally, a multi-conditional encoder integrates the input conditions into a unified representation, guiding the generation of FOA waveforms from noise. Through extensive evaluations on two datasets, we demonstrate that our method consistently outperforms traditional simulation-based baselines across both objective and subjective metrics.
【3】 Investigation of Speaker Representation for Target-Speaker Speech Processing
标题: 目标说话人语音处理中说话人表示的研究
作者: Takanori Ashihara, Takafumi Moriya, Shota Horiguchi, Junyi Peng, Tsubasa Ochiai, Marc Delcroix, Kohei Matsuura, Hiroshi Sato
备注:Accepted at IEEE SLT 2024
链接:点击下载PDF文件
摘要:目标说话人语音处理(TS)任务,诸如目标说话人自动语音识别(TS-ASR)、目标语音提取(TSE)和个人语音活动检测(p-VAD),对于提取关于期望说话人的语音的信息是重要的,即使当其被干扰说话人破坏时。虽然大多数研究都集中在每个特定任务的训练方案或系统架构,但用于嵌入目标说话人线索的辅助网络尚未在统一的跨任务评估中进行全面研究。因此,本文旨在解决一个基本问题:什么是首选的说话人嵌入TS任务?为此,对于TS-ASR、TSE和p-VAD任务,我们比较预先训练的说话者编码器(即,自监督或说话者识别模型),其从目标说话者的预先记录的登记语音计算说话者嵌入,其中理想说话者嵌入以独热向量的形式直接从目标说话者的身份导出。为了进一步了解理想说话人嵌入的特性,我们使用基于梯度的方法对其进行优化,以提高TS任务的性能。我们的分析表明,说话人验证性能是有点无关TS任务的性能,一个热向量优于注册为基础的,和最佳的嵌入依赖于输入的混合。摘要:Target-speaker speech processing (TS) tasks, such as target-speaker automatic speech recognition (TS-ASR), target speech extraction (TSE), and personal voice activity detection (p-VAD), are important for extracting information about a desired speaker's speech even when it is corrupted by interfering speakers. While most studies have focused on training schemes or system architectures for each specific task, the auxiliary network for embedding target-speaker cues has not been investigated comprehensively in a unified cross-task evaluation. Therefore, this paper aims to address a fundamental question: what is the preferred speaker embedding for TS tasks? To this end, for the TS-ASR, TSE, and p-VAD tasks, we compare pre-trained speaker encoders (i.e., self-supervised or speaker recognition models) that compute speaker embeddings from pre-recorded enrollment speech of the target speaker with ideal speaker embeddings derived directly from the target speaker's identity in the form of a one-hot vector. To further understand the properties of ideal speaker embedding, we optimize it using a gradient-based approach to improve performance on the TS task. Our analysis reveals that speaker verification performance is somewhat unrelated to TS task performances, the one-hot vector outperforms enrollment-based ones, and the optimal embedding depends on the input mixture.
【4】 Audio-based Kinship Verification Using Age Domain Conversion
标题: 使用年龄域转换的音频亲属关系验证
作者: Qiyang Sun, Alican Akman, Xin Jing, Manuel Milling, Björn W. Schuller
备注:4 pages, 2 figures, submitted to IEEE Signal Processing Letters
链接:点击下载PDF文件
摘要:基于音频的亲属关系验证(AKV)在许多领域都很重要,例如家庭安全监控,法医鉴定和社交网络分析。任务中的一个关键挑战来自于来自不同个体的样本之间的年龄差异,这可以被解释为跨域验证任务中的域偏差。为了解决这个问题,我们设计了“年龄标准化域”的概念,其中我们利用优化的CycleGAN-VC 3网络来执行年龄音频转换以生成域内音频。生成的音频数据集用于提取一系列特征,然后将其输入到度量学习架构中以验证亲属关系。实验是在KAN_AV音频数据集上进行的,该数据集包含年龄和亲属标签。实验结果表明,该方法显著提高了亲属关系验证的准确性,同时也为未来的亲属关系验证研究提供了新的见解。摘要:Audio-based kinship verification (AKV) is important in many domains, such as home security monitoring, forensic identification, and social network analysis. A key challenge in the task arises from differences in age across samples from different individuals, which can be interpreted as a domain bias in a cross-domain verification task. To address this issue, we design the notion of an "age-standardised domain" wherein we utilise the optimised CycleGAN-VC3 network to perform age-audio conversion to generate the in-domain audio. The generated audio dataset is employed to extract a range of features, which are then fed into a metric learning architecture to verify kinship. Experiments are conducted on the KAN_AV audio dataset, which contains age and kinship labels. The results demonstrate that the method markedly enhances the accuracy of kinship verification, while also offering novel insights for future kinship verification research.
【5】 Character-aware audio-visual subtitling in context
标题: 背景下的用户感知视听字幕
作者: Jaesung Huh, Andrew Zisserman
备注:ACCV 2024
链接:点击下载PDF文件
摘要:本文提出了一种改进的框架,在电视节目中的字符感知视听字幕。我们的方法集成了语音识别,扬声器diarisation,字符识别,利用音频和视觉线索。这个整体解决方案解决了说什么、什么时候说以及谁在说话的问题,为电视节目提供了更全面、更准确的角色感知字幕。我们的方法在两个方面带来了改进:首先,我们表明,视听同步可以用来挑选出在视频剪辑中存在的其他人之间的说话的脸,并分配相应的语音段的身份。与当前方法相比,这种视听方法提高了识别准确性和产量。其次,我们表明,短片段的扬声器可以通过使用场景内的对话的时间上下文来确定。我们提出了一种方法,使用本地语音嵌入的音频,和大语言模型推理的文本转录。这克服了现有方法的局限性,即它们不能准确地将扬声器分配给短时间段。我们验证了12个电视节目的数据集上的方法,表现出优越的性能相比,现有的方法在扬声器diarisation和字符识别的准确性。项目页面:https: www.robots.ox.ac.uk ~vgg research llr-context 摘要:This paper presents an improved framework for character-aware audio-visual subtitling in TV shows. Our approach integrates speech recognition, speaker diarisation, and character recognition, utilising both audio and visual cues. This holistic solution addresses what is said, when it's said, and who is speaking, providing a more comprehensive and accurate character-aware subtitling for TV shows. Our approach brings improvements on two fronts: first, we show that audio-visual synchronisation can be used to pick out the talking face amongst others present in a video clip, and assign an identity to the corresponding speech segment. This audio-visual approach improves recognition accuracy and yield over current methods. Second, we show that the speaker of short segments can be determined by using the temporal context of the dialogue within a scene. We propose an approach using local voice embeddings of the audio, and large language model reasoning on the text transcription. This overcomes a limitation of existing methods that they are unable to accurately assign speakers to short temporal segments. We validate the method on a dataset with 12 TV shows, demonstrating superior performance in speaker diarisation and character recognition accuracy compared to existing approaches. Project page : https: www.robots.ox.ac.uk ~vgg research llr-context
【6】 CleanUMamba: A Compact Mamba Network for Speech Denoising using Channel Pruning
标题: CleanUmamba:一个紧凑的Mamba网络,用于使用通道修剪进行语音降噪
作者: Sjoerd Groot, Qinyu Chen, Jan C. van Gemert, Chang Gao
备注:5 pages, 4 figures
链接:点击下载PDF文件
摘要:本文介绍了CleanUAmba,一种时域神经网络架构,用于直接应用于原始波形的实时因果音频去噪。CleanUAmba利用了U-Net编码器-解码器结构,在瓶颈层中结合了Mamba状态空间模型。通过用Mamba替换传统的自注意力和LSTM机制,我们的架构提供了卓越的去噪性能,同时保持了恒定的内存占用,从而实现了流操作。为了提高效率,我们应用了结构化通道修剪,在不影响音频质量的情况下将模型大小减少了8倍。我们的模型在Interspeech 2020深度噪声抑制挑战中表现出了很好的结果。具体来说,CleanUAmba仅用442 K参数和468 M MAC就实现了2.42的PESQ得分和95.1%的STOI,在实时性能方面与大型模型相匹配或优于大型模型。代码将在:https: github.com lab-emi CleanUMamba摘要:This paper presents CleanUMamba, a time-domain neural network architecture designed for real-time causal audio denoising directly applied to raw waveforms. CleanUMamba leverages a U-Net encoder-decoder structure, incorporating the Mamba state-space model in the bottleneck layer. By replacing conventional self-attention and LSTM mechanisms with Mamba, our architecture offers superior denoising performance while maintaining a constant memory footprint, enabling streaming operation. To enhance efficiency, we applied structured channel pruning, achieving an 8X reduction in model size without compromising audio quality. Our model demonstrates strong results in the Interspeech 2020 Deep Noise Suppression challenge. Specifically, CleanUMamba achieves a PESQ score of 2.42 and STOI of 95.1% with only 442K parameters and 468M MACs, matching or outperforming larger models in real-time performance. Code will be available at: https: github.com lab-emi CleanUMamba
【7】 GraFPrint: A GNN-Based Approach for Audio Identification
标题: GraFPrint:一种基于GNN的音频识别方法
作者: Aditya Bhattacharjee, Shubhr Singh, Emmanouil Benetos
备注:Submitted to IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2025)
链接:点击下载PDF文件
摘要:本文介绍了GraFPrint,这是一个音频识别框架,它利用图神经网络(GNN)的结构学习能力来创建鲁棒的音频指纹。我们的方法从时间-频率表示构造k-最近邻(k-NN)图,并应用最大相对图卷积来编码局部和全局信息。该网络使用自监督对比方法进行训练,该方法通过优化特征表示来增强对环境失真的适应能力。GraFPrint在各种粒度级别的大规模数据集上展示了卓越的性能,证明了它是轻量级和可扩展的,使其适合于具有广泛参考数据库的实际应用。摘要:This paper introduces GraFPrint, an audio identification framework that leverages the structural learning capabilities of Graph Neural Networks (GNNs) to create robust audio fingerprints. Our method constructs a k-nearest neighbor (k-NN) graph from time-frequency representations and applies max-relative graph convolutions to encode local and global information. The network is trained using a self-supervised contrastive approach, which enhances resilience to ambient distortions by optimizing feature representation. GraFPrint demonstrates superior performance on large-scale datasets at various levels of granularity, proving to be both lightweight and scalable, making it suitable for real-world applications with extensive reference databases.
【8】 Audio Captioning via Generative Pair-to-Pair Retrieval with Refined Knowledge Base
标题: 通过具有精细知识库的生成性成对检索的音频字幕
作者: Choi Changin, Lim Sungjun, Rhee Wonjong
链接:点击下载PDF文件
摘要:音频理解任务的最新进展利用了LLM的推理能力。然而,调整LLM以学习音频概念需要大量的训练数据和大量的计算资源。为了解决这些挑战,检索增强生成(RAG)检索音频文本对从知识库(KB)和增强他们与查询音频生成准确的文本响应。在RAG中,检索信息的相关性在有效处理输入中起着至关重要的作用。在本文中,我们分析了不同的检索方法和知识库如何影响音频文本对的相关性和性能的音频字幕与RAG。我们提出了生成式对对检索,它使用生成的字幕作为文本查询,以准确地找到相关的音频文本对查询音频,从而提高检索信息的相关性和准确性。此外,我们改进了大规模的知识库,只保留与上下文意图一致的音频文本对。我们的方法在包括AudioCaps,Clotho和Auto-ACD在内的基准测试中取得了最先进的结果,详细的消融研究验证了我们的检索和KB构建方法的有效性。摘要:Recent advances in audio understanding tasks leverage the reasoning capabilities of LLMs. However, adapting LLMs to learn audio concepts requires massive training data and substantial computational resources. To address these challenges, Retrieval-Augmented Generation (RAG) retrieves audio-text pairs from a knowledge base (KB) and augments them with query audio to generate accurate textual responses. In RAG, the relevance of the retrieved information plays a crucial role in effectively processing the input. In this paper, we analyze how different retrieval methods and knowledge bases impact the relevance of audio-text pairs and the performance of audio captioning with RAG. We propose generative pair-to-pair retrieval, which uses the generated caption as a text query to accurately find relevant audio-text pairs to the query audio, thereby improving the relevance and accuracy of retrieved information. Additionally, we refine the large-scale knowledge base to retain only audio-text pairs that align with the contextualized intents. Our approach achieves state-of-the-art results on benchmarks including AudioCaps, Clotho, and Auto-ACD, with detailed ablation studies validating the effectiveness of our retrieval and KB construction methods.
【9】 LLM Gesticulator: Leveraging Large Language Models for Scalable and Controllable Co-Speech Gesture Synthesis
标题: LLM Gesticulator:利用大型语言模型进行可扩展和可控的同声手势合成
作者: Haozhou Pang, Tianwei Ding, Lanshan He, Qi Gan
链接:点击下载PDF文件
摘要:在这项工作中,我们提出了LLM Gesticulator,一个基于LLM的音频驱动的协同语音手势生成框架,该框架合成全身动画,这些动画与输入音频有节奏地对齐,同时表现出自然的运动和可编辑性。与以前的工作相比,我们的模型表现出很大的可扩展性。随着主干LLM模型的大小增加,我们的框架显示出评估指标(也称为scaling law)。我们的方法也表现出很强的可控性,生成的手势的内容,风格可以控制的文本提示。据我们所知,LLM gesticulator是第一个使用LLM的合作语音生成任务的工作。现有的客观指标和用户研究的评估表明,我们的框架优于以前的作品。摘要:In this work, we present LLM Gesticulator, an LLM-based audio-driven co-speech gesture generation framework that synthesizes full-body animations that are rhythmically aligned with the input audio while exhibiting natural movements and editability. Compared to previous work, our model demonstrates substantial scalability. As the size of the backbone LLM model increases, our framework shows proportional improvements in evaluation metrics (a.k.a. scaling law). Our method also exhibits strong controllability where the content, style of the generated gestures can be controlled by text prompt. To the best of our knowledge, LLM gesticulator is the first work that use LLM on the co-speech generation task. Evaluation with existing objective metrics and user studies indicate that our framework outperforms prior works.
【10】 The importance of spatial and spectral information in multiple speaker tracking
标题: 空间和频谱信息在多说话人跟踪中的重要性
作者: Hanan Beit-On, Vladimir Tourbabin, Boaz Rafaely
链接:点击下载PDF文件
摘要:利用麦克风阵列录音进行多说话人定位和跟踪有着广泛的应用。多说话人跟踪的挑战之一是将方向估计与正确的说话人相关联。大多数现有的关联方法仅依赖于空间或频谱信息,导致性能下降时,这些信息通道之一是部分已知或丢失。本文研究了一种基于联合概率数据关联(JPDA)的方法,该方法基于联合空间-光谱信息进行关联。这是通过在关联概率计算中整合基于频谱信息估计的说话者时间-频率(TF)掩码来实现的。一项实验研究,测试所提出的方法从LOCATA的挑战记录证明了通过使用联合空间光谱信息的关联所获得的增强性能。摘要:Multi-speaker localization and tracking using microphone array recording is of importance in a wide range of applications. One of the challenges with multi-speaker tracking is to associate direction estimates with the correct speaker. Most existing association approaches rely on spatial or spectral information alone, leading to performance degradation when one of these information channels is partially known or missing. This paper studies a joint probability data association (JPDA)-based method that facilitates association based on joint spatial-spectral information. This is achieved by integrating speaker time-frequency (TF) masks, estimated based on spectral information, in the association probabilities calculation. An experimental study that tested the proposed method on recordings from the LOCATA challenge demonstrates the enhanced performance obtained by using joint spatial-spectral information in the association.
【11】 Mini-Omni2: Towards Open-source GPT-4o Model with Vision, Speech and Duplex
标题: Mini-Omni 2:迈向具有视觉、语音和双重功能的开源GPT-4 o模型
作者: Zhifei Xie, Changqiao Wu
备注:13 pages, 6 figures
链接:点击下载PDF文件
摘要:GPT 4 o是一款包罗万象的车型,代表了多模式大型车型发展的里程碑。它可以理解视觉,听觉和文本形式,直接输出音频,并支持灵活的双工交互。然而,它的技术框架并不是开源的。来自开源社区的模型通常实现GPT 4 o的一些功能,例如视觉理解和语音对话。然而,由于多模态数据、复杂的模型架构和训练过程的复杂性,训练一个包含所有模态的统一模型具有挑战性。在本文中,我们介绍了Mini-Omni 2,这是一种视听助手,能够为用户的视频和语音查询提供实时的端到端语音响应,同时还具有听觉功能。通过集成预先训练的视觉和听觉编码器,Mini-Omni 2在各个模式中保持强大的性能。我们提出了一个三阶段的训练过程来调整模态,允许语言模型在有限的数据集上训练后处理多模态输入和输出。对于交互,我们引入了一个基于语义的中断机制,使更灵活的对话与用户。所有建模方法和数据构建方法都将是开源的。据我们所知,Mini-Omni 2是功能上最接近GPT 4 o的模型之一,我们希望它能为后续研究提供有价值的见解。摘要:GPT4o, an all-encompassing model, represents a milestone in the development of multi-modal large models. It can understand visual, auditory, and textual modalities, directly output audio, and support flexible duplex interaction. However, its technical framework is not open-sourced. Models from the open-source community often achieve some functionalities of GPT4o, such as visual understanding and voice dialogue. Nevertheless, training a unified model that incorporates all modalities is challenging due to the complexities of multi-modal data, intricate model architectures, and training processes. In this paper, we introduce Mini-Omni2, a visual-audio assistant capable of providing real-time, end-to-end voice responses to user video and voice queries, while also incorporating auditory capabilities. By integrating pretrained visual and auditory encoders, Mini-Omni2 maintains strong performance in individual modalities. We propose a three-stage training process to align modalities, allowing the language model to handle multi-modal inputs and outputs after training on a limited dataset. For interaction, we introduce a semantic-based interruption mechanism, enabling more flexible dialogues with users. All modeling approaches and data construction methods will be open-sourced. To the best of our knowledge, Mini-Omni2 is one of the models closest to GPT4o in functionality, and we hope it can offer valuable insights for subsequent research.
【12】 DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection
标题: DARNet:具有时空结构的双重注意细化网络,用于听觉注意检测
作者: Sheng Yan, Cunhang fan, Hongyu Zhang, Xiaoke Yang, Jianhua Tao, Zhao Lv
链接:点击下载PDF文件
摘要:在鸡尾酒会上,人类表现出令人印象深刻的注意力转移能力。听觉注意检测(AAD)方法试图通过分析大脑信号(如EEG信号)来识别参与的说话者。然而,目前的AAD算法忽略了EEG信号中的空间分布信息,并且缺乏捕获长距离潜在依赖性的能力,从而限制了模型解码大脑活动的能力。针对这些问题,提出了一种基于时空结构的双注意力精化网络DARNet,该网络由时空结构模块、双注意力精化模块和特征融合与分类模块组成.具体而言,时空构建模块旨在通过捕获EEG信号的空间分布特征来构建更具表达力的时空特征表示。双注意力细化模块旨在提取EEG信号中不同级别的时间模式,并增强模型捕获长距离潜在依赖关系的能力。特征融合与分类器模块的目标是从不同的层次上聚合时间模式和依赖关系,并获得最终的分类结果。实验结果表明,在DTU数据集上,DARNet与现有模型相比,在0.1s、1s和2s时的平均分类准确率分别提高了5.9%、4.6%和3.9%。在保持出色分类性能的同时,DARNet显著减少了所需参数的数量。与最先进的模型相比,DARNet将参数数量减少了91%。代码可从以下网址获得:https: github.com fchest DARNet.git。摘要:At a cocktail party, humans exhibit an impressive ability to direct their attention. The auditory attention detection (AAD) approach seeks to identify the attended speaker by analyzing brain signals, such as EEG signals. However, current AAD algorithms overlook the spatial distribution information within EEG signals and lack the ability to capture long-range latent dependencies, limiting the model's ability to decode brain activity. To address these issues, this paper proposes a dual attention refinement network with spatiotemporal construction for AAD, named DARNet, which consists of the spatiotemporal construction module, dual attention refinement module, and feature fusion & classifier module. Specifically, the spatiotemporal construction module aims to construct more expressive spatiotemporal feature representations, by capturing the spatial distribution characteristics of EEG signals. The dual attention refinement module aims to extract different levels of temporal patterns in EEG signals and enhance the model's ability to capture long-range latent dependencies. The feature fusion & classifier module aims to aggregate temporal patterns and dependencies from different levels and obtain the final classification results. The experimental results indicate that compared to the state-of-the-art models, DARNet achieves an average classification accuracy improvement of 5.9 % for 0.1s, 4.6 % for 1s, and 3.9 % for 2s on the DTU dataset. While maintaining excellent classification performance, DARNet significantly reduces the number of required parameters. Compared to the state-of-the-art models, DARNet reduces the parameter count by 91 %. Code is available at: https: github.com fchest DARNet.git.
【13】 DMDSpeech: Distilled Diffusion Model Surpassing The Teacher in Zero-shot Speech Synthesis via Direct Metric Optimization
标题: DMZ Speech:通过直接度量优化在Zero-Shot语音合成中超越教师的蒸馏扩散模型
作者: Yingahao Aaron Li, Rithesh Kumar, Zeyu Jin
链接:点击下载PDF文件
摘要:扩散模型在语音合成任务中表现出巨大的潜力,包括文本到语音(TTS)和语音克隆。然而,它们的迭代去噪过程是低效的,并且阻碍了使用感知度量的端到端优化的应用。在本文中,我们提出了一种新的方法,提取TTS扩散模型与直接端到端的评估指标优化,实现国家的最先进的性能。通过将连接主义时间分类(CTC)损失和说话人验证(SV)损失,我们的方法优化了感知评估指标,导致显着改善的词错误率和说话人相似性。我们的实验表明,DMDSpeech始终超过现有的国家的最先进的模型在自然和说话人相似性,同时显着更快。此外,我们的合成语音具有更高的水平的语音相似性的提示比地面真理在人类的评价和客观的说话人相似性度量。这项工作突出了语音合成中直接度量优化的潜力,使模型能够更好地与人类听觉偏好保持一致。音频样本可从https: dmdspeech.github.io 获得。摘要:Diffusion models have demonstrated significant potential in speech synthesis tasks, including text-to-speech (TTS) and voice cloning. However, their iterative denoising processes are inefficient and hinder the application of end-to-end optimization with perceptual metrics. In this paper, we propose a novel method of distilling TTS diffusion models with direct end-to-end evaluation metric optimization, achieving state-of-the-art performance. By incorporating Connectionist Temporal Classification (CTC) loss and Speaker Verification (SV) loss, our approach optimizes perceptual evaluation metrics, leading to notable improvements in word error rate and speaker similarity. Our experiments show that DMDSpeech consistently surpasses prior state-of-the-art models in both naturalness and speaker similarity while being significantly faster. Moreover, our synthetic speech has a higher level of voice similarity to the prompt than the ground truth in both human evaluation and objective speaker similarity metric. This work highlights the potential of direct metric optimization in speech synthesis, allowing models to better align with human auditory preferences. The audio samples are available at https: dmdspeech.github.io .
【14】 Code Drift: Towards Idempotent Neural Audio Codecs
标题: 代码漂移:迈向等能神经音频编解码器
作者: Patrick O'Reilly, Prem Seetharaman, Jiaqi Su, Zeyu Jin, Bryan Pardo
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:神经编解码器在低比特率的音频信号的高保真压缩中表现出强大的性能。这些编解码器产生的基于令牌的表示已被证明对生成建模特别有用。虽然很多研究都集中在压缩比和感知透明度的改进,最近的作品在很大程度上忽略了另一个可取的编解码器属性-idemperature,压缩输出的稳定性下多轮编码。我们发现,最先进的神经编解码器表现出不同程度的相同性,一些退化的音频输出后,只有三个编码显着。我们调查可能的原因,低idemperature和设计一种方法,通过微调编解码器模型,以提高idemperature。然后,我们研究了等幂律对简单条件生成建模任务的影响,并发现可以在不对下游建模性能产生负面影响的情况下实现增加的等幂律-可能扩展神经编解码器在实际文件压缩和迭代生成建模工作流中的有用性。摘要:Neural codecs have demonstrated strong performance in high-fidelity compression of audio signals at low bitrates. The token-based representations produced by these codecs have proven particularly useful for generative modeling. While much research has focused on improvements in compression ratio and perceptual transparency, recent works have largely overlooked another desirable codec property -- idempotence, the stability of compressed outputs under multiple rounds of encoding. We find that state-of-the-art neural codecs exhibit varied degrees of idempotence, with some degrading audio outputs significantly after as few as three encodings. We investigate possible causes of low idempotence and devise a method for improving idempotence through fine-tuning a codec model. We then examine the effect of idempotence on a simple conditional generative modeling task, and find that increased idempotence can be achieved without negatively impacting downstream modeling performance -- potentially extending the usefulness of neural codecs for practical file compression and iterative generative modeling workflows.
eess.AS音频处理
【1】 The importance of spatial and spectral information in multiple speaker tracking标题: 空间和频谱信息在多说话人跟踪中的重要性
作者: Hanan Beit-On, Vladimir Tourbabin, Boaz Rafaely
链接:点击下载PDF文件
摘要:利用麦克风阵列录音进行多说话人定位和跟踪有着广泛的应用。多说话人跟踪的挑战之一是将方向估计与正确的说话人相关联。大多数现有的关联方法仅依赖于空间或频谱信息,当这些信息通道之一部分已知或缺失时,会导致性能下降。本文研究了一种基于联合概率数据关联(JPDA)的方法,该方法基于联合空间光谱信息进行关联。这是通过在关联概率计算中整合基于频谱信息估计的说话者时间-频率(TF)掩码来实现的。一项实验研究,测试所提出的方法从LOCATA的挑战记录证明了通过使用联合空间光谱信息的关联所获得的增强性能。摘要:Multi-speaker localization and tracking using microphone array recording is of importance in a wide range of applications. One of the challenges with multi-speaker tracking is to associate direction estimates with the correct speaker. Most existing association approaches rely on spatial or spectral information alone, leading to performance degradation when one of these information channels is partially known or missing. This paper studies a joint probability data association (JPDA)-based method that facilitates association based on joint spatial-spectral information. This is achieved by integrating speaker time-frequency (TF) masks, estimated based on spectral information, in the association probabilities calculation. An experimental study that tested the proposed method on recordings from the LOCATA challenge demonstrates the enhanced performance obtained by using joint spatial-spectral information in the association.
【2】 Mini-Omni2: Towards Open-source GPT-4o Model with Vision, Speech and Duplex
标题: Mini-Omni 2:迈向具有视觉、语音和双重功能的开源GPT-4 o模型
作者: Zhifei Xie, Changqiao Wu
备注:13 pages, 6 figures
链接:点击下载PDF文件
摘要:GPT 4 o是一款包罗万象的车型,代表了多模式大型车型开发的里程碑。它可以理解视觉,听觉和文本形式,直接输出音频,并支持灵活的双工交互。然而,它的技术框架并不是开源的。来自开源社区的模型通常实现GPT 4 o的一些功能,例如视觉理解和语音对话。然而,由于多模态数据、复杂的模型架构和训练过程的复杂性,训练一个包含所有模态的统一模型具有挑战性。在本文中,我们介绍了Mini-Omni 2,这是一种视听助手,能够为用户的视频和语音查询提供实时的端到端语音响应,同时还具有听觉功能。通过集成预先训练的视觉和听觉编码器,Mini-Omni 2在各个模式中保持强大的性能。我们提出了一个三阶段的训练过程来调整模态,允许语言模型在有限的数据集上训练后处理多模态输入和输出。对于交互,我们引入了一个基于语义的中断机制,使更灵活的对话与用户。所有建模方法和数据构建方法都将是开源的。据我们所知,Mini-Omni 2是功能上最接近GPT 4 o的模型之一,我们希望它能为后续研究提供有价值的见解。摘要:GPT4o, an all-encompassing model, represents a milestone in the development of multi-modal large models. It can understand visual, auditory, and textual modalities, directly output audio, and support flexible duplex interaction. However, its technical framework is not open-sourced. Models from the open-source community often achieve some functionalities of GPT4o, such as visual understanding and voice dialogue. Nevertheless, training a unified model that incorporates all modalities is challenging due to the complexities of multi-modal data, intricate model architectures, and training processes. In this paper, we introduce Mini-Omni2, a visual-audio assistant capable of providing real-time, end-to-end voice responses to user video and voice queries, while also incorporating auditory capabilities. By integrating pretrained visual and auditory encoders, Mini-Omni2 maintains strong performance in individual modalities. We propose a three-stage training process to align modalities, allowing the language model to handle multi-modal inputs and outputs after training on a limited dataset. For interaction, we introduce a semantic-based interruption mechanism, enabling more flexible dialogues with users. All modeling approaches and data construction methods will be open-sourced. To the best of our knowledge, Mini-Omni2 is one of the models closest to GPT4o in functionality, and we hope it can offer valuable insights for subsequent research.
【3】 DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection
标题: DARNet:具有时空结构的双重注意细化网络,用于听觉注意检测
作者: Sheng Yan, Cunhang fan, Hongyu Zhang, Xiaoke Yang, Jianhua Tao, Zhao Lv
链接:点击下载PDF文件
摘要:在鸡尾酒会上,人类表现出令人印象深刻的注意力转移能力。听觉注意检测(AAD)方法试图通过分析大脑信号(如EEG信号)来识别参与的说话者。然而,目前的AAD算法忽略了EEG信号中的空间分布信息,并且缺乏捕获长距离潜在依赖性的能力,从而限制了模型解码大脑活动的能力。针对这些问题,提出了一种基于时空结构的双注意力精化网络DARNet,该网络由时空结构模块、双注意力精化模块和特征融合与分类模块组成.具体而言,时空构造模块旨在通过捕获EEG信号的空间分布特征来构造更具表达力的时空特征表示。双注意力细化模块旨在提取EEG信号中不同级别的时间模式,并增强模型捕获长距离潜在依赖关系的能力。特征融合与分类器模块的目标是从不同的层次上聚合时间模式和依赖关系,并获得最终的分类结果。实验结果表明,在DTU数据集上,DARNet与现有模型相比,在0.1s、1s和2s时的平均分类准确率分别提高了5.9%、4.6%和3.9%。在保持出色分类性能的同时,DARNet显著减少了所需参数的数量。与最先进的模型相比,DARNet将参数数量减少了91%。代码可从以下网址获得:https: github.com fchest DARNet.git。摘要:At a cocktail party, humans exhibit an impressive ability to direct their attention. The auditory attention detection (AAD) approach seeks to identify the attended speaker by analyzing brain signals, such as EEG signals. However, current AAD algorithms overlook the spatial distribution information within EEG signals and lack the ability to capture long-range latent dependencies, limiting the model's ability to decode brain activity. To address these issues, this paper proposes a dual attention refinement network with spatiotemporal construction for AAD, named DARNet, which consists of the spatiotemporal construction module, dual attention refinement module, and feature fusion & classifier module. Specifically, the spatiotemporal construction module aims to construct more expressive spatiotemporal feature representations, by capturing the spatial distribution characteristics of EEG signals. The dual attention refinement module aims to extract different levels of temporal patterns in EEG signals and enhance the model's ability to capture long-range latent dependencies. The feature fusion & classifier module aims to aggregate temporal patterns and dependencies from different levels and obtain the final classification results. The experimental results indicate that compared to the state-of-the-art models, DARNet achieves an average classification accuracy improvement of 5.9 % for 0.1s, 4.6 % for 1s, and 3.9 % for 2s on the DTU dataset. While maintaining excellent classification performance, DARNet significantly reduces the number of required parameters. Compared to the state-of-the-art models, DARNet reduces the parameter count by 91 %. Code is available at: https: github.com fchest DARNet.git.
【4】 DMDSpeech: Distilled Diffusion Model Surpassing The Teacher in Zero-shot Speech Synthesis via Direct Metric Optimization
标题: DMZ Speech:通过直接度量优化在Zero-Shot语音合成中超越教师的蒸馏扩散模型
作者: Yingahao Aaron Li, Rithesh Kumar, Zeyu Jin
链接:点击下载PDF文件
摘要:扩散模型在语音合成任务中表现出巨大的潜力,包括文本到语音(TTS)和语音克隆。然而,它们的迭代去噪过程是低效的,并且阻碍了使用感知度量的端到端优化的应用。在本文中,我们提出了一种新的方法,提取TTS扩散模型与直接端到端的评估指标优化,实现国家的最先进的性能。通过将连接主义时间分类(CTC)损失和说话人验证(SV)损失,我们的方法优化了感知评估指标,导致显着改善的词错误率和说话人相似性。我们的实验表明,DMDSpeech始终超过现有的国家的最先进的模型在自然和说话人相似性,同时显着更快。此外,我们的合成语音具有更高的水平的语音相似性的提示比地面真相在人类评估和客观的说话人相似性度量。这项工作突出了语音合成中直接度量优化的潜力,使模型能够更好地与人类听觉偏好保持一致。音频样本可在https: dmdspeech.github.io 上获得。摘要:Diffusion models have demonstrated significant potential in speech synthesis tasks, including text-to-speech (TTS) and voice cloning. However, their iterative denoising processes are inefficient and hinder the application of end-to-end optimization with perceptual metrics. In this paper, we propose a novel method of distilling TTS diffusion models with direct end-to-end evaluation metric optimization, achieving state-of-the-art performance. By incorporating Connectionist Temporal Classification (CTC) loss and Speaker Verification (SV) loss, our approach optimizes perceptual evaluation metrics, leading to notable improvements in word error rate and speaker similarity. Our experiments show that DMDSpeech consistently surpasses prior state-of-the-art models in both naturalness and speaker similarity while being significantly faster. Moreover, our synthetic speech has a higher level of voice similarity to the prompt than the ground truth in both human evaluation and objective speaker similarity metric. This work highlights the potential of direct metric optimization in speech synthesis, allowing models to better align with human auditory preferences. The audio samples are available at https: dmdspeech.github.io .
【5】 Code Drift: Towards Idempotent Neural Audio Codecs
标题: 代码漂移:迈向等能神经音频编解码器
作者: Patrick O'Reilly, Prem Seetharaman, Jiaqi Su, Zeyu Jin, Bryan Pardo
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:神经编解码器在低比特率的音频信号的高保真压缩中表现出强大的性能。这些编解码器产生的基于令牌的表示已被证明对生成建模特别有用。虽然许多研究都集中在压缩比和感知透明度的改进,最近的作品在很大程度上忽略了另一个可取的编解码器属性-idemperature,压缩输出的稳定性下多轮编码。我们发现,最先进的神经编解码器表现出不同程度的相同性,一些退化的音频输出后,只有三个编码显着。我们调查可能的原因,低idemperature和设计一种方法,通过微调编解码器模型,以提高idemperature。然后,我们研究了等幂律对简单条件生成建模任务的影响,并发现可以在不对下游建模性能产生负面影响的情况下实现增加的等幂律-可能扩展神经编解码器在实际文件压缩和迭代生成建模工作流中的有用性。摘要:Neural codecs have demonstrated strong performance in high-fidelity compression of audio signals at low bitrates. The token-based representations produced by these codecs have proven particularly useful for generative modeling. While much research has focused on improvements in compression ratio and perceptual transparency, recent works have largely overlooked another desirable codec property -- idempotence, the stability of compressed outputs under multiple rounds of encoding. We find that state-of-the-art neural codecs exhibit varied degrees of idempotence, with some degrading audio outputs significantly after as few as three encodings. We investigate possible causes of low idempotence and devise a method for improving idempotence through fine-tuning a codec model. We then examine the effect of idempotence on a simple conditional generative modeling task, and find that increased idempotence can be achieved without negatively impacting downstream modeling performance -- potentially extending the usefulness of neural codecs for practical file compression and iterative generative modeling workflows.
【6】 Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction
标题: 利用LLM嵌入进行跨数据集标签对齐和零镜头音乐情感预测
作者: Renhang Liu, Abhinaba Roy, Dorien Herremans
链接:点击下载PDF文件
摘要:在这项工作中,我们提出了一种新的方法,音乐情感识别,利用大语言模型(LLM)嵌入标签对齐跨多个数据集和zero-shot预测新的类别。首先,我们计算情感标签的LLM嵌入,并在包含不相交标签的多个数据集上应用非参数聚类对相似标签进行分组。我们使用这些聚类中心将音乐特征(MERT)映射到LLM嵌入空间。为了进一步增强模型,我们引入了一个对齐正则化,使MERT嵌入从不同的集群分离。这进一步增强了模型更好地适应未知数据集的能力。我们通过对一个新的数据集进行zero-shot推理来证明我们的方法的有效性,展示了它在没有额外训练的情况下推广到看不见的标签的能力。摘要:In this work, we present a novel method for music emotion recognition that leverages Large Language Model (LLM) embeddings for label alignment across multiple datasets and zero-shot prediction on novel categories. First, we compute LLM embeddings for emotion labels and apply non-parametric clustering to group similar labels, across multiple datasets containing disjoint labels. We use these cluster centers to map music features (MERT) to the LLM embedding space. To further enhance the model, we introduce an alignment regularization that enables dissociation of MERT embeddings from different clusters. This further enhances the model's ability to better adaptation to unseen datasets. We demonstrate the effectiveness of our approach by performing zero-shot inference on a new dataset, showcasing its ability to generalize to unseen labels without additional training.
【7】 Diff-SAGe: End-to-End Spatial Audio Generation Using Diffusion Models
标题: Dist-SAGee:使用扩散模型的端到端空间音频生成
作者: Saksham Singh Kushwaha, Jianbo Ma, Mark R. P. Thomas, Yapeng Tian, Avery Bruni
链接:点击下载PDF文件
摘要:空间音频是创造沉浸式体验的关键组成部分。传统的基于模拟的方法来生成空间音频依赖于专业知识,具有有限的可扩展性,并假设语义和空间信息之间的独立性。为了解决这些问题,我们探索端到端的空间音频生成。我们介绍并制定了一个新的任务,产生一阶高保真度立体声(FOA)给定的声音类别和声源的空间位置。我们提出了Diff-SAGe,一个端到端的,基于流的扩散变压器模型,这项任务。Diff-SAGe利用复杂的频谱图表示FOA,保留相位信息的准确空间线索至关重要。此外,多条件编码器将输入条件集成到统一的表示中,指导从噪声中生成FOA波形。通过对两个数据集的广泛评估,我们证明了我们的方法在客观和主观指标上始终优于传统的基于模拟的基线。摘要:Spatial audio is a crucial component in creating immersive experiences. Traditional simulation-based approaches to generate spatial audio rely on expertise, have limited scalability, and assume independence between semantic and spatial information. To address these issues, we explore end-to-end spatial audio generation. We introduce and formulate a new task of generating first-order Ambisonics (FOA) given a sound category and sound source spatial location. We propose Diff-SAGe, an end-to-end, flow-based diffusion-transformer model for this task. Diff-SAGe utilizes a complex spectrogram representation for FOA, preserving the phase information crucial for accurate spatial cues. Additionally, a multi-conditional encoder integrates the input conditions into a unified representation, guiding the generation of FOA waveforms from noise. Through extensive evaluations on two datasets, we demonstrate that our method consistently outperforms traditional simulation-based baselines across both objective and subjective metrics.
【8】 Investigation of Speaker Representation for Target-Speaker Speech Processing
标题: 目标说话人语音处理中说话人表示的研究
作者: Takanori Ashihara, Takafumi Moriya, Shota Horiguchi, Junyi Peng, Tsubasa Ochiai, Marc Delcroix, Kohei Matsuura, Hiroshi Sato
备注:Accepted at IEEE SLT 2024
链接:点击下载PDF文件
摘要:目标说话人语音处理(TS)任务,诸如目标说话人自动语音识别(TS-ASR)、目标语音提取(TSE)和个人语音活动检测(p-VAD),对于提取关于期望说话人的语音的信息是重要的,即使当其被干扰说话人破坏时。虽然大多数研究都集中在每个特定任务的训练方案或系统架构,但用于嵌入目标说话人线索的辅助网络尚未在统一的跨任务评估中进行全面研究。因此,本文旨在解决一个基本问题:什么是首选的说话人嵌入TS任务?为此,对于TS-ASR、TSE和p-VAD任务,我们比较预先训练的说话者编码器(即,自监督或说话者识别模型),其从目标说话者的预先记录的登记语音计算说话者嵌入,其中理想说话者嵌入以独热向量的形式直接从目标说话者的身份导出。为了进一步了解理想说话人嵌入的特性,我们使用基于梯度的方法对其进行优化,以提高TS任务的性能。我们的分析表明,说话人验证性能是有点无关TS任务的性能,一个热向量优于注册为基础的,和最佳的嵌入依赖于输入的混合。摘要:Target-speaker speech processing (TS) tasks, such as target-speaker automatic speech recognition (TS-ASR), target speech extraction (TSE), and personal voice activity detection (p-VAD), are important for extracting information about a desired speaker's speech even when it is corrupted by interfering speakers. While most studies have focused on training schemes or system architectures for each specific task, the auxiliary network for embedding target-speaker cues has not been investigated comprehensively in a unified cross-task evaluation. Therefore, this paper aims to address a fundamental question: what is the preferred speaker embedding for TS tasks? To this end, for the TS-ASR, TSE, and p-VAD tasks, we compare pre-trained speaker encoders (i.e., self-supervised or speaker recognition models) that compute speaker embeddings from pre-recorded enrollment speech of the target speaker with ideal speaker embeddings derived directly from the target speaker's identity in the form of a one-hot vector. To further understand the properties of ideal speaker embedding, we optimize it using a gradient-based approach to improve performance on the TS task. Our analysis reveals that speaker verification performance is somewhat unrelated to TS task performances, the one-hot vector outperforms enrollment-based ones, and the optimal embedding depends on the input mixture.
【9】 Audio-based Kinship Verification Using Age Domain Conversion
标题: 使用年龄域转换的音频亲属关系验证
作者: Qiyang Sun, Alican Akman, Xin Jing, Manuel Milling, Björn W. Schuller
备注:4 pages, 2 figures, submitted to IEEE Signal Processing Letters
链接:点击下载PDF文件
摘要:基于音频的亲属关系验证(AKV)在许多领域都很重要,例如家庭安全监控,法医鉴定和社交网络分析。任务中的一个关键挑战来自于来自不同个体的样本之间的年龄差异,这可以被解释为跨域验证任务中的域偏差。为了解决这个问题,我们设计了“年龄标准化域”的概念,其中我们利用优化的CycleGAN-VC 3网络来执行年龄音频转换以生成域内音频。生成的音频数据集用于提取一系列特征,然后将其输入到度量学习架构中以验证亲属关系。实验是在KAN_AV音频数据集上进行的,该数据集包含年龄和亲属标签。实验结果表明,该方法显著提高了亲属关系验证的准确性,同时也为未来的亲属关系验证研究提供了新的见解。摘要:Audio-based kinship verification (AKV) is important in many domains, such as home security monitoring, forensic identification, and social network analysis. A key challenge in the task arises from differences in age across samples from different individuals, which can be interpreted as a domain bias in a cross-domain verification task. To address this issue, we design the notion of an "age-standardised domain" wherein we utilise the optimised CycleGAN-VC3 network to perform age-audio conversion to generate the in-domain audio. The generated audio dataset is employed to extract a range of features, which are then fed into a metric learning architecture to verify kinship. Experiments are conducted on the KAN_AV audio dataset, which contains age and kinship labels. The results demonstrate that the method markedly enhances the accuracy of kinship verification, while also offering novel insights for future kinship verification research.
【10】 Character-aware audio-visual subtitling in context
标题: 背景下的用户感知视听字幕
作者: Jaesung Huh, Andrew Zisserman
备注:ACCV 2024
链接:点击下载PDF文件
摘要:本文提出了一种改进的框架,在电视节目中的字符感知视听字幕。我们的方法集成了语音识别,扬声器diarisation,字符识别,利用音频和视觉线索。这个整体解决方案解决了说什么、什么时候说以及谁在说话的问题,为电视节目提供了更全面、更准确的角色感知字幕。我们的方法在两个方面带来了改进:首先,我们表明,视听同步可以用来挑选出在视频剪辑中存在的其他人之间的说话的脸,并分配相应的语音段的身份。这种视听方法提高了识别精度和产量超过目前的方法。其次,我们表明,短片段的扬声器可以通过使用场景内的对话的时间上下文来确定。我们提出了一种方法,使用本地语音嵌入的音频,和大语言模型推理的文本转录。这克服了现有方法的局限性,即它们不能准确地将扬声器分配给短时间段。我们验证了12个电视节目的数据集上的方法,表现出优越的性能相比,现有的方法在扬声器diarisation和字符识别的准确性。项目页面:https: www.robots.ox.ac.uk ~vgg research llr-context 摘要:This paper presents an improved framework for character-aware audio-visual subtitling in TV shows. Our approach integrates speech recognition, speaker diarisation, and character recognition, utilising both audio and visual cues. This holistic solution addresses what is said, when it's said, and who is speaking, providing a more comprehensive and accurate character-aware subtitling for TV shows. Our approach brings improvements on two fronts: first, we show that audio-visual synchronisation can be used to pick out the talking face amongst others present in a video clip, and assign an identity to the corresponding speech segment. This audio-visual approach improves recognition accuracy and yield over current methods. Second, we show that the speaker of short segments can be determined by using the temporal context of the dialogue within a scene. We propose an approach using local voice embeddings of the audio, and large language model reasoning on the text transcription. This overcomes a limitation of existing methods that they are unable to accurately assign speakers to short temporal segments. We validate the method on a dataset with 12 TV shows, demonstrating superior performance in speaker diarisation and character recognition accuracy compared to existing approaches. Project page : https: www.robots.ox.ac.uk ~vgg research llr-context
【11】 CleanUMamba: A Compact Mamba Network for Speech Denoising using Channel Pruning
标题: CleanUmamba:一个紧凑的Mamba网络,用于使用通道修剪进行语音降噪
作者: Sjoerd Groot, Qinyu Chen, Jan C. van Gemert, Chang Gao
备注:5 pages, 4 figures
链接:点击下载PDF文件
摘要:本文介绍了CleanUAmba,一种时域神经网络架构,用于直接应用于原始波形的实时因果音频去噪。CleanUAmba利用了U-Net编码器-解码器结构,在瓶颈层中结合了Mamba状态空间模型。通过用Mamba替换传统的自注意力和LSTM机制,我们的架构提供了卓越的去噪性能,同时保持了恒定的内存占用,从而实现了流操作。为了提高效率,我们应用了结构化通道修剪,在不影响音频质量的情况下将模型大小减少了8倍。我们的模型在Interspeech 2020深度噪声抑制挑战中表现出了很好的结果。具体来说,CleanUAmba仅用442 K参数和468 M MAC就实现了2.42的PESQ得分和95.1%的STOI,在实时性能方面与大型模型相匹配或优于大型模型。代码将在:https: github.com lab-emi CleanUMamba摘要:This paper presents CleanUMamba, a time-domain neural network architecture designed for real-time causal audio denoising directly applied to raw waveforms. CleanUMamba leverages a U-Net encoder-decoder structure, incorporating the Mamba state-space model in the bottleneck layer. By replacing conventional self-attention and LSTM mechanisms with Mamba, our architecture offers superior denoising performance while maintaining a constant memory footprint, enabling streaming operation. To enhance efficiency, we applied structured channel pruning, achieving an 8X reduction in model size without compromising audio quality. Our model demonstrates strong results in the Interspeech 2020 Deep Noise Suppression challenge. Specifically, CleanUMamba achieves a PESQ score of 2.42 and STOI of 95.1% with only 442K parameters and 468M MACs, matching or outperforming larger models in real-time performance. Code will be available at: https: github.com lab-emi CleanUMamba
【12】 GraFPrint: A GNN-Based Approach for Audio Identification
标题: GraFPrint:一种基于GNN的音频识别方法
作者: Aditya Bhattacharjee, Shubhr Singh, Emmanouil Benetos
备注:Submitted to IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2025)
链接:点击下载PDF文件
摘要:本文介绍了GraFPrint,这是一个音频识别框架,它利用图神经网络(GNN)的结构学习能力来创建鲁棒的音频指纹。我们的方法从时间-频率表示构造k-最近邻(k-NN)图,并应用最大相对图卷积来编码局部和全局信息。该网络使用自监督对比方法进行训练,该方法通过优化特征表示来增强对环境失真的适应能力。GraFPrint在各种粒度级别的大规模数据集上展示了卓越的性能,证明了它是轻量级和可扩展的,使其适合于具有广泛参考数据库的实际应用。摘要:This paper introduces GraFPrint, an audio identification framework that leverages the structural learning capabilities of Graph Neural Networks (GNNs) to create robust audio fingerprints. Our method constructs a k-nearest neighbor (k-NN) graph from time-frequency representations and applies max-relative graph convolutions to encode local and global information. The network is trained using a self-supervised contrastive approach, which enhances resilience to ambient distortions by optimizing feature representation. GraFPrint demonstrates superior performance on large-scale datasets at various levels of granularity, proving to be both lightweight and scalable, making it suitable for real-world applications with extensive reference databases.
【13】 Audio Captioning via Generative Pair-to-Pair Retrieval with Refined Knowledge Base
标题: 通过具有精细知识库的生成性成对检索的音频字幕
作者: Choi Changin, Lim Sungjun, Rhee Wonjong
链接:点击下载PDF文件
摘要:音频理解任务的最新进展利用了LLM的推理能力。然而,调整LLM以学习音频概念需要大量的训练数据和大量的计算资源。为了解决这些挑战,检索增强生成(RAG)检索音频文本对从知识库(KB)和增强他们与查询音频生成准确的文本响应。在RAG中,检索信息的相关性在有效处理输入中起着至关重要的作用。在本文中,我们分析了不同的检索方法和知识库如何影响音频文本对的相关性和性能的音频字幕与RAG。我们提出了生成式对对检索,它使用生成的字幕作为文本查询,以准确地找到相关的音频文本对查询音频,从而提高检索信息的相关性和准确性。此外,我们改进了大规模的知识库,只保留与上下文意图一致的音频文本对。我们的方法在包括AudioCaps,Clotho和Auto-ACD在内的基准测试中取得了最先进的结果,详细的消融研究验证了我们的检索和KB构建方法的有效性。摘要:Recent advances in audio understanding tasks leverage the reasoning capabilities of LLMs. However, adapting LLMs to learn audio concepts requires massive training data and substantial computational resources. To address these challenges, Retrieval-Augmented Generation (RAG) retrieves audio-text pairs from a knowledge base (KB) and augments them with query audio to generate accurate textual responses. In RAG, the relevance of the retrieved information plays a crucial role in effectively processing the input. In this paper, we analyze how different retrieval methods and knowledge bases impact the relevance of audio-text pairs and the performance of audio captioning with RAG. We propose generative pair-to-pair retrieval, which uses the generated caption as a text query to accurately find relevant audio-text pairs to the query audio, thereby improving the relevance and accuracy of retrieved information. Additionally, we refine the large-scale knowledge base to retain only audio-text pairs that align with the contextualized intents. Our approach achieves state-of-the-art results on benchmarks including AudioCaps, Clotho, and Auto-ACD, with detailed ablation studies validating the effectiveness of our retrieval and KB construction methods.
【14】 LLM Gesticulator: Leveraging Large Language Models for Scalable and Controllable Co-Speech Gesture Synthesis
标题: LLM Gesticulator:利用大型语言模型进行可扩展和可控的同声手势合成
作者: Haozhou Pang, Tianwei Ding, Lanshan He, Qi Gan
链接:点击下载PDF文件
摘要:在这项工作中,我们提出了LLM Gesticulator,一个基于LLM的音频驱动的协同语音手势生成框架,该框架合成全身动画,这些动画与输入音频有节奏地对齐,同时表现出自然的运动和可编辑性。与以前的工作相比,我们的模型表现出很大的可扩展性。随着主干LLM模型的大小增加,我们的框架显示出评估指标(也称为scaling law)。我们的方法也表现出很强的可控性,生成的手势的内容,风格可以控制的文本提示。据我们所知,LLM gesticulator是第一个使用LLM的合作语音生成任务的工作。现有的客观指标和用户研究的评估表明,我们的框架优于以前的作品。摘要:In this work, we present LLM Gesticulator, an LLM-based audio-driven co-speech gesture generation framework that synthesizes full-body animations that are rhythmically aligned with the input audio while exhibiting natural movements and editability. Compared to previous work, our model demonstrates substantial scalability. As the size of the backbone LLM model increases, our framework shows proportional improvements in evaluation metrics (a.k.a. scaling law). Our method also exhibits strong controllability where the content, style of the generated gestures can be controlled by text prompt. To the best of our knowledge, LLM gesticulator is the first work that use LLM on the co-speech generation task. Evaluation with existing objective metrics and user studies indicate that our framework outperforms prior works.
机器翻译,仅供参考
