微信公众号:arXiv_Daily
cs.SD语音
标题: 基于Mamba的网络使用置信度二进制正规化进行半监督歌唱旋律提取
链接:https://arxiv.org/abs/2505.08681
摘要:歌唱旋律提取是音乐信息检索领域的一个关键问题。然而,现有的方法面临着一些限制:首先,以前的模型使用Transformers来捕获上下文依赖,这需要二次计算,导致推理阶段效率低下。其次,现有的作品通常依赖于频率监督方法来估计基频(f0),这忽略了音乐表演实际上是基于音符的。第三,Transformers通常需要大量的标记数据来实现最佳性能,但SME任务缺乏足够的注释数据。为了解决这些问题,在本文中,我们提出了一个基于曼巴的网络,称为SpectMamba,使用置信度二进制正则化的半监督歌唱旋律提取。特别是,我们首先引入视觉曼巴实现计算线性复杂度。然后,我们提出了一种新的音符f0解码器,使模型能够更好地模仿音乐的表现。此外,为了缓解标记数据的稀缺性,我们引入了置信二进制正则化(CBR)模块,通过最大化正确类的概率来利用未标记数据。在几个公共数据集上对所提出的方法进行了评估,并进行了实验,证明了我们所提出的方法的有效性。
摘要:Singing melody extraction (SME) is a key task in the field of music information retrieval. However, existing methods are facing several limitations: firstly, prior models use transformers to capture the contextual dependencies, which requires quadratic computation resulting in low efficiency in the inference stage. Secondly, prior works typically rely on frequencysupervised methods to estimate the fundamental frequency (f0), which ignores that the musical performance is actually based on notes. Thirdly, transformers typically require large amounts of labeled data to achieve optimal performances, but the SME task lacks of sufficient annotated data. To address these issues, in this paper, we propose a mamba-based network, called SpectMamba, for semi-supervised singing melody extraction using confidence binary regularization. In particular, we begin by introducing vision mamba to achieve computational linear complexity. Then, we propose a novel note-f0 decoder that allows the model to better mimic the musical performance. Further, to alleviate the scarcity of the labeled data, we introduce a confidence binary regularization (CBR) module to leverage the unlabeled data by maximizing the probability of the correct classes. The proposed method is evaluated on several public datasets and the conducted experiments demonstrate the effectiveness of our proposed method.
【2】 M3G: Multi-Granular Gesture Generator for Audio-Driven Full-Body Human Motion Synthesis
标题: M3 G:用于音频驱动全身人体运动合成的多颗粒手势生成器链接:https://arxiv.org/abs/2505.08293
备注:9 Pages, 4 figures, submitted to NIPS 2025
摘要:从音频生成包括面部、身体、手和全局运动的全身人类姿态是虚拟化身创建中有价值但具有挑战性的任务。以前的系统专注于逐帧标记人类手势并从输入音频预测每帧的标记。然而,一个观察结果是,定义为粒度的完整表达性人类姿势所需的帧的数量在不同的人类姿势模式之间变化。现有的系统无法建模这些手势模式,由于其手势令牌的固定粒度。为了解决这个问题,我们提出了一个新的框架命名为多粒度手势生成器(M3 G)的音频驱动的整体手势生成。在M3 G中,我们提出了一种新的多粒度VQ-VAE(MGVQ-VAE)来标记运动模式,并从不同的时间粒度重建运动序列。随后,我们提出了一个多粒度令牌预测器,从音频中提取多粒度信息,并预测相应的运动令牌。然后,M3 G使用MGVQ-VAE从预测的令牌中重建人类手势。客观和主观的实验表明,我们提出的M3 G框架优于国家的最先进的方法,在产生自然和富有表现力的全身人体姿态。
摘要:Generating full-body human gestures encompassing face, body, hands, and global movements from audio is a valuable yet challenging task in virtual avatar creation. Previous systems focused on tokenizing the human gestures framewisely and predicting the tokens of each frame from the input audio. However, one observation is that the number of frames required for a complete expressive human gesture, defined as granularity, varies among different human gesture patterns. Existing systems fail to model these gesture patterns due to the fixed granularity of their gesture tokens. To solve this problem, we propose a novel framework named Multi-Granular Gesture Generator (M3G) for audio-driven holistic gesture generation. In M3G, we propose a novel Multi-Granular VQ-VAE (MGVQ-VAE) to tokenize motion patterns and reconstruct motion sequences from different temporal granularities. Subsequently, we proposed a multi-granular token predictor that extracts multi-granular information from audio and predicts the corresponding motion tokens. Then M3G reconstructs the human gestures from the predicted tokens using the MGVQ-VAE. Both objective and subjective experiments demonstrate that our proposed M3G framework outperforms the state-of-the-art methods in terms of generating natural and expressive full-body human gestures.
【3】 Unveiling the Best Practices for Applying Speech Foundation Models to Speech Intelligibility Prediction for Hearing-Impaired People
标题: 揭示将语音基础模型应用于听力障碍者语音可理解度预测的最佳实践链接:https://arxiv.org/abs/2505.08215
摘要:语音基础模型(SFM)在各种下游任务中表现出强大的性能,包括听力受损人群的语音清晰度预测(SIP-HI)。然而,对于SIP-HI的SFM的优化还没有被充分探索。在本文中,我们进行了全面的研究,以确定影响SIP-HI性能的关键设计因素与5 SFM,侧重于编码器层的选择,预测头架构,和合奏配置。我们的研究结果表明,与传统的使用所有层的方法相反,选择单个编码器层会产生更好的结果。此外,时间建模对于有效的预测头至关重要。我们还证明了集成多个SFM可以提高性能,更强大的单个模型可以提供更大的好处。最后,我们探讨了关键的SFM属性和它们对SIP-HI性能的影响之间的关系。我们的研究提供了实际的见解,有效地适应SFM的语音清晰度预测听力受损人群。
摘要:Speech foundation models (SFMs) have demonstrated strong performance across a variety of downstream tasks, including speech intelligibility prediction for hearing-impaired people (SIP-HI). However, optimizing SFMs for SIP-HI has been insufficiently explored. In this paper, we conduct a comprehensive study to identify key design factors affecting SIP-HI performance with 5 SFMs, focusing on encoder layer selection, prediction head architecture, and ensemble configurations. Our findings show that, contrary to traditional use-all-layers methods, selecting a single encoder layer yields better results. Additionally, temporal modeling is crucial for effective prediction heads. We also demonstrate that ensembling multiple SFMs improves performance, with stronger individual models providing greater benefit. Finally, we explore the relationship between key SFM attributes and their impact on SIP-HI performance. Our study offers practical insights into effectively adapting SFMs for speech intelligibility prediction for hearing-impaired populations.
【4】 Not that Groove: Zero-Shot Symbolic Music Editing
标题: 不是那个凹槽:Zero-Shot象征性音乐编辑链接:https://arxiv.org/abs/2505.08203
摘要:人工智能音乐生成的大多数工作都集中在音频上,由于其刚性,音频在音乐制作行业的使用有限。为了最大限度地提高灵活性,同时只假设制作人的文本指令,我们是第一批处理符号音乐编辑的人。我们规避了已知的挑战,缺乏标记的数据证明,LLM与zero-shot提示可以有效地编辑鼓槽。成功的秘诀是一个创造性地设计的格式,接口LLM和音乐,而我们通过提供一个评估数据集与注释的单元测试,高度符合音乐家的判断促进评估。
摘要:Most work in AI music generation focused on audio, which has seen limited use in the music production industry due to its rigidity. To maximize flexibility while assuming only textual instructions from producers, we are among the first to tackle symbolic music editing. We circumvent the known challenge of lack of labeled data by proving that LLMs with zero-shot prompting can effectively edit drum grooves. The recipe of success is a creatively designed format that interfaces LLMs and music, while we facilitate evaluation by providing an evaluation dataset with annotated unit tests that highly aligns with musicians' judgment.
【5】 Fast Text-to-Audio Generation with Adversarial Post-Training
标题: 具有对抗性后训练的快速文本到音频生成链接:https://arxiv.org/abs/2505.08175
摘要:文本到音频系统虽然性能越来越好,但在推理时速度很慢,因此对于许多创造性应用程序来说,它们的延迟是不切实际的。我们提出了对抗性相对论-对比(ARC)后训练,这是第一个不基于蒸馏的扩散/流模型的对抗性加速算法。虽然过去的对抗性后训练方法很难与昂贵的蒸馏方法进行比较,但ARC后训练是一个简单的过程,它(1)将最近的相对论对抗性公式扩展到扩散/流动后训练,(2)将其与一种新的对比性后训练目标相结合,以鼓励更好地及时遵守。我们将ARC后训练与Stable Audio Open的一些优化相结合,并构建了一个能够在H100上以约75毫秒生成约12秒44.1kHz立体声音频的模型,并在移动边缘设备上生成约7秒,这是我们所知的最快的文本到音频模型。
摘要:Text-to-audio systems, while increasingly performant, are slow at inference time, thus making their latency unpractical for many creative applications. We present Adversarial Relativistic-Contrastive (ARC) post-training, the first adversarial acceleration algorithm for diffusion/flow models not based on distillation. While past adversarial post-training methods have struggled to compare against their expensive distillation counterparts, ARC post-training is a simple procedure that (1) extends a recent relativistic adversarial formulation to diffusion/flow post-training and (2) combines it with a novel contrastive discriminator objective to encourage better prompt adherence. We pair ARC post-training with a number optimizations to Stable Audio Open and build a model capable of generating $\approx$12s of 44.1kHz stereo audio in $\approx$75ms on an H100, and $\approx$7s on a mobile edge-device, the fastest text-to-audio model to our knowledge.
【6】 MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
标题: MiniMax-Speech:具有可学习发言人编码器的固有Zero-Shot文本到语音链接:https://arxiv.org/abs/2505.07916
摘要:我们介绍了MiniMax-Speech,一种基于自回归变换器的文本到语音(TTS)模型,可以生成高质量的语音。一个关键的创新是我们的可学习扬声器编码器,它从参考音频中提取音色特征,而不需要转录。这使得MiniMax-Speech能够以zero-shot方式产生具有与参考一致的音色的高度表达的语音,同时还支持与参考语音具有极高相似性的单次语音克隆。此外,通过所提出的Flow-VAE,合成音频的整体质量得到增强。我们的模型支持32种语言,并在多个客观和主观评估指标上表现出出色的性能。值得注意的是,它在客观的语音克隆指标(单词错误率和说话者相似度)上达到了最先进的(SOTA)结果,并在公共TTS Arena排行榜上获得了最高位置。MiniMax-Speech的另一个关键优势是其可扩展性,而无需修改基础模型,例如:通过LoRA进行任意语音情感控制;通过直接从文本描述合成音色特征的文本到语音(T2 V);以及通过使用额外数据微调音色特征的专业语音克隆(PVC)。我们鼓励读者访问https://minimax-ai.github.io/tts_tech_report以获取更多示例。
摘要:We introduce MiniMax-Speech, an autoregressive Transformer-based Text-to-Speech (TTS) model that generates high-quality speech. A key innovation is our learnable speaker encoder, which extracts timbre features from a reference audio without requiring its transcription. This enables MiniMax-Speech to produce highly expressive speech with timbre consistent with the reference in a zero-shot manner, while also supporting one-shot voice cloning with exceptionally high similarity to the reference voice. In addition, the overall quality of the synthesized audio is enhanced through the proposed Flow-VAE. Our model supports 32 languages and demonstrates excellent performance across multiple objective and subjective evaluations metrics. Notably, it achieves state-of-the-art (SOTA) results on objective voice cloning metrics (Word Error Rate and Speaker Similarity) and has secured the top position on the public TTS Arena leaderboard. Another key strength of MiniMax-Speech, granted by the robust and disentangled representations from the speaker encoder, is its extensibility without modifying the base model, enabling various applications such as: arbitrary voice emotion control via LoRA; text to voice (T2V) by synthesizing timbre features directly from text description; and professional voice cloning (PVC) by fine-tuning timbre features with additional data. We encourage readers to visit https://minimax-ai.github.io/tts_tech_report for more examples.
【1】 Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
标题: Granite-speech:开源语音感知LLM,具有强大的英语ASB功能链接:https://arxiv.org/abs/2505.08699
备注:7 pages, 9 figures
摘要:Granite-speech LLM是专为英语ASR和自动语音翻译(AST)设计的紧凑高效的语音语言模型。这些模型是通过将granite-3.3-instruct的2B和8B参数变体与公开可用的开源语料库上的语音进行模态对齐来训练的,这些语料库包含音频输入和文本目标,这些文本目标由ASR的人类转录或自动生成的AST翻译组成。全面的基准测试表明,在英语ASR上,这是我们的主要关注点,它们的表现优于几个竞争对手的模型,这些模型是在数量级更多的专有数据上训练的,并且它们在欧洲主要语言,日语和中文的英语到X AST上保持同步。语音特定的组件是:一个使用块注意力和自我调节训练的构象声学编码器,使用连接主义时间分类,一个窗口查询转换器语音模态适配器,用于对声学嵌入进行时间下采样并将其映射到LLM文本嵌入空间,以及LoRA适配器,用于进一步微调文本LLM。Granite-speech-3.3在两种模式下运行:在语音模式下,它通过激活编码器,投影仪和LoRA适配器来执行ASR和AST;在文本模式下,它直接调用底层的granite-3.3-instruct模型(没有LoRA),基本上保留了所有的文本LLM功能和安全性。这两个模型都可以在HuggingFace(https://www.example.com和https://huggingface.co/ibm-granite/granite-speech-3.3-8b)上免费获得,并且可以在许可的Apache 2.0许可下用于研究和商业目的。huggingface.co/ibm-granite/granite-speech-3.3-2b
摘要:Granite-speech LLMs are compact and efficient speech language models specifically designed for English ASR and automatic speech translation (AST). The models were trained by modality aligning the 2B and 8B parameter variants of granite-3.3-instruct to speech on publicly available open-source corpora containing audio inputs and text targets consisting of either human transcripts for ASR or automatically generated translations for AST. Comprehensive benchmarking shows that on English ASR, which was our primary focus, they outperform several competitors' models that were trained on orders of magnitude more proprietary data, and they keep pace on English-to-X AST for major European languages, Japanese, and Chinese. The speech-specific components are: a conformer acoustic encoder using block attention and self-conditioning trained with connectionist temporal classification, a windowed query-transformer speech modality adapter used to do temporal downsampling of the acoustic embeddings and map them to the LLM text embedding space, and LoRA adapters to further fine-tune the text LLM. Granite-speech-3.3 operates in two modes: in speech mode, it performs ASR and AST by activating the encoder, projector, and LoRA adapters; in text mode, it calls the underlying granite-3.3-instruct model directly (without LoRA), essentially preserving all the text LLM capabilities and safety. Both models are freely available on HuggingFace (https://huggingface.co/ibm-granite/granite-speech-3.3-2b and https://huggingface.co/ibm-granite/granite-speech-3.3-8b) and can be used for both research and commercial purposes under a permissive Apache 2.0 license.
【2】 A Survey of Deep Learning for Complex Speech Spectrograms
标题: 复杂语音谱图深度学习综述链接:https://arxiv.org/abs/2505.08694
摘要:深度学习的最新进展对语音信号处理领域产生了重大影响,特别是在复杂频谱图的分析和操作方面。本调查全面概述了利用深度神经网络处理复杂频谱图的最新技术,这些频谱图封装了幅度和相位信息。我们首先介绍复杂的声谱图及其相关功能,用于各种语音处理任务。接下来,我们将探讨复值神经网络的关键组件和架构,这些组件和架构专门用于处理复值数据,并已应用于复频谱图处理。然后,我们讨论了各种训练策略和损失函数,用于训练神经网络处理和建模复杂的频谱图。该调查进一步研究了关键应用,包括相位检索,语音增强和语音分离,其中深度学习通过利用复杂的频谱图或其衍生的特征表示取得了重大进展。此外,我们还研究了复杂频谱图与生成模型的交叉点。本调查旨在为语音信号处理和复值神经网络领域的研究人员和从业人员提供宝贵的资源。
摘要:Recent advancements in deep learning have significantly impacted the field of speech signal processing, particularly in the analysis and manipulation of complex spectrograms. This survey provides a comprehensive overview of the state-of-the-art techniques leveraging deep neural networks for processing complex spectrograms, which encapsulate both magnitude and phase information. We begin by introducing complex spectrograms and their associated features for various speech processing tasks. Next, we explore the key components and architectures of complex-valued neural networks, which are specifically designed to handle complex-valued data and have been applied for complex spectrogram processing. We then discuss various training strategies and loss functions tailored for training neural networks to process and model complex spectrograms. The survey further examines key applications, including phase retrieval, speech enhancement, and speech separation, where deep learning has achieved significant progress by leveraging complex spectrograms or their derived feature representations. Additionally, we examine the intersection of complex spectrograms with generative models. This survey aims to serve as a valuable resource for researchers and practitioners in the field of speech signal processing and complex-valued neural networks.
【3】 Investigating self-supervised features for expressive, multilingual voice conversion
标题: 研究自我监督功能以实现富有表达力的多语言语音转换链接:https://arxiv.org/abs/2505.08278
备注:Published as a conference paper at ICASSP 2024
摘要:语音转换(VC)系统广泛用于从说话者匿名到个性化语音合成的多种应用。监督方法使用并行数据来学习不同说话者之间的映射,这是昂贵的。无监督方法通常被训练来重建由内容和说话者信息组成的输入信号。解开这些组件是一个挑战,往往会导致扬声器泄漏或韵律信息删除。在本文中,我们通过利用自监督学习(SSL)的潜力来探索语音转换。SSL模型的潜在表示的组合,与扬声器嵌入连接,被馈送到声码器,该声码器被训练以重建输入。Zero-shot语音转换结果表明,该方法允许保持源说话人的韵律和内容,同时匹配基于语音后验图(PPG)的VC系统的说话人相似性。
摘要:Voice conversion (VC) systems are widely used for several applications, from speaker anonymisation to personalised speech synthesis. Supervised approaches learn a mapping between different speakers using parallel data, which is expensive to produce. Unsupervised approaches are typically trained to reconstruct the input signal, which is composed of the content and the speaker information. Disentangling these components is a challenge and often leads to speaker leakage or prosodic information removal. In this paper, we explore voice conversion by leveraging the potential of self-supervised learning (SSL). A combination of the latent representations of SSL models, concatenated with speaker embeddings, is fed to a vocoder which is trained to reconstruct the input. Zero-shot voice conversion results show that this approach allows to keep the prosody and content of the source speaker while matching the speaker similarity of a VC system based on phonetic posteriorgrams (PPGs).
【4】 MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
标题: MiniMax-Speech:具有可学习发言人编码器的固有Zero-Shot文本到语音链接:https://arxiv.org/abs/2505.07916
摘要:我们介绍了MiniMax-Speech,一种基于自回归变换器的文本到语音(TTS)模型,可以生成高质量的语音。一个关键的创新是我们的可学习扬声器编码器,它从参考音频中提取音色特征,而不需要转录。这使得MiniMax-Speech能够以zero-shot方式产生具有与参考一致的音色的高度表达的语音,同时还支持与参考语音具有极高相似性的单次语音克隆。此外,通过所提出的Flow-VAE,合成音频的整体质量得到增强。我们的模型支持32种语言,并在多个客观和主观评估指标上表现出出色的性能。值得注意的是,它在客观的语音克隆指标(单词错误率和说话者相似度)上达到了最先进的(SOTA)结果,并在公共TTS Arena排行榜上获得了最高位置。MiniMax-Speech的另一个关键优势是其可扩展性,而无需修改基础模型,例如:通过LoRA进行任意语音情感控制;通过直接从文本描述合成音色特征的文本到语音(T2 V);以及通过使用额外数据微调音色特征的专业语音克隆(PVC)。我们鼓励读者访问https://minimax-ai.github.io/tts_tech_report以获取更多示例。
摘要:We introduce MiniMax-Speech, an autoregressive Transformer-based Text-to-Speech (TTS) model that generates high-quality speech. A key innovation is our learnable speaker encoder, which extracts timbre features from a reference audio without requiring its transcription. This enables MiniMax-Speech to produce highly expressive speech with timbre consistent with the reference in a zero-shot manner, while also supporting one-shot voice cloning with exceptionally high similarity to the reference voice. In addition, the overall quality of the synthesized audio is enhanced through the proposed Flow-VAE. Our model supports 32 languages and demonstrates excellent performance across multiple objective and subjective evaluations metrics. Notably, it achieves state-of-the-art (SOTA) results on objective voice cloning metrics (Word Error Rate and Speaker Similarity) and has secured the top position on the public TTS Arena leaderboard. Another key strength of MiniMax-Speech, granted by the robust and disentangled representations from the speaker encoder, is its extensibility without modifying the base model, enabling various applications such as: arbitrary voice emotion control via LoRA; text to voice (T2V) by synthesizing timbre features directly from text description; and professional voice cloning (PVC) by fine-tuning timbre features with additional data. We encourage readers to visit https://minimax-ai.github.io/tts_tech_report for more examples.
【5】 Three Tone Networks and a Tessellation
标题: 三音网络和镶嵌链接:https://arxiv.org/abs/2505.08752
备注:29 pages, 12 figures
摘要:我们表明,欧拉tonnetz,其中关联三个小调和弦,每个主要和弦和三个主要和弦,每个小调和弦,可以表示为一个二分图,12个白色顶点表示主要和弦和12个黑色顶点表示次要和弦。这个所谓的列维图唯一地确定了实射影平面上的12个点和12条线的某个显著配置的组合几何,其性质是每条线上有3个点,每条线上有3条线穿过每个点。音调的有趣特征,例如科恩的四个六声周期的存在,对于理解19世纪的声音主导和扩展和声至关重要,可以直接解读为配置的属性。我们表明,类似的音调网络可以构建五声音乐和十二音音乐。
摘要:We show that the Eulerian tonnetz, which associates three minor chords to each major chord and three major chords to each minor chord, can be represented by a bipartite graph with twelve white vertices signifying major chords and twelve black vertices signifying minor chords. This so-called Levi graph uniquely determines the combinatorial geometry of a certain remarkable configuration of twelve points and twelve lines in the real projective plane with the property that three points lie on each line and three lines pass through each point. Interesting features of the tonnetz, such as the existence of Cohn's four hexatonic cycles, crucial for the understanding of nineteenth-century voice leading and extended harmony, can be read off rather directly as properties of the configuration. We show that analogous tone networks can be constructed for pentatonic music and twelve-tone music.
【6】 A Mamba-based Network for Semi-supervised Singing Melody Extraction Using Confidence Binary Regularization
标题: 基于Mamba的网络使用置信度二进制正规化进行半监督歌唱旋律提取链接:https://arxiv.org/abs/2505.08681
摘要:歌唱旋律提取是音乐信息检索领域的一个关键问题。然而,现有的方法面临着一些限制:首先,以前的模型使用Transformers来捕获上下文依赖,这需要二次计算,导致推理阶段效率低下。其次,现有的作品通常依赖于频率监督方法来估计基频(f0),这忽略了音乐表演实际上是基于音符的。第三,Transformers通常需要大量的标记数据来实现最佳性能,但SME任务缺乏足够的注释数据。为了解决这些问题,在本文中,我们提出了一个基于曼巴的网络,称为SpectMamba,使用置信度二进制正则化的半监督歌唱旋律提取。特别是,我们首先引入视觉曼巴实现计算线性复杂度。然后,我们提出了一种新的音符f0解码器,使模型能够更好地模仿音乐的表现。此外,为了缓解标记数据的稀缺性,我们引入了置信二进制正则化(CBR)模块,通过最大化正确类的概率来利用未标记数据。在几个公共数据集上对所提出的方法进行了评估,并进行了实验,证明了我们所提出的方法的有效性。
摘要:Singing melody extraction (SME) is a key task in the field of music information retrieval. However, existing methods are facing several limitations: firstly, prior models use transformers to capture the contextual dependencies, which requires quadratic computation resulting in low efficiency in the inference stage. Secondly, prior works typically rely on frequencysupervised methods to estimate the fundamental frequency (f0), which ignores that the musical performance is actually based on notes. Thirdly, transformers typically require large amounts of labeled data to achieve optimal performances, but the SME task lacks of sufficient annotated data. To address these issues, in this paper, we propose a mamba-based network, called SpectMamba, for semi-supervised singing melody extraction using confidence binary regularization. In particular, we begin by introducing vision mamba to achieve computational linear complexity. Then, we propose a novel note-f0 decoder that allows the model to better mimic the musical performance. Further, to alleviate the scarcity of the labeled data, we introduce a confidence binary regularization (CBR) module to leverage the unlabeled data by maximizing the probability of the correct classes. The proposed method is evaluated on several public datasets and the conducted experiments demonstrate the effectiveness of our proposed method.
【7】 M3G: Multi-Granular Gesture Generator for Audio-Driven Full-Body Human Motion Synthesis
标题: M3 G:用于音频驱动全身人体运动合成的多颗粒手势生成器链接:https://arxiv.org/abs/2505.08293
备注:9 Pages, 4 figures, submitted to NIPS 2025
摘要:从音频生成包括面部、身体、手和全局运动的全身人类姿态是虚拟化身创建中有价值但具有挑战性的任务。以前的系统专注于逐帧标记人类手势并从输入音频预测每帧的标记。然而,一个观察结果是,定义为粒度的完整表达性人类姿势所需的帧的数量在不同的人类姿势模式之间变化。现有的系统无法建模这些手势模式,由于其手势令牌的固定粒度。为了解决这个问题,我们提出了一个新的框架命名为多粒度手势生成器(M3 G)的音频驱动的整体手势生成。在M3 G中,我们提出了一种新的多粒度VQ-VAE(MGVQ-VAE)来标记运动模式,并从不同的时间粒度重建运动序列。随后,我们提出了一个多粒度令牌预测器,从音频中提取多粒度信息,并预测相应的运动令牌。然后,M3 G使用MGVQ-VAE从预测的令牌中重建人类手势。客观和主观的实验表明,我们提出的M3 G框架优于国家的最先进的方法,在产生自然和富有表现力的全身人体姿态。
摘要:Generating full-body human gestures encompassing face, body, hands, and global movements from audio is a valuable yet challenging task in virtual avatar creation. Previous systems focused on tokenizing the human gestures framewisely and predicting the tokens of each frame from the input audio. However, one observation is that the number of frames required for a complete expressive human gesture, defined as granularity, varies among different human gesture patterns. Existing systems fail to model these gesture patterns due to the fixed granularity of their gesture tokens. To solve this problem, we propose a novel framework named Multi-Granular Gesture Generator (M3G) for audio-driven holistic gesture generation. In M3G, we propose a novel Multi-Granular VQ-VAE (MGVQ-VAE) to tokenize motion patterns and reconstruct motion sequences from different temporal granularities. Subsequently, we proposed a multi-granular token predictor that extracts multi-granular information from audio and predicts the corresponding motion tokens. Then M3G reconstructs the human gestures from the predicted tokens using the MGVQ-VAE. Both objective and subjective experiments demonstrate that our proposed M3G framework outperforms the state-of-the-art methods in terms of generating natural and expressive full-body human gestures.
【8】 Unveiling the Best Practices for Applying Speech Foundation Models to Speech Intelligibility Prediction for Hearing-Impaired People
标题: 揭示将语音基础模型应用于听力障碍者语音可理解度预测的最佳实践链接:https://arxiv.org/abs/2505.08215
摘要:语音基础模型(SFM)在各种下游任务中表现出强大的性能,包括听力受损人群的语音清晰度预测(SIP-HI)。然而,对于SIP-HI的SFM的优化还没有被充分探索。在本文中,我们进行了全面的研究,以确定影响SIP-HI性能的关键设计因素与5 SFM,侧重于编码器层的选择,预测头架构,和合奏配置。我们的研究结果表明,与传统的使用所有层的方法相反,选择单个编码器层会产生更好的结果。此外,时间建模对于有效的预测头至关重要。我们还证明了集成多个SFM可以提高性能,更强大的单个模型可以提供更大的好处。最后,我们探讨了关键的SFM属性和它们对SIP-HI性能的影响之间的关系。我们的研究提供了实际的见解,有效地适应SFM的语音清晰度预测听力受损人群。
摘要:Speech foundation models (SFMs) have demonstrated strong performance across a variety of downstream tasks, including speech intelligibility prediction for hearing-impaired people (SIP-HI). However, optimizing SFMs for SIP-HI has been insufficiently explored. In this paper, we conduct a comprehensive study to identify key design factors affecting SIP-HI performance with 5 SFMs, focusing on encoder layer selection, prediction head architecture, and ensemble configurations. Our findings show that, contrary to traditional use-all-layers methods, selecting a single encoder layer yields better results. Additionally, temporal modeling is crucial for effective prediction heads. We also demonstrate that ensembling multiple SFMs improves performance, with stronger individual models providing greater benefit. Finally, we explore the relationship between key SFM attributes and their impact on SIP-HI performance. Our study offers practical insights into effectively adapting SFMs for speech intelligibility prediction for hearing-impaired populations.
【9】 Not that Groove: Zero-Shot Symbolic Music Editing
标题: 不是那个凹槽:Zero-Shot象征性音乐编辑链接:https://arxiv.org/abs/2505.08203
摘要:人工智能音乐生成的大多数工作都集中在音频上,由于其刚性,音频在音乐制作行业的使用有限。为了最大限度地提高灵活性,同时只假设制作人的文本指令,我们是第一批处理符号音乐编辑的人。我们规避了已知的挑战,缺乏标记的数据证明,LLM与zero-shot提示可以有效地编辑鼓槽。成功的秘诀是一个创造性地设计的格式,接口LLM和音乐,而我们通过提供一个评估数据集与注释的单元测试,高度符合音乐家的判断促进评估。
摘要:Most work in AI music generation focused on audio, which has seen limited use in the music production industry due to its rigidity. To maximize flexibility while assuming only textual instructions from producers, we are among the first to tackle symbolic music editing. We circumvent the known challenge of lack of labeled data by proving that LLMs with zero-shot prompting can effectively edit drum grooves. The recipe of success is a creatively designed format that interfaces LLMs and music, while we facilitate evaluation by providing an evaluation dataset with annotated unit tests that highly aligns with musicians' judgment.
【10】 Fast Text-to-Audio Generation with Adversarial Post-Training
标题: 具有对抗性后训练的快速文本到音频生成链接:https://arxiv.org/abs/2505.08175
摘要:文本到音频系统虽然性能越来越好,但在推理时速度很慢,因此对于许多创造性应用程序来说,它们的延迟是不切实际的。我们提出了对抗性相对论-对比(ARC)后训练,这是第一个不基于蒸馏的扩散/流模型的对抗性加速算法。虽然过去的对抗性后训练方法很难与昂贵的蒸馏方法进行比较,但ARC后训练是一个简单的过程,它(1)将最近的相对论对抗性公式扩展到扩散/流动后训练,(2)将其与一种新的对比性后训练目标相结合,以鼓励更好地及时遵守。我们将ARC后训练与Stable Audio Open的一些优化相结合,并构建了一个能够在H100上以约75毫秒生成约12秒44.1kHz立体声音频的模型,并在移动边缘设备上生成约7秒,这是我们所知的最快的文本到音频模型。
摘要:Text-to-audio systems, while increasingly performant, are slow at inference time, thus making their latency unpractical for many creative applications. We present Adversarial Relativistic-Contrastive (ARC) post-training, the first adversarial acceleration algorithm for diffusion/flow models not based on distillation. While past adversarial post-training methods have struggled to compare against their expensive distillation counterparts, ARC post-training is a simple procedure that (1) extends a recent relativistic adversarial formulation to diffusion/flow post-training and (2) combines it with a novel contrastive discriminator objective to encourage better prompt adherence. We pair ARC post-training with a number optimizations to Stable Audio Open and build a model capable of generating $\approx$12s of 44.1kHz stereo audio in $\approx$75ms on an H100, and $\approx$7s on a mobile edge-device, the fastest text-to-audio model to our knowledge.
机器翻译由腾讯交互翻译提供,仅供参考
