微信公众号:arXiv_Daily
cs.SD语音
【1】Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment
标题:Art 2 Mus:通过视觉条件反射和大规模跨模式对齐的艺术作品到音乐的生成
链接:https://arxiv.org/abs/2602.17599
摘要:通过多模态深度学习,音乐生成取得了显着进展,使模型能够从文本合成音频,最近还可以从图像合成音频。然而,现有的图像调节系统有两个基本的局限性:(i)它们通常是在自然照片上训练的,限制了它们捕捉艺术品更丰富的语义、风格和文化内容的能力;(ii)大多数依赖于图像到文本的转换阶段,使用语言作为语义捷径,简化了调节,但阻止了直接的视觉到音频的学习。出于这些差距的动机,我们引入了ArtSound,这是一个大规模的多模态数据集,包含105,884个艺术作品-音乐对,通过扩展ArtGraph和Free Music Archive获得,并添加了双模态字幕。我们进一步提出了ArtToMus,这是第一个明确设计用于直接艺术作品到音乐生成的框架,它将数字化艺术作品映射到音乐,而无需图像到文本的翻译或基于语言的语义监督。该框架将视觉嵌入投射到潜在扩散模型的调节空间中,使音乐合成仅由视觉信息指导。实验结果表明,ArtToMus生成的音乐连贯性和风格一致的输出,反映了显着的视觉线索的源艺术品。虽然绝对对齐分数仍然低于文本条件系统的预期考虑到大幅增加的难度,删除语言监督ArtToMus实现竞争力的感知质量和有意义的跨模态对应。这项工作确立了直接视觉到音乐生成作为一个独特的和具有挑战性的研究方向,并提供资源,支持多媒体艺术,文化遗产和人工智能辅助的创作实践的应用。代码和数据集将在接受后公开发布。
摘要:Music generation has advanced markedly through multimodal deep learning, enabling models to synthesize audio from text and, more recently, from images. However, existing image-conditioned systems suffer from two fundamental limitations: (i) they are typically trained on natural photographs, limiting their ability to capture the richer semantic, stylistic, and cultural content of artworks; and (ii) most rely on an image-to-text conversion stage, using language as a semantic shortcut that simplifies conditioning but prevents direct visual-to-audio learning. Motivated by these gaps, we introduce ArtSound, a large-scale multimodal dataset of 105,884 artwork-music pairs enriched with dual-modality captions, obtained by extending ArtGraph and the Free Music Archive. We further propose ArtToMus, the first framework explicitly designed for direct artwork-to-music generation, which maps digitized artworks to music without image-to-text translation or language-based semantic supervision. The framework projects visual embeddings into the conditioning space of a latent diffusion model, enabling music synthesis guided solely by visual information. Experimental results show that ArtToMus generates musically coherent and stylistically consistent outputs that reflect salient visual cues of the source artworks. While absolute alignment scores remain lower than those of text-conditioned systems-as expected given the substantially increased difficulty of removing linguistic supervision-ArtToMus achieves competitive perceptual quality and meaningful cross-modal correspondence. This work establishes direct visual-to-music generation as a distinct and challenging research direction, and provides resources that support applications in multimedia art, cultural heritage, and AI-assisted creative practice. Code and dataset will be publicly released upon acceptance.
【2】Voice-Driven Semantic Perception for UAV-Assisted Emergency Networks
标题:无人机辅助应急网络的语音驱动语义感知
链接:https://arxiv.org/abs/2602.17394
备注:7 pages, 4 figures
摘要:无人机(UAV)辅助网络越来越被视为一种有前途的应急响应方法,在地面基础设施退化或不可用的环境中提供快速,灵活和弹性的通信。在这种情况下,语音无线电通信由于其鲁棒性对于第一响应者仍然是必不可少的;然而,其非结构化的性质阻止了与自动化无人机辅助网络管理的直接集成。本文提出了SIREN,这是一个AI驱动的框架,可以为无人机辅助网络提供语音驱动的感知。通过将自动语音识别(ASR)与基于大型语言模型(LLM)的语义提取和自然语言处理(NLP)验证相结合,SIREN将紧急语音流量转换为结构化的机器可读信息,包括响应单元,位置参考,紧急情况严重程度和服务质量(QoS)要求。SIREN评估使用合成的紧急情况下,在语言,扬声器计数,背景噪音和信息复杂性的控制变化。结果表明,强大的转录和可靠的语义提取在不同的操作条件下,同时突出扬声器日记和地理模糊的主要限制因素。这些研究结果建立了无人机辅助网络的语音驱动态势感知的可行性,并显示了应急响应操作中的人在回路决策支持和自适应网络管理的实际基础。
摘要:Unmanned Aerial Vehicle (UAV)-assisted networks are increasingly foreseen as a promising approach for emergency response, providing rapid, flexible, and resilient communications in environments where terrestrial infrastructure is degraded or unavailable. In such scenarios, voice radio communications remain essential for first responders due to their robustness; however, their unstructured nature prevents direct integration with automated UAV-assisted network management. This paper proposes SIREN, an AI-driven framework that enables voice-driven perception for UAV-assisted networks. By integrating Automatic Speech Recognition (ASR) with Large Language Model (LLM)-based semantic extraction and Natural Language Processing (NLP) validation, SIREN converts emergency voice traffic into structured, machine-readable information, including responding units, location references, emergency severity, and Quality-of-Service (QoS) requirements. SIREN is evaluated using synthetic emergency scenarios with controlled variations in language, speaker count, background noise, and message complexity. The results demonstrate robust transcription and reliable semantic extraction across diverse operating conditions, while highlighting speaker diarization and geographic ambiguity as the main limiting factors. These findings establish the feasibility of voice-driven situational awareness for UAV-assisted networks and show a practical foundation for human-in-the-loop decision support and adaptive network management in emergency response operations.
【3】AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing
标题:AudioChat:通过Transfusion Forcing统一音频讲故事、编辑和理解
链接:https://arxiv.org/abs/2602.17097
摘要:尽管最近取得了突破,但音频基础模型在处理复杂的多源声学场景方面仍存在困难。我们将这个具有挑战性的领域称为音频故事,它可以有多个扬声器和背景/前景音效。与传统的音频处理任务相比,音频故事引入了语义、时间和物理复杂性的新层。为了应对这一挑战,我们提出了AudioChat,一个框架,用于开发音频基础模型,可以生成,编辑和理解音频故事。AudioChat引入了一种新的范例,其中基于LLM的工具调用代理模拟用户和系统之间的交互,这些模拟对话被用作训练数据。我们还引入了一种新的音频传输强制目标来训练AudioChat模型,使其能够通过结构化的思维链推理同时分解高级指令,并执行交互式多回合音频理解/生成。为了评估生成和编辑性能,我们开发了三个新的指标,直接衡量任务性能,而不是依赖于基于分布的评分。我们强烈建议读者访问我们的演示,以更好地了解AudioChat的功能:https://wanchichen.github.io/audiochat/。
摘要:Despite recent breakthroughs, audio foundation models struggle in processing complex multi-source acoustic scenes. We refer to this challenging domain as audio stories, which can have multiple speakers and background/foreground sound effects. Compared to traditional audio processing tasks, audio stories introduce new layers of semantic, temporal, and physical complexity. To address this challenge, we propose AudioChat, a framework for developing audio foundation models that can generate, edit, and understand audio stories. AudioChat introduces a new paradigm in which LLM-based toolcalling agents simulate interactions between users and the system, and these simulated dialogues are used as training data. We also introduce a novel Audio Transfusion Forcing objective to train the AudioChat model, allowing it to simultaneously decompose high-level instructions via structured chain-of-thought reasoning and perform interactive multi-turn audio understanding/generation. To evaluate generation and editing performance, we develop three new metrics that directly measure task performance instead of relying upon distribution-based scoring. We highly encourage readers to visit our demo to better understand the capabilities of AudioChat: https://wanchichen.github.io/audiochat/.
【4】Generative Audio Extension and Morphing
标题:生成音频扩展和变形
链接:https://arxiv.org/abs/2602.16790
备注:Accepted to ICASSP 2026
摘要:在与音频相关的创作任务中,声音设计师经常寻求从他们的库中扩展和变形不同的声音。生成音频模型能够使用示例作为参考来创建音频,提供了有前途的解决方案。通过掩蔽DiT的噪声潜伏期并在这种掩蔽的潜伏期上应用无分类器指导的新变体,我们证明:(i)给定音频参考,我们可以在指定的持续时间内向前和向后扩展它,以及(ii)给定两个音频参考,我们可以在所需的持续时间内无缝地变形它们。此外,我们表明,通过对不同类型的固定音频数据进行微调,我们可以减轻潜在的幻觉。我们的方法的有效性得到了客观指标的支持,生成的音频达到了与训练数据中的真实样本相当的Fréchet音频距离(FAD)。此外,我们验证我们的结果,通过主观的听众测试,受试者给予积极的评价,提出的模型代。这种技术为更可控和更具表现力的生成声音框架铺平了道路,使声音设计师能够减少对繁琐重复任务的关注,更多地关注他们的实际创作过程。
摘要:In audio-related creative tasks, sound designers often seek to extend and morph different sounds from their libraries. Generative audio models, capable of creating audio using examples as references, offer promising solutions. By masking the noisy latents of a DiT and applying a novel variant of classifier-free guidance on such masked latents, we demonstrate that: (i) given an audio reference, we can extend it both forward and backward for a specified duration, and (ii) given two audio references, we can morph them seamlessly for the desired duration. Furthermore, we show that by fine-tuning the model on different types of stationary audio data we mitigate potential hallucinations. The effectiveness of our method is supported by objective metrics, with the generated audio achieving Fréchet Audio Distances (FADs) comparable to those of real samples from the training data. Additionally, we validate our results through a subjective listener test, where subjects gave positive ratings to the proposed model generations. This technique paves the way for more controllable and expressive generative sound frameworks, enabling sound designers to focus less on tedious, repetitive tasks and more on their actual creative process.
【5】Speech to Speech Synthesis for Voice Impersonation
标题:语音模拟的语音到语音合成
链接:https://arxiv.org/abs/2602.16721
备注:Original work completed in April 2020. This version includes minor formatting updates
摘要:许多模型在语音识别和语音合成领域取得了巨大的成功,但语音到语音处理的模型还没有得到大量的研究。我们提出了语音到语音合成网络(STSSN),一个模型的基础上,融合了这两个学科,以执行有效的语音到语音风格转移的目的,语音模拟的最新系统。我们表明,我们提出的模型是相当强大的,并成功地生成逼真的音频样本,尽管在其容量的一些缺点。我们通过将我们提出的模型与完成类似任务的生成对抗模型进行比较来对我们提出的模型进行基准测试,并表明我们的模型产生了更令人信服的结果。
摘要:Numerous models have shown great success in the fields of speech recognition as well as speech synthesis, but models for speech to speech processing have not been heavily explored. We propose Speech to Speech Synthesis Network (STSSN), a model based on current state of the art systems that fuses the two disciplines in order to perform effective speech to speech style transfer for the purpose of voice impersonation. We show that our proposed model is quite powerful, and succeeds in generating realistic audio samples despite a number of drawbacks in its capacity. We benchmark our proposed model by comparing it with a generative adversarial model which accomplishes a similar task, and show that ours produces more convincing results.
【1】CC-G2PnP: Streaming Grapheme-to-Phoneme and prosody with Conformer-CTC for unsegmented languages
标题:CC-G2 PSYS:使用Conformer-ctc为未分段语言流传输图形到音素和韵律
链接:https://arxiv.org/abs/2602.17157
备注:Accepted by ICASSP 2026
摘要:我们提出了CC-G2 Pstroke,一个流的字形到音素和韵律(G2 Pstroke)模型连接大的语言模型和文本到语音在流的方式。CC-G2 P2P基于Conformer-CTC架构。具体而言,输入的字素标记被逐块处理,这使得能够对音素和韵律(Pestic)标签进行流式推理。通过保证每个输入令牌的最小前瞻大小,所提出的模型可以考虑每个令牌中的未来上下文,这导致稳定的Pestrian标签预测。与之前依赖于显式单词边界的流方法不同,CC-G2 Pencil中的CTC解码器在训练过程中有效地学习了字素和音素之间的对齐,使其适用于未分段的语言。在一个没有明确词边界的日语数据集上的实验表明,CC-G2 P2P模型在P2P标签预测的准确性上明显优于基线流G2 P2P模型。
摘要:We propose CC-G2PnP, a streaming grapheme-to-phoneme and prosody (G2PnP) model to connect large language model and text-to-speech in a streaming manner. CC-G2PnP is based on Conformer-CTC architecture. Specifically, the input grapheme tokens are processed chunk by chunk, which enables streaming inference of phonemic and prosodic (PnP) labels. By guaranteeing minimal look-ahead size to each input token, the proposed model can consider future context in each token, which leads to stable PnP label prediction. Unlike previous streaming methods that depend on explicit word boundaries, the CTC decoder in CC-G2PnP effectively learns the alignment between graphemes and phonemes during training, making it applicable to unsegmented languages. Experiments on a Japanese dataset, which has no explicit word boundaries, show that CC-G2PnP significantly outperforms the baseline streaming G2PnP model in the accuracy of PnP label prediction.
【2】The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR$\rightarrow$LLM Pipelines?
标题:级联等效假设:言语LLM何时表现得像ASB $ightarrow$LLM管道?
链接:https://arxiv.org/abs/2602.17598
备注:10 pages, 6 figures, 7 tables
摘要:当前的语音LLM主要执行隐式ASR:在从成绩单可解决的任务上,它们在行为上和机械上等同于简单的Whisper$\to$LLM级联。我们通过四个语音LLM和六个任务的匹配骨干测试来展示这一点,首次控制LLM骨干。Ultravox在统计上与其匹配的级联($κ{=}0.93$)无法区分; logit透镜揭示了隐藏状态中出现的文字; LEACE概念擦除确认了文本表示在两种测试架构中都是因果必要的,将准确性降低到接近零。Qwen 2-Audio确实存在差异,这表明级联等效依赖于架构,而不是通用的。对于大多数部署的用例,当前的语音LLM是昂贵的级联,并且在噪声下,它们是更差的级联,在0 dB时,清洁条件优势逆转高达7.6%。
摘要:Current speech LLMs largely perform implicit ASR: on tasks solvable from a transcript, they are behaviorally and mechanistically equivalent to simple Whisper$\to$LLM cascades. We show this through matched-backbone testing across four speech LLMs and six tasks, controlling for the LLM backbone for the first time. Ultravox is statistically indistinguishable from its matched cascade ($κ{=}0.93$); logit lens reveals literal text emerging in hidden states; LEACE concept erasure confirms text representations are causally necessary in both architectures tested, collapsing accuracy to near-zero. Qwen2-Audio genuinely diverges, revealing cascade equivalence is architecture-dependent, not universal. For most deployed use cases, current speech LLMs are expensive cascades, and under noise, they are worse ones, with clean-condition advantages reversing by up to 7.6% at 0 dB.
【3】Speech to Speech Synthesis for Voice Impersonation
标题:语音模拟的语音到语音合成
链接:https://arxiv.org/abs/2602.16721
备注:Original work completed in April 2020. This version includes minor formatting updates
摘要:许多模型在语音识别和语音合成领域取得了巨大的成功,但语音到语音处理的模型还没有得到大量的研究。我们提出了语音到语音合成网络(STSSN),一个模型的基础上,融合了这两个学科,以执行有效的语音到语音风格转移的目的,语音模拟的最新系统。我们表明,我们提出的模型是相当强大的,并成功地生成逼真的音频样本,尽管在其容量的一些缺点。我们通过将我们提出的模型与完成类似任务的生成对抗模型进行比较来对我们提出的模型进行基准测试,并表明我们的模型产生了更令人信服的结果。
摘要:Numerous models have shown great success in the fields of speech recognition as well as speech synthesis, but models for speech to speech processing have not been heavily explored. We propose Speech to Speech Synthesis Network (STSSN), a model based on current state of the art systems that fuses the two disciplines in order to perform effective speech to speech style transfer for the purpose of voice impersonation. We show that our proposed model is quite powerful, and succeeds in generating realistic audio samples despite a number of drawbacks in its capacity. We benchmark our proposed model by comparing it with a generative adversarial model which accomplishes a similar task, and show that ours produces more convincing results.
机器翻译由腾讯交互翻译提供,仅供参考
