今日论文合集:CS.SD语音与音频 | 共 6 篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

快速导航

1. 语音合成与声音生成 1 篇

2. 音乐信息检索与音乐生成 3 篇

3. 语音翻译与语音语言模型 1 篇

4. 其他/综合语音音频 1 篇

1. 语音合成与声音生成 | 1 篇

1. ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models

ReGen:用于高效波形扩散模型的分层多提示表示生成

AI 总结:研究针对扩散训练中中间表示正则化问题,提出ReGen框架联合估计向量场,引入GFM提高泛化能力。在波形扩散模型及文本到语音模型上验证,提升了波形生成质量、语音清晰度等,实现高效训练与采样。

链接:https://arxiv.org/abs/2607.09134

机构:Department of Artificial Intelligence, Ajou University, Suwon, Korea(人工智能系,全州大学); KT Corp., Seoul, Korea(KT公司)

作者:Sang-Hoon Lee, Ha-Yeong Choi

英文摘要:Representation alignment (REPA) has been investigated to accelerate diffusion training, but we observe that regularizing intermediate representations in diffusion Transformers (DiT) may implicitly entangle latents and limit generative capacity. To address this issue, we propose ReGen, a hierarchical multi-prompt representation generation framework that jointly estimates multiple vector fields for both representations and data within a single diffusion model. We further introduce generalized flow matching (GFM) to improve the generalization of conditional flow matching (CFM). We validate ReGen on single-stage waveform diffusion models including neural audio codec and Wave-VAE. ReGen significantly improves waveform generation quality from highly compressed latent representations at 12.5 Hz. We also present ReGenVoice, a latent diffusion model (LDM)-based text-to-speech model that achieves strong speech intelligibility (WER) and speaker similarity (SIM) with a small dataset. Moreover, operating the LDM at 6.25 Hz with rich semantic and acoustic latent representation enables efficient training and sampling, requiring only 1 day of training on 4 GPUs and fast inference with an RTF of 0.08. Audio samples are available at this https URL.

2. 音乐信息检索与音乐生成 | 3 篇

2. Tonnetz-Driven Graph Wedgelet for Harmonic Complexity Reduction in Music Scores

用于降低乐谱中和声复杂性的托内兹驱动图楔块

AI 总结:研究如何降低乐谱和声复杂性,提出基于二元楔块划分树的压缩方案,通过完全自适应贪婪算法在六维托内兹嵌入中生成楔块,经实验验证该方法能有效简化乐谱并保留关键信息。

链接:https://arxiv.org/abs/2607.08806

机构:University of Palermo(巴勒莫大学)

作者:Emmanuel Caronna, Elisa Francomano, Silvia Licciardi

英文摘要:Heterogeneous graph built on notes, lyric syllables, and accompaniment events is a natural representation of symbolic music score, providing a substrate for both philological analysis and computational tasks. Music features are therefore well-captured by graph geometry and its properties. This representation has proved effective for analytical tasks as cadence detection, voice separation, and stylistic classification. In the present work, the reduction of harmonic complexity of a music score on graph, by preserving task-relevant information, relation between notes, and graph structure is investigated. A compression scheme for the piano subgraph of vocal-pianistic scores, built on binary wedge partitioning trees, is proposed. The wedges are generated through a fully adaptive greedy algorithm that recursively minimizes the $L^2$-error within a six-dimensional Tonnetz embedding of musical notes. The partitioning process employs a splitting criterion based on harmonic distance, resulting in regions that accurately reflect the intrinsic harmonic relationships among notes. The reconstructed music scores obtained through piecewise-constant functions and the mean values of the notes inside each wedge are used as a new simplified scores human-readable and playable. Some experiments on a corpus of symbolic music scores of three different composers are performed to assess the proposed approach.

3. Clean2FX: Label-conditioned modeling for clean-to-effect guitar audio transformations

Clean2FX:用于清音到效果音吉他音频转换的标签条件建模

AI 总结:研究电吉他音频清音到效果音转换,基于频谱图设置评估四种神经方法,包括两个变分自编码器和两个U-Net模型,U-Net模型表现更佳,不同效果改善情况有差异,还通过演示网站展示了模型应用效果。

链接:https://arxiv.org/abs/2607.08863

作者:Oliverio Bombicci Pontelli, Iran R. Roman

英文摘要:We present Clean2FX, a study and demo of label-conditioned clean-to-effect transformation for electric guitar audio. Given a clean guitar input and a target effect label, the task is to synthesize the corresponding effected signal while preserving the musical content. Training and evaluation pairs are constructed from EGFxSet real, single tone recordings by assembling matched clean/effected chords, melodies, and mixed timelines. This allows for controlled comparison across effects. We evaluate four neural approaches under a common spectrogram-based transformation setting: two variational autoencoders and two U-Net models that differ in whether they operate on linear or log-magnitude representations. Performance is measured using linear-magnitude spectrogram MSE and Fréchet Audio Distance. The U-Net models outperform the variational autoencoder variants. Per-effect results show that distortion effects are most readily improved, whereas delay and reverb effects exhibit weaker FAD gains despite substantial spectral-error reductions. A conditioning-sensitivity diagnostic provides evidence that the best model responds to target labels rather than collapsing to a single transformation. Our demo website compares two models applied on real-world guitar performances outside training and validation data, providing audio and spectrogram examples of the practical clean-to-effect behavior.

4. Event-Based Token Sequences for Audio-Conditioned Music-Game Level Modeling

用于音频条件音乐游戏关卡建模的基于事件的令牌序列

AI 总结:研究音乐游戏关卡程序生成问题,受基于事件的符号音乐建模启发,提出令牌级序列公式,构建Transformer模型,在事件级评估中优于基线,能系统分析音频对节奏对齐事件预测的支持。

链接:https://arxiv.org/abs/2607.09095

机构:Japan Advanced Institute of Science and Technology(日本先进科学技术学院)

作者:Ke Zhang, Chu-Hsuan Hsueh, Kokolo Ikeda

英文摘要:Procedural generation of music game levels is an exciting yet challenging problem, as levels must translate musical structure into interactive sequences of timed gameplay events. Most existing approaches formulate this task by frame-based representations, dividing audio into uniform time grids and predicting events at each frame. This makes gameplay events implicit across many frames. As a result, it is hard to describe event-level timing relations and longer-range structure found in human-authored levels. We use procedural generation as a practical setting to study how musical cues map to interactive event sequences. Inspired by event-based symbolic music modeling, we propose a token-level sequence formulation that casts level generation as a multimodal sequence-to-sequence problem. Conditioned on an audio excerpt and level metadata, the model generates a token sequence alternating gameplay-event and beat-shift tokens. This explicitly represents actions and their relative timing in beat space. Based on this formulation, we build a Transformer model. It outperforms representative frame-level baselines under event-level evaluation. It also enables systematic analysis of how audio supports rhythm-aligned event prediction beyond metadata conditioning.

3. 语音翻译与语音语言模型 | 1 篇

5. Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition

基于最优传输的基于大语言模型的视听语音识别语义对齐

AI 总结:研究基于大语言模型的视听语音识别,提出基于最优传输的语义对齐框架,通过在多模态融合前对齐声学和视觉表征弥合模态差距,经实验验证该方法能有效提升性能,在多种条件下达到最优。

链接:https://arxiv.org/abs/2607.09001

作者:Xugang Lu, Peng Shen, Yu Tsao, Hisashi Kawai

英文摘要:Large language model (LLM)-based audio-visual speech recognition (LLM-AVSR) has recently demonstrated strong robustness in adverse acoustic environments by leveraging complementary audio and visual information. Existing approaches typically employ independently pretrained acoustic and visual encoders, whose outputs are projected and fused as soft prompts to condition an LLM for speech recognition. However, most methods perform multimodal fusion without explicitly addressing the representational discrepancy between audio, visual and text modalities, potentially limiting the effectiveness of cross-modal integration. In this paper, we propose an optimal transport (OT)-based semantic alignment framework for LLM-AVSR. The proposed method explicitly bridges the modality gap by aligning the acoustic and visual representations with reference to the linguistic embedding space of the LLM before multimodal fusion. Specifically, OT is used to estimate probabilistic coupling matrices that characterize structured correspondences between modality-specific features and linguistic embeddings. The resulting OT couplings are further utilized as soft pseudo-labels to supervise contrastive learning, encouraging the extraction of semantically coherent and cross-modal consistent audio-visual representations. By anchoring multimodal features to the linguistic space of the LLM, the proposed framework facilitates more effective multimodal fusion and decoding. We implement the proposed framework using a Whisper-based acoustic encoder, an AV-HuBERT-based visual encoder, and a LLaMA3.2-3B decoder. Experiments conducted on the LRS3-TED benchmark demonstrate consistent improvements over strong baselines and achieve state-of-the-art performance under both clean and noisy evaluation conditions across a wide range of signal-to-noise ratios (SNRs).

4. 其他/综合语音音频 | 1 篇

6. Dual-BEATs: Unlocking Zero-Shot Stereo Audio Perception in Audio Large Language Models via Dithering

双BEATs:通过抖动在音频大语言模型中解锁零样本立体声音频感知

AI 总结:研究针对多模态大语言模型空间感知局限,提出双BEATs架构,通过在编码前注入抖动噪声解决归一化问题,在三元方向分类任务中验证该方法有出色空间分辨率且能零样本泛化,证明标准模型经正则化可实现广义立体声音频理解。

链接:https://arxiv.org/abs/2607.08800

机构:Institute of Information Science, Academia Sinica(台湾中央研究院资讯科学研究所)

作者:Shuo-Chun Lin, Hen-Hsen Huang

英文摘要:Multimodal Large Language Models (LLMs) have remarkable semantic audio understanding, yet they remain "spatially agnostic" due to their reliance on mono-channel audio representations. Currently, spatial audio perception methods mainly focus on complex room simulations and custom-trained, geometry-aware stereo encoders, which limits their accessibility and generalizability. In this paper, we introduce the Dual-BEATs architecture, in which the left and right audio channels are routed independently through two identical semantic encoders as an alternative to specialized spatial modules. To circumvent the architectural bottleneck where internal normalization otherwise erases the inter-channel variance of stereo audio, we inject a static, uncorrelated dithering noise floor prior to encoding. This dithering intervention establishes a macro-variance floor that "smuggles" spatial geometry across the normalization layers. Evaluated on a ternary directional classification task (Left, Center, Right), we demonstrate that dithered models achieve exceptional spatial resolution--reaching up to 97.2% localization accuracy even on subtle 0.5 panning amplitudes--and demonstrates robust, zero-shot generalization to entirely unseen spatial configurations. Our results suggest that with the appropriate acoustic regularization, standard multimodal models are natively capable of generalized stereo audio understanding.