今日论文合集:CS.SD语音与音频 | 共 5 篇。
本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
1. Geometric Iterative Retrieval for Neural Audio Codec Resynthesis
神经音频编解码器重合成的几何迭代检索
AI 总结:该研究针对神经音频编解码器从粗令牌重合成高质量音频的问题,提出几何迭代检索方法,在语音和音乐编解码器恢复任务上优于单次令牌预测和单步回归基线。
链接:https://arxiv.org/abs/2608.19141
机构:ETH-DISCO
作者:Leo Schmidt-Traub, Frédéric Berdoz, Luca A. Lanzendörfer, Roger Wattenhofer
英文摘要:Neural audio codecs based on Residual Vector Quantization (RVQ) have become the dominant discrete representation for token-based general audio generation, yet resynthesizing high-quality audio from coarse codec tokens remains an open problem and bounds the fidelity of every system that generates them. Prior work has framed resynthesis as a choice between discrete token prediction and continuous regression. We argue that this dichotomy is incomplete and introduce geometric iterative retrieval, a paradigm that uses the RVQ layer hierarchy itself as a natural iterative decomposition in continuous codebook space. Rather than classifying over discrete vocabularies or regressing to a single target vector, our method performs contrastive retrieval in the codebook's geometric space. We evaluate our method on codec restoration tasks across speech and music, and show improvements over both single-pass token prediction and one-step regression baselines.
2. Computational Features for Symbolic Melody Analysis
用于符号旋律分析的计算特征
AI 总结:本文提出开源Python工具melody-features,梳理符号旋律的音乐理论与心理学特征并构建分类体系,在Essen民歌集上验证其风格分类性能,为旋律分析提供可解释的特征方案。
链接:https://arxiv.org/abs/2608.19061
机构:University of Cambridge(剑桥大学)
作者:David M. Whyatt, Peter M. C. Harrison
英文摘要:This paper addresses the general problem of extracting music-theoretic and psychological features from symbolically encoded melodies. We review existing melodic feature extraction toolboxes, enumerate their features, and organise them into a common taxonomy. We then describe a new software library that provides implementations of all of these features in a straightforward Python package. We then demonstrate the combined feature set on the Essen Folksong Collection, using the dataset to produce a series of style classification models. These models help us answer key questions about the interpretability and dimensionality of the feature set. Our results show excellent classification accuracy using the full feature set, and promising performance for an eight-dimensional factor-analytic solution that improves the interpretability of the classifier. We distribute our new toolbox as an open-source Python package, $\mathtt{melody-features}$, which can easily be used in various applications within music analysis, music psychology, and music information retrieval.
3. Mitigating Spectral Bias in Neural Operators for Underwater Transmission Loss Prediction
缓解水下传输损耗预测神经算子中的频谱偏差
AI 总结:针对水下传输损耗预测中FNO的频谱偏差问题,提出S2RL框架,通过频谱全局传播器与空间局部细化器的从粗到细结构,在保持毫秒级推理速度的同时,显著优于FNO基线。
链接:https://arxiv.org/abs/2608.18141
作者:Yifan Sun, Shikai Fang, Chao Zhang, Lei Cheng, Jianlong Li, Peter Gerstoft
英文摘要:Predicting underwater acoustic transmission loss rapidly and accurately is crucial for real-time ocean acoustic applications. While Fourier Neural Operators (FNO) have emerged as powerful surrogate models due to their global receptive fields, they suffer from spectral bias. The frequency truncation mechanism in FNO filters out high-frequency components, resulting in over-smoothed predictions that fail to capture fine-grained interference patterns. To overcome this limitation, this paper proposes a Spectral-Spatial Residual Learning (S2RL) framework. S2RL decomposes the prediction task into a coarse-to-fine process: a spectral Global Propagator first generates a globally consistent prediction, and a spatial Local Refiner subsequently recovers the high-frequency residuals. Experimental results on a South China Sea dataset show that the proposed method significantly outperforms FNO baselines while maintaining millisecond-level inference speeds.
4. FM Synthesizer Audio-Parameter Shared Embeddings
FM合成器音频参数共享嵌入
AI 总结:针对合成器预设检索问题,本文设计模仿FM信号处理的图神经网络学习含信号路由的参数表征,结合SLAP多模态目标学习音频与FM合成器参数的联合嵌入,在Yamaha DX7数据集上验证了方法的有效性与泛化性。
链接:https://arxiv.org/abs/2608.18226
作者:David Braun, Adam Finkelstein
英文摘要:Given a target sound, finding the synthesizer preset that best reproduces it remains a core problem in sound design. Existing methods treat synthesis parameters as flat vectors, discarding the signal routing and parameter interactions that produce audio. We make two contributions. First, to learn a representation of parameters including their signal routing, we design a graph neural network whose message passing structure imitates FM signal processing. Second, we adapt the multimodal objective from SLAP to learn joint embeddings of audio and FM synthesizer parameters, enabling preset retrieval from a gallery. We focus on the Yamaha DX7, where six identical sinusoid operators interact according to one of 32 routing topologies. Our graph encoder's message passing weights are shared across all nodes and layers, enabling processing of arbitrary topologies of any size. When every topology is seen during training, the DX7-GNN and two baselines achieve strong audio-to-preset retrieval. When some topologies are held out for testing, the DX7-GNN substantially outperforms both baselines despite having the fewest parameters. Our ablations further support the claim that imitating FM signal flow in a parameter encoder improves generalization to unseen topologies.
5. Finetuning Strategies for Querying Sounds by Vocal Imitation
通过语音模仿查询声音的微调策略
AI 总结:针对AES AIMLA 2025挑战赛的音效语音查询任务,研究人员提出两种微调策略,即基于CED编码器的对比学习和基于MobileNetV3编码器的联合对比-三元组学习,其方案为挑战赛获胜方案。
链接:https://arxiv.org/abs/2608.19174
机构:School of Electronic Engineering and Computer Science, Queen Mary University of London(伦敦玛丽女王大学电子工程与计算机科学学院)
作者:Aditya Bhattacharjee, Christos Plachouras, Sungkyun Chang, Emmanouil Benetos
英文摘要:This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge.
语音与音频学术速递[8.20]
评论 0
文明发言,友善讨论
还没有评论,发表你的看法,来抢沙发~
