今日论文合集:cs.SD语音12篇,eess.AS音频处理0篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation
标题:Vo-Ve:用于说话者身份评估的可解释语音载体
链接:https://arxiv.org/abs/2506.19446
作者:e, Kyogu Lee
备注:Interspeech 2025
摘要:在本文中,我们提出了Vo-Ve,一种新颖的语音矢量嵌入,捕捉说话人的身份。与传统的说话人嵌入不同,Vo-Ve是可解释的,因为它包含显式语音属性类的概率。通过广泛的分析,我们表明,Vo-Ve不仅评估说话人相似性与传统的技术竞争,但也提供了一个可解释的解释的语音属性。我们坚信,Vo-Ve可以增强各种语音任务的评估计划,由于其高层次的可解释性。
摘要:In this paper, we propose Vo-Ve, a novel voice-vector embedding that captures speaker identity. Unlike conventional speaker embeddings, Vo-Ve is explainable, as it contains the probabilities of explicit voice attribute classes. Through extensive analysis, we demonstrate that Vo-Ve not only evaluates speaker similarity competitively with conventional techniques but also provides an interpretable explanation in terms of voice attributes. We strongly believe that Vo-Ve can enhance evaluation schemes across various speech tasks due to its high-level explainability.


【2】TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems

标题:TTSDS2:评估人性化文本到语音系统的资源和基准
链接:https://arxiv.org/abs/2506.19441
作者: Minixhofer, Ondrej Klejch, Peter Bell
摘要:文语转换(TTS)系统的评估具有挑战性和资源密集型。主观指标,如平均意见得分(MOS),不容易在作品之间进行比较。客观指标经常使用,但很少验证主观指标。这两种指标的挑战,最近的TTS系统能够产生合成语音与真实语音难以区分。在这项工作中,我们介绍了文本到语音分布评分2(TTSDS 2),一个更强大和改进的版本TTSDS。在一系列领域和语言中,它是16个比较指标中唯一一个与斯皮尔曼相关性高于0.50的指标。我们还发布了一系列用于评估接近真实语音的合成语音的资源:具有超过11,000个主观意见评分的数据集;用于不断重新创建多语言测试数据集以避免数据泄漏的管道;以及用于14种语言的TTS的不断更新的基准。
摘要:Evaluation of Text to Speech (TTS) systems is challenging and resource-intensive. Subjective metrics such as Mean Opinion Score (MOS) are not easily comparable between works. Objective metrics are frequently used, but rarely validated against subjective ones. Both kinds of metrics are challenged by recent TTS systems capable of producing synthetic speech indistinguishable from real speech. In this work, we introduce Text to Speech Distribution Score 2 (TTSDS2), a more robust and improved version of TTSDS. Across a range of domains and languages, it is the only one out of 16 compared metrics to correlate with a Spearman correlation above 0.50 for every domain and subjective score evaluated. We also release a range of resources for evaluating synthetic speech close to real speech: A dataset with over 11,000 subjective opinion score ratings; a pipeline for continually recreating a multilingual test dataset to avoid data leakage; and a continually updated benchmark for TTS in 14 languages.


【3】ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

标题:ClearerVoice-Studio:连接高级语音处理研究和实际部署
链接:https://arxiv.org/abs/2506.19398
作者:Zhao, Zexu Pan, Bin Ma
备注:accepted by Interspeech 2025, 5 pages, 5 tables
摘要:本文介绍了ClearerVoice-Studio,这是一个开源的人工智能语音处理工具包,旨在连接前沿研究和实际应用。与SpeechBrain和ESPnet等广泛的平台不同,ClearerVoice-Studio专注于语音增强、分离、超分辨率和多模式目标说话人提取等相互关联的语音任务。一个关键的优势是其最先进的预训练模型,包括具有300万次使用的FRCRN和具有250万次使用的MossFormer,针对真实场景进行了优化。它还提供模型优化工具,多格式音频支持,SpeechScore评估工具包和用户友好的界面,以满足研究人员,开发人员和最终用户的需求。它的快速采用吸引了3000个GitHub星和239个分叉,凸显了它的学术和工业影响力。本文详细介绍了ClearerVoice-Studio的功能、架构、培训策略、基准、社区影响和未来计划。源代码可在https://github.com/modelscope/ClearerVoice-Studio上获得。
摘要:This paper introduces ClearerVoice-Studio, an open-source, AI-powered speech processing toolkit designed to bridge cutting-edge research and practical application. Unlike broad platforms like SpeechBrain and ESPnet, ClearerVoice-Studio focuses on interconnected speech tasks of speech enhancement, separation, super-resolution, and multimodal target speaker extraction. A key advantage is its state-of-the-art pretrained models, including FRCRN with 3 million uses and MossFormer with 2.5 million uses, optimized for real-world scenarios. It also offers model optimization tools, multi-format audio support, the SpeechScore evaluation toolkit, and user-friendly interfaces, catering to researchers, developers, and end-users. Its rapid adoption attracting 3000 GitHub stars and 239 forks highlights its academic and industrial impact. This paper details ClearerVoice-Studio's capabilities, architectures, training strategies, benchmarks, community impact, and future plan. Source code is available at https://github.com/modelscope/ClearerVoice-Studio.


【4】Learning to assess subjective impressions from speech

标题:学习评估言语的主观印象
链接:https://arxiv.org/abs/2506.19335
作者:o, Hirokazu Kameoka, Kou Tanaka, Takuhiro Kaneko, Noboru Harada
备注:Accepted on EUSIPCO 2024
摘要:我们解决了一个新的任务,训练神经网络模型,可以评估通过语音传达的主观印象,并相应地分配分数,灵感来自自动语音质量评估(SQA)的工作。人们经常用“可爱的声音”这样的短语来描述言语印象。“我们将这些短语定义为主观语音描述符(SVD)。针对所提出的任务和自动SQA之间的使用场景的差异,我们设计了一个框架,能够容纳个性化的SVDs到每个人,如“我最喜欢的声音。“在这项工作中,我们编译了一个包含来自绝对类别评级(ACR)和比较类别评级(CCR)的语音标签的数据集。   作为评估性能的评价指标,我们引入了预测分数排序的准确性,两个样本的CCR测试样本。除了基于ACR数据的传统模型和学习方法外,我们还研究了使用CCR数据的RankNet学习。我们通过实验发现,即使在非常有限的训练数据下,概率也是中等的。我们还发现CCR训练优于ACR训练。这些结果支持这样的想法,即基于个性化SVD的评估模型,通常必须在有限的数据上进行训练,可以有效地从CCR数据中学习。
摘要:We tackle a new task of training neural network models that can assess subjective impressions conveyed through speech and assign scores accordingly, inspired by the work on automatic speech quality assessment (SQA). Speech impressions are often described using phrases like `cute voice.' We define such phrases as subjective voice descriptors (SVDs). Focusing on the difference in usage scenarios between the proposed task and automatic SQA, we design a framework capable of accommodating SVDs personalized to each individual, such as `my favorite voice.' In this work, we compiled a dataset containing speech labels derived from both abosolute category ratings (ACR) and comparison category ratings (CCR).   As an evaluation metric for assessment performance, we introduce ppref, the accuracy of the predicted score ordering of two samples on CCR test samples. Alongside the conventional model and learning methods based on ACR data, we also investigated RankNet learning using CCR data. We experimentally find that the ppref is moderate even with very limited training data. We also discover the CCR training is superior to the ACR training. These results support the idea that assessment models based on personalized SVDs, which typically must be trained on limited data, can be effectively learned from CCR data.


【5】A Robust Method for Pitch Tracking in the Frequency Following Response using Harmonic Amplitude Summation Filterbank

标题:一种使用和声幅总和过滤器组的频率跟随响应音调跟踪的稳健方法
链接:https://arxiv.org/abs/2506.19253
作者:eghkhani, Maryam Karimi Boroujeni, Hilmi R. Dajani, Saeid R. Seydnejad, Christian Giguère
摘要:频率跟随响应(FFR)反映了大脑对包括语音在内的听觉刺激的神经编码。因为基频(F0)(音高的物理相关)是语音的基本特征之一,所以对表征F0处的FFR特别感兴趣,特别是当F0随时间变化时。在FFR中提取F0的标准方法是自相关函数(ACF)。本文研究了基于谐波结构的F0估计算法,最初是为语音和音乐开发的,并通过两个步骤解决了它们在应用于FFR时的不良性能。首先,考虑到与语音或音乐不同,FFR的刺激F0是已知的,我们引入了一个刺激感知滤波器组,该滤波器组选择性地聚合F0及其谐波处的幅度,同时抑制非谐波频率处的噪声。这种方法称为谐波振幅求和(HAS),仅在以刺激F0为中心的范围内评估F0候选。其次,与选择最高峰的其他音高跟踪方法不同,我们的方法选择最突出的一个,因为它更好地反映了FFR的潜在周期性。据我们所知,这是第一项提出依赖于谐波结构的FFR F0估计算法的研究。分析记录的FFR从16个正常听力受试者4个自然语音刺激与广泛的F0变化从89 Hz到452 Hz表明,这种方法优于ACF减少平均根均方误差(RMSE)内的每个响应和刺激F0轮廓对由8.8%至47.4%,这取决于刺激。
摘要:The Frequency Following Response (FFR) reflects the brain's neural encoding of auditory stimuli including speech. Because the fundamental frequency (F0), a physical correlate of pitch, is one of the essential features of speech, there has been particular interest in characterizing the FFR at F0, especially when F0 varies over time. The standard method for extracting F0 in FFRs has been the Autocorrelation Function (ACF). This paper investigates harmonic-structure-based F0 estimation algorithms, originally developed for speech and music, and resolves their poor performance when applied to FFRs in two steps. Firstly, given that unlike in speech or music, stimulus F0 of FFRs is already known, we introduce a stimulus-aware filterbank that selectively aggregates amplitudes at F0 and its harmonics while suppressing noise at non-harmonic frequencies. This method, called Harmonic Amplitude Summation (HAS), evaluates F0 candidates only within a range centered around the stimulus F0. Secondly, unlike other pitch tracking methods that select the highest peak, our method chooses the most prominent one, as it better reflects the underlying periodicity of FFRs. To the best of our knowledge, this is the first study to propose an F0 estimation algorithm for FFRs that relies on harmonic structure. Analyzing recorded FFRs from 16 normal hearing subjects to 4 natural speech stimuli with a wide F0 variation from 89 Hz to 452 Hz showed that this method outperformed ACF by reducing the average Root-Mean-Square-Error (RMSE) within each response and stimulus F0 contour pair by 8.8% to 47.4%, depending on the stimulus.


【6】Enhanced Hybrid Transducer and Attention Encoder Decoder with Text Data

标题:具有文本数据的增强型混合传感器和注意力编码器解码器
链接:https://arxiv.org/abs/2506.19159
作者: Eesung Kim, Vijendra Raj Apsingekar
备注:Accepted by Interspeech2025
摘要:提出了一种语音和文本联合优化方法,用于混合转换器和基于注意力的编码器解码器(TAED)建模,以利用大量的文本语料,提高ASR的准确性。联合TAED(J-TAED)是用语音和文本输入模态一起训练的,而它在推理过程中只接受语音数据作为输入。训练好的模型可以统一不同模态的内部表示,并进一步扩展到基于文本的领域自适应。由于不需要语音数据,它可以有效地缓解失配域任务的数据稀缺性。实验结果表明,J-TAED模型成功地将语音和语言信息融合到一个模型中,并在Librispeech数据集上将WER降低了5.8%~12.8%。该模型还在两个域外数据集上进行了评估:一个是财务数据集,另一个是命名实体集中的数据集。基于文本的域自适应在这两个数据集上分别带来了15.3%和17.8%的WER减少。
摘要:A joint speech and text optimization method is proposed for hybrid transducer and attention-based encoder decoder (TAED) modeling to leverage large amounts of text corpus and enhance ASR accuracy. The joint TAED (J-TAED) is trained with both speech and text input modalities together, while it only takes speech data as input during inference. The trained model can unify the internal representations from different modalities, and be further extended to text-based domain adaptation. It can effectively alleviate data scarcity for mismatch domain tasks since no speech data is required. Our experiments show J-TAED successfully integrates speech and linguistic information into one model, and reduce the WER by 5.8 ~12.8% on the Librispeech dataset. The model is also evaluated on two out-of-domain datasets: one is finance and another is named entity focused. The text-based domain adaptation brings 15.3% and 17.8% WER reduction on those two datasets respectively.


【7】A Fourier Explanation of AI-music Artifacts

标题:人工智能音乐文物的傅里叶解释
链接:https://arxiv.org/abs/2506.19108
作者:char, Gabriel Meseguer-Brocal, Kamil Akesbi, Romain Hennequin
备注:Accepted at ISMIR 2025
摘要:人工智能的迅速崛起改变了音乐创作,数百万用户参与了人工智能生成的音乐。尽管它很受欢迎,但对侵犯版权、工作岗位流失和道德影响的担忧导致了越来越多的审查和法律挑战。与此同时,人工智能检测服务也出现了,但这些系统在很大程度上仍然不透明,并由私人控制,反映了它们旨在解决的问题。本文探讨了合成内容的基本属性以及如何检测合成内容。具体来说,我们分析了生成模型中常用的反卷积模块,并在数学上证明了它们的输出表现出系统的频率伪影-表现为小而独特的光谱峰。这种现象,相关的众所周知的棋盘工件,被证明是固有的选择模型架构,而不是训练数据或模型权重的后果。我们通过对开源模型以及Suno和Udio等商业AI音乐生成器的广泛实验来验证我们的理论发现。我们利用这些见解为AI生成的音乐提出了一个简单且可解释的检测标准。尽管简单,但我们的方法实现了与基于深度学习的方法相当的检测准确性,在几种情况下的准确率超过99%。
摘要:The rapid rise of generative AI has transformed music creation, with millions of users engaging in AI-generated music. Despite its popularity, concerns regarding copyright infringement, job displacement, and ethical implications have led to growing scrutiny and legal challenges. In parallel, AI-detection services have emerged, yet these systems remain largely opaque and privately controlled, mirroring the very issues they aim to address. This paper explores the fundamental properties of synthetic content and how it can be detected. Specifically, we analyze deconvolution modules commonly used in generative models and mathematically prove that their outputs exhibit systematic frequency artifacts -- manifesting as small yet distinctive spectral peaks. This phenomenon, related to the well-known checkerboard artifact, is shown to be inherent to a chosen model architecture rather than a consequence of training data or model weights. We validate our theoretical findings through extensive experiments on open-source models, as well as commercial AI-music generators such as Suno and Udio. We use these insights to propose a simple and interpretable detection criterion for AI-generated music. Despite its simplicity, our method achieves detection accuracy on par with deep learning-based approaches, surpassing 99% accuracy on several scenarios.


【8】Benchmarking Music Generation Models and Metrics via Human Preference Studies

标题:通过人类偏好研究对音乐生成模型和数据进行基准测试
链接:https://arxiv.org/abs/2506.19085
作者:rötschla, Ahmet Solak, Luca A. Lanzendörfer, Roger Wattenhofer
备注:Accepted at ICASSP 2025
摘要:最近的进步使生成的音乐更接近人类创作的作品,但评估这些模型仍然具有挑战性。虽然人类偏好是评估质量的黄金标准,但将这些主观判断转化为客观指标,特别是文本音频对齐和音乐质量,已被证明是困难的。在这项工作中,我们使用12种最先进的模型生成了6k首歌曲,并与2.5k名人类参与者进行了15k对音频比较的调查,以评估人类偏好与广泛使用的指标之间的相关性。据我们所知,这项工作是第一次根据人类偏好对当前最先进的音乐生成模型和指标进行排名。为了进一步推进主观度量评估领域,我们提供了对生成的音乐和人类评估数据集的开放访问。
摘要:Recent advancements have brought generated music closer to human-created compositions, yet evaluating these models remains challenging. While human preference is the gold standard for assessing quality, translating these subjective judgments into objective metrics, particularly for text-audio alignment and music quality, has proven difficult. In this work, we generate 6k songs using 12 state-of-the-art models and conduct a survey of 15k pairwise audio comparisons with 2.5k human participants to evaluate the correlation between human preferences and widely used metrics. To the best of our knowledge, this work is the first to rank current state-of-the-art music generation models and metrics based on human preference. To further the field of subjective metric evaluation, we provide open access to our dataset of generated music and human evaluations.


【9】IndieFake Dataset: A Benchmark Dataset for Audio Deepfake Detection

标题:IndieFake数据集:音频Deepfake检测的基准数据集
链接:https://arxiv.org/abs/2506.19014
作者:ar, Kunal Verma, Omkar More
摘要:音频深度伪造技术的进步提供了人工智能助手、更好地帮助言语障碍者以及增强娱乐性等好处。然而,它也对数字通信的安全、隐私和信任构成了重大风险。检测和缓解这些威胁需要全面的数据集。现有的数据集缺乏不同的民族口音,这使得它们不适合许多现实世界的场景。因此,在这些数据集上训练的模型很难在不同的语言和文化背景下检测音频deepfake,例如在南亚国家。具有讽刺意味的是,在现有的数据集中,尽管南亚人占世界人口的四分之一,但却严重缺乏南亚人的样本。这项工作介绍了IndieFake数据集(IFD),其中包括来自50名讲英语的印度人的27.17小时的真实和deepfake音频。IFD提供了平衡的数据分布,包括说话者级别的特征,这在ASVspoof 21(DF)等数据集中是不存在的。我们根据现有的ASVspoof 21(DF)和In-The-Wild(ITW)数据集评估了IFD的各种基线。IFD优于ASVspoof 21(DF),并且与基准ITW数据集相比更具挑战性。该数据集将在接受后公开提供。
摘要:Advancements in audio deepfake technology offers benefits like AI assistants, better accessibility for speech impairments, and enhanced entertainment. However, it also poses significant risks to security, privacy, and trust in digital communications. Detecting and mitigating these threats requires comprehensive datasets. Existing datasets lack diverse ethnic accents, making them inadequate for many real-world scenarios. Consequently, models trained on these datasets struggle to detect audio deepfakes in diverse linguistic and cultural contexts such as in South-Asian countries. Ironically, there is a stark lack of South-Asian speaker samples in the existing datasets despite constituting a quarter of the worlds population. This work introduces the IndieFake Dataset (IFD), featuring 27.17 hours of bonafide and deepfake audio from 50 English speaking Indian speakers. IFD offers balanced data distribution and includes speaker-level characterization, absent in datasets like ASVspoof21 (DF). We evaluated various baselines on IFD against existing ASVspoof21 (DF) and In-The-Wild (ITW) datasets. IFD outperforms ASVspoof21 (DF) and proves to be more challenging compared to benchmark ITW dataset. The dataset will be publicly available upon acceptance.


【10】SHAMaNS: Sound Localization with Hybrid Alpha-Stable Spatial Measure and Neural Steerer

标题:SHAMaNS:采用混合Alpha稳定空间测量和神经转向器的声音定位
链接:https://arxiv.org/abs/2506.18954
作者:Carlo (RIKEN AIP), Mathieu Fontaine (LTCI, IP Paris), Aditya Arie Nugraha (RIKEN AIP), Yoshiaki Bando (RIKEN AIP), Kazuyoshi Yoshii
备注:European Signal Processing Conference (EUSIPCO), Sep 2025, Palermo, Italy
摘要:本文介绍了一种声源定位(SSL)技术,它结合了一个$\alpha$-稳定的模型与神经网络为基础的方法建模导向矢量的观测信号。具体而言,一个物理信息的神经网络,被称为神经转向器,用于在固定的麦克风阵列上插入测量的导向矢量(SV)。这允许对所谓的$\alpha$-稳定的空间测量进行更鲁棒的估计,该空间测量表示目标信号的最合理的到达方向(DOA)。由于非高斯情况下的$\alpha$稳定模型($\alpha$ $\in$(0,2))理论上定义了一个唯一的空间度量,因此我们选择利用它来解释下游任务中神经转向器的残余重建误差。客观分数表明,我们提出的技术优于国家的最先进的方法在多个声源的情况下。
摘要:This paper describes a sound source localization (SSL) technique that combines an $\alpha$-stable model for the observed signal with a neural network-based approach for modeling steering vectors. Specifically, a physics-informed neural network, referred to as Neural Steerer, is used to interpolate measured steering vectors (SVs) on a fixed microphone array. This allows for a more robust estimation of the so-called $\alpha$-stable spatial measure, which represents the most plausible direction of arrival (DOA) of a target signal. As an $\alpha$-stable model for the non-Gaussian case ($\alpha$ $\in$ (0, 2)) theoretically defines a unique spatial measure, we choose to leverage it to account for residual reconstruction error of the Neural Steerer in the downstream tasks. The objective scores indicate that our proposed technique outperforms state-of-the-art methods in the case of multiple sound sources.


【11】Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation

标题:Kling-Foley:用于高质量视频到音频生成的多模式扩散Transformer
链接:https://arxiv.org/abs/2506.19774
作者: Xijuan Zeng, Chunyu Qiang, Ruilong Chen, Shiyao Wang, Le Wang, Wangjing Zhou, Pengfei Cai, Jiahui Zhao, Nan Li, Zihan Li, Yuzhe Liang, Xiaopeng Wang, Haorui Zheng, Ming Wen, Kang Yin, Yiran Wang, Nan Li, Feng Deng, Liang Dong, Chen Zhang, Di Zhang, Kun Gai
摘要:我们提出Kling-Foley,一个大规模的多模态视频到音频生成模型,合成高质量的音频与视频内容同步。在Kling-Foley中,我们引入了多模态扩散Transformers来模拟视频,音频和文本模态之间的交互,并将其与视觉语义表示模块和视听同步模块相结合,以增强对齐能力。具体而言,这些模块在帧级将视频条件与潜在音频元素对齐,从而改善语义对齐和视听同步。结合文本条件,这种集成方法可以精确生成视频匹配音效。此外,我们还提出了一种通用的潜在音频编解码器,可以在各种场景中实现高质量的建模,如音效,语音,唱歌和音乐。我们采用立体声渲染方法,使合成音频具有空间存在感。同时,为了弥补开源基准测试的类型和注释不完整,我们还开源了一个工业级基准测试Kling-Audio-Eval。我们的实验表明,Kling-Foley训练的流匹配目标实现了新的视听SOTA性能在公共模型之间的分布匹配,语义对齐,时间对齐和音频质量。
摘要:We propose Kling-Foley, a large-scale multimodal Video-to-Audio generation model that synthesizes high-quality audio synchronized with video content. In Kling-Foley, we introduce multimodal diffusion transformers to model the interactions between video, audio, and text modalities, and combine it with a visual semantic representation module and an audio-visual synchronization module to enhance alignment capabilities. Specifically, these modules align video conditions with latent audio elements at the frame level, thereby improving semantic alignment and audio-visual synchronization. Together with text conditions, this integrated approach enables precise generation of video-matching sound effects. In addition, we propose a universal latent audio codec that can achieve high-quality modeling in various scenarios such as sound effects, speech, singing, and music. We employ a stereo rendering method that imbues synthesized audio with a spatial presence. At the same time, in order to make up for the incomplete types and annotations of the open-source benchmark, we also open-source an industrial-level benchmark Kling-Audio-Eval. Our experiments show that Kling-Foley trained with the flow matching objective achieves new audio-visual SOTA performance among public models in terms of distribution matching, semantic alignment, temporal alignment and audio quality.


【12】Loss functions incorporating auditory spatial perception in deep learning -- a review

标题:深度学习中纳入听觉空间感知的损失函数--评论
链接:https://arxiv.org/abs/2506.19404
作者:ely, Stefan Weinzierl, Or Berebi, Fabian Brinkmann
备注:Submitted to I3DA 2025
摘要:双耳再现旨在通过耳机提供具有高感知现实主义的沉浸式空间音频。损失函数在优化和评估生成双耳信号的算法中起着核心作用。然而,传统的信号相关的差异措施往往无法捕捉的感知属性是必不可少的空间音频质量。本文综述了最近的损失函数,将空间感知线索相关的双耳复制。它侧重于应用于双耳信号的损耗,双耳信号通常来自麦克风录音或高保真度立体声信号,而不包括基于房间脉冲响应的双耳信号。在空间音频质量量表(SAQI)的指导下,该综述强调了与源定位和房间响应相关的感知维度,同时排除了一般的频谱-时间属性。文献调查显示,强烈关注本地化线索,如耳间的时间和电平差异(ITDs,ILD),而混响和其他房间声学属性仍然较少探索损失函数设计。最近的工作,估计房间的声学参数和开发嵌入,捕捉房间的特点表明他们的潜力,未来集成到神经网络训练。最后,本文强调未来的研究方向更感性接地损失函数,更好地捕捉听众的空间体验。
摘要:Binaural reproduction aims to deliver immersive spatial audio with high perceptual realism over headphones. Loss functions play a central role in optimizing and evaluating algorithms that generate binaural signals. However, traditional signal-related difference measures often fail to capture the perceptual properties that are essential to spatial audio quality. This review paper surveys recent loss functions that incorporate spatial perception cues relevant to binaural reproduction. It focuses on losses applied to binaural signals, which are often derived from microphone recordings or Ambisonics signals, while excluding those based on room impulse responses. Guided by the Spatial Audio Quality Inventory (SAQI), the review emphasizes perceptual dimensions related to source localization and room response, while excluding general spectral-temporal attributes. The literature survey reveals a strong focus on localization cues, such as interaural time and level differences (ITDs, ILDs), while reverberation and other room acoustic attributes remain less explored in loss function design. Recent works that estimate room acoustic parameters and develop embeddings that capture room characteristics indicate their potential for future integration into neural network training. The paper concludes by highlighting future research directions toward more perceptually grounded loss functions that better capture the listener's spatial experience.


机器翻译由腾讯交互翻译提供,仅供参考