微信公众号:arXiv_Daily
cs.SD语音
备注:Accepted to WASPAA 2025
摘要:时频(TF)双路径模型是目前性能最好的音频源分离网络架构之一,在语音增强、音乐源分离和电影音频源分离方面实现了最先进的性能。虽然它们的特点是参数计数相对较低,但它们仍然需要相当数量的操作,这意味着更长的执行时间。这个问题由于在大量数据上训练更大模型以解决更一般的任务的趋势而加剧,例如最近引入的任务感知统一源分离(TUSS)模型。TUSS旨在使用单个条件模型解决音频源分离任务,它是建立在TF-Locoformer之上的,TF-Locoformer是一种结合卷积和注意力层的TF双路径模型。任务定义以一系列提示的形式出现,这些提示指定要提取的源的数量和类型。在本文中,我们分析了TUSS的设计选择,其目标是优化其性能和复杂性的权衡。我们得到了两个更有效的模型,FasTUSS-8.3G和FasTUSS-11.7G,减少了81%和73%的原始模型的操作,与轻微的性能下降1.2 dB和0.4 dB的平均在所有的基准测试,分别。此外,我们调查的影响,提示条件反射,推导出一个因果TUSS模型。
摘要:Time-Frequency (TF) dual-path models are currently among the best performing audio source separation network architectures, achieving state-of-the-art performance in speech enhancement, music source separation, and cinematic audio source separation. While they are characterized by a relatively low parameter count, they still require a considerable number of operations, implying a higher execution time. This problem is exacerbated by the trend towards bigger models trained on large amounts of data to solve more general tasks, such as the recently introduced task-aware unified source separation (TUSS) model. TUSS, which aims to solve audio source separation tasks using a single, conditional model, is built upon TF-Locoformer, a TF dual-path model combining convolution and attention layers. The task definition comes in the form of a sequence of prompts that specify the number and type of sources to be extracted. In this paper, we analyze the design choices of TUSS with the goal of optimizing its performance-complexity trade-off. We derive two more efficient models, FasTUSS-8.3G and FasTUSS-11.7G that reduce the original model's operations by 81 % and 73 % with minor performance drops of 1.2~dB and 0.4~dB averaged over all benchmarks, respectively. Additionally, we investigate the impact of prompt conditioning to derive a causal TUSS model.
【2】Improving Neural Pitch Estimation with SWIPE Kernels
备注:Accepted at ISMIR 2025
摘要:神经网络已经成为精确估计音调和周期性的主要技术。虽然很多研究都致力于改进网络架构和训练范式,但大多数方法都直接对原始音频波形或通用时频表示进行操作。我们研究了锯齿启发音高估计(SWIPE)内核作为音频前端的使用,并发现这些手工制作的特定于任务的功能可以使神经音高估计器更准确,对噪声更鲁棒,并且更有效。我们评估了常见数据集上的监督和自监督最先进的架构,并表明SWIPE音频前端允许将网络大小减少一个数量级,而不会降低性能。此外,我们表明SWIPE算法本身比通常报道的更准确,优于最先进的自监督神经音高估计器。
摘要:Neural networks have become the dominant technique for accurate pitch and periodicity estimation. Although a lot of research has gone into improving network architectures and training paradigms, most approaches operate directly on the raw audio waveform or on general-purpose time-frequency representations. We investigate the use of Sawtooth-Inspired Pitch Estimation (SWIPE) kernels as an audio frontend and find that these hand-crafted, task-specific features can make neural pitch estimators more accurate, robust to noise, and more parameter-efficient. We evaluate supervised and self-supervised state-of-the-art architectures on common datasets and show that the SWIPE audio frontend allows for reducing the network size by an order of magnitude without performance degradation. Additionally, we show that the SWIPE algorithm on its own is much more accurate than commonly reported, outperforming state-of-the-art self-supervised neural pitch estimators.
备注:ACL 2025 Findings
摘要:在这项研究中,我们研究了利用交叉注意控制在自回归模型中进行有效的音频编辑。受图像编辑方法的启发,我们开发了一种类似于自适应的方法,通过交叉和自我注意机制来指导编辑。集成基于扩散的策略,受Auffelt的影响,我们扩展了模型的功能,以支持细化编辑,建立了一个基线的音频编辑。此外,我们还引入了一种替代方法,结合了MUSICGEN,一种预训练的冻结自回归模型,并提出了三种编辑机制,基于注意力分数的替换,重新加权和细化。我们采用常用的音乐特定的评估指标和人类的研究,以衡量随时间变化的可控性,坚持全球文本线索,和整体的音频现实主义。自动和人工评估表明,所提出的快速提示指导与自回归生成模型的组合在所生成的音频的旋律、动态和节奏方面显著优于基于扩散的基线。我们的代码可在https: github.com billsioros EditGen上获得。
摘要:In this study, we investigate leveraging cross-attention control for efficient audio editing within auto-regressive models. Inspired by image editing methodologies, we develop a Prompt-to-Prompt-like approach that guides edits through cross and self-attention mechanisms. Integrating a diffusion-based strategy, influenced by Auffusion, we extend the model's functionality to support refinement edits, establishing a baseline for prompt-guided audio editing. Additionally, we introduce an alternative approach by incorporating MUSICGEN, a pre-trained frozen auto-regressive model, and propose three editing mechanisms, based on Replacement, Reweighting, and Refinement of the attention scores. We employ commonly-used music-specific evaluation metrics and a human study, to gauge time-varying controllability, adherence to global text cues, and overall audio realism. The automatic and human evaluations indicate that the proposed combination of prompt-to-prompt guidance with autoregressive generation models significantly outperforms the diffusion-based baseline in terms of melody, dynamics, and tempo of the generated audio. Our code is available at https: github.com billsioros EditGen
备注:
摘要:This paper presents a novel approach for detecting mispronunciations by analyzing deviations between a user's original speech and their voice-cloned counterpart with corrected pronunciation. We hypothesize that regions with maximal acoustic deviation between the original and cloned utterances indicate potential mispronunciations. Our method leverages recent advances in voice cloning to generate a synthetic version of the user's voice with proper pronunciation, then performs frame-by-frame comparisons to identify problematic segments. Experimental results demonstrate the effectiveness of this approach in pinpointing specific pronunciation errors without requiring predefined phonetic rules or extensive training data for each target language.
备注:Accepted by ComputEL-8
摘要:在温哥华岛南部的萨尼奇半岛上使用的SEN OTEN语言,正处于积极的语言复兴努力之中,以扭转由于殖民语言政策而导致的语言丧失的趋势。为了支持这些实地工作,社区正在转向数字技术。自动语音识别(ASR)技术在加速语言文档和创建教育资源方面具有很大的潜力。然而,由于有限的数据和其多合成结构和应力驱动的复分解的显著词汇变化,为SEN OTEN开发ASR系统具有挑战性。为了应对这些挑战,我们提出了一个ASR驱动的文档管道,该管道利用来自文本到语音(TTS)系统的增强语音数据和语音基础模型(SFM)的跨语言迁移学习。一个n-gram语言模型也被纳入通过浅融合或n-best恢复,以最大限度地利用现有的数据。在SEN OTEN数据集上的实验表明,在测试集上,单词错误率(WER)为19.34%,字符错误率(CER)为5.09%,词汇表外(OOV)率为57.02%。在过滤了与cedilla相关的小错误后,WER提高到14.32%(未看到的单词为26.48%),CER提高到3.45%,这表明我们的ASR驱动管道支持SEN OTEN语言文档的潜力。摘要:The SEN '{C}OTEN language, spoken on the Saanich peninsula of southern Vancouver Island, is in the midst of vigorous language revitalization efforts to turn the tide of language loss as a result of colonial language policies. To support these on-the-ground efforts, the community is turning to digital technology. Automatic Speech Recognition (ASR) technology holds great promise for accelerating language documentation and the creation of educational resources. However, developing ASR systems for SEN '{C}OTEN is challenging due to limited data and significant vocabulary variation from its polysynthetic structure and stress-driven metathesis. To address these challenges, we propose an ASR-driven documentation pipeline that leverages augmented speech data from a text-to-speech (TTS) system and cross-lingual transfer learning with Speech Foundation Models (SFMs). An n-gram language model is also incorporated via shallow fusion or n-best restoring to maximize the use of available data. Experiments on the SEN '{C}OTEN dataset show a word error rate (WER) of 19.34% and a character error rate (CER) of 5.09% on the test set with a 57.02% out-of-vocabulary (OOV) rate. After filtering minor cedilla-related errors, WER improves to 14.32% (26.48% on unseen words) and CER to 3.45%, demonstrating the potential of our ASR-driven pipeline to support SEN '{C}OTEN language documentation.
摘要:本文提出了一种新的基于规则的方法,通过改变现有的曲调生成音乐。我们解析每个曲调以找到Pathway Assembly(PA)[ 1],这是一个代表曲调中所有重复的结构。Sequitur算法[2 ]用于此。结果是一个语法。然后,我们对语法进行变异,而不是直接对曲调进行变异。可能有19种类型的变化,例如添加,删除,交换或反转可以应用于语法的语法部分。在这一步中,系统随机使用其中一个变化来自动操作语法。在突变之后,我们需要扩展返回新曲调的语法。一个或多个突变后的输出将是与原始曲调相关的新曲调。我们的研究探讨了曲调如何在多次突变过程中逐渐变化。编辑距离、结构复杂度和曲调的长度被用来显示曲调在多次突变后如何变化。此外,分析了每种突变类型的效应大小。作为最后一点,我们回顾了输出曲调的音乐方面。应该注意的是,该研究仅集中于生成新的音高序列。这项研究是基于爱尔兰传统的曲调数据集和整数列表已被用来代表每个曲调的音高值。
摘要:This paper presents a novel rule-based approach for generating music by varying existing tunes. We parse each tune to find the Pathway Assembly (PA) [ 1], that is a structure representing all repetitions in the tune. The Sequitur algorithm [2 ] is used for this. The result is a grammar. We then carry out mutation on the grammar, rather than on a tune directly. There are potentially 19 types of mutations such as adding, removing, swapping or reversing parts of the grammar that can be applied to the grammars. The system employs one of the mutations randomly in this step to automatically manipulate the grammar. Following the mutation, we need to expand the grammar which returns a new tune. The output after 1 or more mutations will be a new tune related to the original tune. Our study examines how tunes change gradually over the course of multiple mutations. Edit distances, structural complexity and length of the tunes are used to show how a tune is changed after multiple mutations. In addition, the size of effect of each mutation type is analyzed. As a final point, we review the musical aspect of the output tunes. It should be noted that the study only focused on generating new pitch sequences. The study is based on an Irish traditional tune dataset and a list of integers has been used to represent each tune's pitch values.
摘要:
在这项研究中,我们介绍了一个符号数据集组成的非公制伊朗古典音乐,和算法的结构分析这种音乐,并产生变化。该语料库包括来自Radif Mirza Abdollah的Dastgah Shour的数据文件和数据表,这是伊朗古典音乐的基础曲目。此外,我们应用我们先前介绍的算法来解析旋律结构(Kanani等人,2023b)到数据集。与许多西方音乐不同,这种非公制音乐不遵循以小节为中心的组织。我们的解析算法可以很好地捕捉非度量组织。我们将每个曲调(Gusheh)解析为语法,以识别主题和短语。这些语法表示可以用于教育和民族音乐学的目的。我们还进一步开发了一种先前介绍的创建旋律变奏的方法(Kanani等人,第2023段b)。在解析现有的曲调以产生语法之后,通过将突变应用于该语法,我们生成新的语法。扩展这个新版本会产生原曲调的变奏。变化由领域专家听众评估。此外,我们进行了统计分析的突变与不同的表示设置,我们的解析和生成算法。总体结论是,该系统成功地产生了可接受的变异后突变。虽然我们的案例研究侧重于伊朗古典音乐,但该方法可以适用于阿拉伯或土耳其古典音乐。
备注:to appear in IEEE WASPAA 2025
摘要:我们提出了一个用于近场声全息(NAH)中声源重建的迁移学习框架,该框架使用物理信息程序将经过良好训练的数据驱动模型从一种类型的声源调整到另一种类型。该框架包括两个阶段:(1)在大型数据集上对复值卷积神经网络(CV-CNN)进行监督预训练,以及(2)基于Kirchhoff-Helmholtz积分对单个数据样本进行纯粹的物理信息微调。该方法遵循迁移学习的原则,通过物理信息自适应实现跨不同数据集的泛化。通过将预训练模型从矩形板数据集转移到小提琴顶板数据集来验证该方法的有效性,与预训练模型相比,它显示出更高的重建精度,并提供与压缩等效源方法(C-ESM)相当的性能。此外,对于成功的模式,微调模型在准确性上优于预训练模型和C-ESM。摘要:We propose a transfer learning framework for sound source reconstruction in Near-field Acoustic Holography (NAH), which adapts a well-trained data-driven model from one type of sound source to another using a physics-informed procedure. The framework comprises two stages: (1) supervised pre-training of a complex-valued convolutional neural network (CV-CNN) on a large dataset, and (2) purely physics-informed fine-tuning on a single data sample based on the Kirchhoff-Helmholtz integral. This method follows the principles of transfer learning by enabling generalization across different datasets through physics-informed adaptation. The effectiveness of the approach is validated by transferring a pre-trained model from a rectangular plate dataset to a violin top plate dataset, where it shows improved reconstruction accuracy compared to the pre-trained model and delivers performance comparable to that of Compressive-Equivalent Source Method (C-ESM). Furthermore, for successful modes, the fine-tuned model outperforms both the pre-trained model and C-ESM in accuracy.
备注:Accepted for presentation at the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA 2025), 5 pages
摘要:传统的盲源分离评估(BSS-Eval)度量最初被设计用于评估基于时频掩蔽等方法的线性音频源分离模型。然而,最近的生成模型可能会引入分离信号和参考信号之间的非线性关系,限制了这些指标的客观评价的可靠性。为了解决这个问题,我们进行了降级类别评级听力测试,并分析所获得的降级平均意见分数(DMOS)和一组客观的音频质量指标的任务,歌声分离之间的相关性。我们评估了三个国家的最先进的判别模型和两个新的竞争生成模型。对于判别和生成模型,侵入嵌入为基础的指标显示出更高的相关性与DMOS比传统的侵入性指标,如BSS评估。对于判别模型,最高的相关性是通过在Music 2Latent嵌入上计算的MSE来实现的。在生成模型的评估方面,多分辨率STFT损失和MERT-L12嵌入计算的MSE的相关性最强,后者也提供了两种模型类型之间最平衡的相关性。我们的研究结果突出了BSS评估指标的局限性,用于评估生成的歌声分离模型,并强调需要仔细选择和验证的替代评价指标的任务,歌声分离。
摘要:Traditional Blind Source Separation Evaluation (BSS-Eval) metrics were originally designed to evaluate linear audio source separation models based on methods such as time-frequency masking. However, recent generative models may introduce nonlinear relationships between the separated and reference signals, limiting the reliability of these metrics for objective evaluation. To address this issue, we conduct a Degradation Category Rating listening test and analyze correlations between the obtained degradation mean opinion scores (DMOS) and a set of objective audio quality metrics for the task of singing voice separation. We evaluate three state-of-the-art discriminative models and two new competitive generative models. For both discriminative and generative models, intrusive embedding-based metrics show higher correlations with DMOS than conventional intrusive metrics such as BSS-Eval. For discriminative models, the highest correlation is achieved by the MSE computed on Music2Latent embeddings. When it comes to the evaluation of generative models, the strongest correlations are evident for the multi-resolution STFT loss and the MSE calculated on MERT-L12 embeddings, with the latter also providing the most balanced correlation across both model types. Our results highlight the limitations of BSS-Eval metrics for evaluating generative singing voice separation models and emphasize the need for careful selection and validation of alternative evaluation metrics for the task of singing voice separation.
备注:5 pages, 2 figures
摘要:在语音增强系统的语音质量估计中,主观听觉测试一直被认为是最好的标准。考虑到大量新的生成或混合方法涌入该领域,揭示了一些客观指标的问题,这应该更加真实。诸如Interspeech 2025 URGENT语音增强挑战赛等也涉及非英语数据集的努力为测试过程增加了多语言性方面。在本文中,我们简要回顾了ITU-T P.808众包主观听力测试方法。第一个新的贡献是我们提出的本地化的文本和音频组件的Naderi和卡特勒的实施众包的主观绝对类别评级(ACR)的听力测试,涉及文本到语音(TTS)的过程。此外,我们还对URGENT Challenge结果进行了令人惊讶的分析和见解,将ACR主观测试的可靠性作为生成式AI时代的黄金标准。特别是,似乎对于生成SE方法,主观(ACR MOS)和客观(DNSMOS,NISQA)无参考指标应该伴随着客观的电话保真度指标,以可靠地检测幻觉。最后,在已接受的版本中,我们将发布本地化脚本和方法,以便根据ITU-T P.808轻松部署新的多语言语音增强主观评估。
摘要:In speech quality estimation for speech enhancement (SE) systems, subjective listening tests so far are considered as the gold standard. This should be even more true considering the large influx of new generative or hybrid methods into the field, revealing issues of some objective metrics. Efforts such as the Interspeech 2025 URGENT Speech Enhancement Challenge also involving non-English datasets add the aspect of multilinguality to the testing procedure. In this paper, we provide a brief recap of the ITU-T P.808 crowdsourced subjective listening test method. A first novel contribution is our proposed process of localizing both text and audio components of Naderi and Cutler's implementation of crowdsourced subjective absolute category rating (ACR) listening tests involving text-to-speech (TTS). Further, we provide surprising analyses of and insights into URGENT Challenge results, tackling the reliability of (P.808) ACR subjective testing as gold standard in the age of generative AI. Particularly, it seems that for generative SE methods, subjective (ACR MOS) and objective (DNSMOS, NISQA) reference-free metrics should be accompanied by objective phone fidelity metrics to reliably detect hallucinations. Finally, in the accepted version, we will release our localization scripts and methods for easy deployment for new multilingual speech enhancement subjective evaluations according to ITU-T P.808.
摘要:这项工作介绍了一种新的方法,从任意麦克风阵列的双耳再现,基于阵列感知优化的高保真度立体声编码通过头部相关的传递函数(HRTF)预处理。所提出的方法将阵列特定的信息集成到HRTF处理管道中,从而提高了双耳渲染的空间精度。客观的评估表明,优越的性能下模拟可穿戴阵列和头部旋转相比,传统的高保真度立体声编码方法。听觉实验进一步证实,该方法实现了显着更高的音质和空间质量的感知评级。该方法与标准高保真度立体声完全兼容,为虚拟现实、增强现实和可穿戴音频捕获等应用中的空间音频渲染提供了一种实用的解决方案。
摘要:This work introduces a novel method for binaural reproduction from arbitrary microphone arrays, based on array-aware optimization of Ambisonics encoding through Head-Related Transfer Function (HRTF) pre-processing. The proposed approach integrates array-specific information into the HRTF processing pipeline, leading to improved spatial accuracy in binaural rendering. Objective evaluations demonstrate superior performance under simulated wearable-array and head rotations compared to conventional Ambisonics encoding method. A listening experiment further confirms that the method achieves significantly higher perceptual ratings in both timbre and spatial quality. Fully compatible with standard Ambisonics, the proposed method offers a practical solution for spatial audio rendering in applications such as virtual reality, augmented reality, and wearable audio capture.
备注:to appear in IEEE WASPAA 2025
摘要:我们提出了一个用于近场声全息(NAH)中声源重建的迁移学习框架,该框架使用物理信息程序将经过良好训练的数据驱动模型从一种类型的声源调整到另一种类型。该框架包括两个阶段:(1)在大型数据集上对复值卷积神经网络(CV-CNN)进行监督预训练,以及(2)基于Kirchhoff-Helmholtz积分对单个数据样本进行纯粹的物理信息微调。该方法遵循迁移学习的原则,通过物理信息自适应实现跨不同数据集的泛化。通过将预训练模型从矩形板数据集转移到小提琴顶板数据集来验证该方法的有效性,与预训练模型相比,它显示出更高的重建精度,并提供与压缩等效源方法(C-ESM)相当的性能。此外,对于成功的模式,微调模型在准确性上优于预训练模型和C-ESM。摘要:We propose a transfer learning framework for sound source reconstruction in Near-field Acoustic Holography (NAH), which adapts a well-trained data-driven model from one type of sound source to another using a physics-informed procedure. The framework comprises two stages: (1) supervised pre-training of a complex-valued convolutional neural network (CV-CNN) on a large dataset, and (2) purely physics-informed fine-tuning on a single data sample based on the Kirchhoff-Helmholtz integral. This method follows the principles of transfer learning by enabling generalization across different datasets through physics-informed adaptation. The effectiveness of the approach is validated by transferring a pre-trained model from a rectangular plate dataset to a violin top plate dataset, where it shows improved reconstruction accuracy compared to the pre-trained model and delivers performance comparable to that of Compressive-Equivalent Source Method (C-ESM). Furthermore, for successful modes, the fine-tuned model outperforms both the pre-trained model and C-ESM in accuracy.
备注:17 pages, 7 figures, 7 tables
摘要:动机心音描记术可以提供胎儿心率以及直接心音数据,并且完全是被动的,不使用任何类型的辐射。Approach.我们讨论了目前可用的胎儿心音检测和心率估计的方法,并使用一个共同的基准平台和预先选择的测试数据集进行比较。与以前的综述相比,我们以标准化的方式对所讨论的方法进行了评估,以便进行公平的比较。我们的测试包括基于公差的检测准确性,标签插入,删除和替换的错误率,以及心率均方误差的统计测量。结果根据我们的结果,没有明确的最佳方法可以在所有测试中获得最高分数,简单的方法可以执行更复杂的方法。第一心音检测的最佳模型达到97.6%F1评分、97.4%阳性预测值和12.2 ± 8.0 ms平均绝对误差。在第二心音检测方面,最佳模型具有91.4%的F1评分、91.3%的阳性预测值和17.3 ± 12.2 ms的平均绝对误差。对于胎心率,最佳方法的均方误差为0.644。意义我们的主要结论是,需要进一步规范胎儿心率和心音检测方法的评价。测试和算法实现可在https: github.com mulkr standard-fpcg-evaluation上公开获得。
摘要:Motivation. Phonocardiography can give access to the fetal heart rate as well as direct heart sound data, and is entirely passive, using no radiation of any kind. Approach. We discuss the currently available methods for fetal heart sound detection and heart rate estimation and compare them using a common benchmarking platform and a pre-selected testing dataset. Compared to previous reviews, we evaluated the discussed methods in a standardized manner for a fair comparison. Our tests included tolerance-based detection accuracy, error rates for label insertions, deletions, and substitutions, and statistical measures for heart rate mean square error. Results. Based on our results, there is no definite best method that can achieve the highest scores in all of the tests, and simpler methods could perform comparably to more complex ones. The best model for first heart sound detection achieved 97.6% F1-score, 97.4% positive predictive value, and 12.2+-8.0 ms mean absolute error. In terms of second heart sound detection the best model had 91.4% F1-score, 91.3% positive predictive value, and 17.3+-12.2 ms mean absolute error. For fetal heart rate a 0.644 mean square error was achieved by the best method. Significance. Our main conclusion is that further standardization is required in fetal heart rate and heart sound detection method evaluation. The tests and algorithm implementations are openly available at: https: github.com mulkr standard-fpcg-evaluation.
备注:Accepted to WASPAA 2025
摘要:时频(TF)双路径模型是目前性能最好的音频源分离网络架构之一,在语音增强、音乐源分离和电影音频源分离方面实现了最先进的性能。虽然它们的特点是参数计数相对较低,但它们仍然需要相当数量的操作,这意味着更长的执行时间。这个问题由于在大量数据上训练更大模型以解决更一般的任务的趋势而加剧,例如最近引入的任务感知统一源分离(TUSS)模型。TUSS旨在使用单个条件模型解决音频源分离任务,它是建立在TF-Locoformer之上的,TF-Locoformer是一种结合卷积和注意力层的TF双路径模型。任务定义以一系列提示的形式出现,这些提示指定要提取的源的数量和类型。在本文中,我们分析了TUSS的设计选择,其目标是优化其性能和复杂性的权衡。我们得到了两个更有效的模型,FasTUSS-8.3G和FasTUSS-11.7G,减少了81%和73%的原始模型的操作,与轻微的性能下降1.2 dB和0.4 dB的平均在所有的基准测试,分别。此外,我们调查的影响,提示条件反射,推导出一个因果TUSS模型。
摘要:Time-Frequency (TF) dual-path models are currently among the best performing audio source separation network architectures, achieving state-of-the-art performance in speech enhancement, music source separation, and cinematic audio source separation. While they are characterized by a relatively low parameter count, they still require a considerable number of operations, implying a higher execution time. This problem is exacerbated by the trend towards bigger models trained on large amounts of data to solve more general tasks, such as the recently introduced task-aware unified source separation (TUSS) model. TUSS, which aims to solve audio source separation tasks using a single, conditional model, is built upon TF-Locoformer, a TF dual-path model combining convolution and attention layers. The task definition comes in the form of a sequence of prompts that specify the number and type of sources to be extracted. In this paper, we analyze the design choices of TUSS with the goal of optimizing its performance-complexity trade-off. We derive two more efficient models, FasTUSS-8.3G and FasTUSS-11.7G that reduce the original model's operations by 81 % and 73 % with minor performance drops of 1.2~dB and 0.4~dB averaged over all benchmarks, respectively. Additionally, we investigate the impact of prompt conditioning to derive a causal TUSS model.
备注:
摘要:This paper presents a novel approach for detecting mispronunciations by analyzing deviations between a user's original speech and their voice-cloned counterpart with corrected pronunciation. We hypothesize that regions with maximal acoustic deviation between the original and cloned utterances indicate potential mispronunciations. Our method leverages recent advances in voice cloning to generate a synthetic version of the user's voice with proper pronunciation, then performs frame-by-frame comparisons to identify problematic segments. Experimental results demonstrate the effectiveness of this approach in pinpointing specific pronunciation errors without requiring predefined phonetic rules or extensive training data for each target language.
摘要:本文提出了一种新的基于规则的方法,通过改变现有的曲调生成音乐。我们解析每个曲调以找到Pathway Assembly(PA)[ 1],这是一个代表曲调中所有重复的结构。Sequitur算法[2 ]用于此。结果是一个语法。然后,我们对语法进行变异,而不是直接对曲调进行变异。可能有19种类型的变化,例如添加,删除,交换或反转可以应用于语法的语法部分。在这一步中,系统随机使用其中一个变化来自动操作语法。在突变之后,我们需要扩展返回新曲调的语法。一个或多个突变后的输出将是与原始曲调相关的新曲调。我们的研究探讨了曲调如何在多次突变过程中逐渐变化。编辑距离、结构复杂度和曲调的长度被用来显示曲调在多次突变后如何变化。此外,分析了每种突变类型的效应大小。作为最后一点,我们回顾了输出曲调的音乐方面。应该注意的是,该研究仅集中于生成新的音高序列。这项研究是基于爱尔兰传统的曲调数据集和整数列表已被用来代表每个曲调的音高值。
摘要:This paper presents a novel rule-based approach for generating music by varying existing tunes. We parse each tune to find the Pathway Assembly (PA) [ 1], that is a structure representing all repetitions in the tune. The Sequitur algorithm [2 ] is used for this. The result is a grammar. We then carry out mutation on the grammar, rather than on a tune directly. There are potentially 19 types of mutations such as adding, removing, swapping or reversing parts of the grammar that can be applied to the grammars. The system employs one of the mutations randomly in this step to automatically manipulate the grammar. Following the mutation, we need to expand the grammar which returns a new tune. The output after 1 or more mutations will be a new tune related to the original tune. Our study examines how tunes change gradually over the course of multiple mutations. Edit distances, structural complexity and length of the tunes are used to show how a tune is changed after multiple mutations. In addition, the size of effect of each mutation type is analyzed. As a final point, we review the musical aspect of the output tunes. It should be noted that the study only focused on generating new pitch sequences. The study is based on an Irish traditional tune dataset and a list of integers has been used to represent each tune's pitch values.
摘要:
在这项研究中,我们介绍了一个符号数据集组成的非公制伊朗古典音乐,和算法的结构分析这种音乐,并产生变化。该语料库包括来自Radif Mirza Abdollah的Dastgah Shour的数据文件和数据表,这是伊朗古典音乐的基础曲目。此外,我们应用我们先前介绍的算法来解析旋律结构(Kanani等人,2023b)到数据集。与许多西方音乐不同,这种非公制音乐不遵循以小节为中心的组织。我们的解析算法可以很好地捕捉非度量组织。我们将每个曲调(Gusheh)解析为语法,以识别主题和短语。这些语法表示可以用于教育和民族音乐学的目的。我们还进一步开发了一种先前介绍的创建旋律变奏的方法(Kanani等人,第2023段b)。在解析现有的曲调以产生语法之后,通过将突变应用于该语法,我们生成新的语法。扩展这个新版本会产生原曲调的变奏。变化由领域专家听众评估。此外,我们进行了统计分析的突变与不同的表示设置,我们的解析和生成算法。总体结论是,该系统成功地产生了可接受的变异后突变。虽然我们的案例研究侧重于伊朗古典音乐,但该方法可以适用于阿拉伯或土耳其古典音乐。
机器翻译由腾讯交互翻译提供,仅供参考
