今日论文合集:cs.SD语音8篇,eess.AS音频处理8篇。本文经arXiv每日学术速递授权转载
【1】Proactive Detection of Voice Cloning with Localized Watermarking
标题:基于局部水印的语音克隆主动检测
链接:https://arxiv.org/abs/2401.17264
作者:Robin San Roman,Pierre Fernandez,Alexandre Défossez,Teddy Furon,Tuan Tran,Hady Elsahar备注:Code at this https URL
摘要:在快速发展的语音生成模型领域,迫切需要确保音频的真实性,以防止语音克隆的风险。我们提出了AudioSeal,这是第一个专门为AI生成的语音的本地化检测而设计的音频水印技术。AudioSeal采用了一个生成器/检测器架构,与本地化损失一起训练,使本地化水印检测达到样本水平,以及一种由听觉掩蔽启发的新感知损失,使AudioSeal能够实现更好的不可感知性。AudioSeal在基于自动和人工评估指标的真实音频操作和不可感知性方面实现了最先进的性能。此外,AudioSeal还设计了一个快速的单通道检测器,在速度上大大超过现有型号-检测速度快了两个数量级,使其成为大规模和实时应用的理想选择。
摘要:In the rapidly evolving field of speech generative models, there is a pressing need to ensure audio authenticity against the risks of voice cloning. We present AudioSeal, the first audio watermarking technique designed specifically for localized detection of AI-generated speech. AudioSeal employs a generator/detector architecture trained jointly with a localization loss to enable localized watermark detection up to the sample level, and a novel perceptual loss inspired by auditory masking, that enables AudioSeal to achieve better imperceptibility. AudioSeal achieves state-of-the-art performance in terms of robustness to real life audio manipulations and imperceptibility based on automatic and human evaluation metrics. Additionally, AudioSeal is designed with a fast, single-pass detector, that significantly surpasses existing models in speed - achieving detection up to two orders of magnitude faster, making it ideal for large-scale and real-time applications.
【2】 ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models标题:ESPnet-SPK:全流水线扬声器嵌入工具包,具有可复制的配方、自我监督的前端和现成的模型作者:Jee-weon Jung,Wangyou Zhang,Jiatong Shi,Zakaria Aldeneh,Takuya Higuchi,Barry-John Theobald,Ahmed Hussen Abdelaziz,Shinji Watanabe备注:5 pages, 3 figures, 7 tables摘要:本文介绍了ESPnet-SPK,一个多目标的训练说话人嵌入提取器的工具包。首先,我们为说话人识别社区的研究人员提供了一个开源平台,可以轻松地构建模型。我们提供了几种模型,从x向量到最近的SKA-TDNN。通过模块化的架构设计,可以很容易地开发各种变体。我们还希望将开发的模型与其他领域联系起来,促进广泛的研究社区毫不费力地整合最先进的嵌入提取器。预训练的嵌入提取器可以以现成的方式访问,我们通过展示其与两个任务的集成来展示该工具包的多功能性。另一个目标是整合各种自监督学习功能。我们发布了一个可重复的配方,使用WavLM-Large和ECAPA-TDNN在Vox 1-O评估协议上实现了0.39%的相等错误率。摘要:This paper introduces ESPnet-SPK, a toolkit designed with several objectives for training speaker embedding extractors. First, we provide an open-source platform for researchers in the speaker recognition community to effortlessly build models. We provide several models, ranging from x-vector to recent SKA-TDNN. Through the modularized architecture design, variants can be developed easily. We also aspire to bridge developed models with other domains, facilitating the broad research community to effortlessly incorporate state-of-the-art embedding extractors. Pre-trained embedding extractors can be accessed in an off-the-shelf manner and we demonstrate the toolkit's versatility by showcasing its integration with two tasks. Another goal is to integrate with diverse self-supervised learning features. We release a reproducible recipe that achieves an equal error rate of 0.39% on the Vox1-O evaluation protocol using WavLM-Large with ECAPA-TDNN.【3】 A Proactive and Dual Prevention Mechanism against Illegal Song Covers empowered by Singing Voice Conversion作者:Guangke Chen,Yedi Zhang,Fu Song,Ting Wang,Xiaoning Du,Yang Liu摘要:歌声转换(SVC)通过将一个歌手的歌声转换为另一个目标歌手的歌声,并带有原始歌词和旋律,从而自动完成歌曲封面。然而,它引起了对多个实体侵犯版权和民事权利的严重关切。这项工作提出了SongBsAb,第一个主动的方法,以减轻未经授权的SVC为基础的非法歌曲封面。SongBsAb在释放歌声之前,将人类无法察觉的扰动引入歌声,这样当它们被使用时,SVC的生成过程就会受到干扰,从而产生意想不到的歌声。SongBsAb具有双重预防效果,即造成(歌手)身份破坏和歌词破坏,即SVC覆盖的演唱声音既不模仿目标歌手,也不保留原始歌词。为了提高扰动的不可感知性,我们改进了一个基于心理声学模型的损失与后盾轨道作为一个额外的掩蔽,一个独特的伴随元素相比,普通的语音唱歌的声音。为了提高可转移性,我们建议利用帧级的相互作用减少为基础的损失。我们使用客观和基于人类研究的主观指标,在三个SVC模型和两个数据集上展示了SongBsAb的预防有效性、实用性和鲁棒性。我们的工作促进了一个新兴的研究方向,以减轻非法自动歌曲封面。摘要:Singing voice conversion (SVC) automates song covers by converting one singer's singing voice into another target singer's singing voice with the original lyrics and melody. However, it raises serious concerns about copyright and civil right infringements to multiple entities. This work proposes SongBsAb, the first proactive approach to mitigate unauthorized SVC-based illegal song covers. SongBsAb introduces human-imperceptible perturbations to singing voices before releasing them, so that when they are used, the generation process of SVC will be interfered, resulting in unexpected singing voices. SongBsAb features a dual prevention effect by causing both (singer) identity disruption and lyric disruption, namely, the SVC-covered singing voice neither imitates the target singer nor preserves the original lyrics. To improve the imperceptibility of perturbations, we refine a psychoacoustic model-based loss with the backing track as an additional masker, a unique accompanying element for singing voices compared to ordinary speech voices. To enhance the transferability, we propose to utilize a frame-level interaction reduction-based loss. We demonstrate the prevention effectiveness, utility, and robustness of SongBsAb on three SVC models and two datasets using both objective and human study-based subjective metrics. Our work fosters an emerging research direction for mitigating illegal automated song covers.
【4】 Enhanced Sound Event Localization and Detection in Real 360-degree audio-visual soundscapes标题:在真实360度视听音景中增强的声音事件定位和检测作者:Adrian S. Roman,Baladithya Balamurugan,Rithik Pothuganti摘要:本技术报告详细介绍了我们在构建增强型视听声音事件定位和检测(SELD)网络方面的工作。我们建立在仅音频SELDnet23模型的基础上,并通过在仅音频网络的门控递归单元(GRU)之前合并音频和视频信息来将其调整为视听。我们的模型利用了YOLO和DETIC对象检测器。我们还建立了一个框架,实现视听数据增强和视听合成数据生成。我们提供的视听SELDnet系统性能优于现有的视听SELD基线。摘要:This technical report details our work towards building an enhanced audio-visual sound event localization and detection (SELD) network. We build on top of the audio-only SELDnet23 model and adapt it to be audio-visual by merging both audio and video information prior to the gated recurrent unit (GRU) of the audio-only network. Our model leverages YOLO and DETIC object detectors. We also build a framework that implements audio-visual data augmentation and audio-visual synthetic data generation. We deliver an audio-visual SELDnet system that outperforms the existing audio-visual SELD baseline.【5】 SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics标题:SpeechBERTScore:利用NLP评估度量的参考感知语音生成自动评估作者:Takaaki Saeki,Soumi Maiti,Shinnosuke Takamichi,Shinji Watanabe,Hiroshi Saruwatari摘要:虽然主观评估一直是评估语音生成的黄金标准,但由于其成本效率,越来越需要与人类主观判断高度相关的客观度量。本文提出了参考感知的自动评估方法的语音生成的自然语言处理中的评价指标的启发。所提出的SpeechBERTScore计算BERTScore用于生成的和参考语音的自监督密集语音特征,其可以具有不同的序列长度。我们还提出了SpeechBLEU和SpeechTokenDistance,它们是在语音离散令牌上计算的。合成语音的评价表明,我们的方法更好地与人类的主观评级比梅尔倒谱失真和最近的平均意见评分预测模型。此外,他们是有效的噪声语音评价,并具有跨语言的适用性。摘要:While subjective assessments have been the gold standard for evaluating speech generation, there is a growing need for objective metrics that are highly correlated with human subjective judgments due to their cost efficiency. This paper proposes reference-aware automatic evaluation methods for speech generation inspired by evaluation metrics in natural language processing. The proposed SpeechBERTScore computes the BERTScore for self-supervised dense speech features of the generated and reference speech, which can have different sequential lengths. We also propose SpeechBLEU and SpeechTokenDistance, which are computed on speech discrete tokens. The evaluations on synthesized speech show that our method correlates better with human subjective ratings than mel cepstral distortion and a recent mean opinion score prediction model. Also, they are effective in noisy speech evaluation and have cross-lingual applicability.
【6】 PBSCSR: The Piano Bootleg Score Composer Style Recognition Dataset标题:PBSCSR:钢琴盗版乐谱作曲家风格识别数据集作者:Arhan Jain,Alec Bunn,TJ Tsai摘要:本文动机,描述,并提出了PBSCSR数据集研究作曲家风格识别的钢琴乐谱。我们的首要目标是创建一个用于研究作曲家风格识别的数据集,它“与MNIST一样容易访问,与ImageNet一样具有挑战性。“为了实现这一目标,我们从IMSLP上的钢琴乐谱图像中抽取了固定长度的盗版乐谱片段。数据集本身包含40,000张用于9向分类任务的62x64盗版分数图像,100,000张用于100向分类任务的62x64盗版分数图像,以及29,310张用于预训练的未标记可变长度盗版分数图像。标记的数据以反映MNIST图像的形式呈现,以便以有效的方式非常容易地可视化、操作和训练模型。此外,我们还包括相关的元数据,以允许访问IMSLP上的原始乐谱图像和其他相关数据。我们描述了几个研究任务,可以研究的数据集,包括作曲家风格识别的变化,在一个Few-Shot或zero-shot设置。对于之前已经提出模型的任务,我们发布代码和基线结果,以便将来的工作进行比较。我们还讨论了开放的研究问题,PBSCSR数据是特别适合促进研究和富有成效的探索领域在未来的工作。摘要:This article motivates, describes, and presents the PBSCSR dataset for studying composer style recognition of piano sheet music. Our overarching goal was to create a dataset for studying composer style recognition that is "as accessible as MNIST and as challenging as ImageNet." To achieve this goal, we sample fixed-length bootleg score fragments from piano sheet music images on IMSLP. The dataset itself contains 40,000 62x64 bootleg score images for a 9-way classification task, 100,000 62x64 bootleg score images for a 100-way classification task, and 29,310 unlabeled variable-length bootleg score images for pretraining. The labeled data is presented in a form that mirrors MNIST images, in order to make it extremely easy to visualize, manipulate, and train models in an efficient manner. Additionally, we include relevant metadata to allow access to the underlying raw sheet music images and other related data on IMSLP. We describe several research tasks that could be studied with the dataset, including variations of composer style recognition in a few-shot or zero-shot setting. For tasks that have previously proposed models, we release code and baseline results for future works to compare against. We also discuss open research questions that the PBSCSR data is especially well suited to facilitate research on and areas of fruitful exploration in future work.【7】 Spatial-Temporal Activity-Informed Diarization and Separation作者:Yicheng Hsu,Ssuhan Chen,Mingsian R. Bai摘要:利用说话人的时空活动性,提出了一种鲁棒的多通道说话人日志化和分离系统。该系统以混合架构实现,该架构结合了阵列信号处理单元和深度学习单元。对于扬声器日志化,基于麦克风阵列的白化相对传递函数(wRTF)计算跨时间帧的空间相干矩阵。这用作后续机器学习的鲁棒特征,而不需要阵列配置的先验知识。一个计算效率高的空间活动驱动的说话人日记网络(SASDnet)被构造成直接从空间相干矩阵估计说话人活动。对于说话人分离,我们提出了全球和本地活动驱动的说话人提取网络(GLASEnet)分离说话人信号通过特定的全球和本地空间活动功能。局部空间活动函数取决于每个时频区间的wRTF与目标说话者主导区间之间的相干性。基于频率平均的局部空间活动函数,从全局空间相干函数计算全局空间活动函数。实验结果表明,优越的扬声器,日记,计数,和分离性能所提出的系统相比,预先选定的基线具有较低的计算复杂度。摘要:A robust multichannel speaker diarization and separation system is proposed by exploiting the spatio-temporal activity of the speakers. The system is realized in a hybrid architecture that combines the array signal processing units and the deep learning units. For speaker diarization, a spatial coherence matrix across time frames is computed based on the whitened relative transfer functions (wRTFs) of the microphone array. This serves as a robust feature for subsequent machine learning without the need for prior knowledge of the array configuration. A computationally efficient Spatial Activity-driven Speaker Diarization network (SASDnet) is constructed to estimate the speaker activity directly from the spatial coherence matrix. For speaker separation, we propose the Global and Local Activity-driven Speaker Extraction network (GLASEnet) to separate speaker signals via speaker-specific global and local spatial activity functions. The local spatial activity functions depend on the coherence between the wRTFs of each time-frequency bin and the target speaker-dominant bins. The global spatial activity functions are computed from the global spatial coherence functions based on frequency-averaged local spatial activity functions. Experimental results have demonstrated superior speaker, diarization, counting, and separation performance achieved by the proposed system with low computational complexity compared to the pre-selected baselines.【8】 Localizing uniformly moving mono-frequent sources using an inverse 2.5D approach作者:Christian H. Kasess,Wolfgang Kreuzer,Prateek Soni,Holger Waubke摘要:使用麦克风阵列定位线性移动的声源特别具有挑战性,因为信号的瞬态性质导致相对较短的观察周期。通常,使用移动焦点,并且大多数方法至少部分地在时域中操作。与此相反,这里的逆源定位算法的单频均匀移动源,完全在频域中的行为。为此,利用2.5D方法,并导出源和麦克风网格之间的传递函数。通过使用麦克风网格处的数据求解最小二乘问题,可以确定运动框架中的未知源分布。为此,需要使用加窗离散傅里叶变换(DFT)将测量的时间信号变换到频域中,这导致诸如取决于时间间隔的长度和所使用的分析窗口的频谱泄漏的影响。为了将这些效应包括在数值模型中,使用分析窗口的傅里叶变换来修改传递矩阵的计算。目前,这种方法仅限于单频源,因为这允许简化计算并减少计算工作量。最小二乘问题的解决使用吉洪诺夫正则化采用L曲线的方法来确定一个合适的正则化参数。作为一个移动的源被认为是,多普勒效应允许通过组合的多个频率在测量信号的传递函数,以提高系统的稳定性。利用有无地面反射的移动点源的模拟数据对该方法的性能进行了验证。数值实验表明,在接收器频谱中的频率的选择,DFT的效果,源的频率,和源和接收器的距离的效果。摘要:Localizing linearly moving sound sources using microphone arrays is particularly challenging as the transient nature of the signal leads to relatively short observation periods. Commonly, a moving focus is used and most methods operate at least partially in the time domain. In contrast, here an inverse source localization algorithm for mono-frequent uniformly moving sources that acts entirely in the frequency domain is presented. For this, a 2.5D approach is utilized and a transfer function between sources and a microphone grid is derived. By solving a least squares problem using the data at the microphone grid, the unknown source distribution in the moving frame can be determined. For that the measured time signals need to be transformed into the frequency domain using a windowed discrete Fourier transform (DFT), which leads to effects such as spectral leakage that depends on the length of the time interval and the analysis window used. To include these effects in the numerical model, the calculation of the transfer matrix is modified using the Fourier transform of the analysis window. Currently, this approach is limited to mono-frequent sources as this allows a simplification of the calculation and reduces the computational effort. The least squares problem is solved using a Tikhonov regularization employing an L-curve approach to determine a suitable regularization parameter. As a moving source is considered, the Doppler effect allows to enhance the stability of the system by combining the transfer functions for multiple frequencies in the measured signals. The performance of the approach is validated using simulated data of a moving point source with or without a reflecting ground. Numerical experiments are performed to show the effect of the choice of frequencies in the receiver spectrum, the effect of the DFT, the frequency of the source, and the distance of source and receiver.【1】 Spatial-Temporal Activity-Informed Diarization and Separation作者:Yicheng Hsu,Ssuhan Chen,Mingsian R. Bai摘要:利用说话人的时空活动性,提出了一种鲁棒的多通道说话人日志化和分离系统。该系统以混合架构实现,该架构结合了阵列信号处理单元和深度学习单元。对于扬声器日志化,基于麦克风阵列的白化相对传递函数(wRTF)计算跨时间帧的空间相干矩阵。这用作后续机器学习的鲁棒特征,而不需要阵列配置的先验知识。一个计算效率高的空间活动驱动的说话人日记网络(SASDnet)被构造成直接从空间相干矩阵估计说话人活动。对于说话人分离,我们提出了全球和本地活动驱动的说话人提取网络(GLASEnet)分离说话人信号通过特定的全球和本地空间活动功能。局部空间活动函数取决于每个时频区间的wRTF与目标说话者主导区间之间的相干性。基于频率平均的局部空间活动函数,从全局空间相干函数计算全局空间活动函数。实验结果表明,优越的扬声器,日记,计数,和分离性能所提出的系统相比,预先选定的基线具有较低的计算复杂度。摘要:A robust multichannel speaker diarization and separation system is proposed by exploiting the spatio-temporal activity of the speakers. The system is realized in a hybrid architecture that combines the array signal processing units and the deep learning units. For speaker diarization, a spatial coherence matrix across time frames is computed based on the whitened relative transfer functions (wRTFs) of the microphone array. This serves as a robust feature for subsequent machine learning without the need for prior knowledge of the array configuration. A computationally efficient Spatial Activity-driven Speaker Diarization network (SASDnet) is constructed to estimate the speaker activity directly from the spatial coherence matrix. For speaker separation, we propose the Global and Local Activity-driven Speaker Extraction network (GLASEnet) to separate speaker signals via speaker-specific global and local spatial activity functions. The local spatial activity functions depend on the coherence between the wRTFs of each time-frequency bin and the target speaker-dominant bins. The global spatial activity functions are computed from the global spatial coherence functions based on frequency-averaged local spatial activity functions. Experimental results have demonstrated superior speaker, diarization, counting, and separation performance achieved by the proposed system with low computational complexity compared to the pre-selected baselines.
【2】 Localizing uniformly moving mono-frequent sources using an inverse 2.5D approach作者:Christian H. Kasess,Wolfgang Kreuzer,Prateek Soni,Holger Waubke摘要:使用麦克风阵列定位线性移动的声源特别具有挑战性,因为信号的瞬态性质导致相对较短的观察周期。通常,使用移动焦点,并且大多数方法至少部分地在时域中操作。与此相反,这里的逆源定位算法的单频均匀移动源,完全在频域中的行为。为此,利用2.5D方法,并导出源和麦克风网格之间的传递函数。通过使用麦克风网格处的数据求解最小二乘问题,可以确定运动框架中的未知源分布。为此,需要使用加窗离散傅里叶变换(DFT)将测量的时间信号变换到频域中,这导致诸如取决于时间间隔的长度和所使用的分析窗口的频谱泄漏的影响。为了将这些效应包括在数值模型中,使用分析窗口的傅里叶变换来修改传递矩阵的计算。目前,这种方法仅限于单频源,因为这允许简化计算并减少计算工作量。最小二乘问题的解决使用吉洪诺夫正则化采用L曲线的方法来确定一个合适的正则化参数。作为一个移动的源被认为是,多普勒效应允许通过组合的多个频率在测量信号的传递函数,以提高系统的稳定性。利用有无地面反射的移动点源的模拟数据对该方法的性能进行了验证。数值实验表明,在接收器频谱中的频率的选择,DFT的效果,源的频率,和源和接收器的距离的效果。摘要:Localizing linearly moving sound sources using microphone arrays is particularly challenging as the transient nature of the signal leads to relatively short observation periods. Commonly, a moving focus is used and most methods operate at least partially in the time domain. In contrast, here an inverse source localization algorithm for mono-frequent uniformly moving sources that acts entirely in the frequency domain is presented. For this, a 2.5D approach is utilized and a transfer function between sources and a microphone grid is derived. By solving a least squares problem using the data at the microphone grid, the unknown source distribution in the moving frame can be determined. For that the measured time signals need to be transformed into the frequency domain using a windowed discrete Fourier transform (DFT), which leads to effects such as spectral leakage that depends on the length of the time interval and the analysis window used. To include these effects in the numerical model, the calculation of the transfer matrix is modified using the Fourier transform of the analysis window. Currently, this approach is limited to mono-frequent sources as this allows a simplification of the calculation and reduces the computational effort. The least squares problem is solved using a Tikhonov regularization employing an L-curve approach to determine a suitable regularization parameter. As a moving source is considered, the Doppler effect allows to enhance the stability of the system by combining the transfer functions for multiple frequencies in the measured signals. The performance of the approach is validated using simulated data of a moving point source with or without a reflecting ground. Numerical experiments are performed to show the effect of the choice of frequencies in the receiver spectrum, the effect of the DFT, the frequency of the source, and the distance of source and receiver.
【3】 ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models标题:ESPnet-SPK:全流水线扬声器嵌入工具包,具有可复制的配方、自我监督的前端和现成的模型作者:Jee-weon Jung,Wangyou Zhang,Jiatong Shi,Zakaria Aldeneh,Takuya Higuchi,Barry-John Theobald,Ahmed Hussen Abdelaziz,Shinji Watanabe备注:5 pages, 3 figures, 7 tables摘要:本文介绍了ESPnet-SPK,一个多目标的训练说话人嵌入提取器的工具包。首先,我们为说话人识别社区的研究人员提供了一个开源平台,可以轻松地构建模型。我们提供了几种模型,从x向量到最近的SKA-TDNN。通过模块化的架构设计,可以很容易地开发各种变体。我们还希望将开发的模型与其他领域联系起来,促进广泛的研究社区毫不费力地整合最先进的嵌入提取器。预训练的嵌入提取器可以以现成的方式访问,我们通过展示其与两个任务的集成来展示该工具包的多功能性。另一个目标是整合各种自监督学习功能。我们发布了一个可重复的配方,使用WavLM-Large和ECAPA-TDNN在Vox 1-O评估协议上实现了0.39%的相等错误率。摘要:This paper introduces ESPnet-SPK, a toolkit designed with several objectives for training speaker embedding extractors. First, we provide an open-source platform for researchers in the speaker recognition community to effortlessly build models. We provide several models, ranging from x-vector to recent SKA-TDNN. Through the modularized architecture design, variants can be developed easily. We also aspire to bridge developed models with other domains, facilitating the broad research community to effortlessly incorporate state-of-the-art embedding extractors. Pre-trained embedding extractors can be accessed in an off-the-shelf manner and we demonstrate the toolkit's versatility by showcasing its integration with two tasks. Another goal is to integrate with diverse self-supervised learning features. We release a reproducible recipe that achieves an equal error rate of 0.39% on the Vox1-O evaluation protocol using WavLM-Large with ECAPA-TDNN.
【4】 A Proactive and Dual Prevention Mechanism against Illegal Song Covers empowered by Singing Voice Conversion作者:Guangke Chen,Yedi Zhang,Fu Song,Ting Wang,Xiaoning Du,Yang Liu摘要:歌声转换(SVC)通过将一个歌手的歌声转换为另一个目标歌手的歌声,并带有原始歌词和旋律,从而自动完成歌曲封面。然而,它引起了对多个实体侵犯版权和民事权利的严重关切。这项工作提出了SongBsAb,第一个主动的方法,以减轻未经授权的SVC为基础的非法歌曲封面。SongBsAb在释放歌声之前,将人类无法察觉的扰动引入歌声,这样当它们被使用时,SVC的生成过程就会受到干扰,从而产生意想不到的歌声。SongBsAb具有双重预防效果,即造成(歌手)身份破坏和歌词破坏,即SVC覆盖的演唱声音既不模仿目标歌手,也不保留原始歌词。为了提高扰动的不可感知性,我们改进了一个基于心理声学模型的损失与后盾轨道作为一个额外的掩蔽,一个独特的伴随元素相比,普通的语音唱歌的声音。为了提高可转移性,我们建议利用帧级的相互作用减少为基础的损失。我们使用客观和基于人类研究的主观指标,在三个SVC模型和两个数据集上展示了SongBsAb的预防有效性、实用性和鲁棒性。我们的工作促进了一个新兴的研究方向,以减轻非法自动歌曲封面。摘要:Singing voice conversion (SVC) automates song covers by converting one singer's singing voice into another target singer's singing voice with the original lyrics and melody. However, it raises serious concerns about copyright and civil right infringements to multiple entities. This work proposes SongBsAb, the first proactive approach to mitigate unauthorized SVC-based illegal song covers. SongBsAb introduces human-imperceptible perturbations to singing voices before releasing them, so that when they are used, the generation process of SVC will be interfered, resulting in unexpected singing voices. SongBsAb features a dual prevention effect by causing both (singer) identity disruption and lyric disruption, namely, the SVC-covered singing voice neither imitates the target singer nor preserves the original lyrics. To improve the imperceptibility of perturbations, we refine a psychoacoustic model-based loss with the backing track as an additional masker, a unique accompanying element for singing voices compared to ordinary speech voices. To enhance the transferability, we propose to utilize a frame-level interaction reduction-based loss. We demonstrate the prevention effectiveness, utility, and robustness of SongBsAb on three SVC models and two datasets using both objective and human study-based subjective metrics. Our work fosters an emerging research direction for mitigating illegal automated song covers.
【5】 Enhanced Sound Event Localization and Detection in Real 360-degree audio-visual soundscapes标题:在真实360度视听声景中增强的声音事件定位和检测作者:Adrian S. Roman,Baladithya Balamurugan,Rithik Pothuganti摘要:本技术报告详细介绍了我们在构建增强型视听声音事件定位和检测(SELD)网络方面的工作。我们建立在仅音频SELDnet23模型的基础上,并通过在仅音频网络的门控递归单元(GRU)之前合并音频和视频信息来将其调整为视听。我们的模型利用了YOLO和DETIC对象检测器。我们还建立了一个框架,实现视听数据增强和视听合成数据生成。我们提供的视听SELDnet系统性能优于现有的视听SELD基线。摘要:This technical report details our work towards building an enhanced audio-visual sound event localization and detection (SELD) network. We build on top of the audio-only SELDnet23 model and adapt it to be audio-visual by merging both audio and video information prior to the gated recurrent unit (GRU) of the audio-only network. Our model leverages YOLO and DETIC object detectors. We also build a framework that implements audio-visual data augmentation and audio-visual synthetic data generation. We deliver an audio-visual SELDnet system that outperforms the existing audio-visual SELD baseline.
【6】 SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics标题:SpeechBERTScore:利用NLP评估度量的参考感知语音生成自动评估作者:Takaaki Saeki,Soumi Maiti,Shinnosuke Takamichi,Shinji Watanabe,Hiroshi Saruwatari摘要:虽然主观评估一直是评估语音生成的黄金标准,但由于其成本效率,越来越需要与人类主观判断高度相关的客观度量。本文提出了参考感知的自动评估方法的语音生成的自然语言处理中的评价指标的启发。所提出的SpeechBERTScore计算BERTScore用于生成的和参考语音的自监督密集语音特征,其可以具有不同的序列长度。我们还提出了SpeechBLEU和SpeechTokenDistance,它们是在语音离散令牌上计算的。合成语音的评价表明,我们的方法更好地与人类的主观评级比梅尔倒谱失真和最近的平均意见评分预测模型。此外,他们是有效的噪声语音评价,并具有跨语言的适用性。摘要:While subjective assessments have been the gold standard for evaluating speech generation, there is a growing need for objective metrics that are highly correlated with human subjective judgments due to their cost efficiency. This paper proposes reference-aware automatic evaluation methods for speech generation inspired by evaluation metrics in natural language processing. The proposed SpeechBERTScore computes the BERTScore for self-supervised dense speech features of the generated and reference speech, which can have different sequential lengths. We also propose SpeechBLEU and SpeechTokenDistance, which are computed on speech discrete tokens. The evaluations on synthesized speech show that our method correlates better with human subjective ratings than mel cepstral distortion and a recent mean opinion score prediction model. Also, they are effective in noisy speech evaluation and have cross-lingual applicability.【7】 PBSCSR: The Piano Bootleg Score Composer Style Recognition Dataset标题:PBSCSR:钢琴盗版乐谱作曲家风格识别数据集作者:Arhan Jain,Alec Bunn,TJ Tsai摘要:本文动机,描述,并提出了PBSCSR数据集研究作曲家风格识别的钢琴乐谱。我们的首要目标是创建一个用于研究作曲家风格识别的数据集,它“与MNIST一样容易访问,与ImageNet一样具有挑战性。“为了实现这一目标,我们从IMSLP上的钢琴乐谱图像中抽取了固定长度的盗版乐谱片段。数据集本身包含40,000张用于9向分类任务的62x64盗版分数图像,100,000张用于100向分类任务的62x64盗版分数图像,以及29,310张用于预训练的未标记可变长度盗版分数图像。标记的数据以反映MNIST图像的形式呈现,以便以有效的方式非常容易地可视化、操作和训练模型。此外,我们还包括相关的元数据,以允许访问IMSLP上的原始乐谱图像和其他相关数据。我们描述了几个研究任务,可以研究的数据集,包括作曲家风格识别的变化,在一个Few-Shot或zero-shot设置。对于之前已经提出模型的任务,我们发布代码和基线结果,以便将来的工作进行比较。我们还讨论了开放的研究问题,PBSCSR数据是特别适合促进研究和富有成效的探索领域在未来的工作。摘要:This article motivates, describes, and presents the PBSCSR dataset for studying composer style recognition of piano sheet music. Our overarching goal was to create a dataset for studying composer style recognition that is "as accessible as MNIST and as challenging as ImageNet." To achieve this goal, we sample fixed-length bootleg score fragments from piano sheet music images on IMSLP. The dataset itself contains 40,000 62x64 bootleg score images for a 9-way classification task, 100,000 62x64 bootleg score images for a 100-way classification task, and 29,310 unlabeled variable-length bootleg score images for pretraining. The labeled data is presented in a form that mirrors MNIST images, in order to make it extremely easy to visualize, manipulate, and train models in an efficient manner. Additionally, we include relevant metadata to allow access to the underlying raw sheet music images and other related data on IMSLP. We describe several research tasks that could be studied with the dataset, including variations of composer style recognition in a few-shot or zero-shot setting. For tasks that have previously proposed models, we release code and baseline results for future works to compare against. We also discuss open research questions that the PBSCSR data is especially well suited to facilitate research on and areas of fruitful exploration in future work.【8】 OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer标题:OWSM v3.1:基于E-Branchformer的更好更快的开放语音式语音模型作者:Yifan Peng,Jinchuan Tian,William Chen,Siddhant Arora,Brian Yan,Yui Sudo,Muhammad Shakeel,Kwanghee Choi,Jiatong Shi,Xuankai Chang,Jee-weon Jung,Shinji Watanabe备注:Project webpage: this https URL摘要:最近的研究提倡完全开放的基金会模式,以促进透明度和开放科学。作为第一步,开放耳语风格语音模型(OWSM)使用公开可用的数据和开源工具包复制了OpenAI的耳语。为了再现Whisper,以前的OWSM v1到v3模型仍然基于Transformer,这可能导致与其他最先进的语音编码器相比性能较差。在这项工作中,我们的目标是提高性能和效率的OWSM没有额外的训练数据。我们提出了基于E-Branchformer的OWSM v3.1模型在两个尺度,即,100 M和1B。1B模型是已经公开的最大的基于E-Branchformer的语音模型。它在绝大多数评估基准测试中的表现优于之前的OWSM v3,同时推理速度提高了25%。我们公开发布数据准备脚本、预训练模型和训练日志。摘要:Recent studies have advocated for fully open foundation models to promote transparency and open science. As an initial step, the Open Whisper-style Speech Model (OWSM) reproduced OpenAI's Whisper using publicly available data and open-source toolkits. With the aim of reproducing Whisper, the previous OWSM v1 through v3 models were still based on Transformer, which might lead to inferior performance compared to other state-of-the-art speech encoders. In this work, we aim to improve the performance and efficiency of OWSM without extra training data. We present E-Branchformer based OWSM v3.1 models at two scales, i.e., 100M and 1B. The 1B model is the largest E-Branchformer based speech model that has been made publicly available. It outperforms the previous OWSM v3 in a vast majority of evaluation benchmarks, while demonstrating up to 25% faster inference speed. We publicly release the data preparation scripts, pre-trained models and training logs.