今日论文合集:cs.SD语音9篇,eess.AS音频处理8篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Device Feature based on Graph Fourier Transformation with Logarithmic  Processing For Detection of Replay Speech Attacks
标题:基于图形傅里叶变换和算术处理的设备特征检测重播语音攻击
链接:https://arxiv.org/abs/2404.17280
作者:Mingrui He,Longting Xu,Han Wang,Mingjun Zhang,Rohan Kumar Das
摘要:对自动说话人确认系统最常见的欺骗攻击是重放语音攻击。重放语音的检测严重依赖于重放配置信息。先前的研究表明,图形傅立叶变换衍生的功能可以有效地检测重放语音,但忽略设备和环境噪声的影响。在这项工作中,我们提出了一个新的功能,图形频率设备倒谱系数,来自图形频域使用设备相关的线性变换。我们还介绍了两种新的表示:图频率对数系数和图频率对数设备系数。我们使用传统的高斯混合模型和光卷积神经网络系统作为分类器来评估我们的方法。在ASVspoof 2017 V2,ASVspoof 2019物理访问和ASVspoof 2021物理访问数据集上,我们提出的功能优于已知的前端,证明了它们对重放语音检测的有效性。
摘要:The most common spoofing attacks on automatic speaker verification systems are replay speech attacks. Detection of replay speech heavily relies on replay configuration information. Previous studies have shown that graph Fourier transform-derived features can effectively detect replay speech but ignore device and environmental noise effects. In this work, we propose a new feature, the graph frequency device cepstral coefficient, derived from the graph frequency domain using a device-related linear transformation. We also introduce two novel representations: graph frequency logarithmic coefficient and graph frequency logarithmic device coefficient. We evaluate our methods using traditional Gaussian mixture model and light convolutional neural network systems as classifiers. On the ASVspoof 2017 V2, ASVspoof 2019 physical access, and ASVspoof 2021 physical access datasets, our proposed features outperform known front-ends, demonstrating their effectiveness for replay speech detection.


【2】 Comparison of self-supervised in-domain and supervised out-domain  transfer learning for bird species recognition
标题:自监督域内和监督域外迁移学习在鸟类识别中的比较
链接:https://arxiv.org/abs/2404.17252
作者:Houtan Ghaffari,Paul Devos
摘要:将预先训练好的模型的权重转移到另一个任务中,已经成为现代深度学习的重要组成部分,特别是在数据稀缺的情况下。预训练是指在当前感兴趣的任务之外训练模型的初始步骤,通常是在另一个数据集上。它可以通过使用人类注释数据集的监督模型或在未标记数据集上训练的自监督模型来完成。在这两种情况下,许多预先训练的模型可用于微调感兴趣的任务。有趣的是,研究表明,来自ImageNet的预训练模型可以帮助音频任务,尽管是在图像数据集上训练的。因此,目前还不清楚域内模型与合格的域外模型(例如ImageNet的卷积神经网络)相比是否具有优势。我们的实验将通过利用VICReg(一种最近的强大的自监督方法)来证明域内模型和数据集对于鸟类物种识别的有用性。
摘要:Transferring the weights of a pre-trained model to assist another task has become a crucial part of modern deep learning, particularly in data-scarce scenarios. Pre-training refers to the initial step of training models outside the current task of interest, typically on another dataset. It can be done via supervised models using human-annotated datasets or self-supervised models trained on unlabeled datasets. In both cases, many pre-trained models are available to fine-tune for the task of interest. Interestingly, research has shown that pre-trained models from ImageNet can be helpful for audio tasks despite being trained on image datasets. Hence, it's unclear whether in-domain models would be advantageous compared to competent out-domain models, such as convolutional neural networks from ImageNet. Our experiments will demonstrate the usefulness of in-domain models and datasets for bird species recognition by leveraging VICReg, a recent and powerful self-supervised method.

【3】 An Investigation of Time-Frequency Representation Discriminators for  High-Fidelity Vocoder
标题:高保真声码器时频表示鉴别器的研究
链接:https://arxiv.org/abs/2404.17161
作者:Yicheng Gu,Xueyao Zhang,Liumeng Xue,Haizhou Li,Zhizheng Wu
备注:arXiv admin note: text overlap with arXiv:2311.14957
摘要:当从声学表示重建可听波形时,基于生成对抗网络(GAN)的声码器在推理速度和合成质量方面都是优越的。本研究的重点是改善基于GAN的声码器的可扩展性。大多数现有的基于时频表示(TFR)的鉴别器植根于短时傅里叶变换(STFT),其具有恒定的时频(TF)分辨率,线性缩放的中心频率和固定的分解基础,使得其与需要动态注意不同频带和不同时间间隔的歌声等信号不兼容。基于此,本文提出了多尺度子带恒Q变换CQT(MS—SB—CQT)和多尺度时间压缩连续小波变换CWT(MS—TC—CWT)。CQT和CWT都具有针对不同频带的动态TF分辨率。相比之下,CQT对基音信息的建模能力更强,而CWT对短时瞬态的建模能力更强。对语音和歌声进行的实验证实了我们提出的鉴别器的有效性。此外,基于STFT、CQT和CWT的鉴别器可以联合使用以获得更好的性能。所提出的鉴别器可以提高各种最先进的基于GAN的声码器的合成质量,包括HiFi—GAN,BigVGAN和APNet。
摘要:Generative Adversarial Network (GAN) based vocoders are superior in both inference speed and synthesis quality when reconstructing an audible waveform from an acoustic representation. This study focuses on improving the discriminator for GAN-based vocoders. Most existing Time-Frequency Representation (TFR)-based discriminators are rooted in Short-Time Fourier Transform (STFT), which owns a constant Time-Frequency (TF) resolution, linearly scaled center frequencies, and a fixed decomposition basis, making it incompatible with signals like singing voices that require dynamic attention for different frequency bands and different time intervals. Motivated by that, we propose a Multi-Scale Sub-Band Constant-Q Transform CQT (MS-SB-CQT) discriminator and a Multi-Scale Temporal-Compressed Continuous Wavelet Transform CWT (MS-TC-CWT) discriminator. Both CQT and CWT have a dynamic TF resolution for different frequency bands. In contrast, CQT has a better modeling ability in pitch information, and CWT has a better modeling ability in short-time transients. Experiments conducted on both speech and singing voices confirm the effectiveness of our proposed discriminators. Moreover, the STFT, CQT, and CWT-based discriminators can be used jointly for better performance. The proposed discriminators can boost the synthesis quality of various state-of-the-art GAN-based vocoders, including HiFi-GAN, BigVGAN, and APNet.


【4】 Investigating differences in lab-quality and remote recording methods  with dynamic acoustic measures
标题:通过动态声学测量调查实验室质量和远程记录方法的差异
链接:https://arxiv.org/abs/2404.17022
作者:Cong Zhang,Kathleen Jepson,Yu-Ying Chuang
摘要:语音研究越来越多地利用从参与者那里收集的数据,这些参与者在现成的设备上记录自己。虽然这样的记录是方便的,它们是否适合声学分析仍然是一个悬而未决的问题,特别是关于如何随着时间的推移,个别方法影响声学措施。我们使用分位数广义加性混合模型(QGAMM)来分析F0、强度以及第一和第二共振峰的测量,比较使用实验室标准记录方法记录的文件(带有外部麦克风的Zoom H6录音机),三种远程录音方法,(1)智能手机(AVR)上的Awesome语音录音机应用程序,(2)带有默认设置的Zoom会议应用程序(缩放默认),以及(3)具有“打开原始声音”设置的缩放会议应用程序(缩放原始)。在长时间记录会话文件的过程中,观察到缩放方法的线性时间对齐问题。然而,差异是不显着的话语长度文件。使用所有方法可靠地测量了F0。强度和共振峰呈现出无法简单校正的方法之间的非线性差异。总体而言,AVR文件与H6的最相似,因此AVR被认为是比Zoom-default或Zoom-raw更可靠的记录方法。
摘要:Increasingly, phonetic research utilizes data collected from participants who record themselves on readily available devices. Though such recordings are convenient, their suitability for acoustic analysis remains an open question, especially regarding how the individual methods affect acoustic measures over time. We used Quantile Generalized Additive Mixed Models (QGAMMs) to analyze measures of F0, intensity, and the first and second formants, comparing files recorded using a laboratory-standard recording method (Zoom H6 Recorder with an external microphone), to three remote recording methods, (1) the Awesome Voice Recorder application on a smartphone (AVR), (2) the Zoom meeting application with default settings (Zoom-default), and (3) the Zoom meeting application with the "Turn on Original Sound" setting (Zoom-raw). A linear temporal alignment issue was observed for the Zoom methods over the course of the long, recording session files. However, the difference was not significant for utterance-length files. F0 was reliably measured using all methods. Intensity and formants presented non-linear differences across methods that could not be corrected for simply. Overall, the AVR files were most similar to the H6's, and so AVR is deemed to be a more reliable recording method than either Zoom-default or Zoom-raw.

【5】 COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio  Representations
标题:COCOLA:以连贯性为导向的音乐音频表示对比学习
链接:https://arxiv.org/abs/2404.16969
作者:Ruben Ciranni,Emilian Postolache,Giorgio Mariani,Michele Mancusi,Luca Cosmo,Emanuele Rodolà
备注:Demo page: this https URL
摘要:我们提出了COCOLA(Coherence-Oriented Contrastive Learning for Audio),这是一种用于音乐音频表示的对比学习方法,可以捕获样本之间的谐波和节奏一致性。我们的方法操作的水平上的茎(或其组合)组成的音乐曲目,并允许客观评价的作曲模型的音乐伴奏生成的任务。我们还介绍了一个新的基准组成的音乐生成称为CompoNet,基于ControlNet \cite{zhang 2023 adding},概括的MSDM的任务,并量化它对后者使用COCOLA。我们发布了在包含单独词干的公共数据集(MUSDB 18-HQ,MoisesDB,Slakh 2100和CocoChorales)上训练的所有模型。
摘要:We present COCOLA (Coherence-Oriented Contrastive Learning for Audio), a contrastive learning method for musical audio representations that captures the harmonic and rhythmic coherence between samples. Our method operates at the level of stems (or their combinations) composing music tracks and allows the objective evaluation of compositional models for music in the task of accompaniment generation. We also introduce a new baseline for compositional music generation called CompoNet, based on ControlNet \cite{zhang2023adding}, generalizing the tasks of MSDM, and quantify it against the latter using COCOLA. We release all models trained on public datasets containing separate stems (MUSDB18-HQ, MoisesDB, Slakh2100, and CocoChorales).


【6】 Samsung Research China-Beijing at SemEval-2024 Task 3: A multi-stage  framework for Emotion-Cause Pair Extraction in Conversations
标题:三星研究中心中国-北京SemEval-2024任务3:对话中描述-原因对提取的多阶段框架
链接:https://arxiv.org/abs/2404.16905
作者:Shen Zhang,Haojie Zhang,Jing Zhang,Xudong Zhang,Yimeng Zhuang,Jinting Wu
摘要:在人机交互中,智能体通过理解人的情感来响应人是至关重要的。解开情绪的原因更具挑战性。会话中的多模态推理-原因对抽取是一个新的任务,负责识别情感和识别因果表达式。在这项研究中,我们提出了一个多阶段的框架来产生情感和提取情感因果对给定的目标情绪。在第一阶段,基于Llama-2的InstructERC被用来提取会话中每个话语的情感类别。在情感识别之后,采用双流注意力模型来提取子任务2的目标情感的情感因果对,而MuTEC用于提取子任务1的因果跨度。我们的方法在比赛中的两个子任务中都获得了第一名。
摘要:In human-computer interaction, it is crucial for agents to respond to human by understanding their emotions. Unraveling the causes of emotions is more challenging. A new task named Multimodal Emotion-Cause Pair Extraction in Conversations is responsible for recognizing emotion and identifying causal expressions. In this study, we propose a multi-stage framework to generate emotion and extract the emotion causal pairs given the target emotion. In the first stage, Llama-2-based InstructERC is utilized to extract the emotion category of each utterance in a conversation. After emotion recognition, a two-stream attention model is employed to extract the emotion causal pairs given the target emotion for subtask 2 while MuTEC is employed to extract causal span for subtask 1. Our approach achieved first place for both of the two subtasks in the competition.


【7】 A Semi-Automatic Approach to Create Large Gender- and Age-Balanced  Speaker Corpora: Usefulness of Speaker Diarization & Identification
标题:一种半自动方法创建大型性别和性别平衡的说话者库:说话者二元化和识别的有用性
链接:https://arxiv.org/abs/2404.17552
作者:Rémi Uro,David Doukhan,Albert Rilliard,Laëtitia Larcher,Anissa-Claire Adgharouamane,Marie Tahon,Antoine Laurent
备注:None
摘要:本文提出了一种半自动的方法来创建一个历时语料库的声音平衡的说话人的年龄,性别和录音期间,根据32个类别(2个性别,4个年龄段和4个录音期间)。语料库是在法国国家视听研究所(INA)选出的,每个类别至少有30名发言人(总共有960名发言人;目前只找到874名)。对于每个扬声器,语音摘录从视听文件中提取使用自动流水线组成的语音检测,背景音乐和重叠的语音去除和扬声器日记,用于呈现干净的扬声器段识别目标扬声器的人类注释。事实证明,这种管道非常有效,将手动处理减少了十倍。提供了对自动处理和最终输出的质量的评估。它显示了自动处理与最新处理的比较,并且输出为大多数选定的摘录提供了高质量的语音。该方法显示出创建已知目标说话人的大型语料库的希望。
摘要:This paper presents a semi-automatic approach to create a diachronic corpus of voices balanced for speaker's age, gender, and recording period, according to 32 categories (2 genders, 4 age ranges and 4 recording periods). Corpora were selected at French National Institute of Audiovisual (INA) to obtain at least 30 speakers per category (a total of 960 speakers; only 874 have be found yet). For each speaker, speech excerpts were extracted from audiovisual documents using an automatic pipeline consisting of speech detection, background music and overlapped speech removal and speaker diarization, used to present clean speaker segments to human annotators identifying target speakers. This pipeline proved highly effective, cutting down manual processing by a factor of ten. Evaluation of the quality of the automatic processing and of the final output is provided. It shows the automatic processing compare to up-to-date process, and that the output provides high quality speech for most of the selected excerpts. This method shows promise for creating large corpora of known target speakers.

【8】 The CARFAC v2 Cochlear Model in Matlab, NumPy, and JAX
标题:Matlab、NumPy和JAX中的CARFAC v2可卡因模型
链接:https://arxiv.org/abs/2404.17490
作者:Richard F. Lyon,Rob Schonberger,Malcolm Slaney,Mihajlo Velimirović,Honglin Yu
摘要:开源CARFAC(Cascade of Asymmetric Resonators with Fast-Acting Compression)耳蜗模型升级到第2版,改进了Matlab实现,并采用了新的Python/NumPy和JAX实现-但C++版本更改仍悬而未决。一个变化解决了先前报告的DC(直流或零频率)二次失真异常;另一个降低了高频下的神经同步;其他变化在默认配置中几乎没有或没有明显的影响。一个新的功能允许模拟耳蜗放大器功能的减少,作为迈向听力损伤的可微分参数化模型的一步。此外,集成到听觉模型识别器(AMT)已经得到了广泛的改进,因为之前的集成存在缺陷,不适合将CARFAC纳入多模型比较。
摘要:The open-source CARFAC (Cascade of Asymmetric Resonators with Fast-Acting Compression) cochlear model is upgraded to version 2, with improvements to the Matlab implementation, and with new Python/NumPy and JAX implementations -- but C++ version changes are still pending. One change addresses the DC (direct current, or zero frequency) quadratic distortion anomaly previously reported; another reduces the neural synchrony at high frequencies; the others have little or no noticeable effect in the default configuration. A new feature allows modeling a reduction of cochlear amplifier function, as a step toward a differentiable parameterized model of hearing impairment. In addition, the integration into the Auditory Model Toolbox (AMT) has been extensively improved, as the prior integration had bugs that made it unsuitable for including CARFAC in multi-model comparisons.

【9】 Exploring Pre-trained General-purpose Audio Representations for Heart  Murmur Detection
标题:探索预训练的通用音频表示以进行心脏杂音检测
链接:https://arxiv.org/abs/2404.17107
作者:Daisuke Niizumi,Daiki Takeuchi,Yasunori Ohishi,Noboru Harada,Kunio Kashino
备注:4 pages, 1 figure, and 4 tables. Accepted by IEEE EMBC 2024
摘要:为了减少对熟练的临床医生在心音解释方面的需求,最近关于自动心脏听诊的研究探索了深度学习方法。然而,尽管深度学习需要大量数据,但心音数据集的大小有限,并且没有预先训练的模型可用。相反,许多用于一般音频任务的预训练模型可用作通用音频表示。这项研究探讨了在大规模数据集上预训练的通用音频表示在心脏杂音检测中进行迁移学习的潜力。在CirCor DigiScope心音数据集上的实验表明,最近的自监督学习Masked Modeling Duo(M2D)优于以前的方法,其加权准确率为0.832,未加权平均召回率为0.713。实验进一步证实了通过将M2D与其他模型集成来提高性能。这些结果证明了通用音频表示在处理心音中的有效性,并为进一步的应用开辟了道路。我们的代码可以在https://github.com/nttcslab/m2d/tree/master/app/circor上在线获得,它运行在24 GB的消费级GPU上
摘要:To reduce the need for skilled clinicians in heart sound interpretation, recent studies on automating cardiac auscultation have explored deep learning approaches. However, despite the demands for large data for deep learning, the size of the heart sound datasets is limited, and no pre-trained model is available. On the contrary, many pre-trained models for general audio tasks are available as general-purpose audio representations. This study explores the potential of general-purpose audio representations pre-trained on large-scale datasets for transfer learning in heart murmur detection. Experiments on the CirCor DigiScope heart sound dataset show that the recent self-supervised learning Masked Modeling Duo (M2D) outperforms previous methods with the results of a weighted accuracy of 0.832 and an unweighted average recall of 0.713. Experiments further confirm improved performance by ensembling M2D with other models. These results demonstrate the effectiveness of general-purpose audio representation in processing heart sounds and open the way for further applications. Our code is available online which runs on a 24 GB consumer GPU at https://github.com/nttcslab/m2d/tree/master/app/circor


eess.AS音频处理
【1】 A Semi-Automatic Approach to Create Large Gender- and Age-Balanced  Speaker Corpora: Usefulness of Speaker Diarization & Identification
标题:一种半自动方法创建大型性别和性别平衡的说话者库:说话者二元化和识别的有用性
链接:https://arxiv.org/abs/2404.17552
作者:Rémi Uro,David Doukhan,Albert Rilliard,Laëtitia Larcher,Anissa-Claire Adgharouamane,Marie Tahon,Antoine Laurent
备注:None
摘要:本文提出了一种半自动的方法来创建一个历时语料库的声音平衡的说话人的年龄,性别和录音期间,根据32个类别(2个性别,4个年龄段和4个录音期间)。语料库是在法国国家视听研究所(INA)选出的,每个类别至少有30名发言人(总共有960名发言人;目前只找到874名)。对于每个扬声器,语音摘录从视听文件中提取使用自动流水线组成的语音检测,背景音乐和重叠的语音去除和扬声器日记,用于呈现干净的扬声器段识别目标扬声器的人类注释。事实证明,这种管道非常有效,将手动处理减少了十倍。提供了对自动处理和最终输出的质量的评估。它显示了自动处理与最新处理的比较,并且输出为大多数选定的摘录提供了高质量的语音。该方法显示出创建已知目标说话人的大型语料库的希望。
摘要:This paper presents a semi-automatic approach to create a diachronic corpus of voices balanced for speaker's age, gender, and recording period, according to 32 categories (2 genders, 4 age ranges and 4 recording periods). Corpora were selected at French National Institute of Audiovisual (INA) to obtain at least 30 speakers per category (a total of 960 speakers; only 874 have be found yet). For each speaker, speech excerpts were extracted from audiovisual documents using an automatic pipeline consisting of speech detection, background music and overlapped speech removal and speaker diarization, used to present clean speaker segments to human annotators identifying target speakers. This pipeline proved highly effective, cutting down manual processing by a factor of ten. Evaluation of the quality of the automatic processing and of the final output is provided. It shows the automatic processing compare to up-to-date process, and that the output provides high quality speech for most of the selected excerpts. This method shows promise for creating large corpora of known target speakers.

【2】 The CARFAC v2 Cochlear Model in Matlab, NumPy, and JAX
标题:Matlab、NumPy和JAX中的CARFAC v2可卡因模型
链接:https://arxiv.org/abs/2404.17490
作者:Richard F. Lyon,Rob Schonberger,Malcolm Slaney,Mihajlo Velimirović,Honglin Yu
摘要:开源CARFAC(Cascade of Asymmetric Resonators with Fast-Acting Compression)耳蜗模型升级到第2版,改进了Matlab实现,并采用了新的Python/NumPy和JAX实现-但C++版本更改仍悬而未决。一个变化解决了先前报告的DC(直流或零频率)二次失真异常;另一个降低了高频下的神经同步;其他变化在默认配置中几乎没有或没有明显的影响。一个新的功能允许模拟耳蜗放大器功能的减少,作为迈向听力损伤的可微分参数化模型的一步。此外,集成到听觉模型识别器(AMT)已经得到了广泛的改进,因为之前的集成存在缺陷,不适合将CARFAC纳入多模型比较。
摘要:The open-source CARFAC (Cascade of Asymmetric Resonators with Fast-Acting Compression) cochlear model is upgraded to version 2, with improvements to the Matlab implementation, and with new Python/NumPy and JAX implementations -- but C++ version changes are still pending. One change addresses the DC (direct current, or zero frequency) quadratic distortion anomaly previously reported; another reduces the neural synchrony at high frequencies; the others have little or no noticeable effect in the default configuration. A new feature allows modeling a reduction of cochlear amplifier function, as a step toward a differentiable parameterized model of hearing impairment. In addition, the integration into the Auditory Model Toolbox (AMT) has been extensively improved, as the prior integration had bugs that made it unsuitable for including CARFAC in multi-model comparisons.

【3】 Exploring Pre-trained General-purpose Audio Representations for Heart  Murmur Detection
标题:探索预训练的通用音频表示以进行心脏杂音检测
链接:https://arxiv.org/abs/2404.17107
作者:Daisuke Niizumi,Daiki Takeuchi,Yasunori Ohishi,Noboru Harada,Kunio Kashino
备注:4 pages, 1 figure, and 4 tables. Accepted by IEEE EMBC 2024
摘要:为了减少对熟练的临床医生在心音解释方面的需求,最近关于自动心脏听诊的研究探索了深度学习方法。然而,尽管深度学习需要大量数据,但心音数据集的大小有限,并且没有预先训练的模型可用。相反,许多用于一般音频任务的预训练模型可用作通用音频表示。这项研究探讨了在大规模数据集上预训练的通用音频表示在心脏杂音检测中进行迁移学习的潜力。在CirCor DigiScope心音数据集上的实验表明,最近的自监督学习Masked Modeling Duo(M2D)优于以前的方法,其加权准确率为0.832,未加权平均召回率为0.713。实验进一步证实了通过将M2D与其他模型集成来提高性能。这些结果证明了通用音频表示在处理心音中的有效性,并为进一步的应用开辟了道路。我们的代码可在https://github.com/nttcslab/m2d/tree/master/app/circor上在线获得,它运行在24 GB的消费级GPU上
摘要:To reduce the need for skilled clinicians in heart sound interpretation, recent studies on automating cardiac auscultation have explored deep learning approaches. However, despite the demands for large data for deep learning, the size of the heart sound datasets is limited, and no pre-trained model is available. On the contrary, many pre-trained models for general audio tasks are available as general-purpose audio representations. This study explores the potential of general-purpose audio representations pre-trained on large-scale datasets for transfer learning in heart murmur detection. Experiments on the CirCor DigiScope heart sound dataset show that the recent self-supervised learning Masked Modeling Duo (M2D) outperforms previous methods with the results of a weighted accuracy of 0.832 and an unweighted average recall of 0.713. Experiments further confirm improved performance by ensembling M2D with other models. These results demonstrate the effectiveness of general-purpose audio representation in processing heart sounds and open the way for further applications. Our code is available online which runs on a 24 GB consumer GPU at https://github.com/nttcslab/m2d/tree/master/app/circor

【4】 Device Feature based on Graph Fourier Transformation with Logarithmic  Processing For Detection of Replay Speech Attacks
标题:基于图形傅里叶变换和算术处理的设备特征检测重播语音攻击
链接:https://arxiv.org/abs/2404.17280
作者:Mingrui He,Longting Xu,Han Wang,Mingjun Zhang,Rohan Kumar Das
摘要:对自动说话人确认系统最常见的欺骗攻击是重放语音攻击。重放语音的检测严重依赖于重放配置信息。先前的研究表明,图形傅立叶变换衍生的功能可以有效地检测重放语音,但忽略设备和环境噪声的影响。在这项工作中,我们提出了一个新的功能,图形频率设备倒谱系数,来自图形频域使用设备相关的线性变换。我们还介绍了两种新的表示:图频率对数系数和图频率对数设备系数。我们使用传统的高斯混合模型和光卷积神经网络系统作为分类器来评估我们的方法。在ASVspoof 2017 V2,ASVspoof 2019物理访问和ASVspoof 2021物理访问数据集上,我们提出的功能优于已知的前端,证明了它们对重放语音检测的有效性。
摘要:The most common spoofing attacks on automatic speaker verification systems are replay speech attacks. Detection of replay speech heavily relies on replay configuration information. Previous studies have shown that graph Fourier transform-derived features can effectively detect replay speech but ignore device and environmental noise effects. In this work, we propose a new feature, the graph frequency device cepstral coefficient, derived from the graph frequency domain using a device-related linear transformation. We also introduce two novel representations: graph frequency logarithmic coefficient and graph frequency logarithmic device coefficient. We evaluate our methods using traditional Gaussian mixture model and light convolutional neural network systems as classifiers. On the ASVspoof 2017 V2, ASVspoof 2019 physical access, and ASVspoof 2021 physical access datasets, our proposed features outperform known front-ends, demonstrating their effectiveness for replay speech detection.

【5】 An Investigation of Time-Frequency Representation Discriminators for  High-Fidelity Vocoder
标题:高保真声码器时频表示鉴别器的研究
链接:https://arxiv.org/abs/2404.17161
作者:Yicheng Gu,Xueyao Zhang,Liumeng Xue,Haizhou Li,Zhizheng Wu
备注:arXiv admin note: text overlap with arXiv:2311.14957
摘要:当从声学表示重建可听波形时,基于生成对抗网络(GAN)的声码器在推理速度和合成质量方面都是优越的。本研究的重点是改善基于GAN的声码器的可扩展性。大多数现有的基于时频表示(TFR)的鉴别器植根于短时傅里叶变换(STFT),其具有恒定的时频(TF)分辨率,线性缩放的中心频率和固定的分解基础,使得其与需要动态注意不同频带和不同时间间隔的歌声等信号不兼容。基于此,本文提出了多尺度子带恒Q变换CQT(MS-SB-CQT)和多尺度时间压缩连续小波变换CWT(MS-TC-CWT)。CQT和CWT都具有针对不同频带的动态TF分辨率。相比之下,CQT对基音信息的建模能力更强,而CWT对短时瞬态的建模能力更强。对语音和歌声进行的实验证实了我们提出的鉴别器的有效性。此外,基于STFT、CQT和CWT的鉴别器可以联合使用以获得更好的性能。所提出的鉴别器可以提高各种最先进的基于GAN的声码器的合成质量,包括HiFi-GAN,BigVGAN和APNet。
摘要:Generative Adversarial Network (GAN) based vocoders are superior in both inference speed and synthesis quality when reconstructing an audible waveform from an acoustic representation. This study focuses on improving the discriminator for GAN-based vocoders. Most existing Time-Frequency Representation (TFR)-based discriminators are rooted in Short-Time Fourier Transform (STFT), which owns a constant Time-Frequency (TF) resolution, linearly scaled center frequencies, and a fixed decomposition basis, making it incompatible with signals like singing voices that require dynamic attention for different frequency bands and different time intervals. Motivated by that, we propose a Multi-Scale Sub-Band Constant-Q Transform CQT (MS-SB-CQT) discriminator and a Multi-Scale Temporal-Compressed Continuous Wavelet Transform CWT (MS-TC-CWT) discriminator. Both CQT and CWT have a dynamic TF resolution for different frequency bands. In contrast, CQT has a better modeling ability in pitch information, and CWT has a better modeling ability in short-time transients. Experiments conducted on both speech and singing voices confirm the effectiveness of our proposed discriminators. Moreover, the STFT, CQT, and CWT-based discriminators can be used jointly for better performance. The proposed discriminators can boost the synthesis quality of various state-of-the-art GAN-based vocoders, including HiFi-GAN, BigVGAN, and APNet.


【6】 Investigating differences in lab-quality and remote recording methods  with dynamic acoustic measures
标题:通过动态声学测量调查实验室质量和远程记录方法的差异
链接:https://arxiv.org/abs/2404.17022
作者:Cong Zhang,Kathleen Jepson,Yu-Ying Chuang
摘要:语音研究越来越多地利用从参与者那里收集的数据,这些参与者在现成的设备上记录自己。虽然这样的记录是方便的,它们是否适合声学分析仍然是一个悬而未决的问题,特别是关于如何随着时间的推移,个别方法影响声学措施。我们使用分位数广义加性混合模型(QGAMM)来分析F0、强度以及第一和第二共振峰的测量,比较使用实验室标准记录方法记录的文件(带有外部麦克风的Zoom H6录音机),三种远程录音方法,(1)智能手机(AVR)上的Awesome语音录音机应用程序,(2)带有默认设置的Zoom会议应用程序(缩放默认),以及(3)具有“打开原始声音”设置的缩放会议应用程序(缩放原始)。在长时间记录会话文件的过程中,观察到缩放方法的线性时间对齐问题。然而,差异是不显着的话语长度文件。使用所有方法可靠地测量了F0。强度和共振峰呈现出无法简单校正的方法之间的非线性差异。总体而言,AVR文件与H6的最相似,因此AVR被认为是比Zoom-default或Zoom-raw更可靠的记录方法。
摘要:Increasingly, phonetic research utilizes data collected from participants who record themselves on readily available devices. Though such recordings are convenient, their suitability for acoustic analysis remains an open question, especially regarding how the individual methods affect acoustic measures over time. We used Quantile Generalized Additive Mixed Models (QGAMMs) to analyze measures of F0, intensity, and the first and second formants, comparing files recorded using a laboratory-standard recording method (Zoom H6 Recorder with an external microphone), to three remote recording methods, (1) the Awesome Voice Recorder application on a smartphone (AVR), (2) the Zoom meeting application with default settings (Zoom-default), and (3) the Zoom meeting application with the "Turn on Original Sound" setting (Zoom-raw). A linear temporal alignment issue was observed for the Zoom methods over the course of the long, recording session files. However, the difference was not significant for utterance-length files. F0 was reliably measured using all methods. Intensity and formants presented non-linear differences across methods that could not be corrected for simply. Overall, the AVR files were most similar to the H6's, and so AVR is deemed to be a more reliable recording method than either Zoom-default or Zoom-raw.


【7】 COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio  Representations
标题:COCOLA:以连贯性为导向的音乐音频表示对比学习
链接:https://arxiv.org/abs/2404.16969
作者:Ruben Ciranni,Emilian Postolache,Giorgio Mariani,Michele Mancusi,Luca Cosmo,Emanuele Rodolà
备注:Demo page: this https URL
摘要:我们提出了COCOLA(Coherence—Oriented Contrastive Learning for Audio),这是一种用于音乐音频表示的对比学习方法,可以捕获样本之间的谐波和节奏一致性。我们的方法操作的水平上的茎(或其组合)组成的音乐曲目,并允许客观评价的作曲模型的音乐伴奏生成的任务。我们还介绍了一个新的基准组成的音乐生成称为CompoNet,基于ControlNet\cite {zhang2023adding},概括的MSDM的任务,并量化它对后者使用COCOLA。我们发布了在包含单独词干的公共数据集(MUSDB18—HQ,MoisesDB,Slakh2100和CocoChorales)上训练的所有模型。
摘要:We present COCOLA (Coherence-Oriented Contrastive Learning for Audio), a contrastive learning method for musical audio representations that captures the harmonic and rhythmic coherence between samples. Our method operates at the level of stems (or their combinations) composing music tracks and allows the objective evaluation of compositional models for music in the task of accompaniment generation. We also introduce a new baseline for compositional music generation called CompoNet, based on ControlNet \cite{zhang2023adding}, generalizing the tasks of MSDM, and quantify it against the latter using COCOLA. We release all models trained on public datasets containing separate stems (MUSDB18-HQ, MoisesDB, Slakh2100, and CocoChorales).

【8】 Samsung Research China-Beijing at SemEval-2024 Task 3: A multi-stage  framework for Emotion-Cause Pair Extraction in Conversations
标题:三星研究中心中国-北京SemEval-2024任务3:对话中描述-原因对提取的多阶段框架
链接:https://arxiv.org/abs/2404.16905
作者:Shen Zhang,Haojie Zhang,Jing Zhang,Xudong Zhang,Yimeng Zhuang,Jinting Wu
摘要:在人机交互中,智能体通过理解人的情感来响应人是至关重要的。解开情绪的原因更具挑战性。会话中的多模态推理-原因对抽取是一个新的任务,负责识别情感和识别因果表达式。在这项研究中,我们提出了一个多阶段的框架来产生情感和提取情感因果对给定的目标情绪。在第一阶段,基于Llama-2的InstructERC被用来提取会话中每个话语的情感类别。在情感识别之后,采用双流注意力模型来提取子任务2的目标情感的情感因果对,而MuTEC用于提取子任务1的因果跨度。我们的方法在比赛中的两个子任务中都获得了第一名。
摘要:In human-computer interaction, it is crucial for agents to respond to human by understanding their emotions. Unraveling the causes of emotions is more challenging. A new task named Multimodal Emotion-Cause Pair Extraction in Conversations is responsible for recognizing emotion and identifying causal expressions. In this study, we propose a multi-stage framework to generate emotion and extract the emotion causal pairs given the target emotion. In the first stage, Llama-2-based InstructERC is utilized to extract the emotion category of each utterance in a conversation. After emotion recognition, a two-stream attention model is employed to extract the emotion causal pairs given the target emotion for subtask 2 while MuTEC is employed to extract causal span for subtask 1. Our approach achieved first place for both of the two subtasks in the competition.

机器翻译由腾讯交互翻译提供,仅供参考