本文经arXiv每日学术速递授权转载
【1】 SELD-Mamba: Selective State-Space Model for Sound Event Localization and Detection with Source Distance Estimation
标题: SELD-Mamba:利用源距离估计进行声音事件定位和检测的选择性状态空间模型
作者:Da Mu,Zhicheng Zhang,Haobo Yue,Zehao Wang,Jin Tang,Jianqin Yin
链接:点击下载PDF文件
【2】 MIDI-to-Tab: Guitar Tablature Inference via Masked Language Modeling
标题: MIDI-to-Tab:通过掩蔽语言建模的吉他谱推理
作者:Drew Edwards,Xavier Riley,Pedro Sarmento,Simon Dixon
备注:Reviewed pre-print accepted for publication at ISMIR 2024
链接:点击下载PDF文件
【3】 AcousAF: Acoustic Sensing-Based Atrial Fibrillation Detection System for Mobile Phones
标题: AcoustAF:基于声学传感的手机心房颤动检测系统
作者:Xuanyu Liu,Haoxian Liu,Jiao Li,Zongqi Yang,Yi Huang,Jin Zhang
备注:Accepted for publication in Companion of the 2024 ACM International Joint Conference on Pervasive and Ubiquitous Computing (UbiComp Companion '24)
链接:点击下载PDF文件
【4】 TEAdapter: Supply abundant guidance for controllable text-to-music generation
标题: TEAdaptor:为可控的文本到音乐生成提供丰富的指导
作者:Jialing Zou,Jiahao Mei,Xudong Nan,Jinghua Li,Daoguo Dong,Liang He
Journal-ref:2024 IEEE International Conference on Multimedia and Expo (ICME 2024)
链接:点击下载PDF文件
【5】 Hyper Recurrent Neural Network: Condition Mechanisms for Black-box Audio Effect Modeling
标题: 超回归神经网络:黑匣子音效建模的条件机制
作者:Yen-Tung Yeh,Wen-Yi Hsiao,Yi-Hsuan Yang
备注:Accepted to DAFx24
链接:点击下载PDF文件
【6】 Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
标题: 利用一致性保持损失和感知对比度扩展来促进基于SSL的语音增强
作者:Muhammad Salman Khan,Moreno La Quatra,Kuo-Hsuan Hung,Szu-Wei Fu,Sabato Marco Siniscalchi,Yu Tsao
链接:点击下载PDF文件
【7】 Quantifying the Corpus Bias Problem in Automatic Music Transcription Systems
标题: 自动音乐转录系统中的数据库偏差问题的量化
作者:Lukáš Samuel Marták,Patricia Hu,Gerhard Widmer
备注:2 pages, 1 figure, presented in the 1st International Workshop on Sound Signal Processing Applications (IWSSPA) 2024
链接:点击下载PDF文件
【8】 MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
标题: MulliVC:具有循环一致性的多语言语音转换
作者:Jiawei Huang,Chen Zhang,Yi Ren,Ziyue Jiang,Zhenhui Ye,Jinglin Liu,Jinzheng He,Xiang Yin,Zhou Zhao
链接:点击下载PDF文件
【9】 ADD 2023: Towards Audio Deepfake Detection and Analysis in the Wild
标题: ADD 2023:走向野外音频Deepfake检测和分析
作者:Jiangyan Yi,Chu Yuan Zhang,Jianhua Tao,Chenglong Wang,Xinrui Yan,Yong Ren,Hao Gu,Junzuo Zhou
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
标题: ADD 2023:走向野外音频Deepfake检测和分析
作者:Jiangyan Yi,Chu Yuan Zhang,Jianhua Tao,Chenglong Wang,Xinrui Yan,Yong Ren,Hao Gu,Junzuo Zhou
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
【2】 SELD-Mamba: Selective State-Space Model for Sound Event Localization and Detection with Source Distance Estimation
标题: SELD-Mamba:利用源距离估计进行声音事件定位和检测的选择性状态空间模型
作者:Da Mu,Zhicheng Zhang,Haobo Yue,Zehao Wang,Jin Tang,Jianqin Yin
链接:点击下载PDF文件
【3】 AcousAF: Acoustic Sensing-Based Atrial Fibrillation Detection System for Mobile Phones
标题: AcoustAF:基于声学传感的手机心房颤动检测系统
作者:Xuanyu Liu,Haoxian Liu,Jiao Li,Zongqi Yang,Yi Huang,Jin Zhang
备注:Accepted for publication in Companion of the 2024 ACM International Joint Conference on Pervasive and Ubiquitous Computing (UbiComp Companion '24)
链接:点击下载PDF文件
【4】 TEAdapter: Supply abundant guidance for controllable text-to-music generation
标题: TEAdaptor:为可控的文本到音乐生成提供丰富的指导
作者:Jialing Zou,Jiahao Mei,Xudong Nan,Jinghua Li,Daoguo Dong,Liang He
Journal-ref:2024 IEEE International Conference on Multimedia and Expo (ICME 2024)
链接:点击下载PDF文件
【5】 Hyper Recurrent Neural Network: Condition Mechanisms for Black-box Audio Effect Modeling
标题: 超回归神经网络:黑匣子音效建模的条件机制
作者:Yen-Tung Yeh,Wen-Yi Hsiao,Yi-Hsuan Yang
备注:Accepted to DAFx24
链接:点击下载PDF文件
【6】 Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
标题: 利用一致性保持损失和感知对比度扩展来促进基于SSL的语音增强
作者:Muhammad Salman Khan,Moreno La Quatra,Kuo-Hsuan Hung,Szu-Wei Fu,Sabato Marco Siniscalchi,Yu Tsao
链接:点击下载PDF文件
【7】 Quantifying the Corpus Bias Problem in Automatic Music Transcription Systems
标题: 自动音乐转录系统中的数据库偏差问题的量化
作者:Lukáš Samuel Marták,Patricia Hu,Gerhard Widmer
备注:2 pages, 1 figure, presented in the 1st International Workshop on Sound Signal Processing Applications (IWSSPA) 2024
链接:点击下载PDF文件
【8】 MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
标题: MulliVC:具有循环一致性的多语言语音转换
作者:Jiawei Huang,Chen Zhang,Yi Ren,Ziyue Jiang,Zhenhui Ye,Jinglin Liu,Jinzheng He,Xiang Yin,Zhou Zhao
链接:点击下载PDF文件
【9】 Abstractive summarization from Audio Transcription
标题: 音频转录的抽象摘要
作者:Ilia Derkach
备注:36 pages, Master's thesis, 14 figures
链接:点击下载PDF文件
标题: SELD-Mamba:利用源距离估计进行声音事件定位和检测的选择性状态空间模型
作者:Da Mu,Zhicheng Zhang,Haobo Yue,Zehao Wang,Jin Tang,Jianqin Yin
链接:点击下载PDF文件
摘要:在声音事件定位和检测(SELD)任务中,基于Transformer的模型已经展示了令人印象深刻的功能。然而,Transformer的自注意机制的二次复杂性导致计算效率低下。在本文中,我们提出了一种名为SELD-Mamba的SELD网络架构,它利用了Mamba(一种选择性状态空间模型)。我们采用事件独立网络V2(EINV 2)作为基础框架,并将其Conformer块替换为双向Mamba块,以捕获更广泛的上下文信息,同时保持计算效率。此外,我们实现了一个两阶段的训练方法,第一阶段专注于声音事件检测(SED)和到达方向(DoA)估计损失,第二阶段重新引入源距离估计(DOA)损失。我们在2024 DCASE Challenge Task 3数据集上的实验结果证明了选择性状态空间模型在SELD中的有效性,并突出了两阶段训练方法在增强SELD性能方面的优势。摘要:In the Sound Event Localization and Detection (SELD) task, Transformer-based models have demonstrated impressive capabilities. However, the quadratic complexity of the Transformer's self-attention mechanism results in computational inefficiencies. In this paper, we propose a network architecture for SELD called SELD-Mamba, which utilizes Mamba, a selective state-space model. We adopt the Event-Independent Network V2 (EINV2) as the foundational framework and replace its Conformer blocks with bidirectional Mamba blocks to capture a broader range of contextual information while maintaining computational efficiency. Additionally, we implement a two-stage training method, with the first stage focusing on Sound Event Detection (SED) and Direction of Arrival (DoA) estimation losses, and the second stage reintroducing the Source Distance Estimation (SDE) loss. Our experimental results on the 2024 DCASE Challenge Task3 dataset demonstrate the effectiveness of the selective state-space model in SELD and highlight the benefits of the two-stage training approach in enhancing SELD performance.
【2】 MIDI-to-Tab: Guitar Tablature Inference via Masked Language Modeling
标题: MIDI-to-Tab:通过掩蔽语言建模的吉他谱推理
作者:Drew Edwards,Xavier Riley,Pedro Sarmento,Simon Dixon
备注:Reviewed pre-print accepted for publication at ISMIR 2024
链接:点击下载PDF文件
摘要:吉他乐谱丰富了传统音乐记谱法的结构,将每个音符分配给特定调音中的吉他弦和音品,精确地指示在乐器上演奏音符的位置。从符号音乐表示生成指谱的问题涉及在整个作曲或演奏中推断每个音符的这种弦和音品分配。在吉他上,大多数音高都可能有多个弦-品分配,这导致了一个很大的组合空间,阻止了穷举搜索方法。大多数现代方法使用基于约束的动态规划来最小化某些成本函数(例如,手位置移动)。在这项工作中,我们介绍了一种新的深度学习解决方案,以符号吉他指法估计。我们训练一个编码器-解码器Transformer模型,在一个掩码语言建模范例中将音符分配给字符串。该模型首先在DadaGP上进行预训练,DadaGP是一个超过25 K表格的数据集,然后在一组专业转录的吉他表演上进行微调。鉴于评估指法质量的主观性质,我们在吉他手中进行了一项用户研究,其中我们要求参与者对同一个四小节摘录的多个版本的指法的可玩性进行评价。结果表明,我们的系统显着优于竞争算法。摘要:Guitar tablatures enrich the structure of traditional music notation by assigning each note to a string and fret of a guitar in a particular tuning, indicating precisely where to play the note on the instrument. The problem of generating tablature from a symbolic music representation involves inferring this string and fret assignment per note across an entire composition or performance. On the guitar, multiple string-fret assignments are possible for most pitches, which leads to a large combinatorial space that prevents exhaustive search approaches. Most modern methods use constraint-based dynamic programming to minimize some cost function (e.g. hand position movement). In this work, we introduce a novel deep learning solution to symbolic guitar tablature estimation. We train an encoder-decoder Transformer model in a masked language modeling paradigm to assign notes to strings. The model is first pre-trained on DadaGP, a dataset of over 25K tablatures, and then fine-tuned on a curated set of professionally transcribed guitar performances. Given the subjective nature of assessing tablature quality, we conduct a user study amongst guitarists, wherein we ask participants to rate the playability of multiple versions of tablature for the same four-bar excerpt. The results indicate our system significantly outperforms competing algorithms.
【3】 AcousAF: Acoustic Sensing-Based Atrial Fibrillation Detection System for Mobile Phones
标题: AcoustAF:基于声学传感的手机心房颤动检测系统
作者:Xuanyu Liu,Haoxian Liu,Jiao Li,Zongqi Yang,Yi Huang,Jin Zhang
备注:Accepted for publication in Companion of the 2024 ACM International Joint Conference on Pervasive and Ubiquitous Computing (UbiComp Companion '24)
链接:点击下载PDF文件
摘要:心房颤动(AF)的特征是起源于心房的不规则电脉冲,其可导致严重的并发症甚至死亡。由于AF的间歇性,早期和及时监测AF对于患者预防病情进一步恶化至关重要。虽然动态心电动态监测仪提供准确的监测,这些设备的高成本阻碍了他们更广泛的采用。当前基于移动设备的AF检测系统提供便携式解决方案。然而,这些系统具有各种适用性问题,例如容易受环境因素影响并且需要大量用户努力。为了克服上述限制,我们提出了AcoustAF,一种基于智能手机声学传感器的新型AF检测系统。特别是,我们探索的脉搏波采集从手腕使用智能手机扬声器和麦克风的潜力。此外,我们提出了一个精心设计的框架,包括脉搏波探测,脉搏波提取和AF检测,以确保准确和可靠的AF检测。我们使用智能手机上的自定义数据收集应用程序收集20名参与者的数据。大量的实验结果表明,我们的系统的高性能,92.8%的准确率,86.9%的精度,87.4%的召回率,和87.1%的F1分数。摘要:Atrial fibrillation (AF) is characterized by irregular electrical impulses originating in the atria, which can lead to severe complications and even death. Due to the intermittent nature of the AF, early and timely monitoring of AF is critical for patients to prevent further exacerbation of the condition. Although ambulatory ECG Holter monitors provide accurate monitoring, the high cost of these devices hinders their wider adoption. Current mobile-based AF detection systems offer a portable solution. However, these systems have various applicability issues, such as being easily affected by environmental factors and requiring significant user effort. To overcome the above limitations, we present AcousAF, a novel AF detection system based on acoustic sensors of smartphones. Particularly, we explore the potential of pulse wave acquisition from the wrist using smartphone speakers and microphones. In addition, we propose a well-designed framework comprised of pulse wave probing, pulse wave extraction, and AF detection to ensure accurate and reliable AF detection. We collect data from 20 participants utilizing our custom data collection application on the smartphone. Extensive experimental results demonstrate the high performance of our system, with 92.8% accuracy, 86.9% precision, 87.4% recall, and 87.1% F1 Score.
【4】 TEAdapter: Supply abundant guidance for controllable text-to-music generation
标题: TEAdaptor:为可控的文本到音乐生成提供丰富的指导
作者:Jialing Zou,Jiahao Mei,Xudong Nan,Jinghua Li,Daoguo Dong,Liang He
Journal-ref:2024 IEEE International Conference on Multimedia and Expo (ICME 2024)
链接:点击下载PDF文件
摘要:虽然目前的文本引导音乐生成技术可以应对简单的创意场景,但随着用户需求变得更加复杂,实现对单个文本模态条件的细粒度控制仍然具有挑战性。因此,我们介绍了TEAcher适配器(TEAdapter),一个紧凑的插件,旨在指导生成过程中,由用户提供的各种控制信息。此外,我们探讨了可控生成的扩展音乐,利用TEAdapter控制组训练的数据不同的结构功能。一般来说,我们考虑全局、元素和结构级别的控制。实验结果表明,所提出的TEAdapter,使多个精确的控制,并确保高品质的音乐生成。我们的模块也是轻量级的,可转移到任何扩散模型架构。可用的代码和演示将很快在https: github.com Ashley1101 TEAdapter上找到。摘要:Although current text-guided music generation technology can cope with simple creative scenarios, achieving fine-grained control over individual text-modality conditions remains challenging as user demands become more intricate. Accordingly, we introduce the TEAcher Adapter (TEAdapter), a compact plugin designed to guide the generation process with diverse control information provided by users. In addition, we explore the controllable generation of extended music by leveraging TEAdapter control groups trained on data of distinct structural functionalities. In general, we consider controls over global, elemental, and structural levels. Experimental results demonstrate that the proposed TEAdapter enables multiple precise controls and ensures high-quality music generation. Our module is also lightweight and transferable to any diffusion model architecture. Available code and demos will be found soon at https: github.com Ashley1101 TEAdapter.
【5】 Hyper Recurrent Neural Network: Condition Mechanisms for Black-box Audio Effect Modeling
标题: 超回归神经网络:黑匣子音效建模的条件机制
作者:Yen-Tung Yeh,Wen-Yi Hsiao,Yi-Hsuan Yang
备注:Accepted to DAFx24
链接:点击下载PDF文件
摘要:递归神经网络(RNN)在音频效果的虚拟模拟建模方面取得了令人印象深刻的结果。这些网络使用一系列矩阵乘法和非线性激活函数来处理时域音频信号,以准确地模拟目标设备的行为。为了对基于RNN的模型的旋钮的影响进行额外建模,现有方法通过将控制参数与输入信号的某种中间表示逐通道地级联来集成控制参数。虽然该方法是参数有效的,但是存在进一步提高所生成的音频的质量的空间,因为基于级联的调节方法在调制信号方面具有有限的能力。在本文中,我们提出了三种新的RNN调节机制,为黑盒虚拟模拟建模量身定制。这些先进的条件反射机制基于控制参数调节模型,在各种评估指标上产生优于现有基于RNN和CNN的架构的结果。摘要:Recurrent neural networks (RNNs) have demonstrated impressive results for virtual analog modeling of audio effects. These networks process time-domain audio signals using a series of matrix multiplication and nonlinear activation functions to emulate the behavior of the target device accurately. To additionally model the effect of the knobs for an RNN-based model, existing approaches integrate control parameters by concatenating them channel-wisely with some intermediate representation of the input signal. While this method is parameter-efficient, there is room to further improve the quality of generated audio because the concatenation-based conditioning method has limited capacity in modulating signals. In this paper, we propose three novel conditioning mechanisms for RNNs, tailored for black-box virtual analog modeling. These advanced conditioning mechanisms modulate the model based on control parameters, yielding superior results to existing RNN- and CNN-based architectures across various evaluation metrics.
【6】 Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
标题: 利用一致性保持损失和感知对比度扩展来促进基于SSL的语音增强
作者:Muhammad Salman Khan,Moreno La Quatra,Kuo-Hsuan Hung,Szu-Wei Fu,Sabato Marco Siniscalchi,Yu Tsao
链接:点击下载PDF文件
摘要:自监督表示学习(SSL)已经在几个下游语音任务上取得了SOTA结果,但基于SSL的语音增强(SE)解决方案仍然落后。为了解决这个问题,我们利用三个主要的想法:(i)基于变换的掩蔽生成,(ii)一致性保持损失,以及(iii)感知对比度拉伸(PCS)。详细地,构象层,利用注意力机制,被引入到有效地建模帧级表示,并获得理想的比例掩码(SNR)的SE。此外,我们将一致性的损失函数,处理输入,以考虑从频谱图的信号重建的不一致性影响。最后,PCS被用来提高输入和目标特征的对比度,根据感知的重要性。在VoiceBank-DEMAND任务上进行评估,在几个客观指标上进行测试时,所提出的解决方案优于以前基于SSL的SE解决方案,达到SOTA PESQ得分3.54。摘要:Self-supervised representation learning (SSL) has attained SOTA results on several downstream speech tasks, but SSL-based speech enhancement (SE) solutions still lag behind. To address this issue, we exploit three main ideas: (i) Transformer-based masking generation, (ii) consistency-preserving loss, and (iii) perceptual contrast stretching (PCS). In detail, conformer layers, leveraging an attention mechanism, are introduced to effectively model frame-level representations and obtain the Ideal Ratio Mask (IRM) for SE. Moreover, we incorporate consistency in the loss function, which processes the input to account for the inconsistency effects of signal reconstruction from the spectrogram. Finally, PCS is employed to improve the contrast of input and target features according to perceptual importance. Evaluated on the VoiceBank-DEMAND task, the proposed solution outperforms previously SSL-based SE solutions when tested on several objective metrics, attaining a SOTA PESQ score of 3.54.
【7】 Quantifying the Corpus Bias Problem in Automatic Music Transcription Systems
标题: 自动音乐转录系统中的数据库偏差问题的量化
作者:Lukáš Samuel Marták,Patricia Hu,Gerhard Widmer
备注:2 pages, 1 figure, presented in the 1st International Workshop on Sound Signal Processing Applications (IWSSPA) 2024
链接:点击下载PDF文件
摘要:自动音乐转录(AMT)是识别音乐音频记录中的音符的任务。最先进的(SotA)基准测试一直由深度学习系统主导。由于缺乏高质量的数据,他们通常只接受或主要接受古典钢琴音乐的训练和评估。不幸的是,这阻碍了我们理解它们如何推广到其他音乐的能力。以前的工作已经揭示了这些系统中记忆和过拟合的几个方面。我们确定了两个主要来源的分布变化:音乐和声音。补充最近的结果声轴(即声学,音色),我们调查的音乐之一(即音符组合,动态,流派)。我们评估了两个新的实验测试集,我们精心构建,以模拟不同层次的音乐分布转移的几个SotA AMT系统的性能。我们的研究结果揭示了一个明显的性能差距,进一步揭示了语料库偏见问题,以及它继续困扰这些系统的程度。摘要:Automatic Music Transcription (AMT) is the task of recognizing notes in audio recordings of music. The State-of-the-Art (SotA) benchmarks have been dominated by deep learning systems. Due to the scarcity of high quality data, they are usually trained and evaluated exclusively or predominantly on classical piano music. Unfortunately, that hinders our ability to understand how they generalize to other music. Previous works have revealed several aspects of memorization and overfitting in these systems. We identify two primary sources of distribution shift: the music, and the sound. Complementing recent results on the sound axis (i.e. acoustics, timbre), we investigate the musical one (i.e. note combinations, dynamics, genre). We evaluate the performance of several SotA AMT systems on two new experimental test sets which we carefully construct to emulate different levels of musical distribution shift. Our results reveal a stark performance gap, shedding further light on the Corpus Bias problem, and the extent to which it continues to trouble these systems.
【8】 MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
标题: MulliVC:具有循环一致性的多语言语音转换
作者:Jiawei Huang,Chen Zhang,Yi Ren,Ziyue Jiang,Zhenhui Ye,Jinglin Liu,Jinzheng He,Xiang Yin,Zhou Zhao
链接:点击下载PDF文件
摘要:语音转换的目的是修改源说话者的语音,使其与目标说话者相似,同时保留原始语音内容。尽管目前语音转换取得了显着的进步,但多语言语音转换(包括单语和跨语言场景)尚未得到广泛研究。它面临两个主要挑战:1)不同语言之间韵律和发音习惯的相当大的差异; 2)来自同一说话者的配对多语言数据集的罕见性。在本文中,我们提出了MulliVC,一种新颖的语音转换系统,只转换音色,并保持原始内容和源语言的韵律没有多语言配对数据。具体来说,MulliVC的每个训练步骤包含三个子步骤:在第一步中,模型使用单语语音数据进行训练;然后,第二步和第三步从反向翻译中获得灵感,构建一个循环过程,在缺乏来自同一说话人的多语言数据的情况下解开音色和其他信息(内容,韵律和其他语言相关信息)。客观和主观的结果表明,MulliVC显着优于其他方法在单语和跨语言的背景下,证明了系统的有效性和可行性的三步方法与周期的一致性。音频样本可以在我们的演示页面(mullivc.github.io)上找到。摘要:Voice conversion aims to modify the source speaker's voice to resemble the target speaker while preserving the original speech content. Despite notable advancements in voice conversion these days, multi-lingual voice conversion (including both monolingual and cross-lingual scenarios) has yet to be extensively studied. It faces two main challenges: 1) the considerable variability in prosody and articulation habits across languages; and 2) the rarity of paired multi-lingual datasets from the same speaker. In this paper, we propose MulliVC, a novel voice conversion system that only converts timbre and keeps original content and source language prosody without multi-lingual paired data. Specifically, each training step of MulliVC contains three substeps: In step one the model is trained with monolingual speech data; then, steps two and three take inspiration from back translation, construct a cyclical process to disentangle the timbre and other information (content, prosody, and other language-related information) in the absence of multi-lingual data from the same speaker. Both objective and subjective results indicate that MulliVC significantly surpasses other methods in both monolingual and cross-lingual contexts, demonstrating the system's efficacy and the viability of the three-step approach with cycle consistency. Audio samples can be found on our demo page (mullivc.github.io).
【9】 ADD 2023: Towards Audio Deepfake Detection and Analysis in the Wild
标题: ADD 2023:走向野外音频Deepfake检测和分析
作者:Jiangyan Yi,Chu Yuan Zhang,Jianhua Tao,Chenglong Wang,Xinrui Yan,Yong Ren,Hao Gu,Junzuo Zhou
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
摘要:音频deepfake检测领域的日益突出是由于其广泛的应用,特别是在保护公众免受潜在欺诈和其他恶意活动的影响方面,这促使人们需要在这一领域进行更多的关注和研究。ADD 2023挑战超越了二进制真 假分类,通过模拟真实世界的场景,例如识别部分假音频中的操纵间隔,并确定负责生成任何假音频的源,这两者都具有现实意义,特别是在音频取证,执法以及可靠和可信证据的构建方面。为了进一步促进这一领域的研究,在这篇文章中,我们描述了在假游戏中使用的数据集,操纵区域定位和挑战的deepfake算法识别轨迹。我们还重点分析了每项任务中表现最好的参与者对技术方法的分析,并注意到他们方法的共同点和差异。最后,我们讨论了目前的技术限制,通过技术分析,并提供了未来的研究方向的路线图。数据集可供下载。摘要:The growing prominence of the field of audio deepfake detection is driven by its wide range of applications, notably in protecting the public from potential fraud and other malicious activities, prompting the need for greater attention and research in this area. The ADD 2023 challenge goes beyond binary real fake classification by emulating real-world scenarios, such as the identification of manipulated intervals in partially fake audio and determining the source responsible for generating any fake audio, both with real-life implications, notably in audio forensics, law enforcement, and construction of reliable and trustworthy evidence. To further foster research in this area, in this article, we describe the dataset that was used in the fake game, manipulation region location and deepfake algorithm recognition tracks of the challenge. We also focus on the analysis of the technical methodologies by the top-performing participants in each task and note the commonalities and differences in their approaches. Finally, we discuss the current technical limitations as identified through the technical analysis, and provide a roadmap for future research directions. The dataset is available for download.
eess.AS音频处理
【1】 ADD 2023: Towards Audio Deepfake Detection and Analysis in the Wild标题: ADD 2023:走向野外音频Deepfake检测和分析
作者:Jiangyan Yi,Chu Yuan Zhang,Jianhua Tao,Chenglong Wang,Xinrui Yan,Yong Ren,Hao Gu,Junzuo Zhou
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
摘要:音频deepfake检测领域的日益突出是由于其广泛的应用,特别是在保护公众免受潜在欺诈和其他恶意活动的影响方面,这促使人们需要在这一领域进行更多的关注和研究。ADD 2023挑战超越了二进制真 假分类,通过模拟真实世界的场景,例如识别部分假音频中的操纵间隔,并确定负责生成任何假音频的源,这两者都具有现实意义,特别是在音频取证,执法以及可靠和可信证据的构建方面。为了进一步促进这一领域的研究,在这篇文章中,我们描述了在假游戏中使用的数据集,操纵区域定位和挑战的deepfake算法识别轨迹。我们还重点分析了每项任务中表现最好的参与者对技术方法的分析,并注意到他们方法的共同点和差异。最后,我们讨论了目前的技术限制,通过技术分析,并提供了未来的研究方向的路线图。数据集可供下载。摘要:The growing prominence of the field of audio deepfake detection is driven by its wide range of applications, notably in protecting the public from potential fraud and other malicious activities, prompting the need for greater attention and research in this area. The ADD 2023 challenge goes beyond binary real fake classification by emulating real-world scenarios, such as the identification of manipulated intervals in partially fake audio and determining the source responsible for generating any fake audio, both with real-life implications, notably in audio forensics, law enforcement, and construction of reliable and trustworthy evidence. To further foster research in this area, in this article, we describe the dataset that was used in the fake game, manipulation region location and deepfake algorithm recognition tracks of the challenge. We also focus on the analysis of the technical methodologies by the top-performing participants in each task and note the commonalities and differences in their approaches. Finally, we discuss the current technical limitations as identified through the technical analysis, and provide a roadmap for future research directions. The dataset is available for download.
【2】 SELD-Mamba: Selective State-Space Model for Sound Event Localization and Detection with Source Distance Estimation
标题: SELD-Mamba:利用源距离估计进行声音事件定位和检测的选择性状态空间模型
作者:Da Mu,Zhicheng Zhang,Haobo Yue,Zehao Wang,Jin Tang,Jianqin Yin
链接:点击下载PDF文件
摘要:在声音事件定位和检测(SELD)任务中,基于Transformer的模型已经展示了令人印象深刻的功能。然而,Transformer的自注意机制的二次复杂性导致计算效率低下。在本文中,我们提出了一个网络架构SELD称为SELD-Mamba,它利用Mamba,一个选择性的状态空间模型。我们采用事件独立网络V2(EINV 2)作为基础框架,并将其Conformer块替换为双向Mamba块,以捕获更广泛的上下文信息,同时保持计算效率。此外,我们实现了一个两阶段的训练方法,第一阶段专注于声音事件检测(SED)和到达方向(DoA)估计损失,第二阶段重新引入源距离估计(DOA)损失。我们在2024 DCASE Challenge Task 3数据集上的实验结果证明了选择性状态空间模型在SELD中的有效性,并突出了两阶段训练方法在增强SELD性能方面的优势。摘要:In the Sound Event Localization and Detection (SELD) task, Transformer-based models have demonstrated impressive capabilities. However, the quadratic complexity of the Transformer's self-attention mechanism results in computational inefficiencies. In this paper, we propose a network architecture for SELD called SELD-Mamba, which utilizes Mamba, a selective state-space model. We adopt the Event-Independent Network V2 (EINV2) as the foundational framework and replace its Conformer blocks with bidirectional Mamba blocks to capture a broader range of contextual information while maintaining computational efficiency. Additionally, we implement a two-stage training method, with the first stage focusing on Sound Event Detection (SED) and Direction of Arrival (DoA) estimation losses, and the second stage reintroducing the Source Distance Estimation (SDE) loss. Our experimental results on the 2024 DCASE Challenge Task3 dataset demonstrate the effectiveness of the selective state-space model in SELD and highlight the benefits of the two-stage training approach in enhancing SELD performance.
【3】 AcousAF: Acoustic Sensing-Based Atrial Fibrillation Detection System for Mobile Phones
标题: AcoustAF:基于声学传感的手机心房颤动检测系统
作者:Xuanyu Liu,Haoxian Liu,Jiao Li,Zongqi Yang,Yi Huang,Jin Zhang
备注:Accepted for publication in Companion of the 2024 ACM International Joint Conference on Pervasive and Ubiquitous Computing (UbiComp Companion '24)
链接:点击下载PDF文件
摘要:心房颤动(AF)的特征是起源于心房的不规则电脉冲,其可导致严重的并发症甚至死亡。由于房颤的间歇性,早期及时监测房颤对于患者防止病情进一步恶化至关重要。虽然动态心电动态监测仪提供准确的监测,这些设备的高成本阻碍了他们更广泛的采用。当前基于移动设备的AF检测系统提供便携式解决方案。然而,这些系统存在各种适用性问题,例如容易受到环境因素的影响以及需要用户付出大量努力。为了克服上述限制,我们提出了AcoustAF,一种基于智能手机声学传感器的新型AF检测系统。特别是,我们探索的脉搏波采集从手腕使用智能手机扬声器和麦克风的潜力。此外,我们提出了一个精心设计的框架,包括脉搏波探测,脉搏波提取和AF检测,以确保准确和可靠的AF检测。我们使用智能手机上的自定义数据收集应用程序收集20名参与者的数据。大量的实验结果表明,我们的系统的高性能,92.8%的准确率,86.9%的精度,87.4%的召回率,和87.1%的F1分数。摘要:Atrial fibrillation (AF) is characterized by irregular electrical impulses originating in the atria, which can lead to severe complications and even death. Due to the intermittent nature of the AF, early and timely monitoring of AF is critical for patients to prevent further exacerbation of the condition. Although ambulatory ECG Holter monitors provide accurate monitoring, the high cost of these devices hinders their wider adoption. Current mobile-based AF detection systems offer a portable solution. However, these systems have various applicability issues, such as being easily affected by environmental factors and requiring significant user effort. To overcome the above limitations, we present AcousAF, a novel AF detection system based on acoustic sensors of smartphones. Particularly, we explore the potential of pulse wave acquisition from the wrist using smartphone speakers and microphones. In addition, we propose a well-designed framework comprised of pulse wave probing, pulse wave extraction, and AF detection to ensure accurate and reliable AF detection. We collect data from 20 participants utilizing our custom data collection application on the smartphone. Extensive experimental results demonstrate the high performance of our system, with 92.8% accuracy, 86.9% precision, 87.4% recall, and 87.1% F1 Score.
【4】 TEAdapter: Supply abundant guidance for controllable text-to-music generation
标题: TEAdaptor:为可控的文本到音乐生成提供丰富的指导
作者:Jialing Zou,Jiahao Mei,Xudong Nan,Jinghua Li,Daoguo Dong,Liang He
Journal-ref:2024 IEEE International Conference on Multimedia and Expo (ICME 2024)
链接:点击下载PDF文件
摘要:虽然目前的文本引导音乐生成技术可以应对简单的创意场景,但随着用户需求变得更加复杂,实现对单个文本模态条件的细粒度控制仍然具有挑战性。因此,我们介绍了TEAcher适配器(TEAdapter),一个紧凑的插件,旨在指导生成过程中,由用户提供的各种控制信息。此外,我们探讨了可控生成的扩展音乐,利用TEAdapter控制组训练的数据不同的结构功能。一般来说,我们考虑全局、元素和结构级别的控制。实验结果表明,所提出的TEAdapter,使多个精确的控制,并确保高品质的音乐生成。我们的模块也是轻量级的,可转移到任何扩散模型架构。可用的代码和演示将很快在https: github.com Ashley1101 TEAdapter上找到。摘要:Although current text-guided music generation technology can cope with simple creative scenarios, achieving fine-grained control over individual text-modality conditions remains challenging as user demands become more intricate. Accordingly, we introduce the TEAcher Adapter (TEAdapter), a compact plugin designed to guide the generation process with diverse control information provided by users. In addition, we explore the controllable generation of extended music by leveraging TEAdapter control groups trained on data of distinct structural functionalities. In general, we consider controls over global, elemental, and structural levels. Experimental results demonstrate that the proposed TEAdapter enables multiple precise controls and ensures high-quality music generation. Our module is also lightweight and transferable to any diffusion model architecture. Available code and demos will be found soon at https: github.com Ashley1101 TEAdapter.
【5】 Hyper Recurrent Neural Network: Condition Mechanisms for Black-box Audio Effect Modeling
标题: 超回归神经网络:黑匣子音效建模的条件机制
作者:Yen-Tung Yeh,Wen-Yi Hsiao,Yi-Hsuan Yang
备注:Accepted to DAFx24
链接:点击下载PDF文件
摘要:递归神经网络(RNN)在音频效果的虚拟模拟建模方面取得了令人印象深刻的结果。这些网络使用一系列矩阵乘法和非线性激活函数来处理时域音频信号,以准确地模拟目标设备的行为。为了对基于RNN的模型的旋钮的影响进行额外建模,现有方法通过将控制参数与输入信号的某种中间表示按通道连接来集成控制参数。虽然该方法是参数有效的,但是存在进一步提高所生成的音频的质量的空间,因为基于级联的调节方法在调制信号方面具有有限的能力。在本文中,我们提出了三种新的RNN调节机制,为黑盒虚拟模拟建模量身定制。这些先进的条件反射机制根据控制参数调节模型,在各种评估指标上产生优于现有基于RNN和CNN的架构的结果。摘要:Recurrent neural networks (RNNs) have demonstrated impressive results for virtual analog modeling of audio effects. These networks process time-domain audio signals using a series of matrix multiplication and nonlinear activation functions to emulate the behavior of the target device accurately. To additionally model the effect of the knobs for an RNN-based model, existing approaches integrate control parameters by concatenating them channel-wisely with some intermediate representation of the input signal. While this method is parameter-efficient, there is room to further improve the quality of generated audio because the concatenation-based conditioning method has limited capacity in modulating signals. In this paper, we propose three novel conditioning mechanisms for RNNs, tailored for black-box virtual analog modeling. These advanced conditioning mechanisms modulate the model based on control parameters, yielding superior results to existing RNN- and CNN-based architectures across various evaluation metrics.
【6】 Exploiting Consistency-Preserving Loss and Perceptual Contrast Stretching to Boost SSL-based Speech Enhancement
标题: 利用一致性保持损失和感知对比度扩展来促进基于SSL的语音增强
作者:Muhammad Salman Khan,Moreno La Quatra,Kuo-Hsuan Hung,Szu-Wei Fu,Sabato Marco Siniscalchi,Yu Tsao
链接:点击下载PDF文件
摘要:自监督表示学习(SSL)已经在几个下游语音任务上取得了SOTA结果,但基于SSL的语音增强(SE)解决方案仍然落后。为了解决这个问题,我们利用三个主要的想法:(i)基于变换的掩蔽生成,(ii)一致性保持损失,以及(iii)感知对比度拉伸(PCS)。详细地,构象层,利用注意力机制,被引入到有效地建模帧级表示,并获得理想的比例掩码(SNR)的SE。此外,我们将一致性纳入损失函数中,该函数处理输入以考虑从谱图重建信号的不一致性影响。最后,PCS被用来提高输入和目标特征的对比度,根据感知的重要性。在VoiceBank-DEMAND任务上进行评估,在几个客观指标上进行测试时,所提出的解决方案优于以前基于SSL的SE解决方案,达到SOTA PESQ得分3.54。摘要:Self-supervised representation learning (SSL) has attained SOTA results on several downstream speech tasks, but SSL-based speech enhancement (SE) solutions still lag behind. To address this issue, we exploit three main ideas: (i) Transformer-based masking generation, (ii) consistency-preserving loss, and (iii) perceptual contrast stretching (PCS). In detail, conformer layers, leveraging an attention mechanism, are introduced to effectively model frame-level representations and obtain the Ideal Ratio Mask (IRM) for SE. Moreover, we incorporate consistency in the loss function, which processes the input to account for the inconsistency effects of signal reconstruction from the spectrogram. Finally, PCS is employed to improve the contrast of input and target features according to perceptual importance. Evaluated on the VoiceBank-DEMAND task, the proposed solution outperforms previously SSL-based SE solutions when tested on several objective metrics, attaining a SOTA PESQ score of 3.54.
【7】 Quantifying the Corpus Bias Problem in Automatic Music Transcription Systems
标题: 自动音乐转录系统中的数据库偏差问题的量化
作者:Lukáš Samuel Marták,Patricia Hu,Gerhard Widmer
备注:2 pages, 1 figure, presented in the 1st International Workshop on Sound Signal Processing Applications (IWSSPA) 2024
链接:点击下载PDF文件
摘要:自动音乐转录(AMT)是识别音乐音频记录中的音符的任务。最先进的(SotA)基准测试一直由深度学习系统主导。由于缺乏高质量的数据,他们通常只接受或主要接受古典钢琴音乐的训练和评估。不幸的是,这阻碍了我们理解它们如何推广到其他音乐的能力。以前的工作已经揭示了这些系统中记忆和过拟合的几个方面。我们确定了两个主要来源的分布变化:音乐和声音。补充最近的结果声轴(即声学,音色),我们调查的音乐之一(即音符组合,动态,流派)。我们评估了两个新的实验测试集,我们精心构建,以模拟不同层次的音乐分布转移的几个SotA AMT系统的性能。我们的研究结果揭示了一个明显的性能差距,进一步揭示了语料库偏见问题,以及它继续困扰这些系统的程度。摘要:Automatic Music Transcription (AMT) is the task of recognizing notes in audio recordings of music. The State-of-the-Art (SotA) benchmarks have been dominated by deep learning systems. Due to the scarcity of high quality data, they are usually trained and evaluated exclusively or predominantly on classical piano music. Unfortunately, that hinders our ability to understand how they generalize to other music. Previous works have revealed several aspects of memorization and overfitting in these systems. We identify two primary sources of distribution shift: the music, and the sound. Complementing recent results on the sound axis (i.e. acoustics, timbre), we investigate the musical one (i.e. note combinations, dynamics, genre). We evaluate the performance of several SotA AMT systems on two new experimental test sets which we carefully construct to emulate different levels of musical distribution shift. Our results reveal a stark performance gap, shedding further light on the Corpus Bias problem, and the extent to which it continues to trouble these systems.
【8】 MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
标题: MulliVC:具有循环一致性的多语言语音转换
作者:Jiawei Huang,Chen Zhang,Yi Ren,Ziyue Jiang,Zhenhui Ye,Jinglin Liu,Jinzheng He,Xiang Yin,Zhou Zhao
链接:点击下载PDF文件
摘要:语音转换的目的是修改源说话者的语音,使其与目标说话者相似,同时保留原始语音内容。尽管目前语音转换取得了显着的进步,但多语言语音转换(包括单语和跨语言场景)尚未得到广泛研究。它面临两个主要挑战:1)不同语言之间韵律和发音习惯的相当大的差异; 2)来自同一说话者的配对多语言数据集的罕见性。在本文中,我们提出了MulliVC,一种新颖的语音转换系统,只转换音色,并保持原始内容和源语言的韵律没有多语言配对数据。具体来说,MulliVC的每个训练步骤包含三个子步骤:在第一步中,模型使用单语语音数据进行训练;然后,第二步和第三步从反向翻译中获得灵感,构建一个循环过程,在缺乏来自同一说话人的多语言数据的情况下解开音色和其他信息(内容,韵律和其他语言相关信息)。客观和主观的结果表明,MulliVC显着优于其他方法在单语和跨语言的背景下,证明了系统的有效性和可行性的三步方法与周期的一致性。音频样本可以在我们的演示页面(mullivc.github.io)上找到。摘要:Voice conversion aims to modify the source speaker's voice to resemble the target speaker while preserving the original speech content. Despite notable advancements in voice conversion these days, multi-lingual voice conversion (including both monolingual and cross-lingual scenarios) has yet to be extensively studied. It faces two main challenges: 1) the considerable variability in prosody and articulation habits across languages; and 2) the rarity of paired multi-lingual datasets from the same speaker. In this paper, we propose MulliVC, a novel voice conversion system that only converts timbre and keeps original content and source language prosody without multi-lingual paired data. Specifically, each training step of MulliVC contains three substeps: In step one the model is trained with monolingual speech data; then, steps two and three take inspiration from back translation, construct a cyclical process to disentangle the timbre and other information (content, prosody, and other language-related information) in the absence of multi-lingual data from the same speaker. Both objective and subjective results indicate that MulliVC significantly surpasses other methods in both monolingual and cross-lingual contexts, demonstrating the system's efficacy and the viability of the three-step approach with cycle consistency. Audio samples can be found on our demo page (mullivc.github.io).
【9】 Abstractive summarization from Audio Transcription
标题: 音频转录的抽象摘要
作者:Ilia Derkach
备注:36 pages, Master's thesis, 14 figures
链接:点击下载PDF文件
摘要:目前,大型语言模型越来越受欢迎,其成果被用于许多领域,从文本翻译到生成查询答案。然而,这些新的机器学习算法的主要问题是,训练这样的模型需要大量的计算资源,而这些资源只有大型IT公司才拥有。为了避免这个问题,已经提出了许多方法(LoRA,量化),以便可以针对特定任务有效地微调现有模型。在本文中,我们提出了一个E2E(端到端)的音频摘要模型,使用这些技术。此外,本文还研究了这些方法的有效性所考虑的问题,并得出结论,这些方法的适用性。摘要:Currently, large language models are gaining popularity, their achievements are used in many areas, ranging from text translation to generating answers to queries. However, the main problem with these new machine learning algorithms is that training such models requires large computing resources that only large IT companies have. To avoid this problem, a number of methods (LoRA, quantization) have been proposed so that existing models can be effectively fine-tuned for specific tasks. In this paper, we propose an E2E (end to end) audio summarization model using these techniques. In addition, this paper examines the effectiveness of these approaches to the problem under consideration and draws conclusions about the applicability of these methods.
机器翻译,仅供参考
