今日论文合集:cs.SD语音4篇,eess.AS音频处理7篇。本文经arXiv每日学术速递授权转载
【1】An LSTM-Based Chord Generation System Using Chroma Histogram Representations
备注:6 pages, 4 figures, 1 table摘要:本文提出了一个系统的和弦生成单声道符号旋律使用基于LSTM的模型训练的色度直方图表示的和弦。色度表示承诺比基于和弦标签的方法更和谐丰富的生成,同时在数据集中保持少量的维度。该系统被证明是适合于有限的实时使用。虽然它不符合连贯的长期生成的最新技术水平,但它确实显示了具有韵律和弦关系的全音阶生成。需要进一步研究色度直方图作为和弦生成任务中提取的特征。摘要:This paper proposes a system for chord generation to monophonic symbolic melodies using an LSTM-based model trained on chroma histogram representations of chords. Chroma representations promise more harmonically rich generation than chord label-based approaches, whilst maintaining a small number of dimensions in the dataset. This system is shown to be suitable for limited real-time use. While it does not meet the state-of-the-art for coherent long-term generation, it does show diatonic generation with cadential chord relationships. The need for further study into chroma histograms as an extracted feature in chord generation tasks is highlighted.
【2】 Exploring Speech Pattern Disorders in Autism using Machine Learning作者:Chuanbo Hu,Jacob Thrasher,Wenqi Li,Mindi Ruan,Xiangxu Yu,Lynn K Paul,Shuo Wang,Xin Li摘要:通过从检查者-患者对话中识别异常言语模式来诊断自闭症谱系障碍(ASD),由于受影响个体中言语相关症状的微妙和多样的表现而提出了重大挑战。本研究提出了一个全面的方法来识别独特的语音模式,通过分析检查者-病人的对话。利用记录对话的数据集,我们提取了40个语音相关特征,分类为频率,过零率,能量,频谱特征,梅尔频率倒谱系数(MFCC)和平衡。这些特征包括语音的各个方面,如语调,音量,节奏和语速,反映了ASD交际行为的复杂性。我们采用机器学习进行分类和回归任务来分析这些语音特征。该分类模型旨在区分ASD和非ASD病例,准确率为87.75%。回归模型被开发来预测语音模式相关变量和来自所有变量的综合得分,促进对与ASD相关的语音动态的更深入的理解。机器学习在解释复杂语音模式方面的有效性和高分类准确性强调了计算方法在支持ASD诊断过程中的潜力。这种方法不仅有助于早期检测,而且还有助于通过提供对ASD患者的语音和沟通概况的见解来制定个性化的治疗计划。摘要:Diagnosing autism spectrum disorder (ASD) by identifying abnormal speech patterns from examiner-patient dialogues presents significant challenges due to the subtle and diverse manifestations of speech-related symptoms in affected individuals. This study presents a comprehensive approach to identify distinctive speech patterns through the analysis of examiner-patient dialogues. Utilizing a dataset of recorded dialogues, we extracted 40 speech-related features, categorized into frequency, zero-crossing rate, energy, spectral characteristics, Mel Frequency Cepstral Coefficients (MFCCs), and balance. These features encompass various aspects of speech such as intonation, volume, rhythm, and speech rate, reflecting the complex nature of communicative behaviors in ASD. We employed machine learning for both classification and regression tasks to analyze these speech features. The classification model aimed to differentiate between ASD and non-ASD cases, achieving an accuracy of 87.75%. Regression models were developed to predict speech pattern related variables and a composite score from all variables, facilitating a deeper understanding of the speech dynamics associated with ASD. The effectiveness of machine learning in interpreting intricate speech patterns and the high classification accuracy underscore the potential of computational methods in supporting the diagnostic processes for ASD. This approach not only aids in early detection but also contributes to personalized treatment planning by providing insights into the speech and communication profiles of individuals with ASD.【3】 The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio标题:Deepfake音频普遍检测的Codecfake数据集和对策作者:Yuankun Xie,Yi Lu,Ruibo Fu,Zhengqi Wen,Zhiyong Wang,Jianhua Tao,Xin Qi,Xiaopeng Wang,Yukun Liu,Haonan Cheng,Long Ye,Yi Sun摘要:随着基于音频语言模型(ALM)的deepfake音频的激增,迫切需要有效的检测方法。与传统的deepfake音频生成不同,它通常涉及最终使用声码器的多步过程,ALM直接利用神经编解码器方法将离散代码解码为音频。此外,在大规模数据的驱动下,ALM表现出显着的鲁棒性和通用性,对当前的音频深度伪造检测(ADD)模型构成了重大挑战。为了有效地检测基于ALM的deepfake音频,我们重点研究了基于ALM的音频生成方法的机制,即从神经编解码器到波形的转换。我们首先构建Codecfake数据集,这是一个开源的大规模数据集,包括两种语言,数百万个音频样本和各种测试条件,为基于ALM的音频检测量身定制。此外,为了实现deepfake音频的通用检测并解决原始SAM的域上升偏差问题,我们提出了CSAM策略来学习域平衡和广义最小值。实验结果表明,与基线模型相比,在Codecfake数据集和Vocoded数据集上使用CSAM策略进行联合训练,在所有测试条件下的平均等错误率(EER)最低,为0.616%。摘要:With the proliferation of Audio Language Model (ALM) based deepfake audio, there is an urgent need for effective detection methods. Unlike traditional deepfake audio generation, which often involves multi-step processes culminating in vocoder usage, ALM directly utilizes neural codec methods to decode discrete codes into audio. Moreover, driven by large-scale data, ALMs exhibit remarkable robustness and versatility, posing a significant challenge to current audio deepfake detection (ADD) models. To effectively detect ALM-based deepfake audio, we focus on the mechanism of the ALM-based audio generation method, the conversion from neural codec to waveform. We initially construct the Codecfake dataset, an open-source large-scale dataset, including two languages, millions of audio samples, and various test conditions, tailored for ALM-based audio detection. Additionally, to achieve universal detection of deepfake audio and tackle domain ascent bias issue of original SAM, we propose the CSAM strategy to learn a domain balanced and generalized minima. Experiment results demonstrate that co-training on Codecfake dataset and vocoded dataset with CSAM strategy yield the lowest average Equal Error Rate (EER) of 0.616% across all test conditions compared to baseline models.【4】 SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge Evaluation Plan标题:SDDD挑战2024:歌声Deepfake检测挑战评估计划作者:You Zhang,Yongyi Zang,Jiatong Shi,Ryuichi Yamamoto,Jionghao Han,Yuxun Tang,Tomoki Toda,Zhiyao Duan备注:Evaluation plan of the SVDD Challenge @ SLT 2024摘要:人工智能生成的歌声的快速发展,现在可以很好地模仿自然的人类歌唱,并与乐谱无缝匹配,这引起了艺术家和音乐行业的高度关注。与口语不同,由于其音乐性质和强烈背景音乐的存在,歌声提出了独特的挑战,使歌声深度假检测(SVDD)成为需要重点关注的专业领域。为了促进SVDD研究,我们最近提出了“SVDD挑战赛”,这是第一个专注于实验室控制和野外真实和deepfake唱歌录音的SVDD研究挑战赛。该挑战赛将与2024年IEEE口语技术研讨会(IEEE Spoken Language Technology Workshop,简称IEEE2024)一起举办。摘要:The rapid advancement of AI-generated singing voices, which now closely mimic natural human singing and align seamlessly with musical scores, has led to heightened concerns for artists and the music industry. Unlike spoken voice, singing voice presents unique challenges due to its musical nature and the presence of strong background music, making singing voice deepfake detection (SVDD) a specialized field requiring focused attention. To promote SVDD research, we recently proposed the "SVDD Challenge," the very first research challenge focusing on SVDD for lab-controlled and in-the-wild bonafide and deepfake singing voice recordings. The challenge will be held in conjunction with the 2024 IEEE Spoken Language Technology Workshop (SLT 2024).【1】 SVDD Challenge 2024: A Singing Voice Deepfake Detection Challenge Evaluation Plan标题:SDDD挑战2024:歌声Deepfake检测挑战评估计划作者:You Zhang,Yongyi Zang,Jiatong Shi,Ryuichi Yamamoto,Jionghao Han,Yuxun Tang,Tomoki Toda,Zhiyao Duan备注:Evaluation plan of the SVDD Challenge @ SLT 2024摘要:人工智能生成的歌声的快速发展,现在可以很好地模仿自然的人类歌唱,并与乐谱无缝匹配,这引起了艺术家和音乐行业的高度关注。与口语不同,由于其音乐性质和强烈背景音乐的存在,歌声提出了独特的挑战,使歌声深度假检测(SVDD)成为需要重点关注的专业领域。为了促进SVDD研究,我们最近提出了“SVDD挑战赛”,这是第一个专注于实验室控制和野外真实和deepfake唱歌录音的SVDD研究挑战赛。该挑战赛将与2024年IEEE口语技术研讨会(IEEE Spoken Language Technology Workshop,简称IEEE2024)一起举办。摘要:The rapid advancement of AI-generated singing voices, which now closely mimic natural human singing and align seamlessly with musical scores, has led to heightened concerns for artists and the music industry. Unlike spoken voice, singing voice presents unique challenges due to its musical nature and the presence of strong background music, making singing voice deepfake detection (SVDD) a specialized field requiring focused attention. To promote SVDD research, we recently proposed the "SVDD Challenge," the very first research challenge focusing on SVDD for lab-controlled and in-the-wild bonafide and deepfake singing voice recordings. The challenge will be held in conjunction with the 2024 IEEE Spoken Language Technology Workshop (SLT 2024).
【2】 HILCodec: High Fidelity and Lightweight Neural Audio Codec标题:HILCodec:高保真和轻量级神经音频编解码器作者:Sunghwan Ahn,Beom Jun Woo,Min Hyun Han,Chanyeong Moon,Nam Soo Kim摘要:端到端神经音频编解码器的最新进展使得能够以非常低的比特率压缩音频,同时以高保真度重建输出音频。尽管如此,这种改进往往是以增加模型复杂性为代价的。在本文中,我们确定和解决现有的神经音频编解码器的问题。我们发现,Wave—U—Net的性能并不随着网络深度的增加而持续增加。我们分析了这种现象的根本原因,并提出了方差约束设计。此外,我们揭示了各种失真在以前的波形域鉴别器,并提出了一种新的无失真鉴别器。生成的模型\textit {HILCodec}是一个实时流音频编解码器,可在各种比特率和音频类型中展示最先进的质量。摘要:The recent advancement of end-to-end neural audio codecs enables compressing audio at very low bitrates while reconstructing the output audio with high fidelity. Nonetheless, such improvements often come at the cost of increased model complexity. In this paper, we identify and address the problems of existing neural audio codecs. We show that the performance of Wave-U-Net does not increase consistently as the network depth increases. We analyze the root cause of such a phenomenon and suggest a variance-constrained design. Also, we reveal various distortions in previous waveform domain discriminators and propose a novel distortion-free discriminator. The resulting model, \textit{HILCodec}, is a real-time streaming audio codec that demonstrates state-of-the-art quality across various bitrates and audio types.
【3】 SingIt! Singer Voice Transformation作者:Amit Eliav,Aaron Taub,Renana Opochinsky,Sharon Gannot摘要:在本文中,我们提出了一个模型,可以产生一个歌唱的声音从正常的语音话语利用zero-shot,多对多风格的迁移学习。我们的目标是让任何人都有机会及时唱任何歌曲。我们提出了一个系统,包括几个可用的块,以及一个修改后的自动编码器,并显示如何通过定制而简单的解决方案一起实现这个高度复杂的挑战。我们证明了所提出的系统使用一组25个非专家听众的适用性。提供了从我们的模型生成的数据样本。摘要:In this paper, we propose a model which can generate a singing voice from normal speech utterance by harnessing zero-shot, many-to-many style transfer learning. Our goal is to give anyone the opportunity to sing any song in a timely manner. We present a system comprising several available blocks, as well as a modified auto-encoder, and show how this highly-complex challenge can be achieved by tailoring rather simple solutions together. We demonstrate the applicability of the proposed system using a group of 25 non-expert listeners. Samples of the data generated from our model are provided.【4】 An LSTM-Based Chord Generation System Using Chroma Histogram Representations备注:6 pages, 4 figures, 1 table摘要:本文提出了一个系统的和弦生成单声道符号旋律使用基于LSTM的模型训练的色度直方图表示的和弦。色度表示承诺比基于和弦标签的方法更和谐丰富的生成,同时在数据集中保持少量的维度。该系统被证明是适合于有限的实时使用。虽然它不符合连贯的长期生成的最新技术水平,但它确实显示了具有韵律和弦关系的全音阶生成。需要进一步研究色度直方图作为和弦生成任务中提取的特征。摘要:This paper proposes a system for chord generation to monophonic symbolic melodies using an LSTM-based model trained on chroma histogram representations of chords. Chroma representations promise more harmonically rich generation than chord label-based approaches, whilst maintaining a small number of dimensions in the dataset. This system is shown to be suitable for limited real-time use. While it does not meet the state-of-the-art for coherent long-term generation, it does show diatonic generation with cadential chord relationships. The need for further study into chroma histograms as an extracted feature in chord generation tasks is highlighted.【5】 Exploring Speech Pattern Disorders in Autism using Machine Learning作者:Chuanbo Hu,Jacob Thrasher,Wenqi Li,Mindi Ruan,Xiangxu Yu,Lynn K Paul,Shuo Wang,Xin Li摘要:通过从检查者—患者对话中识别异常言语模式来诊断自闭症谱系障碍(ASD),由于受影响个体中言语相关症状的微妙和多样的表现而提出了重大挑战。本研究提出了一个全面的方法来识别独特的语音模式,通过分析检查者—病人的对话。利用记录对话的数据集,我们提取了40个语音相关特征,分类为频率,过零率,能量,频谱特征,梅尔频率倒谱系数(MFCC)和平衡。这些特征包括语音的各个方面,如语调,音量,节奏和语速,反映了ASD交际行为的复杂性。我们采用机器学习进行分类和回归任务来分析这些语音特征。该分类模型旨在区分ASD和非ASD病例,准确率为87.75%。回归模型被开发来预测语音模式相关变量和来自所有变量的综合得分,促进对与ASD相关的语音动态的更深入的理解。机器学习在解释复杂语音模式方面的有效性和高分类准确性强调了计算方法在支持ASD诊断过程中的潜力。这种方法不仅有助于早期检测,而且还有助于通过提供对ASD患者的语音和沟通概况的见解来制定个性化的治疗计划。摘要:Diagnosing autism spectrum disorder (ASD) by identifying abnormal speech patterns from examiner-patient dialogues presents significant challenges due to the subtle and diverse manifestations of speech-related symptoms in affected individuals. This study presents a comprehensive approach to identify distinctive speech patterns through the analysis of examiner-patient dialogues. Utilizing a dataset of recorded dialogues, we extracted 40 speech-related features, categorized into frequency, zero-crossing rate, energy, spectral characteristics, Mel Frequency Cepstral Coefficients (MFCCs), and balance. These features encompass various aspects of speech such as intonation, volume, rhythm, and speech rate, reflecting the complex nature of communicative behaviors in ASD. We employed machine learning for both classification and regression tasks to analyze these speech features. The classification model aimed to differentiate between ASD and non-ASD cases, achieving an accuracy of 87.75%. Regression models were developed to predict speech pattern related variables and a composite score from all variables, facilitating a deeper understanding of the speech dynamics associated with ASD. The effectiveness of machine learning in interpreting intricate speech patterns and the high classification accuracy underscore the potential of computational methods in supporting the diagnostic processes for ASD. This approach not only aids in early detection but also contributes to personalized treatment planning by providing insights into the speech and communication profiles of individuals with ASD.【6】 The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio标题:Deepfake音频普遍检测的Codecfake数据集和对策作者:Yuankun Xie,Yi Lu,Ruibo Fu,Zhengqi Wen,Zhiyong Wang,Jianhua Tao,Xin Qi,Xiaopeng Wang,Yukun Liu,Haonan Cheng,Long Ye,Yi Sun摘要:随着基于音频语言模型(ALM)的deepfake音频的激增,迫切需要有效的检测方法。与传统的deepfake音频生成不同,它通常涉及最终使用声码器的多步过程,ALM直接利用神经编解码器方法将离散代码解码为音频。此外,在大规模数据的驱动下,ALM表现出显着的鲁棒性和通用性,对当前的音频深度伪造检测(ADD)模型构成了重大挑战。为了有效地检测基于ALM的deepfake音频,我们重点研究了基于ALM的音频生成方法的机制,即从神经编解码器到波形的转换。我们首先构建Codecfake数据集,这是一个开源的大规模数据集,包括两种语言,数百万个音频样本和各种测试条件,为基于ALM的音频检测量身定制。此外,为了实现deepfake音频的通用检测并解决原始SAM的域上升偏差问题,我们提出了CSAM策略来学习域平衡和广义最小值。实验结果表明,与基线模型相比,在Codecfake数据集和Vocoded数据集上使用CSAM策略进行联合训练,在所有测试条件下的平均等错误率(EER)最低,为0.616%。摘要:With the proliferation of Audio Language Model (ALM) based deepfake audio, there is an urgent need for effective detection methods. Unlike traditional deepfake audio generation, which often involves multi-step processes culminating in vocoder usage, ALM directly utilizes neural codec methods to decode discrete codes into audio. Moreover, driven by large-scale data, ALMs exhibit remarkable robustness and versatility, posing a significant challenge to current audio deepfake detection (ADD) models. To effectively detect ALM-based deepfake audio, we focus on the mechanism of the ALM-based audio generation method, the conversion from neural codec to waveform. We initially construct the Codecfake dataset, an open-source large-scale dataset, including two languages, millions of audio samples, and various test conditions, tailored for ALM-based audio detection. Additionally, to achieve universal detection of deepfake audio and tackle domain ascent bias issue of original SAM, we propose the CSAM strategy to learn a domain balanced and generalized minima. Experiment results demonstrate that co-training on Codecfake dataset and vocoded dataset with CSAM strategy yield the lowest average Equal Error Rate (EER) of 0.616% across all test conditions compared to baseline models.【7】 Dichotic harmony for the musical practice备注:14 pages, in Russian, links added摘要:听声音的两耳分听法适用于音乐和声的区域。提出了一种将不和谐音分离成若干组的算法。为了增加和弦的愉悦感,通过耳机的不同通道听到不同的声音。是为PC创建的两个演示程序。关键词:音乐,和声,和弦,双耳分听,不协和音,协和音,耳机,愉悦感,midi。摘要:The dichotic method of hearing sound adapts in the region of musical harmony. The algorithm of the separation of the being dissonant voices into several separate groups is proposed. For an increase in the pleasantness of chords the different groups of voices are heard out through the different channels of headphones. Is created two demonstration program for PC. Keywords: music, harmony, chord, dichotic listening, dissonance, consonance, headphones, pleasantness, midi.