今日论文合集:cs.SD语音11篇,eess.AS音频处理16篇。本文经arXiv每日学术速递授权转载
【1】A Novel Audio Representation for Music Genre Identification in MIR
标题:一种新的MIR音乐流派识别音频表示方法
链接:https://arxiv.org/abs/2404.01058
作者:Navin Kamuni,Mayank Jindal,Arpita Soni,Sukender Reddy Mallreddy,Sharath Chandra Macha
摘要:对于音乐信息检索下游任务,最常见的音频表示是基于时间-频率的,例如Mel频谱图。为了识别音乐类型,本研究探讨了一种新形式的音频表示的可能性,最常见的MIR下游任务之一。因此,为了使用深度矢量量化对音乐进行离散编码,为创新的生成音乐模型(即Juan)创建了一种新的音频表示。使用几乎等同于最新技术水平(SOTA)的数据集和几乎相同的Transformer设计,将Jupiter的音频表示的有效性与Mel频谱图进行比较。这项研究的结果意味着,至少当Transformers使用20 k音轨的非常适度的数据集进行预训练时,Juvente的音频表示并不优于Mel频谱图。这可能是因为Jujant的音频表示没有充分考虑到人类听觉感知的特殊性。另一方面,Mel声谱图是专门根据人类听觉创建的。
摘要:For Music Information Retrieval downstream tasks, the most common audio representation is time-frequency-based, such as Mel spectrograms. In order to identify musical genres, this study explores the possibilities of a new form of audio representation one of the most usual MIR downstream tasks. Therefore, to discretely encoding music using deep vector quantization; a novel audio representation was created for the innovative generative music model i.e. Jukebox. The effectiveness of Jukebox's audio representation is compared to Mel spectrograms using a dataset that is almost equivalent to State-of-the-Art (SOTA) and an almost same transformer design. The results of this study imply that, at least when the transformers are pretrained using a very modest dataset of 20k tracks, Jukebox's audio representation is not superior to Mel spectrograms. This could be explained by the fact that Jukebox's audio representation does not sufficiently take into account the peculiarities of human hearing perception. On the other hand, Mel spectrograms are specifically created with the human auditory sense in mind.
【2】 360+x: A Panoptic Multi-modal Scene Understanding Dataset作者:Hao Chen,Yuqi Hou,Chenyuan Qu,Irene Testini,Xiaohan Hong,Jianbo Jiao摘要:人类对世界的感知是由多种观点和模式塑造的。虽然许多现有的数据集专注于从某个角度(例如自我中心或第三人称视图)进行场景理解,但我们的数据集提供了全景视角(即具有多种数据模态的多个视角)。具体来说,我们封装第三人称全景和前视图,以及以自我为中心的单目/双目视图,具有丰富的模态,包括视频,多通道音频,定向双耳延迟,位置数据和每个场景中捕获的文本场景描述,呈现对世界的全面观察。图1提供了我们360+x数据集的所有28个场景类别的一瞥。据我们所知,这是第一个涵盖多个观点和多种数据模式的数据库,以模拟现实世界中如何访问日常信息。通过我们的基准分析,我们在建议的360+x数据集上提出了5个不同的场景理解任务,以评估每个数据模态和视角在全景场景理解中的影响和好处。我们希望这个独特的数据集可以扩大全面了解场景的范围,并鼓励社区从更多样化的角度来处理这些问题。摘要:Human perception of the world is shaped by a multitude of viewpoints and modalities. While many existing datasets focus on scene understanding from a certain perspective (e.g. egocentric or third-person views), our dataset offers a panoptic perspective (i.e. multiple viewpoints with multiple data modalities). Specifically, we encapsulate third-person panoramic and front views, as well as egocentric monocular/binocular views with rich modalities including video, multi-channel audio, directional binaural delay, location data and textual scene descriptions within each scene captured, presenting comprehensive observation of the world. Figure 1 offers a glimpse of all 28 scene categories of our 360+x dataset. To the best of our knowledge, this is the first database that covers multiple viewpoints with multiple data modalities to mimic how daily information is accessed in the real world. Through our benchmark analysis, we presented 5 different scene understanding tasks on the proposed 360+x dataset to evaluate the impact and benefit of each data modality and perspective in panoptic scene understanding. We hope this unique dataset could broaden the scope of comprehensive scene understanding and encourage the community to approach these problems from more diverse perspectives.【3】 Removing Speaker Information from Speech Representation using Variable-Length Soft Pooling标题:基于变长软池的语音表示去除说话人信息作者:Injune Hwang,Kyogu Lee摘要:最近,已经有人努力使用语音合成的自监督框架来编码语音的语言信息。然而,从周围表示预测表示可能无意中使说话者信息纠缠在语音表示中。本文旨在通过利用语音的结构化特性来去除说话人信息,语音由具有清晰边界的音素等离散单元组成。神经网络预测这些边界,使可变长度池化用于基于事件的表示提取,而不是固定速率方法。边界预测器输出0和1之间的边界的概率,使池化变软。该模型经过训练,以最大限度地减少与通过时间拉伸和音高偏移增强的数据的合并表示的差异。为了证实所学习的表征包括内容信息但独立于说话人信息,用libri-light的语音ABX任务和SUPERB的说话人识别任务对模型进行了评估。摘要:Recently, there have been efforts to encode the linguistic information of speech using a self-supervised framework for speech synthesis. However, predicting representations from surrounding representations can inadvertently entangle speaker information in the speech representation. This paper aims to remove speaker information by exploiting the structured nature of speech, composed of discrete units like phonemes with clear boundaries. A neural network predicts these boundaries, enabling variable-length pooling for event-based representation extraction instead of fixed-rate methods. The boundary predictor outputs a probability for the boundary between 0 and 1, making pooling soft. The model is trained to minimize the difference with the pooled representation of the data augmented by time-stretch and pitch-shift. To confirm that the learned representation includes contents information but is independent of speaker information, the model was evaluated with libri-light's phonetic ABX task and SUPERB's speaker identification task.【4】 Personalized Neural Speech Codec作者:Inseon Jang,Haici Yang,Wootaek Lim,Seungkwon Beack,Minje Kim摘要:在本文中,我们提出了一个个性化的神经语音编解码器,设想个性化可以降低模型的复杂性或提高感知语音质量。尽管语音编解码器的常见用法是在通信的每一侧仅涉及单个说话者,但是在文献中很少探索针对特定用户的个性化编解码器。首先,我们假设扬声器可以分组为较小的子集的基础上,他们的感知相似性。然后,我们还假设,一个特定于组的编解码器可以专注于组的语音特性,以提高其感知质量和计算效率。为此,我们首先开发了一个Siamese网络,该网络从LibriSpeech数据集中学习说话人嵌入,然后将其分组为底层说话人集群。最后,我们在每个扬声器集群上重新训练基于LPCNet的语音编解码器基线。主观听力测试表明,所提出的个性化方案,同时保持语音质量的模型压缩。换句话说,在相同的模型复杂度下,个性化编解码器产生更好的语音质量。摘要:In this paper, we propose a personalized neural speech codec, envisioning that personalization can reduce the model complexity or improve perceptual speech quality. Despite the common usage of speech codecs where only a single talker is involved on each side of the communication, personalizing a codec for the specific user has rarely been explored in the literature. First, we assume speakers can be grouped into smaller subsets based on their perceptual similarity. Then, we also postulate that a group-specific codec can focus on the group's speech characteristics to improve its perceptual quality and computational efficiency. To this end, we first develop a Siamese network that learns the speaker embeddings from the LibriSpeech dataset, which are then grouped into underlying speaker clusters. Finally, we retrain the LPCNet-based speech codec baselines on each of the speaker clusters. Subjective listening tests show that the proposed personalization scheme introduces model compression while maintaining speech quality. In other words, with the same model complexity, personalized codecs produce better speech quality.【5】 A Comparative Analysis of Poetry Reading Audio: Singing, Narrating, or Somewhere In Between?标题:诗歌阅读音频的比较分析:歌唱、叙事还是介于两者之间?摘要:本文提供了一个大规模的诗歌朗读音频信号的计算分析,以揭示专业朗读诗歌中的音乐性。虽然其他类型的口语的声学特征已被广泛研究,但大多数文献仅限于叙事语音或歌声,讨论它们彼此之间的差异。在这项工作中,我们开发的信号处理方法,这是量身定制的,以捕捉独特的声学特性的诗歌阅读的基础上,他们的沉默模式,时间变化的本地音高,和节拍的稳定性。我们对三个大语料库的大规模统计分析,每个语料库都由叙述(LibriSpeech),歌唱声音(Intonation)和诗歌阅读(来自The Poetry Foundation)组成,发现诗歌阅读确实与歌唱声音共享一些音乐特征,尽管它也可能类似于叙述性演讲。摘要:This paper provides a computational analysis of poetry reading audio signals at a large scale to unveil the musicality within professionally-read poems. Although the acoustic characteristics of other types of spoken language have been extensively studied, most of the literature is limited to narrative speech or singing voice, discussing how different they are from each other. In this work, we develop signal processing methods, which are tailored to capture the unique acoustic characteristics of poetry reading based on their silence patterns, temporal variations of local pitch, and beat stability. Our large-scale statistical analyses on three big corpora, each of which consists of narration (LibriSpeech), singing voice (Intonation), and poetry reading (from The Poetry Foundation), discover that poetry reading does share some musical characteristics with singing voice, although it may also resemble narrative speech.【6】 Measuring audio prompt adherence with distribution-based embedding distances摘要:越来越多的生成音乐模型可以以音频提示为条件,该音频提示用作模型要为其创建伴奏的音乐上下文(通常使用文本提示进一步指定)。对模型输出如何遵守音频提示的评估通常以模型或问题特定的方式进行,大概是因为还没有出现用于音频提示遵守的通用评估方法。这种方法在新模型的开发和训练中都很有用,并且可以使模型之间的性能具有可比性。在本文中,我们调查是否常用的分布为基础的距离,如Fr\'echet音频距离(FAD),可以用来衡量音频提示遵守。我们提出了一个简单的程序,基于少量的成分(嵌入模型,投影,嵌入距离和数据融合方法),我们系统地评估使用基线验证。在后续的实验中,我们测试的灵敏度,建议的音频坚持措施,音调和时移扰动。结果表明,所提出的措施是敏感的扰动,即使参考和候选分布是来自不同的音乐收藏。虽然需要更多的实验来回答未解决的问题,如测量不影响音频提示遵守的声学伪影的鲁棒性,但目前的结果表明,基于分布的嵌入距离提供了一种可行的方法来测量音频提示遵守。建议措施的python/pytorch实现作为github存储库公开提供。摘要:An increasing number of generative music models can be conditioned on an audio prompt that serves as musical context for which the model is to create an accompaniment (often further specified using a text prompt). Evaluation of how well model outputs adhere to the audio prompt is often done in a model or problem specific manner, presumably because no generic evaluation method for audio prompt adherence has emerged. Such a method could be useful both in the development and training of new models, and to make performance comparable across models. In this paper we investigate whether commonly used distribution-based distances like Fr\'echet Audio Distance (FAD), can be used to measure audio prompt adherence. We propose a simple procedure based on a small number of constituents (an embedding model, a projection, an embedding distance, and a data fusion method), that we systematically assess using a baseline validation. In a follow-up experiment we test the sensitivity of the proposed audio adherence measure to pitch and time shift perturbations. The results show that the proposed measure is sensitive to such perturbations, even when the reference and candidate distributions are from different music collections. Although more experimentation is needed to answer unaddressed questions like the robustness of the measure to acoustic artifacts that do not affect the audio prompt adherence, the current results suggest that distribution-based embedding distances provide a viable way of measuring audio prompt adherence. An python/pytorch implementation of the proposed measure is publicly available as a github repository.
【7】 WavLLM: Towards Robust and Adaptive Speech Large Language Model作者:Shujie Hu,Long Zhou,Shujie Liu,Sanyuan Chen,Hongkun Hao,Jing Pan,Xunying Liu,Jinyu Li,Sunit Sivasankaran,Linquan Liu,Furu Wei摘要:大型语言模型(LLM)的最新进展彻底改变了自然语言处理领域,逐步将其范围扩大到多模态感知和生成。然而,有效地将听力能力整合到LLM中带来了重大挑战,特别是在不同背景下的概括和执行复杂的听觉任务方面。在这项工作中,我们介绍了WavLLM,一个强大的和自适应的语音大语言模型与双编码器,和一个可感知的LoRA权重适配器,优化了两阶段的课程学习方法。利用双编码器,我们解耦不同类型的语音信息,利用Whisper编码器来处理语音的语义内容,并利用WavLM编码器来捕获说话者身份的独特特征。在课程学习框架内,WavLLM首先通过优化混合的基本单一任务来构建其基础能力,然后对更复杂的任务进行高级多任务培训,例如基本任务的组合。为了提高灵活性和对不同任务和指令的坚持,在第二个高级多任务训练阶段引入了一个可感知的LoRA权重适配器。我们验证了所提出的模型的通用语音基准,包括任务,如ASR,ST,SV,ER,并将其应用到专门的数据集,如高考英语听力理解集SQA,和语音思维链(CoT)评估集。实验表明,该模型在相同的模型大小的语音任务范围内实现了最先进的性能,表现出强大的泛化能力,在执行复杂的任务,使用CoT方法。此外,我们的模型在没有专门训练的情况下成功完成了高考任务。代码、模型、音频和高考评估集可在\url{aka.ms/wavllm}访问。摘要:The recent advancements in large language models (LLMs) have revolutionized the field of natural language processing, progressively broadening their scope to multimodal perception and generation. However, effectively integrating listening capabilities into LLMs poses significant challenges, particularly with respect to generalizing across varied contexts and executing complex auditory tasks. In this work, we introduce WavLLM, a robust and adaptive speech large language model with dual encoders, and a prompt-aware LoRA weight adapter, optimized by a two-stage curriculum learning approach. Leveraging dual encoders, we decouple different types of speech information, utilizing a Whisper encoder to process the semantic content of speech, and a WavLM encoder to capture the unique characteristics of the speaker's identity. Within the curriculum learning framework, WavLLM first builds its foundational capabilities by optimizing on mixed elementary single tasks, followed by advanced multi-task training on more complex tasks such as combinations of the elementary tasks. To enhance the flexibility and adherence to different tasks and instructions, a prompt-aware LoRA weight adapter is introduced in the second advanced multi-task training stage. We validate the proposed model on universal speech benchmarks including tasks such as ASR, ST, SV, ER, and also apply it to specialized datasets like Gaokao English listening comprehension set for SQA, and speech Chain-of-Thought (CoT) evaluation set. Experiments demonstrate that the proposed model achieves state-of-the-art performance across a range of speech tasks on the same model size, exhibiting robust generalization capabilities in executing complex tasks using CoT approach. Furthermore, our model successfully completes Gaokao tasks without specialized training. The codes, models, audio, and Gaokao evaluation set can be accessed at \url{aka.ms/wavllm}.
【8】 CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers and Consistency Models标题:CM—TTS:通过加权采样器和一致性模型提高实时文本到语音合成效率作者:Xiang Li,Fan Bu,Ambuj Mehrish,Yingting Li,Jiale Han,Bo Cheng,Soujanya Poria备注:Accepted by Findings of NAACL 2024. Code is available at this https URL摘要:神经文本到语音(TTS)系统在语音助手,电子学习和有声读物创作中有广泛的应用。追求现代模型,如扩散模型(DM),有望实现高保真,实时语音合成。然而,扩散模型中多步采样的效率提出了挑战。已经努力将GANs与DM集成,通过近似去噪分布来加速推理,但由于对抗训练,这会引入模型收敛的问题。为了克服这一点,我们引入CM-TTS,一种新的架构接地一致性模型(CM)。从连续时间扩散模型中汲取灵感,CM-TTS以更少的步骤实现了高质量的语音合成,而无需对抗性训练或预训练模型依赖性。我们进一步设计了加权采样器,将不同的采样位置与动态概率结合到模型训练中,确保整个训练过程中的无偏学习。我们提出了一个实时梅尔频谱生成一致性模型,通过全面的评估进行验证。实验结果强调CM-TTS的优越性比现有的单步语音合成系统,代表了该领域的重大进步。摘要:Neural Text-to-Speech (TTS) systems find broad applications in voice assistants, e-learning, and audiobook creation. The pursuit of modern models, like Diffusion Models (DMs), holds promise for achieving high-fidelity, real-time speech synthesis. Yet, the efficiency of multi-step sampling in Diffusion Models presents challenges. Efforts have been made to integrate GANs with DMs, speeding up inference by approximating denoising distributions, but this introduces issues with model convergence due to adversarial training. To overcome this, we introduce CM-TTS, a novel architecture grounded in consistency models (CMs). Drawing inspiration from continuous-time diffusion models, CM-TTS achieves top-quality speech synthesis in fewer steps without adversarial training or pre-trained model dependencies. We further design weighted samplers to incorporate different sampling positions into model training with dynamic probabilities, ensuring unbiased learning throughout the entire training process. We present a real-time mel-spectrogram generation consistency model, validated through comprehensive evaluations. Experimental results underscore CM-TTS's superiority over existing single-step speech synthesis systems, representing a significant advancement in the field.【9】 Classification of Short Segment Pediatric Heart Sounds Based on a Transformer-Based Convolutional Neural Network作者:Md Hassanuzzaman,Nurul Akhtar Hasan,Mohammad Abdullah Al Mamun,Khawza I Ahmed,Ahsan H Khandoker,Raqibul Mostafa摘要:由于心脏和大血管结构缺陷而引起的先天性异常被称为先天性心脏病或CHD。PCG可以提供有关心脏机械传导系统的基本细节,并指出与不同类型CHD相关的特定模式。本研究旨在探讨心音自动分类所需的最小信号持续时间。本研究还探讨了最佳信号质量评估指标(连续差的均方根)RMSSD和(过零率)ZCR值。基于Mel频率倒谱系数(MFCC)的特征被用作输入,以构建基于变换器的残差一维卷积神经网络,然后将其用于心音分类。研究表明,0.4是获得RMSSD和ZCR指标合适信号的理想阈值。此外,有效的心音分类需要5s的最小信号长度。它还表明,较短的信号(3 s心音)没有足够的信息来准确地分类心音,较长的信号(15 s心音)可能包含更多的噪声。对于5s信号,获得了最好的准确性,为93.69%,以区分心音。摘要:Congenital anomalies arising as a result of a defect in the structure of the heart and great vessels are known as congenital heart diseases or CHDs. A PCG can provide essential details about the mechanical conduction system of the heart and point out specific patterns linked to different kinds of CHD. This study aims to investigate the minimum signal duration required for the automatic classification of heart sounds. This study also investigated the optimum signal quality assessment indicator (Root Mean Square of Successive Differences) RMSSD and (Zero Crossings Rate) ZCR value. Mel-frequency cepstral coefficients (MFCCs) based feature is used as an input to build a Transformer-Based residual one-dimensional convolutional neural network, which is then used for classifying the heart sound. The study showed that 0.4 is the ideal threshold for getting suitable signals for the RMSSD and ZCR indicators. Moreover, a minimum signal length of 5s is required for effective heart sound classification. It also shows that a shorter signal (3 s heart sound) does not have enough information to categorize heart sounds accurately, and the longer signal (15 s heart sound) may contain more noise. The best accuracy, 93.69%, is obtained for the 5s signal to distinguish the heart sound.【10】 Where Are You From? Let Me Guess! Subdialect Recognition of Speeches in Sorani Kurdish标题:你从哪里来?让我猜猜!索拉尼库尔德语语音的亚方言识别作者:Sana Isam,Hossein Hassani备注:30 pages, 25 figures, 6 tables摘要:由于需要公开可用的数据集或可靠的资源(如社交媒体或网站)来收集数据,因此对索拉尼库尔德语亚方言进行分类是一项挑战。为了解决这个问题,我们对各个城市和村庄进行了实地考察,与来自不同年龄组、性别、学术背景和职业的母语人士进行了交流。我们记录了他们的声音,同时参与了涵盖生活方式,背景历史,爱好,兴趣,假期和生活课程等各种主题的对话。研究的目标地区是伊拉克库尔德斯坦地区。因此,我们从107次采访中收集了29小时16分40秒的录音,构成了一个包含六种方言的不平衡数据集。随后,我们采用了三种深度学习模型:ANN,CNN和RNN-LSTM。我们探索了各种配置,包括不同的跟踪持续时间,数据集分割和不平衡的数据集处理技术,如过采样和欠采样。进行了225项实验,并对结果进行了评价。结果表明,RNN-LSTM的准确率达到96%,优于其他方法。CNN达到了93%的准确率,ANN达到了75%。当应用于平衡数据集时,这三个模型都表现出了更好的性能,主要是当我们遵循过采样方法时。未来的研究可以探索其他未来的研究方向,包括其他库尔德方言。摘要:Classifying Sorani Kurdish subdialects poses a challenge due to the need for publicly available datasets or reliable resources like social media or websites for data collection. We conducted field visits to various cities and villages to address this issue, connecting with native speakers from different age groups, genders, academic backgrounds, and professions. We recorded their voices while engaging in conversations covering diverse topics such as lifestyle, background history, hobbies, interests, vacations, and life lessons. The target area of the research was the Kurdistan Region of Iraq. As a result, we accumulated 29 hours, 16 minutes, and 40 seconds of audio recordings from 107 interviews, constituting an unbalanced dataset encompassing six subdialects. Subsequently, we adapted three deep learning models: ANN, CNN, and RNN-LSTM. We explored various configurations, including different track durations, dataset splitting, and imbalanced dataset handling techniques such as oversampling and undersampling. Two hundred and twenty-five(225) experiments were conducted, and the outcomes were evaluated. The results indicated that the RNN-LSTM outperforms the other methods by achieving an accuracy of 96%. CNN achieved an accuracy of 93%, and ANN 75%. All three models demonstrated improved performance when applied to balanced datasets, primarily when we followed the oversampling approach. Future studies can explore additional future research directions to include other Kurdish dialects.
【11】 Data-Driven Room Acoustic Modeling Via Differentiable Feedback Delay Networks With Learnable Delay Lines标题:基于可学习延迟线的可微分反馈延迟网络的数据驱动房间声学建模作者:Alessandro Ilic Mezza,Riccardo Giampiccolo,Enzo De Sena,Alberto Bernardini备注:The article has been submitted to EURASIP Journal on Audio, Speech, and Music Processing on Jan 02, 2024 and is currently under review摘要:在过去的几十年里,广泛的研究一直致力于设计人工混响算法,旨在模拟物理环境的室内声学。尽管取得了重大进展,延迟网络模型的自动参数调整仍然是一个开放的挑战。我们介绍了一种新的方法,用于找到一个反馈延迟网络(FDN)的参数,使其输出呈现的感知品质的测量房间脉冲响应。所提出的方法涉及到可微分FDN与可训练的延迟线,这是第一次,使我们能够同时学习每一个延迟网络参数通过反向传播的实施。迭代优化过程寻求最小化时域损失函数,该时域损失函数包含考虑能量衰减和回波密度的可微项。通过实验验证,我们表明,所提出的方法产生的时不变的频率无关FDNs能够紧密匹配所需的声学特性,并优于现有的方法的基础上遗传算法和分析滤波器设计。摘要:Over the past few decades, extensive research has been devoted to the design of artificial reverberation algorithms aimed at emulating the room acoustics of physical environments. Despite significant advancements, automatic parameter tuning of delay-network models remains an open challenge. We introduce a novel method for finding the parameters of a Feedback Delay Network (FDN) such that its output renders the perceptual qualities of a measured room impulse response. The proposed approach involves the implementation of a differentiable FDN with trainable delay lines, which, for the first time, allows us to simultaneously learn each and every delay-network parameter via backpropagation. The iterative optimization process seeks to minimize a time-domain loss function incorporating differentiable terms accounting for energy decay and echo density. Through experimental validation, we show that the proposed method yields time-invariant frequency-independent FDNs capable of closely matching the desired acoustical characteristics, and outperforms existing methods based on genetic algorithms and analytical filter design.【1】 KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis标题:哈萨克语情感语合成数据集Kazatim TTS作者:Adal Abilbekov,Saida Mussakhojayeva,Rustem Yeshpanov,Huseyin Atakan Varol摘要:本研究的重点是创建哈萨克文语音转换(TTS)数据集,专为情感哈萨克文语音转换(TTS)应用程序。《哈萨克斯坦文语》共收录了54,760对语音文本,总时长为74.85小时,其中34.23小时由一名女性叙述者讲述,40.62小时由两名男性叙述者讲述。所考虑的情绪列表包括“中性”、“生气”、“高兴”、“悲伤”、“害怕”和“惊讶”。我们还开发了一个在Kazalek TTS数据集上训练的TTS模型。客观和主观的评价,以评估合成语音的质量,产生一个MCD分数范围内的6.02至7.67,以及MOS跨越3.51至3.57。为了促进可重复性并激发进一步的研究,我们在GitHub存储库中提供了我们的代码,预训练模型和数据集。摘要:This study focuses on the creation of the KazEmoTTS dataset, designed for emotional Kazakh text-to-speech (TTS) applications. KazEmoTTS is a collection of 54,760 audio-text pairs, with a total duration of 74.85 hours, featuring 34.23 hours delivered by a female narrator and 40.62 hours by two male narrators. The list of the emotions considered include "neutral", "angry", "happy", "sad", "scared", and "surprised". We also developed a TTS model trained on the KazEmoTTS dataset. Objective and subjective evaluations were employed to assess the quality of synthesized speech, yielding an MCD score within the range of 6.02 to 7.67, alongside a MOS that spanned from 3.51 to 3.57. To facilitate reproducibility and inspire further research, we have made our code, pre-trained model, and dataset accessible in our GitHub repository.
【2】 Voice Conversion Augmentation for Speaker Recognition on Defective Datasets作者:Ruijie Tao,Zhan Shi,Yidi Jiang,Tianchi Liu,Haizhou Li摘要:现代说话人识别系统依赖于丰富而均衡的数据集进行分类训练。然而,不同的缺陷数据集,如部分标记,小规模和不平衡的数据集,在现实世界中的应用是常见的。以前的作品通常从算法的角度研究每个场景的特定解决方案。然而,这些问题的根本原因在于数据集的不完善。为了用统一的解决方案来应对这些挑战,我们提出了语音转换增强(VCA)策略来从训练集中获得伪语音。此外,为了保证生成质量,我们设计了VCA-NN~(nearest neighbors)策略来从表示空间中与目标语音接近的话语中选择源语音。我们在三个创建的数据集上的实验结果表明,VCA-NN有效地缓解了这些数据集的问题,这为从数据方面处理说话人识别问题提供了一个新的方向。摘要:Modern speaker recognition system relies on abundant and balanced datasets for classification training. However, diverse defective datasets, such as partially-labelled, small-scale, and imbalanced datasets, are common in real-world applications. Previous works usually studied specific solutions for each scenario from the algorithm perspective. However, the root cause of these problems lies in dataset imperfections. To address these challenges with a unified solution, we propose the Voice Conversion Augmentation (VCA) strategy to obtain pseudo speech from the training set. Furthermore, to guarantee generation quality, we designed the VCA-NN~(nearest neighbours) strategy to select source speech from utterances that are close to the target speech in the representation space. Our experimental results on three created datasets demonstrated that VCA-NN effectively mitigates these dataset problems, which provides a new direction for handling the speaker recognition problems from the data aspect.【3】 Enhancing Real-World Active Speaker Detection with Multi-Modal Extraction Pre-Training作者:Ruijie Tao,Xinyuan Qian,Rohan Kumar Das,Xiaoxue Gao,Jiadong Wang,Haizhou Li摘要:视听主动说话人检测(AV-ASD)的目的是识别在一个或多个人的场景中哪个可见人脸正在说话。大多数现有的AV-ASD方法优先捕获语音唇对应。然而,在解决现实世界AV-ASD场景的挑战方面存在明显的差距。由于在这种情况下存在低质量的噪声视频,没有选择性收听能力的AV-ASD系统无法有效地从混合音频输入中滤除干扰性语音分量。在本文中,我们提出了一个名为“MuSED”的多模态说话人提取到检测框架,该框架使用视听目标说话人提取进行预训练以学习去噪能力,然后使用AV-ASD任务进行微调。同时,为了更好地捕捉多模态信息和处理现实世界中的问题,如丢失的模态,MuSED是直接在时域建模,并集成了多模态加减增强策略。我们的实验表明,MuSED大大优于最先进的AV-ASD方法,并分别在AVA-ActiveSpeaker数据集上实现95.6%的mAP,在ASW数据集上实现98.3%的AP,在Columbia AV-ASD数据集上实现97.9%的F1。我们将在适当的时候公开发布代码。摘要:Audio-visual active speaker detection (AV-ASD) aims to identify which visible face is speaking in a scene with one or more persons. Most existing AV-ASD methods prioritize capturing speech-lip correspondence. However, there is a noticeable gap in addressing the challenges from real-world AV-ASD scenarios. Due to the presence of low-quality noisy videos in such cases, AV-ASD systems without a selective listening ability are short of effectively filtering out disruptive voice components from mixed audio inputs. In this paper, we propose a Multi-modal Speaker Extraction-to-Detection framework named `MuSED', which is pre-trained with audio-visual target speaker extraction to learn the denoising ability, then it is fine-tuned with the AV-ASD task. Meanwhile, to better capture the multi-modal information and deal with real-world problems such as missing modality, MuSED is modelled on the time domain directly and integrates the multi-modal plus-and-minus augmentation strategy. Our experiments demonstrate that MuSED substantially outperforms the state-of-the-art AV-ASD methods and achieves 95.6% mAP on the AVA-ActiveSpeaker dataset, 98.3% AP on the ASW dataset, and 97.9% F1 on the Columbia AV-ASD dataset, respectively. We will publicly release the code in due course.【4】 Heterogeneity over Homogeneity: Investigating Multilingual Speech Pre-Trained Models for Detecting Audio Deepfake标题:同质性上的异质性:研究用于检测音频深度伪造的多语言语音预训练模型作者:Orchid Chetia Phukan,Gautam Siddharth Kashyap,Arun Balaji Buduru,Rajesh Sharma备注:Accepted to NAACL (Findings) 2024摘要:在这项工作中,我们研究了用于音频deepfake检测(ADD)的多语言语音预训练模型(PTM)。我们假设,在大规模多样化的多语言数据上训练的多语言PTM在训练前阶段获得了关于不同音高、口音和音调的知识,并使它们对变化更加鲁棒。因此,它们将更有效地检测音频deepfake。为了验证我们的假设,我们从最先进的(SOTA)PTM中提取表示,包括单语,多语言以及针对说话者和情感识别训练的PTM,并在ASVSpoof 2019(ASV),In-the-Wild(ITW)和DECRO基准数据库上对其进行评估。我们表明,表示从多语言PTM,简单的下游网络,获得最佳性能的ADD相比,其他PTM表示,这验证了我们的假设。我们还探讨了融合选定的PTM表示进一步改善ADD的可能性,我们提出了一个框架,MiO(合并为一)为此目的。通过MiO,我们在ASV和ITW上实现了SOTA性能,并在DECRO上实现了与当前SOTA工作相当的性能。摘要:In this work, we investigate multilingual speech Pre-Trained models (PTMs) for Audio deepfake detection (ADD). We hypothesize that multilingual PTMs trained on large-scale diverse multilingual data gain knowledge about diverse pitches, accents, and tones, during their pre-training phase and making them more robust to variations. As a result, they will be more effective for detecting audio deepfakes. To validate our hypothesis, we extract representations from state-of-the-art (SOTA) PTMs including monolingual, multilingual as well as PTMs trained for speaker and emotion recognition, and evaluated them on ASVSpoof 2019 (ASV), In-the-Wild (ITW), and DECRO benchmark databases. We show that representations from multilingual PTMs, with simple downstream networks, attain the best performance for ADD compared to other PTM representations, which validates our hypothesis. We also explore the possibility of fusion of selected PTM representations for further improvements in ADD, and we propose a framework, MiO (Merge into One) for this purpose. With MiO, we achieve SOTA performance on ASV and ITW and comparable performance on DECRO with current SOTA works.
【5】 Scaling Properties of Speech Language Models作者:Santiago Cuervo,Ricard Marxer摘要:语音语言模型(SLM)旨在从原始音频中学习语言,而无需文本资源。尽管取得了重大进展,我们目前的模型表现出较弱的语法和语义能力。然而,如果神经语言模型的缩放特性适用于语音模态,则这些能力将随着用于训练的计算量的增加而提高。在本文中,我们使用这种缩放行为的模型来估计我们目前的方法将产生一个SLM与基于文本的大型语言模型(LLM)的英语熟练程度的规模。我们在SLM和LLM中建立了预训练损失与下游句法和语义性能之间的强相关性,从而导致语言性能的可预测缩放。我们发现,语言性能的SLM规模高达三个数量级更慢,比基于文本的LLM。此外,我们还研究了旨在提高语义理解的合成数据的好处以及粗语音标记化的效果。摘要:Speech Language Models (SLMs) aim to learn language from raw audio, without textual resources. Despite significant advances, our current models exhibit weak syntax and semantic abilities. However, if the scaling properties of neural language models hold for the speech modality, these abilities will improve as the amount of compute used for training increases. In this paper, we use models of this scaling behavior to estimate the scale at which our current methods will yield a SLM with the English proficiency of text-based Large Language Models (LLMs). We establish a strong correlation between pre-training loss and downstream syntactic and semantic performance in SLMs and LLMs, which results in predictable scaling of linguistic performance. We show that the linguistic performance of SLMs scales up to three orders of magnitude more slowly than that of text-based LLMs. Additionally, we study the benefits of synthetic data designed to boost semantic understanding and the effects of coarser speech tokenization.【6】 Data-Driven Room Acoustic Modeling Via Differentiable Feedback Delay Networks With Learnable Delay Lines标题:基于可学习延迟线的可微分反馈延迟网络的数据驱动房间声学建模作者:Alessandro Ilic Mezza,Riccardo Giampiccolo,Enzo De Sena,Alberto Bernardini备注:The article has been submitted to EURASIP Journal on Audio, Speech, and Music Processing on Jan 02, 2024 and is currently under review摘要:在过去的几十年里,广泛的研究一直致力于设计人工混响算法,旨在模拟物理环境的室内声学。尽管取得了重大进展,延迟网络模型的自动参数调整仍然是一个开放的挑战。我们介绍了一种新的方法,用于找到一个反馈延迟网络(FDN)的参数,使其输出呈现的感知品质的测量房间脉冲响应。所提出的方法涉及到可微分FDN与可训练的延迟线,这是第一次,使我们能够同时学习每一个延迟网络参数通过反向传播的实施。迭代优化过程寻求最小化时域损失函数,该时域损失函数包含考虑能量衰减和回波密度的可微项。通过实验验证,我们表明,所提出的方法产生的时不变的频率无关FDNs能够紧密匹配所需的声学特性,并优于现有的方法的基础上遗传算法和分析滤波器设计。摘要:Over the past few decades, extensive research has been devoted to the design of artificial reverberation algorithms aimed at emulating the room acoustics of physical environments. Despite significant advancements, automatic parameter tuning of delay-network models remains an open challenge. We introduce a novel method for finding the parameters of a Feedback Delay Network (FDN) such that its output renders the perceptual qualities of a measured room impulse response. The proposed approach involves the implementation of a differentiable FDN with trainable delay lines, which, for the first time, allows us to simultaneously learn each and every delay-network parameter via backpropagation. The iterative optimization process seeks to minimize a time-domain loss function incorporating differentiable terms accounting for energy decay and echo density. Through experimental validation, we show that the proposed method yields time-invariant frequency-independent FDNs capable of closely matching the desired acoustical characteristics, and outperforms existing methods based on genetic algorithms and analytical filter design.【7】 A Novel Audio Representation for Music Genre Identification in MIR作者:Navin Kamuni,Mayank Jindal,Arpita Soni,Sukender Reddy Mallreddy,Sharath Chandra Macha摘要:对于音乐信息检索下游任务,最常见的音频表示是基于时间-频率的,例如Mel频谱图。为了识别音乐类型,本研究探讨了一种新形式的音频表示的可能性,最常见的MIR下游任务之一。因此,为了使用深度矢量量化对音乐进行离散编码,为创新的生成音乐模型(即Juan)创建了一种新的音频表示。使用几乎等同于最新技术水平(SOTA)的数据集和几乎相同的Transformer设计,将Jupiter的音频表示的有效性与Mel频谱图进行比较。这项研究的结果意味着,至少当Transformers使用20 k音轨的非常适度的数据集进行预训练时,Juvente的音频表示并不优于Mel频谱图。这可能是因为Jujant的音频表示没有充分考虑到人类听觉感知的特殊性。另一方面,Mel声谱图是专门根据人类听觉创建的。摘要:For Music Information Retrieval downstream tasks, the most common audio representation is time-frequency-based, such as Mel spectrograms. In order to identify musical genres, this study explores the possibilities of a new form of audio representation one of the most usual MIR downstream tasks. Therefore, to discretely encoding music using deep vector quantization; a novel audio representation was created for the innovative generative music model i.e. Jukebox. The effectiveness of Jukebox's audio representation is compared to Mel spectrograms using a dataset that is almost equivalent to State-of-the-Art (SOTA) and an almost same transformer design. The results of this study imply that, at least when the transformers are pretrained using a very modest dataset of 20k tracks, Jukebox's audio representation is not superior to Mel spectrograms. This could be explained by the fact that Jukebox's audio representation does not sufficiently take into account the peculiarities of human hearing perception. On the other hand, Mel spectrograms are specifically created with the human auditory sense in mind.
【8】 360+x: A Panoptic Multi-modal Scene Understanding Dataset作者:Hao Chen,Yuqi Hou,Chenyuan Qu,Irene Testini,Xiaohan Hong,Jianbo Jiao摘要:人类对世界的感知是由多种观点和模式塑造的。虽然许多现有的数据集专注于从某个角度(例如自我中心或第三人称视图)进行场景理解,但我们的数据集提供了全景视角(即具有多种数据模态的多个视角)。具体来说,我们封装第三人称全景和前视图,以及以自我为中心的单目/双目视图,具有丰富的模态,包括视频,多通道音频,定向双耳延迟,位置数据和每个场景中捕获的文本场景描述,呈现对世界的全面观察。图1提供了我们360+x数据集的所有28个场景类别的一瞥。据我们所知,这是第一个涵盖多个观点和多种数据模式的数据库,以模拟现实世界中如何访问日常信息。通过我们的基准分析,我们在建议的360+x数据集上提出了5个不同的场景理解任务,以评估每个数据模态和视角在全景场景理解中的影响和好处。我们希望这个独特的数据集可以扩大全面了解场景的范围,并鼓励社区从更多样化的角度来处理这些问题。摘要:Human perception of the world is shaped by a multitude of viewpoints and modalities. While many existing datasets focus on scene understanding from a certain perspective (e.g. egocentric or third-person views), our dataset offers a panoptic perspective (i.e. multiple viewpoints with multiple data modalities). Specifically, we encapsulate third-person panoramic and front views, as well as egocentric monocular/binocular views with rich modalities including video, multi-channel audio, directional binaural delay, location data and textual scene descriptions within each scene captured, presenting comprehensive observation of the world. Figure 1 offers a glimpse of all 28 scene categories of our 360+x dataset. To the best of our knowledge, this is the first database that covers multiple viewpoints with multiple data modalities to mimic how daily information is accessed in the real world. Through our benchmark analysis, we presented 5 different scene understanding tasks on the proposed 360+x dataset to evaluate the impact and benefit of each data modality and perspective in panoptic scene understanding. We hope this unique dataset could broaden the scope of comprehensive scene understanding and encourage the community to approach these problems from more diverse perspectives.【9】 Removing Speaker Information from Speech Representation using Variable-Length Soft Pooling作者:Injune Hwang,Kyogu Lee摘要:最近,已经有人努力使用语音合成的自监督框架来编码语音的语言信息。然而,从周围表示预测表示可能无意中使说话者信息纠缠在语音表示中。本文旨在通过利用语音的结构化特性来去除说话人信息,语音由具有清晰边界的音素等离散单元组成。神经网络预测这些边界,使可变长度池化用于基于事件的表示提取,而不是固定速率方法。边界预测器输出0和1之间的边界的概率,使池化变软。该模型经过训练,以最大限度地减少与通过时间拉伸和音高偏移增强的数据的合并表示的差异。为了证实所学习的表征包括内容信息但独立于说话人信息,用libri-light的语音ABX任务和SUPERB的说话人识别任务对模型进行了评估。摘要:Recently, there have been efforts to encode the linguistic information of speech using a self-supervised framework for speech synthesis. However, predicting representations from surrounding representations can inadvertently entangle speaker information in the speech representation. This paper aims to remove speaker information by exploiting the structured nature of speech, composed of discrete units like phonemes with clear boundaries. A neural network predicts these boundaries, enabling variable-length pooling for event-based representation extraction instead of fixed-rate methods. The boundary predictor outputs a probability for the boundary between 0 and 1, making pooling soft. The model is trained to minimize the difference with the pooled representation of the data augmented by time-stretch and pitch-shift. To confirm that the learned representation includes contents information but is independent of speaker information, the model was evaluated with libri-light's phonetic ABX task and SUPERB's speaker identification task.【10】 Personalized Neural Speech Codec作者:Inseon Jang,Haici Yang,Wootaek Lim,Seungkwon Beack,Minje Kim摘要:在本文中,我们提出了一个个性化的神经语音编解码器,设想个性化可以降低模型的复杂性或提高感知语音质量。尽管语音编解码器的常见用法是在通信的每一侧仅涉及单个说话者,但是在文献中很少探索针对特定用户的个性化编解码器。首先,我们假设扬声器可以分组为较小的子集的基础上,他们的感知相似性。然后,我们还假设,一个特定于组的编解码器可以专注于组的语音特性,以提高其感知质量和计算效率。为此,我们首先开发了一个Siamese网络,该网络从LibriSpeech数据集中学习说话人嵌入,然后将其分组为底层说话人集群。最后,我们在每个扬声器集群上重新训练基于LPCNet的语音编解码器基线。主观听力测试表明,所提出的个性化方案,同时保持语音质量的模型压缩。换句话说,在相同的模型复杂度下,个性化编解码器产生更好的语音质量。摘要:In this paper, we propose a personalized neural speech codec, envisioning that personalization can reduce the model complexity or improve perceptual speech quality. Despite the common usage of speech codecs where only a single talker is involved on each side of the communication, personalizing a codec for the specific user has rarely been explored in the literature. First, we assume speakers can be grouped into smaller subsets based on their perceptual similarity. Then, we also postulate that a group-specific codec can focus on the group's speech characteristics to improve its perceptual quality and computational efficiency. To this end, we first develop a Siamese network that learns the speaker embeddings from the LibriSpeech dataset, which are then grouped into underlying speaker clusters. Finally, we retrain the LPCNet-based speech codec baselines on each of the speaker clusters. Subjective listening tests show that the proposed personalization scheme introduces model compression while maintaining speech quality. In other words, with the same model complexity, personalized codecs produce better speech quality.【11】 A Comparative Analysis of Poetry Reading Audio: Singing, Narrating, or Somewhere In Between?标题:诗歌阅读音频的比较分析:歌唱、叙事还是介于两者之间?摘要:本文提供了一个大规模的诗歌朗读音频信号的计算分析,以揭示专业朗读诗歌中的音乐性。虽然其他类型的口语的声学特征已被广泛研究,但大多数文献仅限于叙事语音或歌声,讨论它们彼此之间的差异。在这项工作中,我们开发的信号处理方法,这是量身定制的,以捕捉独特的声学特性的诗歌阅读的基础上,他们的沉默模式,时间变化的本地音高,和节拍的稳定性。我们对三个大语料库的大规模统计分析,每个语料库都由叙述(LibriSpeech),歌唱声音(Intonation)和诗歌阅读(来自The Poetry Foundation)组成,发现诗歌阅读确实与歌唱声音共享一些音乐特征,尽管它也可能类似于叙述性演讲。摘要:This paper provides a computational analysis of poetry reading audio signals at a large scale to unveil the musicality within professionally-read poems. Although the acoustic characteristics of other types of spoken language have been extensively studied, most of the literature is limited to narrative speech or singing voice, discussing how different they are from each other. In this work, we develop signal processing methods, which are tailored to capture the unique acoustic characteristics of poetry reading based on their silence patterns, temporal variations of local pitch, and beat stability. Our large-scale statistical analyses on three big corpora, each of which consists of narration (LibriSpeech), singing voice (Intonation), and poetry reading (from The Poetry Foundation), discover that poetry reading does share some musical characteristics with singing voice, although it may also resemble narrative speech.
【12】 Measuring audio prompt adherence with distribution-based embedding distances摘要:越来越多的生成音乐模型可以以音频提示为条件,该音频提示用作模型要为其创建伴奏的音乐上下文(通常使用文本提示进一步指定)。对模型输出如何遵守音频提示的评估通常以模型或问题特定的方式进行,大概是因为还没有出现用于音频提示遵守的通用评估方法。这种方法在新模型的开发和训练中都很有用,并且可以使模型之间的性能具有可比性。在本文中,我们调查是否常用的分布为基础的距离,如Fr\'echet音频距离(FAD),可以用来衡量音频提示遵守。我们提出了一个简单的程序,基于少量的成分(嵌入模型,投影,嵌入距离和数据融合方法),我们系统地评估使用基线验证。在后续的实验中,我们测试的灵敏度,建议的音频坚持措施,音调和时移扰动。结果表明,所提出的措施是敏感的扰动,即使参考和候选分布是来自不同的音乐收藏。虽然需要更多的实验来回答未解决的问题,如测量不影响音频提示遵守的声学伪影的鲁棒性,但目前的结果表明,基于分布的嵌入距离提供了一种可行的方法来测量音频提示遵守。建议措施的python/pytorch实现作为github存储库公开提供。摘要:An increasing number of generative music models can be conditioned on an audio prompt that serves as musical context for which the model is to create an accompaniment (often further specified using a text prompt). Evaluation of how well model outputs adhere to the audio prompt is often done in a model or problem specific manner, presumably because no generic evaluation method for audio prompt adherence has emerged. Such a method could be useful both in the development and training of new models, and to make performance comparable across models. In this paper we investigate whether commonly used distribution-based distances like Fr\'echet Audio Distance (FAD), can be used to measure audio prompt adherence. We propose a simple procedure based on a small number of constituents (an embedding model, a projection, an embedding distance, and a data fusion method), that we systematically assess using a baseline validation. In a follow-up experiment we test the sensitivity of the proposed audio adherence measure to pitch and time shift perturbations. The results show that the proposed measure is sensitive to such perturbations, even when the reference and candidate distributions are from different music collections. Although more experimentation is needed to answer unaddressed questions like the robustness of the measure to acoustic artifacts that do not affect the audio prompt adherence, the current results suggest that distribution-based embedding distances provide a viable way of measuring audio prompt adherence. An python/pytorch implementation of the proposed measure is publicly available as a github repository.
【13】 WavLLM: Towards Robust and Adaptive Speech Large Language Model作者:Shujie Hu,Long Zhou,Shujie Liu,Sanyuan Chen,Hongkun Hao,Jing Pan,Xunying Liu,Jinyu Li,Sunit Sivasankaran,Linquan Liu,Furu Wei摘要:大型语言模型(LLM)的最新进展彻底改变了自然语言处理领域,逐步将其范围扩大到多模态感知和生成。然而,有效地将听力能力整合到LLM中带来了重大挑战,特别是在不同背景下的概括和执行复杂的听觉任务方面。在这项工作中,我们介绍了WavLLM,一个强大的和自适应的语音大语言模型与双编码器,和一个可感知的LoRA权重适配器,优化了两阶段的课程学习方法。利用双编码器,我们解耦不同类型的语音信息,利用Whisper编码器来处理语音的语义内容,并利用WavLM编码器来捕获说话者身份的独特特征。在课程学习框架内,WavLLM首先通过优化混合的基本单一任务来构建其基础能力,然后对更复杂的任务进行高级多任务培训,例如基本任务的组合。为了提高灵活性和对不同任务和指令的坚持,在第二个高级多任务训练阶段引入了一个可感知的LoRA权重适配器。我们验证了所提出的模型的通用语音基准,包括任务,如ASR,ST,SV,ER,并将其应用到专门的数据集,如高考英语听力理解集SQA,和语音思维链(CoT)评估集。实验表明,该模型在相同的模型大小的语音任务范围内实现了最先进的性能,表现出强大的泛化能力,在执行复杂的任务,使用CoT方法。此外,我们的模型在没有专门训练的情况下成功完成了高考任务。代码、模型、音频和高考评估集可在\url{aka.ms/wavllm}访问。摘要:The recent advancements in large language models (LLMs) have revolutionized the field of natural language processing, progressively broadening their scope to multimodal perception and generation. However, effectively integrating listening capabilities into LLMs poses significant challenges, particularly with respect to generalizing across varied contexts and executing complex auditory tasks. In this work, we introduce WavLLM, a robust and adaptive speech large language model with dual encoders, and a prompt-aware LoRA weight adapter, optimized by a two-stage curriculum learning approach. Leveraging dual encoders, we decouple different types of speech information, utilizing a Whisper encoder to process the semantic content of speech, and a WavLM encoder to capture the unique characteristics of the speaker's identity. Within the curriculum learning framework, WavLLM first builds its foundational capabilities by optimizing on mixed elementary single tasks, followed by advanced multi-task training on more complex tasks such as combinations of the elementary tasks. To enhance the flexibility and adherence to different tasks and instructions, a prompt-aware LoRA weight adapter is introduced in the second advanced multi-task training stage. We validate the proposed model on universal speech benchmarks including tasks such as ASR, ST, SV, ER, and also apply it to specialized datasets like Gaokao English listening comprehension set for SQA, and speech Chain-of-Thought (CoT) evaluation set. Experiments demonstrate that the proposed model achieves state-of-the-art performance across a range of speech tasks on the same model size, exhibiting robust generalization capabilities in executing complex tasks using CoT approach. Furthermore, our model successfully completes Gaokao tasks without specialized training. The codes, models, audio, and Gaokao evaluation set can be accessed at \url{aka.ms/wavllm}.【14】 CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers and Consistency Models标题:CM—TTS:通过加权采样器和一致性模型提高实时文本到语音合成效率作者:Xiang Li,Fan Bu,Ambuj Mehrish,Yingting Li,Jiale Han,Bo Cheng,Soujanya Poria备注:Accepted by Findings of NAACL 2024. Code is available at this https URL摘要:神经文本到语音(TTS)系统在语音助手,电子学习和有声读物创作中有广泛的应用。追求现代模型,如扩散模型(DM),有望实现高保真,实时语音合成。然而,扩散模型中多步采样的效率提出了挑战。已经努力将GANs与DM集成,通过近似去噪分布来加速推理,但由于对抗训练,这会引入模型收敛的问题。为了克服这一点,我们引入CM-TTS,一种新的架构接地一致性模型(CM)。从连续时间扩散模型中汲取灵感,CM-TTS以更少的步骤实现了高质量的语音合成,而无需对抗性训练或预训练模型依赖性。我们进一步设计了加权采样器,将不同的采样位置与动态概率结合到模型训练中,确保整个训练过程中的无偏学习。我们提出了一个实时梅尔频谱生成一致性模型,通过全面的评估进行验证。实验结果强调CM-TTS的优越性比现有的单步语音合成系统,代表了该领域的重大进步。摘要:Neural Text-to-Speech (TTS) systems find broad applications in voice assistants, e-learning, and audiobook creation. The pursuit of modern models, like Diffusion Models (DMs), holds promise for achieving high-fidelity, real-time speech synthesis. Yet, the efficiency of multi-step sampling in Diffusion Models presents challenges. Efforts have been made to integrate GANs with DMs, speeding up inference by approximating denoising distributions, but this introduces issues with model convergence due to adversarial training. To overcome this, we introduce CM-TTS, a novel architecture grounded in consistency models (CMs). Drawing inspiration from continuous-time diffusion models, CM-TTS achieves top-quality speech synthesis in fewer steps without adversarial training or pre-trained model dependencies. We further design weighted samplers to incorporate different sampling positions into model training with dynamic probabilities, ensuring unbiased learning throughout the entire training process. We present a real-time mel-spectrogram generation consistency model, validated through comprehensive evaluations. Experimental results underscore CM-TTS's superiority over existing single-step speech synthesis systems, representing a significant advancement in the field.
【15】 Classification of Short Segment Pediatric Heart Sounds Based on a Transformer-Based Convolutional Neural Network作者:Md Hassanuzzaman,Nurul Akhtar Hasan,Mohammad Abdullah Al Mamun,Khawza I Ahmed,Ahsan H Khandoker,Raqibul Mostafa摘要:由于心脏和大血管结构缺陷而引起的先天性异常被称为先天性心脏病或CHD。PCG可以提供有关心脏机械传导系统的基本细节,并指出与不同类型CHD相关的特定模式。本研究旨在探讨心音自动分类所需的最小信号持续时间。本研究还探讨了最佳信号质量评估指标(连续差的均方根)RMSSD和(过零率)ZCR值。基于Mel频率倒谱系数(MFCC)的特征被用作输入,以构建基于变换器的残差一维卷积神经网络,然后将其用于心音分类。研究表明,0.4是获得RMSSD和ZCR指标合适信号的理想阈值。此外,有效的心音分类需要5s的最小信号长度。它还表明,较短的信号(3 s心音)没有足够的信息来准确地分类心音,较长的信号(15 s心音)可能包含更多的噪声。对于5s信号,获得了最好的准确性,为93.69%,以区分心音。摘要:Congenital anomalies arising as a result of a defect in the structure of the heart and great vessels are known as congenital heart diseases or CHDs. A PCG can provide essential details about the mechanical conduction system of the heart and point out specific patterns linked to different kinds of CHD. This study aims to investigate the minimum signal duration required for the automatic classification of heart sounds. This study also investigated the optimum signal quality assessment indicator (Root Mean Square of Successive Differences) RMSSD and (Zero Crossings Rate) ZCR value. Mel-frequency cepstral coefficients (MFCCs) based feature is used as an input to build a Transformer-Based residual one-dimensional convolutional neural network, which is then used for classifying the heart sound. The study showed that 0.4 is the ideal threshold for getting suitable signals for the RMSSD and ZCR indicators. Moreover, a minimum signal length of 5s is required for effective heart sound classification. It also shows that a shorter signal (3 s heart sound) does not have enough information to categorize heart sounds accurately, and the longer signal (15 s heart sound) may contain more noise. The best accuracy, 93.69%, is obtained for the 5s signal to distinguish the heart sound.【16】 Where Are You From? Let Me Guess! Subdialect Recognition of Speeches in Sorani Kurdish标题:你从哪里来?让我猜猜!索拉尼库尔德语语音的亚方言识别作者:Sana Isam,Hossein Hassani备注:30 pages, 25 figures, 6 tables摘要:由于需要公开可用的数据集或可靠的资源(如社交媒体或网站)来收集数据,因此对索拉尼库尔德语亚方言进行分类是一项挑战。为了解决这个问题,我们对各个城市和村庄进行了实地考察,与来自不同年龄组、性别、学术背景和职业的母语人士进行了交流。我们记录了他们的声音,同时参与了涵盖生活方式,背景历史,爱好,兴趣,假期和生活课程等各种主题的对话。研究的目标地区是伊拉克库尔德斯坦地区。因此,我们从107次采访中收集了29小时16分40秒的录音,构成了一个包含六种方言的不平衡数据集。随后,我们采用了三种深度学习模型:ANN,CNN和RNN-LSTM。我们探索了各种配置,包括不同的跟踪持续时间,数据集分割和不平衡的数据集处理技术,如过采样和欠采样。进行了225项实验,并对结果进行了评价。结果表明,RNN-LSTM的准确率达到96%,优于其他方法。CNN达到了93%的准确率,ANN达到了75%。当应用于平衡数据集时,这三个模型都表现出了更好的性能,主要是当我们遵循过采样方法时。未来的研究可以探索其他未来的研究方向,包括其他库尔德方言。摘要:Classifying Sorani Kurdish subdialects poses a challenge due to the need for publicly available datasets or reliable resources like social media or websites for data collection. We conducted field visits to various cities and villages to address this issue, connecting with native speakers from different age groups, genders, academic backgrounds, and professions. We recorded their voices while engaging in conversations covering diverse topics such as lifestyle, background history, hobbies, interests, vacations, and life lessons. The target area of the research was the Kurdistan Region of Iraq. As a result, we accumulated 29 hours, 16 minutes, and 40 seconds of audio recordings from 107 interviews, constituting an unbalanced dataset encompassing six subdialects. Subsequently, we adapted three deep learning models: ANN, CNN, and RNN-LSTM. We explored various configurations, including different track durations, dataset splitting, and imbalanced dataset handling techniques such as oversampling and undersampling. Two hundred and twenty-five(225) experiments were conducted, and the outcomes were evaluated. The results indicated that the RNN-LSTM outperforms the other methods by achieving an accuracy of 96%. CNN achieved an accuracy of 93%, and ANN 75%. All three models demonstrated improved performance when applied to balanced datasets, primarily when we followed the oversampling approach. Future studies can explore additional future research directions to include other Kurdish dialects.