今日论文合集:cs.SD语音18篇,eess.AS音频处理18篇。

本文经arXiv每日学术速递授权转载

cs.SD语音
【1】 Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
标题: Mini-Omni:语言模型可以在流媒体中一边思考一边听、说话
作者:Zhifei Xie,Changqiao Wu
备注:10 pages
链接:点击下载PDF文件
摘要:语言模型的最新进展取得了重大进展。GPT-4 o作为一个新的里程碑,实现了与人类的实时对话,展示了接近人类的自然流畅性。这种人机交互需要具有直接使用音频模态执行推理并在流中生成输出的能力的模型。然而,这仍然超出了当前学术模型的范围,因为它们通常依赖于额外的TTS系统进行语音合成,从而导致不期望的延迟。本文介绍了Mini-Omni,一个基于音频的端到端会话模型,能够实时语音交互。为了实现这一功能,我们提出了一种文本指导的语音生成方法,以及推理过程中的批处理并行策略,以进一步提高性能。我们的方法还有助于保留原始模型的语言能力,最小的退化,使其他作品建立实时交互能力。我们称这种训练方法为“任何模型都能说话”。我们还引入了VoiceAssistant-400 K数据集来微调针对语音输出优化的模型。据我们所知,Mini-Omni是第一个完全端到端的实时语音交互开源模型,为未来的研究提供了宝贵的潜力。摘要:Recent advances in language models have achieved significant progress. GPT-4o, as a new milestone, has enabled real-time conversations with humans, demonstrating near-human natural fluency. Such human-computer interaction necessitates models with the capability to perform reasoning directly with the audio modality and generate output in streaming. However, this remains beyond the reach of current academic models, as they typically depend on extra TTS systems for speech synthesis, resulting in undesirable latency. This paper introduces the Mini-Omni, an audio-based end-to-end conversational model, capable of real-time speech interaction. To achieve this capability, we propose a text-instructed speech generation method, along with batch-parallel strategies during inference to further boost the performance. Our method also helps to retain the original model's language capabilities with minimal degradation, enabling other works to establish real-time interaction capabilities. We call this training method "Any Model Can Talk". We also introduce the VoiceAssistant-400K dataset to fine-tune models optimized for speech output. To our best knowledge, Mini-Omni is the first fully end-to-end, open-source model for real-time speech interaction, offering valuable potential for future research.

【2】 Towards Efficient Modelling of String Dynamics: A Comparison of State Space and Koopman based Deep Learning Methods
标题: 实现字符串动力学的高效建模:状态空间和基于Koopman的深度学习方法的比较
作者:Rodrigo Diaz,Carlos De La Vega Martin,Mark Sandler
备注:Accepted to DAFx2024
链接:点击下载PDF文件
摘要:本文介绍了状态空间模型(SSM)和基于Koopman的深度学习方法,用于对线性和非线性刚性弦的动力学进行建模。通过对不同初始条件和采样率下生成的数据集进行实验,我们评估了这些模型准确模拟弦动力学中观察到的复杂行为的能力。我们的研究结果表明,我们提出的Koopman为基础的模型执行以及或优于其他现有的方法在非线性情况下的长序列建模。 我们通知这些架构的设计与手头的问题的结构。尽管在将模型预测扩展到训练范围之外(即,外推),我们研究的重点在于模型在训练时间间隔内跨不同初始条件进行概括的能力。这项研究有助于深入了解动力系统的物理建模(特别是那些解决音乐声学),通过提供这些和以前的方法的比较概述,并引入模型改进的创新策略。我们的研究结果突出了这些模型在模拟非线性动力学的有效性,并强调其广泛的适用性,在扩展序列的动力系统准确建模。摘要:This paper presents an examination of State Space Models (SSM) and Koopman-based deep learning methods for modelling the dynamics of both linear and non-linear stiff strings. Through experiments with datasets generated under different initial conditions and sample rates, we assess the capacity of these models to accurately model the complex behaviours observed in string dynamics. Our findings indicate that our proposed Koopman-based model performs as well as or better than other existing approaches in non-linear cases for long-sequence modelling. We inform the design of these architectures with the structure of the problems at hand. Although challenges remain in extending model predictions beyond the training horizon (i.e., extrapolation), the focus of our investigation lies in the models' ability to generalise across different initial conditions within the training time interval. This research contributes insights into the physical modelling of dynamical systems (in particular those addressing musical acoustics) by offering a comparative overview of these and previous methods and introducing innovative strategies for model improvement. Our results highlight the efficacy of these models in simulating non-linear dynamics and emphasise their wide-ranging applicability in accurately modelling dynamical systems over extended sequences.

【3】 Audio xLSTMs: Learning Self-supervised audio representations with xLSTMs
标题: 音频xLSTM:使用xLSTM学习自我监督的音频表示
作者:Sarthak Yadav,Sergios Theodoridis,Zheng-Hua Tan
备注:Under review at ICASSP 2025. arXiv admin note: text overlap with arXiv:2406.02178
链接:点击下载PDF文件
摘要:虽然Transformer已经成为杰出的神经架构,但已经出现了几个独立的研究路线来解决其局限性。循环神经方法也引起了很多新的兴趣,包括扩展的长短期记忆(xLSTM)架构,它重振了原始的LSTM架构。然而,虽然xLSTM与Transformer相比表现出了竞争力,但它们用于学习自监督通用音频表示的可行性尚未得到评估。这项工作提出了Audio xLSTM(AxLSTM),这是一种在自监督设置中从掩蔽的频谱图补丁中学习音频表示的方法。在AudioSet数据集上进行预训练后,所提出的AxLSTM模型在一组10个不同的下游任务中的相对性能比可比的自监督音频频谱图Transformer(SSAST)基线高出20%,同时参数减少了45%。摘要:While the transformer has emerged as the eminent neural architecture, several independent lines of research have emerged to address its limitations. Recurrent neural approaches have also observed a lot of renewed interest, including the extended long short-term memory (xLSTM) architecture, which reinvigorates the original LSTM architecture. However, while xLSTMs have shown competitive performance compared to the transformer, their viability for learning self-supervised general-purpose audio representations has not yet been evaluated. This work proposes Audio xLSTM (AxLSTM), an approach to learn audio representations from masked spectrogram patches in a self-supervised setting. Pretrained on the AudioSet dataset, the proposed AxLSTM models outperform comparable self-supervised audio spectrogram transformer (SSAST) baselines by up to 20% in relative performance across a set of ten diverse downstream tasks while having up to 45% fewer parameters.

【4】 Human-Inspired Audio-Visual Speech Recognition: Spike Activity, Cueing Interaction and Causal Processing
标题: 受人类启发的视听语音识别:尖峰活动、线索交互和因果处理
作者:Qianhui Liu,Jiadong Wang,Yang Wang,Xin Yang,Gang Pan,Haizhou Li
链接:点击下载PDF文件
摘要:人类自然地执行视听语音识别(AVSR),通过整合听觉和视觉信息来提高准确性和鲁棒性。尖峰神经网络(SNN)模仿大脑的信息处理机制,非常适合模拟人类的AVSR能力。尽管SNN具有潜力,但针对AVSR的SNN研究却很少,大多数现有的视听多模式方法都集中在对象或数字识别上。这些模型简单地整合了两种模式的特征,忽略了它们的独特特征和相互作用。此外,它们通常依赖于未来的信息进行当前处理,这增加了识别延迟并限制了实时适用性。受人类语音感知的启发,本文提出了一种新的人类启发的SNN命名为HI-AVSNN的AVSR,结合了三个关键特征:提示交互,因果处理和尖峰活动。对于线索交互,我们提出了一个视觉线索听觉注意模块(VCA2M),利用视觉线索来引导注意听觉功能。我们实现因果关系的处理对齐SNN的时间维度与视觉和听觉功能,并应用时间掩蔽,只利用过去和当前的信息。为了实现尖峰活动,除了使用SNN之外,我们还利用事件相机来捕获嘴唇运动作为尖峰,模仿人类视网膜并提供有效的视觉数据。我们评估HI-AVSNN的视听语音识别数据集相结合的DVS-Lip数据集与其相应的音频样本。实验结果表明,我们提出的融合方法的优越性,优于现有的视听SNN融合方法,并实现了2.27%的提高精度比现有的基于SNN的AVSR方法。摘要:Humans naturally perform audiovisual speech recognition (AVSR), enhancing the accuracy and robustness by integrating auditory and visual information. Spiking neural networks (SNNs), which mimic the brain's information-processing mechanisms, are well-suited for emulating the human capability of AVSR. Despite their potential, research on SNNs for AVSR is scarce, with most existing audio-visual multimodal methods focused on object or digit recognition. These models simply integrate features from both modalities, neglecting their unique characteristics and interactions. Additionally, they often rely on future information for current processing, which increases recognition latency and limits real-time applicability. Inspired by human speech perception, this paper proposes a novel human-inspired SNN named HI-AVSNN for AVSR, incorporating three key characteristics: cueing interaction, causal processing and spike activity. For cueing interaction, we propose a visual-cued auditory attention module (VCA2M) that leverages visual cues to guide attention to auditory features. We achieve causal processing by aligning the SNN's temporal dimension with that of visual and auditory features and applying temporal masking to utilize only past and current information. To implement spike activity, in addition to using SNNs, we leverage the event camera to capture lip movement as spikes, mimicking the human retina and providing efficient visual data. We evaluate HI-AVSNN on an audiovisual speech recognition dataset combining the DVS-Lip dataset with its corresponding audio samples. Experimental results demonstrate the superiority of our proposed fusion method, outperforming existing audio-visual SNN fusion methods and achieving a 2.27% improvement in accuracy over the only existing SNN-based AVSR method.

【5】 RAVE for Speech: Efficient Voice Conversion at High Sampling Rates
标题: RAVE for Speech:高采样率下的高效语音转换
作者:Anders R. Bargum,Simon Lajboschitz,Cumhur Erkut
备注:Accepted for publication in Proceedings of the 27th International Conference on Digital Audio Effects (DAFx24), Guildford, United Kingdom, 3 - 7 September 2024
链接:点击下载PDF文件
摘要:语音转换在音频处理和语音合成领域中越来越受欢迎。通常,主要目标是将输入身份转换为目标说话者的身份,而不改变其语言内容。虽然目前的工作提供了高保真的解决方案,他们很少关注模型的简单性,高采样率的环境或流的能力。通过将语音表示学习到一个生成的音色转移模型,传统上创建的音乐目的,我们调查的领域直接在时域中产生的高采样率的语音转换。更具体地说,我们将基线模型的潜在空间引导到语言相关的表示,并将其置于外部说话者信息上。通过客观和主观评估,我们证明提出的解决方案可以达到与已知说话者的最先进解决方案相当的自然度、质量和可懂度水平,同时显着减少推理时间。然而,尽管转换后的输出中存在目标说话者的特征,但与未见过的说话者的实际相似性仍然是一个挑战。摘要:Voice conversion has gained increasing popularity within the field of audio manipulation and speech synthesis. Often, the main objective is to transfer the input identity to that of a target speaker without changing its linguistic content. While current work provides high-fidelity solutions they rarely focus on model simplicity, high-sampling rate environments or stream-ability. By incorporating speech representation learning into a generative timbre transfer model, traditionally created for musical purposes, we investigate the realm of voice conversion generated directly in the time domain at high sampling rates. More specifically, we guide the latent space of a baseline model towards linguistically relevant representations and condition it on external speaker information. Through objective and subjective assessments, we demonstrate that the proposed solution can attain levels of naturalness, quality, and intelligibility comparable to those of a state-of-the-art solution for seen speakers, while significantly decreasing inference time. However, despite the presence of target speaker characteristics in the converted output, the actual similarity to unseen speakers remains a challenge.

【6】 SALSA: Speedy ASR-LLM Synchronous Aggregation
标题: SALSA:快速ASR-LLM同步聚合
作者:Ashish Mittal,Darshan Prabhu,Sunita Sarawagi,Preethi Jyothi
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:利用预先训练的LLM来改善ASR系统,特别是低资源语言,现在是一个新兴的研究领域。现有方法的范围从使用LLM进行ASR纠错到用LLM代替ASR解码器的紧密耦合系统。这些方法要么增加解码时间,要么需要对交叉注意层进行昂贵的训练。我们提出了SALSA,它将ASR的解码器层耦合到LLM解码器,同时同步推进两个解码器。这样的耦合是用最后一个解码器状态的简单投影来执行的,因此比早期的方法明显更有训练效率。我们提出的耦合的一个挑战是处理LLM和ASR系统的标记器之间的不匹配。我们使用与LLM和ASR词汇表相关的级联标记化来处理这种不匹配。我们在FLEURS基准测试中对8种低资源语言进行了SALSA评估,产生了高达38%的WER大幅降低。摘要:Harnessing pre-trained LLMs to improve ASR systems, particularly for low-resource languages, is now an emerging area of research. Existing methods range from using LLMs for ASR error correction to tightly coupled systems that replace the ASR decoder with the LLM. These approaches either increase decoding time or require expensive training of the cross-attention layers. We propose SALSA, which couples the decoder layers of the ASR to the LLM decoder, while synchronously advancing both decoders. Such coupling is performed with a simple projection of the last decoder state, and is thus significantly more training efficient than earlier approaches. A challenge of our proposed coupling is handling the mismatch between the tokenizers of the LLM and ASR systems. We handle this mismatch using cascading tokenization with respect to the LLM and ASR vocabularies. We evaluate SALSA on 8 low-resource languages in the FLEURS benchmark, yielding substantial WER reductions of up to 38%.

【7】 Enabling Beam Search for Language Model-Based Text-to-Speech Synthesis
标题: 启用基于语言模型的文本到语音合成的梁搜索
作者:Zehai Tu,Guangyan Zhang,Yiting Lu,Adaeze Adigwe,Simon King,Yiwen Guo
链接:点击下载PDF文件
摘要:将连续语音标记为离散标记序列并使用语言模型(LM)对其进行建模已经在文本到语音(TTS)合成中取得了重大成功。尽管这些模型可以生成高质量和自然度的语音,但它们的合成样本仍然可能存在伪影、发音错误、单词重复等问题。在本文中,我们认为这些不良特性可能部分是由LM自回归解码期间基于采样的策略的随机性引起的。因此,我们着眼于基于最大化的解码方法,并提出时间重复感知的多样化波束搜索(TRAD-BS),以找到最可能的序列生成的语音令牌。两个国家的最先进的LM为基础的TTS模型的实验表明,我们提出的最大化为基础的解码策略生成的语音具有更少的发音错误和提高说话人的一致性。摘要:Tokenising continuous speech into sequences of discrete tokens and modelling them with language models (LMs) has led to significant success in text-to-speech (TTS) synthesis. Although these models can generate speech with high quality and naturalness, their synthesised samples can still suffer from artefacts, mispronunciation, word repeating, etc. In this paper, we argue these undesirable properties could partly be caused by the randomness of sampling-based strategies during the autoregressive decoding of LMs. Therefore, we look at maximisation-based decoding approaches and propose Temporal Repetition Aware Diverse Beam Search (TRAD-BS) to find the most probable sequences of the generated speech tokens. Experiments with two state-of-the-art LM-based TTS models demonstrate that our proposed maximisation-based decoding strategy generates speech with fewer mispronunciations and improved speaker consistency.

【8】 Measuring the Accuracy of Automatic Speech Recognition Solutions
标题: 衡量自动语音识别解决方...性
作者:Korbinian Kuhn,Verena Kersken,Benedikt Reuter,Niklas Egger,Gottfried Zimmermann
Journal-ref:ACM Transactions on Accessible Computing, Volume 16, Issue 4, Article 25 (2023), 1-23
链接:点击下载PDF文件
摘要:d 对于聋人和重听人来说,字幕是一个必不可少的无障碍工具。人工智能(AI)的重大发展意味着自动语音识别(ASR)现在是许多流行应用的一部分。这使得创建字幕变得容易和广泛可用-但转录需要高水平的准确性才能访问。科学出版物和行业报告的错误率非常低,声称人工智能已经达到了人类的水平,甚至超过了人工转录。与此同时,DHH社区报告了ASR准确性和可靠性的严重问题。对于依赖转录的人来说,技术创新和现实生活体验之间似乎存在不匹配。需要独立和全面的数据来捕获ASR的状态。我们测量了11个常见的ASR服务与高等教育讲座录音的性能。我们评估了技术条件的影响,如流,词汇的使用和语言之间的差异。我们的研究结果表明,准确性范围广泛的供应商和个人的音频样本。我们还测量了用于直播活动的流媒体ASR的质量明显较低。我们的研究表明,尽管最近的ASR的改进,公共服务缺乏准确性的可靠性。摘要:For d Deaf and hard of hearing (DHH) people, captioning is an essential accessibility tool. Significant developments in artificial intelligence (AI) mean that Automatic Speech Recognition (ASR) is now a part of many popular applications. This makes creating captions easy and broadly available - but transcription needs high levels of accuracy to be accessible. Scientific publications and industry report very low error rates, claiming AI has reached human parity or even outperforms manual transcription. At the same time the DHH community reports serious issues with the accuracy and reliability of ASR. There seems to be a mismatch between technical innovations and the real-life experience for people who depend on transcription. Independent and comprehensive data is needed to capture the state of ASR. We measured the performance of eleven common ASR services with recordings of Higher Education lectures. We evaluated the influence of technical conditions like streaming, the use of vocabularies, and differences between languages. Our results show that accuracy ranges widely between vendors and for the individual audio samples. We also measured a significant lower quality for streaming ASR, which is used for live events. Our study shows that despite the recent improvements of ASR, common services lack reliability in accuracy.

【9】 Revisit Micro-batch Clipping: Adaptive Data Pruning via Gradient Manipulation
标题: 重温微批剪辑:通过梯度操纵进行自适应数据修剪
作者:Lun Wang
链接:点击下载PDF文件
摘要:微批量裁剪,梯度裁剪方法,最近显示出潜在的提高自动语音识别(ASR)模型的性能。然而,这种改进背后的潜在机制仍然是神秘的,特别是观察到只有某些微批量是有益的。在本文中,我们首次尝试解释这一现象。受最近数据修剪研究的启发,我们假设特定的训练样本可能会在某些训练阶段阻碍模型收敛。在此假设下,收敛性分析表明,微批量裁剪可以以不随训练迭代次数增加而减小的额外常数偏差为代价,渐进地提高收敛速度。偏倚取决于几个因素,并且可以在特定的微批量下最小化,从而阐明先前观察到的最佳点微批量的存在。我们还验证了语音模型之外的视觉和语言模型上的微批量裁剪的有效性,并在这些领域显示出有前途的性能增益。对潜在限制的探索表明,当训练数据来自多个不同的领域时,微批量裁剪的效果较差。摘要:Micro-batch clipping, a gradient clipping method, has recently shown potential in enhancing auto-speech recognition (ASR) model performance. However, the underlying mechanism behind this improvement remains mysterious, particularly the observation that only certain micro-batch sizes are beneficial. In this paper, we make the first attempt to explain this phenomenon. Inspired by recent data pruning research, we assume that specific training samples may impede model convergence during certain training phases. Under this assumption, the convergence analysis shows that micro-batch clipping can improve the convergence rate asymptotically at the cost of an additional constant bias that does not diminish with more training iterations. The bias is dependent on a few factors and can be minimized at specific micro-batch size, thereby elucidating the existence of the sweet-spot micro-batch size observed previously. We also verify the effectiveness of micro-batch clipping beyond speech models on vision and language models, and show promising performance gains in these domains. An exploration of potential limitations shows that micro-batch clipping is less effective when training data originates from multiple distinct domains.

【10】 Improving Generalization of Speech Separation in Real-World Scenarios: Strategies in Simulation, Optimization, and Evaluation
标题: 改善现实场景中语音分离的推广:模拟、优化和评估策略
作者:Ke Chen,Jiaqi Su,Taylor Berg-Kirkpatrick,Shlomo Dubnov,Zeyu Jin
备注:In Proceedings of the 25th Annual Conference of the International Speech Communication Association, Interspeech 2024
链接:点击下载PDF文件
摘要:在具有噪声和混响的各种声学环境中实现对重叠扬声器的鲁棒语音分离仍然是一个公开的挑战。尽管现有的数据集可用于针对特定场景训练分离器,但是它们不能有效地在不同的现实世界场景中推广。在本文中,我们提出了一种新的数据模拟管道,从一系列声学环境和内容中产生不同的训练数据,并提出了新的训练范例,以提高一般语音分离模型的质量。具体来说,我们首先介绍AC-SIM,这是一种数据模拟管道,它包含了内容和声学的广泛变化。然后,我们将多个训练目标集成到排列不变训练(PIT)中,以提高训练模型的分离质量和泛化能力。最后,我们在分离架构和基准测试中进行了全面的客观和人类听觉实验,以验证我们的方法,证明了非同源和真实世界测试集的泛化能力有了实质性的提高。摘要:Achieving robust speech separation for overlapping speakers in various acoustic environments with noise and reverberation remains an open challenge. Although existing datasets are available to train separators for specific scenarios, they do not effectively generalize across diverse real-world scenarios. In this paper, we present a novel data simulation pipeline that produces diverse training data from a range of acoustic environments and content, and propose new training paradigms to improve quality of a general speech separation model. Specifically, we first introduce AC-SIM, a data simulation pipeline that incorporates broad variations in both content and acoustics. Then we integrate multiple training objectives into the permutation invariant training (PIT) to enhance separation quality and generalization of the trained model. Finally, we conduct comprehensive objective and human listening experiments across separation architectures and benchmarks to validate our methods, demonstrating substantial improvement of generalization on both non-homologous and real-world test sets.

【11】 A Deep Learning Approach to Localizing Multi-level Airway Collapse Based on Snoring Sounds
标题: 基于打鼾声音定位多层气道塌陷的深度学习方法
作者:Ying-Chieh Hsu,Stanley Yung-Chuan Liu,Chao-Jung Huang,Chi-Wei Wu,Ren-Kai Cheng,Jane Yung-Jen Hsu,Shang-Ran Huang,Yuan-Ren Cheng,Fu-Shun Hsu
链接:点击下载PDF文件
摘要:本研究利用药物诱导睡眠内窥镜检查(DISE)的数据,研究了机器 深度学习在阻塞性睡眠呼吸暂停(OSA)患者上气道不同水平激发的打鼾声音分类中的应用。根据腭、口咽、舌根和会厌(VOTE)分类系统,对39名受试者的鼾声进行分析和标记。该数据集包括5,173个一秒片段,用于训练和测试模型,包括支持向量机(SVM),双向长短期记忆(BiLSTM)和ResNet-50。ResNet-50是一种卷积神经网络(CNN),在对打鼾声音进行分类方面表现出最佳的整体性能,特别是在识别多级障碍物方面。该研究强调了将打鼾声学与深度学习相结合以改善OSA诊断和治疗的潜力。然而,注意到诸如样本量有限、数据不平衡以及打鼾引起的声音与自然打鼾声音之间的差异等挑战,这表明需要进一步研究以提高模型的准确性和可推广性。摘要:This study investigates the application of machine deep learning to classify snoring sounds excited at different levels of the upper airway in patients with obstructive sleep apnea (OSA) using data from drug-induced sleep endoscopy (DISE). The snoring sounds of 39 subjects were analyzed and labeled according to the Velum, Oropharynx, Tongue Base, and Epiglottis (VOTE) classification system. The dataset, comprising 5,173 one-second segments, was used to train and test models, including Support Vector Machine (SVM), Bidirectional Long Short-Term Memory (BiLSTM), and ResNet-50. The ResNet-50, a convolutional neural network (CNN), showed the best overall performance in classifying snoring acoustics, particularly in identifying multi-level obstructions. The study emphasizes the potential of integrating snoring acoustics with deep learning to improve the diagnosis and treatment of OSA. However, challenges such as limited sample size, data imbalance, and differences between pharmacologically induced and natural snoring sounds were noted, suggesting further research to enhance model accuracy and generalizability.

【12】 Automatic detection of Mild Cognitive Impairment using high-dimensional acoustic features in spontaneous speech
标题: 使用自发言语中的多维声学特征自动检测轻度认知障碍
作者:Cong Zhang,Wenxing Guo,Hongsheng Dai
链接:点击下载PDF文件
摘要:本研究解决了TAUKADIAL挑战,重点是轻度认知障碍(MCI)和神经典型对照人群的语音分类。我们进行了三个实验,比较了五种机器学习方法:随机森林,稀疏逻辑回归,k最近邻,稀疏支持向量机和决策树,利用openSMILE自动提取的1076个声学特征。在实验1中,整个数据集被用来训练语言无关的模型。实验2引入了语言检测步骤,导致每种语言的单独模型训练。实验3进一步增强了实验1的语言不可知模型,特别关注使用样本外测试数据评估模型的鲁棒性。在所有三个实验中,结果一致有利于能够处理高维数据的模型,如随机森林和稀疏逻辑回归,在分类语音MCI和控制。摘要:This study addresses the TAUKADIAL challenge, focusing on the classification of speech from people with Mild Cognitive Impairment (MCI) and neurotypical controls. We conducted three experiments comparing five machine-learning methods: Random Forests, Sparse Logistic Regression, k-Nearest Neighbors, Sparse Support Vector Machine, and Decision Tree, utilizing 1076 acoustic features automatically extracted using openSMILE. In Experiment 1, the entire dataset was used to train a language-agnostic model. Experiment 2 introduced a language detection step, leading to separate model training for each language. Experiment 3 further enhanced the language-agnostic model from Experiment 1, with a specific focus on evaluating the robustness of the models using out-of-sample test data. Across all three experiments, results consistently favored models capable of handling high-dimensional data, such as Random Forest and Sparse Logistic Regression, in classifying speech from MCI and controls.

【13】 WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
标题: WavTokenizer:一种用于音频语言建模的高效声学离散编解码器Tokenizer
作者:Shengpeng Ji,Ziyue Jiang,Xize Cheng,Yifu Chen,Minghui Fang,Jialong Zuo,Qian Yang,Ruiqi Li,Ziang Zhang,Xiaoda Yang,Rongjie Huang,Yidi Jiang,Qian Chen,Siqi Zheng,Wen Wang,Zhou Zhao
备注:Working in progress. arXiv admin note: text overlap with arXiv:2402.12208
链接:点击下载PDF文件
摘要:语言模型已被有效地应用于对自然信号(诸如图像、视频、语音和音频)进行建模。这些模型的一个重要组成部分是编解码器标记器,它将高维自然信号压缩成低维离散标记。在本文中,我们介绍了WavTokenizer,它在音频域中提供了几个优于以前SOTA声学编解码器模型的优点:1)极端压缩。通过压缩量化器的层和离散编解码器的时间维度,24kHz采样率的一秒音频仅需要具有40或75个令牌的单个量化器。2)改善主观质量。尽管减少了令牌的数量,WavTokenizer仍以出色的UTMOS分数实现了最先进的重建质量,并且内在地包含了更丰富的语义信息。具体来说,我们通过设计更广泛的VQ空间,扩展的上下文窗口和改进的注意力网络,以及引入强大的多尺度搜索和逆傅立叶变换结构来实现这些结果。我们在语音、音频和音乐领域进行了广泛的重建实验。与最先进的模型相比,WavTokenizer在各种客观和主观指标上表现出强大的性能。我们还测试了语义信息,VQ的利用率和适应性生成模型。全面的消融研究证实了WavTokenizer中每个模块的必要性。相关代码、演示和预训练模型可在https: github.com jishengpeng WavTokenizer上获得。摘要:Language models have been effectively applied to modeling natural signals, such as images, video, speech, and audio. A crucial component of these models is the codec tokenizer, which compresses high-dimensional natural signals into lower-dimensional discrete tokens. In this paper, we introduce WavTokenizer, which offers several advantages over previous SOTA acoustic codec models in the audio domain: 1)extreme compression. By compressing the layers of quantizers and the temporal dimension of the discrete codec, one-second audio of 24kHz sampling rate requires only a single quantizer with 40 or 75 tokens. 2)improved subjective quality. Despite the reduced number of tokens, WavTokenizer achieves state-of-the-art reconstruction quality with outstanding UTMOS scores and inherently contains richer semantic information. Specifically, we achieve these results by designing a broader VQ space, extended contextual windows, and improved attention networks, as well as introducing a powerful multi-scale discriminator and an inverse Fourier transform structure. We conducted extensive reconstruction experiments in the domains of speech, audio, and music. WavTokenizer exhibited strong performance across various objective and subjective metrics compared to state-of-the-art models. We also tested semantic information, VQ utilization, and adaptability to generative models. Comprehensive ablation studies confirm the necessity of each module in WavTokenizer. The related code, demos, and pre-trained models are available at https: github.com jishengpeng WavTokenizer.

【14】 WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
标题: WHISMA:执行Zero-Shot口语理解的Speech-LLM
作者:Mohan Li,Cong-Thanh Do,Simon Keizer,Youmna Farag,Svetlana Stoyanchev,Rama Doddipatla
备注:accepted to SLT 2024
链接:点击下载PDF文件
摘要:语音大语言模型(Speech-LLM)集成了基于语音和文本的基础模型,为处理各种下游任务提供了统一的框架。在本文中,我们介绍WHISMA,一个语音LLM口语理解(SLU),在各种zero-shot设置中表现出强大的性能。WHISMA将Whisper的语音编码器与Llama-3 LLM相结合,并以参数高效的方式在全面收集的SLU相关数据集上进行微调。我们的实验表明,WHISMA显着提高了zero-shot插槽填充性能的SLURP基准,实现了26.6%的相对增益相比,目前的国家的最先进的模型。此外,为了评估WHISMA的泛化能力看不见的领域,我们开发了一个新的任务无关的基准命名为SLU-GLUE。评估结果表明,WHISMA优于现有的语音LLM(Qwen-Audio),相对增益为33.0%。摘要:Speech large language models (speech-LLMs) integrate speech and text-based foundation models to provide a unified framework for handling a wide range of downstream tasks. In this paper, we introduce WHISMA, a speech-LLM tailored for spoken language understanding (SLU) that demonstrates robust performance in various zero-shot settings. WHISMA combines the speech encoder from Whisper with the Llama-3 LLM, and is fine-tuned in a parameter-efficient manner on a comprehensive collection of SLU-related datasets. Our experiments show that WHISMA significantly improves the zero-shot slot filling performance on the SLURP benchmark, achieving a relative gain of 26.6% compared to the current state-of-the-art model. Furthermore, to evaluate WHISMA's generalisation capabilities to unseen domains, we develop a new task-agnostic benchmark named SLU-GLUE. The evaluation results indicate that WHISMA outperforms an existing speech-LLM (Qwen-Audio) with a relative gain of 33.0%.

【15】 Denoising of Photogrammetric Dummy Head Ear Point Clouds for Individual Head-Related Transfer Functions Computation
标题: 摄影测量假人头耳点云去噪以计算个体头部相关传递函数
作者:Fabio Di Giusto,Francesc Lluís,Sjoerd van Ophem,Elke Deckers
链接:点击下载PDF文件
摘要:个体头部相关传递函数(HRTF)对于逼真的虚拟音频渲染至关重要,可以在精确的三维头部和耳部扫描上有效地进行数值计算。虽然摄影测量扫描是有前途的,但它通常缺乏准确性,导致HRTF显示出与参考数据的显着感知偏差,由于扫描误差主要影响最闭塞的耳廓结构。本文分析了深度神经网络(DNN)用于摄影测量耳部扫描去噪的使用。各种DNN,微调耳廓样本损坏与建模的合成误差模仿摄影测量假人头部耳朵扫描中观察到的,测试和基准对一个经典的去噪方法。一个DNN被进一步修改和重新训练,以提高其去噪性能。原始和去噪扫描计算的HRTF与参考扫描的HRTF进行比较,表明性能最佳的DNN通常能够将摄影测量假人头部HRTF的偏差降低到用精确测量的个体数据获得的水平。在扫描的点云上计算的几何度量与相关的HRTF之间的相关性分析用于识别最相关的度量,以根据在目标扫描和参考扫描上计算的HRTF的相似性来评估目标扫描和参考扫描之间的几何偏差。摘要:Individual Head Related Transfer Functions (HRTFs), crucial for realistic virtual audio rendering, can be efficiently numerically computed on precise three-dimensional head and ear scans. While photogrammetry scanning is promising, it generally lacks in accuracy, leading to HRTFs showing significant perceptual deviation from reference data, owing to the scanning error mainly affecting the most occluded pinna structures. This papers analyses the use of Deep Neural Networks (DNNs) for denoising photogrammetric ear scans. Various DNNs, fine-tuned on pinna samples corrupted with modelled synthetic error mimicking that observed in photogrammetric dummy head ear scans, are tested and benchmarked against a classical denoising approach. One DNN is further modified and retrained to increase its denoising performance. The HRTFs computed on original and denoised scans are compared to those of a reference scan, showing that the best-performing DNN is capable of generally decreasing the deviation of photogrammetric dummy head HRTFs to levels obtained with accurately measured individual data. Correlation analysis between the geometrical metrics, computed on the scanned point clouds, and the related HRTFs is used to identify the most relevant metrics to assess the geometrical deviation between target and reference scans, in terms of the similarity of the HRTFs computed on them.

【16】 SSDM: Scalable Speech Dysfluency Modeling
标题: SSDP:可扩展语音流畅性建模
作者:Jiachen Lian,Xuanru Zhou,Zoe Ezzes,Jet Vonk,Brittany Morin,David Baquirin,Zachary Mille,Maria Luisa Gorno Tempini,Gopala Anumanchipalli
链接:点击下载PDF文件
摘要:言语不流利建模是口语学习和言语治疗的核心模块。然而,有三个挑战。首先,当前最先进的解决方案具有较差的可扩展性。第二,缺乏大规模的不流利语料库。第三,没有一个有效的学习框架。在本文中,我们提出了 textit{SSDM:可扩展的语音不流利建模},它(1)采用发音手势作为可扩展的强制对齐;(2)引入连接主义子序列对齐器(CSA)来实现不流利对齐;(3)引入一个名为Libri-Dys的大规模模拟不流利语料库;(4)通过利用大型语言模型(LLM)的功能开发一个端到端系统。我们希望SSDM能够成为流畅障碍建模领域的标准。演示可在 url{https: www.example.com}上获得。eureka235.github.io摘要:Speech dysfluency modeling is the core module for spoken language learning, and speech therapy. However, there are three challenges. First, current state-of-the-art solutions suffer from poor scalability. Second, there is a lack of a large-scale dysfluency corpus. Third, there is not an effective learning framework. In this paper, we propose textit{SSDM: Scalable Speech Dysfluency Modeling}, which (1) adopts articulatory gestures as scalable forced alignment; (2) introduces connectionist subsequence aligner (CSA) to achieve dysfluency alignment; (3) introduces a large-scale simulated dysfluency corpus called Libri-Dys; and (4) develops an end-to-end system by leveraging the power of large language models (LLMs). We expect SSDM to serve as a standard in the area of dysfluency modeling. Demo is available at url{https: eureka235.github.io}.

【17】 Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
标题: 在ASR-LLM设置上对日语语音识别进行基准测试,具有多遍增强生成式错误纠正
作者:Yuka Ko,Sheng Li,Chao-Han Huck Yang,Tatsuya Kawahara
备注:submitted to SLT2024
链接:点击下载PDF文件
摘要:利用大型语言模型(LLM)的强大代表性,自动语音识别(ASR)的生成纠错(GER)旨在提供语义和语音改进以解决ASR错误。这项工作探讨了基于LLM的GER如何增强和扩展日语语言处理的能力,提出了第一个GER基准日语ASR与0.9-2.6k文本话语。我们还介绍了一种新的多通道增强生成误差校正(MPA GER),通过将输入侧的多个系统假设与输出侧的多个LLM的校正相结合,然后将它们合并。据我们所知,这是第一次对LLM用于日语GER进行调查,其中涉及对ASR系统生成的输出transmittance进行二次语言建模(例如,N-best hypotheses)。我们的实验表明,在SPREDS-U1-JA和CSJ数据中,所提出的ASR质量和泛化方法的性能都有所提高。摘要:With the strong representational power of large language models (LLMs), generative error correction (GER) for automatic speech recognition (ASR) aims to provide semantic and phonetic refinements to address ASR errors. This work explores how LLM-based GER can enhance and expand the capabilities of Japanese language processing, presenting the first GER benchmark for Japanese ASR with 0.9-2.6k text utterances. We also introduce a new multi-pass augmented generative error correction (MPA GER) by integrating multiple system hypotheses on the input side with corrections from multiple LLMs on the output side and then merging them. To the best of our knowledge, this is the first investigation of the use of LLMs for Japanese GER, which involves second-pass language modeling on the output transcriptions generated by the ASR system (e.g., N-best hypotheses). Our experiments demonstrated performance improvement in the proposed methods of ASR quality and generalization both in SPREDS-U1-ja and CSJ data.

【18】 SVDD 2024: The Inaugural Singing Voice Deepfake Detection Challenge
标题: SDDD 2024:首届歌唱声音Deepfake检测挑战赛
作者:You Zhang,Yongyi Zang,Jiatong Shi,Ryuichi Yamamoto,Tomoki Toda,Zhiyao Duan
链接:点击下载PDF文件
摘要:随着歌声生成技术的进步以及人工智能歌手在媒体平台上的出现越来越多,首届歌声Deepfake检测(SVDD)挑战赛旨在推进识别人工智能生成的真实歌手歌声的研究。这个挑战有两个轨道:一个控制设置轨道(CtrSVDD)和一个在野外场景轨道(WildSVDD)。CtrSVDD音轨利用公开可用的歌唱声音数据,使用最先进的歌唱声音合成和转换系统生成deepfake。同时,WildSVDD轨道扩展了现有的SingFake数据集,其中包括来自流行的用户生成内容网站的数据。对于CtrSVDD赛道,我们收到了来自47个团队的提交,其中37个超过了我们的基线,最高团队的平均错误率为1.65%。对于WildSVDD跟踪,我们对基线进行了基准测试。本文回顾了这些结果,讨论了关键的发现,并概述了未来的方向SVDD研究。摘要:With the advancements in singing voice generation and the growing presence of AI singers on media platforms, the inaugural Singing Voice Deepfake Detection (SVDD) Challenge aims to advance research in identifying AI-generated singing voices from authentic singers. This challenge features two tracks: a controlled setting track (CtrSVDD) and an in-the-wild scenario track (WildSVDD). The CtrSVDD track utilizes publicly available singing vocal data to generate deepfakes using state-of-the-art singing voice synthesis and conversion systems. Meanwhile, the WildSVDD track expands upon the existing SingFake dataset, which includes data sourced from popular user-generated content websites. For the CtrSVDD track, we received submissions from 47 teams, with 37 surpassing our baselines and the top team achieving a 1.65% equal error rate. For the WildSVDD track, we benchmarked the baselines. This paper reviews these results, discusses key findings, and outlines future directions for SVDD research.


eess.AS音频处理
【1】 WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
标题: WavTokenizer:一种用于音频语言建模的高效声学离散编解码器Tokenizer
作者:Shengpeng Ji,Ziyue Jiang,Xize Cheng,Yifu Chen,Minghui Fang,Jialong Zuo,Qian Yang,Ruiqi Li,Ziang Zhang,Xiaoda Yang,Rongjie Huang,Yidi Jiang,Qian Chen,Siqi Zheng,Wen Wang,Zhou Zhao
备注:Working in progress. arXiv admin note: text overlap with arXiv:2402.12208
链接:点击下载PDF文件
摘要:语言模型已被有效地应用于对自然信号(诸如图像、视频、语音和音频)进行建模。这些模型的一个重要组成部分是编解码器标记器,它将高维自然信号压缩成低维离散标记。在本文中,我们介绍了WavTokenizer,它在音频域中提供了几个优于以前SOTA声学编解码器模型的优点:1)极端压缩。通过压缩量化器的层和离散编解码器的时间维度,24kHz采样率的一秒音频仅需要具有40或75个令牌的单个量化器。2)改善主观质量。尽管减少了令牌的数量,WavTokenizer仍以出色的UTMOS分数实现了最先进的重建质量,并且内在地包含了更丰富的语义信息。具体来说,我们通过设计更广泛的VQ空间,扩展的上下文窗口和改进的注意力网络,以及引入强大的多尺度搜索和逆傅立叶变换结构来实现这些结果。我们在语音、音频和音乐领域进行了广泛的重建实验。与最先进的模型相比,WavTokenizer在各种客观和主观指标上表现出强大的性能。我们还测试了语义信息,VQ的利用率和适应性生成模型。全面的消融研究证实了WavTokenizer中每个模块的必要性。相关代码、演示和预训练模型可在https: github.com jishengpeng WavTokenizer上获得。摘要:Language models have been effectively applied to modeling natural signals, such as images, video, speech, and audio. A crucial component of these models is the codec tokenizer, which compresses high-dimensional natural signals into lower-dimensional discrete tokens. In this paper, we introduce WavTokenizer, which offers several advantages over previous SOTA acoustic codec models in the audio domain: 1)extreme compression. By compressing the layers of quantizers and the temporal dimension of the discrete codec, one-second audio of 24kHz sampling rate requires only a single quantizer with 40 or 75 tokens. 2)improved subjective quality. Despite the reduced number of tokens, WavTokenizer achieves state-of-the-art reconstruction quality with outstanding UTMOS scores and inherently contains richer semantic information. Specifically, we achieve these results by designing a broader VQ space, extended contextual windows, and improved attention networks, as well as introducing a powerful multi-scale discriminator and an inverse Fourier transform structure. We conducted extensive reconstruction experiments in the domains of speech, audio, and music. WavTokenizer exhibited strong performance across various objective and subjective metrics compared to state-of-the-art models. We also tested semantic information, VQ utilization, and adaptability to generative models. Comprehensive ablation studies confirm the necessity of each module in WavTokenizer. The related code, demos, and pre-trained models are available at https: github.com jishengpeng WavTokenizer.

【2】 WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
标题: WHISMA:执行Zero-Shot口语理解的Speech-LLM
作者:Mohan Li,Cong-Thanh Do,Simon Keizer,Youmna Farag,Svetlana Stoyanchev,Rama Doddipatla
备注:accepted to SLT 2024
链接:点击下载PDF文件
摘要:语音大语言模型(Speech-LLM)集成了基于语音和文本的基础模型,为处理各种下游任务提供了统一的框架。在本文中,我们介绍WHISMA,一个语音LLM口语理解(SLU),在各种zero-shot设置中表现出强大的性能。WHISMA将Whisper的语音编码器与Llama-3 LLM相结合,并以参数高效的方式在全面收集的SLU相关数据集上进行微调。我们的实验表明,WHISMA显着提高了zero-shot插槽填充性能的SLURP基准,实现了26.6%的相对增益相比,目前的国家的最先进的模型。此外,为了评估WHISMA的泛化能力看不见的领域,我们开发了一个新的任务无关的基准命名为SLU-GLUE。评估结果表明,WHISMA优于现有的语音LLM(Qwen-Audio),相对增益为33.0%。摘要:Speech large language models (speech-LLMs) integrate speech and text-based foundation models to provide a unified framework for handling a wide range of downstream tasks. In this paper, we introduce WHISMA, a speech-LLM tailored for spoken language understanding (SLU) that demonstrates robust performance in various zero-shot settings. WHISMA combines the speech encoder from Whisper with the Llama-3 LLM, and is fine-tuned in a parameter-efficient manner on a comprehensive collection of SLU-related datasets. Our experiments show that WHISMA significantly improves the zero-shot slot filling performance on the SLURP benchmark, achieving a relative gain of 26.6% compared to the current state-of-the-art model. Furthermore, to evaluate WHISMA's generalisation capabilities to unseen domains, we develop a new task-agnostic benchmark named SLU-GLUE. The evaluation results indicate that WHISMA outperforms an existing speech-LLM (Qwen-Audio) with a relative gain of 33.0%.

【3】 Denoising of Photogrammetric Dummy Head Ear Point Clouds for Individual Head-Related Transfer Functions Computation
标题: 摄影测量假人头耳点云去噪以计算个体头部相关传递函数
作者:Fabio Di Giusto,Francesc Lluís,Sjoerd van Ophem,Elke Deckers
链接:点击下载PDF文件
摘要:个体头部相关传递函数(HRTF)对于逼真的虚拟音频渲染至关重要,可以在精确的三维头部和耳部扫描上有效地进行数值计算。虽然摄影测量扫描是有前途的,但它通常缺乏准确性,导致HRTF显示出与参考数据的显着感知偏差,由于扫描误差主要影响最闭塞的耳廓结构。本文分析了深度神经网络(DNN)用于对摄影测量耳扫描进行去噪的使用。各种DNN,微调耳廓样本损坏与建模的合成误差模仿摄影测量假人头部耳朵扫描中观察到的,测试和基准对一个经典的去噪方法。一个DNN被进一步修改和重新训练,以提高其去噪性能。原始和去噪扫描上计算的HRTF与参考扫描的HRTF进行比较,表明性能最佳的DNN能够将摄影测量假人头部HRTF的偏差降低到用精确测量的个体数据获得的水平。在扫描的点云上计算的几何度量与相关的HRTF之间的相关性分析用于识别最相关的度量,以根据在目标扫描和参考扫描上计算的HRTF的相似性来评估目标扫描和参考扫描之间的几何偏差。摘要:Individual Head Related Transfer Functions (HRTFs), crucial for realistic virtual audio rendering, can be efficiently numerically computed on precise three-dimensional head and ear scans. While photogrammetry scanning is promising, it generally lacks in accuracy, leading to HRTFs showing significant perceptual deviation from reference data, owing to the scanning error mainly affecting the most occluded pinna structures. This papers analyses the use of Deep Neural Networks (DNNs) for denoising photogrammetric ear scans. Various DNNs, fine-tuned on pinna samples corrupted with modelled synthetic error mimicking that observed in photogrammetric dummy head ear scans, are tested and benchmarked against a classical denoising approach. One DNN is further modified and retrained to increase its denoising performance. The HRTFs computed on original and denoised scans are compared to those of a reference scan, showing that the best-performing DNN is capable of generally decreasing the deviation of photogrammetric dummy head HRTFs to levels obtained with accurately measured individual data. Correlation analysis between the geometrical metrics, computed on the scanned point clouds, and the related HRTFs is used to identify the most relevant metrics to assess the geometrical deviation between target and reference scans, in terms of the similarity of the HRTFs computed on them.

【4】 SSDM: Scalable Speech Dysfluency Modeling
标题: SSDP:可扩展语音流畅性建模
作者:Jiachen Lian,Xuanru Zhou,Zoe Ezzes,Jet Vonk,Brittany Morin,David Baquirin,Zachary Mille,Maria Luisa Gorno Tempini,Gopala Anumanchipalli
链接:点击下载PDF文件
摘要:言语不流利建模是口语学习和言语治疗的核心模块。然而,有三个挑战。首先,当前最先进的解决方案具有较差的可扩展性。第二,缺乏大规模的不流利语料库。第三,没有一个有效的学习框架。在本文中,我们提出了 textit{SSDM:可扩展的语音不流利建模},它(1)采用发音手势作为可扩展的强制对齐;(2)引入连接主义子序列对齐器(CSA)来实现不流利对齐;(3)引入一个名为Libri-Dys的大规模模拟不流利语料库;(4)通过利用大型语言模型(LLM)的功能开发一个端到端系统。我们希望SSDM能够作为不流畅建模领域的标准。演示可在 url{https: www.example.com}上获得。eureka235.github.io摘要:Speech dysfluency modeling is the core module for spoken language learning, and speech therapy. However, there are three challenges. First, current state-of-the-art solutions suffer from poor scalability. Second, there is a lack of a large-scale dysfluency corpus. Third, there is not an effective learning framework. In this paper, we propose textit{SSDM: Scalable Speech Dysfluency Modeling}, which (1) adopts articulatory gestures as scalable forced alignment; (2) introduces connectionist subsequence aligner (CSA) to achieve dysfluency alignment; (3) introduces a large-scale simulated dysfluency corpus called Libri-Dys; and (4) develops an end-to-end system by leveraging the power of large language models (LLMs). We expect SSDM to serve as a standard in the area of dysfluency modeling. Demo is available at url{https: eureka235.github.io}.

【5】 Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
标题: 在ASR-LLM设置上对日语语音识别进行基准测试,具有多遍增强生成式错误纠正
作者:Yuka Ko,Sheng Li,Chao-Han Huck Yang,Tatsuya Kawahara
备注:submitted to SLT2024
链接:点击下载PDF文件
摘要:利用大型语言模型(LLM)的强大代表性,自动语音识别(ASR)的生成纠错(GER)旨在提供语义和语音改进以解决ASR错误。这项工作探讨了基于LLM的GER如何增强和扩展日语语言处理的能力,提出了第一个GER基准日语ASR与0.9-2.6k文本话语。我们还介绍了一种新的多通道增强生成误差校正(MPA GER),通过将输入侧的多个系统假设与输出侧的多个LLM的校正相结合,然后将它们合并。据我们所知,这是第一次对LLM用于日语GER进行调查,其中涉及对ASR系统生成的输出transmittance进行二次语言建模(例如,N-best hypotheses)。我们的实验表明,在SPREDS-U1-JA和CSJ数据中,所提出的ASR质量和泛化方法的性能都有所提高。摘要:With the strong representational power of large language models (LLMs), generative error correction (GER) for automatic speech recognition (ASR) aims to provide semantic and phonetic refinements to address ASR errors. This work explores how LLM-based GER can enhance and expand the capabilities of Japanese language processing, presenting the first GER benchmark for Japanese ASR with 0.9-2.6k text utterances. We also introduce a new multi-pass augmented generative error correction (MPA GER) by integrating multiple system hypotheses on the input side with corrections from multiple LLMs on the output side and then merging them. To the best of our knowledge, this is the first investigation of the use of LLMs for Japanese GER, which involves second-pass language modeling on the output transcriptions generated by the ASR system (e.g., N-best hypotheses). Our experiments demonstrated performance improvement in the proposed methods of ASR quality and generalization both in SPREDS-U1-ja and CSJ data.

【6】 SVDD 2024: The Inaugural Singing Voice Deepfake Detection Challenge
标题: SDDD 2024:首届歌唱声音Deepfake检测挑战赛
作者:You Zhang,Yongyi Zang,Jiatong Shi,Ryuichi Yamamoto,Tomoki Toda,Zhiyao Duan
链接:点击下载PDF文件
摘要:随着歌声生成技术的进步以及人工智能歌手在媒体平台上的出现越来越多,首届歌声Deepfake检测(SVDD)挑战赛旨在推进识别人工智能生成的真实歌手歌声的研究。这个挑战有两个轨道:一个控制设置轨道(CtrSVDD)和一个在野外场景轨道(WildSVDD)。CtrSVDD音轨利用公开可用的歌唱声音数据,使用最先进的歌唱声音合成和转换系统生成deepfake。同时,WildSVDD轨道扩展了现有的SingFake数据集,其中包括来自流行的用户生成内容网站的数据。对于CtrSVDD赛道,我们收到了来自47个团队的提交,其中37个超过了我们的基线,最高团队的平均错误率为1.65%。对于WildSVDD跟踪,我们对基线进行了基准测试。本文回顾了这些结果,讨论了关键的发现,并概述了未来的方向SVDD研究。摘要:With the advancements in singing voice generation and the growing presence of AI singers on media platforms, the inaugural Singing Voice Deepfake Detection (SVDD) Challenge aims to advance research in identifying AI-generated singing voices from authentic singers. This challenge features two tracks: a controlled setting track (CtrSVDD) and an in-the-wild scenario track (WildSVDD). The CtrSVDD track utilizes publicly available singing vocal data to generate deepfakes using state-of-the-art singing voice synthesis and conversion systems. Meanwhile, the WildSVDD track expands upon the existing SingFake dataset, which includes data sourced from popular user-generated content websites. For the CtrSVDD track, we received submissions from 47 teams, with 37 surpassing our baselines and the top team achieving a 1.65% equal error rate. For the WildSVDD track, we benchmarked the baselines. This paper reviews these results, discusses key findings, and outlines future directions for SVDD research.

【7】 Automatic detection of Mild Cognitive Impairment using high-dimensional acoustic features in spontaneous speech
标题: 使用自发言语中的多维声学特征自动检测轻度认知障碍
作者:Cong Zhang,Wenxing Guo,Hongsheng Dai
链接:点击下载PDF文件
摘要:本研究解决了TAUKADIAL挑战,重点是轻度认知障碍(MCI)和神经典型对照人群的语音分类。我们进行了三个实验,比较了五种机器学习方法:随机森林,稀疏逻辑回归,k最近邻,稀疏支持向量机和决策树,利用openSMILE自动提取的1076个声学特征。在实验1中,整个数据集被用来训练语言无关的模型。实验2引入了语言检测步骤,导致每种语言的单独模型训练。实验3进一步增强了实验1的语言不可知模型,特别关注使用样本外测试数据评估模型的鲁棒性。在所有三个实验中,结果一致有利于能够处理高维数据的模型,如随机森林和稀疏逻辑回归,在分类语音MCI和控制。摘要:This study addresses the TAUKADIAL challenge, focusing on the classification of speech from people with Mild Cognitive Impairment (MCI) and neurotypical controls. We conducted three experiments comparing five machine-learning methods: Random Forests, Sparse Logistic Regression, k-Nearest Neighbors, Sparse Support Vector Machine, and Decision Tree, utilizing 1076 acoustic features automatically extracted using openSMILE. In Experiment 1, the entire dataset was used to train a language-agnostic model. Experiment 2 introduced a language detection step, leading to separate model training for each language. Experiment 3 further enhanced the language-agnostic model from Experiment 1, with a specific focus on evaluating the robustness of the models using out-of-sample test data. Across all three experiments, results consistently favored models capable of handling high-dimensional data, such as Random Forest and Sparse Logistic Regression, in classifying speech from MCI and controls.

【8】 Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
标题: Mini-Omni:语言模型可以在流媒体中一边思考一边听、说话
作者:Zhifei Xie,Changqiao Wu
备注:10 pages
链接:点击下载PDF文件
摘要:语言模型的最新进展取得了重大进展。GPT-4 o作为一个新的里程碑,实现了与人类的实时对话,展示了接近人类的自然流畅性。这种人机交互需要具有直接使用音频模态执行推理并在流中生成输出的能力的模型。然而,这仍然超出了当前学术模型的范围,因为它们通常依赖于额外的TTS系统进行语音合成,从而导致不期望的延迟。本文介绍了Mini-Omni,一个基于音频的端到端会话模型,能够实时语音交互。为了实现这种能力,我们提出了一种文本指导的语音生成方法,以及在推理过程中的批并行策略,以进一步提高性能。我们的方法还有助于保留原始模型的语言能力,最小的退化,使其他作品建立实时交互能力。我们称这种训练方法为“任何模型都能说话”。我们还引入了VoiceAssistant-400 K数据集来微调针对语音输出优化的模型。据我们所知,Mini-Omni是第一个完全端到端的实时语音交互开源模型,为未来的研究提供了宝贵的潜力。摘要:Recent advances in language models have achieved significant progress. GPT-4o, as a new milestone, has enabled real-time conversations with humans, demonstrating near-human natural fluency. Such human-computer interaction necessitates models with the capability to perform reasoning directly with the audio modality and generate output in streaming. However, this remains beyond the reach of current academic models, as they typically depend on extra TTS systems for speech synthesis, resulting in undesirable latency. This paper introduces the Mini-Omni, an audio-based end-to-end conversational model, capable of real-time speech interaction. To achieve this capability, we propose a text-instructed speech generation method, along with batch-parallel strategies during inference to further boost the performance. Our method also helps to retain the original model's language capabilities with minimal degradation, enabling other works to establish real-time interaction capabilities. We call this training method "Any Model Can Talk". We also introduce the VoiceAssistant-400K dataset to fine-tune models optimized for speech output. To our best knowledge, Mini-Omni is the first fully end-to-end, open-source model for real-time speech interaction, offering valuable potential for future research.

【9】 Towards Efficient Modelling of String Dynamics: A Comparison of State Space and Koopman based Deep Learning Methods
标题: 实现字符串动力学的高效建模:状态空间和基于Koopman的深度学习方法的比较
作者:Rodrigo Diaz,Carlos De La Vega Martin,Mark Sandler
备注:Accepted to DAFx2024
链接:点击下载PDF文件
摘要:本文介绍了状态空间模型(SSM)和基于Koopman的深度学习方法,用于对线性和非线性刚性弦的动力学进行建模。通过对不同初始条件和采样率下生成的数据集进行实验,我们评估了这些模型准确模拟弦动力学中观察到的复杂行为的能力。我们的研究结果表明,我们提出的Koopman为基础的模型执行以及或优于其他现有的方法在非线性情况下的长序列建模。 我们通知这些架构的设计与手头的问题的结构。尽管在将模型预测扩展到训练范围之外(即,外推),我们研究的重点在于模型在训练时间间隔内跨不同初始条件进行概括的能力。这项研究有助于深入了解动力系统的物理建模(特别是那些解决音乐声学),通过提供这些和以前的方法的比较概述,并引入模型改进的创新策略。我们的研究结果突出了这些模型在模拟非线性动力学的有效性,并强调其广泛的适用性,在扩展序列的动力系统准确建模。摘要:This paper presents an examination of State Space Models (SSM) and Koopman-based deep learning methods for modelling the dynamics of both linear and non-linear stiff strings. Through experiments with datasets generated under different initial conditions and sample rates, we assess the capacity of these models to accurately model the complex behaviours observed in string dynamics. Our findings indicate that our proposed Koopman-based model performs as well as or better than other existing approaches in non-linear cases for long-sequence modelling. We inform the design of these architectures with the structure of the problems at hand. Although challenges remain in extending model predictions beyond the training horizon (i.e., extrapolation), the focus of our investigation lies in the models' ability to generalise across different initial conditions within the training time interval. This research contributes insights into the physical modelling of dynamical systems (in particular those addressing musical acoustics) by offering a comparative overview of these and previous methods and introducing innovative strategies for model improvement. Our results highlight the efficacy of these models in simulating non-linear dynamics and emphasise their wide-ranging applicability in accurately modelling dynamical systems over extended sequences.

【10】 Audio xLSTMs: Learning Self-supervised audio representations with xLSTMs
标题: 音频xLSTM:使用xLSTM学习自我监督的音频表示
作者:Sarthak Yadav,Sergios Theodoridis,Zheng-Hua Tan
备注:Under review at ICASSP 2025. arXiv admin note: text overlap with arXiv:2406.02178
链接:点击下载PDF文件
摘要:虽然Transformer已经成为杰出的神经架构,但已经出现了几个独立的研究路线来解决其局限性。循环神经方法也引起了很多新的兴趣,包括扩展的长短期记忆(xLSTM)架构,它重振了原始的LSTM架构。然而,虽然xLSTM与Transformer相比表现出了竞争力,但它们用于学习自监督通用音频表示的可行性尚未得到评估。这项工作提出了Audio xLSTM(AxLSTM),这是一种在自监督设置中从掩蔽的频谱图补丁中学习音频表示的方法。在AudioSet数据集上进行预训练后,所提出的AxLSTM模型在一组10个不同的下游任务中的相对性能比可比的自监督音频频谱图Transformer(SSAST)基线高出20%,同时参数减少了45%。摘要:While the transformer has emerged as the eminent neural architecture, several independent lines of research have emerged to address its limitations. Recurrent neural approaches have also observed a lot of renewed interest, including the extended long short-term memory (xLSTM) architecture, which reinvigorates the original LSTM architecture. However, while xLSTMs have shown competitive performance compared to the transformer, their viability for learning self-supervised general-purpose audio representations has not yet been evaluated. This work proposes Audio xLSTM (AxLSTM), an approach to learn audio representations from masked spectrogram patches in a self-supervised setting. Pretrained on the AudioSet dataset, the proposed AxLSTM models outperform comparable self-supervised audio spectrogram transformer (SSAST) baselines by up to 20% in relative performance across a set of ten diverse downstream tasks while having up to 45% fewer parameters.

【11】 Human-Inspired Audio-Visual Speech Recognition: Spike Activity, Cueing Interaction and Causal Processing
标题: 受人类启发的视听语音识别:尖峰活动、线索交互和因果处理
作者:Qianhui Liu,Jiadong Wang,Yang Wang,Xin Yang,Gang Pan,Haizhou Li
链接:点击下载PDF文件
摘要:人类自然地执行视听语音识别(AVSR),通过整合听觉和视觉信息来提高准确性和鲁棒性。尖峰神经网络(SNN)模仿大脑的信息处理机制,非常适合模拟人类的AVSR能力。尽管SNN具有潜力,但针对AVSR的SNN研究却很少,大多数现有的视听多模式方法都集中在对象或数字识别上。这些模型简单地整合了两种模式的特征,忽略了它们的独特特征和相互作用。此外,它们通常依赖于未来的信息进行当前处理,这增加了识别延迟并限制了实时适用性。受人类语音感知的启发,本文提出了一种新的人类启发的SNN命名为HI-AVSNN的AVSR,结合了三个关键特征:提示交互,因果处理和尖峰活动。对于线索交互,我们提出了一个视觉线索听觉注意模块(VCA 2 M),利用视觉线索来引导注意听觉功能。我们实现因果关系的处理对齐SNN的时间维度与视觉和听觉功能,并应用时间掩蔽,只利用过去和当前的信息。为了实现尖峰活动,除了使用SNN之外,我们还利用事件相机来捕获嘴唇运动作为尖峰,模仿人类视网膜并提供有效的视觉数据。我们评估HI-AVSNN的视听语音识别数据集相结合的DVS-Lip数据集与其相应的音频样本。实验结果表明,我们提出的融合方法的优越性,优于现有的视听SNN融合方法,并实现了2.27%的提高精度比现有的基于SNN的AVSR方法。摘要:Humans naturally perform audiovisual speech recognition (AVSR), enhancing the accuracy and robustness by integrating auditory and visual information. Spiking neural networks (SNNs), which mimic the brain's information-processing mechanisms, are well-suited for emulating the human capability of AVSR. Despite their potential, research on SNNs for AVSR is scarce, with most existing audio-visual multimodal methods focused on object or digit recognition. These models simply integrate features from both modalities, neglecting their unique characteristics and interactions. Additionally, they often rely on future information for current processing, which increases recognition latency and limits real-time applicability. Inspired by human speech perception, this paper proposes a novel human-inspired SNN named HI-AVSNN for AVSR, incorporating three key characteristics: cueing interaction, causal processing and spike activity. For cueing interaction, we propose a visual-cued auditory attention module (VCA2M) that leverages visual cues to guide attention to auditory features. We achieve causal processing by aligning the SNN's temporal dimension with that of visual and auditory features and applying temporal masking to utilize only past and current information. To implement spike activity, in addition to using SNNs, we leverage the event camera to capture lip movement as spikes, mimicking the human retina and providing efficient visual data. We evaluate HI-AVSNN on an audiovisual speech recognition dataset combining the DVS-Lip dataset with its corresponding audio samples. Experimental results demonstrate the superiority of our proposed fusion method, outperforming existing audio-visual SNN fusion methods and achieving a 2.27% improvement in accuracy over the only existing SNN-based AVSR method.

【12】 RAVE for Speech: Efficient Voice Conversion at High Sampling Rates
标题: RAVE for Speech:高采样率下的高效语音转换
作者:Anders R. Bargum,Simon Lajboschitz,Cumhur Erkut
备注:Accepted for publication in Proceedings of the 27th International Conference on Digital Audio Effects (DAFx24), Guildford, United Kingdom, 3 - 7 September 2024
链接:点击下载PDF文件
摘要:语音转换在音频处理和语音合成领域中越来越受欢迎。通常,主要目标是将输入身份转换为目标说话者的身份,而不改变其语言内容。虽然目前的工作提供了高保真的解决方案,他们很少关注模型的简单性,高采样率的环境或流的能力。通过将语音表示学习到一个生成的音色转移模型,传统上创建的音乐目的,我们调查的领域直接在时域中产生的高采样率的语音转换。更具体地说,我们将基线模型的潜在空间引导到语言相关的表示,并将其置于外部说话者信息上。通过客观和主观的评估,我们表明,所提出的解决方案可以达到的自然度,质量和可理解性的水平相比,一个国家的最先进的解决方案看到的扬声器,同时显着减少推理时间。然而,尽管转换后的输出中存在目标说话者的特征,但与未见过的说话者的实际相似性仍然是一个挑战。摘要:Voice conversion has gained increasing popularity within the field of audio manipulation and speech synthesis. Often, the main objective is to transfer the input identity to that of a target speaker without changing its linguistic content. While current work provides high-fidelity solutions they rarely focus on model simplicity, high-sampling rate environments or stream-ability. By incorporating speech representation learning into a generative timbre transfer model, traditionally created for musical purposes, we investigate the realm of voice conversion generated directly in the time domain at high sampling rates. More specifically, we guide the latent space of a baseline model towards linguistically relevant representations and condition it on external speaker information. Through objective and subjective assessments, we demonstrate that the proposed solution can attain levels of naturalness, quality, and intelligibility comparable to those of a state-of-the-art solution for seen speakers, while significantly decreasing inference time. However, despite the presence of target speaker characteristics in the converted output, the actual similarity to unseen speakers remains a challenge.

【13】 SALSA: Speedy ASR-LLM Synchronous Aggregation
标题: SALSA:快速ASR-LLM同步聚合
作者:Ashish Mittal,Darshan Prabhu,Sunita Sarawagi,Preethi Jyothi
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:利用预先训练的LLM来改善ASR系统,特别是低资源语言,现在是一个新兴的研究领域。现有方法的范围从使用LLM进行ASR纠错到用LLM代替ASR解码器的紧密耦合系统。这些方法要么增加解码时间,要么需要对交叉注意层进行昂贵的训练。我们提出了SALSA,它将ASR的解码器层耦合到LLM解码器,同时同步推进两个解码器。这样的耦合是用最后一个解码器状态的简单投影来执行的,因此比早期的方法明显更有训练效率。我们提出的耦合的一个挑战是处理LLM和ASR系统的标记器之间的不匹配。我们使用与LLM和ASR词汇表相关的级联标记化来处理这种不匹配。我们在FLEURS基准测试中对8种低资源语言进行了SALSA评估,产生了高达38%的WER大幅降低。摘要:Harnessing pre-trained LLMs to improve ASR systems, particularly for low-resource languages, is now an emerging area of research. Existing methods range from using LLMs for ASR error correction to tightly coupled systems that replace the ASR decoder with the LLM. These approaches either increase decoding time or require expensive training of the cross-attention layers. We propose SALSA, which couples the decoder layers of the ASR to the LLM decoder, while synchronously advancing both decoders. Such coupling is performed with a simple projection of the last decoder state, and is thus significantly more training efficient than earlier approaches. A challenge of our proposed coupling is handling the mismatch between the tokenizers of the LLM and ASR systems. We handle this mismatch using cascading tokenization with respect to the LLM and ASR vocabularies. We evaluate SALSA on 8 low-resource languages in the FLEURS benchmark, yielding substantial WER reductions of up to 38%.

【14】 Enabling Beam Search for Language Model-Based Text-to-Speech Synthesis
标题: 启用基于语言模型的文本到语音合成的梁搜索
作者:Zehai Tu,Guangyan Zhang,Yiting Lu,Adaeze Adigwe,Simon King,Yiwen Guo
链接:点击下载PDF文件
摘要:将连续语音标记为离散标记序列并使用语言模型(LM)对其进行建模已经在文本到语音(TTS)合成中取得了重大成功。虽然这些模型可以生成高质量和自然的语音,他们的合成样本仍然可以遭受人工制品,发音错误,单词重复等,在本文中,我们认为这些不良属性可能部分是由随机性的采样为基础的策略在自回归解码的LM。因此,我们着眼于基于最大化的解码方法,并提出时间重复感知的多样化波束搜索(TRAD-BS),以找到最可能的序列生成的语音令牌。两个国家的最先进的LM为基础的TTS模型的实验表明,我们提出的最大化为基础的解码策略生成的语音具有更少的发音错误和提高说话人的一致性。摘要:Tokenising continuous speech into sequences of discrete tokens and modelling them with language models (LMs) has led to significant success in text-to-speech (TTS) synthesis. Although these models can generate speech with high quality and naturalness, their synthesised samples can still suffer from artefacts, mispronunciation, word repeating, etc. In this paper, we argue these undesirable properties could partly be caused by the randomness of sampling-based strategies during the autoregressive decoding of LMs. Therefore, we look at maximisation-based decoding approaches and propose Temporal Repetition Aware Diverse Beam Search (TRAD-BS) to find the most probable sequences of the generated speech tokens. Experiments with two state-of-the-art LM-based TTS models demonstrate that our proposed maximisation-based decoding strategy generates speech with fewer mispronunciations and improved speaker consistency.

【15】 Measuring the Accuracy of Automatic Speech Recognition Solutions
标题: 衡量自动语音识别解决方...性
作者:Korbinian Kuhn,Verena Kersken,Benedikt Reuter,Niklas Egger,Gottfried Zimmermann
Journal-ref:ACM Transactions on Accessible Computing, Volume 16, Issue 4, Article 25 (2023), 1-23
链接:点击下载PDF文件
摘要:d 对于聋人和重听人来说,字幕是一个必不可少的无障碍工具。人工智能(AI)的重大发展意味着自动语音识别(ASR)现在是许多流行应用的一部分。这使得创建字幕变得容易和广泛可用-但转录需要高水平的准确性才能访问。科学出版物和行业报告的错误率非常低,声称人工智能已经达到了人类的水平,甚至超过了人工转录。与此同时,DHH社区报告了ASR准确性和可靠性的严重问题。对于依赖转录的人来说,技术创新和现实生活体验之间似乎存在不匹配。需要独立和全面的数据来捕获ASR的状态。我们测量了11个常见的ASR服务与高等教育讲座录音的性能。我们评估了技术条件的影响,如流,词汇的使用和语言之间的差异。我们的研究结果表明,准确性范围广泛的供应商和个人的音频样本。我们还测量了用于直播活动的流媒体ASR的质量明显较低。我们的研究表明,尽管最近的ASR的改进,公共服务缺乏准确性的可靠性。摘要:For d Deaf and hard of hearing (DHH) people, captioning is an essential accessibility tool. Significant developments in artificial intelligence (AI) mean that Automatic Speech Recognition (ASR) is now a part of many popular applications. This makes creating captions easy and broadly available - but transcription needs high levels of accuracy to be accessible. Scientific publications and industry report very low error rates, claiming AI has reached human parity or even outperforms manual transcription. At the same time the DHH community reports serious issues with the accuracy and reliability of ASR. There seems to be a mismatch between technical innovations and the real-life experience for people who depend on transcription. Independent and comprehensive data is needed to capture the state of ASR. We measured the performance of eleven common ASR services with recordings of Higher Education lectures. We evaluated the influence of technical conditions like streaming, the use of vocabularies, and differences between languages. Our results show that accuracy ranges widely between vendors and for the individual audio samples. We also measured a significant lower quality for streaming ASR, which is used for live events. Our study shows that despite the recent improvements of ASR, common services lack reliability in accuracy.

【16】 Revisit Micro-batch Clipping: Adaptive Data Pruning via Gradient Manipulation
标题: 重温微批剪辑:通过梯度操纵进行自适应数据修剪
作者:Lun Wang
链接:点击下载PDF文件
摘要:微批量裁剪,梯度裁剪方法,最近显示出潜在的提高自动语音识别(ASR)模型的性能。然而,这种改进背后的潜在机制仍然是神秘的,特别是观察到只有某些微批量是有益的。在本文中,我们首次尝试解释这一现象。受最近数据修剪研究的启发,我们假设特定的训练样本可能会在某些训练阶段阻碍模型收敛。在此假设下,收敛性分析表明,微批量裁剪可以以不随训练迭代次数增加而减小的额外常数偏差为代价,渐进地提高收敛速度。偏倚取决于几个因素,并且可以在特定的微批量下最小化,从而阐明先前观察到的最佳点微批量的存在。我们还验证了语音模型之外的视觉和语言模型上的微批量裁剪的有效性,并在这些领域显示出有前途的性能增益。对潜在限制的探索表明,当训练数据来自多个不同的领域时,微批量裁剪的效果较差。摘要:Micro-batch clipping, a gradient clipping method, has recently shown potential in enhancing auto-speech recognition (ASR) model performance. However, the underlying mechanism behind this improvement remains mysterious, particularly the observation that only certain micro-batch sizes are beneficial. In this paper, we make the first attempt to explain this phenomenon. Inspired by recent data pruning research, we assume that specific training samples may impede model convergence during certain training phases. Under this assumption, the convergence analysis shows that micro-batch clipping can improve the convergence rate asymptotically at the cost of an additional constant bias that does not diminish with more training iterations. The bias is dependent on a few factors and can be minimized at specific micro-batch size, thereby elucidating the existence of the sweet-spot micro-batch size observed previously. We also verify the effectiveness of micro-batch clipping beyond speech models on vision and language models, and show promising performance gains in these domains. An exploration of potential limitations shows that micro-batch clipping is less effective when training data originates from multiple distinct domains.

【17】 Improving Generalization of Speech Separation in Real-World Scenarios: Strategies in Simulation, Optimization, and Evaluation
标题: 改善现实场景中语音分离的推广:模拟、优化和评估策略
作者:Ke Chen,Jiaqi Su,Taylor Berg-Kirkpatrick,Shlomo Dubnov,Zeyu Jin
备注:In Proceedings of the 25th Annual Conference of the International Speech Communication Association, Interspeech 2024
链接:点击下载PDF文件
摘要:在具有噪声和混响的各种声学环境中实现对重叠扬声器的鲁棒语音分离仍然是一个公开的挑战。尽管现有的数据集可用于针对特定场景训练分离器,但是它们不能有效地在不同的现实世界场景中推广。在本文中,我们提出了一种新的数据模拟管道,从一系列声学环境和内容中产生不同的训练数据,并提出了新的训练范例,以提高一般语音分离模型的质量。具体来说,我们首先介绍AC-SIM,这是一种数据模拟管道,它包含了内容和声学的广泛变化。然后,我们将多个训练目标集成到排列不变训练(PIT)中,以提高训练模型的分离质量和泛化能力。最后,我们在分离架构和基准测试中进行了全面的客观和人类听觉实验,以验证我们的方法,证明了非同源和真实世界测试集的泛化能力有了实质性的提高。摘要:Achieving robust speech separation for overlapping speakers in various acoustic environments with noise and reverberation remains an open challenge. Although existing datasets are available to train separators for specific scenarios, they do not effectively generalize across diverse real-world scenarios. In this paper, we present a novel data simulation pipeline that produces diverse training data from a range of acoustic environments and content, and propose new training paradigms to improve quality of a general speech separation model. Specifically, we first introduce AC-SIM, a data simulation pipeline that incorporates broad variations in both content and acoustics. Then we integrate multiple training objectives into the permutation invariant training (PIT) to enhance separation quality and generalization of the trained model. Finally, we conduct comprehensive objective and human listening experiments across separation architectures and benchmarks to validate our methods, demonstrating substantial improvement of generalization on both non-homologous and real-world test sets.

【18】 A Deep Learning Approach to Localizing Multi-level Airway Collapse Based on Snoring Sounds
标题: 基于打鼾声音定位多层气道塌陷的深度学习方法
作者:Ying-Chieh Hsu,Stanley Yung-Chuan Liu,Chao-Jung Huang,Chi-Wei Wu,Ren-Kai Cheng,Jane Yung-Jen Hsu,Shang-Ran Huang,Yuan-Ren Cheng,Fu-Shun Hsu
链接:点击下载PDF文件
摘要:本研究利用药物诱导睡眠内窥镜检查(DISE)的数据,研究了机器 深度学习在阻塞性睡眠呼吸暂停(OSA)患者上气道不同水平激发的打鼾声音分类中的应用。根据腭、口咽、舌根和会厌(VOTE)分类系统,对39名受试者的鼾声进行分析和标记。该数据集包括5,173个一秒片段,用于训练和测试模型,包括支持向量机(SVM),双向长短期记忆(BiLSTM)和ResNet-50。ResNet-50是一种卷积神经网络(CNN),在对打鼾声音进行分类方面表现出最佳的整体性能,特别是在识别多级障碍物方面。该研究强调了将打鼾声学与深度学习相结合以改善阻塞性睡眠呼吸暂停综合症诊断和治疗的潜力。然而,注意到诸如样本量有限、数据不平衡以及打鼾引起的声音与自然打鼾声音之间的差异等挑战,这表明需要进一步研究以提高模型的准确性和可推广性。摘要:This study investigates the application of machine deep learning to classify snoring sounds excited at different levels of the upper airway in patients with obstructive sleep apnea (OSA) using data from drug-induced sleep endoscopy (DISE). The snoring sounds of 39 subjects were analyzed and labeled according to the Velum, Oropharynx, Tongue Base, and Epiglottis (VOTE) classification system. The dataset, comprising 5,173 one-second segments, was used to train and test models, including Support Vector Machine (SVM), Bidirectional Long Short-Term Memory (BiLSTM), and ResNet-50. The ResNet-50, a convolutional neural network (CNN), showed the best overall performance in classifying snoring acoustics, particularly in identifying multi-level obstructions. The study emphasizes the potential of integrating snoring acoustics with deep learning to improve the diagnosis and treatment of OSA. However, challenges such as limited sample size, data imbalance, and differences between pharmacologically induced and natural snoring sounds were noted, suggesting further research to enhance model accuracy and generalizability.


机器翻译,仅供参考