今日论文合集:cs.SD语音11篇,eess.AS音频处理7篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】SpeechAlign: Aligning Speech Generation to Human Preferences
标题:SpeechAlign:将语音生成与人类偏好相匹配
链接:https://arxiv.org/abs/2404.05600
作者:Dong Zhang,Zhaowei Li,Shimin Li,Xin Zhang,Pengyu Wang,Yaqian Zhou,Xipeng Qiu
备注:Work in progress
摘要:语音语言模型在生成真实语音方面取得了显着进步,其中神经编解码器语言模型脱颖而出。然而,人类反馈的整合,以调整语音输出人类的喜好往往被忽视。本文通过首先分析编解码器语言模型中的分布差距来解决这一差距,强调它如何导致训练和推理阶段之间的差异,从而对性能产生负面影响。然后,我们探索利用人类反馈的学习来弥合分配差距。我们引入SpeechAlign,这是一种迭代的自我改进策略,可以将语音语言模型与人类偏好相匹配。SpeechAlign涉及构建一个偏好编解码器数据集,将黄金编解码器令牌与合成令牌进行对比,然后进行偏好优化以改进编解码器语言模型。这个改进周期是迭代进行的,以稳定地将弱模型转换为强模型。通过主观和客观的评价,我们表明SpeechAlign可以弥合分布差距,促进语音语言模型的不断自我改进。此外,SpeechAlign具有强大的泛化能力,适用于较小的模型。代码和模型将在https://github.com/0nutation/SpeechGPT上提供。
摘要:Speech language models have significantly advanced in generating realistic speech, with neural codec language models standing out. However, the integration of human feedback to align speech outputs to human preferences is often neglected. This paper addresses this gap by first analyzing the distribution gap in codec language models, highlighting how it leads to discrepancies between the training and inference phases, which negatively affects performance. Then we explore leveraging learning from human feedback to bridge the distribution gap. We introduce SpeechAlign, an iterative self-improvement strategy that aligns speech language models to human preferences. SpeechAlign involves constructing a preference codec dataset contrasting golden codec tokens against synthetic tokens, followed by preference optimization to improve the codec language model. This cycle of improvement is carried out iteratively to steadily convert weak models to strong ones. Through both subjective and objective evaluations, we show that SpeechAlign can bridge the distribution gap and facilitating continuous self-improvement of the speech language model. Moreover, SpeechAlign exhibits robust generalization capabilities and works for smaller models. Code and models will be available at https://github.com/0nutation/SpeechGPT.


【2】 SoundingActions: Learning How Actions Sound from Narrated Egocentric  Videos
标题:SoundingAction:学习如何从叙事自我中心视频中听到动作的声音
链接:https://arxiv.org/abs/2404.05206
作者:Changan Chen,Kumar Ashutosh,Rohit Girdhar,David Harwath,Kristen Grauman
备注:Accepted at CVPR 2024. Project page: this https URL
摘要:我们提出了一种新的自我监督嵌入,以学习如何从叙述在野生自我中心的视频动作的声音。虽然现有的方法依赖于具有已知视听对应关系的策划数据,但我们的多模态对比一致性编码(MC 3)嵌入在所有模态对一致时加强了音频,语言和视觉之间的关联,同时在任何一对不一致时减少了这些关联。我们证明了我们的方法可以成功地发现人类行为的长尾如何从以自我为中心的视频中发出声音,在两个数据集(Ego 4D和EPIC-Sounds)和多个跨模态任务上的表现优于最近的一系列多模态嵌入技术。
摘要:We propose a novel self-supervised embedding to learn how actions sound from narrated in-the-wild egocentric videos. Whereas existing methods rely on curated data with known audio-visual correspondence, our multimodal contrastive-consensus coding (MC3) embedding reinforces the associations between audio, language, and vision when all modality pairs agree, while diminishing those associations when any one pair does not. We show our approach can successfully discover how the long tail of human actions sound from egocentric video, outperforming an array of recent multimodal embedding techniques on two datasets (Ego4D and EPIC-Sounds) and multiple cross-modal tasks.

【3】 Test-Time Training for Depression Detection
标题:抑郁症检测的测试时间训练
链接:https://arxiv.org/abs/2404.05071
作者:Sri Harsha Dumpala,Chandramouli Shama Sastry,Rudolf Uher,Sageev Oore
摘要:先前关于抑郁检测的工作使用在类似环境中收集的数据集来训练和测试模型。然而,在实践中,不能保证训练和测试分布相同。由于诸如记录环境(例如,背景噪声)和人口统计(例如,性别、年龄等)。这种分布变化可能会令人惊讶地导致抑郁症检测模型的严重性能下降。在本文中,我们分析了测试时间训练(TTT)的应用,以提高抑郁症检测训练模型的鲁棒性。当与模型的常规测试相比时,我们发现TTT可以显着提高模型在由于以下原因引入的各种分布偏移下的鲁棒性:(a)背景噪声,(b)性别偏见,以及(c)数据收集和策展程序(即,训练和测试样本来自不同数据集)。
摘要:Previous works on depression detection use datasets collected in similar environments to train and test the models. In practice, however, the train and test distributions cannot be guaranteed to be identical. Distribution shifts can be introduced due to variations such as recording environment (e.g., background noise) and demographics (e.g., gender, age, etc). Such distributional shifts can surprisingly lead to severe performance degradation of the depression detection models. In this paper, we analyze the application of test-time training (TTT) to improve robustness of models trained for depression detection. When compared to regular testing of the models, we find TTT can significantly improve the robustness of the model under a variety of distributional shifts introduced due to: (a) background-noise, (b) gender-bias, and (c) data collection and curation procedure (i.e., train and test samples are from separate datasets).


【4】 Cross-Domain Audio Deepfake Detection: Dataset and Analysis
标题:跨域音频Deepfake检测:数据集和分析
链接:https://arxiv.org/abs/2404.04904
作者:Yuang Li,Min Zhang,Mengxin Ren,Miaomiao Ma,Daimeng Wei,Hao Yang
摘要:音频深度伪造检测(ADD)对于防止可能侵犯个人权利和隐私的合成语音的滥用至关重要。最近的zero-shot文本到语音(TTS)模型带来了更高的风险,因为它们可以用一个单一的话语克隆声音。然而,现有的ADD数据集是过时的,导致检测模型的次优泛化。在本文中,我们构建了一个新的跨域ADD数据集,包括超过300小时的语音数据,由五个先进的zero-shot TTS模型生成。为了模拟真实世界的场景,我们采用了不同的攻击方法和来自不同数据集的音频提示。实验结果表明,通过新的攻击增强训练,Wav 2 Vec 2-large和Whisper-medium模型的错误率分别为4.1%和6.5%。此外,我们通过仅用一分钟的目标域数据进行微调,展示了我们的模型出色的Few-Shot ADD能力。然而,神经编解码器压缩器极大地影响检测精度,需要进一步的研究。
摘要:Audio deepfake detection (ADD) is essential for preventing the misuse of synthetic voices that may infringe on personal rights and privacy. Recent zero-shot text-to-speech (TTS) models pose higher risks as they can clone voices with a single utterance. However, the existing ADD datasets are outdated, leading to suboptimal generalization of detection models. In this paper, we construct a new cross-domain ADD dataset comprising over 300 hours of speech data that is generated by five advanced zero-shot TTS models. To simulate real-world scenarios, we employ diverse attack methods and audio prompts from different datasets. Experiments show that, through novel attack-augmented training, the Wav2Vec2-large and Whisper-medium models achieve equal error rates of 4.1\% and 6.5\% respectively. Additionally, we demonstrate our models' outstanding few-shot ADD ability by fine-tuning with just one minute of target-domain data. Nonetheless, neural codec compressors greatly affect the detection accuracy, necessitating further research.

【5】 Mathematics of the MML functional quantizer modules for VCV Rack  software synthesizer
标题:VCV Rack软件合成器MML函数量化器模块的数学
链接:https://arxiv.org/abs/2404.04739
作者:Michael G. Maxwell,C. McCarthy,Joshua Pfeffer,Maxwell Schneider,Robert Schneider,Andrew V. Sills
备注:4 pages, submitted for publication
摘要:我们详细介绍了数学和音乐实验室(MML)在密歇根理工大学开发的“功能量化器”模块的数学公式,为VCV机架软件模块化合成器平台,它允许合成器播放器调谐振荡器到新的音乐音阶的基础上的数学函数。例如,我们描述了最近发布的MML对数量化器(MML Logarithmic Quantizer,简称QNT)模块,该模块将合成器振荡器调谐到流行乐队The Apples in Stereo推出的非毕达哥拉斯音阶。
摘要:We detail the mathematical formulation of the line of "functional quantizer" modules developed by the Mathematics and Music Lab (MML) at Michigan Technological University, for the VCV Rack software modular synthesizer platform, which allow synthesizer players to tune oscillators to new musical scales based on mathematical functions. For example, we describe the recently-released MML Logarithmic Quantizer (LOG QNT) module that tunes synthesizer oscillators to the non-Pythagorean musical scale introduced by pop band The Apples in Stereo.


【6】 HyperTTS: Parameter Efficient Adaptation in Text to Speech using  Hypernetworks
标题:HyperTTS:使用超网络实现文本到语音的参数高效自适应
链接:https://arxiv.org/abs/2404.04645
作者:Yingting Li,Rishabh Bhardwaj,Ambuj Mehrish,Bo Cheng,Soujanya Poria
摘要:神经语音合成,或文本到语音(TTS),旨在将信号从文本域转换到语音域。虽然开发在同一组扬声器上进行训练和测试的TTS架构已经有了显著的改进,但域外扬声器性能仍然面临着巨大的限制。新的一组扬声器上的域自适应可以通过针对每个新域微调整个模型来实现,从而使其参数效率低下。这个问题可以通过适配器来解决,适配器为域适配提供了一个参数高效的替代方案。虽然在NLP中很有名,但语音合成并没有从Adapters中看到太大的改进。在这项工作中,我们提出了HyperTTS,它包括一个小的可学习的网络,“超网络”,生成适配器块的参数,使我们能够对扬声器表示适配器的条件,并使它们动态。两个域的适应设置的广泛评估表明其有效性,在实现国家的最先进的性能参数有效的制度。我们还比较了HyperTTS的不同变体,将它们与不同研究中的基线进行比较。使用超网络动态适配适配器参数的有希望的结果开辟了新的途径域通用的多扬声器TTS系统。音频示例和代码可在https://github.com/declare-lab/HyperTTS上获得。
摘要:Neural speech synthesis, or text-to-speech (TTS), aims to transform a signal from the text domain to the speech domain. While developing TTS architectures that train and test on the same set of speakers has seen significant improvements, out-of-domain speaker performance still faces enormous limitations. Domain adaptation on a new set of speakers can be achieved by fine-tuning the whole model for each new domain, thus making it parameter-inefficient. This problem can be solved by Adapters that provide a parameter-efficient alternative to domain adaptation. Although famous in NLP, speech synthesis has not seen much improvement from Adapters. In this work, we present HyperTTS, which comprises a small learnable network, "hypernetwork", that generates parameters of the Adapter blocks, allowing us to condition Adapters on speaker representations and making them dynamic. Extensive evaluations of two domain adaptation settings demonstrate its effectiveness in achieving state-of-the-art performance in the parameter-efficient regime. We also compare different variants of HyperTTS, comparing them with baselines in different studies. Promising results on the dynamic adaptation of adapter parameters using hypernetworks open up new avenues for domain-generic multi-speaker TTS systems. The audio samples and code are available at https://github.com/declare-lab/HyperTTS.


【7】 The NES Video-Music Database: A Dataset of Symbolic Video Game Music  Paired with Gameplay Videos
标题:NES视频音乐数据库:符号游戏音乐与游戏视频配对的数据集
链接:https://arxiv.org/abs/2404.04420
作者:Igor Cardoso,Rubens O. Moraes,Lucas N. Ferreira
备注:Accepted for publication at the 19th International Conference on the Foundations of Digital Games
摘要:神经模型是最流行的音乐生成方法之一,但没有标准的大型数据集可以直接从游戏数据中学习音乐。为了解决这一研究空白,我们引入了一个名为NES-VMDB的新数据集,其中包含来自389个NES游戏的98,940个游戏视频,每个视频都与其符号格式(Symbolic Format,缩写为NES)的原声音乐配对。NES-VMDB建立在任天堂娱乐系统音乐数据库(NES-MDB)的基础上,包含来自397款NES游戏的5,278首音乐。我们的方法包括收集原始数据集中389场比赛的长时间播放视频,将它们切成15秒长的片段,并从每个片段中提取音频。随后,我们应用音频指纹算法(类似于Shazam)来自动识别NES-MDB数据集中的相应片段。此外,我们引入了一个基线方法的基础上可控音乐Transformer生成NES音乐的游戏剪辑的条件。我们评估了这种方法与客观的指标,结果表明,有条件的CMT相比,其无条件的对应时,提高了音乐的结构质量。此外,我们使用神经分类器来预测生成作品的游戏类型。结果表明,CMT生成器可以学习游戏视频和游戏类型之间的相关性,但还需要进一步的研究才能实现人类水平的性能。
摘要:Neural models are one of the most popular approaches for music generation, yet there aren't standard large datasets tailored for learning music directly from game data. To address this research gap, we introduce a novel dataset named NES-VMDB, containing 98,940 gameplay videos from 389 NES games, each paired with its original soundtrack in symbolic format (MIDI). NES-VMDB is built upon the Nintendo Entertainment System Music Database (NES-MDB), encompassing 5,278 music pieces from 397 NES games. Our approach involves collecting long-play videos for 389 games of the original dataset, slicing them into 15-second-long clips, and extracting the audio from each clip. Subsequently, we apply an audio fingerprinting algorithm (similar to Shazam) to automatically identify the corresponding piece in the NES-MDB dataset. Additionally, we introduce a baseline method based on the Controllable Music Transformer to generate NES music conditioned on gameplay clips. We evaluated this approach with objective metrics, and the results showed that the conditional CMT improves musical structural quality when compared to its unconditional counterpart. Moreover, we used a neural classifier to predict the game genre of the generated pieces. Results showed that the CMT generator can learn correlations between gameplay videos and game genres, but further research has to be conducted to achieve human-level performance.


【8】 "It is okay to be uncommon": Quantizing Sound Event Detection Networks  on Hardware Accelerators with Uncommon Sub-Byte Support
链接:https://arxiv.org/abs/2404.04386
作者:Yushu Wu,Xiao Quan,Mohammad Rasool Izadi,Chuan-Che Huang
备注:5 pages, 2 figures, Accepted to ICASSP 2024
摘要:如果我们的降噪耳机能够理解我们的音频环境,它们就可以通知我们重要的声音事件,根据我们收听的内容类型调整均衡,并根据音频场景动态调整降噪参数,以进一步减少分心。然而,在具有有限能量预算和片上存储器的耳机上运行多个音频理解模型仍然是一项具有挑战性的任务。在这项工作中,我们确定了一类新的神经网络加速器(例如,GAP 9上的NE 16),其允许网络权重被量化为不同的公共(例如,8位)和不常见的位宽(例如,3比特)。然后,我们应用可微分神经架构搜索来在两个不同的声音事件检测任务上搜索网络的最佳位宽,这两个不同的声音事件检测任务对量化和预测粒度具有潜在的不同要求(即,分类与用于Few-Shot学习的嵌入)。我们进一步在实际硬件上评估了我们的量化模型,结果表明,与8位模型相比,我们的内存使用、推理延迟和能耗分别平均降低了62%、46%和61%,同时保持了浮点性能。我们的工作揭示了这种加速器的声音事件检测任务的好处时,结合适当的搜索方法。
摘要:If our noise-canceling headphones can understand our audio environments, they can then inform us of important sound events, tune equalization based on the types of content we listen to, and dynamically adjust noise cancellation parameters based on audio scenes to further reduce distraction. However, running multiple audio understanding models on headphones with a limited energy budget and on-chip memory remains a challenging task. In this work, we identify a new class of neural network accelerators (e.g., NE16 on GAP9) that allows network weights to be quantized to different common (e.g., 8 bits) and uncommon bit-widths (e.g., 3 bits). We then applied a differentiable neural architecture search to search over the optimal bit-widths of a network on two different sound event detection tasks with potentially different requirements on quantization and prediction granularity (i.e., classification vs. embeddings for few-shot learning). We further evaluated our quantized models on actual hardware, showing that we reduce memory usage, inference latency, and energy consumption by an average of 62%, 46%, and 61% respectively compared to 8-bit models while maintaining floating point performance. Our work sheds light on the benefits of such accelerators on sound event detection tasks when combined with an appropriate search method.

【9】 Transducers with Pronunciation-aware Embeddings for Automatic Speech  Recognition
标题:语音识别中的语音识别嵌入传感器
链接:https://arxiv.org/abs/2404.04295
作者:Hainan Xu,Zhehuai Chen,Fei Jia,Boris Ginsburg
备注:accepted at the ICASSP 2024 conference
摘要:本文提出了具有语音感知嵌入(PET)的换能器。与传统的传感器不同,不同标记的解码器嵌入是独立训练的,PET模型的解码器嵌入包含具有相同或相似发音的文本标记的共享组件。通过在汉语普通话和韩语的多个数据集上进行的实验,我们表明,与传统的换能器相比,PET模型始终提高了语音识别的准确性。我们的调查还揭示了一种现象,我们称之为错误连锁反应。识别错误并不是均匀地分布在整个话语中,而是倾向于聚集在一起,随后的错误通常紧随着先前的错误。我们的分析表明,PET模型通过大幅降低模型在先前错误之后产生额外错误的可能性,有效地缓解了这一问题。我们的实现将通过NeMo工具包开源。
摘要:This paper proposes Transducers with Pronunciation-aware Embeddings (PET). Unlike conventional Transducers where the decoder embeddings for different tokens are trained independently, the PET model's decoder embedding incorporates shared components for text tokens with the same or similar pronunciations. With experiments conducted in multiple datasets in Mandarin Chinese and Korean, we show that PET models consistently improve speech recognition accuracy compared to conventional Transducers. Our investigation also uncovers a phenomenon that we call error chain reactions. Instead of recognition errors being evenly spread throughout an utterance, they tend to group together, with subsequent errors often following earlier ones. Our analysis shows that PET models effectively mitigate this issue by substantially reducing the likelihood of the model generating additional errors following a prior one. Our implementation will be open-sourced with the NeMo toolkit.


【10】 Gull: A Generative Multifunctional Audio Codec
标题:Gull:一种生成式多功能音频编解码器
链接:https://arxiv.org/abs/2404.04947
作者:Yi Luo,Jianwei Yu,Hangting Chen,Rongzhi Gu,Chao Weng
备注:Demo page: this https URL
摘要:我们介绍Gull,一个生成式多功能音频编解码器。Gull是一个通用的神经音频压缩和解压缩模型,可应用于各种任务和应用,如实时通信,音频超分辨率和编解码器语言模型。Gull的关键组件包括(1)由音频源分离的最新进展激发的通过子带建模方案的通用采样率建模,(2)由传统音频编解码器激发的增益形状表示,(3)用于更简单训练的改进的残差矢量量化模块,(4)在推理时间期间实现用户定义模型大小和复杂性的弹性解码器网络,(5)在不增加比特率的情况下,内置音频超分辨率能力。我们将Gull与现有的传统和神经音频编解码器进行了比较,并表明Gull能够在各种采样率,比特率和模型复杂性方面在主观和客观评估指标上实现同等或更好的性能。
摘要:We introduce Gull, a generative multifunctional audio codec. Gull is a general purpose neural audio compression and decompression model which can be applied to a wide range of tasks and applications such as real-time communication, audio super-resolution, and codec language models. The key components of Gull include (1) universal-sample-rate modeling via subband modeling schemes motivated by recent progress in audio source separation, (2) gain-shape representations motivated by traditional audio codecs, (3) improved residual vector quantization modules for simpler training, (4) elastic decoder network that enables user-defined model size and complexity during inference time, (5) built-in ability for audio super-resolution without the increase of bitrate. We compare Gull with existing traditional and neural audio codecs and show that Gull is able to achieve on par or better performance across various sample rates, bitrates and model complexities in both subjective and objective evaluation metrics.


【11】 Rethinking Non-Negative Matrix Factorization with Implicit Neural  Representations
标题:用隐式神经表示重新思考非负矩阵分解
链接:https://arxiv.org/abs/2404.04439
作者:Krishna Subramani,Paris Smaragdis,Takuya Higuchi,Mehrez Souden
备注:Submitted to IEEE SPL, Code: this https URL
摘要:非负矩阵分解(NMF)是一种用于分析常规采样数据的强大技术,即,可以存储在矩阵中的数据。对于音频,这导致了许多使用时频(TF)表示的应用,如短时傅立叶变换。然而,将这些应用扩展到不规则间隔的TF表示,如恒定Q变换,小波或正弦分析模型,一直是不可能的,因为这些表示不能直接以矩阵形式存储。在本文中,我们制定NMF的连续函数(而不是固定向量),并表明NMF可以扩展到更广泛的各种信号类,不需要定期采样。
摘要:Non-negative Matrix Factorization (NMF) is a powerful technique for analyzing regularly-sampled data, i.e., data that can be stored in a matrix. For audio, this has led to numerous applications using time-frequency (TF) representations like the Short-Time Fourier Transform. However extending these applications to irregularly-spaced TF representations, like the Constant-Q transform, wavelets, or sinusoidal analysis models, has not been possible since these representations cannot be directly stored in matrix form. In this paper, we formulate NMF in terms of continuous functions (instead of fixed vectors) and show that NMF can be extended to a wider variety of signal classes that need not be regularly sampled.

eess.AS音频处理
【1】 Gull: A Generative Multifunctional Audio Codec
标题:Gull:一种生成式多功能音频编解码器
链接:https://arxiv.org/abs/2404.04947
作者:Yi Luo,Jianwei Yu,Hangting Chen,Rongzhi Gu,Chao Weng
备注:Demo page: this https URL
摘要:我们介绍Gull,一个生成式多功能音频编解码器。Gull是一个通用的神经音频压缩和解压缩模型,可应用于各种任务和应用,如实时通信,音频超分辨率和编解码器语言模型。Gull的关键组件包括(1)由音频源分离的最新进展激发的通过子带建模方案的通用采样率建模,(2)由传统音频编解码器激发的增益形状表示,(3)用于更简单训练的改进的残差矢量量化模块,(4)在推理时间期间实现用户定义模型大小和复杂性的弹性解码器网络,(5)在不增加比特率的情况下,内置音频超分辨率能力。我们将Gull与现有的传统和神经音频编解码器进行了比较,并表明Gull能够在各种采样率,比特率和模型复杂性方面在主观和客观评估指标上实现同等或更好的性能。
摘要:We introduce Gull, a generative multifunctional audio codec. Gull is a general purpose neural audio compression and decompression model which can be applied to a wide range of tasks and applications such as real-time communication, audio super-resolution, and codec language models. The key components of Gull include (1) universal-sample-rate modeling via subband modeling schemes motivated by recent progress in audio source separation, (2) gain-shape representations motivated by traditional audio codecs, (3) improved residual vector quantization modules for simpler training, (4) elastic decoder network that enables user-defined model size and complexity during inference time, (5) built-in ability for audio super-resolution without the increase of bitrate. We compare Gull with existing traditional and neural audio codecs and show that Gull is able to achieve on par or better performance across various sample rates, bitrates and model complexities in both subjective and objective evaluation metrics.


【2】 Rethinking Non-Negative Matrix Factorization with Implicit Neural  Representations
标题:用隐式神经表示重新思考非负矩阵分解
链接:https://arxiv.org/abs/2404.04439
作者:Krishna Subramani,Paris Smaragdis,Takuya Higuchi,Mehrez Souden
备注:Submitted to IEEE SPL, Code: this https URL
摘要:非负矩阵分解(NMF)是一种用于分析常规采样数据的强大技术,即,可以存储在矩阵中的数据。对于音频,这导致了许多使用时频(TF)表示的应用,如短时傅立叶变换。然而,将这些应用扩展到不规则间隔的TF表示,如恒定Q变换,小波或正弦分析模型,一直是不可能的,因为这些表示不能直接以矩阵形式存储。在本文中,我们制定NMF的连续函数(而不是固定向量),并表明NMF可以扩展到更广泛的各种信号类,不需要定期采样。
摘要:Non-negative Matrix Factorization (NMF) is a powerful technique for analyzing regularly-sampled data, i.e., data that can be stored in a matrix. For audio, this has led to numerous applications using time-frequency (TF) representations like the Short-Time Fourier Transform. However extending these applications to irregularly-spaced TF representations, like the Constant-Q transform, wavelets, or sinusoidal analysis models, has not been possible since these representations cannot be directly stored in matrix form. In this paper, we formulate NMF in terms of continuous functions (instead of fixed vectors) and show that NMF can be extended to a wider variety of signal classes that need not be regularly sampled.


【3】 Uformer: A UNet-Transformer fused robust end-to-end deep learning  framework for real-time denoising of lung sounds
标题:Uformer:一个融合了UNet—Transformer的强大端到端深度学习框架,用于实时对肺音进行降噪
链接:https://arxiv.org/abs/2404.04365
作者:Samiul Based Shuvo,Syed Samiul Alam,Taufiq Hasan
摘要:目的:肺部听诊是诊断和监测各种呼吸系统疾病的重要手段。然而,肺音(LS)受到许多污染源的显著影响,特别是在真实世界的临床环境中记录时。传统的去噪模型被证明对于LS去噪是不切实际的,主要是由于不同噪声源引起的频谱重叠复杂性。为了解决这个问题,我们提出了一个专门的深度学习模型(Uformer)用于肺音去噪。研究方法:所提出的Uformer模型由三个模块组成:卷积神经网络(CNN)编码器模块,专用于提取潜在特征; Transformer编码器模块,用于进一步增强独特LS特征的编码,并有效地捕获复杂的长程依赖关系;以及CNN解码器模块,用于生成去噪信号。进行了消融研究,以找到最佳结构。结果:所提出的Uformer模型的性能进行了评估肺音与不同类型的合成和真实世界的噪音。在测试实验中考虑了-12dB至15 dB信噪比(SNR)的肺音信号。当使用-12 dB LS信号进行评估时,所提出的模型显示出16.51 dB的平均SNR改善。我们的端到端模型,平均信噪比提高了19.31 dB,优于现有的模型时,环境噪声和更少的参数进行评估。结论:基于本研究中的定性和定量结果,可以说Uformer具有稳健性,可广泛用于辅助监测呼吸状况。
摘要:Objective: Lung auscultation is a valuable tool in diagnosing and monitoring various respiratory diseases. However, lung sounds (LS) are significantly affected by numerous sources of contamination, especially when recorded in real-world clinical settings. Conventional denoising models prove impractical for LS denoising, primarily owing to spectral overlap complexities arising from diverse noise sources. To address this issue, we propose a specialized deep-learning model (Uformer) for lung sound denoising. Methods: The proposed Uformer model is constituted of three modules: a Convolutional Neural Network (CNN) encoder module, dedicated to extracting latent features; a Transformer encoder module, employed to further enhance the encoding of unique LS features and effectively capture intricate long-range dependencies; and a CNN decoder module, employed to generate the denoised signals. An ablation study was performed in order to find the most optimal architecture. Results: The performance of the proposed Uformer model was evaluated on lung sounds induced with different types of synthetic and real-world noises. Lung sound signals of -12 dB to 15 dB signal-to-noise ratio (SNR) were considered in testing experiments. The proposed model showed an average SNR improvement of 16.51 dB when evaluated with -12 dB LS signals. Our end-to-end model, with an average SNR improvement of 19.31 dB, outperforms the existing model when evaluated with ambient noise and fewer parameters. Conclusion: Based on the qualitative and quantitative findings in this study, it can be stated that Uformer is robust and generalized to be used in assisting the monitoring of respiratory conditions.


【4】 VietMed: A Dataset and Benchmark for Automatic Speech Recognition of  Vietnamese in the Medical Domain
标题:VietMed:医学领域越南语自动语音识别的数据集和基准
链接:https://arxiv.org/abs/2404.05659
作者:Khai Le-Duc
备注:LREC-COLING 2024
摘要:由于隐私限制,医疗领域缺乏公开的语音识别数据集。在这项工作中,我们提出了VietMed -一个越南语语音识别数据集在医疗领域包括16小时的标记医疗语音,1000小时的未标记医疗语音和1200小时的未标记一般域语音。据我们所知,VietMed是迄今为止世界上最大的公共医疗语音识别数据集,包括7个方面:总时长,说话者数量,疾病,记录条件,说话者角色,独特的医疗术语和口音。VietMed也是迄今为止最大的公共越南语语音数据集。此外,我们是第一个提供涵盖所有ICD-10疾病组和一个国家内所有口音的医学ASR数据集的公司。此外,我们还发布了第一个用于越南语ASR的公共大规模预训练模型w2 v2-Viet和XLSR-53-Viet,以及第一个用于医疗ASR的公共大规模微调模型。即使在无监督预训练中没有任何医疗数据,我们最好的预训练模型XLSR-53-Viet也能很好地推广到医疗领域,表现优于最先进的XLSR-53,在测试集上的WER从51.8%降低到29.6%(相对降低超过40%)。所有代码、数据和模型都在这里公开:https://github.com/leduckhai/MultiMed。
摘要:Due to privacy restrictions, there's a shortage of publicly available speech recognition datasets in the medical domain. In this work, we present VietMed - a Vietnamese speech recognition dataset in the medical domain comprising 16h of labeled medical speech, 1000h of unlabeled medical speech and 1200h of unlabeled general-domain speech. To our best knowledge, VietMed is by far the world's largest public medical speech recognition dataset in 7 aspects: total duration, number of speakers, diseases, recording conditions, speaker roles, unique medical terms and accents. VietMed is also by far the largest public Vietnamese speech dataset in terms of total duration. Additionally, we are the first to present a medical ASR dataset covering all ICD-10 disease groups and all accents within a country. Moreover, we release the first public large-scale pre-trained models for Vietnamese ASR, w2v2-Viet and XLSR-53-Viet, along with the first public large-scale fine-tuned models for medical ASR. Even without any medical data in unsupervised pre-training, our best pre-trained model XLSR-53-Viet generalizes very well to the medical domain by outperforming state-of-the-art XLSR-53, from 51.8% to 29.6% WER on test set (a relative reduction of more than 40%). All code, data and models are made publicly available here: https://github.com/leduckhai/MultiMed.


【5】 Linguistic Changes in Spontaneous Speech for Detecting Parkinsons  Disease Using Large Language Models
标题:使用大语言模型检测帕金森病自发语音的语言变化
链接:https://arxiv.org/abs/2404.05160
作者:Jonathan Crawford
备注:12 pages, 3 figures
摘要:帕金森病是第二大流行的神经退行性疾病,全球有超过一千万活跃病例,每年有一百万新诊断。检测和随后诊断疾病是具有挑战性的,因为症状的异质性方面的复杂性,以及类型和时间表型表现。通常,语言障碍可以出现在前驱期和运动症状之前,这表明基于语言的方法可以作为早期帕金森病的诊断方法。此外,改进的语言模型可以通过集成技术增强其他方法。大型语言模型领域正在迅速发展,为探索使用这些新模型检测帕金森病提供了机会,并通过语言学的高维表示来改进当前的语言学方法。我们评估了应用最先进的大型语言模型从自发语音中自动检测帕金森病,准确率高达73%。
摘要:Parkinsons disease is the second most prevalent neurodegenerative disorder with over ten million active cases worldwide and one million new diagnoses per year. Detecting and subsequently diagnosing the disease is challenging because of symptom heterogeneity with respect to complexity, as well as the type and timing of phenotypic manifestations. Typically, language impairment can present in the prodromal phase and precede motor symptoms suggesting that a linguistic-based approach could serve as a diagnostic method for incipient Parkinsons disease. Additionally, improved linguistic models may enhance other approaches through ensemble techniques. The field of large language models is advancing rapidly, presenting the opportunity to explore the use of these new models for detecting Parkinsons disease and to improve on current linguistic approaches with high-dimensional representations of linguistics. We evaluate the application of state-of-the-art large language models to detect Parkinsons disease automatically from spontaneous speech with up to 73% accuracy.


【6】 Mathematics of the MML functional quantizer modules for VCV Rack  software synthesizer
标题:VCV Rack软件合成器MML函数量化器模块的数学
链接:https://arxiv.org/abs/2404.04739
作者:Michael G. Maxwell,C. McCarthy,Joshua Pfeffer,Maxwell Schneider,Robert Schneider,Andrew V. Sills
备注:4 pages, submitted for publication
摘要:我们详细介绍了数学和音乐实验室(MML)在密歇根理工大学开发的“功能量化器”模块的数学公式,为VCV机架软件模块化合成器平台,它允许合成器播放器调谐振荡器到新的音乐音阶的基础上的数学函数。例如,我们描述了最近发布的MML对数量化器(MML Logarithmic Quantizer,简称QNT)模块,该模块将合成器振荡器调谐到流行乐队The Apples in Stereo推出的非毕达哥拉斯音阶。
摘要:We detail the mathematical formulation of the line of "functional quantizer" modules developed by the Mathematics and Music Lab (MML) at Michigan Technological University, for the VCV Rack software modular synthesizer platform, which allow synthesizer players to tune oscillators to new musical scales based on mathematical functions. For example, we describe the recently-released MML Logarithmic Quantizer (LOG QNT) module that tunes synthesizer oscillators to the non-Pythagorean musical scale introduced by pop band The Apples in Stereo.

【7】 "It is okay to be uncommon": Quantizing Sound Event Detection Networks  on Hardware Accelerators with Uncommon Sub-Byte Support
链接:https://arxiv.org/abs/2404.04386
作者:Yushu Wu,Xiao Quan,Mohammad Rasool Izadi,Chuan-Che Huang
备注:5 pages, 2 figures, Accepted to ICASSP 2024
摘要:如果我们的降噪耳机能够理解我们的音频环境,它们就可以通知我们重要的声音事件,根据我们收听的内容类型调整均衡,并根据音频场景动态调整降噪参数,以进一步减少分心。然而,在具有有限能量预算和片上存储器的耳机上运行多个音频理解模型仍然是一项具有挑战性的任务。在这项工作中,我们确定了一类新的神经网络加速器(例如,GAP 9上的NE 16),其允许网络权重被量化为不同的公共(例如,8位)和不常见的位宽(例如,3比特)。然后,我们应用可微分神经架构搜索来在两个不同的声音事件检测任务上搜索网络的最佳位宽,这两个不同的声音事件检测任务对量化和预测粒度具有潜在的不同要求(即,分类与用于Few-Shot学习的嵌入)。我们进一步在实际硬件上评估了我们的量化模型,结果表明,与8位模型相比,我们的内存使用、推理延迟和能耗分别平均降低了62%、46%和61%,同时保持了浮点性能。我们的工作揭示了这种加速器的声音事件检测任务的好处时,结合适当的搜索方法。
摘要:If our noise-canceling headphones can understand our audio environments, they can then inform us of important sound events, tune equalization based on the types of content we listen to, and dynamically adjust noise cancellation parameters based on audio scenes to further reduce distraction. However, running multiple audio understanding models on headphones with a limited energy budget and on-chip memory remains a challenging task. In this work, we identify a new class of neural network accelerators (e.g., NE16 on GAP9) that allows network weights to be quantized to different common (e.g., 8 bits) and uncommon bit-widths (e.g., 3 bits). We then applied a differentiable neural architecture search to search over the optimal bit-widths of a network on two different sound event detection tasks with potentially different requirements on quantization and prediction granularity (i.e., classification vs. embeddings for few-shot learning). We further evaluated our quantized models on actual hardware, showing that we reduce memory usage, inference latency, and energy consumption by an average of 62%, 46%, and 61% respectively compared to 8-bit models while maintaining floating point performance. Our work sheds light on the benefits of such accelerators on sound event detection tasks when combined with an appropriate search method.


机器翻译由腾讯交互翻译提供,仅供参考