微信公众号:arXiv_Daily
cs.SD语音
标题: 通过局部李群变换实现稳健ASB的构音障碍规范化
链接:https://arxiv.org/abs/2504.12279
备注:Preprint. 11 pages, 3 figures, 2 tables, 8 appendices. Code and data available upon request
摘要:我们提出了一个几何驱动的方法正常化构音障碍的语音使用局部李群变换的频谱图。时间,频率和幅度失真被建模为平滑,可逆的变形,参数化的标量场和应用通过指数映射。一个神经网络经过训练,可以从典型语音的合成失真中推断出这些场,而不需要使用任何病理数据。在测试时,该模型将近似逆应用于真实构音障碍输入。尽管zero-shot泛化,我们观察到大量的ASR增益,包括高达16个百分点的WER减少具有挑战性的TORGO样本,没有退化干净的语音。这项工作介绍了一个原则性的,可解释的方法,强大的语音识别下运动言语障碍
摘要:We present a geometry-driven method for normalizing dysarthric speech using local Lie group transformations of spectrograms. Time, frequency, and amplitude distortions are modeled as smooth, invertible deformations, parameterized by scalar fields and applied via exponential maps. A neural network is trained to infer these fields from synthetic distortions of typical speech-without using any pathological data. At test time, the model applies an approximate inverse to real dysarthric inputs. Despite zero-shot generalization, we observe substantial ASR gains, including up to 16 percentage points WER reduction on challenging TORGO samples, with no degradation on clean speech. This work introduces a principled, interpretable approach for robust speech recognition under motor speech disorders
【2】 Edge Intelligence for Wildlife Conservation: Real-Time Hornbill Call Classification Using TinyML
标题: 野生动物保护的边缘智能:使用TinyML进行实时犀鸟叫声分类链接:https://arxiv.org/abs/2504.12272
备注:This is a preprint version of a paper accepted and published in Springer Lecture Notes in Networks and Systems. The final version is available at https://doi.org/10.1007/978-981-96-3949-6_40
摘要:犀鸟是马来西亚生物多样性的标志性物种,面临着栖息地丧失、偷猎和环境变化的威胁,需要准确和实时的种群监测,这在传统上是具有挑战性的,而且是资源密集型的。微型机器学习(TinyML)的出现提供了一个机会,通过直接在边缘设备上实现高效,实时的数据分析来改变野生动物监测。为了应对野生动物保护的挑战,这篇研究论文探讨了机器学习,特别是TinyML在马来西亚犀鸟叫声分类和监测中的关键作用。利用Xeno-Canto数据库中的音频数据,该研究旨在开发一个能够识别和分类犀鸟发声的语音识别系统。所提出的方法包括预处理音频数据,使用Mel频率能量(MFE)提取特征,并将模型部署在擅长边缘计算的Arduino Nano 33 BLE上。本研究包含基础性工作,包括一个全面的介绍,文献综述和方法。该模型使用边缘脉冲进行训练,并通过真实世界的测试进行验证,在犀鸟物种识别中实现了较高的准确性。该项目强调了TinyML在环境监测方面的潜力及其在生态保护工作中的更广泛应用,为TinyML和野生动物保护领域做出了贡献。
摘要:Hornbills, an iconic species of Malaysia's biodiversity, face threats from habi-tat loss, poaching, and environmental changes, necessitating accurate and real-time population monitoring that is traditionally challenging and re-source intensive. The emergence of Tiny Machine Learning (TinyML) offers a chance to transform wildlife monitoring by enabling efficient, real-time da-ta analysis directly on edge devices. Addressing the challenge of wildlife conservation, this research paper explores the pivotal role of machine learn-ing, specifically TinyML, in the classification and monitoring of hornbill calls in Malaysia. Leveraging audio data from the Xeno-canto database, the study aims to develop a speech recognition system capable of identifying and classifying hornbill vocalizations. The proposed methodology involves pre-processing the audio data, extracting features using Mel-Frequency Energy (MFE), and deploying the model on an Arduino Nano 33 BLE, which is adept at edge computing. The research encompasses foundational work, in-cluding a comprehensive introduction, literature review, and methodology. The model is trained using Edge Impulse and validated through real-world tests, achieving high accuracy in hornbill species identification. The project underscores the potential of TinyML for environmental monitoring and its broader application in ecological conservation efforts, contributing to both the field of TinyML and wildlife conservation.
【3】 Voice Conversion with Diverse Intonation using Conditional Variational Auto-Encoder
标题: 使用条件变分自动编码器实现多样化语调的语音转换链接:https://arxiv.org/abs/2504.12005
备注:2 pages, Machine Learning in Speech and Language Processing Workshop (MLSLP) 2018
摘要:语音转换是指在保持源语音语言信息的前提下,将源语音合成为目标语音。虽然说话者可以从具有不同语调的单个脚本产生不同的话语,但是传统的语音转换模型限于每个源输入仅产生一个结果。为了克服这一限制,我们提出了一种新的方法,语音转换与不同的语调使用条件变分自动编码器(CVAE)。实验表明,说话人的风格特征可以映射到一个高斯分布的隐空间中。我们也已经能够通过使用逆自回归流(IAF)使潜在空间的后部更加复杂来转换具有更多样化语调的声音。结果,转换后的语音不仅具有语调的多样性,而且具有比没有CVAE的模型更好的音质。
摘要:Voice conversion is a task of synthesizing an utterance with target speaker's voice while maintaining linguistic information of the source utterance. While a speaker can produce varying utterances from a single script with different intonations, conventional voice conversion models were limited to producing only one result per source input. To overcome this limitation, we propose a novel approach for voice conversion with diverse intonations using conditional variational autoencoder (CVAE). Experiments have shown that the speaker's style feature can be mapped into a latent space with Gaussian distribution. We have also been able to convert voices with more diverse intonation by making the posterior of the latent space more complex with inverse autoregressive flow (IAF). As a result, the converted voice not only has a diversity of intonations, but also has better sound quality than the model without CVAE.
【4】 Making Acoustic Side-Channel Attacks on Noisy Keyboards Viable with LLM-Assisted Spectrograms' "Typo" Correction
链接:https://arxiv.org/abs/2504.11622备注:Length: 13 pages Figures: 5 figures Tables: 7 tables Keywords: Acoustic side-channel attacks, machine learning, Visual Transformers, Large Language Models (LLMs), security Conference: Accepted at the 19th USENIX WOOT Conference on Offensive Technologies (WOOT '25). Licensing: This paper is submitted under the CC BY Creative Commons Attribution license. arXiv admin note: text overlap with arXiv:2502.09782
摘要:将麦克风大量集成到设备中增加了声学侧信道攻击(ASCA)的机会,因为这些可以用于捕获可能泄露敏感信息的麦克风音频信号。然而,ASCA的当前最先进(SOTA)模型,包括卷积神经网络(CNN)和混合模型,如CoAtNet,在现实的噪声条件下仍然表现出有限的鲁棒性。解决这个问题需要:(i)增加模型从较长序列中推断上下文信息的能力,允许模型学习最初有噪声输入的单词与将来收集的无噪声单词相同,或者(ii)修复上下文中错误识别信息的方法,因为人们不会输入随机单词,而是最适合对话上下文的单词。在本文中,我们证明了这两种策略是可行的和互补的解决方案,使ASCAs实用。我们观察到,没有现有的解决方案利用先进的Transformer架构的权力,这些任务,并提出:(i)视觉Transformers(VT)是捕获长期上下文信息的候选解决方案和(ii)transformer-powered大语言模型(LLM)的候选解决方案,以修复“typos”(误预测)的模型可能会。因此,我们在这里提出了第一种方法,将VT和LLM集成为ASCA。 我们首先表明,与以前的CNN基准相比,VT实现SOTA的性能分类的神经网络。其次,我们证明了LLM可以减轻现实世界噪声的影响。对自然句子的评估显示:(i)将LLM(例如,GPT-4 o)在我们的ASCA管道中提高了纠错任务的性能;(ii)可以通过轻量级,微调的较小LLM(比GPT-4 o小67倍)实现类似的性能,使用...
摘要:The large integration of microphones into devices increases the opportunities for Acoustic Side-Channel Attacks (ASCAs), as these can be used to capture keystrokes' audio signals that might reveal sensitive information. However, the current State-Of-The-Art (SOTA) models for ASCAs, including Convolutional Neural Networks (CNNs) and hybrid models, such as CoAtNet, still exhibit limited robustness under realistic noisy conditions. Solving this problem requires either: (i) an increased model's capacity to infer contextual information from longer sequences, allowing the model to learn that an initially noisily typed word is the same as a futurely collected non-noisy word, or (ii) an approach to fix misidentified information from the contexts, as one does not type random words, but the ones that best fit the conversation context. In this paper, we demonstrate that both strategies are viable and complementary solutions for making ASCAs practical. We observed that no existing solution leverages advanced transformer architectures' power for these tasks and propose that: (i) Visual Transformers (VTs) are the candidate solutions for capturing long-term contextual information and (ii) transformer-powered Large Language Models (LLMs) are the candidate solutions to fix the ``typos'' (mispredictions) the model might make. Thus, we here present the first-of-its-kind approach that integrates VTs and LLMs for ASCAs. We first show that VTs achieve SOTA performance in classifying keystrokes when compared to the previous CNN benchmark. Second, we demonstrate that LLMs can mitigate the impact of real-world noise. Evaluations on the natural sentences revealed that: (i) incorporating LLMs (e.g., GPT-4o) in our ASCA pipeline boosts the performance of error-correction tasks; and (ii) the comparable performance can be attained by a lightweight, fine-tuned smaller LLM (67 times smaller than GPT-4o), using...
【1】 Dysarthria Normalization via Local Lie Group Transformations for Robust ASR
标题: 通过局部李群变换实现稳健ASB的构音障碍规范化链接:https://arxiv.org/abs/2504.12279
备注:Preprint. 11 pages, 3 figures, 2 tables, 8 appendices. Code and data available upon request
摘要:我们提出了一个几何驱动的方法正常化构音障碍的语音使用局部李群变换的频谱图。时间,频率和幅度失真被建模为平滑,可逆的变形,参数化的标量场和应用通过指数映射。一个神经网络经过训练,可以从典型语音的合成失真中推断出这些场,而不需要使用任何病理数据。在测试时,该模型将近似逆应用于真实构音障碍输入。尽管zero-shot泛化,我们观察到大量的ASR增益,包括高达16个百分点的WER减少具有挑战性的TORGO样本,没有退化干净的语音。这项工作介绍了一个原则性的,可解释的方法,强大的语音识别下运动言语障碍
摘要:We present a geometry-driven method for normalizing dysarthric speech using local Lie group transformations of spectrograms. Time, frequency, and amplitude distortions are modeled as smooth, invertible deformations, parameterized by scalar fields and applied via exponential maps. A neural network is trained to infer these fields from synthetic distortions of typical speech-without using any pathological data. At test time, the model applies an approximate inverse to real dysarthric inputs. Despite zero-shot generalization, we observe substantial ASR gains, including up to 16 percentage points WER reduction on challenging TORGO samples, with no degradation on clean speech. This work introduces a principled, interpretable approach for robust speech recognition under motor speech disorders
【2】 Edge Intelligence for Wildlife Conservation: Real-Time Hornbill Call Classification Using TinyML
标题: 野生动物保护的边缘智能:使用TinyML进行实时犀鸟叫声分类链接:https://arxiv.org/abs/2504.12272
备注:This is a preprint version of a paper accepted and published in Springer Lecture Notes in Networks and Systems. The final version is available at https://doi.org/10.1007/978-981-96-3949-6_40
摘要:犀鸟是马来西亚生物多样性的标志性物种,面临着栖息地丧失、偷猎和环境变化的威胁,需要准确和实时的种群监测,这在传统上是具有挑战性的,而且是资源密集型的。微型机器学习(TinyML)的出现提供了一个机会,通过直接在边缘设备上实现高效,实时的数据分析来改变野生动物监测。为了应对野生动物保护的挑战,这篇研究论文探讨了机器学习,特别是TinyML在马来西亚犀鸟叫声分类和监测中的关键作用。利用Xeno-Canto数据库中的音频数据,该研究旨在开发一个能够识别和分类犀鸟发声的语音识别系统。所提出的方法包括预处理音频数据,使用Mel频率能量(MFE)提取特征,并将模型部署在擅长边缘计算的Arduino Nano 33 BLE上。本研究包含基础性工作,包括一个全面的介绍,文献综述和方法。该模型使用边缘脉冲进行训练,并通过真实世界的测试进行验证,在犀鸟物种识别中实现了较高的准确性。该项目强调了TinyML在环境监测方面的潜力及其在生态保护工作中的更广泛应用,为TinyML和野生动物保护领域做出了贡献。
摘要:Hornbills, an iconic species of Malaysia's biodiversity, face threats from habi-tat loss, poaching, and environmental changes, necessitating accurate and real-time population monitoring that is traditionally challenging and re-source intensive. The emergence of Tiny Machine Learning (TinyML) offers a chance to transform wildlife monitoring by enabling efficient, real-time da-ta analysis directly on edge devices. Addressing the challenge of wildlife conservation, this research paper explores the pivotal role of machine learn-ing, specifically TinyML, in the classification and monitoring of hornbill calls in Malaysia. Leveraging audio data from the Xeno-canto database, the study aims to develop a speech recognition system capable of identifying and classifying hornbill vocalizations. The proposed methodology involves pre-processing the audio data, extracting features using Mel-Frequency Energy (MFE), and deploying the model on an Arduino Nano 33 BLE, which is adept at edge computing. The research encompasses foundational work, in-cluding a comprehensive introduction, literature review, and methodology. The model is trained using Edge Impulse and validated through real-world tests, achieving high accuracy in hornbill species identification. The project underscores the potential of TinyML for environmental monitoring and its broader application in ecological conservation efforts, contributing to both the field of TinyML and wildlife conservation.
【3】 Voice Conversion with Diverse Intonation using Conditional Variational Auto-Encoder
标题: 使用条件变分自动编码器实现多样化语调的语音转换链接:https://arxiv.org/abs/2504.12005
备注:2 pages, Machine Learning in Speech and Language Processing Workshop (MLSLP) 2018
摘要:语音转换是指在保持源语音语言信息的前提下,将源语音合成为目标语音。虽然说话者可以从具有不同语调的单个脚本产生不同的话语,但是传统的语音转换模型限于每个源输入仅产生一个结果。为了克服这一限制,我们提出了一种新的方法,语音转换与不同的语调使用条件变分自动编码器(CVAE)。实验表明,说话人的风格特征可以映射到一个高斯分布的隐空间中。我们也已经能够通过使用逆自回归流(IAF)使潜在空间的后部更加复杂来转换具有更多样化语调的声音。结果,转换后的语音不仅具有语调的多样性,而且具有比没有CVAE的模型更好的音质。
摘要:Voice conversion is a task of synthesizing an utterance with target speaker's voice while maintaining linguistic information of the source utterance. While a speaker can produce varying utterances from a single script with different intonations, conventional voice conversion models were limited to producing only one result per source input. To overcome this limitation, we propose a novel approach for voice conversion with diverse intonations using conditional variational autoencoder (CVAE). Experiments have shown that the speaker's style feature can be mapped into a latent space with Gaussian distribution. We have also been able to convert voices with more diverse intonation by making the posterior of the latent space more complex with inverse autoregressive flow (IAF). As a result, the converted voice not only has a diversity of intonations, but also has better sound quality than the model without CVAE.
【4】 Making Acoustic Side-Channel Attacks on Noisy Keyboards Viable with LLM-Assisted Spectrograms' "Typo" Correction
链接:https://arxiv.org/abs/2504.11622备注:Length: 13 pages Figures: 5 figures Tables: 7 tables Keywords: Acoustic side-channel attacks, machine learning, Visual Transformers, Large Language Models (LLMs), security Conference: Accepted at the 19th USENIX WOOT Conference on Offensive Technologies (WOOT '25). Licensing: This paper is submitted under the CC BY Creative Commons Attribution license. arXiv admin note: text overlap with arXiv:2502.09782
摘要:将麦克风大量集成到设备中增加了声学侧信道攻击(ASCA)的机会,因为这些可以用于捕获可能泄露敏感信息的麦克风音频信号。然而,ASCA的当前最先进(SOTA)模型,包括卷积神经网络(CNN)和混合模型,如CoAtNet,在现实的噪声条件下仍然表现出有限的鲁棒性。解决这个问题需要:(i)增加模型从较长序列中推断上下文信息的能力,允许模型学习最初有噪声输入的单词与将来收集的无噪声单词相同,或者(ii)修复上下文中错误识别信息的方法,因为人们不会输入随机单词,而是最适合对话上下文的单词。在本文中,我们证明了这两种策略是可行的和互补的解决方案,使ASCAs实用。我们观察到,没有现有的解决方案利用先进的Transformer架构的权力,这些任务,并提出:(i)视觉Transformers(VT)是捕获长期上下文信息的候选解决方案和(ii)transformer-powered大语言模型(LLM)的候选解决方案,以修复“typos”(误预测)的模型可能会。因此,我们在这里提出了第一种方法,将VT和LLM集成为ASCA。 我们首先表明,与以前的CNN基准相比,VT实现SOTA的性能分类的神经网络。其次,我们证明了LLM可以减轻现实世界噪声的影响。对自然句子的评估显示:(i)将LLM(例如,GPT-4 o)在我们的ASCA管道中提高了纠错任务的性能;(ii)可以通过轻量级,微调的较小LLM(比GPT-4 o小67倍)实现类似的性能,使用...
摘要:The large integration of microphones into devices increases the opportunities for Acoustic Side-Channel Attacks (ASCAs), as these can be used to capture keystrokes' audio signals that might reveal sensitive information. However, the current State-Of-The-Art (SOTA) models for ASCAs, including Convolutional Neural Networks (CNNs) and hybrid models, such as CoAtNet, still exhibit limited robustness under realistic noisy conditions. Solving this problem requires either: (i) an increased model's capacity to infer contextual information from longer sequences, allowing the model to learn that an initially noisily typed word is the same as a futurely collected non-noisy word, or (ii) an approach to fix misidentified information from the contexts, as one does not type random words, but the ones that best fit the conversation context. In this paper, we demonstrate that both strategies are viable and complementary solutions for making ASCAs practical. We observed that no existing solution leverages advanced transformer architectures' power for these tasks and propose that: (i) Visual Transformers (VTs) are the candidate solutions for capturing long-term contextual information and (ii) transformer-powered Large Language Models (LLMs) are the candidate solutions to fix the ``typos'' (mispredictions) the model might make. Thus, we here present the first-of-its-kind approach that integrates VTs and LLMs for ASCAs. We first show that VTs achieve SOTA performance in classifying keystrokes when compared to the previous CNN benchmark. Second, we demonstrate that LLMs can mitigate the impact of real-world noise. Evaluations on the natural sentences revealed that: (i) incorporating LLMs (e.g., GPT-4o) in our ASCA pipeline boosts the performance of error-correction tasks; and (ii) the comparable performance can be attained by a lightweight, fine-tuned smaller LLM (67 times smaller than GPT-4o), using...
机器翻译由腾讯交互翻译提供,仅供参考
