今日论文合集:cs.SD语音13篇,eess.AS音频处理13篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models
标题: 超越单音频:推进音频大型语言模型中的多音频处理
作者:Yiming Chen,Xianghu Yue,Xiaoxue Gao,Chen Zhang,Luis Fernando D'Haro,Robby T. Tan,Haizhou Li
备注:EMNLP24 Findings
链接:点击下载PDF文件
摘要:最近已经探索了各种音频LLM(ALLM),用于使用单个统一模型同时处理不同的音频任务。虽然ALLM的现有评估主要集中在单音频任务上,但实际应用通常涉及同时处理多个音频流。为了弥合这一差距,我们提出了第一个多音频评估(MAE)基准,该基准由来自11个多音频任务的20个数据集组成,包括语音和声音场景。对MAE的综合实验表明,现有的ALLM虽然在理解单个音频输入中的主要音频元素方面很强大,但难以处理多音频场景。为此,我们提出了一种新的多音频LLM(MALLM),使用我们提出的合成数据上的判别学习来捕获多个相似音频之间的音频上下文。结果表明,所提出的MALLM优于所有基线,并实现了高数据效率,使用合成数据,而无需人工注释。提出的MALLM为ALLM打开了通往多音频处理时代的大门,并使我们更接近于在机器中复制人类听觉能力。摘要:Various audio-LLMs (ALLMs) have been explored recently for tackling different audio tasks simultaneously using a single, unified model. While existing evaluations of ALLMs primarily focus on single-audio tasks, real-world applications often involve processing multiple audio streams simultaneously. To bridge this gap, we propose the first multi-audio evaluation (MAE) benchmark that consists of 20 datasets from 11 multi-audio tasks encompassing both speech and sound scenarios. Comprehensive experiments on MAE demonstrate that the existing ALLMs, while being powerful in comprehending primary audio elements in individual audio inputs, struggling to handle multi-audio scenarios. To this end, we propose a novel multi-audio-LLM (MALLM) to capture audio context among multiple similar audios using discriminative learning on our proposed synthetic data. The results demonstrate that the proposed MALLM outperforms all baselines and achieves high data efficiency using synthetic data without requiring human annotations. The proposed MALLM opens the door for ALLMs towards multi-audio processing era and brings us closer to replicating human auditory capabilities in machines.

【2】 Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech
标题: 基于音频的语言特征提取用于增强多语言和低资源的文本到语音
作者:Youngjae Kim,Yejin Jeon,Gary Geunbae Lee
备注:EMNLP 2024 Findings
链接:点击下载PDF文件
摘要:获取丰富、高质量数据的困难,特别是在多语言环境中,引发了人们对解决低资源情景的兴趣。此外,目前的文献依赖于固定的表达式从语言ID,这导致在语言表示的学习不足,并生成语音在看不见的语言失败。为了解决这些挑战,我们提出了一种新的方法,直接从音频输入中提取语言特征,同时有效地过滤掉杂项声学信息,包括扬声器特定的属性,如音色。主观和客观的评价肯定了我们的方法对多语言文本到语音的有效性,并突出了其在以前看不见的语言的低资源迁移学习的优势。摘要:The difficulty of acquiring abundant, high-quality data, especially in multi-lingual contexts, has sparked interest in addressing low-resource scenarios. Moreover, current literature rely on fixed expressions from language IDs, which results in the inadequate learning of language representations, and the failure to generate speech in unseen languages. To address these challenges, we propose a novel method that directly extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre. Subjective and objective evaluations affirm the effectiveness of our approach for multi-lingual text-to-speech, and highlight its superiority in low-resource transfer learning for previously unseen language.

【3】 ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5
标题: ChildMandarin:针对3-5岁幼儿的全面普通话语音数据集
作者:Jiaming Zhou,Shiyao Wang,Shiwan Zhao,Jiabei He,Haoqin Sun,Hui Wang,Cheng Liu,Aobo Kong,Yujie Guo,Yong Qin
链接:点击下载PDF文件
摘要:自动语音识别(ASR)系统已经通过Whisper、Conformer等模型以及Wav 2 vec 2.0和HuBERT等自监督框架取得了显着进步。然而,开发强大的ASR模型,幼儿的语音仍然具有挑战性,由于发音,语调和速度相比,成人语音的差异。在本文中,我们介绍了一个新的普通话语音数据集,专注于3至5岁的儿童,解决了这方面的资源稀缺。该数据集包括41.25小时的演讲,经过精心制作的人工翻译,收集了来自中国各省的397名演讲者,性别代表性均衡。我们提供了一个全面的分析,扬声器人口统计,演讲持续时间分布和地理覆盖范围。此外,我们还评估了从头开始训练的模型(如Conformer)以及微调的预训练模型(如HuBERT和Whisper)的ASR性能,其中微调显示了显着的性能改进。此外,我们在我们的数据集上评估了说话人验证(SV),结果表明,尽管幼儿独特的声音特征带来了挑战,但该数据集有效地支持ASR和SV任务。该数据集对汉语儿童语音研究是一个有价值的贡献,并具有在教育技术和儿童-计算机交互中的应用潜力。它将是开源的,可免费用于所有学术目的。摘要:Automatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However, developing robust ASR models for young children's speech remains challenging due to differences in pronunciation, tone, and pace compared to adult speech. In this paper, we introduce a new Mandarin speech dataset focused on children aged 3 to 5, addressing the scarcity of resources in this area. The dataset comprises 41.25 hours of speech with carefully crafted manual transcriptions, collected from 397 speakers across various provinces in China, with balanced gender representation. We provide a comprehensive analysis of speaker demographics, speech duration distribution and geographic coverage. Additionally, we evaluate ASR performance on models trained from scratch, such as Conformer, as well as fine-tuned pre-trained models like HuBERT and Whisper, where fine-tuning demonstrates significant performance improvements. Furthermore, we assess speaker verification (SV) on our dataset, showing that, despite the challenges posed by the unique vocal characteristics of young children, the dataset effectively supports both ASR and SV tasks. This dataset is a valuable contribution to Mandarin child speech research and holds potential for applications in educational technology and child-computer interaction. It will be open-source and freely available for all academic purposes.

【4】 XWSB: A Blend System Utilizing XLS-R and WavLM with SLS Classifier detection system for SVDD 2024 Challenge
标题: XWSB:一个利用XLS-R和WavLM以及SLS分类器检测系统的混合系统,用于SDDD 2024挑战赛
作者:Qishan Zhang,Shuangbing Wen,Fangke Yan,Tao Hu,Jun Li
链接:点击下载PDF文件
摘要:本文介绍了SVDD 2024挑战赛中使用的模型结构。SVDD 2024挑战今年首次推出。歌唱声音深度假检测(SVDD),由于非正式的语音语调和不同的语音速率而面临复杂性。在本文中,我们提出了XWSB系统,它实现了SOTA per-percent在SVDD的挑战。XWSB代表XLS-R、WavLM和SLS Blend,代表这些技术的集成,用于SVDD。具体来说,我们使用了ASVspoof DF数据集中性能最好的模型结构XLS-R&SLS,并将SLS应用于WavLM以形成WavLM&SLS结构。最后,我们将两个模型集成在一起,形成了XWSB系统。实验结果表明,我们的系统表现出先进的识别能力,在SVDD的挑战,特别是实现了2.32%的EER在CtrSVDD轨道。代码和数据可以在https: github.com QiShanZhang XWSB_for_ SVDD 2024上找到。摘要:This paper introduces the model structure used in the SVDD 2024 Challenge. The SVDD 2024 challenge has been introduced this year for the first time. Singing voice deepfake detection (SVDD) which faces complexities due to informal speech intonations and varying speech rates. In this paper, we propose the XWSB system, which achieved SOTA per-formance in the SVDD challenge. XWSB stands for XLS-R, WavLM, and SLS Blend, representing the integration of these technologies for the purpose of SVDD. Specifically, we used the best performing model structure XLS-R&SLS from the ASVspoof DF dataset, and applied SLS to WavLM to form the WavLM&SLS structure. Finally, we integrated two models to form the XWSB system. Experimental results show that our system demonstrates advanced recognition capabilities in the SVDD challenge, specifically achieving an EER of 2.32% in the CtrSVDD track. The code and data can be found at https: github.com QiShanZhang XWSB_for_ SVDD2024.

【5】 EmoPro: A Prompt Selection Strategy for Emotional Expression in LM-based Speech Synthesis
标题: DeliverPro:基于LM的语音合成中情感表达的即时选择策略
作者:Haoyu Wang,Chunyu Qiang,Tianrui Wang,Cheng Gong,Qiuyu Liu,Yu Jiang,Xiaobao Wang,Chenyang Wang,Chen Zhang
链接:点击下载PDF文件
摘要:在广泛的数据集上训练的语音合成模型的最新进展已经证明了显着的zero-shot能力。这些模型可以根据提示输入控制生成语音中的内容、音色和情感。尽管有这些进步,提示的选择显着影响输出质量,但大多数现有的选择方案没有充分解决情绪强度的控制。针对这一问题,本文提出了一种两阶段提示选择策略,专门用于情感可控的语音合成。该策略的重点是选择高表达性和高质量的提示,从四个方面进行评估:情感表达强度,语音质量,文本情感一致性和模型生成性能。实验结果表明,使用所提出的方法选择的提示,结果在情感上更有表现力和从事合成语音相比,通过基线。音频样本和代码将在https: whyrrrrun.github.io EmoPro 上提供。摘要:Recent advancements in speech synthesis models, trained on extensive datasets, have demonstrated remarkable zero-shot capabilities. These models can control content, timbre, and emotion in generated speech based on prompt inputs. Despite these advancements, the choice of prompts significantly impacts the output quality, yet most existing selection schemes do not adequately address the control of emotional intensity. To address this question, this paper proposes a two-stage prompt selection strategy EmoPro, which is specifically designed for emotionally controllable speech synthesis. This strategy focuses on selecting highly expressive and high-quality prompts by evaluating them from four perspectives: emotional expression strength, speech quality, text-emotion consistency, and model generation performance. Experimental results show that prompts selected using the proposed method result in more emotionally expressive and engaging synthesized speech compared to those obtained through baseline. Audio samples and codes will be available at https: whyrrrrun.github.io EmoPro .

【6】 Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking
标题: 使用简单的N最佳重新排名在野外改善多语言ASB
作者:Brian Yan,Vineel Pratap,Shinji Watanabe,Michael Auli
链接:点击下载PDF文件
摘要:多语言自动语音识别(ASR)模型通常在语音话语的真实语言已知的环境中进行评估,然而,对于大多数实际环境来说,情况往往并非如此。自动口语识别(SLID)模型并不完美,错误分类对最终的ASR准确性有很大影响。在本文中,我们提出了一个简单而有效的N-最好的重新排名的方法,以提高多语言ASR的准确性,几个突出的声学模型,采用外部功能,如语言模型和基于文本的语言识别模型。我们使用MMS和Whisper模型对FLEURS的结果显示,口语识别准确率分别提高了8.7%和6.1%,单词错误率分别降低了3.3%和2.0%。摘要:Multilingual Automatic Speech Recognition (ASR) models are typically evaluated in a setting where the ground-truth language of the speech utterance is known, however, this is often not the case for most practical settings. Automatic Spoken Language Identification (SLID) models are not perfect and misclassifications have a substantial impact on the final ASR accuracy. In this paper, we present a simple and effective N-best re-ranking approach to improve multilingual ASR accuracy for several prominent acoustic models by employing external features such as language models and text-based language identification models. Our results on FLEURS using the MMS and Whisper models show spoken language identification accuracy improvements of 8.7% and 6.1%, respectively and word error rates which are 3.3% and 2.0% lower on these benchmarks.

【7】 Towards sub-millisecond latency real-time speech enhancement models on hearables
标题: 面向可听设备上的亚毫秒延迟实时语音增强模型
作者:Artem Dementyev,Chandan K. A. Reddy,Scott Wisdom,Navin Chatlani,John R. Hershey,Richard F. Lyon
链接:点击下载PDF文件
摘要:低延迟模型对于实时语音增强应用(例如助听器和可听设备)至关重要。然而,资源受限的听觉设备的亚毫秒级延迟空间仍然没有得到充分利用。我们演示了语音增强使用计算效率最低相位FIR滤波器,使样本的样本处理,以实现平均算法延迟0.32毫秒至1.25毫秒。用一个麦克风,我们观察到的平均SI-SDRi为4.1分贝。该方法显示了泛化,在看不见的音频记录上DNSMOS增加了0.2。我们使用一个轻量级的基于LSTM的644 k参数模型来生成FIR抽头。我们的基准测试,我们的系统可以运行在低功耗DSP与388 MIPS和平均端到端的延迟为3.35毫秒。我们提供了一个与基线低延迟频谱掩蔽技术的比较。我们希望这项工作能够更好地理解延迟,并可用于提高听觉设备的舒适性和可用性。摘要:Low latency models are critical for real-time speech enhancement applications, such as hearing aids and hearables. However, the sub-millisecond latency space for resource-constrained hearables remains underexplored. We demonstrate speech enhancement using a computationally efficient minimum-phase FIR filter, enabling sample-by-sample processing to achieve mean algorithmic latency of 0.32 ms to 1.25 ms. With a single microphone, we observe a mean SI-SDRi of 4.1 dB. The approach shows generalization with a DNSMOS increase of 0.2 on unseen audio recordings. We use a lightweight LSTM-based model of 644k parameters to generate FIR taps. We benchmark that our system can run on low-power DSP with 388 MIPS and mean end-to-end latency of 3.35 ms. We provide a comparison with baseline low-latency spectral masking techniques. We hope this work will enable a better understanding of latency and can be used to improve the comfort and usability of hearables.

【8】 A Fly on the Wall -- Exploiting Acoustic Side-Channels in Differential Pressure Sensors
标题: 墙上的苍蝇--利用压差传感器中的声学侧通道
作者:Yonatan Gizachew Achamyeleh,Mohamad Habib Fakih,Gabriel Garcia,Anomadarshi Barua,Mohammad Al Faruque
备注:Accepted to ACSAC 2024
链接:点击下载PDF文件
摘要:压差传感器被广泛用于监测关键环境。然而,我们的研究揭示了一个以前被忽视的漏洞:它们对压力变化的高度敏感性使它们容易受到声学侧信道攻击。我们证明了DPS中的压力传感膜片可以无意中捕获由语音引起的细微空气振动,这些振动通过传感器的组件传播并影响压力读数。利用这一发现,我们介绍了 textbf{BaroVox},一种新颖的攻击,从DPS读数重建语音,有效地把DPS变成了“墙上的苍蝇”。“我们模拟声音对DPS的影响,探索声泄漏的限制和挑战。为了克服这些挑战,我们提出了两种解决方案:使用独特的谱减法的信号处理方法和基于深度学习的关键字分类方法。在各种条件下的评估表明BaroVox的有效性,实现了0.29的人工识别和90.51%的自动识别的准确率的单词错误率。我们的研究结果强调了这种漏洞对隐私的重大影响。我们还讨论了潜在的防御策略,以减轻BaroVox带来的风险。摘要:Differential Pressure Sensors are widely deployed to monitor critical environments. However, our research unveils a previously overlooked vulnerability: their high sensitivity to pressure variations makes them susceptible to acoustic side-channel attacks. We demonstrate that the pressure-sensing diaphragms in DPS can inadvertently capture subtle air vibrations caused by speech, which propagate through the sensor's components and affect the pressure readings. Exploiting this discovery, we introduce textbf{BaroVox}, a novel attack that reconstructs speech from DPS readings, effectively turning DPS into a "fly on the wall." We model the effect of sound on DPS, exploring the limits and challenges of acoustic leakage. To overcome these challenges, we propose two solutions: a signal-processing approach using a unique spectral subtraction method and a deep learning-based approach for keyword classification. Evaluations under various conditions demonstrate BaroVox's effectiveness, achieving a word error rate of 0.29 for manual recognition and 90.51 % accuracy for automatic recognition. Our findings highlight the significant privacy implications of this vulnerability. We also discuss potential defense strategies to mitigate the risks posed by BaroVox.

【9】 Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects
标题: 文本2FX:利用CLAP嵌入实现文本引导音效
作者:Annie Chu,Patrick O'Reilly,Julia Barnett,Bryan Pardo
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:这项工作介绍了Text 2FX,这是一种利用CLAP嵌入和可微分数字信号处理来控制音频效果的方法,例如均衡和混响,使用开放词汇的自然语言提示(例如,“Make this sound in-your-face and bold”)。Text 2FX无需重新训练任何模型,而是依赖于现有嵌入空间内的单实例优化。我们表明,CLAP编码有价值的信息控制音频效果,并提出了两种优化方法,使用CLAP映射文本的音频效果参数。虽然我们用CLAP进行了演示,但这种方法适用于任何共享的文本音频嵌入空间。类似地,虽然我们用均衡和混响来演示,但是可以控制任何可区分的音频效果。我们进行了一个听众的研究与不同的文本提示和源音频,以评估这些方法与人类感知的质量和对齐。摘要:This work introduces Text2FX, a method that leverages CLAP embeddings and differentiable digital signal processing to control audio effects, such as equalization and reverberation, using open-vocabulary natural language prompts (e.g., "make this sound in-your-face and bold"). Text2FX operates without retraining any models, relying instead on single-instance optimization within the existing embedding space. We show that CLAP encodes valuable information for controlling audio effects and propose two optimization approaches using CLAP to map text to audio effect parameters. While we demonstrate with CLAP, this approach is applicable to any shared text-audio embedding space. Similarly, while we demonstrate with equalization and reverberation, any differentiable audio effect may be controlled. We conduct a listener study with diverse text prompts and source audio to evaluate the quality and alignment of these methods with human perception.

【10】 Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
标题: 语音增强:TWS耳机的低延迟实时语音增强
作者:Hanbin Bae,Pavel Andreev,Azat Saginbaev,Nicholas Babaev,Won-Jun Lee,Hosang Sung,Hoon-Young Cho
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:本文介绍了一种专为设备上使用的真正无线立体声(TWS)耳塞量身定制的语音增强解决方案。该解决方案专为支持嘈杂环境中的对话而设计,并激活了主动噪声消除(ANC)。在这种情况下,语音增强模型的主要挑战来自于限制设备上使用的计算复杂性和必须小于3 ms才能保留实时对话的延迟。为了解决这些问题,我们评估了几个关键的设计元素,包括网络架构和域,损失函数的设计,修剪方法和硬件特定的优化。因此,与基线模型相比,我们证明了语音增强质量的显著改善,同时降低了计算复杂度和算法延迟。摘要:This paper introduces a speech enhancement solution tailored for true wireless stereo (TWS) earbuds on-device usage. The solution was specifically designed to support conversations in noisy environments, with active noise cancellation (ANC) activated. The primary challenges for speech enhancement models in this context arise from computational complexity that limits on-device usage and latency that must be less than 3 ms to preserve a live conversation. To address these issues, we evaluated several crucial design elements, including the network architecture and domain, design of loss functions, pruning method, and hardware-specific optimization. Consequently, we demonstrated substantial improvements in speech enhancement quality compared with that in baseline models, while simultaneously reducing the computational complexity and algorithmic latency.

【11】 Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
标题: Speech-Mamba:使用选择性状态空间模型的长上下文语音识别
作者:Xiaoxue Gao,Nancy F. Chen
备注:8 pages; SLT 2024
链接:点击下载PDF文件
摘要:由于基于transformer的模型的高二次复杂度,当前的自动语音识别系统难以对长语音序列进行建模。选择性状态空间模型(如Mamba)在自然语言处理和计算机视觉任务中的长序列建模方面表现良好。然而,语音技术任务的研究工作一直未得到充分的探索。我们提出了Speech-Mamba,它在Transformer神经架构中集成了选择性状态空间建模。Speech-Mamba中具有选择性状态空间模型的长序列表示与基于transformer的建模的较低级别表示相辅相成。Speech-mamba实现了更好的能力来建模长距离依赖关系,因为它与序列长度近似线性地缩放。摘要:Current automatic speech recognition systems struggle with modeling long speech sequences due to high quadratic complexity of Transformer-based models. Selective state space models such as Mamba has performed well on long-sequence modeling in natural language processing and computer vision tasks. However, research endeavors in speech technology tasks has been under-explored. We propose Speech-Mamba, which incorporates selective state space modeling in Transformer neural architectures. Long sequence representations with selective state space models in Speech-Mamba is complemented with lower-level representations from Transformer-based modeling. Speech-mamba achieves better capacity to model long-range dependencies, as it scales near-linearly with sequence length.

【12】 The IEEE-IS2 2024 Music Packet Loss Concealment Challenge
标题: IEEE-IS 2 2024音乐数据包丢失隐藏挑战
作者:Alessandro Ilic Mezza,Alberto Bernardini
备注:8 pages, 4 figures, 3 tables. Official report of the IEEE-IS2 2024 Music Packet Loss Concealment Challenge, part of the 2nd International Workshop on Networked Immersive Audio
链接:点击下载PDF文件
摘要:IEEE-IS 2 2024音乐丢包隐藏挑战赛我们首先详细介绍挑战规则,然后概述所提供的基线系统、盲测集和用于确定最终排名的评估方法。该第一版旨在促进信号处理、机器学习和网络音乐表演领域的研究人员和从业人员之间的合作,同时也为音乐信号的丢包隐藏的未来发展奠定基础。摘要:We present the IEEE-IS2 2024 Music Packet Loss Concealment Challenge. We begin by detailing the challenge rules, followed by an overview of the provided baseline system, the blind test set, and the evaluation methodology used to determine the final ranking. This inaugural edition aimed to foster collaboration between researchers and practitioners from the fields of signal processing, machine learning, and networked music performance, while also laying the groundwork for future advancements in packet loss concealment for music signals.

【13】 MIMII-Gen: Generative Modeling Approach for Simulated Evaluation of Anomalous Sound Detection System
标题: MIII-Gen:异常声音检测系统模拟评估的生成式建模方法
作者:Harsh Purohit,Tomoya Nishida,Kota Dohi,Takashi Endo,Yohei Kawaguchi
链接:点击下载PDF文件
摘要:记录不足和异常的稀缺性对开发和验证机器声音的鲁棒异常检测系统提出了重大挑战。为了解决这些限制,我们提出了一种新的方法,使用集成编码器-解码器框架的基于潜在扩散的模型来生成机器声音中的各种异常。我们的方法利用Flan-T5模型对来自音频文件元数据的字幕进行编码,通过精心设计的U-Net架构实现条件生成。这种方法有助于我们的模型在EnCodec潜在空间内生成音频信号,确保高上下文相关性和质量。我们使用Fr 'echet音频距离(FAD)评分和其他指标客观地评估了我们生成的声音的质量,证明我们的方法在生成与实际异常情况非常相似的可靠机器音频方面优于现有模型。使用我们生成的数据对异常检测系统进行的评估显示出很强的相关性,曲线下面积(AUC)评分与原始值相差4.8%,验证了我们生成的数据的有效性。这些结果表明,我们的方法,以提高评估和鲁棒性的异常检测系统在不同的和以前看不见的条件的潜力。音频示例可以在 url{https: hpworkhub.github.io MIMII-Gen.github.io }找到。摘要:Insufficient recordings and the scarcity of anomalies present significant challenges in developing and validating robust anomaly detection systems for machine sounds. To address these limitations, we propose a novel approach for generating diverse anomalies in machine sound using a latent diffusion-based model that integrates an encoder-decoder framework. Our method utilizes the Flan-T5 model to encode captions derived from audio file metadata, enabling conditional generation through a carefully designed U-Net architecture. This approach aids our model in generating audio signals within the EnCodec latent space, ensuring high contextual relevance and quality. We objectively evaluated the quality of our generated sounds using the Fr 'echet Audio Distance (FAD) score and other metrics, demonstrating that our approach surpasses existing models in generating reliable machine audio that closely resembles actual abnormal conditions. The evaluation of the anomaly detection system using our generated data revealed a strong correlation, with the area under the curve (AUC) score differing by 4.8 % from the original, validating the effectiveness of our generated data. These results demonstrate the potential of our approach to enhance the evaluation and robustness of anomaly detection systems across varied and previously unseen conditions. Audio samples can be found at url{https: hpworkhub.github.io MIMII-Gen.github.io }.


eess.AS音频处理
【1】 Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects
标题: 文本2FX:利用CLAP嵌入实现文本引导音效
作者:Annie Chu,Patrick O'Reilly,Julia Barnett,Bryan Pardo
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:这项工作介绍了Text 2FX,这是一种利用CLAP嵌入和可微分数字信号处理来控制音频效果的方法,例如均衡和混响,使用开放词汇的自然语言提示(例如,“Make this sound in-your-face and bold”)。Text 2FX无需重新训练任何模型,而是依赖于现有嵌入空间内的单实例优化。我们表明,CLAP编码有价值的信息控制音频效果,并提出了两种优化方法,使用CLAP映射文本的音频效果参数。虽然我们用CLAP进行了演示,但这种方法适用于任何共享的文本音频嵌入空间。类似地,虽然我们用均衡和混响来演示,但是可以控制任何可区分的音频效果。我们进行了一个听众的研究与不同的文本提示和源音频,以评估这些方法与人类感知的质量和对齐。摘要:This work introduces Text2FX, a method that leverages CLAP embeddings and differentiable digital signal processing to control audio effects, such as equalization and reverberation, using open-vocabulary natural language prompts (e.g., "make this sound in-your-face and bold"). Text2FX operates without retraining any models, relying instead on single-instance optimization within the existing embedding space. We show that CLAP encodes valuable information for controlling audio effects and propose two optimization approaches using CLAP to map text to audio effect parameters. While we demonstrate with CLAP, this approach is applicable to any shared text-audio embedding space. Similarly, while we demonstrate with equalization and reverberation, any differentiable audio effect may be controlled. We conduct a listener study with diverse text prompts and source audio to evaluate the quality and alignment of these methods with human perception.

【2】 Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
标题: 语音增强:TWS耳机的低延迟实时语音增强
作者:Hanbin Bae,Pavel Andreev,Azat Saginbaev,Nicholas Babaev,Won-Jun Lee,Hosang Sung,Hoon-Young Cho
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:本文介绍了一种专为设备上使用的真正无线立体声(TWS)耳塞量身定制的语音增强解决方案。该解决方案专为支持嘈杂环境中的对话而设计,并激活了主动噪声消除(ANC)。在这种情况下,语音增强模型的主要挑战来自于限制设备上使用的计算复杂性和必须小于3 ms才能保留实时对话的延迟。为了解决这些问题,我们评估了几个关键的设计元素,包括网络架构和域,损失函数的设计,修剪方法和硬件特定的优化。因此,与基线模型相比,我们证明了语音增强质量的显著改善,同时降低了计算复杂度和算法延迟。摘要:This paper introduces a speech enhancement solution tailored for true wireless stereo (TWS) earbuds on-device usage. The solution was specifically designed to support conversations in noisy environments, with active noise cancellation (ANC) activated. The primary challenges for speech enhancement models in this context arise from computational complexity that limits on-device usage and latency that must be less than 3 ms to preserve a live conversation. To address these issues, we evaluated several crucial design elements, including the network architecture and domain, design of loss functions, pruning method, and hardware-specific optimization. Consequently, we demonstrated substantial improvements in speech enhancement quality compared with that in baseline models, while simultaneously reducing the computational complexity and algorithmic latency.

【3】 Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
标题: Speech-Mamba:使用选择性状态空间模型的长上下文语音识别
作者:Xiaoxue Gao,Nancy F. Chen
备注:8 pages; SLT 2024
链接:点击下载PDF文件
摘要:由于基于transformer的模型的高二次复杂度,当前的自动语音识别系统难以对长语音序列进行建模。选择性状态空间模型(如Mamba)在自然语言处理和计算机视觉任务中的长序列建模方面表现良好。然而,语音技术任务的研究工作一直未得到充分的探索。我们提出了Speech-Mamba,它在Transformer神经架构中集成了选择性状态空间建模。Speech-Mamba中具有选择性状态空间模型的长序列表示与基于Transformer的建模的较低级别表示进行了补充。Speech-mamba实现了更好的能力来建模长距离依赖关系,因为它与序列长度近似线性地缩放。摘要:Current automatic speech recognition systems struggle with modeling long speech sequences due to high quadratic complexity of Transformer-based models. Selective state space models such as Mamba has performed well on long-sequence modeling in natural language processing and computer vision tasks. However, research endeavors in speech technology tasks has been under-explored. We propose Speech-Mamba, which incorporates selective state space modeling in Transformer neural architectures. Long sequence representations with selective state space models in Speech-Mamba is complemented with lower-level representations from Transformer-based modeling. Speech-mamba achieves better capacity to model long-range dependencies, as it scales near-linearly with sequence length.

【4】 The IEEE-IS2 2024 Music Packet Loss Concealment Challenge
标题: IEEE-IS 2 2024音乐数据包丢失隐藏挑战
作者:Alessandro Ilic Mezza,Alberto Bernardini
备注:8 pages, 4 figures, 3 tables. Official report of the IEEE-IS2 2024 Music Packet Loss Concealment Challenge, part of the 2nd International Workshop on Networked Immersive Audio
链接:点击下载PDF文件
摘要:IEEE-IS 2 2024音乐丢包隐藏挑战赛我们首先详细介绍挑战规则,然后概述所提供的基线系统、盲测集和用于确定最终排名的评估方法。该第一版旨在促进信号处理、机器学习和网络音乐表演领域的研究人员和从业人员之间的合作,同时也为音乐信号的丢包隐藏的未来发展奠定基础。摘要:We present the IEEE-IS2 2024 Music Packet Loss Concealment Challenge. We begin by detailing the challenge rules, followed by an overview of the provided baseline system, the blind test set, and the evaluation methodology used to determine the final ranking. This inaugural edition aimed to foster collaboration between researchers and practitioners from the fields of signal processing, machine learning, and networked music performance, while also laying the groundwork for future advancements in packet loss concealment for music signals.

【5】 MIMII-Gen: Generative Modeling Approach for Simulated Evaluation of Anomalous Sound Detection System
标题: MIII-Gen:异常声音检测系统模拟评估的生成式建模方法
作者:Harsh Purohit,Tomoya Nishida,Kota Dohi,Takashi Endo,Yohei Kawaguchi
链接:点击下载PDF文件
摘要:记录不足和异常的稀缺性对开发和验证机器声音的鲁棒异常检测系统提出了重大挑战。为了解决这些限制,我们提出了一种新的方法,用于生成不同的异常机器声音使用潜在的扩散为基础的模型,集成了编码器-解码器框架。我们的方法利用Flan-T5模型对来自音频文件元数据的字幕进行编码,通过精心设计的U-Net架构实现条件生成。这种方法有助于我们的模型在EnCodec潜在空间内生成音频信号,确保高上下文相关性和质量。我们使用Fr 'echet音频距离(FAD)评分和其他指标客观地评估了我们生成的声音的质量,证明我们的方法在生成与实际异常情况非常相似的可靠机器音频方面优于现有模型。使用我们生成的数据对异常检测系统进行的评估显示出很强的相关性,曲线下面积(AUC)评分与原始值相差4.8%,验证了我们生成的数据的有效性。这些结果表明,我们的方法,以提高评估和鲁棒性的异常检测系统在不同的和以前看不见的条件的潜力。音频示例可以在 url{https: hpworkhub.github.io MIMII-Gen.github.io }找到。摘要:Insufficient recordings and the scarcity of anomalies present significant challenges in developing and validating robust anomaly detection systems for machine sounds. To address these limitations, we propose a novel approach for generating diverse anomalies in machine sound using a latent diffusion-based model that integrates an encoder-decoder framework. Our method utilizes the Flan-T5 model to encode captions derived from audio file metadata, enabling conditional generation through a carefully designed U-Net architecture. This approach aids our model in generating audio signals within the EnCodec latent space, ensuring high contextual relevance and quality. We objectively evaluated the quality of our generated sounds using the Fr 'echet Audio Distance (FAD) score and other metrics, demonstrating that our approach surpasses existing models in generating reliable machine audio that closely resembles actual abnormal conditions. The evaluation of the anomaly detection system using our generated data revealed a strong correlation, with the area under the curve (AUC) score differing by 4.8 % from the original, validating the effectiveness of our generated data. These results demonstrate the potential of our approach to enhance the evaluation and robustness of anomaly detection systems across varied and previously unseen conditions. Audio samples can be found at url{https: hpworkhub.github.io MIMII-Gen.github.io }.

【6】 Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models
标题: 超越单音频:推进音频大型语言模型中的多音频处理
作者:Yiming Chen,Xianghu Yue,Xiaoxue Gao,Chen Zhang,Luis Fernando D'Haro,Robby T. Tan,Haizhou Li
备注:EMNLP24 Findings
链接:点击下载PDF文件
摘要:最近已经探索了各种音频LLM(ALLM),用于使用单个统一模型同时处理不同的音频任务。虽然ALLM的现有评估主要集中在单音频任务上,但实际应用通常涉及同时处理多个音频流。为了弥合这一差距,我们提出了第一个多音频评估(MAE)基准,该基准由来自11个多音频任务的20个数据集组成,包括语音和声音场景。对MAE的综合实验表明,现有的ALLM虽然在理解单个音频输入中的主要音频元素方面很强大,但难以处理多音频场景。为此,我们提出了一种新的多音频LLM(MALLM),使用我们提出的合成数据上的判别学习来捕获多个相似音频之间的音频上下文。结果表明,所提出的MALLM优于所有基线,并实现了高数据效率,使用合成数据,而无需人工注释。提出的MALLM为ALLM打开了通往多音频处理时代的大门,并使我们更接近于在机器中复制人类听觉能力。摘要:Various audio-LLMs (ALLMs) have been explored recently for tackling different audio tasks simultaneously using a single, unified model. While existing evaluations of ALLMs primarily focus on single-audio tasks, real-world applications often involve processing multiple audio streams simultaneously. To bridge this gap, we propose the first multi-audio evaluation (MAE) benchmark that consists of 20 datasets from 11 multi-audio tasks encompassing both speech and sound scenarios. Comprehensive experiments on MAE demonstrate that the existing ALLMs, while being powerful in comprehending primary audio elements in individual audio inputs, struggling to handle multi-audio scenarios. To this end, we propose a novel multi-audio-LLM (MALLM) to capture audio context among multiple similar audios using discriminative learning on our proposed synthetic data. The results demonstrate that the proposed MALLM outperforms all baselines and achieves high data efficiency using synthetic data without requiring human annotations. The proposed MALLM opens the door for ALLMs towards multi-audio processing era and brings us closer to replicating human auditory capabilities in machines.

【7】 Audio-Based Linguistic Feature Extraction for Enhancing Multi-lingual and Low-Resource Text-to-Speech
标题: 基于音频的语言特征提取用于增强多语言和低资源的文本到语音
作者:Youngjae Kim,Yejin Jeon,Gary Geunbae Lee
备注:EMNLP 2024 Findings
链接:点击下载PDF文件
摘要:获取丰富、高质量数据的困难,特别是在多语言环境中,引发了人们对解决低资源情景的兴趣。此外,目前的文献依赖于固定的表达式从语言ID,这导致在语言表示的学习不足,并生成语音在看不见的语言失败。为了解决这些挑战,我们提出了一种新的方法,直接从音频输入中提取语言特征,同时有效地过滤掉杂项声学信息,包括扬声器特定的属性,如音色。主观和客观的评价肯定了我们的方法对多语言文本到语音的有效性,并突出了其在以前看不见的语言的低资源迁移学习的优势。摘要:The difficulty of acquiring abundant, high-quality data, especially in multi-lingual contexts, has sparked interest in addressing low-resource scenarios. Moreover, current literature rely on fixed expressions from language IDs, which results in the inadequate learning of language representations, and the failure to generate speech in unseen languages. To address these challenges, we propose a novel method that directly extracts linguistic features from audio input while effectively filtering out miscellaneous acoustic information including speaker-specific attributes like timbre. Subjective and objective evaluations affirm the effectiveness of our approach for multi-lingual text-to-speech, and highlight its superiority in low-resource transfer learning for previously unseen language.

【8】 ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5
标题: ChildMandarin:针对3-5岁幼儿的全面普通话语音数据集
作者:Jiaming Zhou,Shiyao Wang,Shiwan Zhao,Jiabei He,Haoqin Sun,Hui Wang,Cheng Liu,Aobo Kong,Yujie Guo,Yong Qin
链接:点击下载PDF文件
摘要:自动语音识别(ASR)系统已经通过Whisper、Conformer等模型以及Wav 2 vec 2.0和HuBERT等自监督框架取得了显着进步。然而,开发强大的ASR模型,幼儿的语音仍然具有挑战性,由于发音,语调和速度相比,成人语音的差异。在本文中,我们介绍了一个新的普通话语音数据集,专注于3至5岁的儿童,解决了这方面的资源稀缺。该数据集包括41.25小时的演讲,经过精心制作的人工翻译,收集了来自中国各省的397名演讲者,性别代表性均衡。我们提供了一个全面的分析,扬声器人口统计,演讲持续时间分布和地理覆盖范围。此外,我们还评估了从头开始训练的模型(如Conformer)以及微调的预训练模型(如HuBERT和Whisper)的ASR性能,其中微调显示了显着的性能改进。此外,我们在我们的数据集上评估了说话人验证(SV),结果表明,尽管幼儿独特的声音特征带来了挑战,但该数据集有效地支持ASR和SV任务。该数据集对汉语儿童语音研究是一个有价值的贡献,并具有在教育技术和儿童-计算机交互中的应用潜力。它将是开源的,可免费用于所有学术目的。摘要:Automatic speech recognition (ASR) systems have advanced significantly with models like Whisper, Conformer, and self-supervised frameworks such as Wav2vec 2.0 and HuBERT. However, developing robust ASR models for young children's speech remains challenging due to differences in pronunciation, tone, and pace compared to adult speech. In this paper, we introduce a new Mandarin speech dataset focused on children aged 3 to 5, addressing the scarcity of resources in this area. The dataset comprises 41.25 hours of speech with carefully crafted manual transcriptions, collected from 397 speakers across various provinces in China, with balanced gender representation. We provide a comprehensive analysis of speaker demographics, speech duration distribution and geographic coverage. Additionally, we evaluate ASR performance on models trained from scratch, such as Conformer, as well as fine-tuned pre-trained models like HuBERT and Whisper, where fine-tuning demonstrates significant performance improvements. Furthermore, we assess speaker verification (SV) on our dataset, showing that, despite the challenges posed by the unique vocal characteristics of young children, the dataset effectively supports both ASR and SV tasks. This dataset is a valuable contribution to Mandarin child speech research and holds potential for applications in educational technology and child-computer interaction. It will be open-source and freely available for all academic purposes.

【9】 XWSB: A Blend System Utilizing XLS-R and WavLM with SLS Classifier detection system for SVDD 2024 Challenge
标题: XWSB:一个利用XLS-R和WavLM以及SLS分类器检测系统的混合系统,用于SDDD 2024挑战赛
作者:Qishan Zhang,Shuangbing Wen,Fangke Yan,Tao Hu,Jun Li
链接:点击下载PDF文件
摘要:本文介绍了SVDD 2024挑战赛中使用的模型结构。SVDD 2024挑战今年首次推出。歌唱声音深度假检测(SVDD),由于非正式的语音语调和不同的语音速率而面临复杂性。在本文中,我们提出了XWSB系统,它实现了SOTA per-percent在SVDD的挑战。XWSB代表XLS-R、WavLM和SLS Blend,代表这些技术的集成,用于SVDD。具体来说,我们使用了ASVspoof DF数据集中性能最好的模型结构XLS-R&SLS,并将SLS应用于WavLM以形成WavLM&SLS结构。最后,我们将两个模型集成在一起,形成了XWSB系统。实验结果表明,我们的系统表现出先进的识别能力,在SVDD的挑战,特别是实现了2.32%的EER在CtrSVDD轨道。代码和数据可以在https: github.com QiShanZhang XWSB_for_ SVDD 2024上找到。摘要:This paper introduces the model structure used in the SVDD 2024 Challenge. The SVDD 2024 challenge has been introduced this year for the first time. Singing voice deepfake detection (SVDD) which faces complexities due to informal speech intonations and varying speech rates. In this paper, we propose the XWSB system, which achieved SOTA per-formance in the SVDD challenge. XWSB stands for XLS-R, WavLM, and SLS Blend, representing the integration of these technologies for the purpose of SVDD. Specifically, we used the best performing model structure XLS-R&SLS from the ASVspoof DF dataset, and applied SLS to WavLM to form the WavLM&SLS structure. Finally, we integrated two models to form the XWSB system. Experimental results show that our system demonstrates advanced recognition capabilities in the SVDD challenge, specifically achieving an EER of 2.32% in the CtrSVDD track. The code and data can be found at https: github.com QiShanZhang XWSB_for_ SVDD2024.

【10】 EmoPro: A Prompt Selection Strategy for Emotional Expression in LM-based Speech Synthesis
标题: DeliverPro:基于LM的语音合成中情感表达的即时选择策略
作者:Haoyu Wang,Chunyu Qiang,Tianrui Wang,Cheng Gong,Qiuyu Liu,Yu Jiang,Xiaobao Wang,Chenyang Wang,Chen Zhang
链接:点击下载PDF文件
摘要:在广泛的数据集上训练的语音合成模型的最新进展已经证明了显着的zero-shot能力。这些模型可以根据提示输入控制生成语音中的内容、音色和情感。尽管有这些进步,提示的选择显着影响输出质量,但大多数现有的选择方案没有充分解决情绪强度的控制。针对这一问题,本文提出了一种两阶段提示选择策略,专门用于情感可控的语音合成。该策略的重点是选择高表达性和高质量的提示,从四个方面进行评估:情感表达强度,语音质量,文本情感一致性和模型生成性能。实验结果表明,使用所提出的方法选择的提示,结果在情感上更有表现力和从事合成语音相比,通过基线。音频样本和代码将在https: whyrrrrun.github.io EmoPro 上提供。摘要:Recent advancements in speech synthesis models, trained on extensive datasets, have demonstrated remarkable zero-shot capabilities. These models can control content, timbre, and emotion in generated speech based on prompt inputs. Despite these advancements, the choice of prompts significantly impacts the output quality, yet most existing selection schemes do not adequately address the control of emotional intensity. To address this question, this paper proposes a two-stage prompt selection strategy EmoPro, which is specifically designed for emotionally controllable speech synthesis. This strategy focuses on selecting highly expressive and high-quality prompts by evaluating them from four perspectives: emotional expression strength, speech quality, text-emotion consistency, and model generation performance. Experimental results show that prompts selected using the proposed method result in more emotionally expressive and engaging synthesized speech compared to those obtained through baseline. Audio samples and codes will be available at https: whyrrrrun.github.io EmoPro .

【11】 Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking
标题: 使用简单的N最佳重新排名在野外改善多语言ASB
作者:Brian Yan,Vineel Pratap,Shinji Watanabe,Michael Auli
链接:点击下载PDF文件
摘要:多语言自动语音识别(ASR)模型通常在语音话语的地面真实语言已知的设置中进行评估,然而,对于大多数实际设置来说,情况往往并非如此。自动口语识别(SLID)模型并不完美,错误分类对最终的ASR准确性有很大影响。在本文中,我们提出了一个简单而有效的N-最好的重新排名的方法,以提高多语言ASR的准确性,几个突出的声学模型,采用外部功能,如语言模型和基于文本的语言识别模型。我们使用MMS和Whisper模型对FLEURS的结果显示,口语识别准确率分别提高了8.7%和6.1%,单词错误率分别降低了3.3%和2.0%。摘要:Multilingual Automatic Speech Recognition (ASR) models are typically evaluated in a setting where the ground-truth language of the speech utterance is known, however, this is often not the case for most practical settings. Automatic Spoken Language Identification (SLID) models are not perfect and misclassifications have a substantial impact on the final ASR accuracy. In this paper, we present a simple and effective N-best re-ranking approach to improve multilingual ASR accuracy for several prominent acoustic models by employing external features such as language models and text-based language identification models. Our results on FLEURS using the MMS and Whisper models show spoken language identification accuracy improvements of 8.7% and 6.1%, respectively and word error rates which are 3.3% and 2.0% lower on these benchmarks.

【12】 Towards sub-millisecond latency real-time speech enhancement models on hearables
标题: 面向可听设备上的亚毫秒延迟实时语音增强模型
作者:Artem Dementyev,Chandan K. A. Reddy,Scott Wisdom,Navin Chatlani,John R. Hershey,Richard F. Lyon
链接:点击下载PDF文件
摘要:低延迟模型对于实时语音增强应用(例如助听器和可听设备)至关重要。然而,资源受限的听觉设备的亚毫秒级延迟空间仍然没有得到充分利用。我们演示了语音增强使用计算效率最低相位FIR滤波器,使样本的样本处理,以实现平均算法延迟0.32毫秒至1.25毫秒。用一个麦克风,我们观察到的平均SI-SDRi为4.1分贝。该方法显示了泛化,在看不见的音频记录上DNSMOS增加了0.2。我们使用一个轻量级的基于LSTM的644 k参数模型来生成FIR抽头。我们的基准测试表明,我们的系统可以在388 MIPS的低功耗DSP上运行,平均端到端延迟为3.35 ms。我们提供了与基线低延迟频谱掩蔽技术的比较。我们希望这项工作能够更好地理解延迟,并可用于提高听觉设备的舒适性和可用性。摘要:Low latency models are critical for real-time speech enhancement applications, such as hearing aids and hearables. However, the sub-millisecond latency space for resource-constrained hearables remains underexplored. We demonstrate speech enhancement using a computationally efficient minimum-phase FIR filter, enabling sample-by-sample processing to achieve mean algorithmic latency of 0.32 ms to 1.25 ms. With a single microphone, we observe a mean SI-SDRi of 4.1 dB. The approach shows generalization with a DNSMOS increase of 0.2 on unseen audio recordings. We use a lightweight LSTM-based model of 644k parameters to generate FIR taps. We benchmark that our system can run on low-power DSP with 388 MIPS and mean end-to-end latency of 3.35 ms. We provide a comparison with baseline low-latency spectral masking techniques. We hope this work will enable a better understanding of latency and can be used to improve the comfort and usability of hearables.

【13】 A Fly on the Wall -- Exploiting Acoustic Side-Channels in Differential Pressure Sensors
标题: 墙上的苍蝇--利用压差传感器中的声学侧通道
作者:Yonatan Gizachew Achamyeleh,Mohamad Habib Fakih,Gabriel Garcia,Anomadarshi Barua,Mohammad Al Faruque
备注:Accepted to ACSAC 2024
链接:点击下载PDF文件
摘要:压差传感器被广泛用于监测关键环境。然而,我们的研究揭示了一个以前被忽视的漏洞:它们对压力变化的高度敏感性使它们容易受到声学侧信道攻击。我们证明了DPS中的压力传感膜片可以无意中捕获由语音引起的细微空气振动,这些振动通过传感器的组件传播并影响压力读数。利用这一发现,我们介绍了 textbf{BaroVox},一种新颖的攻击,从DPS读数重建语音,有效地把DPS变成了“墙上的苍蝇”。“我们模拟声音对DPS的影响,探索声泄漏的限制和挑战。为了克服这些挑战,我们提出了两种解决方案:使用独特的谱减法的信号处理方法和基于深度学习的关键字分类方法。在各种条件下的评估表明BaroVox的有效性,实现了0.29的人工识别和90.51%的自动识别的准确率的单词错误率。我们的研究结果强调了这种漏洞对隐私的重大影响。我们还讨论了潜在的防御策略,以减轻BaroVox带来的风险。摘要:Differential Pressure Sensors are widely deployed to monitor critical environments. However, our research unveils a previously overlooked vulnerability: their high sensitivity to pressure variations makes them susceptible to acoustic side-channel attacks. We demonstrate that the pressure-sensing diaphragms in DPS can inadvertently capture subtle air vibrations caused by speech, which propagate through the sensor's components and affect the pressure readings. Exploiting this discovery, we introduce textbf{BaroVox}, a novel attack that reconstructs speech from DPS readings, effectively turning DPS into a "fly on the wall." We model the effect of sound on DPS, exploring the limits and challenges of acoustic leakage. To overcome these challenges, we propose two solutions: a signal-processing approach using a unique spectral subtraction method and a deep learning-based approach for keyword classification. Evaluations under various conditions demonstrate BaroVox's effectiveness, achieving a word error rate of 0.29 for manual recognition and 90.51 % accuracy for automatic recognition. Our findings highlight the significant privacy implications of this vulnerability. We also discuss potential defense strategies to mitigate the risks posed by BaroVox.


机器翻译,仅供参考