今日论文合集:cs.SD语音5篇,eess.AS音频处理9篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Vivo : une approche multimodale de la synthese concatenative par corpus  dans le cadre d'une oeuvre audiovisuelle immersive
标题:Vivo:沉浸式视听素材中的合成多模式
链接:https://arxiv.org/abs/2404.10578
作者:Mateo Fayet
备注:in French language
摘要:哪些视觉描述符适用于多模态交互,以及如何通过实时视频数据分析将它们集成到基于语料库的拼接合成声音系统中。
摘要:Which visual descriptors are suitable for multi-modal interaction and how to integrate them via real-time video data analysis into a corpus-based concatenative synthesis sound system.

【2】 Multiple Mobile Target Detection and Tracking in Active Sonar Array  Using a Track-Before-Detect Approach
标题:采用先检测前跟踪方法的主动声纳阵中多移动目标检测和跟踪
链接:https://arxiv.org/abs/2404.10316
作者:Avi Abu,Nikola Miskovic,Oleg Chebotar,Neven Cukrov,Roee Diamant
备注:10 pages, 10 figures
摘要:本文提出了一种利用水听器阵列接收宽带线性调频信号的主动声发射来检测和跟踪水下运动目标的算法。该方法通过对接收到的回波序列进行检测前跟踪处理,克服了虚警率高的问题。通过执行延迟和求和波束成形以及脉冲压缩,为从每个发射的探测信号接收的混响创建2D时空矩阵。结果通过2D恒虚警率(CFAR)检测器进行滤波,以识别与潜在目标对应的反射模式。用于多个探测器传输的紧密间隔的信号被组合成斑点,以避免对单个对象的多次检测。针对多目标跟踪问题,提出了一种基于近等速模型的检测前跟踪方法。位置和速度的估计由去偏转换测量卡尔曼滤波器。结果进行了分析,模拟场景和实验在海上,GPS标记的镀金头海鱼进行了跟踪。与两个基准方案相比,结果表明,良好的跟踪连续性和准确性,是鲁棒的检测阈值的选择。
摘要:We present an algorithm for detecting and tracking underwater mobile objects using active acoustic transmission of broadband chirp signals whose reflections are received by a hydrophone array. The method overcomes the problem of high false alarm rate by applying a track-before-detect ap- proach to the sequence of received reflections. A 2D time- space matrix is created for the reverberations received from each transmitted probe signal by performing delay and sum beamforming and pulse compression. The result is filtered by a 2D constant false alarm rate (CFAR) detector to identify reflection patterns corresponding to potential targets. Closely spaced signals for multiple probe transmissions are combined into blobs to avoid multiple detections of a single object. A track- before-detect method using a Nearly Constant Velocity (NCV) model is employed to track multiple objects. The position and velocity is estimated by the debiased converted measurement Kalman filter. Results are analyzed for simulated scenarios and for experiments at sea, where GPS tagged gilt-head seabream fish were tracked. Compared to two benchmark schemes, the results show a favorable track continuity and accuracy that is robust to the choice of detection threshold.


【3】 Long-form music generation with latent diffusion
标题:具有潜在扩散性的长篇音乐一代
链接:https://arxiv.org/abs/2404.10301
作者:Zach Evans,Julian D. Parker,CJ Carr,Zack Zukowski,Josiah Taylor,Jordi Pons
摘要:基于音频的音乐生成模型最近取得了很大的进步,但到目前为止还没有成功地生成具有连贯音乐结构的完整音乐曲目。我们表明,通过在长时间背景下训练生成模型,可以产生长达4m45s的长形式音乐。我们的模型由一个扩散变压器操作高度下采样的连续潜在表示(潜在率为21.5Hz)。根据音频质量和即时对齐的指标,它获得了最先进的几代产品,主观测试表明,它可以产生具有连贯结构的全长音乐。
摘要:Audio-based generative models for music have seen great strides recently, but so far have not managed to produce full-length music tracks with coherent musical structure. We show that by training a generative model on long temporal contexts it is possible to produce long-form music of up to 4m45s. Our model consists of a diffusion-transformer operating on a highly downsampled continuous latent representation (latent rate of 21.5Hz). It obtains state-of-the-art generations according to metrics on audio quality and prompt alignment, and subjective tests reveal that it produces full-length music with coherent structure.

【4】 Clustering and Data Augmentation to Improve Accuracy of Sleep Assessment  and Sleep Individuality Analysis
标题:集群和数据增强提高睡眠评估和睡眠个性分析的准确性
链接:https://arxiv.org/abs/2404.10299
作者:Shintaro Tamai,Masayuki Numao,Ken-ichi Fukui
摘要:最近,越来越多的健康意识,新的方法允许个人在家里监测睡眠。利用睡眠声音提供了传统方法的优势,如智能手表,是非侵入性的,并且能够检测各种生理活动。该研究旨在构建一个基于机器学习的睡眠评估模型,提供基于证据的评估,例如由于睡眠开始期间频繁运动而导致的睡眠不佳。提取睡眠声音事件,使用VAE导出潜在表示,使用GMM进行聚类,并训练LSTM进行主观睡眠评估,在区分睡眠满意度方面达到了94.8%的高准确率。此外,TimeSHAP揭示了不同个体在有影响力的声音事件类型和时间上的差异。
摘要:Recently, growing health awareness, novel methods allow individuals to monitor sleep at home. Utilizing sleep sounds offers advantages over conventional methods like smartwatches, being non-intrusive, and capable of detecting various physiological activities. This study aims to construct a machine learning-based sleep assessment model providing evidence-based assessments, such as poor sleep due to frequent movement during sleep onset. Extracting sleep sound events, deriving latent representations using VAE, clustering with GMM, and training LSTM for subjective sleep assessment achieved a high accuracy of 94.8% in distinguishing sleep satisfaction. Moreover, TimeSHAP revealed differences in impactful sound event types and timings for different individuals.

【5】 PRODIS - a speech database and a phoneme-based language model for the  study of predictability effects in Polish
标题:PRODIS -一个语音数据库和基于音素的语言模型,用于研究波兰语的可预测性影响
链接:https://arxiv.org/abs/2404.10112
作者:Zofia Malisz,Jan Foremski,Małgorzata Kul
备注:To appear in the proceedings of LREC2024: Language Resources and Evaluation Conference 2024, Turin, Italy
摘要:我们提出了一个语音数据库和波兰语的音素级语言模型。该数据库和模型的目的是分析韵律和话语因素及其对声学参数的影响与可预测性的影响。该数据库也是第一个大型的、公开可用的波兰语语音语料库,具有出色的声学质量,可用于语音分析和多说话人语音技术系统的训练。数据库中的语音在管道中处理,实现了90%的自动化程度。它包含了最先进的、免费提供的工具,使数据库能够扩展或适应其他语言。
摘要:We present a speech database and a phoneme-level language model of Polish. The database and model are designed for the analysis of prosodic and discourse factors and their impact on acoustic parameters in interaction with predictability effects. The database is also the first large, publicly available Polish speech corpus of excellent acoustic quality that can be used for phonetic analysis and training of multi-speaker speech technology systems. The speech in the database is processed in a pipeline that achieves a 90% degree of automation. It incorporates state-of-the-art, freely available tools enabling database expansion or adaptation to additional languages.


eess.AS音频处理
【1】 MAD Speech: Measures of Acoustic Diversity of Speech
标题:MAD语音:语音声学多样性的测量
链接:https://arxiv.org/abs/2404.10419
作者:Matthieu Futeral,Andrea Agostinelli,Marco Tagliasacchi,Neil Zeghidour,Eugene Kharitonov
摘要:生成式口语模型在各种声音、韵律和记录条件下产生语音,似乎接近自然语音的多样性。然而,在何种程度上产生的语音是声学多样性仍然不清楚,由于缺乏适当的指标。我们通过开发声学多样性的轻量级指标来解决这一差距,我们统称为MAD语音。我们专注于测量声学多样性的五个方面:声音,性别,情感,口音和背景噪音。我们构建的指标作为一个组合的专门的,每方面的嵌入模型和聚合函数,衡量嵌入空间内的多样性。接下来,我们构建一系列数据集,每个方面都有先验已知的多样性偏好。使用这些数据集,我们证明了我们提出的指标比基线与地面真实多样性的一致性更强。最后,我们展示了我们提出的指标在几个现实生活中的评估方案的适用性。MAD演讲将公开访问。
摘要:Generative spoken language models produce speech in a wide range of voices, prosody, and recording conditions, seemingly approaching the diversity of natural speech. However, the extent to which generated speech is acoustically diverse remains unclear due to a lack of appropriate metrics. We address this gap by developing lightweight metrics of acoustic diversity, which we collectively refer to as MAD Speech. We focus on measuring five facets of acoustic diversity: voice, gender, emotion, accent, and background noise. We construct the metrics as a composition of specialized, per-facet embedding models and an aggregation function that measures diversity within the embedding space. Next, we build a series of datasets with a priori known diversity preferences for each facet. Using these datasets, we demonstrate that our proposed metrics achieve a stronger agreement with the ground-truth diversity than baselines. Finally, we showcase the applicability of our proposed metrics across several real-life evaluation scenarios. MAD Speech will be made publicly accessible.

【2】 Wireless Earphone-based Real-Time Monitoring of Breathing Exercises: A  Deep Learning Approach
标题:基于无线耳机的呼吸练习实时监控:深度学习方法
链接:https://arxiv.org/abs/2404.10310
作者:Hassam Khan Wazir,Zaid Waghoo,Vikram Kapila
备注:4 pages, 2 figures. Paper accepted at IEEE International Conference on Engineering in Medicine & Biology Society, 2024
摘要:几种常规治疗需要深呼吸练习作为关键组成部分,接受这种治疗的患者必须定期进行这些练习。评估治疗的结果和定制其过程需要监测患者对治疗的依从性。虽然治疗依从性监测在临床环境中是常规的,但在家庭环境中进行具有挑战性。这是因为家庭环境缺乏对有效地监测患者的治疗例程的执行所需的专门设备和熟练专业人员的访问。对于某些类型的治疗,这些挑战可以通过使用消费级硬件来解决,例如耳机和智能手机,作为实用的解决方案。为了准确地监测呼吸练习使用无线耳机,本文提出了一个框架,有可能评估病人的依从性在家里的治疗。该系统通过两个卷积神经网络处理$\mathbf{500}$ ms的音频信号,以高精度进行呼吸相位和通道的实时检测。第一个网络称为通道分类器,区分鼻呼吸和口呼吸以及停顿。第二个网络称为相位分类器,确定音频片段是来自吸气还是呼气。根据$k$倍交叉验证,通道和相位分类器分别获得了$\mathbf{97.99\%}$和$\mathbf{89.46\%}$的最大F1得分。结果表明,使用商品耳机进行呼吸治疗依从性监测的实时呼吸通道和相位检测的潜力。
摘要:Several therapy routines require deep breathing exercises as a key component and patients undergoing such therapies must perform these exercises regularly. Assessing the outcome of a therapy and tailoring its course necessitates monitoring a patient's compliance with the therapy. While therapy compliance monitoring is routine in a clinical environment, it is challenging to do in an at-home setting. This is so because a home setting lacks access to specialized equipment and skilled professionals needed to effectively monitor the performance of a therapy routine by a patient. For some types of therapies, these challenges can be addressed with the use of consumer-grade hardware, such as earphones and smartphones, as practical solutions. To accurately monitor breathing exercises using wireless earphones, this paper proposes a framework that has the potential for assessing a patient's compliance with an at-home therapy. The proposed system performs real-time detection of breathing phases and channels with high accuracy by processing a $\mathbf{500}$ ms audio signal through two convolutional neural networks. The first network, called a channel classifier, distinguishes between nasal and oral breathing, and a pause. The second network, called a phase classifier, determines whether the audio segment is from inhalation or exhalation. According to $k$-fold cross-validation, the channel and phase classifiers achieved a maximum F1 score of $\mathbf{97.99\%}$ and $\mathbf{89.46\%}$, respectively. The results demonstrate the potential of using commodity earphones for real-time breathing channel and phase detection for breathing therapy compliance monitoring.

【3】 Vivo : une approche multimodale de la synthese concatenative par corpus  dans le cadre d'une oeuvre audiovisuelle immersive
标题:Vivo:沉浸式视听素材中的合成多模式
链接:https://arxiv.org/abs/2404.10578
作者:Mateo Fayet
备注:in French language
摘要:哪些视觉描述符适用于多模态交互,以及如何通过实时视频数据分析将它们集成到基于语料库的拼接合成声音系统中。
摘要:Which visual descriptors are suitable for multi-modal interaction and how to integrate them via real-time video data analysis into a corpus-based concatenative synthesis sound system.

【4】 Language Proficiency and F0 Entrainment: A Study of L2 English Imitation  in Italian, French, and Slovak Speakers
标题:语言能力和F0融入:意大利语、法语和斯洛伐克语使用者的L2英语模仿研究
链接:https://arxiv.org/abs/2404.10440
作者:Zheng Yuan,Štefan Beňuš,Alessandro D'Ausilio
备注:Accepted at Speech Prosody 2024
摘要:本研究探讨了交替阅读任务中第二语言(L2)英语语音模仿中的F0夹带。意大利,法国和斯洛伐克的母语的参与者模仿英语的话语,和他们的F0夹带量化使用动态时间规整(DTW)的距离之间的参数化F0轮廓的模仿的话语和那些的模型话语。结果表明,二语英语水平和夹带之间的微妙关系:具有较高水平的扬声器一般表现出较少的夹带在音高变化和decompression。然而,在二人组中,更熟练的说话者表现出更大的模仿音高范围的能力,从而增加了夹带。这表明,熟练程度的影响夹带不同的个人和二元水平,突出了复杂的语言技能和韵律适应之间的相互作用。
摘要:This study explores F0 entrainment in second language (L2) English speech imitation during an Alternating Reading Task (ART). Participants with Italian, French, and Slovak native languages imitated English utterances, and their F0 entrainment was quantified using the Dynamic Time Warping (DTW) distance between the parameterized F0 contours of the imitated utterances and those of the model utterances. Results indicate a nuanced relationship between L2 English proficiency and entrainment: speakers with higher proficiency generally exhibit less entrainment in pitch variation and declination. However, within dyads, the more proficient speakers demonstrate a greater ability to mimic pitch range, leading to increased entrainment. This suggests that proficiency influences entrainment differently at individual and dyadic levels, highlighting the complex interplay between language skill and prosodic adaptation.


【5】 Multiple Mobile Target Detection and Tracking in Active Sonar Array  Using a Track-Before-Detect Approach
标题:采用先检测前跟踪方法的主动声纳阵中多移动目标检测和跟踪
链接:https://arxiv.org/abs/2404.10316
作者:Avi Abu,Nikola Miskovic,Oleg Chebotar,Neven Cukrov,Roee Diamant
备注:10 pages, 10 figures
摘要:本文提出了一种利用水听器阵列接收宽带线性调频信号的主动声发射来检测和跟踪水下运动目标的算法。该方法通过对接收到的回波序列进行检测前跟踪处理,克服了虚警率高的问题。通过执行延迟和求和波束成形以及脉冲压缩,为从每个发射的探测信号接收的混响创建2D时空矩阵。结果通过2D恒虚警率(CFAR)检测器进行滤波,以识别与潜在目标对应的反射模式。用于多个探测器传输的紧密间隔的信号被组合成斑点,以避免对单个对象的多次检测。针对多目标跟踪问题,提出了一种基于近等速模型的检测前跟踪方法。位置和速度的估计由去偏转换测量卡尔曼滤波器。结果进行了分析,模拟场景和实验在海上,GPS标记的镀金头海鱼进行了跟踪。与两个基准方案相比,结果表明,良好的跟踪连续性和准确性,是鲁棒的检测阈值的选择。
摘要:We present an algorithm for detecting and tracking underwater mobile objects using active acoustic transmission of broadband chirp signals whose reflections are received by a hydrophone array. The method overcomes the problem of high false alarm rate by applying a track-before-detect ap- proach to the sequence of received reflections. A 2D time- space matrix is created for the reverberations received from each transmitted probe signal by performing delay and sum beamforming and pulse compression. The result is filtered by a 2D constant false alarm rate (CFAR) detector to identify reflection patterns corresponding to potential targets. Closely spaced signals for multiple probe transmissions are combined into blobs to avoid multiple detections of a single object. A track- before-detect method using a Nearly Constant Velocity (NCV) model is employed to track multiple objects. The position and velocity is estimated by the debiased converted measurement Kalman filter. Results are analyzed for simulated scenarios and for experiments at sea, where GPS tagged gilt-head seabream fish were tracked. Compared to two benchmark schemes, the results show a favorable track continuity and accuracy that is robust to the choice of detection threshold.


【6】 Long-form music generation with latent diffusion
标题:具有潜在扩散性的长篇音乐一代
链接:https://arxiv.org/abs/2404.10301
作者:Zach Evans,Julian D. Parker,CJ Carr,Zack Zukowski,Josiah Taylor,Jordi Pons
摘要:基于音频的音乐生成模型最近取得了很大的进步,但到目前为止还没有成功地生成具有连贯音乐结构的完整音乐曲目。我们表明,通过在长时间背景下训练生成模型,可以产生长达4m45s的长形式音乐。我们的模型由一个扩散变压器操作高度下采样的连续潜在表示(潜在率为21.5Hz)。根据音频质量和即时对齐的指标,它获得了最先进的几代产品,主观测试表明,它可以产生具有连贯结构的全长音乐。
摘要:Audio-based generative models for music have seen great strides recently, but so far have not managed to produce full-length music tracks with coherent musical structure. We show that by training a generative model on long temporal contexts it is possible to produce long-form music of up to 4m45s. Our model consists of a diffusion-transformer operating on a highly downsampled continuous latent representation (latent rate of 21.5Hz). It obtains state-of-the-art generations according to metrics on audio quality and prompt alignment, and subjective tests reveal that it produces full-length music with coherent structure.


【7】 Clustering and Data Augmentation to Improve Accuracy of Sleep Assessment  and Sleep Individuality Analysis
标题:集群和数据增强提高睡眠评估和睡眠个性分析的准确性
链接:https://arxiv.org/abs/2404.10299
作者:Shintaro Tamai,Masayuki Numao,Ken-ichi Fukui
摘要:最近,越来越多的健康意识,新的方法允许个人在家里监测睡眠。利用睡眠声音提供了传统方法的优势,如智能手表,是非侵入性的,并且能够检测各种生理活动。该研究旨在构建一个基于机器学习的睡眠评估模型,提供基于证据的评估,例如由于睡眠开始期间频繁运动而导致的睡眠不佳。提取睡眠声音事件,使用VAE导出潜在表示,使用GMM进行聚类,并训练LSTM进行主观睡眠评估,在区分睡眠满意度方面达到了94.8%的高准确率。此外,TimeSHAP揭示了不同个体在有影响力的声音事件类型和时间上的差异。
摘要:Recently, growing health awareness, novel methods allow individuals to monitor sleep at home. Utilizing sleep sounds offers advantages over conventional methods like smartwatches, being non-intrusive, and capable of detecting various physiological activities. This study aims to construct a machine learning-based sleep assessment model providing evidence-based assessments, such as poor sleep due to frequent movement during sleep onset. Extracting sleep sound events, deriving latent representations using VAE, clustering with GMM, and training LSTM for subjective sleep assessment achieved a high accuracy of 94.8% in distinguishing sleep satisfaction. Moreover, TimeSHAP revealed differences in impactful sound event types and timings for different individuals.

【8】 Deferred NAM: Low-latency Top-K Context Injection via DeferredContext  Encoding for Non-Streaming ASR
标题:延迟NAM:通过针对非流媒体ASB的延迟上下文编码进行低延迟Top-K上下文注入
链接:https://arxiv.org/abs/2404.10180
作者:Zelin Wu,Gan Song,Christopher Li,Pat Rondon,Zhong Meng,Xavier Velez,Weiran Wang,Diamantino Caseiro,Golan Pundak,Tsendsuren Munkhdalai,Angad Chandorkar,Rohit Prabhavalkar
备注:None
摘要:上下文偏置使语音识别器能够转录说话者上下文中的重要短语,例如联系人姓名,即使它们在训练数据中很少或不存在。基于注意力的偏置是一种领先的方法,它允许识别器和偏置系统的完全端到端的协同训练,并且不需要单独的推理时间组件。这种偏置器通常由上下文编码器组成;其次是上下文过滤器,它缩小了要应用的上下文,提高了每步推理时间;最后,通过交叉注意应用上下文。虽然在优化每帧性能方面已经做了很多工作,但上下文编码器至少同样重要:识别不能在上下文编码结束之前开始。在这里,我们展示了轻量级短语选择通道可以在上下文编码之前移动,从而实现高达16.1倍的加速,并使偏置能够扩展到20 K短语,最大预解码延迟低于33 ms。通过添加短语和单词级交叉熵损失,我们的技术在没有损失和轻量级短语选择通道的情况下,在基线上还实现了高达37.5%的相对WER降低。
摘要:Contextual biasing enables speech recognizers to transcribe important phrases in the speaker's context, such as contact names, even if they are rare in, or absent from, the training data. Attention-based biasing is a leading approach which allows for full end-to-end cotraining of the recognizer and biasing system and requires no separate inference-time components. Such biasers typically consist of a context encoder; followed by a context filter which narrows down the context to apply, improving per-step inference time; and, finally, context application via cross attention. Though much work has gone into optimizing per-frame performance, the context encoder is at least as important: recognition cannot begin before context encoding ends. Here, we show the lightweight phrase selection pass can be moved before context encoding, resulting in a speedup of up to 16.1 times and enabling biasing to scale to 20K phrases with a maximum pre-decoding delay under 33ms. With the addition of phrase- and wordpiece-level cross-entropy losses, our technique also achieves up to a 37.5% relative WER reduction over the baseline without the losses and lightweight phrase selection pass.

【9】 PRODIS - a speech database and a phoneme-based language model for the  study of predictability effects in Polish
标题:PRODIS -一个语音数据库和基于音素的语言模型,用于研究波兰语的可预测性影响
链接:https://arxiv.org/abs/2404.10112
作者:Zofia Malisz,Jan Foremski,Małgorzata Kul
备注:To appear in the proceedings of LREC2024: Language Resources and Evaluation Conference 2024, Turin, Italy
摘要:我们提出了一个语音数据库和波兰语的音素级语言模型。该数据库和模型的目的是分析韵律和话语因素及其对声学参数的影响与可预测性的影响。该数据库也是第一个大型的、公开可用的波兰语语音语料库,具有出色的声学质量,可用于语音分析和多说话人语音技术系统的训练。数据库中的语音在管道中处理,实现了90%的自动化程度。它包含了最先进的、免费提供的工具,使数据库能够扩展或适应其他语言。
摘要:We present a speech database and a phoneme-level language model of Polish. The database and model are designed for the analysis of prosodic and discourse factors and their impact on acoustic parameters in interaction with predictability effects. The database is also the first large, publicly available Polish speech corpus of excellent acoustic quality that can be used for phonetic analysis and training of multi-speaker speech technology systems. The speech in the database is processed in a pipeline that achieves a 90% degree of automation. It incorporates state-of-the-art, freely available tools enabling database expansion or adaptation to additional languages.


机器翻译由腾讯交互翻译提供,仅供参考