今日论文合集:cs.SD语音4篇,eess.AS音频处理4篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】 UniSync: A Unified Framework for Audio-Visual Synchronization
标题:UniSec:视听同步的统一框架
链接:https://arxiv.org/abs/2503.16357
作者:Tao Feng,  Yifan Xie,  Xun Guan,  Jiyuan Song,  Zhou Liu,  Fei Ma,  Fei Yu
备注:7 pages, 3 figures, accepted by ICME 2025
摘要:语音视频中精确的视听同步对于内容质量和观众理解至关重要。现有方法通过基于规则的方法和端到端学习技术在应对这一挑战方面取得了重大进展。然而,这些方法往往依赖于有限的视听表示和次优的学习策略,可能会限制其在更复杂的情况下的有效性。为了解决这些限制,我们提出了UniSync,一种新的方法,用于评估视听同步使用嵌入相似性。UniSync提供与各种音频表示(例如,Mel频谱图,HuBERT)和视觉表示(例如,RGB图像、面部解析图、面部标志、3DMM),有效地处理它们的显著尺寸差异。我们增强了对比学习框架,基于边缘的损失分量和跨扬声器不同步对,提高了区分能力。UniSync在标准数据集上的性能优于现有方法,并在各种视听表示中展示了多功能性。它集成到说话面部生成框架中,提高了自然内容和人工智能生成内容的同步质量。
摘要:Precise audio-visual synchronization in speech videos is crucial for contentquality and viewer comprehension. Existing methods have made significantstrides in addressing this challenge through rule-based approaches andend-to-end learning techniques. However, these methods often rely on limitedaudio-visual representations and suboptimal learning strategies, potentiallyconstraining their effectiveness in more complex scenarios. To address theselimitations, we present UniSync, a novel approach for evaluating audio-visualsynchronization using embedding similarities. UniSync offers broadcompatibility with various audio representations (e.g., Mel spectrograms,HuBERT) and visual representations (e.g., RGB images, face parsing maps, faciallandmarks, 3DMM), effectively handling their significant dimensionaldifferences. We enhance the contrastive learning framework with a margin-basedloss component and cross-speaker unsynchronized pairs, improving discriminativecapabilities. UniSync outperforms existing methods on standard datasets anddemonstrates versatility across diverse audio-visual representations. Itsintegration into talking face generation frameworks enhances synchronizationquality in both natural and AI-generated content.

【2】 Structured-Noise Masked Modeling for Video, Audio and Beyond
标题:视频、音频及其他领域的结构噪音掩蔽建模
链接:https://arxiv.org/abs/2503.16311
作者:Aritra Bhowmik,  Fida Mohammad Thoker,  Carlos Hinojosa,  Bernard Ghanem,  Cees G. M. Snoek
摘要:掩蔽建模已经成为一个强大的自监督学习框架,但现有的方法在很大程度上依赖于随机掩蔽,忽略了不同模态的结构特性。在这项工作中,我们介绍了结构化的基于噪声的掩蔽,一个简单而有效的方法,自然与视频和音频数据的空间,时间和频谱特性。通过将白噪声过滤成不同的颜色噪声分布,我们生成了结构化的掩模,该掩模可以保留特定于模态的模式,而不需要手工制作的算法或访问数据。我们的方法在没有任何计算开销的情况下提高了掩蔽视频和音频建模框架的性能。大量实验表明,对于标准和高级掩蔽建模方法,结构化噪声掩蔽相对于随机掩蔽实现了一致的改进,凸显了模态感知掩蔽策略对于表示学习的重要性。
摘要:Masked modeling has emerged as a powerful self-supervised learning framework,but existing methods largely rely on random masking, disregarding thestructural properties of different modalities. In this work, we introducestructured noise-based masking, a simple yet effective approach that naturallyaligns with the spatial, temporal, and spectral characteristics of video andaudio data. By filtering white noise into distinct color noise distributions,we generate structured masks that preserve modality-specific patterns withoutrequiring handcrafted heuristics or access to the data. Our approach improvesthe performance of masked video and audio modeling frameworks without anycomputational overhead. Extensive experiments demonstrate that structured noisemasking achieves consistent improvement over random masking for standard andadvanced masked modeling methods, highlighting the importance of modality-awaremasking strategies for representation learning.

【3】 A Bird Song Detector for improving bird identification through Deep  Learning: a case study from Doñana
标题:通过深度学习改进鸟类识别的鸟鸣检测器:Doñana的案例研究
链接:https://arxiv.org/abs/2503.15576
作者:Alba Márquez-Rodríguez,  Miguel Ángel Mohedano-Munoz,  Manuel J. Marín-Jiménez,  Eduardo Santamaría-García,  Giulia Bastianelli,  Pedro Jordano,  Irene Mendoza
备注:20 pages, 13 images, for associated dataset see this https URL , for associated code see this https URL and this https URL
摘要:使用自动录音机的被动声学监测对于生态系统保护至关重要,但会产生大量无人监督的音频数据,这对提取有意义的信息构成了挑战。深度学习技术提供了一个很有前途的解决方案。BirdNET是一种广泛使用的鸟类识别模型,在许多研究系统中取得了成功,但由于其训练数据中的偏差,在某些地区受到限制。鸟类物种检测的一个关键挑战是,许多记录要么缺乏目标物种,要么包含重叠的发声。为了克服这些问题,我们开发了一个多阶段的管道,自动鸟类发声识别在多纳纳国家公园(西南西班牙),一个地区面临着重大的保护威胁。我们的方法包括一个鸟鸣检测器来隔离发声和使用BirdNET嵌入训练的自定义分类器。我们手动注释了来自9个地点的3个栖息地的461分钟音频,为34个类别产生了3,749个注释。光谱图促进了图像处理技术的使用。在分类之前应用Bird Song Detector改进了物种识别,因为所有分类模型在仅分析检测到鸟类的片段时表现更好。具体来说,鸟鸣探测器和微调BirdNET的组合与没有鸟鸣探测器的基线相比。我们的方法证明了有效的集成鸟鸣探测器与微调分类模型的鸟类识别在当地的音景。这些研究结果突出表明,需要适应特定的生态挑战的通用工具,在Do nana证明。鉴于鸟类对环境变化的敏感性,自动检测鸟类物种有助于跟踪这一受威胁生态系统的健康状况,并有助于设计保护措施,以减少生物多样性的损失
摘要:Passive Acoustic Monitoring with automatic recorders is essential forecosystem conservation but generates vast unsupervised audio data, posingchallenges for extracting meaningful information. Deep Learning techniquesoffer a promising solution. BirdNET, a widely used model for birdidentification, has shown success in many study systems but is limited in someregions due to biases in its training data. A key challenge in bird speciesdetection is that many recordings either lack target species or containoverlapping vocalizations. To overcome these problems, we developed amulti-stage pipeline for automatic bird vocalization identification in Do\~nanaNational Park (SW Spain), a region facing significant conservation threats. Ourapproach included a Bird Song Detector to isolate vocalizations and customclassifiers trained with BirdNET embeddings. We manually annotated 461 minutesof audio from three habitats across nine locations, yielding 3,749 annotationsfor 34 classes. Spectrograms facilitated the use of image processingtechniques. Applying the Bird Song Detector before classification improvedspecies identification, as all classification models performed better whenanalyzing only the segments where birds were detected. Specifically, thecombination of the Bird Song Detector and fine-tuned BirdNET compared to thebaseline without the Bird Song Detector. Our approach demonstrated theeffectiveness of integrating a Bird Song Detector with fine-tunedclassification models for bird identification at local soundscapes. Thesefindings highlight the need to adapt general-purpose tools for specificecological challenges, as demonstrated in Do\~nana. Automatically detectingbird species serves for tracking the health status of this threatenedecosystem, given the sensitivity of birds to environmental changes, and helpsin the design of conservation measures for reducing biodiversity loss

【4】 Revival: Collaborative Artistic Creation through Human-AI Interactions  in Musical Creativity
标题:复兴:通过音乐创意中的人机交互进行协同艺术创作
链接:https://arxiv.org/abs/2503.15498
作者:Keon Ju M. Lee,  Philippe Pasquier,  Jun Yuri
备注:Keon Ju M. Lee, Philippe Pasquier and Jun Yuri. 2024. In Proceedings of the Creativity and Generative AI NIPS (Neural Information Processing Systems) Workshop
摘要:Revival是由我们的艺术家集体K-Phi-A进行的创新现场视听表演和音乐即兴创作,融合了人类和人工智能的音乐才能,创造出具有音频反应视觉效果的电子音乐。该表演的特点是一位音乐家、一位电子音乐艺术家和人工智能音乐代理之间的实时共同创作即兴表演。在已故作曲家的作品和集体作品的训练下,这些代理人动态地响应人类输入并模仿复杂的音乐风格。一个由人工智能驱动的视觉合成器,由人类VJ引导,产生随着音乐景观而发展的视觉效果。Revival展示了人工智能和人类合作在即兴艺术创作中的潜力。
摘要:Revival is an innovative live audiovisual performance and music improvisationby our artist collective K-Phi-A, blending human and AI musicianship to createelectronic music with audio-reactive visuals. The performance featuresreal-time co-creative improvisation between a percussionist, an electronicmusic artist, and AI musical agents. Trained in works by deceased composers andthe collective's compositions, these agents dynamically respond to human inputand emulate complex musical styles. An AI-driven visual synthesizer, guided bya human VJ, produces visuals that evolve with the musical landscape. Revivalshowcases the potential of AI and human collaboration in improvisationalartistic creation.

eess.AS音频处理

【1】 A Speech Production Model for Radar: Connecting Speech Acoustics with  Radar-Measured Vibrations
标题:雷达语音产生模型:将语音声学与雷达测量的振动连接起来
链接:https://arxiv.org/abs/2503.15627
作者:Isabella Lenz,  Yu Rong,  Daniel Bliss,  Julie Liss,  Visar Berisha
备注:5 pages, 6 figure, InterSpeech Conference
摘要:毫米波(mmWave)雷达已经成为一种有前途的语音传感模式,提供了传统麦克风的优势。先前的工作已经表明,雷达捕获的运动信号与声音振动,但有一个差距,在雷达测量的振动和声学语音信号之间的分析连接的理解。我们建立了一个数学框架连接雷达捕获的颈部振动语音声学。我们推导出颈部表面位移和语音之间的解析关系。我们使用来自66名人类参与者的数据和统计光谱距离分析来实证评估该模型。我们的研究结果表明,雷达测量的信号更紧密地与我们的模型过滤来自语音的振动信号,而不是与原始语音本身。这些发现为改进基于雷达的语音处理在语音增强、编码、监视和认证中的应用提供了基础。
摘要:Millimeter Wave (mmWave) radar has emerged as a promising modality for speechsensing, offering advantages over traditional microphones. Prior works havedemonstrated that radar captures motion signals related to vocal vibrations,but there is a gap in the understanding of the analytical connection betweenradar-measured vibrations and acoustic speech signals. We establish amathematical framework linking radar-captured neck vibrations to speechacoustics. We derive an analytical relationship between neck surfacedisplacements and speech. We use data from 66 human participants, andstatistical spectral distance analysis to empirically assess the model. Ourresults show that the radar-measured signal aligns more closely with our modelfiltered vibration signal derived from speech than with raw speech itself.These findings provide a foundation for improved radar-based speech processingfor applications in speech enhancement, coding, surveillance, andauthentication.

【2】 UniSync: A Unified Framework for Audio-Visual Synchronization
标题:UniSec:视听同步的统一框架
链接:https://arxiv.org/abs/2503.16357
作者:Tao Feng,  Yifan Xie,  Xun Guan,  Jiyuan Song,  Zhou Liu,  Fei Ma,  Fei Yu
备注:7 pages, 3 figures, accepted by ICME 2025
摘要:语音视频中精确的视听同步对于内容质量和观众理解至关重要。现有方法通过基于规则的方法和端到端学习技术在应对这一挑战方面取得了重大进展。然而,这些方法往往依赖于有限的视听表示和次优的学习策略,可能会限制其在更复杂的情况下的有效性。为了解决这些限制,我们提出了UniSync,一种新的方法,用于评估视听同步使用嵌入相似性。UniSync提供与各种音频表示(例如,Mel频谱图,HuBERT)和视觉表示(例如,RGB图像、面部解析图、面部标志、3DMM),有效地处理它们的显著尺寸差异。我们增强了对比学习框架,基于边缘的损失分量和跨扬声器不同步对,提高了区分能力。UniSync在标准数据集上的性能优于现有方法,并在不同的视听表示中展示了多功能性。它集成到说话面部生成框架中,提高了自然内容和人工智能生成内容的同步质量。
摘要:Precise audio-visual synchronization in speech videos is crucial for contentquality and viewer comprehension. Existing methods have made significantstrides in addressing this challenge through rule-based approaches andend-to-end learning techniques. However, these methods often rely on limitedaudio-visual representations and suboptimal learning strategies, potentiallyconstraining their effectiveness in more complex scenarios. To address theselimitations, we present UniSync, a novel approach for evaluating audio-visualsynchronization using embedding similarities. UniSync offers broadcompatibility with various audio representations (e.g., Mel spectrograms,HuBERT) and visual representations (e.g., RGB images, face parsing maps, faciallandmarks, 3DMM), effectively handling their significant dimensionaldifferences. We enhance the contrastive learning framework with a margin-basedloss component and cross-speaker unsynchronized pairs, improving discriminativecapabilities. UniSync outperforms existing methods on standard datasets anddemonstrates versatility across diverse audio-visual representations. Itsintegration into talking face generation frameworks enhances synchronizationquality in both natural and AI-generated content.

【3】 Development of an Inclusive Educational Platform Using Open Technologies  and Machine Learning: A Case Study on Accessibility Enhancement
标题:使用开放技术和机器学习开发包容性教育平台:无障碍增强案例研究
链接:https://arxiv.org/abs/2503.15501
作者:Jimi Togni
备注:14 pages, 1 figure
摘要:本研究通过提出和开发一个包容性教育平台,解决了为有特殊需要的学生提供教育包容性的紧迫挑战。该平台集成了机器学习、自然语言处理和跨平台接口,具有关键功能,如语音识别功能,支持语音命令和通过语音输入生成文本;使用YOLOv 5模型进行实时对象识别,适用于教育环境;使用带有注意力的seq 2seq模型的文本到语音系统的字形到音素(G2 P)转换,确保自然和流畅的语音合成;以及在Flutter中开发跨平台移动应用程序,并使用TensorFlow Lite在设备上执行推理。结果表明,高准确性,可用性和积极的影响,在教育场景中,验证该提案作为教育包容性的有效工具。该项目强调了开放和无障碍技术在促进包容性和优质教育方面的重要性。
摘要:This study addresses the pressing challenge of educational inclusion forstudents with special needs by proposing and developing an inclusiveeducational platform. Integrating machine learning, natural languageprocessing, and cross-platform interfaces, the platform features keyfunctionalities such as speech recognition functionality to support voicecommands and text generation via voice input; real-time object recognitionusing the YOLOv5 model, adapted for educational environments;Grapheme-to-Phoneme (G2P) conversion for Text-to-Speech systems using seq2seqmodels with attention, ensuring natural and fluent voice synthesis; and thedevelopment of a cross-platform mobile application in Flutter with on-deviceinference execution using TensorFlow Lite. The results demonstrated highaccuracy, usability, and positive impact in educational scenarios, validatingthe proposal as an effective tool for educational inclusion. This projectunderscores the importance of open and accessible technologies in promotinginclusive and quality education.

【4】 Revival: Collaborative Artistic Creation through Human-AI Interactions  in Musical Creativity
标题:复兴:通过音乐创意中的人机交互进行协同艺术创作
链接:https://arxiv.org/abs/2503.15498
作者:Keon Ju M. Lee,  Philippe Pasquier,  Jun Yuri
备注:Keon Ju M. Lee, Philippe Pasquier and Jun Yuri. 2024. In Proceedings of the Creativity and Generative AI NIPS (Neural Information Processing Systems) Workshop
摘要:Revival是由我们的艺术家集体K-Phi-A进行的创新现场视听表演和音乐即兴创作,融合了人类和人工智能的音乐才能,创造出具有音频反应视觉效果的电子音乐。该表演的特点是一位音乐家、一位电子音乐艺术家和人工智能音乐代理之间的实时共同创作即兴表演。在已故作曲家的作品和集体作品的训练下,这些代理人动态地响应人类输入并模仿复杂的音乐风格。一个由人工智能驱动的视觉合成器,由人类VJ引导,产生随着音乐景观而发展的视觉效果。Revival展示了人工智能和人类合作在即兴艺术创作中的潜力。
摘要:Revival is an innovative live audiovisual performance and music improvisationby our artist collective K-Phi-A, blending human and AI musicianship to createelectronic music with audio-reactive visuals. The performance featuresreal-time co-creative improvisation between a percussionist, an electronicmusic artist, and AI musical agents. Trained in works by deceased composers andthe collective's compositions, these agents dynamically respond to human inputand emulate complex musical styles. An AI-driven visual synthesizer, guided bya human VJ, produces visuals that evolve with the musical landscape. Revivalshowcases the potential of AI and human collaboration in improvisationalartistic creation.

机器翻译由腾讯交互翻译提供,仅供参考