本文经arXiv每日学术速递授权转载
【1】 Embodied Exploration of Latent Spaces and Explainable AI
标题: 潜在空间和可解释人工智能的有序探索
作者: Elizabeth Wilson, Mika Satomi, Alex McLean, Deva Schubert, Juan Felipe Amaya Gonzalez
备注:In Proceedings of Explainable AI for the Arts Workshop 2024 (XAIxArts 2024) arXiv:2406.14485
链接:点击下载PDF文件
【2】 SNAC: Multi-Scale Neural Audio Codec
标题: SNAC:多尺度神经音频编解码器
作者: Hubert Siuzdak, Florian Grötschla, Luca A. Lanzendörfer
链接:点击下载PDF文件
【3】 Towards Robust Transcription: Exploring Noise Injection Strategies for Training Data Augmentation
标题: 迈向稳健转录:探索用于训练数据增强的噪音注入策略
作者: Yonghyun Kim, Alexander Lerch
备注:Accepted to the Late-Breaking Demo Session of the 25th International Society for Music Information Retrieval (ISMIR) Conference, 2024
链接:点击下载PDF文件
【4】 Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech
标题: 沉浸式视觉文本到语音的多源空间知识理解
作者: Shuwei He, Rui Liu, Haizhou Li
备注:5 pages, 1 figure
链接:点击下载PDF文件
标题: 收集22种印度语言文本到语音合成数据集的统一框架
作者: Sujitha Sathiyamoorthy (1), N Mohana (1), Anusha Prakash (3), Hema A Murthy (1 and 2) ((1) Dept of Computer Science & Engineering, Indian Institute of Technology Madras, Chennai, India (2) Shiv Nadar University Chennai, India, (3) Independent Researcher Bengaluru, India)
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【2】 Embodied Exploration of Latent Spaces and Explainable AI
标题: 潜在空间和可解释人工智能的有序探索
作者: Elizabeth Wilson, Mika Satomi, Alex McLean, Deva Schubert, Juan Felipe Amaya Gonzalez
备注:In Proceedings of Explainable AI for the Arts Workshop 2024 (XAIxArts 2024) arXiv:2406.14485
链接:点击下载PDF文件
【3】 SNAC: Multi-Scale Neural Audio Codec
标题: SNAC:多尺度神经音频编解码器
作者: Hubert Siuzdak, Florian Grötschla, Luca A. Lanzendörfer
链接:点击下载PDF文件
【4】 Towards Robust Transcription: Exploring Noise Injection Strategies for Training Data Augmentation
标题: 迈向稳健转录:探索用于训练数据增强的噪音注入策略
作者: Yonghyun Kim, Alexander Lerch
备注:Accepted to the Late-Breaking Demo Session of the 25th International Society for Music Information Retrieval (ISMIR) Conference, 2024
链接:点击下载PDF文件
【5】 Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech
标题: 沉浸式视觉文本到语音的多源空间知识理解
作者: Shuwei He, Rui Liu, Haizhou Li
备注:5 pages, 1 figure
链接:点击下载PDF文件
标题: 潜在空间和可解释人工智能的有序探索
作者: Elizabeth Wilson, Mika Satomi, Alex McLean, Deva Schubert, Juan Felipe Amaya Gonzalez
备注:In Proceedings of Explainable AI for the Arts Workshop 2024 (XAIxArts 2024) arXiv:2406.14485
链接:点击下载PDF文件
摘要:在本文中,我们探讨了如何执行者的体现与神经音频合成模型的相互作用,允许这样一个模型的潜在空间的探索,通过电子纺织品感测到的运动介导。我们提供了性能的背景和上下文,突出了体现实践的潜力,有助于开发可解释的人工智能系统。通过将各种艺术领域与可解释的人工智能原则相结合,我们的跨学科探索有助于对艺术,体现和人工智能的论述,为通过身体表达发现的直观方法提供见解。摘要:In this paper, we explore how performers' embodied interactions with a Neural Audio Synthesis model allow the exploration of the latent space of such a model, mediated through movements sensed by e-textiles. We provide background and context for the performance, highlighting the potential of embodied practices to contribute to developing explainable AI systems. By integrating various artistic domains with explainable AI principles, our interdisciplinary exploration contributes to the discourse on art, embodiment, and AI, offering insights into intuitive approaches found through bodily expression.
【2】 SNAC: Multi-Scale Neural Audio Codec
标题: SNAC:多尺度神经音频编解码器
作者: Hubert Siuzdak, Florian Grötschla, Luca A. Lanzendörfer
链接:点击下载PDF文件
摘要:神经音频编解码器最近越来越受欢迎,因为它们可以在非常低的比特率下以高保真度表示音频信号,使得使用语言建模方法进行音频生成和理解变得可行。残差矢量量化(RVQ)已经成为使用VQ码本级联的神经音频压缩的标准技术。本文提出了多尺度神经音频编解码器,一个简单的扩展RVQ的量化器可以在不同的时间分辨率。通过以可变帧速率应用量化器的层次结构,编解码器适应跨多个时间尺度的音频结构。这导致更有效的压缩,如广泛的客观和主观评价所证明的。代码和模型权重在https: github.com hubertsiuzdak snac上开源。摘要:Neural audio codecs have recently gained popularity because they can represent audio signals with high fidelity at very low bitrates, making it feasible to use language modeling approaches for audio generation and understanding. Residual Vector Quantization (RVQ) has become the standard technique for neural audio compression using a cascade of VQ codebooks. This paper proposes the Multi-Scale Neural Audio Codec, a simple extension of RVQ where the quantizers can operate at different temporal resolutions. By applying a hierarchy of quantizers at variable frame rates, the codec adapts to the audio structure across multiple timescales. This leads to more efficient compression, as demonstrated by extensive objective and subjective evaluations. The code and model weights are open-sourced at https: github.com hubertsiuzdak snac.
【3】 Towards Robust Transcription: Exploring Noise Injection Strategies for Training Data Augmentation
标题: 迈向稳健转录:探索用于训练数据增强的噪音注入策略
作者: Yonghyun Kim, Alexander Lerch
备注:Accepted to the Late-Breaking Demo Session of the 25th International Society for Music Information Retrieval (ISMIR) Conference, 2024
链接:点击下载PDF文件
摘要:自动钢琴转录(APT)的最新进展显着提高了系统的性能,但噪声环境对系统性能的影响仍然在很大程度上未被探索。本研究调查了各种信噪比(SNR)水平下的白噪声对最先进的APT模型的影响,并评估了在噪声增强数据上训练时发作和帧模型的性能。我们希望这项研究提供有价值的见解,作为开发转录模型的初步工作,这些模型在一系列声学条件下保持一致的性能。摘要:Recent advancements in Automatic Piano Transcription (APT) have significantly improved system performance, but the impact of noisy environments on the system performance remains largely unexplored. This study investigates the impact of white noise at various Signal-to-Noise Ratio (SNR) levels on state-of-the-art APT models and evaluates the performance of the Onsets and Frames model when trained on noise-augmented data. We hope this research provides valuable insights as preliminary work toward developing transcription models that maintain consistent performance across a range of acoustic conditions.
【4】 Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech
标题: 沉浸式视觉文本到语音的多源空间知识理解
作者: Shuwei He, Rui Liu, Haizhou Li
备注:5 pages, 1 figure
链接:点击下载PDF文件
摘要:可视文语转换(VTTS)技术是以空间环境图像为提示,对语音内容合成混响语音。以前的研究集中在RGB模式的全球环境建模,忽视了多源空间知识的潜力,如深度,扬声器位置和环境语义。为了解决这些问题,我们提出了一种新的多源空间知识理解方案沉浸式VTTS,称为MS$^2$KU-VTTS。具体来说,我们首先优先考虑RGB图像作为主要来源,并考虑深度图像,来自对象检测的扬声器位置知识以及来自图像理解LLM的语义字幕作为补充来源。之后,我们提出了一个系列的互动机制,深入参与主导和补充来源。实验结果表明,MS$^2$KU-VTTS在生成沉浸式空间语音方面优于现有的基准。演示和代码可在:https: github.com MS2KU-VTTS MS2KU-VTTS。摘要:Visual Text-to-Speech (VTTS) aims to take the spatial environmental image as the prompt to synthesize the reverberation speech for the spoken content. Previous research focused on the RGB modality for global environmental modeling, overlooking the potential of multi-source spatial knowledge like depth, speaker position, and environmental semantics. To address the issues, we propose a novel multi-source spatial knowledge understanding scheme for immersive VTTS, termed MS$^2$KU-VTTS. Specifically, we first prioritize RGB image as the dominant source and consider depth image, speaker position knowledge from object detection, and semantic captions from image understanding LLM as supplementary sources. Afterwards, we propose a serial interaction mechanism to deeply engage with both dominant and supplementary sources. The resulting multi-source knowledge is dynamically integrated based on their contributions.This enriched interaction and integration of multi-source spatial knowledge guides the speech generation model, enhancing the immersive spatial speech experience.Experimental results demonstrate that the MS$^2$KU-VTTS surpasses existing baselines in generating immersive speech. Demos and code are available at: https: github.com MS2KU-VTTS MS2KU-VTTS.
eess.AS音频处理
【1】 A Unified Framework for Collecting Text-to-Speech Synthesis Datasets for 22 Indian Languages标题: 收集22种印度语言文本到语音合成数据集的统一框架
作者: Sujitha Sathiyamoorthy (1), N Mohana (1), Anusha Prakash (3), Hema A Murthy (1 and 2) ((1) Dept of Computer Science & Engineering, Indian Institute of Technology Madras, Chennai, India (2) Shiv Nadar University Chennai, India, (3) Independent Researcher Bengaluru, India)
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:文语转换(TTS)合成模型的性能取决于多种因素,其中训练数据的质量至关重要。全球各地收集了数百万种语言的数据,但印度语言的资源很少。虽然有许多工作涉及数据收集,一套共同的协议,数据收集成为必要的建设TTS系统在印度语言,主要是因为需要一个统一的发展TTS系统跨语言。在本文中,我们提出了我们的学习数据收集工作的印度语言超过15年。这些数据库已被用于单元选择合成,隐马尔可夫模型的基础上,和端到端的框架,并产生韵律丰富的TTS系统。所收集的数据的最显著特征是数据纯度使得能够使用与欧洲 中国语言相比相对较小的数据集来构建高质量的TTS系统。摘要:The performance of a text-to-speech (TTS) synthesis model depends on various factors, of which the quality of the training data is of utmost importance. Millions of data are collected around the globe for various languages, but resources for Indian languages are few. Although there are many efforts involved in data collection, a common set of protocols for data collection becomes necessary for building TTS systems in Indian languages primarily because of the need for a uniform development of TTS systems across languages. In this paper, we present our learnings on data collection efforts' for Indic languages over 15 years. These databases have been used in unit selection synthesis, hidden Markov model based, and end-to-end frameworks, and for generating prosodically rich TTS systems. The most significant feature of the data collected is that data purity enables building high-quality TTS systems with a comparatively small dataset compared to that of European Chinese languages.
【2】 Embodied Exploration of Latent Spaces and Explainable AI
标题: 潜在空间和可解释人工智能的有序探索
作者: Elizabeth Wilson, Mika Satomi, Alex McLean, Deva Schubert, Juan Felipe Amaya Gonzalez
备注:In Proceedings of Explainable AI for the Arts Workshop 2024 (XAIxArts 2024) arXiv:2406.14485
链接:点击下载PDF文件
摘要:在本文中,我们探讨了如何执行者的体现与神经音频合成模型的相互作用,允许这样一个模型的潜在空间的探索,通过电子纺织品感测到的运动介导。我们提供了性能的背景和上下文,突出了体现实践的潜力,有助于开发可解释的人工智能系统。通过将各种艺术领域与可解释的人工智能原则相结合,我们的跨学科探索有助于对艺术,体现和人工智能的论述,为通过身体表达发现的直观方法提供见解。摘要:In this paper, we explore how performers' embodied interactions with a Neural Audio Synthesis model allow the exploration of the latent space of such a model, mediated through movements sensed by e-textiles. We provide background and context for the performance, highlighting the potential of embodied practices to contribute to developing explainable AI systems. By integrating various artistic domains with explainable AI principles, our interdisciplinary exploration contributes to the discourse on art, embodiment, and AI, offering insights into intuitive approaches found through bodily expression.
【3】 SNAC: Multi-Scale Neural Audio Codec
标题: SNAC:多尺度神经音频编解码器
作者: Hubert Siuzdak, Florian Grötschla, Luca A. Lanzendörfer
链接:点击下载PDF文件
摘要:神经音频编解码器最近越来越受欢迎,因为它们可以在非常低的比特率下以高保真度表示音频信号,使得使用语言建模方法进行音频生成和理解变得可行。残差矢量量化(RVQ)已经成为使用VQ码本级联的神经音频压缩的标准技术。本文提出了多尺度神经音频编解码器,一个简单的扩展RVQ的量化器可以在不同的时间分辨率。通过以可变帧速率应用量化器的层次结构,编解码器适应跨多个时间尺度的音频结构。这导致更有效的压缩,如广泛的客观和主观评价所证明的。代码和模型权重在https: github.com hubertsiuzdak snac上开源。摘要:Neural audio codecs have recently gained popularity because they can represent audio signals with high fidelity at very low bitrates, making it feasible to use language modeling approaches for audio generation and understanding. Residual Vector Quantization (RVQ) has become the standard technique for neural audio compression using a cascade of VQ codebooks. This paper proposes the Multi-Scale Neural Audio Codec, a simple extension of RVQ where the quantizers can operate at different temporal resolutions. By applying a hierarchy of quantizers at variable frame rates, the codec adapts to the audio structure across multiple timescales. This leads to more efficient compression, as demonstrated by extensive objective and subjective evaluations. The code and model weights are open-sourced at https: github.com hubertsiuzdak snac.
【4】 Towards Robust Transcription: Exploring Noise Injection Strategies for Training Data Augmentation
标题: 迈向稳健转录:探索用于训练数据增强的噪音注入策略
作者: Yonghyun Kim, Alexander Lerch
备注:Accepted to the Late-Breaking Demo Session of the 25th International Society for Music Information Retrieval (ISMIR) Conference, 2024
链接:点击下载PDF文件
摘要:自动钢琴转录(APT)的最新进展显着提高了系统性能,但噪音环境对系统性能的影响在很大程度上仍未得到探讨。本研究调查了各种信噪比(SNR)水平下的白噪声对最先进的APT模型的影响,并评估了在噪声增强数据上训练时发作和帧模型的性能。我们希望这项研究提供有价值的见解,作为开发转录模型的初步工作,这些模型在一系列声学条件下保持一致的性能。摘要:Recent advancements in Automatic Piano Transcription (APT) have significantly improved system performance, but the impact of noisy environments on the system performance remains largely unexplored. This study investigates the impact of white noise at various Signal-to-Noise Ratio (SNR) levels on state-of-the-art APT models and evaluates the performance of the Onsets and Frames model when trained on noise-augmented data. We hope this research provides valuable insights as preliminary work toward developing transcription models that maintain consistent performance across a range of acoustic conditions.
【5】 Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech
标题: 沉浸式视觉文本到语音的多源空间知识理解
作者: Shuwei He, Rui Liu, Haizhou Li
备注:5 pages, 1 figure
链接:点击下载PDF文件
摘要:可视文语转换(VTTS)技术是以空间环境图像为提示,对语音内容合成混响语音。以前的研究集中在RGB模式的全球环境建模,忽视了多源空间知识的潜力,如深度,扬声器位置和环境语义。为了解决这些问题,我们提出了一种新的多源空间知识理解方案沉浸式VTTS,称为MS$^2$KU-VTTS。具体来说,我们首先优先考虑RGB图像作为主要来源,并考虑深度图像,来自对象检测的扬声器位置知识以及来自图像理解LLM的语义字幕作为补充来源。之后,我们提出了一个系列的互动机制,深入参与主导和补充来源。实验结果表明,MS$^2$KU-VTTS在生成沉浸式空间语音方面优于现有的基准。演示和代码可在:https: github.com MS2KU-VTTS MS2KU-VTTS。摘要:Visual Text-to-Speech (VTTS) aims to take the spatial environmental image as the prompt to synthesize the reverberation speech for the spoken content. Previous research focused on the RGB modality for global environmental modeling, overlooking the potential of multi-source spatial knowledge like depth, speaker position, and environmental semantics. To address the issues, we propose a novel multi-source spatial knowledge understanding scheme for immersive VTTS, termed MS$^2$KU-VTTS. Specifically, we first prioritize RGB image as the dominant source and consider depth image, speaker position knowledge from object detection, and semantic captions from image understanding LLM as supplementary sources. Afterwards, we propose a serial interaction mechanism to deeply engage with both dominant and supplementary sources. The resulting multi-source knowledge is dynamically integrated based on their contributions.This enriched interaction and integration of multi-source spatial knowledge guides the speech generation model, enhancing the immersive spatial speech experience.Experimental results demonstrate that the MS$^2$KU-VTTS surpasses existing baselines in generating immersive speech. Demos and code are available at: https: github.com MS2KU-VTTS MS2KU-VTTS.
机器翻译,仅供参考
