本文经arXiv每日学术速递授权转载
【1】 Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond
标题: Speech-MASSIVE:适用于SL及其他领域的多语言语音数据集
作者:Beomseok Lee,Ioan Calapodescu,Marco Gaido,Matteo Negri,Laurent Besacier
备注:Accepted at INTERSPEECH 2024. This version includes the same content but with additional appendices
链接:点击下载PDF文件
【2】 Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
标题: 中央库尔德语文本到语音合成,采用新颖的端到端Transformer训练
作者:Hawraz A. Ahmad,Tarik A. Rashid
备注:19 pages
链接:点击下载PDF文件
【3】 Speech privacy-preserving methods using secret key for convolutional neural network models and their robustness evaluation
标题: 卷积神经网络模型使用密钥的语音隐私保护方法及其鲁棒性评估
作者:Shoko Niwa,Sayaka Shiota,Hitoshi Kiya
链接:点击下载PDF文件
【4】 Feasibility of iMagLS-BSM -- ILD Informed Binaural Signal Matching with Arbitrary Microphone Arrays
标题: iMagLS-BSM -- ILD知情双耳信号与任意麦克风阵列匹配的可行性
作者:Or Berebi,Zamir Ben-Hur,David Lou Alon,Boaz Rafaely
备注:Paper accepted for publication in IWAENC 2024, 4 pages, 2 figures
链接:点击下载PDF文件
【5】 Facing the Music: Tackling Singing Voice Separation in Cinematic Audio Source Separation
标题: 面对音乐:解决电影音频源分离中的歌声分离问题
作者:Karn N. Watcharasupat,Chih-Wei Wu,Iroro Orife
备注:Submitted to the Late-Breaking Demo Session of the 25th International Society for Music Information Retrieval (ISMIR) Conference, 2024
链接:点击下载PDF文件
【6】 TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement
标题: TF-Locoformer:通过卷积进行局部建模的Transformer,用于语音分离和增强
作者:Kohei Saijo,Gordon Wichern,François G. Germain,Zexu Pan,Jonathan Le Roux
备注:Accepted to IWAENC 2024
链接:点击下载PDF文件
【7】 Enhanced Reverberation as Supervision for Unsupervised Speech Separation
标题: 增强的回响作为无监督语音分离的监督
作者:Kohei Saijo,Gordon Wichern,François G. Germain,Zexu Pan,Jonathan Le Roux
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
标题: 卷积神经网络模型使用密钥的语音隐私保护方法及其鲁棒性评估
作者:Shoko Niwa,Sayaka Shiota,Hitoshi Kiya
链接:点击下载PDF文件
【2】 One-Shot Distributed Node-Specific Signal Estimation with Non-Overlapping Latent Subspaces in Acoustic Sensor Networks
标题: 声传感器网络中具有非重叠潜在子空间的单次分布式节点特定信号估计
作者:Paul Didier,Pourya Behmandpoor,Toon van Waterschoot,Marc Moonen
链接:点击下载PDF文件
【3】 Feasibility of iMagLS-BSM -- ILD Informed Binaural Signal Matching with Arbitrary Microphone Arrays
标题: iMagLS-BSM -- ILD知情双耳信号与任意麦克风阵列匹配的可行性
作者:Or Berebi,Zamir Ben-Hur,David Lou Alon,Boaz Rafaely
备注:Paper accepted for publication in IWAENC 2024, 4 pages, 2 figures
链接:点击下载PDF文件
【4】 Bridging the Gap between Audio and Text using Parallel-attention for User-defined Keyword Spotting
标题: 使用用户定义的关键词定位的纵向注意力弥合音频和文本之间的差距
作者:Youkyum Kim,Jaemin Jung,Jihwan Park,Byeong-Yeol Kim,Joon Son Chung
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
【5】 Facing the Music: Tackling Singing Voice Separation in Cinematic Audio Source Separation
标题: 面对音乐:解决电影音频源分离中的歌声分离问题
作者:Karn N. Watcharasupat,Chih-Wei Wu,Iroro Orife
备注:Submitted to the Late-Breaking Demo Session of the 25th International Society for Music Information Retrieval (ISMIR) Conference, 2024
链接:点击下载PDF文件
【6】 Design and Analysis of Binaural Signal Matching with Arbitrary Microphone Arrays
标题: 任意麦克风阵列的双耳信号匹配设计与分析
作者:Lior Madmoni,Zamir Ben-Hur,Jacob Donley,Vladimir Tourbabin,Boaz Rafaely
备注:Submitted to EURASIP Journal on audio speech and music processing
链接:点击下载PDF文件
【7】 TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement
标题: TF-Locoformer:通过卷积进行局部建模的Transformer,用于语音分离和增强
作者:Kohei Saijo,Gordon Wichern,François G. Germain,Zexu Pan,Jonathan Le Roux
备注:Accepted to IWAENC 2024
链接:点击下载PDF文件
【8】 Enhanced Reverberation as Supervision for Unsupervised Speech Separation
标题: 增强的回响作为无监督语音分离的监督
作者:Kohei Saijo,Gordon Wichern,François G. Germain,Zexu Pan,Jonathan Le Roux
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【9】 Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond
标题: Speech-MASSIVE:适用于SL及其他领域的多语言语音数据集
作者:Beomseok Lee,Ioan Calapodescu,Marco Gaido,Matteo Negri,Laurent Besacier
备注:Accepted at INTERSPEECH 2024. This version includes the same content but with additional appendices
链接:点击下载PDF文件
【10】 Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
标题: 中央库尔德语文本到语音合成,采用新颖的端到端Transformer训练
作者:Hawraz A. Ahmad,Tarik A. Rashid
备注:19 pages
链接:点击下载PDF文件
标题: Speech-MASSIVE:适用于SL及其他领域的多语言语音数据集
作者:Beomseok Lee,Ioan Calapodescu,Marco Gaido,Matteo Negri,Laurent Besacier
备注:Accepted at INTERSPEECH 2024. This version includes the same content but with additional appendices
链接:点击下载PDF文件
摘要:我们提出了Speech-MASSIVE,一个多语言口语理解(SLU)数据集,包括一部分MASSIVE文本语料库的语音对应物。Speech-MASSIVE涵盖了来自不同家族的12种语言,并从MASSIVE继承了用于意图预测和插槽填充任务的注释。我们的扩展是由大量多语言SLU数据集的稀缺性和对多功能语音数据集的日益增长的需求,以评估跨语言和任务的基础模型(LLM,语音编码器)。我们在各种训练场景(zero-shot、Few-Shot和完全微调)中使用级联和端到端架构提供多模式、多任务、多语言数据集和报告SLU基线。此外,我们证明了适合的Speech-MASSIVE的基准其他任务,如语音转录,语言识别和语音翻译。数据集、模型和代码可在https: github.com hlt-mt Speech-MASSIVE上公开获取摘要:We present Speech-MASSIVE, a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages from different families and inherits from MASSIVE the annotations for the intent prediction and slot-filling tasks. Our extension is prompted by the scarcity of massively multilingual SLU datasets and the growing need for versatile speech datasets to assess foundation models (LLMs, speech encoders) across languages and tasks. We provide a multimodal, multitask, multilingual dataset and report SLU baselines using both cascaded and end-to-end architectures in various training scenarios (zero-shot, few-shot, and full fine-tune). Furthermore, we demonstrate the suitability of Speech-MASSIVE for benchmarking other tasks such as speech transcription, language identification, and speech translation. The dataset, models, and code are publicly available at: https: github.com hlt-mt Speech-MASSIVE
【2】 Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
标题: 中央库尔德语文本到语音合成,采用新颖的端到端Transformer训练
作者:Hawraz A. Ahmad,Tarik A. Rashid
备注:19 pages
链接:点击下载PDF文件
摘要:文本到语音(TTS)模型的最新进展旨在将两阶段过程简化为单阶段训练方法。然而,许多单级模型在音频质量方面仍然落后,特别是在处理库尔德语文本和语音时。迫切需要加强库尔德语的文本到语音转换,特别是索拉尼方言,该方言相对被忽视,在最近的文本到语音转换进展中代表性不足。这项研究介绍了一个端到端的TTS模型,有效地产生高品质的库尔德音频。所提出的方法利用变分自动编码器(VAE),该变分自动编码器(VAE)被预先训练用于音频波形重构,并且通过对抗训练来增强。这涉及将预训练编码器建立的先验分布与潜在变量内文本编码器的后验分布进行对齐。此外,一个随机的持续时间预测器被纳入灌输不同节奏的合成库尔德语语音。通过调整潜在分布和整合随机持续时间预测器,所提出的方法有利于实时生成自然库尔德语音音频,提供灵活的音高和节奏。通过自定义数据集上的平均意见得分(MOS)进行的实证评估证实了我们的方法(MOS为3.94)与通过主观人类评估评估的单阶段系统和其他两阶段系统相比具有更好的性能。摘要:Recent advancements in text-to-speech (TTS) models have aimed to streamline the two-stage process into a single-stage training approach. However, many single-stage models still lag behind in audio quality, particularly when handling Kurdish text and speech. There is a critical need to enhance text-to-speech conversion for the Kurdish language, particularly for the Sorani dialect, which has been relatively neglected and is underrepresented in recent text-to-speech advancements. This study introduces an end-to-end TTS model for efficiently generating high-quality Kurdish audio. The proposed method leverages a variational autoencoder (VAE) that is pre-trained for audio waveform reconstruction and is augmented by adversarial training. This involves aligning the prior distribution established by the pre-trained encoder with the posterior distribution of the text encoder within latent variables. Additionally, a stochastic duration predictor is incorporated to imbue synthesized Kurdish speech with diverse rhythms. By aligning latent distributions and integrating the stochastic duration predictor, the proposed method facilitates the real-time generation of natural Kurdish speech audio, offering flexibility in pitches and rhythms. Empirical evaluation via the mean opinion score (MOS) on a custom dataset confirms the superior performance of our approach (MOS of 3.94) compared with that of a one-stage system and other two-staged systems as assessed through a subjective human evaluation.
【3】 Speech privacy-preserving methods using secret key for convolutional neural network models and their robustness evaluation
标题: 卷积神经网络模型使用密钥的语音隐私保护方法及其鲁棒性评估
作者:Shoko Niwa,Sayaka Shiota,Hitoshi Kiya
链接:点击下载PDF文件
摘要:在本文中,我们提出了一个秘密的密钥的卷积神经网络(CNN)为基础的模型在语音处理任务的隐私保护方法。在不受信任的第三方(如云服务器)提供基于CNN的系统的环境中,确保语音查询的隐私至关重要。本文提出了加密方法的语音查询使用密钥和模型结构,允许加密查询被接受而不解密。我们的方法引入了三种类型的密钥:洗牌,翻转和随机正交矩阵(ROM)。在实验中,我们证明了当所提出的方法与正确的密钥一起使用时,识别性能不会降低。相反,当使用不正确的密钥时,性能显著下降。特别是,使用ROM,我们表明,即使有一个相对较小的密钥空间,高隐私保护性能可以保持许多语音处理任务。此外,我们还证明了在各种鲁棒性评估中从加密查询中恢复原始语音的困难。摘要:In this paper, we propose privacy-preserving methods with a secret key for convolutional neural network (CNN)-based models in speech processing tasks. In environments where untrusted third parties, like cloud servers, provide CNN-based systems, ensuring the privacy of speech queries becomes essential. This paper proposes encryption methods for speech queries using secret keys and a model structure that allows for encrypted queries to be accepted without decryption. Our approach introduces three types of secret keys: Shuffling, Flipping, and random orthogonal matrix (ROM). In experiments, we demonstrate that when the proposed methods are used with the correct key, identification performance did not degrade. Conversely, when an incorrect key is used, the performance significantly decreased. Particularly, with the use of ROM, we show that even with a relatively small key space, high privacy-preserving performance can be maintained many speech processing tasks. Furthermore, we also demonstrate the difficulty of recovering original speech from encrypted queries in various robustness evaluations.
【4】 Feasibility of iMagLS-BSM -- ILD Informed Binaural Signal Matching with Arbitrary Microphone Arrays
标题: iMagLS-BSM -- ILD知情双耳信号与任意麦克风阵列匹配的可行性
作者:Or Berebi,Zamir Ben-Hur,David Lou Alon,Boaz Rafaely
备注:Paper accepted for publication in IWAENC 2024, 4 pages, 2 figures
链接:点击下载PDF文件
摘要:用于以耳机为中心的收听的双耳再现已成为正在进行的研究的焦点,特别是在诸如增强现实和虚拟现实(AR和VR)等先进技术的领域内。这些应用对高质量空间音频的需求对于保持无缝沉浸感至关重要。然而,由于设计约束,仅配备有有限数量的麦克风和不规则的麦克风放置的可穿戴记录设备带来了挑战。与高阶麦克风阵列捕获的参考信号相比,这些因素导致有限的再现质量。本文介绍了一种新的优化损失量身定制的波束成形为基础的,信号无关的双耳再现方案。这种方法,命名为iMagLS-BSM将一个耳间电平差(ILD)的误差项到先前提出的双耳信号匹配(BSM)幅度最小二乘(MagLS)渲染损失的横向平面角度。该方法利用非线性规划来最小化引入的损失。初步结果表明,ILD误差大幅减少,同时保持双耳幅度误差与MagLS BSM解决方案实现的。这些发现为增强所得到的双耳信号的整体空间质量提供了希望。摘要:Binaural reproduction for headphone-centric listening has become a focal point in ongoing research, particularly within the realm of advancing technologies such as augmented and virtual reality (AR and VR). The demand for high-quality spatial audio in these applications is essential to uphold a seamless sense of immersion. However, challenges arise from wearable recording devices equipped with only a limited number of microphones and irregular microphone placements due to design constraints. These factors contribute to limited reproduction quality compared to reference signals captured by high-order microphone arrays. This paper introduces a novel optimization loss tailored for a beamforming-based, signal-independent binaural reproduction scheme. This method, named iMagLS-BSM incorporates an interaural level difference (ILD) error term into the previously proposed binaural signal matching (BSM) magnitude least squares (MagLS) rendering loss for lateral plane angles. The method leverages nonlinear programming to minimize the introduced loss. Preliminary results show a substantial reduction in ILD error, while maintaining a binaural magnitude error comparable to that achieved with a MagLS BSM solution. These findings hold promise for enhancing the overall spatial quality of resultant binaural signals.
【5】 Facing the Music: Tackling Singing Voice Separation in Cinematic Audio Source Separation
标题: 面对音乐:解决电影音频源分离中的歌声分离问题
作者:Karn N. Watcharasupat,Chih-Wei Wu,Iroro Orife
备注:Submitted to the Late-Breaking Demo Session of the 25th International Society for Music Information Retrieval (ISMIR) Conference, 2024
链接:点击下载PDF文件
摘要:电影音频源分离(CASS)是音频源分离的一个较新的子任务。CASS的一个典型设置是一个三干问题,目的是将混合物分离为对话干(DX),音乐干(MX)和效果干(FX)。然而,在实践中,存在几种边缘情况,因为一些声源不完全适合这三个柄中的任何一个,需要在生产中使用额外的辅助柄。一个非常常见的边缘情况是电影音频中的歌声,它可能属于DX或MX,这在很大程度上取决于电影背景。在这项工作中,我们展示了一个非常简单的扩展专用的解码器Bandit和基于查询的单解码器宴会模型的四个干的问题,治疗非音乐对话,器乐,歌声和效果作为单独的干。有趣的是,基于查询的Banquet模型优于专用解码器Bandit模型。我们假设这是由于在瓶颈处的更好的特征对齐,如由与频带无关的薄膜层所实施的。数据集和模型实施将在https: github.com kwatcharasupat source-separation-landing上提供。摘要:Cinematic audio source separation (CASS) is a fairly new subtask of audio source separation. A typical setup of CASS is a three-stem problem, with the aim of separating the mixture into the dialogue stem (DX), music stem (MX), and effects stem (FX). In practice, however, several edge cases exist as some sound sources do not fit neatly in either of these three stems, necessitating the use of additional auxiliary stems in production. One very common edge case is the singing voice in film audio, which may belong in either the DX or MX, depending heavily on the cinematic context. In this work, we demonstrate a very straightforward extension of the dedicated-decoder Bandit and query-based single-decoder Banquet models to a four-stem problem, treating non-musical dialogue, instrumental music, singing voice, and effects as separate stems. Interestingly, the query-based Banquet model outperformed the dedicated-decoder Bandit model. We hypothesized that this is due to a better feature alignment at the bottleneck as enforced by the band-agnostic FiLM layer. Dataset and model implementation will be made available at https: github.com kwatcharasupat source-separation-landing.
【6】 TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement
标题: TF-Locoformer:通过卷积进行局部建模的Transformer,用于语音分离和增强
作者:Kohei Saijo,Gordon Wichern,François G. Germain,Zexu Pan,Jonathan Le Roux
备注:Accepted to IWAENC 2024
链接:点击下载PDF文件
摘要:时频(TF)域双径模型实现高保真语音分离。虽然以前的一些最先进的(SoTA)模型依赖于RNN,但这种依赖意味着它们缺乏Transformer块的并行性、可扩展性和多功能性。鉴于纯基于transformer的架构在其他领域取得了广泛的成功,在这项工作中,我们专注于从TF域双路径模型中删除RNN,同时保持SoTA性能。本文提出了TF-Locoformer,一种基于变换器的模型,通过卷积进行局部建模。该模型使用具有卷积层的前馈网络(FFN),而不是线性层,来捕获局部信息,让自我注意力集中在捕获全局模式上。我们将两个这样的FFN之前和之后的自我注意,以提高本地建模能力。我们还介绍了一种新的TF域双路径模型的归一化。分离和增强数据集上的实验表明,该模型在多个基准测试中达到或超过SoTA,具有RNN自由架构。摘要:Time-frequency (TF) domain dual-path models achieve high-fidelity speech separation. While some previous state-of-the-art (SoTA) models rely on RNNs, this reliance means they lack the parallelizability, scalability, and versatility of Transformer blocks. Given the wide-ranging success of pure Transformer-based architectures in other fields, in this work we focus on removing the RNN from TF-domain dual-path models, while maintaining SoTA performance. This work presents TF-Locoformer, a Transformer-based model with LOcal-modeling by COnvolution. The model uses feed-forward networks (FFNs) with convolution layers, instead of linear layers, to capture local information, letting the self-attention focus on capturing global patterns. We place two such FFNs before and after self-attention to enhance the local-modeling capability. We also introduce a novel normalization for TF-domain dual-path models. Experiments on separation and enhancement datasets show that the proposed model meets or exceeds SoTA in multiple benchmarks with an RNN-free architecture.
【7】 Enhanced Reverberation as Supervision for Unsupervised Speech Separation
标题: 增强的回响作为无监督语音分离的监督
作者:Kohei Saijo,Gordon Wichern,François G. Germain,Zexu Pan,Jonathan Le Roux
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:混响监督(RAS)是一个框架,它允许以无监督的方式从多声道混合中训练单声道语音分离模型。在RAS中,训练模型,使得可以映射从输入通道处的混合预测的源,以在目标通道处重建混合。然而,到目前为止,稳定的无监督训练仅在超定源信道条件下实现,而关键的确定情况尚未解决。这项工作提出了增强RAS(ERAS)解决这个问题。通过定性分析,我们发现稳定的训练可以通过利用损失项来缓解频率置换问题。分离性能还通过添加新的损失项来提高,其中将映射回其自身输入混合的分离信号用作从其他通道分离并映射到相同通道的信号的伪目标。实验结果表明,ERAS具有较高的稳定性和性能。摘要:Reverberation as supervision (RAS) is a framework that allows for training monaural speech separation models from multi-channel mixtures in an unsupervised manner. In RAS, models are trained so that sources predicted from a mixture at an input channel can be mapped to reconstruct a mixture at a target channel. However, stable unsupervised training has so far only been achieved in over-determined source-channel conditions, leaving the key determined case unsolved. This work proposes enhanced RAS (ERAS) for solving this problem. Through qualitative analysis, we found that stable training can be achieved by leveraging the loss term to alleviate the frequency-permutation problem. Separation performance is also boosted by adding a novel loss term where separated signals mapped back to their own input mixture are used as pseudo-targets for the signals separated from other channels and mapped to the same channel. Experimental results demonstrate high stability and performance of ERAS.
eess.AS音频处理
【1】 Speech privacy-preserving methods using secret key for convolutional neural network models and their robustness evaluation标题: 卷积神经网络模型使用密钥的语音隐私保护方法及其鲁棒性评估
作者:Shoko Niwa,Sayaka Shiota,Hitoshi Kiya
链接:点击下载PDF文件
摘要:在本文中,我们提出了一个秘密的密钥的卷积神经网络(CNN)为基础的模型在语音处理任务的隐私保护方法。在不受信任的第三方(如云服务器)提供基于CNN的系统的环境中,确保语音查询的隐私至关重要。本文提出了使用密钥的语音查询加密方法和模型结构,允许加密查询被接受而不解密。我们的方法引入了三种类型的密钥:洗牌,翻转和随机正交矩阵(ROM)。在实验中,我们证明了当所提出的方法与正确的密钥一起使用时,识别性能不会降低。相反,当使用不正确的密钥时,性能显著下降。特别是,使用ROM,我们表明,即使有一个相对较小的密钥空间,高隐私保护性能可以保持许多语音处理任务。此外,我们还证明了在各种鲁棒性评估中从加密查询中恢复原始语音的困难。摘要:In this paper, we propose privacy-preserving methods with a secret key for convolutional neural network (CNN)-based models in speech processing tasks. In environments where untrusted third parties, like cloud servers, provide CNN-based systems, ensuring the privacy of speech queries becomes essential. This paper proposes encryption methods for speech queries using secret keys and a model structure that allows for encrypted queries to be accepted without decryption. Our approach introduces three types of secret keys: Shuffling, Flipping, and random orthogonal matrix (ROM). In experiments, we demonstrate that when the proposed methods are used with the correct key, identification performance did not degrade. Conversely, when an incorrect key is used, the performance significantly decreased. Particularly, with the use of ROM, we show that even with a relatively small key space, high privacy-preserving performance can be maintained many speech processing tasks. Furthermore, we also demonstrate the difficulty of recovering original speech from encrypted queries in various robustness evaluations.
【2】 One-Shot Distributed Node-Specific Signal Estimation with Non-Overlapping Latent Subspaces in Acoustic Sensor Networks
标题: 声传感器网络中具有非重叠潜在子空间的单次分布式节点特定信号估计
作者:Paul Didier,Pourya Behmandpoor,Toon van Waterschoot,Marc Moonen
链接:点击下载PDF文件
摘要:介绍了一种称为无迭代DANSE(iDANSE)的一次性算法,用于在部署在潜在信号子空间不重叠的环境中的全连接无线声学传感器网络(WASN)中执行分布式自适应特定节点信号估计(DANSE)。iDANSE算法在单个处理周期中与集中式算法的性能相匹配,同时设备交换其多通道本地麦克风信号的融合版本。iDANSE相对于当前可用解决方案的主要优势在于其无迭代特性,这有利于在实时应用中部署,以及设备可以交换比环境中潜在源数量更少的融合信号。所提出的方法进行了验证,包括语音增强方案的数值模拟。摘要:A one-shot algorithm called iterationless DANSE (iDANSE) is introduced to perform distributed adaptive node-specific signal estimation (DANSE) in a fully connected wireless acoustic sensor network (WASN) deployed in an environment with non-overlapping latent signal subspaces. The iDANSE algorithm matches the performance of a centralized algorithm in a single processing cycle while devices exchange fused versions of their multichannel local microphone signals. Key advantages of iDANSE over currently available solutions are its iterationless nature, which favors deployment in real-time applications, and the fact that devices can exchange fewer fused signals than the number of latent sources in the environment. The proposed method is validated in numerical simulations including a speech enhancement scenario.
【3】 Feasibility of iMagLS-BSM -- ILD Informed Binaural Signal Matching with Arbitrary Microphone Arrays
标题: iMagLS-BSM -- ILD知情双耳信号与任意麦克风阵列匹配的可行性
作者:Or Berebi,Zamir Ben-Hur,David Lou Alon,Boaz Rafaely
备注:Paper accepted for publication in IWAENC 2024, 4 pages, 2 figures
链接:点击下载PDF文件
摘要:用于以耳机为中心的收听的双耳再现已成为正在进行的研究的焦点,特别是在诸如增强现实和虚拟现实(AR和VR)等先进技术的领域内。这些应用对高质量空间音频的需求对于保持无缝沉浸感至关重要。然而,由于设计约束,仅配备有有限数量的麦克风和不规则的麦克风放置的可穿戴记录设备带来了挑战。与高阶麦克风阵列捕获的参考信号相比,这些因素导致有限的再现质量。本文介绍了一种新的优化损失量身定制的波束成形为基础的,信号无关的双耳再现方案。这种方法,命名为iMagLS-BSM将一个耳间电平差(ILD)的误差项到先前提出的双耳信号匹配(BSM)幅度最小二乘(MagLS)渲染损失的横向平面角度。该方法利用非线性规划来最小化引入的损失。初步结果表明,ILD误差大幅减少,同时保持双耳幅度误差与MagLS BSM解决方案实现的。这些发现为增强所得到的双耳信号的整体空间质量提供了希望。摘要:Binaural reproduction for headphone-centric listening has become a focal point in ongoing research, particularly within the realm of advancing technologies such as augmented and virtual reality (AR and VR). The demand for high-quality spatial audio in these applications is essential to uphold a seamless sense of immersion. However, challenges arise from wearable recording devices equipped with only a limited number of microphones and irregular microphone placements due to design constraints. These factors contribute to limited reproduction quality compared to reference signals captured by high-order microphone arrays. This paper introduces a novel optimization loss tailored for a beamforming-based, signal-independent binaural reproduction scheme. This method, named iMagLS-BSM incorporates an interaural level difference (ILD) error term into the previously proposed binaural signal matching (BSM) magnitude least squares (MagLS) rendering loss for lateral plane angles. The method leverages nonlinear programming to minimize the introduced loss. Preliminary results show a substantial reduction in ILD error, while maintaining a binaural magnitude error comparable to that achieved with a MagLS BSM solution. These findings hold promise for enhancing the overall spatial quality of resultant binaural signals.
【4】 Bridging the Gap between Audio and Text using Parallel-attention for User-defined Keyword Spotting
标题: 使用用户定义的关键词定位的纵向注意力弥合音频和文本之间的差距
作者:Youkyum Kim,Jaemin Jung,Jihwan Park,Byeong-Yeol Kim,Joon Son Chung
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
摘要:本文提出了一种新颖的用户定义关键字识别框架,可以基于文本注册准确检测音频关键字。由于音频数据与文本相比具有额外的声学信息,因此这两种模态之间存在差异。为了应对这一挑战,我们提出了一种新的KWS,它利用自我和交叉注意力在一个并行的架构,以有效地捕捉信息内和跨两种方式。我们进一步提出了一个音素持续时间为基础的对齐损失,强制执行音频和文本功能之间的顺序对应。大量的实验结果表明,我们提出的方法在可见和不可见的领域中的几个基准数据集上实现了最先进的性能,而无需将先前研究中使用的数据集之外的额外数据合并。摘要:This paper proposes a novel user-defined keyword spotting framework that accurately detects audio keywords based on text enrollment. Since audio data possesses additional acoustic information compared to text, there are discrepancies between these two modalities. To address this challenge, we present ParallelKWS, which utilises self- and cross-attention in a parallel architecture to effectively capture information both within and across the two modalities. We further propose a phoneme duration-based alignment loss that enforces the sequential correspondence between audio and text features. Extensive experimental results demonstrate that our proposed method achieves state-of-the-art performance on several benchmark datasets in both seen and unseen domains, without incorporating extra data beyond the dataset used in previous studies.
【5】 Facing the Music: Tackling Singing Voice Separation in Cinematic Audio Source Separation
标题: 面对音乐:解决电影音频源分离中的歌声分离问题
作者:Karn N. Watcharasupat,Chih-Wei Wu,Iroro Orife
备注:Submitted to the Late-Breaking Demo Session of the 25th International Society for Music Information Retrieval (ISMIR) Conference, 2024
链接:点击下载PDF文件
摘要:电影音频源分离(CASS)是音频源分离的一个较新的子任务。CASS的一个典型设置是一个三干问题,目的是将混合物分离为对话干(DX),音乐干(MX)和效果干(FX)。然而,在实践中,存在几种边缘情况,因为一些声源不完全适合这三个柄中的任何一个,需要在生产中使用额外的辅助柄。一个非常常见的边缘情况是电影音频中的歌声,它可能属于DX或MX,这在很大程度上取决于电影背景。在这项工作中,我们展示了一个非常简单的扩展专用的解码器Bandit和基于查询的单解码器宴会模型的四个干的问题,治疗非音乐对话,器乐,歌声和效果作为单独的干。有趣的是,基于查询的Banquet模型优于专用解码器Bandit模型。我们假设这是由于在瓶颈处的更好的特征对齐,如由与频带无关的薄膜层所实施的。数据集和模型实施将在https: github.com kwatcharasupat source-separation-landing上提供。摘要:Cinematic audio source separation (CASS) is a fairly new subtask of audio source separation. A typical setup of CASS is a three-stem problem, with the aim of separating the mixture into the dialogue stem (DX), music stem (MX), and effects stem (FX). In practice, however, several edge cases exist as some sound sources do not fit neatly in either of these three stems, necessitating the use of additional auxiliary stems in production. One very common edge case is the singing voice in film audio, which may belong in either the DX or MX, depending heavily on the cinematic context. In this work, we demonstrate a very straightforward extension of the dedicated-decoder Bandit and query-based single-decoder Banquet models to a four-stem problem, treating non-musical dialogue, instrumental music, singing voice, and effects as separate stems. Interestingly, the query-based Banquet model outperformed the dedicated-decoder Bandit model. We hypothesized that this is due to a better feature alignment at the bottleneck as enforced by the band-agnostic FiLM layer. Dataset and model implementation will be made available at https: github.com kwatcharasupat source-separation-landing.
【6】 Design and Analysis of Binaural Signal Matching with Arbitrary Microphone Arrays
标题: 任意麦克风阵列的双耳信号匹配设计与分析
作者:Lior Madmoni,Zamir Ben-Hur,Jacob Donley,Vladimir Tourbabin,Boaz Rafaely
备注:Submitted to EURASIP Journal on audio speech and music processing
链接:点击下载PDF文件
摘要:双耳再现正迅速成为研究界非常感兴趣的话题,特别是随着虚拟现实耳机、智能眼镜和头戴式耳机等新的流行设备的激增。为了使收听者沉浸在具有此类设备的虚拟或远程环境中,必须生成逼真且准确的双耳信号。这是具有挑战性的,特别是因为安装在这些设备上的麦克风阵列通常由任意布置的少量麦克风组成,这阻碍了标准音频格式(如高保真度立体声)的使用,并且提供有限的空间分辨率。双耳信号匹配(BSM)方法是最近开发的,以克服这些挑战。虽然它使用相对简单的阵列产生具有低误差的双耳信号,但当引入头部旋转时,其性能显着下降。本文旨在进一步发展BSM方法,克服其局限性。为此,该方法首先进行了详细分析,并提出了一个设计框架,保证准确的双耳再现相对复杂的声学环境。接下来,它表明,BSM的准确性可能会显着降低在高频率,因此,提出了一个感知动机的扩展的方法,基于幅度最小二乘(MagLS)制定。这些见解和发展,然后分析了一个简单的六麦克风半圆形阵列的广泛的仿真研究的帮助下。它进一步表明,BSM-MagLS方法可以是非常有用的,在补偿头部旋转与此阵列。最后,一个听实验进行了一个四麦克风阵列上的一副眼镜在混响语音环境中,包括头部旋转,它表明,BSM-MagLS确实可以产生双耳信号具有高感知质量。摘要:Binaural reproduction is rapidly becoming a topic of great interest in the research community, especially with the surge of new and popular devices, such as virtual reality headsets, smart glasses, and head-tracked headphones. In order to immerse the listener in a virtual or remote environment with such devices, it is essential to generate realistic and accurate binaural signals. This is challenging, especially since the microphone arrays mounted on these devices are typically composed of an arbitrarily-arranged small number of microphones, which impedes the use of standard audio formats like Ambisonics, and provides limited spatial resolution. The binaural signal matching (BSM) method was developed recently to overcome these challenges. While it produced binaural signals with low error using relatively simple arrays, its performance degraded significantly when head rotation was introduced. This paper aims to develop the BSM method further and overcome its limitations. For this purpose, the method is first analyzed in detail, and a design framework that guarantees accurate binaural reproduction for relatively complex acoustic environments is presented. Next, it is shown that the BSM accuracy may significantly degrade at high frequencies, and thus, a perceptually motivated extension to the method is proposed, based on a magnitude least-squares (MagLS) formulation. These insights and developments are then analyzed with the help of an extensive simulation study of a simple six-microphone semi-circular array. It is further shown that the BSM-MagLS method can be very useful in compensating for head rotations with this array. Finally, a listening experiment is conducted with a four-microphone array on a pair of glasses in a reverberant speech environment and including head rotations, where it is shown that BSM-MagLS can indeed produce binaural signals with a high perceived quality.
【7】 TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement
标题: TF-Locoformer:通过卷积进行局部建模的Transformer,用于语音分离和增强
作者:Kohei Saijo,Gordon Wichern,François G. Germain,Zexu Pan,Jonathan Le Roux
备注:Accepted to IWAENC 2024
链接:点击下载PDF文件
摘要:时频(TF)域双径模型实现高保真语音分离。虽然以前的一些最先进的(SoTA)模型依赖于RNN,但这种依赖意味着它们缺乏Transformer块的并行性、可扩展性和多功能性。鉴于纯基于transformer的架构在其他领域取得了广泛的成功,在这项工作中,我们专注于从TF域双路径模型中删除RNN,同时保持SoTA性能。本文提出了TF-Locoformer,一种基于变换器的模型,通过卷积进行局部建模。该模型使用具有卷积层的前馈网络(FFN),而不是线性层,来捕获局部信息,让自我注意力集中在捕获全局模式上。我们将两个这样的FFN之前和之后的自我注意,以提高本地建模能力。我们还介绍了一种新的TF域双路径模型的归一化。分离和增强数据集上的实验表明,该模型在多个基准测试中达到或超过SoTA,具有RNN自由架构。摘要:Time-frequency (TF) domain dual-path models achieve high-fidelity speech separation. While some previous state-of-the-art (SoTA) models rely on RNNs, this reliance means they lack the parallelizability, scalability, and versatility of Transformer blocks. Given the wide-ranging success of pure Transformer-based architectures in other fields, in this work we focus on removing the RNN from TF-domain dual-path models, while maintaining SoTA performance. This work presents TF-Locoformer, a Transformer-based model with LOcal-modeling by COnvolution. The model uses feed-forward networks (FFNs) with convolution layers, instead of linear layers, to capture local information, letting the self-attention focus on capturing global patterns. We place two such FFNs before and after self-attention to enhance the local-modeling capability. We also introduce a novel normalization for TF-domain dual-path models. Experiments on separation and enhancement datasets show that the proposed model meets or exceeds SoTA in multiple benchmarks with an RNN-free architecture.
【8】 Enhanced Reverberation as Supervision for Unsupervised Speech Separation
标题: 增强的回响作为无监督语音分离的监督
作者:Kohei Saijo,Gordon Wichern,François G. Germain,Zexu Pan,Jonathan Le Roux
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:混响监督(RAS)是一个框架,它允许以无监督的方式从多声道混合中训练单声道语音分离模型。在RAS中,训练模型,使得可以映射从输入通道处的混合预测的源,以在目标通道处重建混合。然而,到目前为止,稳定的无监督训练仅在超定源信道条件下实现,而关键的确定情况尚未解决。这项工作提出了增强型RAS(ERAS)来解决这个问题。通过定性分析,我们发现稳定的训练可以通过利用损失项来缓解频率置换问题。分离性能还通过添加新的损失项来提高,其中将映射回其自身输入混合的分离信号用作从其他通道分离并映射到相同通道的信号的伪目标。实验结果表明,ERAS具有较高的稳定性和性能。摘要:Reverberation as supervision (RAS) is a framework that allows for training monaural speech separation models from multi-channel mixtures in an unsupervised manner. In RAS, models are trained so that sources predicted from a mixture at an input channel can be mapped to reconstruct a mixture at a target channel. However, stable unsupervised training has so far only been achieved in over-determined source-channel conditions, leaving the key determined case unsolved. This work proposes enhanced RAS (ERAS) for solving this problem. Through qualitative analysis, we found that stable training can be achieved by leveraging the loss term to alleviate the frequency-permutation problem. Separation performance is also boosted by adding a novel loss term where separated signals mapped back to their own input mixture are used as pseudo-targets for the signals separated from other channels and mapped to the same channel. Experimental results demonstrate high stability and performance of ERAS.
【9】 Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond
标题: Speech-MASSIVE:适用于SL及其他领域的多语言语音数据集
作者:Beomseok Lee,Ioan Calapodescu,Marco Gaido,Matteo Negri,Laurent Besacier
备注:Accepted at INTERSPEECH 2024. This version includes the same content but with additional appendices
链接:点击下载PDF文件
摘要:我们提出了Speech-MASSIVE,一个多语言口语理解(SLU)数据集,包括一部分MASSIVE文本语料库的语音对应物。Speech-MASSIVE涵盖了来自不同家族的12种语言,并从MASSIVE继承了用于意图预测和插槽填充任务的注释。我们的扩展是由大量多语言SLU数据集的稀缺性和对多功能语音数据集的日益增长的需求,以评估跨语言和任务的基础模型(LLM,语音编码器)。我们在各种训练场景(zero-shot、Few-Shot和完全微调)中使用级联和端到端架构提供多模式、多任务、多语言数据集和报告SLU基线。此外,我们证明了适合的Speech-MASSIVE的基准其他任务,如语音转录,语言识别和语音翻译。数据集、模型和代码可在https: github.com hlt-mt Speech-MASSIVE上公开获取摘要:We present Speech-MASSIVE, a multilingual Spoken Language Understanding (SLU) dataset comprising the speech counterpart for a portion of the MASSIVE textual corpus. Speech-MASSIVE covers 12 languages from different families and inherits from MASSIVE the annotations for the intent prediction and slot-filling tasks. Our extension is prompted by the scarcity of massively multilingual SLU datasets and the growing need for versatile speech datasets to assess foundation models (LLMs, speech encoders) across languages and tasks. We provide a multimodal, multitask, multilingual dataset and report SLU baselines using both cascaded and end-to-end architectures in various training scenarios (zero-shot, few-shot, and full fine-tune). Furthermore, we demonstrate the suitability of Speech-MASSIVE for benchmarking other tasks such as speech transcription, language identification, and speech translation. The dataset, models, and code are publicly available at: https: github.com hlt-mt Speech-MASSIVE
【10】 Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
标题: 中央库尔德语文本到语音合成,采用新颖的端到端Transformer训练
作者:Hawraz A. Ahmad,Tarik A. Rashid
备注:19 pages
链接:点击下载PDF文件
摘要:文本到语音(TTS)模型的最新进展旨在将两阶段过程简化为单阶段训练方法。然而,许多单级模型在音频质量方面仍然落后,特别是在处理库尔德语文本和语音时。迫切需要加强库尔德语的文本到语音转换,特别是索拉尼方言,该方言相对被忽视,在最近的文本到语音转换进展中代表性不足。这项研究介绍了一个端到端的TTS模型,有效地产生高品质的库尔德音频。所提出的方法利用变分自动编码器(VAE),该变分自动编码器(VAE)被预先训练用于音频波形重构,并且通过对抗训练来增强。这涉及将预训练编码器建立的先验分布与潜在变量内文本编码器的后验分布进行对齐。此外,一个随机的持续时间预测器被纳入灌输不同节奏的合成库尔德语语音。通过调整潜在分布和整合随机持续时间预测器,所提出的方法有利于实时生成自然库尔德语音音频,提供灵活的音高和节奏。通过自定义数据集上的平均意见得分(MOS)进行的实证评估证实了我们的方法(MOS为3.94)与通过主观人类评估评估的单阶段系统和其他两阶段系统相比具有更好的性能。摘要:Recent advancements in text-to-speech (TTS) models have aimed to streamline the two-stage process into a single-stage training approach. However, many single-stage models still lag behind in audio quality, particularly when handling Kurdish text and speech. There is a critical need to enhance text-to-speech conversion for the Kurdish language, particularly for the Sorani dialect, which has been relatively neglected and is underrepresented in recent text-to-speech advancements. This study introduces an end-to-end TTS model for efficiently generating high-quality Kurdish audio. The proposed method leverages a variational autoencoder (VAE) that is pre-trained for audio waveform reconstruction and is augmented by adversarial training. This involves aligning the prior distribution established by the pre-trained encoder with the posterior distribution of the text encoder within latent variables. Additionally, a stochastic duration predictor is incorporated to imbue synthesized Kurdish speech with diverse rhythms. By aligning latent distributions and integrating the stochastic duration predictor, the proposed method facilitates the real-time generation of natural Kurdish speech audio, offering flexibility in pitches and rhythms. Empirical evaluation via the mean opinion score (MOS) on a custom dataset confirms the superior performance of our approach (MOS of 3.94) compared with that of a one-stage system and other two-staged systems as assessed through a subjective human evaluation.
机器翻译,仅供参考
