本文经arXiv每日学术速递授权转载
【1】 NeuralMultiling: A Novel Neural Architecture Search for Smartphone based Multilingual Speaker Verification
标题: NeuralMultiling:一种基于智能手机的多语言说话者验证的新型神经架构搜索
作者:Aravinda Reddy PN,Raghavendra Ramachandra,K. Sreenivasa Rao,Pabitra Mitra
链接:点击下载PDF文件
【2】 TheGlueNote: Learned Representations for Robust and Flexible Note Alignment
标题: TheGluge Note:学习表示,可实现稳健且灵活的注释对齐
作者:Silvan David Peter,Gerhard Widmer
备注:to be published in Proceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR), 2024
链接:点击下载PDF文件
【3】 Distil-DCCRN: A Small-footprint DCCRN Leveraging Feature-based Knowledge Distillation in Speech Enhancement
标题: Distil-DCCRN:一个在语音增强中利用基于环境的知识提炼的小占地DCCRN
作者:Runduo Han,Weiming Xu,Zihan Zhang,Mingshuai Liu,Lei Xie
备注:Accepted by IEEE Signal Processing Letters
链接:点击下载PDF文件
【4】 wav2graph: A Framework for Supervised Learning Knowledge Graph from Speech
标题: wav 2graph:语音监督学习知识图谱的框架
作者:Khai Le-Duc,Quy-Anh Dang,Tan-Hanh Pham,Truong-Son Hy
备注:Preprint, 32 pages
链接:点击下载PDF文件
【5】 Speaker Adaptation for Quantised End-to-End ASR Models
标题: 量化端到端ASB模型的说话者自适应
作者:Qiuming Zhao,Guangzhi Sun,Chao Zhang,Mingxing Xu,Thomas Fang Zheng
备注:submitted to ASRU 2023 Workshop
链接:点击下载PDF文件
【6】 Articulatory Configurations across Genders and Periods in French Radio and TV archives
标题: 法国广播电视档案中跨性别和时期的发音
作者:Benjamin Elie,David Doukhan,Rémi Uro,Lucas Ondel-Yang,Albert Rilliard,Simon Devauchelle
备注:accepted to InterSpeech 2024, Kos Island, Greece keywords : acoustic to articulatory inversion, diachrony, gender, French, media
链接:点击下载PDF文件
【7】 Assessing the Potential Impact of Direction-Dependent HRTF Selection on Sound Localization Accuracy
标题: 评估方向相关HRTI选择对声音定位准确性的潜在影响
作者:Sapir Goldring,Zamir Ben Hur,David Lou Alon,Boaz Rafaely
备注:Accepted for publication in the 2024 AES International Conference on Audio for Virtual and Augmented Reality, 5 pages, 4 figures
链接:点击下载PDF文件
标题: 法国广播电视档案中跨性别和时期的发音
作者:Benjamin Elie,David Doukhan,Rémi Uro,Lucas Ondel-Yang,Albert Rilliard,Simon Devauchelle
备注:accepted to InterSpeech 2024, Kos Island, Greece keywords : acoustic to articulatory inversion, diachrony, gender, French, media
链接:点击下载PDF文件
【2】 Simulating Articulatory Trajectories with Phonological Feature Interpolation
标题: 利用音素特征插值模拟关节语轨迹
作者:Angelo Ortiz Tandazo,Thomas Schatz,Thomas Hueber,Emmanuel Dupoux
备注:accepted at Interspeech 2024
链接:点击下载PDF文件
【3】 HydraFormer: One Encoder For All Subsampling Rates
标题: HydraFormer:适用于所有子采样率的一个编码器
作者:Yaoxun Xu,Xingchen Song,Zhiyong Wu,Di Wu,Zhendong Peng,Binbin Zhang
备注:accepted by ICME 2024
链接:点击下载PDF文件
【4】 Preserving spoken content in voice anonymisation with character-level vocoder conditioning
标题: 利用字符级声码器条件反射在语音匿名化中保留口语内容
作者:Michele Panariello,Massimiliano Todisco,Nicholas Evans
备注:Accepted at SIG-SPSC 2024 Symposium
链接:点击下载PDF文件
【5】 Assessing the Potential Impact of Direction-Dependent HRTF Selection on Sound Localization Accuracy
标题: 评估方向相关HRTI选择对声音定位准确性的潜在影响
作者:Sapir Goldring,Zamir Ben Hur,David Lou Alon,Boaz Rafaely
备注:Accepted for publication in the 2024 AES International Conference on Audio for Virtual and Augmented Reality, 5 pages, 4 figures
链接:点击下载PDF文件
【6】 NeuralMultiling: A Novel Neural Architecture Search for Smartphone based Multilingual Speaker Verification
标题: NeuralMultiling:一种基于智能手机的多语言说话者验证的新型神经架构搜索
作者:Aravinda Reddy PN,Raghavendra Ramachandra,K. Sreenivasa Rao,Pabitra Mitra
链接:点击下载PDF文件
【7】 TheGlueNote: Learned Representations for Robust and Flexible Note Alignment
标题: TheGluge Note:学习表示,可实现稳健且灵活的注释对齐
作者:Silvan David Peter,Gerhard Widmer
备注:to be published in Proceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR), 2024
链接:点击下载PDF文件
【8】 Distil-DCCRN: A Small-footprint DCCRN Leveraging Feature-based Knowledge Distillation in Speech Enhancement
标题: Distil-DCCRN:一个在语音增强中利用基于环境的知识提炼的小占地DCCRN
作者:Runduo Han,Weiming Xu,Zihan Zhang,Mingshuai Liu,Lei Xie
备注:Accepted by IEEE Signal Processing Letters
链接:点击下载PDF文件
【9】 wav2graph: A Framework for Supervised Learning Knowledge Graph from Speech
标题: wav 2graph:语音监督学习知识图谱的框架
作者:Khai Le-Duc,Quy-Anh Dang,Tan-Hanh Pham,Truong-Son Hy
备注:Preprint, 32 pages
链接:点击下载PDF文件
【10】 Speaker Adaptation for Quantised End-to-End ASR Models
标题: 量化端到端ASB模型的说话者自适应
作者:Qiuming Zhao,Guangzhi Sun,Chao Zhang,Mingxing Xu,Thomas Fang Zheng
备注:submitted to ASRU 2023 Workshop
链接:点击下载PDF文件
标题: NeuralMultiling:一种基于智能手机的多语言说话者验证的新型神经架构搜索
作者:Aravinda Reddy PN,Raghavendra Ramachandra,K. Sreenivasa Rao,Pabitra Mitra
链接:点击下载PDF文件
摘要:多语言说话人验证引入了用多种语言验证说话人的挑战。现有的系统是使用i-vector x-vector方法以及Bi-LSTM构建的,这些方法经过训练可以区分说话者,而不管语言如何。我们提出了一种适合移动设备的多语言说话人验证的神经架构搜索,而不是手动探索设计空间,称为 textbf{NeuralMultiling}。首先,我们的算法搜索一个最佳的操作组合的神经细胞与不同的架构正常细胞和减少细胞,然后推导出一个CNN模型堆叠神经细胞。使用衍生的架构,我们进行了两个不同的研究:1)语言不可知条件和2)语言和设备之间的互操作性公开可用的多语言视听智能手机(MAVS)数据集。实验结果表明,所推导的体系结构在使用更少的模型参数的情况下,等错误率(EER)降低了5- 6%,显着优于现有的自动语音方法。摘要:Multilingual speaker verification introduces the challenge of verifying a speaker in multiple languages. Existing systems were built using i-vector x-vector approaches along with Bi-LSTMs, which were trained to discriminate speakers, irrespective of the language. Instead of exploring the design space manually, we propose a neural architecture search for multilingual speaker verification suitable for mobile devices, called textbf{NeuralMultiling}. First, our algorithm searches for an optimal operational combination of neural cells with different architectures for normal cells and reduction cells and then derives a CNN model by stacking neural cells. Using the derived architecture, we performed two different studies:1) language agnostic condition and 2) interoperability between languages and devices on the publicly available Multilingual Audio-Visual Smartphone (MAVS) dataset. The experimental results suggest that the derived architecture significantly outperforms the existing Autospeech method by a 5-6 % reduction in the Equal Error Rate (EER) with fewer model parameters.
【2】 TheGlueNote: Learned Representations for Robust and Flexible Note Alignment
标题: TheGluge Note:学习表示,可实现稳健且灵活的注释对齐
作者:Silvan David Peter,Gerhard Widmer
备注:to be published in Proceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR), 2024
链接:点击下载PDF文件
摘要:音符对齐指的是将同一个符号编码作品的两个版本的单个音符进行匹配的任务。解决该任务的方法通常依赖于序列比对算法,例如直接应用于音符或起始序列的隐马尔可夫模型或动态时间规整(DTW)。虽然在许多情况下是成功的,但这些方法在版本之间存在很大的不匹配。在这项工作中,我们学习注意到明智的表示与各种复杂的不匹配的情况下,例如,重复,跳过,块插入,和长颤音增强的数据。在我们的方法的核心是一个Transformer编码器网络-TheGesterNote-预测两个512音符的相似性。我们使用weightedDTW和音高分离的onsetDTW的风味后处理预测的相似性检索注意匹配两个序列的任意长度。我们的方法在音符对齐精度方面与最先进的方法不相上下,对版本不匹配的鲁棒性更强,并且直接适用于任何一对双文件。摘要:Note alignment refers to the task of matching individual notes of two versions of the same symbolically encoded piece. Methods addressing this task commonly rely on sequence alignment algorithms such as Hidden Markov Models or Dynamic Time Warping (DTW) applied directly to note or onset sequences. While successful in many cases, such methods struggle with large mismatches between the versions. In this work, we learn note-wise representations from data augmented with various complex mismatch cases, e.g. repeats, skips, block insertions, and long trills. At the heart of our approach lies a transformer encoder network - TheGlueNote - which predicts pairwise note similarities for two 512 note subsequences. We postprocess the predicted similarities using flavors of weightedDTW and pitch-separated onsetDTW to retrieve note matches for two sequences of arbitrary length. Our approach performs on par with the state of the art in terms of note alignment accuracy, is considerably more robust to version mismatches, and works directly on any pair of MIDI files.
【3】 Distil-DCCRN: A Small-footprint DCCRN Leveraging Feature-based Knowledge Distillation in Speech Enhancement
标题: Distil-DCCRN:一个在语音增强中利用基于环境的知识提炼的小占地DCCRN
作者:Runduo Han,Weiming Xu,Zihan Zhang,Mingshuai Liu,Lei Xie
备注:Accepted by IEEE Signal Processing Letters
链接:点击下载PDF文件
摘要:深度复卷积递归网络(DCCRN)利用音频频谱的复特性,实现了良好的语音增强性能。然而,它有大量的模型参数。我们提出了一个更小的模型,蒸馏DCCRN,它只有30%的参数相比,DCCRN。为了确保Distil-DCCRN的性能与DCCRN的性能相匹配,我们采用知识蒸馏(KD)方法来使用较大的教师模型来帮助训练较小的学生模型。设计了一种知识提取方法,将注意力转移和Kullback-Leibler发散(AT-KL)相结合,对学生模型Distil-DCCRN进行训练。此外,我们使用一个模型,具有更好的性能和更复杂的结构,Uformer,作为教师模型。与之前主要关注模型输出的KD方法不同,我们的方法还利用了模型中间层的中间特征,促进了不同结构化模型之间丰富的知识传递,尽管中间特征的层配置存在变化以及通道和时间维度的差异。使用我们的AT-KL方法,Distil-DCCRN在DNS测试集上的PESQ和SI-SNR指标方面优于DCCRN以及其他几个竞争模型,并在DNSMOS中实现了与DCCRN相当的结果。摘要:The deep complex convolution recurrent network (DCCRN) achieves excellent speech enhancement performance by utilizing the audio spectrum's complex features. However, it has a large number of model parameters. We propose a smaller model, Distil-DCCRN, which has only 30% of the parameters compared to the DCCRN. To ensure that the performance of Distil-DCCRN matches that of the DCCRN, we employ the knowledge distillation (KD) method to use a larger teacher model to help train a smaller student model. We design a knowledge distillation (KD) method, integrating attention transfer and Kullback-Leibler divergence (AT-KL) to train the student model Distil-DCCRN. Additionally, we use a model with better performance and a more complicated structure, Uformer, as the teacher model. Unlike previous KD approaches that mainly focus on model outputs, our method also leverages the intermediate features from the models' middle layers, facilitating rich knowledge transfer across different structured models despite variations in layer configurations and discrepancies in the channel and time dimensions of intermediate features. Employing our AT-KL approach, Distil-DCCRN outperforms DCCRN as well as several other competitive models in both PESQ and SI-SNR metrics on the DNS test set and achieves comparable results to DCCRN in DNSMOS.
【4】 wav2graph: A Framework for Supervised Learning Knowledge Graph from Speech
标题: wav 2graph:语音监督学习知识图谱的框架
作者:Khai Le-Duc,Quy-Anh Dang,Tan-Hanh Pham,Truong-Son Hy
备注:Preprint, 32 pages
链接:点击下载PDF文件
摘要:知识图(KGs)通过提供结构化的互连数据来提高大型语言模型(LLM)和搜索引擎的性能,从而提高推理和上下文感知能力。然而,知识库只关注文本数据,从而忽略了其他形式,如语音。在这项工作中,我们介绍了wav2graph,第一个从语音数据中监督学习知识图的框架。我们的管道很简单:(1)基于转录的口语和命名实体数据库构建KG,(2)将KG转换为嵌入向量,以及(3)训练图神经网络(GNN)用于节点分类和链接预测任务。通过使用最先进的GNN模型在归纳和转导学习环境中进行的广泛实验,我们为人类成绩单和自动语音识别(ASR)成绩单的节点分类和链接预测任务提供了基线结果和错误分析,包括使用基于编码器和基于解码器的节点嵌入以及单语和多语言声学预训练模型的评估。所有相关的代码、数据和模型都在线发布。摘要:Knowledge graphs (KGs) enhance the performance of large language models (LLMs) and search engines by providing structured, interconnected data that improves reasoning and context-awareness. However, KGs only focus on text data, thereby neglecting other modalities such as speech. In this work, we introduce wav2graph, the first framework for supervised learning knowledge graph from speech data. Our pipeline are straightforward: (1) constructing a KG based on transcribed spoken utterances and a named entity database, (2) converting KG into embedding vectors, and (3) training graph neural networks (GNNs) for node classification and link prediction tasks. Through extensive experiments conducted in inductive and transductive learning contexts using state-of-the-art GNN models, we provide baseline results and error analysis for node classification and link prediction tasks on human transcripts and automatic speech recognition (ASR) transcripts, including evaluations using both encoder-based and decoder-based node embeddings, as well as monolingual and multilingual acoustic pre-trained models. All related code, data, and models are published online.
【5】 Speaker Adaptation for Quantised End-to-End ASR Models
标题: 量化端到端ASB模型的说话者自适应
作者:Qiuming Zhao,Guangzhi Sun,Chao Zhang,Mingxing Xu,Thomas Fang Zheng
备注:submitted to ASRU 2023 Workshop
链接:点击下载PDF文件
摘要:端到端模型在自动语音识别(ASR)方面表现出了卓越的性能。然而,这样的模型通常尺寸非常大,因此在资源受限的边缘设备上部署具有挑战性。虽然量化可以减少模型大小,但它可能会导致字错误率(WER)增加。虽然提出了改进的量化方法来解决性能下降的问题,但部署在边缘设备上的量化模型通常只针对一小群用户的事实尚未得到充分探讨。为此,我们提出了个性化的量化模型(P4 Q),一种新的策略,使用扬声器自适应(SA),以改善量化的端到端ASR模型,通过拟合他们的目标扬声器的特性。在本文中,我们研究了基于Whisper和Conformer基于注意力的编码器-解码器(AED)端到端ASR模型的P4 Q策略,该模型利用4位块式NormalFloat 4(NF 4)方法进行量化,并利用低秩自适应(LoRA)方法进行SA。在LibriSpeech和TED-LIUM 3语料库上的实验结果表明,与全精度模型相比,量化Whisper和Conformer AED模型在模型大小减少7倍和增加1%说话人特定参数的情况下,分别实现了15.1%和23.3%的相对WER降低.摘要:End-to-end models have shown superior performance for automatic speech recognition (ASR). However, such models are often very large in size and thus challenging to deploy on resource-constrained edge devices. While quantisation can reduce model sizes, it can lead to increased word error rates (WERs). Although improved quantisation methods were proposed to address the issue of performance degradation, the fact that quantised models deployed on edge devices often target only on a small group of users is under-explored. To this end, we propose personalisation for quantised models (P4Q), a novel strategy that uses speaker adaptation (SA) to improve quantised end-to-end ASR models by fitting them to the characteristics of the target speakers. In this paper, we study the P4Q strategy based on Whisper and Conformer attention-based encoder-decoder (AED) end-to-end ASR models, which leverages a 4-bit block-wise NormalFloat4 (NF4) approach for quantisation and the low-rank adaptation (LoRA) approach for SA. Experimental results on the LibriSpeech and the TED-LIUM 3 corpora show that, with a 7-time reduction in model size and 1% extra speaker-specific parameters, 15.1% and 23.3% relative WER reductions were achieved on quantised Whisper and Conformer AED models respectively, comparing to the full precision models.
【6】 Articulatory Configurations across Genders and Periods in French Radio and TV archives
标题: 法国广播电视档案中跨性别和时期的发音
作者:Benjamin Elie,David Doukhan,Rémi Uro,Lucas Ondel-Yang,Albert Rilliard,Simon Devauchelle
备注:accepted to InterSpeech 2024, Kos Island, Greece keywords : acoustic to articulatory inversion, diachrony, gender, French, media
链接:点击下载PDF文件
摘要:本文研究发音配置的变化,跨性别和时期使用的声学发音参数的反演。从基于1955年至2015年60年法国媒体档案的历时语料库中,自动转录和强制对齐允许提取每个元音的中心框架。超过100万帧来自1000多个性别和年龄类别的发言者。他们的共振峰被用来从这些声乐框架,以适应前田的发音模型的参数。提供了这些过程的质量评价。在这里,我们专注于两个参数的前田的模型与总声道长度:喉的相对位置(女性较高)和嘴唇突出(男性更突出)。语音质量跨性别的影响进行了讨论。跨时期的影响似乎性别独立,因此,断言女性降低了他们的音调随着时间的推移不支持。摘要:This paper studies changes in articulatory configurations across genders and periods using an inversion from acoustic to articulatory parameters. From a diachronic corpus based on French media archives spanning 60 years from 1955 to 2015, automatic transcription and forced alignment allowed extracting the central frame of each vowel. More than one million frames were obtained from over a thousand speakers across gender and age categories. Their formants were used from these vocalic frames to fit the parameters of Maeda's articulatory model. Evaluations of the quality of these processes are provided. We focus here on two parameters of Maeda's model linked to total vocal tract length: the relative position of the larynx (higher for females) and the lips protrusion (more protruded for males). Implications for voice quality across genders are discussed. The effect across periods seems gender independent; thus, the assertion that females lowered their pitch with time is not supported.
【7】 Assessing the Potential Impact of Direction-Dependent HRTF Selection on Sound Localization Accuracy
标题: 评估方向相关HRTI选择对声音定位准确性的潜在影响
作者:Sapir Goldring,Zamir Ben Hur,David Lou Alon,Boaz Rafaely
备注:Accepted for publication in the 2024 AES International Conference on Audio for Virtual and Augmented Reality, 5 pages, 4 figures
链接:点击下载PDF文件
摘要:本研究探讨了头相关传递函数(HRTF)的方向依赖性选择方法及其对声音定位精度的影响。对于虚拟现实(VR)和电话会议等应用,获得个性化的HRTF可能是有益的,但具有挑战性,因此,这项工作的目的是评估是否将HRTF以方向相关的方式可以提高定位精度,而不需要获得个性化的HRTF。使用VR耳机进行的定位实验评估了定位误差,比较了来自一组的整体最佳HRTF,并根据每个方向的平均性能选择最佳HRTF。结果表明,与方向相关的HRTF选择激励的方法在仰角定位误差的实质性改善,而揭示方位角误差的差异不明显。摘要:This study investigates the approach of direction-dependent selection of Head-Related Transfer Functions (HRTFs) and its impact on sound localization accuracy. For applications such as virtual reality (VR) and teleconferencing, obtaining individualized HRTFs can be beneficial yet challenging, the objective of this work is therefore to assess whether incorporating HRTFs in a direction-dependent manner could improve localization precision without the need to obtain individualized HRTFs. A localization experiment conducted with a VR headset assessed localization errors, comparing an overall best HRTF from a set, against selecting the best HRTF based on average performance in each direction. The results demonstrate a substantial improvement in elevation localization error with the method motivated by direction-dependent HRTF selection, while revealing insignificant differences in azimuth errors.
eess.AS音频处理
【1】 Articulatory Configurations across Genders and Periods in French Radio and TV archives标题: 法国广播电视档案中跨性别和时期的发音
作者:Benjamin Elie,David Doukhan,Rémi Uro,Lucas Ondel-Yang,Albert Rilliard,Simon Devauchelle
备注:accepted to InterSpeech 2024, Kos Island, Greece keywords : acoustic to articulatory inversion, diachrony, gender, French, media
链接:点击下载PDF文件
摘要:本文研究发音配置的变化,跨性别和时期使用的声学发音参数的反演。从基于1955年至2015年60年法国媒体档案的历时语料库中,自动转录和强制对齐允许提取每个元音的中心框架。超过100万帧来自1000多个性别和年龄类别的发言者。他们的共振峰被用来从这些声乐框架,以适应前田的发音模型的参数。提供了这些过程的质量评价。在这里,我们专注于两个参数的前田的模型与总声道长度:喉的相对位置(女性较高)和嘴唇突出(男性更突出)。语音质量跨性别的影响进行了讨论。跨时期的影响似乎性别独立,因此,断言女性降低了他们的音调随着时间的推移不支持。摘要:This paper studies changes in articulatory configurations across genders and periods using an inversion from acoustic to articulatory parameters. From a diachronic corpus based on French media archives spanning 60 years from 1955 to 2015, automatic transcription and forced alignment allowed extracting the central frame of each vowel. More than one million frames were obtained from over a thousand speakers across gender and age categories. Their formants were used from these vocalic frames to fit the parameters of Maeda's articulatory model. Evaluations of the quality of these processes are provided. We focus here on two parameters of Maeda's model linked to total vocal tract length: the relative position of the larynx (higher for females) and the lips protrusion (more protruded for males). Implications for voice quality across genders are discussed. The effect across periods seems gender independent; thus, the assertion that females lowered their pitch with time is not supported.
【2】 Simulating Articulatory Trajectories with Phonological Feature Interpolation
标题: 利用音素特征插值模拟关节语轨迹
作者:Angelo Ortiz Tandazo,Thomas Schatz,Thomas Hueber,Emmanuel Dupoux
备注:accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:作为第一步,对一个完整的计算模型的语音学习,涉及感知生产循环,我们研究了伪电机命令和发音轨迹之间的正向映射。两个语音特征集,分别基于生成音系学和发音音系学,用于编码语音目标序列。比较不同的插值技术,以在这些特征空间中生成平滑的轨迹,并对目标值和时间进行潜在的优化,以捕获协同发音效果。我们报告皮尔逊相关性的线性投影生成的轨迹和发音数据来自多扬声器数据集的电磁关节造影(EMA)记录。基于生成音系学和线性插值技术的扩展特征集获得了0.67的相关性。我们讨论了我们的研究结果的影响,我们的生物运动的动力学的理解。摘要:As a first step towards a complete computational model of speech learning involving perception-production loops, we investigate the forward mapping between pseudo-motor commands and articulatory trajectories. Two phonological feature sets, based respectively on generative and articulatory phonology, are used to encode a phonetic target sequence. Different interpolation techniques are compared to generate smooth trajectories in these feature spaces, with a potential optimisation of the target value and timing to capture co-articulation effects. We report the Pearson correlation between a linear projection of the generated trajectories and articulatory data derived from a multi-speaker dataset of electromagnetic articulography (EMA) recordings. A correlation of 0.67 is obtained with an extended feature set based on generative phonology and a linear interpolation technique. We discuss the implications of our results for our understanding of the dynamics of biological motion.
【3】 HydraFormer: One Encoder For All Subsampling Rates
标题: HydraFormer:适用于所有子采样率的一个编码器
作者:Yaoxun Xu,Xingchen Song,Zhiyong Wu,Di Wu,Zhendong Peng,Binbin Zhang
备注:accepted by ICME 2024
链接:点击下载PDF文件
摘要:在自动语音识别中,子采样对于处理不同的场景至关重要。然而,单一子采样率不足以解决各种现实情况,往往需要训练和部署多个模型,从而增加了相关成本。为了解决这个问题,我们提出了HydraFormer,包括HydraSub,基于Conformer的编码器和基于BiTransformer的解码器。HydraSub包含多个分支,每个分支代表不同的子采样率,允许在基于特定用例的推理过程中灵活选择任何分支。HydraFormer可有效管理不同的二次采样率,显著降低培训和部署成本。在AISHELL-1和LibriSpeech数据集上的实验表明,HydraFormer有效地适应了各种子采样率和语言,同时保持了较高的识别性能。此外,HydraFormer具有出色的稳定性,在各种初始化条件下保持一致的性能,并通过从预训练的单次采样率自动语音识别模型中学习,显示出强大的可移植性 footnote{模型代码和脚本:https: github.com HydraFormer hydraformer}。摘要:In automatic speech recognition, subsampling is essential for tackling diverse scenarios. However, the inadequacy of a single subsampling rate to address various real-world situations often necessitates training and deploying multiple models, consequently increasing associated costs. To address this issue, we propose HydraFormer, comprising HydraSub, a Conformer-based encoder, and a BiTransformer-based decoder. HydraSub encompasses multiple branches, each representing a distinct subsampling rate, allowing for the flexible selection of any branch during inference based on the specific use case. HydraFormer can efficiently manage different subsampling rates, significantly reducing training and deployment expenses. Experiments on AISHELL-1 and LibriSpeech datasets reveal that HydraFormer effectively adapts to various subsampling rates and languages while maintaining high recognition performance. Additionally, HydraFormer showcases exceptional stability, sustaining consistent performance under various initialization conditions, and exhibits robust transferability by learning from pretrained single subsampling rate automatic speech recognition models footnote{Model code and scripts: https: github.com HydraFormer hydraformer}.
【4】 Preserving spoken content in voice anonymisation with character-level vocoder conditioning
标题: 利用字符级声码器条件反射在语音匿名化中保留口语内容
作者:Michele Panariello,Massimiliano Todisco,Nicholas Evans
备注:Accepted at SIG-SPSC 2024 Symposium
链接:点击下载PDF文件
摘要:当语音数据与不受信任的其他人共享时,语音匿名可用于帮助保护说话者隐私。在大多数实际应用中,虽然语音身份应该被净化,但其他属性(如语音内容)应该被保留。总有一个权衡;迄今为止报道的所有方法都为了匿名性能而牺牲了口语内容。据我们所知,我们报告了第一次尝试在语音匿名化中积极保存口语内容。我们展示了辅助自动语音识别模型的输出如何用于使用一组可学习的嵌入词典来调节匿名化系统的声码器模块,以保留口语内容。相对于基线方法,并且在匿名化性能方面仅花费适度的成本,该技术成功地将从匿名话语计算的单词错误率降低了近60%。摘要:Voice anonymisation can be used to help protect speaker privacy when speech data is shared with untrusted others. In most practical applications, while the voice identity should be sanitised, other attributes such as the spoken content should be preserved. There is always a trade-off; all approaches reported thus far sacrifice spoken content for anonymisation performance. We report what is, to the best of our knowledge, the first attempt to actively preserve spoken content in voice anonymisation. We show how the output of an auxiliary automatic speech recognition model can be used to condition the vocoder module of an anonymisation system using a set of learnable embedding dictionaries in order to preserve spoken content. Relative to a baseline approach, and for only a modest cost in anonymisation performance, the technique is successful in decreasing the word error rate computed from anonymised utterances by almost 60%.
【5】 Assessing the Potential Impact of Direction-Dependent HRTF Selection on Sound Localization Accuracy
标题: 评估方向相关HRTI选择对声音定位准确性的潜在影响
作者:Sapir Goldring,Zamir Ben Hur,David Lou Alon,Boaz Rafaely
备注:Accepted for publication in the 2024 AES International Conference on Audio for Virtual and Augmented Reality, 5 pages, 4 figures
链接:点击下载PDF文件
摘要:本研究探讨了头相关传递函数(HRTF)的方向依赖性选择方法及其对声音定位精度的影响。对于虚拟现实(VR)和电话会议等应用,获得个性化的HRTF可能是有益的,但具有挑战性,因此,这项工作的目的是评估是否将HRTF以方向相关的方式可以提高定位精度,而不需要获得个性化的HRTF。使用VR耳机进行的定位实验评估了定位误差,比较了来自一组的整体最佳HRTF,并根据每个方向的平均性能选择最佳HRTF。结果表明,与方向相关的HRTF选择激励的方法在仰角定位误差的实质性改善,而揭示方位角误差的差异不明显。摘要:This study investigates the approach of direction-dependent selection of Head-Related Transfer Functions (HRTFs) and its impact on sound localization accuracy. For applications such as virtual reality (VR) and teleconferencing, obtaining individualized HRTFs can be beneficial yet challenging, the objective of this work is therefore to assess whether incorporating HRTFs in a direction-dependent manner could improve localization precision without the need to obtain individualized HRTFs. A localization experiment conducted with a VR headset assessed localization errors, comparing an overall best HRTF from a set, against selecting the best HRTF based on average performance in each direction. The results demonstrate a substantial improvement in elevation localization error with the method motivated by direction-dependent HRTF selection, while revealing insignificant differences in azimuth errors.
【6】 NeuralMultiling: A Novel Neural Architecture Search for Smartphone based Multilingual Speaker Verification
标题: NeuralMultiling:一种基于智能手机的多语言说话者验证的新型神经架构搜索
作者:Aravinda Reddy PN,Raghavendra Ramachandra,K. Sreenivasa Rao,Pabitra Mitra
链接:点击下载PDF文件
摘要:多语言说话人验证引入了用多种语言验证说话人的挑战。现有的系统是使用i-vector x-vector方法以及Bi-LSTM构建的,这些方法经过训练可以区分说话者,而不管语言如何。我们提出了一种适合移动设备的多语言说话人验证的神经架构搜索,而不是手动探索设计空间,称为 textbf{NeuralMultiling}。首先,我们的算法搜索一个最佳的操作组合的神经细胞与不同的架构正常细胞和减少细胞,然后推导出一个CNN模型堆叠神经细胞。使用衍生的架构,我们进行了两个不同的研究:1)语言不可知条件和2)语言和设备之间的互操作性公开可用的多语言视听智能手机(MAVS)数据集。实验结果表明,所推导的体系结构在使用更少的模型参数的情况下,等错误率(EER)降低了5- 6%,显着优于现有的自动语音方法。摘要:Multilingual speaker verification introduces the challenge of verifying a speaker in multiple languages. Existing systems were built using i-vector x-vector approaches along with Bi-LSTMs, which were trained to discriminate speakers, irrespective of the language. Instead of exploring the design space manually, we propose a neural architecture search for multilingual speaker verification suitable for mobile devices, called textbf{NeuralMultiling}. First, our algorithm searches for an optimal operational combination of neural cells with different architectures for normal cells and reduction cells and then derives a CNN model by stacking neural cells. Using the derived architecture, we performed two different studies:1) language agnostic condition and 2) interoperability between languages and devices on the publicly available Multilingual Audio-Visual Smartphone (MAVS) dataset. The experimental results suggest that the derived architecture significantly outperforms the existing Autospeech method by a 5-6 % reduction in the Equal Error Rate (EER) with fewer model parameters.
【7】 TheGlueNote: Learned Representations for Robust and Flexible Note Alignment
标题: TheGluge Note:学习表示,可实现稳健且灵活的注释对齐
作者:Silvan David Peter,Gerhard Widmer
备注:to be published in Proceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR), 2024
链接:点击下载PDF文件
摘要:音符对齐指的是将同一个符号编码作品的两个版本的单个音符进行匹配的任务。解决该任务的方法通常依赖于序列比对算法,例如直接应用于音符或起始序列的隐马尔可夫模型或动态时间规整(DTW)。虽然在许多情况下是成功的,但这些方法在版本之间存在很大的不匹配。在这项工作中,我们学习注意到明智的表示与各种复杂的不匹配的情况下,例如,重复,跳过,块插入,和长颤音增强的数据。在我们的方法的核心是一个Transformer编码器网络-TheGesterNote-预测两个512音符的相似性。我们使用weightedDTW和音高分离的onsetDTW的风味后处理预测的相似性检索注意匹配两个序列的任意长度。我们的方法在音符对齐精度方面与最先进的方法不相上下,对版本不匹配的鲁棒性更强,并且直接适用于任何一对双文件。摘要:Note alignment refers to the task of matching individual notes of two versions of the same symbolically encoded piece. Methods addressing this task commonly rely on sequence alignment algorithms such as Hidden Markov Models or Dynamic Time Warping (DTW) applied directly to note or onset sequences. While successful in many cases, such methods struggle with large mismatches between the versions. In this work, we learn note-wise representations from data augmented with various complex mismatch cases, e.g. repeats, skips, block insertions, and long trills. At the heart of our approach lies a transformer encoder network - TheGlueNote - which predicts pairwise note similarities for two 512 note subsequences. We postprocess the predicted similarities using flavors of weightedDTW and pitch-separated onsetDTW to retrieve note matches for two sequences of arbitrary length. Our approach performs on par with the state of the art in terms of note alignment accuracy, is considerably more robust to version mismatches, and works directly on any pair of MIDI files.
【8】 Distil-DCCRN: A Small-footprint DCCRN Leveraging Feature-based Knowledge Distillation in Speech Enhancement
标题: Distil-DCCRN:一个在语音增强中利用基于环境的知识提炼的小占地DCCRN
作者:Runduo Han,Weiming Xu,Zihan Zhang,Mingshuai Liu,Lei Xie
备注:Accepted by IEEE Signal Processing Letters
链接:点击下载PDF文件
摘要:深度复卷积递归网络(DCCRN)利用音频频谱的复特性,实现了良好的语音增强性能。然而,它有大量的模型参数。我们提出了一个更小的模型,蒸馏DCCRN,它只有30%的参数相比,DCCRN。为了确保Distil-DCCRN的性能与DCCRN的性能相匹配,我们采用知识蒸馏(KD)方法来使用较大的教师模型来帮助训练较小的学生模型。设计了一种知识提取方法,将注意力转移和Kullback-Leibler发散(AT-KL)相结合,对学生模型Distil-DCCRN进行训练。此外,我们使用一个模型,具有更好的性能和更复杂的结构,Uformer,作为教师模型。与以前主要关注模型输出的KD方法不同,我们的方法还利用了模型中间层的中间特征,促进了不同结构化模型之间的丰富知识转移,尽管中间特征的层配置和通道和时间维度存在差异。使用我们的AT-KL方法,Distil-DCCRN在DNS测试集上的PESQ和SI-SNR指标方面优于DCCRN以及其他几个竞争模型,并在DNSMOS中实现了与DCCRN相当的结果。摘要:The deep complex convolution recurrent network (DCCRN) achieves excellent speech enhancement performance by utilizing the audio spectrum's complex features. However, it has a large number of model parameters. We propose a smaller model, Distil-DCCRN, which has only 30% of the parameters compared to the DCCRN. To ensure that the performance of Distil-DCCRN matches that of the DCCRN, we employ the knowledge distillation (KD) method to use a larger teacher model to help train a smaller student model. We design a knowledge distillation (KD) method, integrating attention transfer and Kullback-Leibler divergence (AT-KL) to train the student model Distil-DCCRN. Additionally, we use a model with better performance and a more complicated structure, Uformer, as the teacher model. Unlike previous KD approaches that mainly focus on model outputs, our method also leverages the intermediate features from the models' middle layers, facilitating rich knowledge transfer across different structured models despite variations in layer configurations and discrepancies in the channel and time dimensions of intermediate features. Employing our AT-KL approach, Distil-DCCRN outperforms DCCRN as well as several other competitive models in both PESQ and SI-SNR metrics on the DNS test set and achieves comparable results to DCCRN in DNSMOS.
【9】 wav2graph: A Framework for Supervised Learning Knowledge Graph from Speech
标题: wav 2graph:语音监督学习知识图谱的框架
作者:Khai Le-Duc,Quy-Anh Dang,Tan-Hanh Pham,Truong-Son Hy
备注:Preprint, 32 pages
链接:点击下载PDF文件
摘要:知识图(KGs)通过提供结构化的互连数据来提高大型语言模型(LLM)和搜索引擎的性能,从而提高推理和上下文感知能力。然而,知识库只关注文本数据,从而忽略了其他形式,如语音。在这项工作中,我们介绍了wav2graph,第一个从语音数据中监督学习知识图的框架。我们的管道很简单:(1)基于转录的口语和命名实体数据库构建KG,(2)将KG转换为嵌入向量,以及(3)训练图神经网络(GNN)用于节点分类和链接预测任务。通过使用最先进的GNN模型在归纳和转导学习环境中进行的广泛实验,我们为人类成绩单和自动语音识别(ASR)成绩单的节点分类和链接预测任务提供了基线结果和错误分析,包括使用基于编码器和基于解码器的节点嵌入以及单语和多语言声学预训练模型的评估。所有相关的代码、数据和模型都在线发布。摘要:Knowledge graphs (KGs) enhance the performance of large language models (LLMs) and search engines by providing structured, interconnected data that improves reasoning and context-awareness. However, KGs only focus on text data, thereby neglecting other modalities such as speech. In this work, we introduce wav2graph, the first framework for supervised learning knowledge graph from speech data. Our pipeline are straightforward: (1) constructing a KG based on transcribed spoken utterances and a named entity database, (2) converting KG into embedding vectors, and (3) training graph neural networks (GNNs) for node classification and link prediction tasks. Through extensive experiments conducted in inductive and transductive learning contexts using state-of-the-art GNN models, we provide baseline results and error analysis for node classification and link prediction tasks on human transcripts and automatic speech recognition (ASR) transcripts, including evaluations using both encoder-based and decoder-based node embeddings, as well as monolingual and multilingual acoustic pre-trained models. All related code, data, and models are published online.
【10】 Speaker Adaptation for Quantised End-to-End ASR Models
标题: 量化端到端ASB模型的说话者自适应
作者:Qiuming Zhao,Guangzhi Sun,Chao Zhang,Mingxing Xu,Thomas Fang Zheng
备注:submitted to ASRU 2023 Workshop
链接:点击下载PDF文件
摘要:端到端模型在自动语音识别(ASR)方面表现出了卓越的性能。然而,这样的模型通常尺寸非常大,因此在资源受限的边缘设备上部署具有挑战性。虽然量化可以减少模型大小,但它可能会导致字错误率(WER)增加。虽然提出了改进的量化方法来解决性能下降的问题,但部署在边缘设备上的量化模型通常只针对一小群用户的事实尚未得到充分探讨。为此,我们提出了个性化的量化模型(P4 Q),一种新的策略,使用扬声器自适应(SA),以改善量化的端到端ASR模型,通过拟合他们的目标扬声器的特性。在本文中,我们研究了基于Whisper和Conformer基于注意力的编码器-解码器(AED)端到端ASR模型的P4 Q策略,该模型利用4位块式NormalFloat 4(NF 4)方法进行量化,并利用低秩自适应(LoRA)方法进行SA。在LibriSpeech和TED-LIUM 3语料库上的实验结果表明,与全精度模型相比,量化Whisper和Conformer AED模型在模型大小减少7倍和增加1%说话人特定参数的情况下,分别实现了15.1%和23.3%的相对WER降低.摘要:End-to-end models have shown superior performance for automatic speech recognition (ASR). However, such models are often very large in size and thus challenging to deploy on resource-constrained edge devices. While quantisation can reduce model sizes, it can lead to increased word error rates (WERs). Although improved quantisation methods were proposed to address the issue of performance degradation, the fact that quantised models deployed on edge devices often target only on a small group of users is under-explored. To this end, we propose personalisation for quantised models (P4Q), a novel strategy that uses speaker adaptation (SA) to improve quantised end-to-end ASR models by fitting them to the characteristics of the target speakers. In this paper, we study the P4Q strategy based on Whisper and Conformer attention-based encoder-decoder (AED) end-to-end ASR models, which leverages a 4-bit block-wise NormalFloat4 (NF4) approach for quantisation and the low-rank adaptation (LoRA) approach for SA. Experimental results on the LibriSpeech and the TED-LIUM 3 corpora show that, with a 7-time reduction in model size and 1% extra speaker-specific parameters, 15.1% and 23.3% relative WER reductions were achieved on quantised Whisper and Conformer AED models respectively, comparing to the full precision models.
机器翻译,仅供参考
