今日论文合集:cs.SD语音16篇,eess.AS音频处理18篇。

本文经arXiv每日学术速递授权转载

cs.SD语音

【1】T-FOLEY: A Controllable Waveform-Domain Diffusion Model for  Temporal-Event-Guided Foley Sound Synthesis

标题:T-Foley:时间事件引导的Foley声音合成的可控波形域扩散模型
链接:https://arxiv.org/abs/2401.09294
作者:Yoonjin Chung,Junwon Lee,Juhan Nam
摘要:Foley sound是与视频同步插入的音频内容,在多媒体内容的用户体验中起着至关重要的作用。最近,人们对Foley声音合成进行了积极的研究,利用了深度生成模型的进步。然而,这些工作主要集中在复制一个单一的声音类或文本的声音描述,忽略了时间信息,这是至关重要的福利声音的实际应用。我们提出了T-Foley,一个用于Foley声音合成的时间事件引导波形生成模型。T-Foley使用两个条件生成高质量的音频:声音类别和时间事件特征。对于时间条件反射,我们设计了一个时间事件功能和一个新的条件反射技术,名为块电影。T-Foley在客观和主观评估指标方面都取得了优异的性能,并生成与时间事件同步的Foley声音。此外,我们展示T-福利的实际应用,特别是在涉及语音模仿的时间事件控制的情况下。我们在我们的配套网站上展示演示。
摘要:Foley sound, audio content inserted synchronously with videos, plays a critical role in the user experience of multimedia content. Recently, there has been active research in Foley sound synthesis, leveraging the advancements in deep generative models. However, such works mainly focus on replicating a single sound class or a textual sound description, neglecting temporal information, which is crucial in the practical applications of Foley sound. We present T-Foley, a Temporal-event-guided waveform generation model for Foley sound synthesis. T-Foley generates high-quality audio using two conditions: the sound class and temporal event feature. For temporal conditioning, we devise a temporal event feature and a novel conditioning technique named Block-FiLM. T-Foley achieves superior performance in both objective and subjective evaluation metrics and generates Foley sound well-synchronized with the temporal events. Additionally, we showcase T-Foley's practical applications, particularly in scenarios involving vocal mimicry for temporal event control. We show the demo on our companion website.

【2】 A Real-Time Lyrics Alignment System Using Chroma And Phonetic Features  For Classical Vocal Performance
标题:一种基于色度和语音特征的经典声乐实时歌词对齐系统
链接:https://arxiv.org/abs/2401.09200
作者:Jiyun Park,Sangeon Yong,Taegyun Kwon,Juhan Nam
备注:To Appear IEEE ICASSP 2024
摘要:实时歌词对齐的目标是将现场演唱音频作为输入,并在飞行中精确定位给定歌词的确切位置。该任务可以使现实世界的应用程序受益,例如现场音乐会或歌剧的自动字幕。然而,设计一个实时模型提出了一个很大的挑战,由于只使用过去的输入和操作在最小的延迟的约束。此外,由于缺乏用于歌词对齐的实时模型的数据集,以前的研究大多使用私人内部数据集进行评估,导致缺乏标准的评估方法。本文提出了一个实时歌词对齐系统的经典声乐表演有两个贡献。首先,我们改进了歌词对齐算法,找到一个最佳的组合chromagram和语音posteriorgram(PPG),捕捉旋律和语音特征的歌声,分别。其次,我们重铸舒伯特Winterreise数据集(SWD),其中包含多个性能再现相同的作品作为评估集的实时歌词对齐。
摘要:The goal of real-time lyrics alignment is to take live singing audio as input and to pinpoint the exact position within given lyrics on the fly. The task can benefit real-world applications such as the automatic subtitling of live concerts or operas. However, designing a real-time model poses a great challenge due to the constraints of only using past input and operating within a minimal latency. Furthermore, due to the lack of datasets for real-time models for lyrics alignment, previous studies have mostly evaluated with private in-house datasets, resulting in a lack of standard evaluation methods. This paper presents a real-time lyrics alignment system for classical vocal performances with two contributions. First, we improve the lyrics alignment algorithm by finding an optimal combination of chromagram and phonetic posteriorgram (PPG) that capture melodic and phonetics features of the singing voice, respectively. Second, we recast the Schubert Winterreise Dataset (SWD) which contains multiple performance renditions of the same pieces as an evaluation set for the real-time lyrics alignment.


【3】 Efficient Adapter Finetuning for Tail Languages in Streaming  Multilingual ASR
标题:流媒体多语言ASR中尾语言适配器的高效微调
链接:https://arxiv.org/abs/2401.08992
作者:Junwen Bai,Bo Li,Qiujia Li,Tara N. Sainath,Trevor Strohman
备注:Accepted to ICASSP 2024
摘要:在流媒体多语言场景中通常需要端到端ASR模型,因为它更容易部署,并且可以从预先训练的语音模型(如强大的基础模型)中受益。同时,不同语言的异构性和不平衡的数据丰富度可能会导致性能下降,导致不同语言在训练过程中的异步峰值性能,特别是在尾部。有时,由于加强了隐私保护,甚至数据本身也可能变得不可用。现有的工作往往会显着增加模型的大小或学习语言特定的解码器,以分别适应每种语言。在这项研究中,我们探索了简单而有效的依赖于数据库的适配器(LDA)微调下的级联Conformer换能器框架增强教师伪标签的尾语言流多语言ASR。适配器只占每种语言完整模型的0.4%。它被插入到冻结的基础模型中,是微调过程中唯一可训练的模块,具有嘈杂的学生训练。最终的模型合并了来自不同语言的不同检查点的适配器参数。该模型的性能验证了一个具有挑战性的多语言听写数据集,其中包括39尾语言在拉丁语,希腊语,阿拉伯语等,我们提出的方法带来了12.2%的字错误率平均降低,高达37.5%的单一地区。此外,我们表明,我们的参数有效的LDA可以匹配的完整模型微调的质量,从而大大减轻了异步峰值性能的问题。
摘要:The end-to-end ASR model is often desired in the streaming multilingual scenario since it is easier to deploy and can benefit from pre-trained speech models such as powerful foundation models. Meanwhile, the heterogeneous nature and imbalanced data abundance of different languages may cause performance degradation, leading to asynchronous peak performance for different languages during training, especially on tail ones. Sometimes even the data itself may become unavailable as a result of the enhanced privacy protection. Existing work tend to significantly increase the model size or learn language-specific decoders to accommodate each language separately. In this study, we explore simple yet effective Language-Dependent Adapter (LDA) finetuning under a cascaded Conformer transducer framework enhanced by teacher pseudo-labeling for tail languages in the streaming multilingual ASR. The adapter only accounts for 0.4% of the full model per language. It is plugged into the frozen foundation model and is the only trainable module during the finetuning process with noisy student training. The final model merges the adapter parameters from different checkpoints for different languages. The model performance is validated on a challenging multilingual dictation dataset, which includes 39 tail languages across Latin, Greek, Arabic, etc. Our proposed method brings 12.2% word error rate reduction on average and up to 37.5% on a single locale. Furthermore, we show that our parameter-efficient LDA can match the quality of the full model finetuning, thus greatly alleviating the asynchronous peak performance issue.


【4】 DOO-RE: A dataset of ambient sensors in a meeting room for activity  recognition
标题:DOO-RE:会议室环境传感器的数据集,用于活动识别
链接:https://arxiv.org/abs/2401.08962
作者:Hyunju Kim,Geon Kim,Taehoon Lee,Kisoo Kim,Dongman Lee
摘要:随着物联网技术的进步,利用机器学习方法识别用户活动是为用户提供各种智能服务的一种很有前途的方式。具有隐私保护的高质量数据对于在现实世界中部署此类服务至关重要。来自周围环境传感器的数据流非常适合该要求。现有的环境传感器数据集只支持受限的私人空间,而公共空间的环境传感器数据集还有待探索,尽管人们对它们的研究兴趣越来越大。为了满足这一需求,我们构建了一个从配备环境传感器的会议室收集的数据集。数据集DOO-RE包括来自各种环境传感器类型(如声音和投影仪)的数据流。每个传感器数据流被分割成活动单元,多个注释器通过交叉验证注释过程提供活动标签,以提高注释质量。最后,我们得到了9种类型的活动。据我们所知,DOO-RE是第一个支持在真实会议室中识别单个和组活动的数据集,具有可靠的注释。
摘要:With the advancement of IoT technology, recognizing user activities with machine learning methods is a promising way to provide various smart services to users. High-quality data with privacy protection is essential for deploying such services in the real world. Data streams from surrounding ambient sensors are well suited to the requirement. Existing ambient sensor datasets only support constrained private spaces and those for public spaces have yet to be explored despite growing interest in research on them. To meet this need, we build a dataset collected from a meeting room equipped with ambient sensors. The dataset, DOO-RE, includes data streams from various ambient sensor types such as Sound and Projector. Each sensor data stream is segmented into activity units and multiple annotators provide activity labels through a cross-validation annotation process to improve annotation quality. We finally obtain 9 types of activities. To our best knowledge, DOO-RE is the first dataset to support the recognition of both single and group activities in a real meeting room with reliable annotations.


【5】 Similar but Faster: Manipulation of Tempo in Music Audio Embeddings for  Tempo Prediction and Search
标题:相似但速度更快:音乐音频中的节奏操作用于节奏预测和搜索
链接:https://arxiv.org/abs/2401.08902
作者:Matthew C. McCallum,Florian Henkel,Jaehun Kim,Samuel E. Sandberg,Matthew E. P. Davies
备注:Accepted to the International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2024
摘要:音频嵌入使得能够对音频文件的相似性进行大规模比较,以用于诸如搜索和推荐之类的应用。由于音频相似性的主观性,可能期望设计不仅回答音频是否相似,而且回答以何种方式相似(例如,wrt.节奏、情绪或流派)。以前的作品提出了解开嵌入空间,其中表示特定的,但可能相关的,属性的子空间可以被加权,以强调下游任务中的这些属性。然而,没有研究已经进行到这些子空间的独立性,也没有他们的操纵,以检索轨迹相似,但在一个特定的方式不同。在这里,我们探讨嵌入空间的节奏操纵作为实现这一目标的案例研究。我们提出了节奏翻译功能,允许有效地操纵节奏内预先存在的嵌入空间,同时保持其他属性,如流派。由于这种转换是特定于节奏的,因此它使得能够检索相似但具有特定不同节奏的音轨。我们表明,这样的函数可以作为一种有效的数据增强策略,用于下游节奏预测器的训练,以及改进的最近邻检索在很大程度上独立于节奏的属性。
摘要:Audio embeddings enable large scale comparisons of the similarity of audio files for applications such as search and recommendation. Due to the subjectivity of audio similarity, it can be desirable to design systems that answer not only whether audio is similar, but similar in what way (e.g., wrt. tempo, mood or genre). Previous works have proposed disentangled embedding spaces where subspaces representing specific, yet possibly correlated, attributes can be weighted to emphasize those attributes in downstream tasks. However, no research has been conducted into the independence of these subspaces, nor their manipulation, in order to retrieve tracks that are similar but different in a specific way. Here, we explore the manipulation of tempo in embedding spaces as a case-study towards this goal. We propose tempo translation functions that allow for efficient manipulation of tempo within a pre-existing embedding space whilst maintaining other properties such as genre. As this translation is specific to tempo it enables retrieval of tracks that are similar but have specifically different tempi. We show that such a function can be used as an efficient data augmentation strategy for both training of downstream tempo predictors, and improved nearest neighbor retrieval of properties largely independent of tempo.

【6】 Tempo estimation as fully self-supervised binary classification
标题:基于完全自监督二值分类的节奏估计
链接:https://arxiv.org/abs/2401.08891
作者:Florian Henkel,Jaehun Kim,Matthew C. McCallum,Samuel E. Sandberg,Matthew E. P. Davies
备注:Accepted to the International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2024
摘要:本文研究音乐音频中的全局速度估计问题。鉴于注释节奏是耗时的,并且需要一定的音乐专业知识,很少有公开可用的数据源可以为这项任务训练机器学习模型。为了缓解这个问题,我们提出了一种完全自我监督的方法,不依赖于任何人类标记的数据。我们的方法建立在通用(音乐)音频嵌入已经编码了各种属性的事实上,包括关于节奏的信息,使它们很容易适应下游任务。虽然最近的工作在自我监督节奏估计的目的是学习节奏的具体表示,随后用于训练监督分类器,我们重新制定的任务到二进制分类问题的预测目标轨道是否具有相同或不同的节奏相比,参考。虽然前者仍然需要标记的训练数据用于最终的分类模型,但我们的方法使用任意未标记的音乐数据结合时间拉伸进行模型训练,以及一小组综合创建的参考样本用于预测最终的节奏。与最先进的方法相比,我们的方法的评估显示,当找到精确的速度八度音阶的约束放松时,具有很强的竞争力。
摘要:This paper addresses the problem of global tempo estimation in musical audio. Given that annotating tempo is time-consuming and requires certain musical expertise, few publicly available data sources exist to train machine learning models for this task. Towards alleviating this issue, we propose a fully self-supervised approach that does not rely on any human labeled data. Our method builds on the fact that generic (music) audio embeddings already encode a variety of properties, including information about tempo, making them easily adaptable for downstream tasks. While recent work in self-supervised tempo estimation aimed to learn a tempo specific representation that was subsequently used to train a supervised classifier, we reformulate the task into the binary classification problem of predicting whether a target track has the same or a different tempo compared to a reference. While the former still requires labeled training data for the final classification model, our approach uses arbitrary unlabeled music data in combination with time-stretching for model training as well as a small set of synthetically created reference samples for predicting the final tempo. Evaluation of our approach in comparison with the state-of-the-art reveals highly competitive performance when the constraint of finding the precise tempo octave is relaxed.

【7】 On the Effect of Data-Augmentation on Local Embedding Properties in the  Contrastive Learning of Music Audio Representations
标题:音乐音频表征对比学习中数据增强对局部嵌入性的影响
链接:https://arxiv.org/abs/2401.08889
作者:Matthew C. McCallum,Matthew E. P. Davies,Florian Henkel,Jaehun Kim,Samuel E. Sandberg
备注:Accepted to the International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2024
摘要:音频嵌入是理解大型音乐目录的重要工具。通常,嵌入的评估是基于它们在广泛的下游任务中提供的性能,然而很少有研究调查嵌入空间本身的局部属性,这些属性在最近邻算法中很重要,通常用于音乐搜索和推荐。在这项工作中,我们表明,当通过对比学习学习音乐数据集上的音频表示时,通常在曲目内均匀的音乐属性(例如,键和节奏)反映在所得到的嵌入空间中的邻域的局部性中。通过应用适当的数据增强策略,不仅可以减少这些属性的本地化,而且可以增加其他属性的本地化。例如,可以减轻与非专家收听者不太相关的特征(诸如音高和节奏)的局部性,同时改善更显著的特征(诸如流派和情绪)的局部性,从而在最近邻检索准确性方面实现最先进的性能。同样,我们表明,音乐音频嵌入的对比学习的数据增强策略的最佳选择取决于下游任务,强调这是一个重要的嵌入设计决策。
摘要:Audio embeddings are crucial tools in understanding large catalogs of music. Typically embeddings are evaluated on the basis of the performance they provide in a wide range of downstream tasks, however few studies have investigated the local properties of the embedding spaces themselves which are important in nearest neighbor algorithms, commonly used in music search and recommendation. In this work we show that when learning audio representations on music datasets via contrastive learning, musical properties that are typically homogeneous within a track (e.g., key and tempo) are reflected in the locality of neighborhoods in the resulting embedding space. By applying appropriate data augmentation strategies, localisation of such properties can not only be reduced but the localisation of other attributes is increased. For example, locality of features such as pitch and tempo that are less relevant to non-expert listeners, may be mitigated while improving the locality of more salient features such as genre and mood, achieving state-of-the-art performance in nearest neighbor retrieval accuracy. Similarly, we show that the optimal selection of data augmentation strategies for contrastive learning of music audio embeddings is dependent on the downstream task, highlighting this as an important embedding design decision.


【8】 NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant  Meeting Transcription
标题:NOTSOFAR-1挑战:远程会议转录的新数据集、基线和任务
链接:https://arxiv.org/abs/2401.08887
作者:Alon Vinnikov,Amir Ivry,Aviv Hurvitz,Igor Abramovski,Sharon Koubi,Ilya Gurvich,Shai Pe`er,Xiong Xiao,Benjamin Martinez Elizalde,Naoyuki Kanda,Xiaofei Wang,Shalev Shaer,Stav Yagev,Yossi Asher,Sunit Sivasankaran,Yifan Gong,Min Tang,Huaming Wang,Eyal Krupka
备注:preprint
摘要:我们介绍了第一个自然办公室谈话者在设置远场录音(“NOTSOFAR-1”)挑战旁边的数据集和基线系统。该挑战的重点是远场会议场景中的远距离说话者日记和自动语音识别(DASR),具有单通道和已知几何多通道轨道,并作为两个新数据集的发布平台:首先,315次会议的基准数据集,平均每次6分钟,捕获广泛的真实世界声学条件和会话动态。它在30个会议室录制,共有4-8名与会者和35名独特的演讲者。第二,一个1000小时的模拟训练数据集,通过增强真实性进行合成,以实现真实世界的泛化,其中包含15,000个真实的声学传递函数。任务集中在单器件DASR上,其中多通道器件始终共享相同的已知几何结构。这与实际会议室中的常见设置一致,并避免了与多设备任务相关的技术复杂性。它还允许开发特定于几何形状的解决方案。NOTSOFAR-1挑战赛旨在推进远程会话语音识别领域的研究,为释放数据驱动方法的潜力提供关键资源,我们认为这些方法目前受到缺乏全面的高质量训练和基准数据集的限制。
摘要:We introduce the first Natural Office Talkers in Settings of Far-field Audio Recordings (``NOTSOFAR-1'') Challenge alongside datasets and baseline system. The challenge focuses on distant speaker diarization and automatic speech recognition (DASR) in far-field meeting scenarios, with single-channel and known-geometry multi-channel tracks, and serves as a launch platform for two new datasets: First, a benchmarking dataset of 315 meetings, averaging 6 minutes each, capturing a broad spectrum of real-world acoustic conditions and conversational dynamics. It is recorded across 30 conference rooms, featuring 4-8 attendees and a total of 35 unique speakers. Second, a 1000-hour simulated training dataset, synthesized with enhanced authenticity for real-world generalization, incorporating 15,000 real acoustic transfer functions. The tasks focus on single-device DASR, where multi-channel devices always share the same known geometry. This is aligned with common setups in actual conference rooms, and avoids technical complexities associated with multi-device tasks. It also allows for the development of geometry-specific solutions. The NOTSOFAR-1 Challenge aims to advance research in the field of distant conversational speech recognition, providing key resources to unlock the potential of data-driven methods, which we believe are currently constrained by the absence of comprehensive high-quality training and benchmarking datasets.

【9】 Using i-vectors for subject-independent cross-session EEG transfer  learning
标题:利用I向量进行非受试者跨时段脑电信号迁移学习
链接:https://arxiv.org/abs/2401.08851
作者:Jonathan Lasko,Jeff Ma,Mike Nicoletti,Jonathan Sussman-Fort,Sooyoung Jeong,William Hartmann
备注:11 pages
摘要:认知负荷分类是基于诸如脑电图(EEG)的生理测量来自动确定个体在执行任务期间对工作记忆资源的利用的任务。在本文中,我们遵循一种跨学科的方法,使用语音处理的工具和方法来解决这个问题。我们使用的语料库已于2021年公开发布,作为第一届跨会话工作量估计被动脑机接口竞赛的一部分。我们提出了我们的方法,使用基于i向量的神经网络分类器来完成受试者间的跨会话EEG迁移学习,与等效的受试者相关模型相比,实现了18%的相对改善。我们还报告了实验,显示了我们的独立于受试者的模型如何在被拒受试者上表现出竞争力,并通过额外的受试者数据进行改进,这表明有效的认知负荷确定不需要依赖于受试者的训练。
摘要:Cognitive load classification is the task of automatically determining an individual's utilization of working memory resources during performance of a task based on physiologic measures such as electroencephalography (EEG). In this paper, we follow a cross-disciplinary approach, where tools and methodologies from speech processing are used to tackle this problem. The corpus we use was released publicly in 2021 as part of the first passive brain-computer interface competition on cross-session workload estimation. We present our approach which used i-vector-based neural network classifiers to accomplish inter-subject cross-session EEG transfer learning, achieving 18% relative improvement over equivalent subject-dependent models. We also report experiments showing how our subject-independent models perform competitively on held-out subjects and improve with additional subject data, suggesting that subject-dependent training is not required for effective cognitive load determination.


【10】 Robust DOA estimation using deep acoustic imaging
标题:基于深声成像的稳健波达方向估计
链接:https://arxiv.org/abs/2401.08717
作者:Adrian S. Roman,Iran R. Roman,Juan P. Bello
摘要:波达方向估计(DoAE)的目的是在方位角和仰角上跟踪声音。最近的进展包括数据驱动模型,其输入来源于高保真度立体声强度向量或麦克风阵列中通道之间的相关性。球面强度图(SIM)或声图像是一种尚未探索的替代输入表示。SIM受益于高分辨率麦克风阵列,但大多数DoAE数据集使用低分辨率麦克风阵列。因此,我们首先提出了一种超分辨率方法来对低分辨率麦克风进行上采样。接下来,我们对使用SIM作为输入的DoAE模型进行基准测试。我们得到了一个模型,它使用SIM进行DoAE估计,并且优于基线和最先进的模型。我们的研究突出了DoAE任务的声学成像的相关性。
摘要:Direction of arrival estimation (DoAE) aims at tracking a sound in azimuth and elevation. Recent advancements include data-driven models with inputs derived from ambisonics intensity vectors or correlations between channels in a microphone array. A spherical intensity map (SIM), or acoustic image, is an alternative input representation that remains underexplored. SIMs benefit from high-resolution microphone arrays, yet most DoAE datasets use low-resolution ones. Therefore, we first propose a super-resolution method to upsample low-resolution microphones. Next, we benchmark DoAE models that use SIMs as input. We arrive to a model that uses SIMs for DoAE estimation and outperforms a baseline and a state-of-the-art model. Our study highlights the relevance of acoustic imaging for DoAE tasks.


【11】 Transcending Controlled Environments Assessing the Transferability of  ASRRobust NLU Models to Real-World Applications
标题:超越受控环境评估ASR鲁棒NLU模型到实际应用的可移植性
链接:https://arxiv.org/abs/2401.09354
作者:Hania Khan,Aleena Fatima Khalid,Zaryab Hassan
摘要:本研究探讨了自动语音识别(ASR)-鲁棒自然语言理解(NLU)模型从受控实验条件到实际现实应用的可转移性。该研究专注于乌尔都语的智能家居自动化命令,评估了不同噪声配置文件,语言变化和ASR错误场景下的模型性能。利用UrduBERT模型,该研究采用了一种系统的方法,涉及真实世界的数据收集,交叉验证,迁移学习,噪声变化研究和领域适应。评估指标包括特定于任务的准确性、延迟、用户满意度和对ASR错误的鲁棒性。这些发现有助于深入了解ASR鲁棒NLU模型在超越受控环境中的挑战和适应性。
摘要:This research investigates the transferability of Automatic Speech Recognition (ASR)-robust Natural Language Understanding (NLU) models from controlled experimental conditions to practical, real-world applications. Focused on smart home automation commands in Urdu, the study assesses model performance under diverse noise profiles, linguistic variations, and ASR error scenarios. Leveraging the UrduBERT model, the research employs a systematic methodology involving real-world data collection, cross-validation, transfer learning, noise variation studies, and domain adaptation. Evaluation metrics encompass task-specific accuracy, latency, user satisfaction, and robustness to ASR errors. The findings contribute insights into the challenges and adaptability of ASR-robust NLU models in transcending controlled environments.


【12】 Can Synthetic Data Boost the Training of Deep Acoustic Vehicle Counting  Networks?
标题:合成数据能促进深度声学车辆计数网络的训练吗?
链接:https://arxiv.org/abs/2401.09308
作者:Stefano Damiano,Luca Bondi,Shabnam Ghaffarzadegan,Andre Guntoro,Toon van Waterschoot备注:Accepted paper: 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2024)
摘要:在优化城市交通基础设施的交通监控解决方案的设计中,声学车辆计数模型因其成本效益和能源效率而受到关注。虽然深度学习已被证明对视觉交通监控有效,但其在音频领域的使用尚未得到彻底研究,这可能是由于现实世界的数据稀缺。在这项工作中,我们提出了一种新的声学车辆计数方法,通过开发:i)一个交通噪声模拟框架,以合成现实的车辆经过事件; ii)一种策略,将合成数据和真实数据混合起来,训练深度学习模型,用于交通计数。所提出的系统能够同时计数在双车道道路上行驶的汽车和商用车辆,并在中等交通密度条件下识别它们的行驶方向。只有24小时的标记现实世界的交通噪音,我们能够提高计数准确性的现实世界的数据从63\%$到88\%$的汽车和从86\%$到94\%$的商用车。
摘要:In the design of traffic monitoring solutions for optimizing the urban mobility infrastructure, acoustic vehicle counting models have received attention due to their cost effectiveness and energy efficiency. Although deep learning has proven effective for visual traffic monitoring, its use has not been thoroughly investigated in the audio domain, likely due to real-world data scarcity. In this work, we propose a novel approach to acoustic vehicle counting by developing: i) a traffic noise simulation framework to synthesize realistic vehicle pass-by events; ii) a strategy to mix synthetic and real data to train a deep-learning model for traffic counting. The proposed system is capable of simultaneously counting cars and commercial vehicles driving on a two-lane road, and identifying their direction of travel under moderate traffic density conditions. With only 24 hours of labeled real-world traffic noise, we are able to improve counting accuracy on real-world data from $63\%$ to $88\%$ for cars and from $86\%$ to $94\%$ for commercial vehicles.


【13】 Two-pass Endpoint Detection for Speech Recognition
标题:语音识别中的两遍端点检测
链接:https://arxiv.org/abs/2401.08916
作者:Anirudh Raju,Aparna Khare,Di He,Ilya Sklyar,Long Chen,Sam Alptekin,Viet Anh Trinh,Zhe Zhang,Colin Vaz,Venkatesh Ravichandran,Roland Maas,Ariya Rastrow
备注:ASRU 2023
摘要:端点(EP)检测是远场语音识别系统的关键组成部分,通过语音命令帮助用户。端点检测器必须在准确性和延迟之间进行权衡,因为等待更长时间会减少用户被提前切断的情况。我们提出了一种新的两遍端点定位的解决方案,其中从第一遍端点检测到的话语端点由第二遍模型验证,称为EP消除器。我们的方法改善了早期截止和延迟之间的权衡基线端点,测试数据集,包括语音助理事务查询,会话语音,和公共SLURP语料库。我们证明,我们的方法显示出改善,无论第一次通过EP模型使用。
摘要:Endpoint (EP) detection is a key component of far-field speech recognition systems that assist the user through voice commands. The endpoint detector has to trade-off between accuracy and latency, since waiting longer reduces the cases of users being cut-off early. We propose a novel two-pass solution for endpointing, where the utterance endpoint detected from a first pass endpointer is verified by a 2nd-pass model termed EP Arbitrator. Our method improves the trade-off between early cut-offs and latency over a baseline endpointer, as tested on datasets including voice-assistant transactional queries, conversational speech, and the public SLURP corpus. We demonstrate that our method shows improvements regardless of the first-pass EP model used.

【14】 Binaural Angular Separation Network
标题:双耳角度分离网络
链接:https://arxiv.org/abs/2401.08864
作者:Yang Yang,George Sung,Shao-Fu Shih,Hakan Erdogan,Chehung Lee,Matthias Grundmann
备注:Accepted to ICASSP 2024
摘要:我们提出了一个神经网络模型,可以分离目标语音源从干扰源在不同的角度区域使用两个麦克风。该模型使用全向麦克风使用模拟的房间脉冲响应(RIR)进行训练,而无需收集真实的RIR。通过依赖于特定的角度区域和多个房间模拟,该模型利用一致的到达时间差(TDOA)线索,或我们所说的延迟对比,以分离目标和干扰源,同时在各种混响环境中保持稳健。我们证明了该模型不仅可以推广到具有略微不同的麦克风几何形状的市售设备,而且优于我们以前在同一设备上使用一个额外麦克风的工作。该模型在设备上实时运行,适用于电话和视频会议等低延迟流媒体应用。
摘要:We propose a neural network model that can separate target speech sources from interfering sources at different angular regions using two microphones. The model is trained with simulated room impulse responses (RIRs) using omni-directional microphones without needing to collect real RIRs. By relying on specific angular regions and multiple room simulations, the model utilizes consistent time difference of arrival (TDOA) cues, or what we call delay contrast, to separate target and interference sources while remaining robust in various reverberation environments. We demonstrate the model is not only generalizable to a commercially available device with a slightly different microphone geometry, but also outperforms our previous work which uses one additional microphone on the same device. The model runs in real-time on-device and is suitable for low-latency streaming applications such as telephony and video conferencing.

【15】 Revisiting Self-supervised Learning of Speech Representation from a  Mutual Information Perspective
标题:从互信息角度重新审视语音表征的自监督学习
链接:https://arxiv.org/abs/2401.08833
作者:Alexander H. Liu,Sung-Lin Yeh,James Glass
备注:ICASSP 2024
摘要:现有的自监督语音表示学习的研究主要集中在开发新的训练方法和将预训练模型应用于不同的应用。然而,这些模型的质量通常由不同下游任务的性能来衡量。表征如何访问感兴趣的信息的研究较少。在这项工作中,我们从信息理论的角度仔细研究了现有的语音自监督方法。我们的目标是使用互信息来开发度量标准,以帮助解决模型设计和选择等实际问题。我们使用线性探针来估计目标信息和学习表征之间的互信息,显示从语音表征到目标信息的可访问性的另一种见解。此外,我们还探索了以自我监督的方式评估表示的潜力,在这种方式中,我们在不使用任何标签的情况下估计数据不同部分之间的互信息。最后,我们表明,监督和无监督的措施回声层的线性探测和语音识别模型的性能。
摘要:Existing studies on self-supervised speech representation learning have focused on developing new training methods and applying pre-trained models for different applications. However, the quality of these models is often measured by the performance of different downstream tasks. How well the representations access the information of interest is less studied. In this work, we take a closer look into existing self-supervised methods of speech from an information-theoretic perspective. We aim to develop metrics using mutual information to help practical problems such as model design and selection. We use linear probes to estimate the mutual information between the target information and learned representations, showing another insight into the accessibility to the target information from speech representations. Further, we explore the potential of evaluating representations in a self-supervised fashion, where we estimate the mutual information between different parts of the data without using any labels. Finally, we show that both supervised and unsupervised measures echo the performance of the models on layer-wise linear probing and speech recognition.


【16】 Sub-band and Full-band Interactive U-Net with DPRNN for Demixing  Cross-talk Stereo Music
标题:基于DPRNN的子带和全带交互U-Net用于分离相声立体声音乐
链接:https://arxiv.org/abs/2401.08678
作者:Han Yin,Mou Wang,Jisheng Bai,Dongyuan Shi,Woon-Seng Gan,Jianfeng Chen
备注:Submitted to ICASSP 2024
摘要:本文详细介绍了我们提出的ICASSP 2024华彩乐段挑战赛的方法。实验结果表明,该系统可以实现更好的性能比官方基线。
摘要:This paper presents a detailed description of our proposed methods for the ICASSP 2024 Cadenza Challenge. Experimental results show that the proposed system can achieve better performance than official baselines.

eess.AS音频处理
【1】 Transcending Controlled Environments Assessing the Transferability of  ASRRobust NLU Models to Real-World Applications
标题:超越受控环境评估ASRRobust NLU模型向现实应用程序的可转移性
链接:https://arxiv.org/abs/2401.09354
作者:Hania Khan,Aleena Fatima Khalid,Zaryab Hassan
摘要:本研究探讨了自动语音识别(ASR)-鲁棒自然语言理解(NLU)模型从受控实验条件到实际现实应用的可转移性。该研究专注于乌尔都语的智能家居自动化命令,评估了不同噪声配置文件,语言变化和ASR错误场景下的模型性能。利用UrduBERT模型,该研究采用了一种系统的方法,涉及真实世界的数据收集,交叉验证,迁移学习,噪声变化研究和领域适应。评估指标包括特定于任务的准确性、延迟、用户满意度和对ASR错误的鲁棒性。这些发现有助于深入了解ASR鲁棒NLU模型在超越受控环境中的挑战和适应性。
摘要:This research investigates the transferability of Automatic Speech Recognition (ASR)-robust Natural Language Understanding (NLU) models from controlled experimental conditions to practical, real-world applications. Focused on smart home automation commands in Urdu, the study assesses model performance under diverse noise profiles, linguistic variations, and ASR error scenarios. Leveraging the UrduBERT model, the research employs a systematic methodology involving real-world data collection, cross-validation, transfer learning, noise variation studies, and domain adaptation. Evaluation metrics encompass task-specific accuracy, latency, user satisfaction, and robustness to ASR errors. The findings contribute insights into the challenges and adaptability of ASR-robust NLU models in transcending controlled environments.


【2】 On Speech Pre-emphasis as a Simple and Inexpensive Method to Boost  Speech Enhancement标题:语音预加重作为一种简单而廉价的语音增强方法
链接:https://arxiv.org/abs/2401.09315
作者:Iván López-Espejo,Aditya Joglekar,Antonio M. Peinado,Jesper Jensen
摘要:多年来,补偿语音在较高频率处的自然能量衰减的预加重滤波已被认为是许多语音处理任务中的常见预处理步骤。在这项工作中,我们首次证明了预加重滤波也可以用作一种简单且计算成本低廉的方式来利用基于深度神经网络的语音增强性能。特别是,我们研究在损失计算之前预强调估计的和实际的干净语音,以便不同的语音频率分量在训练阶段更好地反映它们的感知重要性。在TIMIT数据集的噪声版本上的实验结果表明,在训练阶段,对于可见和不可见的噪声类型,集成基于预加重的方法可以分别获得高达4.6%和3.4%的相对估计语音质量改进。类似的情况下,预加重被认为是一个默认的预处理步骤,在经典的自动语音识别和语音编码系统,本文中分析的基于预加重的方法可能成为一个默认的附加现代语音增强。
摘要:Pre-emphasis filtering, compensating for the natural energy decay of speech at higher frequencies, has been considered as a common pre-processing step in a number of speech processing tasks over the years. In this work, we demonstrate, for the first time, that pre-emphasis filtering may also be used as a simple and computationally-inexpensive way to leverage deep neural network-based speech enhancement performance. Particularly, we look into pre-emphasizing the estimated and actual clean speech prior to loss calculation so that different speech frequency components better mirror their perceptual importance during the training phase. Experimental results on a noisy version of the TIMIT dataset show that integrating the pre-emphasis-based methodology at hand yields relative estimated speech quality improvements of up to 4.6% and 3.4% for noise types seen and unseen, respectively, during the training phase. Similar to the case of pre-emphasis being considered as a default pre-processing step in classical automatic speech recognition and speech coding systems, the pre-emphasis-based methodology analyzed in this article may potentially become a default add-on for modern speech enhancement.

【3】 Can Synthetic Data Boost the Training of Deep Acoustic Vehicle Counting  Networks?
标题:合成数据能促进深声车辆计数网络的训练吗?
链接:https://arxiv.org/abs/2401.09308
作者:Stefano Damiano,Luca Bondi,Shabnam Ghaffarzadegan,Andre Guntoro,Toon van Waterschoot备注:Accepted paper: 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2024)
摘要:在优化城市交通基础设施的交通监控解决方案的设计中,声学车辆计数模型因其成本效益和能源效率而受到关注。虽然深度学习已被证明对视觉交通监控有效,但其在音频领域的使用尚未得到彻底研究,这可能是由于现实世界的数据稀缺。在这项工作中,我们提出了一种新的声学车辆计数方法,通过开发:i)一个交通噪声模拟框架,以合成现实的车辆经过事件; ii)一种策略,将合成数据和真实数据混合起来,训练深度学习模型,用于交通计数。所提出的系统能够同时计数在双车道道路上行驶的汽车和商用车辆,并在中等交通密度条件下识别它们的行驶方向。只有24小时的标记现实世界的交通噪音,我们能够提高计数准确性的现实世界的数据从63\%$到88\%$的汽车和从86\%$到94\%$的商用车。
摘要:In the design of traffic monitoring solutions for optimizing the urban mobility infrastructure, acoustic vehicle counting models have received attention due to their cost effectiveness and energy efficiency. Although deep learning has proven effective for visual traffic monitoring, its use has not been thoroughly investigated in the audio domain, likely due to real-world data scarcity. In this work, we propose a novel approach to acoustic vehicle counting by developing: i) a traffic noise simulation framework to synthesize realistic vehicle pass-by events; ii) a strategy to mix synthetic and real data to train a deep-learning model for traffic counting. The proposed system is capable of simultaneously counting cars and commercial vehicles driving on a two-lane road, and identifying their direction of travel under moderate traffic density conditions. With only 24 hours of labeled real-world traffic noise, we are able to improve counting accuracy on real-world data from $63\%$ to $88\%$ for cars and from $86\%$ to $94\%$ for commercial vehicles.


【4】 Two-pass Endpoint Detection for Speech Recognition
标题:语音识别中的两遍端点检测
链接:https://arxiv.org/abs/2401.08916
作者:Anirudh Raju,Aparna Khare,Di He,Ilya Sklyar,Long Chen,Sam Alptekin,Viet Anh Trinh,Zhe Zhang,Colin Vaz,Venkatesh Ravichandran,Roland Maas,Ariya Rastrow
备注:ASRU 2023
摘要:端点(EP)检测是远场语音识别系统的关键组成部分,通过语音命令帮助用户。端点检测器必须在准确性和延迟之间进行权衡,因为等待更长时间会减少用户被提前切断的情况。我们提出了一种新的两遍端点定位的解决方案,其中从第一遍端点检测到的话语端点由第二遍模型验证,称为EP消除器。我们的方法改善了早期截止和延迟之间的权衡基线端点,测试数据集,包括语音助理事务查询,会话语音,和公共SLURP语料库。我们证明,我们的方法显示出改善,无论第一次通过EP模型使用。
摘要:Endpoint (EP) detection is a key component of far-field speech recognition systems that assist the user through voice commands. The endpoint detector has to trade-off between accuracy and latency, since waiting longer reduces the cases of users being cut-off early. We propose a novel two-pass solution for endpointing, where the utterance endpoint detected from a first pass endpointer is verified by a 2nd-pass model termed EP Arbitrator. Our method improves the trade-off between early cut-offs and latency over a baseline endpointer, as tested on datasets including voice-assistant transactional queries, conversational speech, and the public SLURP corpus. We demonstrate that our method shows improvements regardless of the first-pass EP model used.

【5】 Binaural Angular Separation Network
标题:双耳角度分离网络
链接:https://arxiv.org/abs/2401.08864
作者:Yang Yang,George Sung,Shao-Fu Shih,Hakan Erdogan,Chehung Lee,Matthias Grundmann
备注:Accepted to ICASSP 2024
摘要:我们提出了一个神经网络模型,可以分离目标语音源从干扰源在不同的角度区域使用两个麦克风。该模型使用全向麦克风使用模拟的房间脉冲响应(RIR)进行训练,而无需收集真实的RIR。通过依赖于特定的角度区域和多个房间模拟,该模型利用一致的到达时间差(TDOA)线索,或我们所说的延迟对比,以分离目标和干扰源,同时在各种混响环境中保持稳健。我们证明了该模型不仅可以推广到具有略微不同的麦克风几何形状的市售设备,而且优于我们以前在同一设备上使用一个额外麦克风的工作。该模型在设备上实时运行,适用于电话和视频会议等低延迟流媒体应用。
摘要:We propose a neural network model that can separate target speech sources from interfering sources at different angular regions using two microphones. The model is trained with simulated room impulse responses (RIRs) using omni-directional microphones without needing to collect real RIRs. By relying on specific angular regions and multiple room simulations, the model utilizes consistent time difference of arrival (TDOA) cues, or what we call delay contrast, to separate target and interference sources while remaining robust in various reverberation environments. We demonstrate the model is not only generalizable to a commercially available device with a slightly different microphone geometry, but also outperforms our previous work which uses one additional microphone on the same device. The model runs in real-time on-device and is suitable for low-latency streaming applications such as telephony and video conferencing.

【6】 Revisiting Self-supervised Learning of Speech Representation from a  Mutual Information Perspective
标题:从互信息角度重新审视语音表征的自监督学习
链接:https://arxiv.org/abs/2401.08833
作者:Alexander H. Liu,Sung-Lin Yeh,James Glass
备注:ICASSP 2024
摘要:现有的自监督语音表示学习的研究主要集中在开发新的训练方法和将预训练模型应用于不同的应用。然而,这些模型的质量通常由不同下游任务的性能来衡量。表征如何访问感兴趣的信息的研究较少。在这项工作中,我们从信息理论的角度仔细研究了现有的语音自监督方法。我们的目标是使用互信息来开发度量标准,以帮助解决模型设计和选择等实际问题。我们使用线性探针来估计目标信息和学习表征之间的互信息,显示从语音表征到目标信息的可访问性的另一种见解。此外,我们还探索了以自我监督的方式评估表示的潜力,在这种方式中,我们在不使用任何标签的情况下估计数据不同部分之间的互信息。最后,我们表明,监督和无监督的措施回声层的线性探测和语音识别模型的性能。
摘要:Existing studies on self-supervised speech representation learning have focused on developing new training methods and applying pre-trained models for different applications. However, the quality of these models is often measured by the performance of different downstream tasks. How well the representations access the information of interest is less studied. In this work, we take a closer look into existing self-supervised methods of speech from an information-theoretic perspective. We aim to develop metrics using mutual information to help practical problems such as model design and selection. We use linear probes to estimate the mutual information between the target information and learned representations, showing another insight into the accessibility to the target information from speech representations. Further, we explore the potential of evaluating representations in a self-supervised fashion, where we estimate the mutual information between different parts of the data without using any labels. Finally, we show that both supervised and unsupervised measures echo the performance of the models on layer-wise linear probing and speech recognition.


【7】 Sub-band and Full-band Interactive U-Net with DPRNN for Demixing  Cross-talk Stereo Music
标题:基于DPRNN的子带和全带交互U-Net用于分离相声立体声音乐
链接:https://arxiv.org/abs/2401.08678

作者:Han Yin,Mou Wang,Jisheng Bai,Dongyuan Shi,Woon-Seng Gan,Jianfeng Chen

备注:Submitted to ICASSP 2024

摘要:本文详细介绍了我们提出的ICASSP 2024华彩乐段挑战赛的方法。实验结果表明,该系统可以实现更好的性能比官方基线。

摘要:This paper presents a detailed description of our proposed methods for the ICASSP 2024 Cadenza Challenge. Experimental results show that the proposed system can achieve better performance than official baselines.


【8】 T-FOLEY: A Controllable Waveform-Domain Diffusion Model for  Temporal-Event-Guided Foley Sound Synthesis
标题:T-Foley:时间事件引导的Foley声音合成的可控波形域扩散模型
链接:https://arxiv.org/abs/2401.09294
作者:Yoonjin Chung,Junwon Lee,Juhan Nam
摘要:Foley sound是与视频同步插入的音频内容,在多媒体内容的用户体验中起着至关重要的作用。最近,人们对Foley声音合成进行了积极的研究,利用了深度生成模型的进步。然而,这些工作主要集中在复制一个单一的声音类或文本的声音描述,忽略了时间信息,这是至关重要的福利声音的实际应用。我们提出了T-Foley,一个用于Foley声音合成的时间事件引导波形生成模型。T-Foley使用两个条件生成高质量的音频:声音类别和时间事件特征。对于时间条件反射,我们设计了一个时间事件功能和一个新的条件反射技术,名为块电影。T-Foley在客观和主观评估指标方面都取得了优异的性能,并生成与时间事件同步的Foley声音。此外,我们展示T-福利的实际应用,特别是在涉及语音模仿的时间事件控制的情况下。我们在我们的配套网站上展示演示。
摘要:Foley sound, audio content inserted synchronously with videos, plays a critical role in the user experience of multimedia content. Recently, there has been active research in Foley sound synthesis, leveraging the advancements in deep generative models. However, such works mainly focus on replicating a single sound class or a textual sound description, neglecting temporal information, which is crucial in the practical applications of Foley sound. We present T-Foley, a Temporal-event-guided waveform generation model for Foley sound synthesis. T-Foley generates high-quality audio using two conditions: the sound class and temporal event feature. For temporal conditioning, we devise a temporal event feature and a novel conditioning technique named Block-FiLM. T-Foley achieves superior performance in both objective and subjective evaluation metrics and generates Foley sound well-synchronized with the temporal events. Additionally, we showcase T-Foley's practical applications, particularly in scenarios involving vocal mimicry for temporal event control. We show the demo on our companion website.

【9】 A Real-Time Lyrics Alignment System Using Chroma And Phonetic Features  For Classical Vocal Performance
标题:一种基于色度和语音特征的经典声乐实时歌词对齐系统
链接:https://arxiv.org/abs/2401.09200
作者:Jiyun Park,Sangeon Yong,Taegyun Kwon,Juhan Nam
备注:To Appear IEEE ICASSP 2024
摘要:实时歌词对齐的目标是将现场演唱音频作为输入,并在飞行中精确定位给定歌词的确切位置。该任务可以使现实世界的应用程序受益,例如现场音乐会或歌剧的自动字幕。然而,设计一个实时模型提出了一个很大的挑战,由于只使用过去的输入和操作在最小的延迟的约束。此外,由于缺乏用于歌词对齐的实时模型的数据集,以前的研究大多使用私人内部数据集进行评估,导致缺乏标准的评估方法。本文提出了一个实时歌词对齐系统的经典声乐表演有两个贡献。首先,我们改进了歌词对齐算法,找到一个最佳的组合chromagram和语音posteriorgram(PPG),捕捉旋律和语音特征的歌声,分别。其次,我们重铸舒伯特Winterreise数据集(SWD),其中包含多个性能再现相同的作品作为评估集的实时歌词对齐。
摘要:The goal of real-time lyrics alignment is to take live singing audio as input and to pinpoint the exact position within given lyrics on the fly. The task can benefit real-world applications such as the automatic subtitling of live concerts or operas. However, designing a real-time model poses a great challenge due to the constraints of only using past input and operating within a minimal latency. Furthermore, due to the lack of datasets for real-time models for lyrics alignment, previous studies have mostly evaluated with private in-house datasets, resulting in a lack of standard evaluation methods. This paper presents a real-time lyrics alignment system for classical vocal performances with two contributions. First, we improve the lyrics alignment algorithm by finding an optimal combination of chromagram and phonetic posteriorgram (PPG) that capture melodic and phonetics features of the singing voice, respectively. Second, we recast the Schubert Winterreise Dataset (SWD) which contains multiple performance renditions of the same pieces as an evaluation set for the real-time lyrics alignment.

【10】 Efficient Adapter Finetuning for Tail Languages in Streaming  Multilingual ASR
标题:流媒体多语言ASR中尾语言的高效适配器精调
链接:https://arxiv.org/abs/2401.08992
作者:Junwen Bai,Bo Li,Qiujia Li,Tara N. Sainath,Trevor Strohman
备注:Accepted to ICASSP 2024
摘要:在流媒体多语言场景中通常需要端到端ASR模型,因为它更容易部署,并且可以从预先训练的语音模型(如强大的基础模型)中受益。同时,不同语言的异构性和不平衡的数据丰富度可能会导致性能下降,导致不同语言在训练过程中的异步峰值性能,特别是在尾部。有时,由于加强了隐私保护,甚至数据本身也可能变得不可用。现有的工作往往会显着增加模型的大小或学习语言特定的解码器,以分别适应每种语言。在这项研究中,我们探索了简单而有效的依赖于数据库的适配器(LDA)微调下的级联Conformer换能器框架增强教师伪标签的尾语言流多语言ASR。适配器只占每种语言完整模型的0.4%。它被插入到冻结的基础模型中,是微调过程中唯一可训练的模块,具有嘈杂的学生训练。最终的模型合并了来自不同语言的不同检查点的适配器参数。该模型的性能验证了一个具有挑战性的多语言听写数据集,其中包括39尾语言在拉丁语,希腊语,阿拉伯语等,我们提出的方法带来了12.2%的字错误率平均降低,高达37.5%的单一地区。此外,我们表明,我们的参数有效的LDA可以匹配的完整模型微调的质量,从而大大减轻了异步峰值性能的问题。
摘要:The end-to-end ASR model is often desired in the streaming multilingual scenario since it is easier to deploy and can benefit from pre-trained speech models such as powerful foundation models. Meanwhile, the heterogeneous nature and imbalanced data abundance of different languages may cause performance degradation, leading to asynchronous peak performance for different languages during training, especially on tail ones. Sometimes even the data itself may become unavailable as a result of the enhanced privacy protection. Existing work tend to significantly increase the model size or learn language-specific decoders to accommodate each language separately. In this study, we explore simple yet effective Language-Dependent Adapter (LDA) finetuning under a cascaded Conformer transducer framework enhanced by teacher pseudo-labeling for tail languages in the streaming multilingual ASR. The adapter only accounts for 0.4% of the full model per language. It is plugged into the frozen foundation model and is the only trainable module during the finetuning process with noisy student training. The final model merges the adapter parameters from different checkpoints for different languages. The model performance is validated on a challenging multilingual dictation dataset, which includes 39 tail languages across Latin, Greek, Arabic, etc. Our proposed method brings 12.2% word error rate reduction on average and up to 37.5% on a single locale. Furthermore, we show that our parameter-efficient LDA can match the quality of the full model finetuning, thus greatly alleviating the asynchronous peak performance issue.

【11】 DOO-RE: A dataset of ambient sensors in a meeting room for activity  recognition
标题:DOO-RE:会议室环境传感器的数据集,用于活动识别
链接:https://arxiv.org/abs/2401.08962
作者:Hyunju Kim,Geon Kim,Taehoon Lee,Kisoo Kim,Dongman Lee
摘要:随着物联网技术的进步,利用机器学习方法识别用户活动是为用户提供各种智能服务的一种很有前途的方式。具有隐私保护的高质量数据对于在现实世界中部署此类服务至关重要。来自周围环境传感器的数据流非常适合该要求。现有的环境传感器数据集只支持受限的私人空间,而公共空间的环境传感器数据集还有待探索,尽管人们对它们的研究兴趣越来越大。为了满足这一需求,我们构建了一个从配备环境传感器的会议室收集的数据集。数据集DOO-RE包括来自各种环境传感器类型(如声音和投影仪)的数据流。每个传感器数据流被分割成活动单元,多个注释器通过交叉验证注释过程提供活动标签,以提高注释质量。最后,我们得到了9种类型的活动。据我们所知,DOO-RE是第一个支持在真实会议室中识别单个和组活动的数据集,具有可靠的注释。
摘要:With the advancement of IoT technology, recognizing user activities with machine learning methods is a promising way to provide various smart services to users. High-quality data with privacy protection is essential for deploying such services in the real world. Data streams from surrounding ambient sensors are well suited to the requirement. Existing ambient sensor datasets only support constrained private spaces and those for public spaces have yet to be explored despite growing interest in research on them. To meet this need, we build a dataset collected from a meeting room equipped with ambient sensors. The dataset, DOO-RE, includes data streams from various ambient sensor types such as Sound and Projector. Each sensor data stream is segmented into activity units and multiple annotators provide activity labels through a cross-validation annotation process to improve annotation quality. We finally obtain 9 types of activities. To our best knowledge, DOO-RE is the first dataset to support the recognition of both single and group activities in a real meeting room with reliable annotations.

【12】 Similar but Faster: Manipulation of Tempo in Music Audio Embeddings for  Tempo Prediction and Search
标题:相似但速度更快:音乐音频中的节奏操作用于节奏预测和搜索
链接:https://arxiv.org/abs/2401.08902
作者:Matthew C. McCallum,Florian Henkel,Jaehun Kim,Samuel E. Sandberg,Matthew E. P. Davies
备注:Accepted to the International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2024
摘要:音频嵌入使得能够对音频文件的相似性进行大规模比较,以用于诸如搜索和推荐之类的应用。由于音频相似性的主观性,可能期望设计不仅回答音频是否相似,而且回答以何种方式相似(例如,wrt.节奏、情绪或流派)。以前的作品提出了解开嵌入空间,其中表示特定的,但可能相关的,属性的子空间可以被加权,以强调下游任务中的这些属性。然而,没有研究已经进行到这些子空间的独立性,也没有他们的操纵,以检索轨迹相似,但在一个特定的方式不同。在这里,我们探讨嵌入空间的节奏操纵作为实现这一目标的案例研究。我们提出了节奏翻译功能,允许有效地操纵节奏内预先存在的嵌入空间,同时保持其他属性,如流派。由于这种转换是特定于节奏的,因此它使得能够检索相似但具有特定不同节奏的音轨。我们表明,这样的函数可以作为一种有效的数据增强策略,用于下游节奏预测器的训练,以及改进的最近邻检索在很大程度上独立于节奏的属性。
摘要:Audio embeddings enable large scale comparisons of the similarity of audio files for applications such as search and recommendation. Due to the subjectivity of audio similarity, it can be desirable to design systems that answer not only whether audio is similar, but similar in what way (e.g., wrt. tempo, mood or genre). Previous works have proposed disentangled embedding spaces where subspaces representing specific, yet possibly correlated, attributes can be weighted to emphasize those attributes in downstream tasks. However, no research has been conducted into the independence of these subspaces, nor their manipulation, in order to retrieve tracks that are similar but different in a specific way. Here, we explore the manipulation of tempo in embedding spaces as a case-study towards this goal. We propose tempo translation functions that allow for efficient manipulation of tempo within a pre-existing embedding space whilst maintaining other properties such as genre. As this translation is specific to tempo it enables retrieval of tracks that are similar but have specifically different tempi. We show that such a function can be used as an efficient data augmentation strategy for both training of downstream tempo predictors, and improved nearest neighbor retrieval of properties largely independent of tempo.

【13】 Tempo estimation as fully self-supervised binary classification
标题:基于完全自监督二值分类的节奏估计
链接:https://arxiv.org/abs/2401.08891
作者:Florian Henkel,Jaehun Kim,Matthew C. McCallum,Samuel E. Sandberg,Matthew E. P. Davies
备注:Accepted to the International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2024
摘要:本文研究音乐音频中的全局速度估计问题。鉴于注释节奏是耗时的,并且需要一定的音乐专业知识,很少有公开可用的数据源可以为这项任务训练机器学习模型。为了缓解这个问题,我们提出了一种完全自我监督的方法,不依赖于任何人类标记的数据。我们的方法建立在通用(音乐)音频嵌入已经编码了各种属性的事实上,包括关于节奏的信息,使它们很容易适应下游任务。虽然最近的工作在自我监督节奏估计的目的是学习节奏的具体表示,随后用于训练监督分类器,我们重新制定的任务到二进制分类问题的预测目标轨道是否具有相同或不同的节奏相比,参考。虽然前者仍然需要标记的训练数据用于最终的分类模型,但我们的方法使用任意未标记的音乐数据结合时间拉伸进行模型训练,以及一小组综合创建的参考样本用于预测最终的节奏。与最先进的方法相比,我们的方法的评估显示,当找到精确的速度八度音阶的约束放松时,具有很强的竞争力。
摘要:This paper addresses the problem of global tempo estimation in musical audio. Given that annotating tempo is time-consuming and requires certain musical expertise, few publicly available data sources exist to train machine learning models for this task. Towards alleviating this issue, we propose a fully self-supervised approach that does not rely on any human labeled data. Our method builds on the fact that generic (music) audio embeddings already encode a variety of properties, including information about tempo, making them easily adaptable for downstream tasks. While recent work in self-supervised tempo estimation aimed to learn a tempo specific representation that was subsequently used to train a supervised classifier, we reformulate the task into the binary classification problem of predicting whether a target track has the same or a different tempo compared to a reference. While the former still requires labeled training data for the final classification model, our approach uses arbitrary unlabeled music data in combination with time-stretching for model training as well as a small set of synthetically created reference samples for predicting the final tempo. Evaluation of our approach in comparison with the state-of-the-art reveals highly competitive performance when the constraint of finding the precise tempo octave is relaxed.

【14】 On the Effect of Data-Augmentation on Local Embedding Properties in the  Contrastive Learning of Music Audio Representations
标题:音乐音频表征对比学习中数据增强对局部嵌入性的影响
链接:https://arxiv.org/abs/2401.08889
作者:Matthew C. McCallum,Matthew E. P. Davies,Florian Henkel,Jaehun Kim,Samuel E. Sandberg
备注:Accepted to the International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2024
摘要:音频嵌入是理解大型音乐目录的重要工具。通常,嵌入的评估是基于它们在广泛的下游任务中提供的性能,然而很少有研究调查嵌入空间本身的局部属性,这些属性在最近邻算法中很重要,通常用于音乐搜索和推荐。在这项工作中,我们表明,当通过对比学习学习音乐数据集上的音频表示时,通常在曲目内均匀的音乐属性(例如,键和节奏)反映在所得到的嵌入空间中的邻域的局部性中。通过应用适当的数据增强策略,不仅可以减少这些属性的本地化,而且可以增加其他属性的本地化。例如,可以减轻与非专家收听者不太相关的特征(诸如音高和节奏)的局部性,同时改善更显著的特征(诸如流派和情绪)的局部性,从而在最近邻检索准确性方面实现最先进的性能。同样,我们表明,音乐音频嵌入的对比学习的数据增强策略的最佳选择取决于下游任务,强调这是一个重要的嵌入设计决策。
摘要:Audio embeddings are crucial tools in understanding large catalogs of music. Typically embeddings are evaluated on the basis of the performance they provide in a wide range of downstream tasks, however few studies have investigated the local properties of the embedding spaces themselves which are important in nearest neighbor algorithms, commonly used in music search and recommendation. In this work we show that when learning audio representations on music datasets via contrastive learning, musical properties that are typically homogeneous within a track (e.g., key and tempo) are reflected in the locality of neighborhoods in the resulting embedding space. By applying appropriate data augmentation strategies, localisation of such properties can not only be reduced but the localisation of other attributes is increased. For example, locality of features such as pitch and tempo that are less relevant to non-expert listeners, may be mitigated while improving the locality of more salient features such as genre and mood, achieving state-of-the-art performance in nearest neighbor retrieval accuracy. Similarly, we show that the optimal selection of data augmentation strategies for contrastive learning of music audio embeddings is dependent on the downstream task, highlighting this as an important embedding design decision.

【15】 NOTSOFAR-1 Challenge: New Datasets, Baseline, and Tasks for Distant  Meeting Transcription
标题:NOTSOFAR-1挑战:远程会议转录的新数据集、基线和任务
链接:https://arxiv.org/abs/2401.08887
作者:Alon Vinnikov,Amir Ivry,Aviv Hurvitz,Igor Abramovski,Sharon Koubi,Ilya Gurvich,Shai Pe`er,Xiong Xiao,Benjamin Martinez Elizalde,Naoyuki Kanda,Xiaofei Wang,Shalev Shaer,Stav Yagev,Yossi Asher,Sunit Sivasankaran,Yifan Gong,Min Tang,Huaming Wang,Eyal Krupka
备注:preprint
摘要:我们介绍了第一个自然办公室谈话者在设置远场录音(“NOTSOFAR-1”)挑战旁边的数据集和基线系统。该挑战的重点是远场会议场景中的远距离说话者日记和自动语音识别(DASR),具有单通道和已知几何多通道轨道,并作为两个新数据集的发布平台:首先,315次会议的基准数据集,平均每次6分钟,捕获广泛的真实世界声学条件和会话动态。它在30个会议室录制,共有4-8名与会者和35名独特的演讲者。第二,一个1000小时的模拟训练数据集,通过增强真实性进行合成,以实现真实世界的泛化,其中包含15,000个真实的声学传递函数。任务集中在单器件DASR上,其中多通道器件始终共享相同的已知几何结构。这与实际会议室中的常见设置一致,并避免了与多设备任务相关的技术复杂性。它还允许开发特定于几何形状的解决方案。NOTSOFAR-1挑战赛旨在推进远程会话语音识别领域的研究,为释放数据驱动方法的潜力提供关键资源,我们认为这些方法目前受到缺乏全面的高质量训练和基准数据集的限制。
摘要:We introduce the first Natural Office Talkers in Settings of Far-field Audio Recordings (``NOTSOFAR-1'') Challenge alongside datasets and baseline system. The challenge focuses on distant speaker diarization and automatic speech recognition (DASR) in far-field meeting scenarios, with single-channel and known-geometry multi-channel tracks, and serves as a launch platform for two new datasets: First, a benchmarking dataset of 315 meetings, averaging 6 minutes each, capturing a broad spectrum of real-world acoustic conditions and conversational dynamics. It is recorded across 30 conference rooms, featuring 4-8 attendees and a total of 35 unique speakers. Second, a 1000-hour simulated training dataset, synthesized with enhanced authenticity for real-world generalization, incorporating 15,000 real acoustic transfer functions. The tasks focus on single-device DASR, where multi-channel devices always share the same known geometry. This is aligned with common setups in actual conference rooms, and avoids technical complexities associated with multi-device tasks. It also allows for the development of geometry-specific solutions. The NOTSOFAR-1 Challenge aims to advance research in the field of distant conversational speech recognition, providing key resources to unlock the potential of data-driven methods, which we believe are currently constrained by the absence of comprehensive high-quality training and benchmarking datasets.

【16】 Using i-vectors for subject-independent cross-session EEG transfer  learning
标题:使用i向量进行主体无关的跨会话EEG迁移学习
链接:https://arxiv.org/abs/2401.08851
作者:Jonathan Lasko,Jeff Ma,Mike Nicoletti,Jonathan Sussman-Fort,Sooyoung Jeong,William Hartmann
备注:11 pages
摘要:认知负荷分类是基于诸如脑电图(EEG)的生理测量来自动确定个体在执行任务期间对工作记忆资源的利用的任务。在本文中,我们遵循一种跨学科的方法,使用语音处理的工具和方法来解决这个问题。我们使用的语料库已于2021年公开发布,作为第一届跨会话工作量估计被动脑机接口竞赛的一部分。我们提出了我们的方法,使用基于i向量的神经网络分类器来完成受试者间的跨会话EEG迁移学习,与等效的受试者相关模型相比,实现了18%的相对改善。我们还报告了实验,显示了我们的独立于受试者的模型如何在被拒受试者上表现出竞争力,并通过额外的受试者数据进行改进,这表明有效的认知负荷确定不需要依赖于受试者的训练。
摘要:Cognitive load classification is the task of automatically determining an individual's utilization of working memory resources during performance of a task based on physiologic measures such as electroencephalography (EEG). In this paper, we follow a cross-disciplinary approach, where tools and methodologies from speech processing are used to tackle this problem. The corpus we use was released publicly in 2021 as part of the first passive brain-computer interface competition on cross-session workload estimation. We present our approach which used i-vector-based neural network classifiers to accomplish inter-subject cross-session EEG transfer learning, achieving 18% relative improvement over equivalent subject-dependent models. We also report experiments showing how our subject-independent models perform competitively on held-out subjects and improve with additional subject data, suggesting that subject-dependent training is not required for effective cognitive load determination.

【17】 Improving ASR Contextual Biasing with Guided Attention
标题:利用引导性注意改善ASR语境偏向
链接:https://arxiv.org/abs/2401.08835
作者:Jiyang Tang,Kwangyoun Kim,Suwon Shon,Felix Wu,Prashant Sridhar,Shinji Watanabe
备注:Accepted at ICASSP 2024
摘要:在本文中,我们提出了一个引导注意(GA)辅助训练损失,提高了自动语音识别(ASR)上下文偏置的有效性和鲁棒性,而无需引入额外的参数。在以往的文献中,一个共同的挑战是,单词错误率(WER)的减少所带来的上下文偏置减少的偏见短语的数量增加。为了解决这一挑战,我们采用GA损失作为除换能器损失之外的额外训练目标。提出的GA损失旨在教导交叉注意如何将偏见短语与文本标记或音频帧对齐。与具有类似动机的研究相比,所提出的损失直接作用于交叉注意权重,并且更容易实现。通过大量的实验,我们证明了该方法不仅会导致较低的WER,而且随着偏见短语数量的增加,仍保持其有效性。具体来说,GA损失减少了罕见的词汇的WER高达19.2%的LibriSpeech相比,上下文偏置基线,并高达49.3%相比,香草换能器。
摘要:In this paper, we propose a Guided Attention (GA) auxiliary training loss, which improves the effectiveness and robustness of automatic speech recognition (ASR) contextual biasing without introducing additional parameters. A common challenge in previous literature is that the word error rate (WER) reduction brought by contextual biasing diminishes as the number of bias phrases increases. To address this challenge, we employ a GA loss as an additional training objective besides the Transducer loss. The proposed GA loss aims to teach the cross attention how to align bias phrases with text tokens or audio frames. Compared to studies with similar motivations, the proposed loss operates directly on the cross attention weights and is easier to implement. Through extensive experiments based on Conformer Transducer with Contextual Adapter, we demonstrate that the proposed method not only leads to a lower WER but also retains its effectiveness as the number of bias phrases increases. Specifically, the GA loss decreases the WER of rare vocabularies by up to 19.2% on LibriSpeech compared to the contextual biasing baseline, and up to 49.3% compared to a vanilla Transducer.

【18】 Robust DOA estimation using deep acoustic imaging
标题:基于深声成像的稳健波达方向估计
链接:https://arxiv.org/abs/2401.08717
作者:Adrian S. Roman,Iran R. Roman,Juan P. Bello
摘要:波达方向估计(DoAE)的目的是在方位角和仰角上跟踪声音。最近的进展包括数据驱动模型,其输入来源于高保真度立体声强度向量或麦克风阵列中通道之间的相关性。球面强度图(SIM)或声图像是一种尚未探索的替代输入表示。SIM受益于高分辨率麦克风阵列,但大多数DoAE数据集使用低分辨率麦克风阵列。因此,我们首先提出了一种超分辨率方法来对低分辨率麦克风进行上采样。接下来,我们对使用SIM作为输入的DoAE模型进行基准测试。我们得到了一个模型,它使用SIM进行DoAE估计,并且优于基线和最先进的模型。我们的研究突出了DoAE任务的声学成像的相关性。
摘要:Direction of arrival estimation (DoAE) aims at tracking a sound in azimuth and elevation. Recent advancements include data-driven models with inputs derived from ambisonics intensity vectors or correlations between channels in a microphone array. A spherical intensity map (SIM), or acoustic image, is an alternative input representation that remains underexplored. SIMs benefit from high-resolution microphone arrays, yet most DoAE datasets use low-resolution ones. Therefore, we first propose a super-resolution method to upsample low-resolution microphones. Next, we benchmark DoAE models that use SIMs as input. We arrive to a model that uses SIMs for DoAE estimation and outperforms a baseline and a state-of-the-art model. Our study highlights the relevance of acoustic imaging for DoAE tasks.


机器翻译由腾讯交互翻译提供,仅供参考