今日论文合集:cs.SD语音26篇,eess.AS音频处理29篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
标题: 利用多令牌预测和推测解码加速基于编解码的语音合成
作者: Tan Dat Nguyen, Ji-Hoon Kim, Jeongsoo Choi, Shukjae Choi, Jinseok Park, Younglo Lee, Joon Son Chung
备注:Submitted to IEEE ICASSP 2025
链接:点击下载PDF文件
摘要:本文的目标是加速基于编解码器的语音合成系统,以最小的牺牲语音质量。我们提出了一种增强的推理方法,允许在推理过程中的速度和质量之间的灵活权衡,而不需要额外的训练。我们的核心思想是使用多个预测头来预测AR模块的每个推理步骤的多个令牌,从而随着头的数量增加而线性减少合成时间。此外,我们引入了一种新的推测性解码技术,利用维特比为基础的算法来选择在每个解码步骤生成的令牌的最佳序列。在我们的实验中,我们证明了预测每个令牌所需的时间与基线模型相比减少了4到5倍,并且在语音清晰度方面具有最小的质量权衡甚至改善。音频样本可在multpletokensprediction.github.io multipletokensprediction.github.io 上获得。摘要:The goal of this paper is to accelerate codec-based speech synthesis systems with minimum sacrifice to speech quality. We propose an enhanced inference method that allows for flexible trade-offs between speed and quality during inference without requiring additional training. Our core idea is to predict multiple tokens per inference step of the AR module using multiple prediction heads, resulting in a linear reduction in synthesis time as the number of heads increases. Furthermore, we introduce a novel speculative decoding technique that utilises a Viterbi-based algorithm to select the optimal sequence of generated tokens at each decoding step. In our experiments, we demonstrate that the time required to predict each token is reduced by a factor of 4 to 5 compared to baseline models, with minimal quality trade-off or even improvement in terms of speech intelligibility. Audio samples are available at: multpletokensprediction.github.io multipletokensprediction.github.io .

【2】 Dynamic Range Compression and Its Effect on Music Genre Classification
标题: 动态范围压缩及其对音乐流派分类的影响
作者: Arlyn Reese Madsen III
链接:点击下载PDF文件
摘要:本文研究了动态范围压缩(DRC)对音乐流派分类精度的影响。通过对200首歌曲的测试集应用各种压缩设置,我们的目标是确定压缩是否可以提高分类器辨别不同音乐流派的能力。支持向量机(SVM)分类器在原始的未压缩数据集上进行训练。研究了阈值、比率、膝宽、攻击时间、释放时间和补偿增益对分类性能的影响。我们的研究结果表明,对测试集进行压缩确实可以将音乐流派分类的准确率平均提高3.1%。最佳压缩设置在实验中有所不同,这表明压缩的有效性取决于模型的训练数据。提供了超过1000个训练和测试分段的最高压缩设置表。总之,这项研究表明,动态范围压缩可以作为一个有价值的预处理技术,提高音乐流派分类。从这项研究中获得的见解可以为开发更准确、更强大的音乐推荐系统提供信息。摘要:This paper investigates the impact of dynamic range compression (DRC) on music genre classification accuracy. By applying various compression settings to the test set of 200 songs, we aim to determine if compression can enhance the classifier's ability to discern distinct musical genres. A support vector machine (SVM) classifier was trained on the original, uncompressed dataset. The study explored the influence of threshold, ratio, knee width, attack time, release time, and makeup gain on classification performance. Our findings indicate that applying compression to the test set can indeed improve music genre classification accuracy on average by 3.1%. The optimal compression settings varied across experiments, suggesting that the effectiveness of compression depends on the training data of the model. A table of the top compression settings over 1000 train and test splits is provided. In conclusion, this research demonstrates that dynamic range compression can serve as a valuable preprocessing technique for enhancing music genre classification. The insights gained from this study can inform the development of more accurate and robust music recommendation systems.

【3】 MeloTrans: A Text to Symbolic Music Generation Model Following Human Composition Habit
标题: MeloTrans:遵循人类作曲习惯的符号音乐生成模型的文本
作者: Yutian Wang, Wanyin Yang, Zhenrong Dai, Yilong Zhang, Kun Zhao, Hui Wang
链接:点击下载PDF文件
摘要:目前,神经网络模型显示出强大的序列预测能力,被广泛应用于自动排版模型中。相比之下,人类作曲的方式与之大不相同,作曲家通常从创作音乐主题开始,然后通过一系列规则将其发展成音乐。这个过程确保了音乐具有特定的结构和变化模式。然而,神经网络模型很难从训练数据中学习这些作曲规则,这导致生成的音乐缺乏音乐性和多样性。本文假设,将神经网络的学习能力与人类来源的知识相结合可能会带来更好的结果。为了存档,我们开发了POP 909 M数据集,这是第一个包含音乐主题及其变体标签的数据集,为模仿人类作曲习惯提供了基础。在此基础上,我们提出了MeloTrans,一个文本到音乐的组成模型,采用主题发展规则的原则。我们的实验表明,MeloTrans超越了现有的音乐生成模型,甚至超过了ChatGPT-4等大型语言模型(LLM)。这凸显了将人类洞察力与神经网络能力相结合以实现卓越的符号音乐生成的重要性。摘要:At present, neural network models show powerful sequence prediction ability and are used in many automatic composition models. In comparison, the way humans compose music is very different from it. Composers usually start by creating musical motifs and then develop them into music through a series of rules. This process ensures that the music has a specific structure and changing pattern. However, it is difficult for neural network models to learn these composition rules from training data, which results in a lack of musicality and diversity in the generated music. This paper posits that integrating the learning capabilities of neural networks with human-derived knowledge may lead to better results. To archive this, we develop the POP909$ _$M dataset, the first to include labels for musical motifs and their variants, providing a basis for mimicking human compositional habits. Building on this, we propose MeloTrans, a text-to-music composition model that employs principles of motif development rules. Our experiments demonstrate that MeloTrans excels beyond existing music generation models and even surpasses Large Language Models (LLMs) like ChatGPT-4. This highlights the importance of merging human insights with neural network capabilities to achieve superior symbolic music generation.

【4】 Enhancing 1-Second 3D SELD Performance with Filter Bank Analysis and SCConv Integration in CST-Former
标题: 通过CST-Former中的过滤器组分析和SCConv集成增强1秒3D SELD性能
作者: Zhehui Zhang
链接:点击下载PDF文件
摘要:最近的SELD研究主要集中在长时间段场景(通常为5到10秒,偶尔为2秒),提高了基准性能,但缺乏真实世界应用所需的时间粒度。为了弥补这一差距,本文研究了短时间段下的SELD与距离估计(3D SELD)系统,特别是针对1秒的窗口,建立了一个新的基线实际3D SELD的适用性。我们进一步探讨了不同滤波器组的影响-巴克,梅尔,和Gammatone音频特征提取,实验结果表明,Gammatone滤波器实现了最高的整体精度在这种情况下。最后,我们建议用SCConv模块替换CST-Former(一种有竞争力的SELD架构)中的卷积模块。这种调整在短片段场景中产生可测量的F分数增益,强调了SCConv改进空间和通道特征表示的潜力。实验结果突出表明,我们的方法是在低延迟约束下实现3D SELD系统在现实世界中部署的重要一步。摘要:Recent SELD research has predominantly focused on long-time segment scenarios (typically 5 to 10 seconds, occasionally 2 seconds), improving benchmark performance but lacking the temporal granularity needed for real-world applications. To bridge this gap, this paper investigates SELD with distance estimation (3D SELD) systems under short-time segments, specifically targeting a 1-second window, establishing a new baseline for practical 3D SELD applicability. We further explore the impact of different filter banks -- Bark, Mel, and Gammatone for audio feature extraction, and experimental results demonstrate that the Gammatone filter achieves the highest overall accuracy in this context. Finally, we propose replacing the convolutional modules within the CST-Former, a competitive SELD architecture, with the SCConv module. This adjustment yields measurable F-score gains in short-segment scenarios, underscoring SCConv's potential to improve spatial and channel feature representation. The experimental results highlight our approach as a significant step towards the real-world deployment of 3D SELD systems under low-latency constraints.

【5】 End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
标题: 使用自我监督学习功能实现语音情感识别与语音活动检测的端到端集成
作者: Natsuo Yamashita, Masaaki Yamamoto, Yohei Kawaguchi
链接:点击下载PDF文件
摘要:语音情感识别(SER)通常对由语音活动检测(VAD)模型检测到的语音段进行操作。然而,VAD模型可能会输出有缺陷的语音段,特别是在嘈杂的环境中,导致后续SER模型的性能下降。为了解决这个问题,我们提出了一个端到端(E2E)的方法,集成VAD和SER使用自监督学习(SSL)功能。VAD模块首先接收SSL特征作为输入,然后将分段的SSL特征馈送到SER模块中。VAD和SER模块都经过联合训练,以优化SER性能。IEMOCAP数据集上的实验结果表明,我们提出的方法提高了SER性能。此外,为了研究我们所提出的方法对VAD和SSL模块的影响,我们对VAD输出和SSL编码器的每一层的权重进行了分析。摘要:Speech Emotion Recognition (SER) often operates on speech segments detected by a Voice Activity Detection (VAD) model. However, VAD models may output flawed speech segments, especially in noisy environments, resulting in degraded performance of subsequent SER models. To address this issue, we propose an end-to-end (E2E) method that integrates VAD and SER using Self-Supervised Learning (SSL) features. The VAD module first receives the SSL features as input, and the segmented SSL features are then fed into the SER module. Both the VAD and SER modules are jointly trained to optimize SER performance. Experimental results on the IEMOCAP dataset demonstrate that our proposed method improves SER performance. Furthermore, to investigate the effect of our proposed method on the VAD and SSL modules, we present an analysis of the VAD outputs and the weights of each layer of the SSL encoder.

【6】 Roadmap towards Superhuman Speech Understanding using Large Language Models
标题: 使用大型语言模型实现超人语音理解的路线图
作者: Fan Bu, Yuhao Zhang, Xidong Wang, Benyou Wang, Qun Liu, Haizhou Li
链接:点击下载PDF文件
摘要:大型语言模型(LLM)的成功促使人们努力整合语音和音频数据,旨在创建能够处理文本和非文本输入的通用基础模型。最近的进展,如GPT-4 o,突出了端到端语音LLM的潜力,它保留了非语义信息和世界知识,以进行更深入的语音理解。为了指导语音LLM的发展,我们提出了一个五级路线图,从基本的自动语音识别(ASR)到高级超人模型,这些模型能够将非语义信息与抽象的声学知识相结合,以完成复杂的任务。此外,我们设计了一个基准,SAGI Bechmark,在这五个级别的各种任务的关键方面,发现在使用抽象的声学知识和能力的完整性的挑战。我们的研究结果揭示了处理非语言线索和抽象声学知识的差距,我们提供了未来的方向。本文概述了推进语音LLM的路线图,介绍了一个评估基准,并提供了其当前的局限性和潜力的关键见解。摘要:The success of large language models (LLMs) has prompted efforts to integrate speech and audio data, aiming to create general foundation models capable of processing both textual and non-textual inputs. Recent advances, such as GPT-4o, highlight the potential for end-to-end speech LLMs, which preserves non-semantic information and world knowledge for deeper speech understanding. To guide the development of speech LLMs, we propose a five-level roadmap, ranging from basic automatic speech recognition (ASR) to advanced superhuman models capable of integrating non-semantic information with abstract acoustic knowledge for complex tasks. Moreover, we design a benchmark, SAGI Bechmark, that standardizes critical aspects across various tasks in these five levels, uncovering challenges in using abstract acoustic knowledge and completeness of capability. Our findings reveal gaps in handling paralinguistic cues and abstract acoustic knowledge, and we offer future directions. This paper outlines a roadmap for advancing speech LLMs, introduces a benchmark for evaluation, and provides key insights into their current limitations and potential.

【7】 CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models
标题: CLaMP 2:使用大型语言模型跨101种语言的多模式音乐信息检索
作者: Shangda Wu, Yashan Wang, Ruibin Yuan, Zhancheng Guo, Xu Tan, Ge Zhang, Monan Zhou, Jing Chen, Xuefeng Mu, Yuejie Gao, Yuanliang Dong, Jiafeng Liu, Xiaobing Li, Feng Yu, Maosong Sun
备注:17 pages, 10 figures, 4 tables
链接:点击下载PDF文件
摘要:当前的音乐信息检索系统面临着管理语言多样性和整合各种音乐形式的挑战。这些限制降低了它们在全球多模态音乐环境中的有效性。为了解决这些问题,我们介绍了CLaMP 2,一个系统兼容101种语言,支持ABC记谱法(基于文本的乐谱格式)和MIDI(乐器数字接口)的音乐信息检索。CLaMP 2在150万个ABC文本三元组上进行了预训练,包括通过对比学习对齐的多语言文本编码器和多模式音乐编码器。通过利用大型语言模型,我们可以大规模地获得精确且一致的多语言描述,从而显著减少文本噪音并平衡语言分布。我们的实验表明,CLaMP 2在多语言语义搜索和跨模态音乐分类方面都取得了最先进的结果,从而建立了包容性和全球音乐信息检索的新标准。摘要:Challenges in managing linguistic diversity and integrating various musical modalities are faced by current music information retrieval systems. These limitations reduce their effectiveness in a global, multimodal music environment. To address these issues, we introduce CLaMP 2, a system compatible with 101 languages that supports both ABC notation (a text-based musical notation format) and MIDI (Musical Instrument Digital Interface) for music information retrieval. CLaMP 2, pre-trained on 1.5 million ABC-MIDI-text triplets, includes a multilingual text encoder and a multimodal music encoder aligned via contrastive learning. By leveraging large language models, we obtain refined and consistent multilingual descriptions at scale, significantly reducing textual noise and balancing language distribution. Our experiments show that CLaMP 2 achieves state-of-the-art results in both multilingual semantic search and music classification across modalities, thus establishing a new standard for inclusive and global music information retrieval.

【8】 EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
标题: EH-MAM:用于自我监督语音表示学习的易于操作的掩蔽声学建模
作者: Ashish Seth, Ramaneswaran Selvakumar, S Sakshi, Sonal Kumar, Sreyan Ghosh, Dinesh Manocha
链接:点击下载PDF文件
摘要:在本文中,我们提出了EH-MAM(易到硬自适应掩蔽声学建模),一种新的自监督学习方法的语音表示学习。在以往的方法,使用随机掩蔽方案掩蔽声学建模(MAM)相比,我们引入了一种新的选择性和自适应掩蔽策略。具体来说,在SSL训练期间,我们逐步将较硬的区域引入模型进行重建。我们的方法自动选择硬区域,并且基于这样的观察:MAM中单个帧的重建丢失可以提供自然信号来判断解决该帧的MAM预文本任务的难度。为了识别这些硬区域,我们采用了一个教师模型,该模型首先预测帧的损失,然后决定屏蔽哪些帧。通过学习创建具有挑战性的问题,例如识别更难的帧并同时解决它们,模型能够学习更有效的表示,从而获得对语音的更全面的理解。在量化方面,EH-MAM在各种低资源语音识别和SUPERB基准测试中的表现优于几个最先进的基线5%-10%。此外,我们进行了全面的分析,表明EH-MAM掩蔽的区域有效地捕捉有用的上下文跨语音帧。摘要:In this paper, we present EH-MAM (Easy-to-Hard adaptive Masked Acoustic Modeling), a novel self-supervised learning approach for speech representation learning. In contrast to the prior methods that use random masking schemes for Masked Acoustic Modeling (MAM), we introduce a novel selective and adaptive masking strategy. Specifically, during SSL training, we progressively introduce harder regions to the model for reconstruction. Our approach automatically selects hard regions and is built on the observation that the reconstruction loss of individual frames in MAM can provide natural signals to judge the difficulty of solving the MAM pre-text task for that frame. To identify these hard regions, we employ a teacher model that first predicts the frame-wise losses and then decides which frames to mask. By learning to create challenging problems, such as identifying harder frames and solving them simultaneously, the model is able to learn more effective representations and thereby acquire a more comprehensive understanding of the speech. Quantitatively, EH-MAM outperforms several state-of-the-art baselines across various low-resource speech recognition and SUPERB benchmarks by 5%-10%. Additionally, we conduct a thorough analysis to show that the regions masked by EH-MAM effectively capture useful context across speech frames.

【9】 Sound Check: Auditing Audio Datasets
标题: 声音检查:审核音频数据集
作者: William Agnew, Julia Barnett, Annie Chu, Rachel Hong, Michael Feffer, Robin Netzorg, Harry H. Jiang, Ezra Awumey, Sauvik Das
链接:点击下载PDF文件
摘要:生成音频模型在功能和公共使用方面都在迅速发展-几个强大的生成音频模型已经提供了开放权重,一些科技公司已经发布了高质量的生成音频产品。然而,虽然先前的工作已经列举了许多来自生成视觉和文本模型训练数据的伦理问题,但我们对生成音频数据集的类似问题知之甚少,包括与偏见,毒性和知识产权有关的问题。为了弥补这一差距,我们对数百个音频数据集进行了文献综述,并选择了其中七个最突出的数据集进行更详细的审计。我们发现这些数据集对女性有偏见,包含对边缘化社区的有害刻板印象,并包含大量受版权保护的作品。为了使艺术家能够看到他们是否在流行的音频数据集中,并促进对这些数据集内容的探索,我们在https: audio-audit.vercel.app上开发了一个Web工具音频数据集探索工具。摘要:Generative audio models are rapidly advancing in both capabilities and public utilization -- several powerful generative audio models have readily available open weights, and some tech companies have released high quality generative audio products. Yet, while prior work has enumerated many ethical issues stemming from the data on which generative visual and textual models have been trained, we have little understanding of similar issues with generative audio datasets, including those related to bias, toxicity, and intellectual property. To bridge this gap, we conducted a literature review of hundreds of audio datasets and selected seven of the most prominent to audit in more detail. We found that these datasets are biased against women, contain toxic stereotypes about marginalized communities, and contain significant amounts of copyrighted work. To enable artists to see if they are in popular audio datasets and facilitate exploration of the contents of these datasets, we developed a web tool audio datasets exploration tool at https: audio-audit.vercel.app.

【10】 AADNet: An End-to-End Deep Learning Model for Auditory Attention Decoding
标题: AADNet:用于听觉注意力解码的端到端深度学习模型
作者: Nhan Duc Thanh Nguyen, Huy Phan, Simon Geirnaert, Kaare Mikkelsen, Preben Kidmose
备注:11 pages, 6 figures
链接:点击下载PDF文件
摘要:听觉注意解码(AAD)是在多人说话环境中使用大脑信号(通常通过脑电图(EEG)记录)识别注意语音的过程。在过去的十年中,AAD经历了不断的发展,这是由于其在神经导向听力设备中的应用前景。与无人值守语音相比,大多数AAD算法依赖于神经夹带到有人值守语音的包络的增加,通常使用两步方法。首先,该算法预测表示的出席语音信号包络,其次,它确定出席语音通过找到最高的预测和实际语音信号的表示之间的相关性。在这项研究中,我们提出了一种新的端到端神经网络架构,命名为AADNet,它将这两个阶段结合起来,直接解决AAD问题。我们将所提出的网络与传统方法进行比较,包括线性刺激重建,典型相关分析,以及使用两个不同数据集的替代非线性刺激重建。AADNet在特定于主题和独立于主题的模型中表现出显着的性能改进。值得注意的是,分析窗口长度分别为1至40秒时,平均独立于受试者的分类准确率为56.1%至82.7%,这表明概括来自未见过受试者的数据的能力显着提高。这些结果突出了深度学习模型在推进AAD方面的潜力,对未来的助听器,辅助设备和临床评估具有重要意义。摘要:Auditory attention decoding (AAD) is the process of identifying the attended speech in a multi-talker environment using brain signals, typically recorded through electroencephalography (EEG). Over the past decade, AAD has undergone continuous development, driven by its promising application in neuro-steered hearing devices. Most AAD algorithms are relying on the increase in neural entrainment to the envelope of attended speech, as compared to unattended speech, typically using a two-step approach. First, the algorithm predicts representations of the attended speech signal envelopes; second, it identifies the attended speech by finding the highest correlation between the predictions and the representations of the actual speech signals. In this study, we proposed a novel end-to-end neural network architecture, named AADNet, which combines these two stages into a direct approach to address the AAD problem. We compare the proposed network against the traditional approaches, including linear stimulus reconstruction, canonical correlation analysis, and an alternative non-linear stimulus reconstruction using two different datasets. AADNet shows a significant performance improvement for both subject-specific and subject-independent models. Notably, the average subject-independent classification accuracies from 56.1 % to 82.7 % with analysis window lengths ranging from 1 to 40 seconds, respectively, show a significantly improved ability to generalize to data from unseen subjects. These results highlight the potential of deep learning models for advancing AAD, with promising implications for future hearing aids, assistive devices, and clinical assessments.

【11】 MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization
标题: MuVi:具有语义对齐和节奏同步的视频到音乐生成
作者: Ruiqi Li, Siqi Zheng, Xize Cheng, Ziang Zhang, Shengpeng Ji, Zhou Zhao
备注:Working in progress
链接:点击下载PDF文件
摘要:生成与视频的视觉内容一致的音乐一直是一项具有挑战性的任务,因为它需要对视觉语义的深刻理解,并且涉及生成旋律,节奏和动态与视觉叙事相协调的音乐。本文介绍了MuVi,一个新的框架,有效地解决了这些挑战,以提高视听内容的凝聚力和沉浸式体验。MuVi通过专门设计的视觉适配器分析视频内容,以提取上下文和时间相关的特征。这些功能用于生成音乐,不仅匹配视频的情绪和主题,还匹配其节奏和节奏。我们还引入了一个对比的音乐视觉预训练计划,以确保同步,基于音乐短语的周期性。此外,我们证明了我们的基于流匹配的音乐生成器具有上下文学习能力,使我们能够控制所生成的音乐的风格和流派。实验结果表明,MuVi在音频质量和时间同步方面表现出优异的性能。生成的音乐视频样本可在https: muvi-v2m.github.io上获得。摘要:Generating music that aligns with the visual content of a video has been a challenging task, as it requires a deep understanding of visual semantics and involves generating music whose melody, rhythm, and dynamics harmonize with the visual narratives. This paper presents MuVi, a novel framework that effectively addresses these challenges to enhance the cohesion and immersive experience of audio-visual content. MuVi analyzes video content through a specially designed visual adaptor to extract contextually and temporally relevant features. These features are used to generate music that not only matches the video's mood and theme but also its rhythm and pacing. We also introduce a contrastive music-visual pre-training scheme to ensure synchronization, based on the periodicity nature of music phrases. In addition, we demonstrate that our flow-matching-based music generator has in-context learning ability, allowing us to control the style and genre of the generated music. Experimental results show that MuVi demonstrates superior performance in both audio quality and temporal synchronization. The generated music video samples are available at https: muvi-v2m.github.io.

【12】 Towards Computational Analysis of Pansori Singing
标题: 潘索里歌唱的计算分析
作者: Sangheon Park, Danbinaerin Han, Dasaem Jeong
备注:Late-Breaking Demo Session of the 25th International Society for Music Information Retrieval (ISMIR) Conference, 2024
链接:点击下载PDF文件
摘要:盘索里是韩国传统音乐中最具代表性的声乐体裁之一,其声乐旋律线细腻,颤音强烈。虽然音乐是口头传播的,没有任何乐谱,但用西方五线谱转录潘索里音乐有几个目的,如音乐文献,教育或研究。在本文中,我们介绍了基于音频和相应的转录的pansori的计算分析,现代音乐信息检索任务如何可以用于分析传统的音乐,以及它如何揭示pansori包含的不同的音频特征。摘要:Pansori is one of the most representative vocal genres of Korean traditional music, which has an elaborated vocal melody line with strong vibrato. Although the music is transmitted orally without any music notation, transcribing pansori music in Western staff notation has been introduced for several purposes, such as documentation of music, education, or research. In this paper, we introduce computational analysis of pansori based on both audio and corresponding transcription, how modern Music Information Retrieval tasks can be used in analyzing traditional music and how it revealed different audio characteristics of what pansori contains.

【13】 What Do Speech Foundation Models Not Learn About Speech?
标题: Speech Foundation模型对Speech不了解什么?
作者: Abdul Waheed, Hanin Atwany, Bhiksha Raj, Rita Singh
备注:20 Pages
链接:点击下载PDF文件
摘要:理解语音基础模型如何捕捉非语言线索对于提高其在不同任务中的可解释性和适应性至关重要。在我们的工作中,我们分析了几个突出的模型,如Whisper,Seamless,Wav 2 Vec,HuBERT和Qwen 2-Audio,重点关注他们在Dynamic-SUPERB基准测试中学习到的语言和非语言任务的表示。我们的研究解决了三个关键问题:(1)什么非语言线索(例如,说话者意图、情感、环境背景)被捕获?(2)这些线索如何在模型的不同层中表示?以及(3)这些表征在多大程度上可以有效地适应下游任务?为了回答这些问题,我们首先在zero-shot设置中评估模型,然后对从这些模型中提取的逐层特征进行微调。我们的研究结果提供了深入了解模型的泛化能力,其分层表示的特点,以及下游任务适应所需的转换程度。我们的研究结果表明,其中一些模型在zero-shot设置中的各种任务上表现良好,尽管没有针对这些任务进行明确的训练。我们还观察到,zero-shot性能与更好地学习表示。逐层特征的分析表明,一些模型在学习表示的可分性和模型深度之间表现出凸关系,不同的层捕获特定于任务的特征。摘要:Understanding how speech foundation models capture non-verbal cues is crucial for improving their interpretability and adaptability across diverse tasks. In our work, we analyze several prominent models such as Whisper, Seamless, Wav2Vec, HuBERT, and Qwen2-Audio focusing on their learned representations in both paralinguistic and non-paralinguistic tasks from the Dynamic-SUPERB benchmark. Our study addresses three key questions: (1) What non-verbal cues (e.g., speaker intent, emotion, environmental context) are captured? (2) How are these cues represented across different layers of the models? and (3) To what extent can these representations be effectively adapted to downstream tasks? To answer these questions, we first evaluate the models in a zero-shot setting, followed by fine-tuning on layer-wise features extracted from these models. Our results provide insights into the models' capacity for generalization, the characteristics of their layer-wise representations, and the degree of transformation required for downstream task adaptation. Our findings suggest that some of these models perform well on various tasks in zero-shot settings, despite not being explicitly trained for those tasks. We also observe that zero-shot performance correlates with better-learned representations. The analysis of layer-wise features demonstrates that some models exhibit a convex relationship between the separability of the learned representations and model depth, with different layers capturing task-specific features.

【14】 Towards Homogeneous Lexical Tone Decoding from Heterogeneous Intracranial Recordings
标题: 从不均匀的脑内记录中实现同质的语调解码
作者: Di Wu, Siyuan Li, Chen Feng, Lu Cao, Yue Zhang, Jie Yang, Mohamad Sawan
备注:Preprint V1 with 10 pages main text
链接:点击下载PDF文件
摘要:脑机接口(BCI)的最新进展已经能够从颅内记录中解码词汇音调,为恢复语音受损的音调语言说话者的沟通能力提供了可能。然而,由生理和仪器因素引起的数据异质性对统一的侵入性脑音解码提出了重大挑战。传统的特定于主题的模型在异构解码范式下运行,无法捕获广义神经表征,并且无法有效地利用跨主题的数据。为了解决这些局限性,我们引入了神经表征的同质性-异质性分解学习(H2 DiLR),这是一种新的框架,可以从多个受试者的颅内记录中分解和学习同质性和异质性。为了评估H2 DiLR,我们收集了来自多个参与者的立体脑电图(sEEG)数据,这些参与者阅读包括407个音节的普通话材料,几乎代表了所有的普通话字符。大量的实验表明,H2 DiLR,作为一个统一的解码范式,显着优于传统的异构解码方法。此外,我们的经验证实,H2 DiLR有效地捕捉神经表征学习过程中的同质性和异质性。摘要:Recent advancements in brain-computer interfaces (BCIs) have enabled the decoding of lexical tones from intracranial recordings, offering the potential to restore the communication abilities of speech-impaired tonal language speakers. However, data heterogeneity induced by both physiological and instrumental factors poses a significant challenge for unified invasive brain tone decoding. Traditional subject-specific models, which operate under a heterogeneous decoding paradigm, fail to capture generalized neural representations and cannot effectively leverage data across subjects. To address these limitations, we introduce Homogeneity-Heterogeneity Disentangled Learning for neural Representations (H2DiLR), a novel framework that disentangles and learns both the homogeneity and heterogeneity from intracranial recordings across multiple subjects. To evaluate H2DiLR, we collected stereoelectroencephalography (sEEG) data from multiple participants reading Mandarin materials comprising 407 syllables, representing nearly all Mandarin characters. Extensive experiments demonstrate that H2DiLR, as a unified decoding paradigm, significantly outperforms the conventional heterogeneous decoding approach. Furthermore, we empirically confirm that H2DiLR effectively captures both homogeneity and heterogeneity during neural representation learning.

【15】 Decoding Emotions: Unveiling Facial Expressions through Acoustic Sensing with Contrastive Attention
标题: 解码情绪:通过对比注意力的声学感知揭示面部表情
作者: Guangjing Wang, Juexing Wang, Ce Zhou, Weikang Ding, Huacheng Zeng, Tianxing Li, Qiben Yan
备注:The extended version of the 2023 IEEE INFOCOM conference paper
链接:点击下载PDF文件
摘要:表情识别通过准确检测用户的情绪状态,为内容推荐和心理健康等应用提供了巨大的希望。传统方法通常依赖于摄像头或可穿戴传感器,这会引发隐私问题并增加额外的设备负担。此外,当训练数据集和推理数据集之间存在分布偏移时,现有的基于声学的方法难以保持令人满意的性能。在本文中,我们介绍了FacER+,一个主动声学面部表情识别系统,它消除了外部麦克风阵列的要求。FacER+通过分析3D面部轮廓和智能手机耳机扬声器之间发射的近超声信号的回波来提取面部表情特征。这种方法不仅减少了背景噪声,而且能够以最少的训练数据识别来自不同用户的不同表情。我们开发了一个基于外部注意力的对比模型,以持续地学习不同用户的表情特征,减少分布差异。涉及20名志愿者的广泛实验表明,FacER+可以准确地识别六种常见的面部表情,在不同的、与用户无关的现实生活场景中的准确率超过90%,超过领先的声学传感方法10%。FacER+为面部表情识别提供了一个强大而实用的解决方案。摘要:Expression recognition holds great promise for applications such as content recommendation and mental healthcare by accurately detecting users' emotional states. Traditional methods often rely on cameras or wearable sensors, which raise privacy concerns and add extra device burdens. In addition, existing acoustic-based methods struggle to maintain satisfactory performance when there is a distribution shift between the training dataset and the inference dataset. In this paper, we introduce FacER+, an active acoustic facial expression recognition system, which eliminates the requirement for external microphone arrays. FacER+ extracts facial expression features by analyzing the echoes of near-ultrasound signals emitted between the 3D facial contour and the earpiece speaker on a smartphone. This approach not only reduces background noise but also enables the identification of different expressions from various users with minimal training data. We develop a contrastive external attention-based model to consistently learn expression features across different users, reducing the distribution differences. Extensive experiments involving 20 volunteers, both with and without masks, demonstrate that FacER+ can accurately recognize six common facial expressions with over 90% accuracy in diverse, user-independent real-life scenarios, surpassing the performance of the leading acoustic sensing methods by 10%. FacER+ offers a robust and practical solution for facial expression recognition.

【16】 Align-ULCNet: Towards Low-Complexity and Robust Acoustic Echo and Noise Reduction
标题: Align-ULCNet:迈向低复杂性和稳健的声学回声和降噪
作者: Shrishti Saha Shetu, Naveen Kumar Desiraju, Wolfgang Mack, Emanuël A. P. Habets
备注:5 pages, 4 figures
链接:点击下载PDF文件
摘要:基于深度学习的声学回声和降噪(AENR)方法在消费类设备中的成功部署,激发了人们对开发低复杂度解决方案的兴趣,同时强调了在现实生活中对强大性能的需求。在这项工作中,我们提出了一种混合方法来增强最先进的(SOTA)ULCNet模型,通过集成时间对齐和并行编码器模块的模型输入,从而实现更好的回声降低和与现有SOTA方法相当的降噪性能。我们还提出了一种基于通道采样的特征重定向方法,在许多具有挑战性的场景中确保鲁棒性能,同时保持整体低计算和内存需求。摘要:The successful deployment of deep learning-based acoustic echo and noise reduction (AENR) methods in consumer devices has spurred interest in developing low-complexity solutions, while emphasizing the need for robust performance in real-life applications. In this work, we propose a hybrid approach to enhance the state-of-the-art (SOTA) ULCNet model by integrating time alignment and parallel encoder blocks for the model inputs, resulting in better echo reduction and comparable noise reduction performance to existing SOTA methods. We also propose a channel-wise sampling-based feature reorientation method, ensuring robust performance across many challenging scenarios, while maintaining overall low computational and memory requirements.

【17】 GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
标题: 使用潜在特征条件处理的基于GAN的低SNR语音增强
作者: Shrishti Saha Shetu, Emanuël A. P. Habets, Andreas Brendel
备注:5 pages, 2 figures
链接:点击下载PDF文件
摘要:在不利的SNR条件下增强语音质量仍然是基于区分性深度神经网络(DNN)的方法的重大挑战。在这项工作中,我们提出了Disco GAN,这是一个时频域生成对抗网络(GAN)的条件下的潜在特征的判别模型预训练的语音增强在低信噪比的情况下。与最先进的判别方法相比,我们提出的方法具有更好的性能,并且还超过了端到端(E2E)训练的GAN模型。我们还调查了各种配置的影响,以调节所提出的GAN模型与判别模型,并评估其对提高语音质量的影响摘要:Enhancing speech quality under adverse SNR conditions remains a significant challenge for discriminative deep neural network (DNN)-based approaches. In this work, we propose DisCoGAN, which is a time-frequency-domain generative adversarial network (GAN) conditioned by the latent features of a discriminative model pre-trained for speech enhancement in low SNR scenarios. Our proposed method achieves superior performance compared to state-of-the-arts discriminative methods and also surpasses end-to-end (E2E) trained GAN models. We also investigate the impact of various configurations for conditioning the proposed GAN model with the discriminative model and assess their influence on enhancing speech quality

【18】 STCON System for the CHiME-8 Challenge
标题: CHiME-8挑战赛的STCON系统
作者: Anton Mitrofanov, Tatiana Prisyach, Tatiana Timofeeva, Sergei Novoselov, Maxim Korenevsky, Yuri Khokhlov, Artem Akulov, Alexander Anikin, Roman Khalili, Iurii Lezhenin, Aleksandr Melnikov, Dmitriy Miroshnichenko, Nikita Mamaev, Ilya Odegov, Olga Rudnitskaya, Aleksei Romanenko
链接:点击下载PDF文件
摘要:本文介绍了CHiME-8挑战任务1(DASR)的STCON系统,旨在远程自动语音转录和日记与多个记录设备。我们主要关注的是经过精心训练和调整的日记管道和发言人计数。这允许显著降低日志化错误率(DER),并获得用于语音分离和识别的更可靠的段。为了改善源分离,我们设计了一个引导目标说话人提取(G-TSE)模型,并将其与传统的引导源分离(GSS)方法结合使用。为了训练我们管道的各个部分,我们研究了几种数据增强和生成技术,这有助于我们提高整体系统质量。摘要:This paper describes the STCON system for the CHiME-8 Challenge Task 1 (DASR) aimed at distant automatic speech transcription and diarization with multiple recording devices. Our main attention was paid to carefully trained and tuned diarization pipeline and speaker counting. This allowed to significantly reduce diarization error rate (DER) and obtain more reliable segments for speech separation and recognition. To improve source separation, we designed a Guided Target speaker Extraction (G-TSE) model and used it in conjunction with the traditional Guided Source Separation (GSS) method. To train various parts of our pipeline, we investigated several data augmentation and generation techniques, which helped us to improve the overall system quality.

【19】 On the Use of Audio to Improve Dialogue Policies
标题: 关于利用音频改进对话政策
作者: Daniel Roncel, Federico Costa, Javier Hernando
备注:IberSpeech 2024
链接:点击下载PDF文件
摘要:随着语音技术的进步,面向目标的口语对话系统越来越受欢迎。对话系统的主要模块之一通常是负责确定系统动作的对话策略。该组件通常仅依赖于音频transmittance,强烈依赖于它们的质量,而忽略了嵌入在用户语音中的非常重要的语言外信息。在本文中,我们提出了新的架构来添加音频信息相结合的语音和文本嵌入使用双多头注意组件。我们的实验表明,音频嵌入感知对话策略优于基于文本的策略,特别是在嘈杂的转录场景中,并且如何将文本和音频嵌入相结合对于提高性能至关重要。与DSTC 2数据集上仅基于文本的对话系统相比,我们在用户请求得分方面获得了9.8%的相对提高。摘要:With the significant progress of speech technologies, spoken goal-oriented dialogue systems are becoming increasingly popular. One of the main modules of a dialogue system is typically the dialogue policy, which is responsible for determining system actions. This component usually relies only on audio transcriptions, being strongly dependent on their quality and ignoring very important extralinguistic information embedded in the user's speech. In this paper, we propose new architectures to add audio information by combining speech and text embeddings using a Double Multi-Head Attention component. Our experiments show that audio embedding-aware dialogue policies outperform text-based ones, particularly in noisy transcription scenarios, and that how text and audio embeddings are combined is crucial to improve performance. We obtained a 9.8% relative improvement in the User Request Score compared to an only-text-based dialogue system on the DSTC2 dataset.

【20】 Enhancing Crowdsourced Audio for Text-to-Speech Models
标题: 增强文本到语音模型的众包音频
作者: José Giraldo, Martí Llopart-Font, Alex Peiró-Lilja, Carme Armentano-Oller, Gerard Sant, Baybars Külebi
备注:Submitted to Iberspeech 2024
链接:点击下载PDF文件
摘要:高质量的音频数据是训练鲁棒的文本到语音模型的关键先决条件,这通常限制了机会主义或众包数据集的使用。本文提出了一种方法来克服这一限制,通过在Commonvoice的加泰罗尼亚语子集上实现去噪管道,Commonvoice是一个以其固有的噪声和可变性而闻名的众包语料库。该流水线结合了音频增强阶段,然后是选择性滤波策略。我们开发了一种自动过滤机制,利用非侵入式语音质量评估(NISQA)模型来识别和保留增强后的最高质量样本。为了评估这种方法的有效性,我们在处理后的数据集上训练了最先进的基于扩散的TTS模型。结果显示出显著的改善,与没有增强的基线数据集相比,UTMOS评分增加了0.4。这种方法有望扩大TTS应用程序中众包数据的实用性,特别是对于像加泰罗尼亚语这样的中低资源语言。摘要:High-quality audio data is a critical prerequisite for training robust text-to-speech models, which often limits the use of opportunistic or crowdsourced datasets. This paper presents an approach to overcome this limitation by implementing a denoising pipeline on the Catalan subset of Commonvoice, a crowd-sourced corpus known for its inherent noise and variability. The pipeline incorporates an audio enhancement phase followed by a selective filtering strategy. We developed an automatic filtering mechanism leveraging Non-Intrusive Speech Quality Assessment (NISQA) models to identify and retain the highest quality samples post-enhancement. To evaluate the efficacy of this approach, we trained a state of the art diffusion-based TTS model on the processed dataset. The results show a significant improvement, with an increase of 0.4 in the UTMOS Score compared to the baseline dataset without enhancement. This methodology shows promise for expanding the utility of crowdsourced data in TTS applications, particularly for mid to low resource languages like Catalan.

【21】 DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
标题: DART:多说话人文本到语音中口音和说话人表示的分离
作者: Jan Melechovsky, Ambuj Mehrish, Berrak Sisman, Dorien Herremans
备注:Accepted in Audio Imagination workshop of NeurIPS 2024
链接:点击下载PDF文件
摘要:文本到语音(TTS)系统的最新进展使得能够从文本输入生成自然且有表达力的语音。重音TTS旨在通过使合成语音与少数群体听众更相关来增强用户体验,并在各种应用程序和上下文中有用。通过允许用户选择说话者身份和口音的任何组合,可以进一步使语音合成更加灵活,从而产生广泛的个性化语音输出。目前的模型很难将说话者和口音表示分开,这使得很难在保持相同说话者特征的同时准确地模仿不同的口音。我们提出了一种新的方法来解开扬声器和口音表示使用多级变分自编码器(ML-VAE)和矢量量化(VQ),以提高灵活性和增强个性化的语音合成。我们提出的方法解决了有效地分离扬声器和口音特征的挑战,使更细粒度的控制合成语音。代码和语音样本是公开的。摘要:Recent advancements in Text-to-Speech (TTS) systems have enabled the generation of natural and expressive speech from textual input. Accented TTS aims to enhance user experience by making the synthesized speech more relatable to minority group listeners, and useful across various applications and context. Speech synthesis can further be made more flexible by allowing users to choose any combination of speaker identity and accent, resulting in a wide range of personalized speech outputs. Current models struggle to disentangle speaker and accent representation, making it difficult to accurately imitate different accents while maintaining the same speaker characteristics. We propose a novel approach to disentangle speaker and accent representations using multi-level variational autoencoders (ML-VAE) and vector quantization (VQ) to improve flexibility and enhance personalization in speech synthesis. Our proposed method addresses the challenge of effectively separating speaker and accent characteristics, enabling more fine-grained control over the synthesized speech. Code and speech samples are publicly available.

【22】 DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis
标题: DuRIAN-E 2:具有自适应变分自动编码器和对抗学习的持续时间知情注意力网络,用于表达性文本到语音合成
作者: Yu Gu, Qiushi Zhu, Guangzhi Lei, Chao Weng, Dan Su
备注:Accepted by ICASSP2024
链接:点击下载PDF文件
摘要:本文提出了一种改进的DuIAN-E(DuIAN-E 2),这也是一个持续时间通知注意神经网络的表达和高保真的文本到语音(TTS)合成。与DuIAN-E模型类似,多个堆叠的基于SwishRNN的Transformer块被用作语言编码器,并且样式自适应实例规范化(SAIN)层也被利用到帧级编码器中,以提高所提出的DuIAN-E 2的表达能力的建模能力。同时,受其他TTS模型(如VITS)的启发,建议的DuriAN-E 2采用了归一化流增强的变分自编码器(VAE)和具有对抗训练策略的BigVGAN波形发生器,进一步提高了合成语音的质量和表现力。客观测试和主观评估结果都证明,所提出的表达性TTS模型DurIAN-E 2比DurIAN-E以外的几种最先进的方法具有更好的性能。摘要:This paper proposes an improved version of DurIAN-E (DurIAN-E 2), which is also a duration informed attention neural network for expressive and high-fidelity text-to-speech (TTS) synthesis. Similar with the DurIAN-E model, multiple stacked SwishRNN-based Transformer blocks are utilized as linguistic encoders and Style-Adaptive Instance Normalization (SAIN) layers are also exploited into frame-level encoders to improve the modeling ability of expressiveness in the proposed the DurIAN-E 2. Meanwhile, motivated by other TTS models using generative models such as VITS, the proposed DurIAN-E 2 utilizes variational autoencoders (VAEs) augmented with normalizing flows and a BigVGAN waveform generator with adversarial training strategy, which further improve the synthesized speech quality and expressiveness. Both objective test and subjective evaluation results prove that the proposed expressive TTS model DurIAN-E 2 can achieve better performance than several state-of-the-art approaches besides DurIAN-E.

【23】 Investigating Effective Speaker Property Privacy Protection in Federated Learning for Speech Emotion Recognition
标题: 研究语音情感识别联邦学习中有效的说话人财产隐私保护
作者: Chao Tan, Sheng Li, Yang Cao, Zhao Ren, Tanja Schultz
链接:点击下载PDF文件
摘要:联合学习(FL)是一种隐私保护方法,允许服务器聚合从本地客户端传输的分布式模型,而不是对用户数据进行训练。最近,FL已应用于语音情感识别(SER),以实现安全的人机交互应用。最近的研究发现,FL仍然容易受到推理攻击。为此,本文重点研究了FL在SER中的安全性。我们提出了一种新的方法来保护语音数据中的属性信息,通过分解声音中的各种属性,并添加扰动这些属性。我们的实验表明,该方法提供了更好的隐私效用权衡比现有的方法。这种权衡可以在保持相似的FL效用水平的同时实现更有效的攻击防御。这项工作可以指导未来的工作在语音处理中的隐私保护方法。摘要:Federated Learning (FL) is a privacy-preserving approach that allows servers to aggregate distributed models transmitted from local clients rather than training on user data. More recently, FL has been applied to Speech Emotion Recognition (SER) for secure human-computer interaction applications. Recent research has found that FL is still vulnerable to inference attacks. To this end, this paper focuses on investigating the security of FL for SER concerning property inference attacks. We propose a novel method to protect the property information in speech data by decomposing various properties in the sound and adding perturbations to these properties. Our experiments show that the proposed method offers better privacy-utility trade-offs than existing methods. The trade-offs enable more effective attack prevention while maintaining similar FL utility levels. This work can guide future work on privacy protection methods in speech processing.

【24】 Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation
标题: 失败向前:通过合成数据和检索增强改进ASB的生成式错误纠正
作者: Sreyan Ghosh, Mohammad Sadegh Rasooli, Michael Levit, Peidong Wang, Jian Xue, Dinesh Manocha, Jinyu Li
备注:Preprint. Under Review
链接:点击下载PDF文件
摘要:生成式纠错(GEC)是一种有效的语音后处理方法,可以提高自动语音识别(ASR)系统的性能。然而,我们发现,GEC模型很难推广到训练过程中遇到的特定类型的错误之外,限制了它们在测试时纠正新的、看不见的错误的能力,特别是在域外(OOD)场景中。这种现象随着命名实体(NE)而放大,其中,除了关于NE的上下文信息或知识不足之外,新的NE不断出现。为了解决这些问题,我们提出了DARAG(数据和检索增强生成纠错),一种新的方法,旨在提高GEC域(ID)和OOD场景中的ASR。我们使用通过提示LLM和文本到语音模型生成的合成数据来增强GEC训练数据集,从而模拟模型可以学习的其他错误。对于面向对象的场景,我们模拟测试时的错误,从新的领域类似,并在一个无监督的方式。此外,为了更好地处理命名实体,我们引入检索增强校正,通过从数据库中检索的实体来增强输入。我们的方法是简单的,可扩展的,域和语言无关。我们在多个数据集和设置上进行了实验,表明DARAG优于我们所有的基线,在ID和OOD设置中分别实现了8%-30%和10%-33%的相对WER改进。摘要:Generative Error Correction (GEC) has emerged as a powerful post-processing method to enhance the performance of Automatic Speech Recognition (ASR) systems. However, we show that GEC models struggle to generalize beyond the specific types of errors encountered during training, limiting their ability to correct new, unseen errors at test time, particularly in out-of-domain (OOD) scenarios. This phenomenon amplifies with named entities (NEs), where, in addition to insufficient contextual information or knowledge about the NEs, novel NEs keep emerging. To address these issues, we propose DARAG (Data- and Retrieval-Augmented Generative Error Correction), a novel approach designed to improve GEC for ASR in in-domain (ID) and OOD scenarios. We augment the GEC training dataset with synthetic data generated by prompting LLMs and text-to-speech models, thereby simulating additional errors from which the model can learn. For OOD scenarios, we simulate test-time errors from new domains similarly and in an unsupervised fashion. Additionally, to better handle named entities, we introduce retrieval-augmented correction by augmenting the input with entities retrieved from a database. Our approach is simple, scalable, and both domain- and language-agnostic. We experiment on multiple datasets and settings, showing that DARAG outperforms all our baselines, achieving 8 % -- 30 % relative WER improvements in ID and 10 % -- 33 % improvements in OOD settings.

【25】 Using RLHF to align speech enhancement approaches to mean-opinion quality scores
标题: 使用WLHF将语音增强方法与平均意见质量分数保持一致
作者: Anurag Kumar, Andrew Perrault, Donald S. Williamson
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:客观语音质量测量通常用于评估语音增强算法,但已经表明,它们作为学习目标是次优的,因为它们并不总是与人类主观评级很好地一致。这种未对准通常会导致明显的失真和伪像,从而导致语音增强无效。为了解决这些问题,我们提出了一个从人类反馈中强化学习(RLHF)框架,通过使用基于平均意见得分(MOS)的奖励模型优化性能来微调现有的语音增强方法。我们的研究结果表明,RLHF微调模型具有最好的性能在不同的基准的客观和基于MOS的语音质量评估指标的语音银行+DEMAND数据集。通过消融研究,我们表明,策略梯度损失和监督MSE损失对于不同指标的平衡优化都很重要。摘要:Objective speech quality measures are typically used to assess speech enhancement algorithms, but it has been shown that they are sub-optimal as learning objectives because they do not always align well with human subjective ratings. This misalignment often results in noticeable distortions and artifacts that cause speech enhancement to be ineffective. To address these issues, we propose a reinforcement learning from human feedback (RLHF) framework to fine-tune an existing speech enhancement approach by optimizing performance using a mean-opinion score (MOS)-based reward model. Our results show that the RLHF-finetuned model has the best performance across different benchmarks for both objective and MOS-based speech quality assessment metrics on the Voicebank+DEMAND dataset. Through ablation studies, we show that both policy gradient loss and supervised MSE loss are important for balanced optimization across the different metrics.

【26】 Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks
标题: 使用语音基础模型的多视图多任务建模用于语音取证任务
作者: Orchid Chetia Phukan, Devyani Koshal, Swarup Ranjan Behera, Arun Balaji Buduru, Rajesh Sharma
链接:点击下载PDF文件
摘要:语音取证任务(SFT),诸如自动说话人识别(ASR)、语音情感识别(SER)、性别识别(GR)和年龄估计(AE),在不同的安全和生物特征应用中找到用途。以前的工作已经应用了各种技术,最近的研究集中在应用语音基础模型(SFM),以提高性能。然而,大多数先前的努力都集中在为每个任务单独构建单独的模型上,尽管这些任务之间具有内在的相似性。这种孤立的方法会导致更高的计算资源需求,增加成本,时间消耗和维护挑战。在这项研究中,我们通过采用多任务学习策略来应对这些挑战。首先,我们探讨了各种国家的最先进的(SOTA)SFM提取他们的表示学习这些SFT和调查他们的有效性,在每个任务具体。其次,我们分析了在多任务学习框架中提取的表征在SFT上的性能。与单独的特定任务模型相比,我们观察到当SFT一起建模时性能会下降,作为补救措施,我们提出了多视图学习(MVL)。视图是来自不同的SFM的表示,通过每个SFM独有的特征转换成不同的抽象空间。通过利用MVL,我们整合了这些不同的表示,以捕获跨任务的互补信息,增强共享学习过程。我们引入了一个新的框架,称为TANGO(任务对齐与Inter-view门控最优运输)来实现这种方法。通过TANGO,我们实现了与单个SFM表示以及基准数据集(如CREMA-D,emo-DB和BAVED)的基线融合技术相比的最高性能。摘要:Speech forensic tasks (SFTs), such as automatic speaker recognition (ASR), speech emotion recognition (SER), gender recognition (GR), and age estimation (AE), find use in different security and biometric applications. Previous works have applied various techniques, with recent studies focusing on applying speech foundation models (SFMs) for improved performance. However, most prior efforts have centered on building individual models for each task separately, despite the inherent similarities among these tasks. This isolated approach results in higher computational resource requirements, increased costs, time consumption, and maintenance challenges. In this study, we address these challenges by employing a multi-task learning strategy. Firstly, we explore the various state-of-the-art (SOTA) SFMs by extracting their representations for learning these SFTs and investigating their effectiveness at each task specifically. Secondly, we analyze the performance of the extracted representations on the SFTs in a multi-task learning framework. We observe a decline in performance when SFTs are modeled together compared to individual task-specific models, and as a remedy, we propose multi-view learning (MVL). Views are representations from different SFMs transformed into distinct abstract spaces by characteristics unique to each SFM. By leveraging MVL, we integrate these diverse representations to capture complementary information across tasks, enhancing the shared learning process. We introduce a new framework called TANGO (Task Alignment with iNter-view Gated Optimal transport) to implement this approach. With TANGO, we achieve the topmost performance in comparison to individual SFM representations as well as baseline fusion techniques across benchmark datasets such as CREMA-D, emo-DB, and BAVED.


eess.AS音频处理
【1】 Align-ULCNet: Towards Low-Complexity and Robust Acoustic Echo and Noise Reduction
标题: Align-ULCNet:迈向低复杂性和稳健的声学回声和降噪
作者: Shrishti Saha Shetu, Naveen Kumar Desiraju, Wolfgang Mack, Emanuël A. P. Habets
备注:5 pages, 4 figures
链接:点击下载PDF文件
摘要:基于深度学习的声学回声和降噪(AENR)方法在消费类设备中的成功部署,激发了人们对开发低复杂度解决方案的兴趣,同时强调了在现实生活中对强大性能的需求。在这项工作中,我们提出了一种混合方法来增强最先进的(SOTA)ULCNet模型,通过集成时间对齐和并行编码器模块的模型输入,从而实现更好的回声降低和与现有SOTA方法相当的降噪性能。我们还提出了一种基于通道采样的特征重定向方法,在许多具有挑战性的场景中确保鲁棒性能,同时保持整体低计算和内存需求。摘要:The successful deployment of deep learning-based acoustic echo and noise reduction (AENR) methods in consumer devices has spurred interest in developing low-complexity solutions, while emphasizing the need for robust performance in real-life applications. In this work, we propose a hybrid approach to enhance the state-of-the-art (SOTA) ULCNet model by integrating time alignment and parallel encoder blocks for the model inputs, resulting in better echo reduction and comparable noise reduction performance to existing SOTA methods. We also propose a channel-wise sampling-based feature reorientation method, ensuring robust performance across many challenging scenarios, while maintaining overall low computational and memory requirements.

【2】 GAN-Based Speech Enhancement for Low SNR Using Latent Feature Conditioning
标题: 使用潜在特征条件处理的基于GAN的低SNR语音增强
作者: Shrishti Saha Shetu, Emanuël A. P. Habets, Andreas Brendel
备注:5 pages, 2 figures
链接:点击下载PDF文件
摘要:在不利的SNR条件下增强语音质量仍然是基于区分性深度神经网络(DNN)的方法的重大挑战。在这项工作中,我们提出了Disco GAN,这是一个时频域生成对抗网络(GAN)的条件下的潜在特征的判别模型预训练的语音增强在低信噪比的情况下。与最先进的判别方法相比,我们提出的方法具有更好的性能,并且还超过了端到端(E2E)训练的GAN模型。我们还调查了各种配置的影响,以调节所提出的GAN模型与判别模型,并评估其对提高语音质量的影响摘要:Enhancing speech quality under adverse SNR conditions remains a significant challenge for discriminative deep neural network (DNN)-based approaches. In this work, we propose DisCoGAN, which is a time-frequency-domain generative adversarial network (GAN) conditioned by the latent features of a discriminative model pre-trained for speech enhancement in low SNR scenarios. Our proposed method achieves superior performance compared to state-of-the-arts discriminative methods and also surpasses end-to-end (E2E) trained GAN models. We also investigate the impact of various configurations for conditioning the proposed GAN model with the discriminative model and assess their influence on enhancing speech quality

【3】 STCON System for the CHiME-8 Challenge
标题: CHiME-8挑战赛的STCON系统
作者: Anton Mitrofanov, Tatiana Prisyach, Tatiana Timofeeva, Sergei Novoselov, Maxim Korenevsky, Yuri Khokhlov, Artem Akulov, Alexander Anikin, Roman Khalili, Iurii Lezhenin, Aleksandr Melnikov, Dmitriy Miroshnichenko, Nikita Mamaev, Ilya Odegov, Olga Rudnitskaya, Aleksei Romanenko
链接:点击下载PDF文件
摘要:本文介绍了CHiME-8挑战任务1(DASR)的STCON系统,旨在远程自动语音转录和日记与多个记录设备。我们主要关注的是经过精心训练和调整的日记管道和发言人计数。这允许显著降低日志化错误率(DER),并获得用于语音分离和识别的更可靠的段。为了改善源分离,我们设计了一个引导目标说话人提取(G-TSE)模型,并将其与传统的引导源分离(GSS)方法结合使用。为了训练我们管道的各个部分,我们研究了几种数据增强和生成技术,这有助于我们提高整体系统质量。摘要:This paper describes the STCON system for the CHiME-8 Challenge Task 1 (DASR) aimed at distant automatic speech transcription and diarization with multiple recording devices. Our main attention was paid to carefully trained and tuned diarization pipeline and speaker counting. This allowed to significantly reduce diarization error rate (DER) and obtain more reliable segments for speech separation and recognition. To improve source separation, we designed a Guided Target speaker Extraction (G-TSE) model and used it in conjunction with the traditional Guided Source Separation (GSS) method. To train various parts of our pipeline, we investigated several data augmentation and generation techniques, which helped us to improve the overall system quality.

【4】 On the Use of Audio to Improve Dialogue Policies
标题: 关于利用音频改进对话政策
作者: Daniel Roncel, Federico Costa, Javier Hernando
备注:IberSpeech 2024
链接:点击下载PDF文件
摘要:随着语音技术的进步,面向目标的口语对话系统越来越受欢迎。对话系统的主要模块之一通常是负责确定系统动作的对话策略。该组件通常仅依赖于音频transmittance,强烈依赖于它们的质量,而忽略了嵌入在用户语音中的非常重要的语言外信息。在本文中,我们提出了新的架构来添加音频信息相结合的语音和文本嵌入使用双多头注意组件。我们的实验表明,音频嵌入感知对话策略优于基于文本的策略,特别是在嘈杂的转录场景中,并且如何将文本和音频嵌入相结合对于提高性能至关重要。与DSTC 2数据集上仅基于文本的对话系统相比,我们在用户请求得分方面获得了9.8%的相对提高。摘要:With the significant progress of speech technologies, spoken goal-oriented dialogue systems are becoming increasingly popular. One of the main modules of a dialogue system is typically the dialogue policy, which is responsible for determining system actions. This component usually relies only on audio transcriptions, being strongly dependent on their quality and ignoring very important extralinguistic information embedded in the user's speech. In this paper, we propose new architectures to add audio information by combining speech and text embeddings using a Double Multi-Head Attention component. Our experiments show that audio embedding-aware dialogue policies outperform text-based ones, particularly in noisy transcription scenarios, and that how text and audio embeddings are combined is crucial to improve performance. We obtained a 9.8% relative improvement in the User Request Score compared to an only-text-based dialogue system on the DSTC2 dataset.

【5】 Enhancing Crowdsourced Audio for Text-to-Speech Models
标题: 增强文本到语音模型的众包音频
作者: José Giraldo, Martí Llopart-Font, Alex Peiró-Lilja, Carme Armentano-Oller, Gerard Sant, Baybars Külebi
备注:Submitted to Iberspeech 2024
链接:点击下载PDF文件
摘要:高质量的音频数据是训练鲁棒的文本到语音模型的关键先决条件,这通常限制了机会主义或众包数据集的使用。本文提出了一种方法来克服这一限制,通过在Commonvoice的加泰罗尼亚语子集上实现去噪管道,Commonvoice是一个以其固有的噪声和可变性而闻名的众包语料库。该流水线结合了音频增强阶段,然后是选择性滤波策略。我们开发了一种自动过滤机制,利用非侵入式语音质量评估(NISQA)模型来识别和保留增强后的最高质量样本。为了评估这种方法的有效性,我们在处理过的数据集上训练了一个最先进的基于扩散的TTS模型。结果显示出显著的改善,与没有增强的基线数据集相比,UTMOS评分增加了0.4。这种方法有望扩大TTS应用程序中众包数据的实用性,特别是对于像加泰罗尼亚语这样的中低资源语言。摘要:High-quality audio data is a critical prerequisite for training robust text-to-speech models, which often limits the use of opportunistic or crowdsourced datasets. This paper presents an approach to overcome this limitation by implementing a denoising pipeline on the Catalan subset of Commonvoice, a crowd-sourced corpus known for its inherent noise and variability. The pipeline incorporates an audio enhancement phase followed by a selective filtering strategy. We developed an automatic filtering mechanism leveraging Non-Intrusive Speech Quality Assessment (NISQA) models to identify and retain the highest quality samples post-enhancement. To evaluate the efficacy of this approach, we trained a state of the art diffusion-based TTS model on the processed dataset. The results show a significant improvement, with an increase of 0.4 in the UTMOS Score compared to the baseline dataset without enhancement. This methodology shows promise for expanding the utility of crowdsourced data in TTS applications, particularly for mid to low resource languages like Catalan.

【6】 DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
标题: DART:多说话人文本到语音中口音和说话人表示的分离
作者: Jan Melechovsky, Ambuj Mehrish, Berrak Sisman, Dorien Herremans
备注:Accepted in Audio Imagination workshop of NeurIPS 2024
链接:点击下载PDF文件
摘要:文本到语音(TTS)系统的最新进展使得能够从文本输入生成自然且有表达力的语音。重音TTS旨在通过使合成语音与少数群体听众更相关来增强用户体验,并在各种应用程序和上下文中有用。通过允许用户选择说话者身份和口音的任何组合,可以进一步使语音合成更加灵活,从而产生广泛的个性化语音输出。目前的模型很难将说话者和口音表示分开,这使得很难在保持相同说话者特征的同时准确地模仿不同的口音。我们提出了一种新的方法来解开扬声器和口音表示使用多级变分自编码器(ML-VAE)和矢量量化(VQ),以提高灵活性和增强个性化的语音合成。我们提出的方法解决了有效地分离扬声器和口音特征的挑战,使更细粒度的控制合成语音。代码和语音样本是公开的。摘要:Recent advancements in Text-to-Speech (TTS) systems have enabled the generation of natural and expressive speech from textual input. Accented TTS aims to enhance user experience by making the synthesized speech more relatable to minority group listeners, and useful across various applications and context. Speech synthesis can further be made more flexible by allowing users to choose any combination of speaker identity and accent, resulting in a wide range of personalized speech outputs. Current models struggle to disentangle speaker and accent representation, making it difficult to accurately imitate different accents while maintaining the same speaker characteristics. We propose a novel approach to disentangle speaker and accent representations using multi-level variational autoencoders (ML-VAE) and vector quantization (VQ) to improve flexibility and enhance personalization in speech synthesis. Our proposed method addresses the challenge of effectively separating speaker and accent characteristics, enabling more fine-grained control over the synthesized speech. Code and speech samples are publicly available.

【7】 DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis
标题: DuRIAN-E 2:具有自适应变分自动编码器和对抗学习的持续时间知情注意力网络,用于表达性文本到语音合成
作者: Yu Gu, Qiushi Zhu, Guangzhi Lei, Chao Weng, Dan Su
备注:Accepted by ICASSP2024
链接:点击下载PDF文件
摘要:本文提出了一种改进的DuIAN-E(DuIAN-E 2),这也是一个持续时间通知注意神经网络的表达和高保真的文本到语音(TTS)合成。与DuIAN-E模型类似,多个堆叠的基于SwishRNN的Transformer块被用作语言编码器,并且样式自适应实例规范化(SAIN)层也被利用到帧级编码器中,以提高所提出的DuIAN-E 2的表达能力的建模能力。同时,受其他TTS模型(如VITS)的启发,建议的DuriAN-E 2采用了归一化流增强的变分自编码器(VAE)和具有对抗训练策略的BigVGAN波形发生器,进一步提高了合成语音的质量和表现力。客观测试和主观评估结果都证明,所提出的表达性TTS模型DurIAN-E 2比DurIAN-E以外的几种最先进的方法具有更好的性能。摘要:This paper proposes an improved version of DurIAN-E (DurIAN-E 2), which is also a duration informed attention neural network for expressive and high-fidelity text-to-speech (TTS) synthesis. Similar with the DurIAN-E model, multiple stacked SwishRNN-based Transformer blocks are utilized as linguistic encoders and Style-Adaptive Instance Normalization (SAIN) layers are also exploited into frame-level encoders to improve the modeling ability of expressiveness in the proposed the DurIAN-E 2. Meanwhile, motivated by other TTS models using generative models such as VITS, the proposed DurIAN-E 2 utilizes variational autoencoders (VAEs) augmented with normalizing flows and a BigVGAN waveform generator with adversarial training strategy, which further improve the synthesized speech quality and expressiveness. Both objective test and subjective evaluation results prove that the proposed expressive TTS model DurIAN-E 2 can achieve better performance than several state-of-the-art approaches besides DurIAN-E.

【8】 Investigating Effective Speaker Property Privacy Protection in Federated Learning for Speech Emotion Recognition
标题: 研究语音情感识别联邦学习中有效的说话人财产隐私保护
作者: Chao Tan, Sheng Li, Yang Cao, Zhao Ren, Tanja Schultz
链接:点击下载PDF文件
摘要:联合学习(FL)是一种隐私保护方法,允许服务器聚合从本地客户端传输的分布式模型,而不是对用户数据进行训练。最近,FL已应用于语音情感识别(SER),以实现安全的人机交互应用。最近的研究发现,FL仍然容易受到推理攻击。为此,本文重点研究了FL在SER中的安全性。我们提出了一种新的方法来保护语音数据中的属性信息,通过分解声音中的各种属性,并添加扰动这些属性。我们的实验表明,该方法提供了更好的隐私效用权衡比现有的方法。这种权衡可以在保持相似的FL效用水平的同时实现更有效的攻击防御。这项工作可以指导未来的工作在语音处理中的隐私保护方法。摘要:Federated Learning (FL) is a privacy-preserving approach that allows servers to aggregate distributed models transmitted from local clients rather than training on user data. More recently, FL has been applied to Speech Emotion Recognition (SER) for secure human-computer interaction applications. Recent research has found that FL is still vulnerable to inference attacks. To this end, this paper focuses on investigating the security of FL for SER concerning property inference attacks. We propose a novel method to protect the property information in speech data by decomposing various properties in the sound and adding perturbations to these properties. Our experiments show that the proposed method offers better privacy-utility trade-offs than existing methods. The trade-offs enable more effective attack prevention while maintaining similar FL utility levels. This work can guide future work on privacy protection methods in speech processing.

【9】 Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation
标题: 失败向前:通过合成数据和检索增强改进ASB的生成式错误纠正
作者: Sreyan Ghosh, Mohammad Sadegh Rasooli, Michael Levit, Peidong Wang, Jian Xue, Dinesh Manocha, Jinyu Li
备注:Preprint. Under Review
链接:点击下载PDF文件
摘要:生成式纠错(GEC)是一种有效的语音后处理方法,可以提高自动语音识别(ASR)系统的性能。然而,我们发现,GEC模型很难推广到训练过程中遇到的特定类型的错误之外,限制了它们在测试时纠正新的、看不见的错误的能力,特别是在域外(OOD)场景中。这种现象随着命名实体(NE)而放大,其中,除了关于NE的上下文信息或知识不足之外,新的NE不断出现。为了解决这些问题,我们提出了DARAG(数据和检索增强生成纠错),一种新的方法,旨在提高GEC域(ID)和OOD场景中的ASR。我们使用通过提示LLM和文本到语音模型生成的合成数据来增强GEC训练数据集,从而模拟模型可以学习的其他错误。对于面向对象的场景,我们模拟测试时的错误,从新的领域类似,并在一个无监督的方式。此外,为了更好地处理命名实体,我们引入检索增强校正,通过从数据库中检索的实体来增强输入。我们的方法是简单的,可扩展的,域和语言无关。我们在多个数据集和设置上进行了实验,表明DARAG优于我们所有的基线,在ID和OOD设置中分别实现了8%-30%和10%-33%的相对WER改进。摘要:Generative Error Correction (GEC) has emerged as a powerful post-processing method to enhance the performance of Automatic Speech Recognition (ASR) systems. However, we show that GEC models struggle to generalize beyond the specific types of errors encountered during training, limiting their ability to correct new, unseen errors at test time, particularly in out-of-domain (OOD) scenarios. This phenomenon amplifies with named entities (NEs), where, in addition to insufficient contextual information or knowledge about the NEs, novel NEs keep emerging. To address these issues, we propose DARAG (Data- and Retrieval-Augmented Generative Error Correction), a novel approach designed to improve GEC for ASR in in-domain (ID) and OOD scenarios. We augment the GEC training dataset with synthetic data generated by prompting LLMs and text-to-speech models, thereby simulating additional errors from which the model can learn. For OOD scenarios, we simulate test-time errors from new domains similarly and in an unsupervised fashion. Additionally, to better handle named entities, we introduce retrieval-augmented correction by augmenting the input with entities retrieved from a database. Our approach is simple, scalable, and both domain- and language-agnostic. We experiment on multiple datasets and settings, showing that DARAG outperforms all our baselines, achieving 8 % -- 30 % relative WER improvements in ID and 10 % -- 33 % improvements in OOD settings.

【10】 Using RLHF to align speech enhancement approaches to mean-opinion quality scores
标题: 使用WLHF将语音增强方法与平均意见质量分数保持一致
作者: Anurag Kumar, Andrew Perrault, Donald S. Williamson
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:客观语音质量测量通常用于评估语音增强算法,但已经表明,它们作为学习目标是次优的,因为它们并不总是与人类主观评级很好地一致。这种未对准通常会导致明显的失真和伪像,从而导致语音增强无效。为了解决这些问题,我们提出了一个从人类反馈中强化学习(RLHF)框架,通过使用基于平均意见得分(MOS)的奖励模型优化性能来微调现有的语音增强方法。我们的研究结果表明,RLHF微调模型具有最好的性能在不同的基准上的客观和基于MOS的语音质量评估指标的语音银行+DEMAND数据集。通过消融研究,我们表明,策略梯度损失和监督MSE损失对于不同指标的平衡优化都很重要。摘要:Objective speech quality measures are typically used to assess speech enhancement algorithms, but it has been shown that they are sub-optimal as learning objectives because they do not always align well with human subjective ratings. This misalignment often results in noticeable distortions and artifacts that cause speech enhancement to be ineffective. To address these issues, we propose a reinforcement learning from human feedback (RLHF) framework to fine-tune an existing speech enhancement approach by optimizing performance using a mean-opinion score (MOS)-based reward model. Our results show that the RLHF-finetuned model has the best performance across different benchmarks for both objective and MOS-based speech quality assessment metrics on the Voicebank+DEMAND dataset. Through ablation studies, we show that both policy gradient loss and supervised MSE loss are important for balanced optimization across the different metrics.

【11】 Multi-View Multi-Task Modeling with Speech Foundation Models for Speech Forensic Tasks
标题: 使用语音基础模型的多视图多任务建模用于语音取证任务
作者: Orchid Chetia Phukan, Devyani Koshal, Swarup Ranjan Behera, Arun Balaji Buduru, Rajesh Sharma
链接:点击下载PDF文件
摘要:语音取证任务(SFT),诸如自动说话人识别(ASR)、语音情感识别(SER)、性别识别(GR)和年龄估计(AE),在不同的安全和生物特征应用中找到用途。以前的工作已经应用了各种技术,最近的研究集中在应用语音基础模型(SFM),以提高性能。然而,大多数先前的努力都集中在为每个任务单独构建单独的模型上,尽管这些任务之间具有内在的相似性。这种孤立的方法会导致更高的计算资源需求,增加成本,时间消耗和维护挑战。在这项研究中,我们通过采用多任务学习策略来应对这些挑战。首先,我们探讨了各种国家的最先进的(SOTA)SFM提取他们的表示学习这些SFT和调查他们的有效性,在每个任务具体。其次,我们分析了在多任务学习框架中提取的表征在SFT上的性能。我们观察到的性能下降时,SFT一起建模相比,个别的任务特定的模型,作为一种补救措施,我们提出了多视图学习(MVL)。视图是来自不同的SFM的表示,通过每个SFM独有的特征转换成不同的抽象空间。通过利用MVL,我们整合了这些不同的表示,以捕获跨任务的互补信息,增强共享学习过程。我们引入了一个新的框架,称为TANGO(任务对齐与Inter-view门控最优运输)来实现这种方法。通过TANGO,我们实现了与单个SFM表示以及基准数据集(如CREMA-D,emo-DB和BAVED)的基线融合技术相比的最高性能。摘要:Speech forensic tasks (SFTs), such as automatic speaker recognition (ASR), speech emotion recognition (SER), gender recognition (GR), and age estimation (AE), find use in different security and biometric applications. Previous works have applied various techniques, with recent studies focusing on applying speech foundation models (SFMs) for improved performance. However, most prior efforts have centered on building individual models for each task separately, despite the inherent similarities among these tasks. This isolated approach results in higher computational resource requirements, increased costs, time consumption, and maintenance challenges. In this study, we address these challenges by employing a multi-task learning strategy. Firstly, we explore the various state-of-the-art (SOTA) SFMs by extracting their representations for learning these SFTs and investigating their effectiveness at each task specifically. Secondly, we analyze the performance of the extracted representations on the SFTs in a multi-task learning framework. We observe a decline in performance when SFTs are modeled together compared to individual task-specific models, and as a remedy, we propose multi-view learning (MVL). Views are representations from different SFMs transformed into distinct abstract spaces by characteristics unique to each SFM. By leveraging MVL, we integrate these diverse representations to capture complementary information across tasks, enhancing the shared learning process. We introduce a new framework called TANGO (Task Alignment with iNter-view Gated Optimal transport) to implement this approach. With TANGO, we achieve the topmost performance in comparison to individual SFM representations as well as baseline fusion techniques across benchmark datasets such as CREMA-D, emo-DB, and BAVED.

【12】 AI-Enhanced Acoustic Analysis for Comprehensive Biodiversity Monitoring and Assessment
标题: 用于全面生物多样性监测和评估的人工智能增强声学分析
作者: Kumar Srinivas Bobba, Kartheeban K, Vamsi Krishna Sai, Dinesh Bugga, Vijaya Mani Surendra Bolla
链接:点击下载PDF文件
摘要:该项目提议开发一个全面的实时生物多样性监测系统,通过声学传感器网络和先进的人工智能算法利用声音数据。该系统分析来自各种生态系统的声音记录,以识别和分类不同的物种,为生态系统健康和生物多样性模式提供有价值的见解,同时有助于检测物种存在和行为随时间的微妙变化。通过应对噪音污染和物种重叠等关键挑战,该系统采用了复杂的过滤和分类技术,以确保准确可靠的监测,区分自然声音和人为噪音。最终,这一举措旨在提高我们对生物多样性动态的理解,并提供必要的信息,以支持有效的保护战略,并为政策决策提供信息,为利益相关者提供可操作的见解,以保护和维护重要的生态系统。摘要:This project proposes the development of a comprehensive real-time biodiversity monitoring system that harnesses sound data through a network of acoustic sensors and advanced artificial intelligence algorithms. The system analyzes sound recordings from various ecosystems to identify and classify different species, providing valuable insights into ecosystem health and biodiversity patterns while facilitating the detection of subtle changes in species presence and behavior over time. By addressing critical challenges such as noise pollution and species overlap, the system employs sophisticated filtering and classification techniques to ensure accurate and reliable monitoring, distinguishing between natural sounds and anthropogenic noise. Ultimately, this initiative aims to enhance our understanding of biodiversity dynamics and provide essential information to support effective conservation strategies and inform policy decisions, empowering stakeholders with actionable insights to protect and preserve vital ecosystems.

【13】 Exploiting Longitudinal Speech Sessions via Voice Assistant Systems for Early Detection of Cognitive Decline
标题: 通过语音助理系统利用纵向语音会话来早期检测认知下降
作者: Kristin Qi, Jiatong Shi, Caroline Summerour, John A. Batsis, Xiaohui Liang
备注:IEEE International Conference on E-health Networking, Application & Services
链接:点击下载PDF文件
摘要:轻度认知障碍(MCI)是阿尔茨海默病(AD)的早期阶段,AD是一种神经退行性疾病。MCI的早期识别对于通过及时干预延缓其进展至关重要。现有的研究已经证明了使用从临床访谈或数字设备收集的语音检测MCI的可行性。然而,这些方法通常分析在有限时间点收集的数据,限制了它们识别随时间推移的认知变化的能力。本文介绍了一项纵向研究,使用语音助理系统(VAS)在18个月内每隔3个月远程收集7次会话语音数据。我们提出了两种方法来提高MCI检测和认知变化的预测。第一种方法结合了历史数据,而第二种方法预测了两个时间点的认知变化。我们的研究结果表明,在结合历史数据时,MCI检测的平均F1分数在声学特征的情况下从58.6%提高到71.2%(提高了12.6%),在语言特征的情况下从62.1%提高到75.1%(提高了13.0%)。此外,在声学特征的情况下,认知变化的预测达到了73.7%的F1分数。这些结果证实了基于VAS的语音会话用于早期检测认知衰退的潜力。摘要:Mild Cognitive Impairment (MCI) is an early stage of Alzheimer's disease (AD), a form of neurodegenerative disorder. Early identification of MCI is crucial for delaying its progression through timely interventions. Existing research has demonstrated the feasibility of detecting MCI using speech collected from clinical interviews or digital devices. However, these approaches typically analyze data collected at limited time points, limiting their ability to identify cognitive changes over time. This paper presents a longitudinal study using voice assistant systems (VAS) to remotely collect seven-session speech data at three-month intervals across 18 months. We propose two methods to improve MCI detection and the prediction of cognitive changes. The first method incorporates historical data, while the second predicts cognitive changes at two time points. Our results indicate improvements when incorporating historical data: the average F1-score for MCI detection improves from 58.6% to 71.2% (by 12.6%) in the case of acoustic features and from 62.1% to 75.1% (by 13.0%) in the case of linguistic features. Additionally, the prediction of cognitive changes achieves an F1-score of 73.7% in the case of acoustic features. These results confirm the potential of VAS-based speech sessions for early detection of cognitive decline.

【14】 Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
标题: 利用多令牌预测和推测解码加速基于编解码的语音合成
作者: Tan Dat Nguyen, Ji-Hoon Kim, Jeongsoo Choi, Shukjae Choi, Jinseok Park, Younglo Lee, Joon Son Chung
备注:Submitted to IEEE ICASSP 2025
链接:点击下载PDF文件
摘要:本文的目标是加速基于编解码器的语音合成系统,以最小的牺牲语音质量。我们提出了一种增强的推理方法,允许在推理过程中的速度和质量之间的灵活权衡,而不需要额外的训练。我们的核心思想是使用多个预测头来预测AR模块的每个推理步骤的多个令牌,从而随着头的数量增加而线性减少合成时间。此外,我们引入了一种新的推测性解码技术,利用维特比为基础的算法来选择在每个解码步骤生成的令牌的最佳序列。在我们的实验中,我们证明了预测每个令牌所需的时间与基线模型相比减少了4到5倍,并且在语音清晰度方面具有最小的质量权衡甚至改善。音频样本可在multpletokensprediction.github.io multipletokensprediction.github.io 上获得。摘要:The goal of this paper is to accelerate codec-based speech synthesis systems with minimum sacrifice to speech quality. We propose an enhanced inference method that allows for flexible trade-offs between speed and quality during inference without requiring additional training. Our core idea is to predict multiple tokens per inference step of the AR module using multiple prediction heads, resulting in a linear reduction in synthesis time as the number of heads increases. Furthermore, we introduce a novel speculative decoding technique that utilises a Viterbi-based algorithm to select the optimal sequence of generated tokens at each decoding step. In our experiments, we demonstrate that the time required to predict each token is reduced by a factor of 4 to 5 compared to baseline models, with minimal quality trade-off or even improvement in terms of speech intelligibility. Audio samples are available at: multpletokensprediction.github.io multipletokensprediction.github.io .

【15】 Dynamic Range Compression and Its Effect on Music Genre Classification
标题: 动态范围压缩及其对音乐流派分类的影响
作者: Arlyn Reese Madsen III
链接:点击下载PDF文件
摘要:本文研究了动态范围压缩(DRC)对音乐流派分类精度的影响。通过对200首歌曲的测试集应用各种压缩设置,我们的目标是确定压缩是否可以提高分类器辨别不同音乐流派的能力。支持向量机(SVM)分类器在原始的未压缩数据集上进行训练。研究了阈值、比率、膝宽、攻击时间、释放时间和补偿增益对分类性能的影响。我们的研究结果表明,对测试集进行压缩确实可以将音乐流派分类的准确率平均提高3.1%。最佳压缩设置在实验中有所不同,这表明压缩的有效性取决于模型的训练数据。提供了超过1000个训练和测试分段的最高压缩设置表。总之,这项研究表明,动态范围压缩可以作为一个有价值的预处理技术,提高音乐流派分类。从这项研究中获得的见解可以为开发更准确、更强大的音乐推荐系统提供信息。摘要:This paper investigates the impact of dynamic range compression (DRC) on music genre classification accuracy. By applying various compression settings to the test set of 200 songs, we aim to determine if compression can enhance the classifier's ability to discern distinct musical genres. A support vector machine (SVM) classifier was trained on the original, uncompressed dataset. The study explored the influence of threshold, ratio, knee width, attack time, release time, and makeup gain on classification performance. Our findings indicate that applying compression to the test set can indeed improve music genre classification accuracy on average by 3.1%. The optimal compression settings varied across experiments, suggesting that the effectiveness of compression depends on the training data of the model. A table of the top compression settings over 1000 train and test splits is provided. In conclusion, this research demonstrates that dynamic range compression can serve as a valuable preprocessing technique for enhancing music genre classification. The insights gained from this study can inform the development of more accurate and robust music recommendation systems.

【16】 Parameter-efficient Adaptation of Multilingual Multimodal Models for Low-resource ASR
标题: 针对低资源ASB的多语言多模式模型的参数高效自适应
作者: Abhishek Gupta, Amruta Parulekar, Sameep Chattopadhyay, Preethi Jyothi
链接:点击下载PDF文件
摘要:低资源语言的自动语音识别(ASR)仍然是一个挑战,由于缺乏标记的训练数据。参数高效微调和纯文本自适应是两种用于解决此类低资源设置的流行方法。在这项工作中,我们研究了如何使用多语言多模态模型(如无障碍M4T)将这些技术有效地结合起来。多模态模型能够通过纯文本自适应来利用未标记的文本,并进一步进行参数高效的ASR微调,从而提高ASR性能。我们还显示了从高资源语言的跨语言迁移,在没有任何标记语音的zero-shot设置中,相对于基线,WER降低了17%。摘要:Automatic speech recognition (ASR) for low-resource languages remains a challenge due to the scarcity of labeled training data. Parameter-efficient fine-tuning and text-only adaptation are two popular methods that have been used to address such low-resource settings. In this work, we investigate how these techniques can be effectively combined using a multilingual multimodal model like SeamlessM4T. Multimodal models are able to leverage unlabeled text via text-only adaptation with further parameter-efficient ASR fine-tuning, thus boosting ASR performance. We also show cross-lingual transfer from a high-resource language, achieving up to a relative 17% WER reduction over a baseline in a zero-shot setting without any labeled speech.

【17】 MeloTrans: A Text to Symbolic Music Generation Model Following Human Composition Habit
标题: MeloTrans:遵循人类作曲习惯的符号音乐生成模型的文本
作者: Yutian Wang, Wanyin Yang, Zhenrong Dai, Yilong Zhang, Kun Zhao, Hui Wang
链接:点击下载PDF文件
摘要:None摘要:At present, neural network models show powerful sequence prediction ability and are used in many automatic composition models. In comparison, the way humans compose music is very different from it. Composers usually start by creating musical motifs and then develop them into music through a series of rules. This process ensures that the music has a specific structure and changing pattern. However, it is difficult for neural network models to learn these composition rules from training data, which results in a lack of musicality and diversity in the generated music. This paper posits that integrating the learning capabilities of neural networks with human-derived knowledge may lead to better results. To archive this, we develop the POP909$ _$M dataset, the first to include labels for musical motifs and their variants, providing a basis for mimicking human compositional habits. Building on this, we propose MeloTrans, a text-to-music composition model that employs principles of motif development rules. Our experiments demonstrate that MeloTrans excels beyond existing music generation models and even surpasses Large Language Models (LLMs) like ChatGPT-4. This highlights the importance of merging human insights with neural network capabilities to achieve superior symbolic music generation.

【18】 Enhancing 1-Second 3D SELD Performance with Filter Bank Analysis and SCConv Integration in CST-Former
标题: 通过CST-Former中的过滤器组分析和SCConv集成增强1秒3D SELD性能
作者: Zhehui Zhang
链接:点击下载PDF文件
摘要:最近的SELD研究主要集中在长时间段场景(通常为5到10秒,偶尔为2秒),提高了基准性能,但缺乏真实世界应用所需的时间粒度。为了弥补这一差距,本文研究了短时间段下的SELD与距离估计(3D SELD)系统,特别是针对1秒的窗口,建立了一个新的基线实际3D SELD的适用性。我们进一步探讨了不同滤波器组的影响-巴克,梅尔,和Gammatone音频特征提取,实验结果表明,Gammatone滤波器实现了最高的整体精度在这种情况下。最后,我们建议用SCConv模块替换CST-Former(一种有竞争力的SELD架构)中的卷积模块。这种调整在短片段场景中产生可测量的F分数增益,强调了SCConv改进空间和通道特征表示的潜力。实验结果突出表明,我们的方法是在低延迟约束下实现3D SELD系统在现实世界中部署的重要一步。摘要:Recent SELD research has predominantly focused on long-time segment scenarios (typically 5 to 10 seconds, occasionally 2 seconds), improving benchmark performance but lacking the temporal granularity needed for real-world applications. To bridge this gap, this paper investigates SELD with distance estimation (3D SELD) systems under short-time segments, specifically targeting a 1-second window, establishing a new baseline for practical 3D SELD applicability. We further explore the impact of different filter banks -- Bark, Mel, and Gammatone for audio feature extraction, and experimental results demonstrate that the Gammatone filter achieves the highest overall accuracy in this context. Finally, we propose replacing the convolutional modules within the CST-Former, a competitive SELD architecture, with the SCConv module. This adjustment yields measurable F-score gains in short-segment scenarios, underscoring SCConv's potential to improve spatial and channel feature representation. The experimental results highlight our approach as a significant step towards the real-world deployment of 3D SELD systems under low-latency constraints.

【19】 End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
标题: 使用自我监督学习功能实现语音情感识别与语音活动检测的端到端集成
作者: Natsuo Yamashita, Masaaki Yamamoto, Yohei Kawaguchi
链接:点击下载PDF文件
摘要:语音情感识别(SER)通常对由语音活动检测(VAD)模型检测到的语音段进行操作。然而,VAD模型可能会输出有缺陷的语音段,特别是在嘈杂的环境中,导致后续SER模型的性能下降。为了解决这个问题,我们提出了一个端到端(E2E)的方法,集成VAD和SER使用自监督学习(SSL)功能。VAD模块首先接收SSL特征作为输入,然后将分段的SSL特征馈送到SER模块中。VAD和SER模块都经过联合训练,以优化SER性能。IEMOCAP数据集上的实验结果表明,我们提出的方法提高了SER性能。此外,为了研究我们所提出的方法对VAD和SSL模块的影响,我们对VAD输出和SSL编码器的每一层的权重进行了分析。摘要:Speech Emotion Recognition (SER) often operates on speech segments detected by a Voice Activity Detection (VAD) model. However, VAD models may output flawed speech segments, especially in noisy environments, resulting in degraded performance of subsequent SER models. To address this issue, we propose an end-to-end (E2E) method that integrates VAD and SER using Self-Supervised Learning (SSL) features. The VAD module first receives the SSL features as input, and the segmented SSL features are then fed into the SER module. Both the VAD and SER modules are jointly trained to optimize SER performance. Experimental results on the IEMOCAP dataset demonstrate that our proposed method improves SER performance. Furthermore, to investigate the effect of our proposed method on the VAD and SSL modules, we present an analysis of the VAD outputs and the weights of each layer of the SSL encoder.

【20】 Roadmap towards Superhuman Speech Understanding using Large Language Models
标题: 使用大型语言模型实现超人语音理解的路线图
作者: Fan Bu, Yuhao Zhang, Xidong Wang, Benyou Wang, Qun Liu, Haizhou Li
链接:点击下载PDF文件
摘要:大型语言模型(LLM)的成功促使人们努力整合语音和音频数据,旨在创建能够处理文本和非文本输入的通用基础模型。最近的进展,如GPT-4 o,突出了端到端语音LLM的潜力,它保留了非语义信息和世界知识,以进行更深入的语音理解。为了指导语音LLM的发展,我们提出了一个五级路线图,从基本的自动语音识别(ASR)到高级超人模型,这些模型能够将非语义信息与抽象的声学知识相结合,以完成复杂的任务。此外,我们设计了一个基准,SAGI Bechmark,在这五个级别的各种任务的关键方面,发现在使用抽象的声学知识和能力的完整性的挑战。我们的研究结果揭示了处理非语言线索和抽象声学知识的差距,我们提供了未来的方向。本文概述了推进语音LLM的路线图,介绍了一个评估基准,并提供了其当前的局限性和潜力的关键见解。摘要:The success of large language models (LLMs) has prompted efforts to integrate speech and audio data, aiming to create general foundation models capable of processing both textual and non-textual inputs. Recent advances, such as GPT-4o, highlight the potential for end-to-end speech LLMs, which preserves non-semantic information and world knowledge for deeper speech understanding. To guide the development of speech LLMs, we propose a five-level roadmap, ranging from basic automatic speech recognition (ASR) to advanced superhuman models capable of integrating non-semantic information with abstract acoustic knowledge for complex tasks. Moreover, we design a benchmark, SAGI Bechmark, that standardizes critical aspects across various tasks in these five levels, uncovering challenges in using abstract acoustic knowledge and completeness of capability. Our findings reveal gaps in handling paralinguistic cues and abstract acoustic knowledge, and we offer future directions. This paper outlines a roadmap for advancing speech LLMs, introduces a benchmark for evaluation, and provides key insights into their current limitations and potential.

【21】 CLaMP 2: Multimodal Music Information Retrieval Across 101 Languages Using Large Language Models
标题: CLaMP 2:使用大型语言模型跨101种语言的多模式音乐信息检索
作者: Shangda Wu, Yashan Wang, Ruibin Yuan, Zhancheng Guo, Xu Tan, Ge Zhang, Monan Zhou, Jing Chen, Xuefeng Mu, Yuejie Gao, Yuanliang Dong, Jiafeng Liu, Xiaobing Li, Feng Yu, Maosong Sun
备注:17 pages, 10 figures, 4 tables
链接:点击下载PDF文件
摘要:当前的音乐信息检索系统面临着管理语言多样性和整合各种音乐形式的挑战。这些限制降低了它们在全球多模态音乐环境中的有效性。为了解决这些问题,我们介绍了CLaMP 2,一个系统兼容101种语言,支持ABC记谱法(基于文本的乐谱格式)和MIDI(乐器数字接口)的音乐信息检索。CLaMP 2在150万个ABC文本三元组上进行了预训练,包括通过对比学习对齐的多语言文本编码器和多模式音乐编码器。通过利用大型语言模型,我们可以大规模地获得精确且一致的多语言描述,从而显著减少文本噪音并平衡语言分布。我们的实验表明,CLaMP 2在多语言语义搜索和跨模态音乐分类方面都取得了最先进的结果,从而建立了包容性和全球音乐信息检索的新标准。摘要:Challenges in managing linguistic diversity and integrating various musical modalities are faced by current music information retrieval systems. These limitations reduce their effectiveness in a global, multimodal music environment. To address these issues, we introduce CLaMP 2, a system compatible with 101 languages that supports both ABC notation (a text-based musical notation format) and MIDI (Musical Instrument Digital Interface) for music information retrieval. CLaMP 2, pre-trained on 1.5 million ABC-MIDI-text triplets, includes a multilingual text encoder and a multimodal music encoder aligned via contrastive learning. By leveraging large language models, we obtain refined and consistent multilingual descriptions at scale, significantly reducing textual noise and balancing language distribution. Our experiments show that CLaMP 2 achieves state-of-the-art results in both multilingual semantic search and music classification across modalities, thus establishing a new standard for inclusive and global music information retrieval.

【22】 EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
标题: EH-MAM:用于自我监督语音表示学习的易于操作的掩蔽声学建模
作者: Ashish Seth, Ramaneswaran Selvakumar, S Sakshi, Sonal Kumar, Sreyan Ghosh, Dinesh Manocha
链接:点击下载PDF文件
摘要:在本文中,我们提出了EH-MAM(易到硬自适应掩蔽声学建模),一种新的自监督学习方法的语音表示学习。在以往的方法,使用随机掩蔽方案掩蔽声学建模(MAM)相比,我们引入了一种新的选择性和自适应掩蔽策略。具体来说,在SSL训练期间,我们逐步将较硬的区域引入模型进行重建。我们的方法自动选择硬区域,并建立在观察到的MAM中的各个帧的重建损失可以提供自然的信号来判断解决该帧的MAM预文本任务的难度。为了识别这些硬区域,我们采用了一个教师模型,该模型首先预测帧的损失,然后决定屏蔽哪些帧。通过学习创建具有挑战性的问题,例如识别更难的帧并同时解决它们,模型能够学习更有效的表示,从而获得对语音的更全面的理解。在量化方面,EH-MAM在各种低资源语音识别和SUPERB基准测试中的表现优于几个最先进的基线5%-10%。此外,我们进行了全面的分析,表明EH-MAM掩蔽的区域有效地捕捉有用的上下文跨语音帧。摘要:In this paper, we present EH-MAM (Easy-to-Hard adaptive Masked Acoustic Modeling), a novel self-supervised learning approach for speech representation learning. In contrast to the prior methods that use random masking schemes for Masked Acoustic Modeling (MAM), we introduce a novel selective and adaptive masking strategy. Specifically, during SSL training, we progressively introduce harder regions to the model for reconstruction. Our approach automatically selects hard regions and is built on the observation that the reconstruction loss of individual frames in MAM can provide natural signals to judge the difficulty of solving the MAM pre-text task for that frame. To identify these hard regions, we employ a teacher model that first predicts the frame-wise losses and then decides which frames to mask. By learning to create challenging problems, such as identifying harder frames and solving them simultaneously, the model is able to learn more effective representations and thereby acquire a more comprehensive understanding of the speech. Quantitatively, EH-MAM outperforms several state-of-the-art baselines across various low-resource speech recognition and SUPERB benchmarks by 5%-10%. Additionally, we conduct a thorough analysis to show that the regions masked by EH-MAM effectively capture useful context across speech frames.

【23】 Sound Check: Auditing Audio Datasets
标题: 声音检查:审核音频数据集
作者: William Agnew, Julia Barnett, Annie Chu, Rachel Hong, Michael Feffer, Robin Netzorg, Harry H. Jiang, Ezra Awumey, Sauvik Das
链接:点击下载PDF文件
摘要:生成音频模型在功能和公共使用方面都在迅速发展-几个强大的生成音频模型已经提供了开放权重,一些科技公司已经发布了高质量的生成音频产品。然而,虽然先前的工作已经列举了许多来自生成视觉和文本模型训练数据的伦理问题,但我们对生成音频数据集的类似问题知之甚少,包括与偏见,毒性和知识产权有关的问题。为了弥补这一差距,我们对数百个音频数据集进行了文献综述,并选择了其中七个最突出的数据集进行更详细的审计。我们发现这些数据集对女性有偏见,包含对边缘化社区的有害刻板印象,并包含大量受版权保护的作品。为了使艺术家能够看到他们是否在流行的音频数据集中,并促进对这些数据集内容的探索,我们在https: audio-audit.vercel.app上开发了一个Web工具音频数据集探索工具。摘要:Generative audio models are rapidly advancing in both capabilities and public utilization -- several powerful generative audio models have readily available open weights, and some tech companies have released high quality generative audio products. Yet, while prior work has enumerated many ethical issues stemming from the data on which generative visual and textual models have been trained, we have little understanding of similar issues with generative audio datasets, including those related to bias, toxicity, and intellectual property. To bridge this gap, we conducted a literature review of hundreds of audio datasets and selected seven of the most prominent to audit in more detail. We found that these datasets are biased against women, contain toxic stereotypes about marginalized communities, and contain significant amounts of copyrighted work. To enable artists to see if they are in popular audio datasets and facilitate exploration of the contents of these datasets, we developed a web tool audio datasets exploration tool at https: audio-audit.vercel.app.

【24】 AADNet: An End-to-End Deep Learning Model for Auditory Attention Decoding
标题: AADNet:用于听觉注意力解码的端到端深度学习模型
作者: Nhan Duc Thanh Nguyen, Huy Phan, Simon Geirnaert, Kaare Mikkelsen, Preben Kidmose
备注:11 pages, 6 figures
链接:点击下载PDF文件
摘要:听觉注意解码(AAD)是在多人说话环境中使用大脑信号(通常通过脑电图(EEG)记录)识别注意语音的过程。在过去的十年中,AAD经历了不断的发展,这是由于其在神经导向听力设备中的应用前景。与无人值守语音相比,大多数AAD算法依赖于神经夹带到有人值守语音的包络的增加,通常使用两步方法。首先,该算法预测表示的出席语音信号包络,其次,它确定出席语音通过找到最高的预测和实际语音信号的表示之间的相关性。在这项研究中,我们提出了一种新的端到端神经网络架构,命名为AADNet,它将这两个阶段结合起来,直接解决AAD问题。我们将所提出的网络与传统方法进行比较,包括线性刺激重建,典型相关分析,以及使用两个不同数据集的替代非线性刺激重建。AADNet在特定于主题和独立于主题的模型中表现出显着的性能改进。值得注意的是,分析窗口长度分别为1秒至40秒的情况下,平均独立于受试者的分类准确度为56.1%至82.7%,这表明概括来自未见过受试者的数据的能力显著提高。这些结果突出了深度学习模型在推进AAD方面的潜力,对未来的助听器,辅助设备和临床评估具有重要意义。摘要:Auditory attention decoding (AAD) is the process of identifying the attended speech in a multi-talker environment using brain signals, typically recorded through electroencephalography (EEG). Over the past decade, AAD has undergone continuous development, driven by its promising application in neuro-steered hearing devices. Most AAD algorithms are relying on the increase in neural entrainment to the envelope of attended speech, as compared to unattended speech, typically using a two-step approach. First, the algorithm predicts representations of the attended speech signal envelopes; second, it identifies the attended speech by finding the highest correlation between the predictions and the representations of the actual speech signals. In this study, we proposed a novel end-to-end neural network architecture, named AADNet, which combines these two stages into a direct approach to address the AAD problem. We compare the proposed network against the traditional approaches, including linear stimulus reconstruction, canonical correlation analysis, and an alternative non-linear stimulus reconstruction using two different datasets. AADNet shows a significant performance improvement for both subject-specific and subject-independent models. Notably, the average subject-independent classification accuracies from 56.1 % to 82.7 % with analysis window lengths ranging from 1 to 40 seconds, respectively, show a significantly improved ability to generalize to data from unseen subjects. These results highlight the potential of deep learning models for advancing AAD, with promising implications for future hearing aids, assistive devices, and clinical assessments.

【25】 MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization
标题: MuVi:具有语义对齐和节奏同步的视频到音乐生成
作者: Ruiqi Li, Siqi Zheng, Xize Cheng, Ziang Zhang, Shengpeng Ji, Zhou Zhao
备注:Working in progress
链接:点击下载PDF文件
摘要:生成与视频的视觉内容一致的音乐一直是一项具有挑战性的任务,因为它需要对视觉语义的深刻理解,并且涉及生成旋律,节奏和动态与视觉叙事相协调的音乐。本文介绍了MuVi,这是一个新颖的框架,可以有效地解决这些挑战,以增强视听内容的凝聚力和沉浸式体验。MuVi通过专门设计的视觉适配器分析视频内容,以提取上下文和时间相关的特征。这些功能用于生成音乐,不仅匹配视频的情绪和主题,还匹配其节奏和节奏。我们还引入了一个对比的音乐视觉预训练计划,以确保同步,基于音乐短语的周期性。此外,我们证明了我们的基于流匹配的音乐生成器具有上下文学习能力,使我们能够控制所生成的音乐的风格和流派。实验结果表明,MuVi在音频质量和时间同步方面表现出优异的性能。生成的音乐视频样本可在https: muvi-v2m.github.io上获得。摘要:Generating music that aligns with the visual content of a video has been a challenging task, as it requires a deep understanding of visual semantics and involves generating music whose melody, rhythm, and dynamics harmonize with the visual narratives. This paper presents MuVi, a novel framework that effectively addresses these challenges to enhance the cohesion and immersive experience of audio-visual content. MuVi analyzes video content through a specially designed visual adaptor to extract contextually and temporally relevant features. These features are used to generate music that not only matches the video's mood and theme but also its rhythm and pacing. We also introduce a contrastive music-visual pre-training scheme to ensure synchronization, based on the periodicity nature of music phrases. In addition, we demonstrate that our flow-matching-based music generator has in-context learning ability, allowing us to control the style and genre of the generated music. Experimental results show that MuVi demonstrates superior performance in both audio quality and temporal synchronization. The generated music video samples are available at https: muvi-v2m.github.io.

【26】 Towards Computational Analysis of Pansori Singing
标题: 潘索里歌唱的计算分析
作者: Sangheon Park, Danbinaerin Han, Dasaem Jeong
备注:Late-Breaking Demo Session of the 25th International Society for Music Information Retrieval (ISMIR) Conference, 2024
链接:点击下载PDF文件
摘要:盘索里是韩国传统音乐中最具代表性的声乐体裁之一,其声乐旋律线细腻,颤音强烈。虽然音乐是口头传播的,没有任何乐谱,但用西方五线谱转录潘索里音乐有几个目的,如音乐文献,教育或研究。在本文中,我们介绍了基于音频和相应的转录的pansori的计算分析,现代音乐信息检索任务如何可以用于分析传统的音乐,以及它如何揭示pansori包含的不同的音频特征。摘要:Pansori is one of the most representative vocal genres of Korean traditional music, which has an elaborated vocal melody line with strong vibrato. Although the music is transmitted orally without any music notation, transcribing pansori music in Western staff notation has been introduced for several purposes, such as documentation of music, education, or research. In this paper, we introduce computational analysis of pansori based on both audio and corresponding transcription, how modern Music Information Retrieval tasks can be used in analyzing traditional music and how it revealed different audio characteristics of what pansori contains.

【27】 What Do Speech Foundation Models Not Learn About Speech?
标题: Speech Foundation模型对Speech不了解什么?
作者: Abdul Waheed, Hanin Atwany, Bhiksha Raj, Rita Singh
备注:20 Pages
链接:点击下载PDF文件
摘要:理解语音基础模型如何捕捉非语言线索对于提高其在不同任务中的可解释性和适应性至关重要。在我们的工作中,我们分析了几个突出的模型,如Whisper,Seamless,Wav 2 Vec,HuBERT和Qwen 2-Audio,重点关注他们在Dynamic-SUPERB基准测试中学习到的语言和非语言任务的表示。我们的研究解决了三个关键问题:(1)什么非语言线索(例如,说话者意图、情感、环境背景)被捕获?(2)这些线索如何在模型的不同层中表示?以及(3)这些表征在多大程度上可以有效地适应下游任务?为了回答这些问题,我们首先在zero-shot设置中评估模型,然后对从这些模型中提取的逐层特征进行微调。我们的研究结果提供了深入了解模型的泛化能力,其分层表示的特点,以及下游任务适应所需的转换程度。我们的研究结果表明,其中一些模型在zero-shot设置中的各种任务上表现良好,尽管没有针对这些任务进行明确的训练。我们还观察到,zero-shot性能与更好地学习表示。逐层特征的分析表明,一些模型在学习表示的可分性和模型深度之间表现出凸关系,不同的层捕获特定于任务的特征。摘要:Understanding how speech foundation models capture non-verbal cues is crucial for improving their interpretability and adaptability across diverse tasks. In our work, we analyze several prominent models such as Whisper, Seamless, Wav2Vec, HuBERT, and Qwen2-Audio focusing on their learned representations in both paralinguistic and non-paralinguistic tasks from the Dynamic-SUPERB benchmark. Our study addresses three key questions: (1) What non-verbal cues (e.g., speaker intent, emotion, environmental context) are captured? (2) How are these cues represented across different layers of the models? and (3) To what extent can these representations be effectively adapted to downstream tasks? To answer these questions, we first evaluate the models in a zero-shot setting, followed by fine-tuning on layer-wise features extracted from these models. Our results provide insights into the models' capacity for generalization, the characteristics of their layer-wise representations, and the degree of transformation required for downstream task adaptation. Our findings suggest that some of these models perform well on various tasks in zero-shot settings, despite not being explicitly trained for those tasks. We also observe that zero-shot performance correlates with better-learned representations. The analysis of layer-wise features demonstrates that some models exhibit a convex relationship between the separability of the learned representations and model depth, with different layers capturing task-specific features.

【28】 Towards Homogeneous Lexical Tone Decoding from Heterogeneous Intracranial Recordings
标题: 从不均匀的脑内记录中实现同质的语调解码
作者: Di Wu, Siyuan Li, Chen Feng, Lu Cao, Yue Zhang, Jie Yang, Mohamad Sawan
备注:Preprint V1 with 10 pages main text
链接:点击下载PDF文件
摘要:脑机接口(BCI)的最新进展已经能够从颅内记录中解码词汇音调,为恢复语音受损的音调语言说话者的沟通能力提供了可能。然而,由生理和仪器因素引起的数据异质性对统一的侵入性脑音解码提出了重大挑战。传统的特定于主题的模型在异构解码范式下运行,无法捕获广义神经表征,并且无法有效地利用跨主题的数据。为了解决这些局限性,我们引入了神经表征的同质性-异质性分解学习(H2 DiLR),这是一种新的框架,可以从多个受试者的颅内记录中分解和学习同质性和异质性。为了评估H2 DiLR,我们收集了来自多名参与者的立体脑电图(sEEG)数据,这些参与者阅读包括407个音节的普通话材料,几乎代表了所有的普通话字符。大量实验表明,H2 DiLR作为一种统一的解码范式,其性能显着优于传统的异构解码方法。此外,我们的经验证实,H2 DiLR有效地捕捉神经表征学习过程中的同质性和异质性。摘要:Recent advancements in brain-computer interfaces (BCIs) have enabled the decoding of lexical tones from intracranial recordings, offering the potential to restore the communication abilities of speech-impaired tonal language speakers. However, data heterogeneity induced by both physiological and instrumental factors poses a significant challenge for unified invasive brain tone decoding. Traditional subject-specific models, which operate under a heterogeneous decoding paradigm, fail to capture generalized neural representations and cannot effectively leverage data across subjects. To address these limitations, we introduce Homogeneity-Heterogeneity Disentangled Learning for neural Representations (H2DiLR), a novel framework that disentangles and learns both the homogeneity and heterogeneity from intracranial recordings across multiple subjects. To evaluate H2DiLR, we collected stereoelectroencephalography (sEEG) data from multiple participants reading Mandarin materials comprising 407 syllables, representing nearly all Mandarin characters. Extensive experiments demonstrate that H2DiLR, as a unified decoding paradigm, significantly outperforms the conventional heterogeneous decoding approach. Furthermore, we empirically confirm that H2DiLR effectively captures both homogeneity and heterogeneity during neural representation learning.

【29】 Decoding Emotions: Unveiling Facial Expressions through Acoustic Sensing with Contrastive Attention
标题: 解码情绪:通过对比注意力的声学感知揭示面部表情
作者: Guangjing Wang, Juexing Wang, Ce Zhou, Weikang Ding, Huacheng Zeng, Tianxing Li, Qiben Yan
备注:The extended version of the 2023 IEEE INFOCOM conference paper
链接:点击下载PDF文件
摘要:表情识别通过准确检测用户的情绪状态,为内容推荐和心理健康等应用提供了巨大的希望。传统方法通常依赖于摄像头或可穿戴传感器,这会引发隐私问题并增加额外的设备负担。此外,当训练数据集和推理数据集之间存在分布偏移时,现有的基于声学的方法难以保持令人满意的性能。在本文中,我们介绍了FacER+,这是一种主动声学面部表情识别系统,它消除了对外部麦克风阵列的需求。FacER+通过分析3D面部轮廓和智能手机耳机扬声器之间发射的近超声信号的回波来提取面部表情特征。这种方法不仅减少了背景噪声,而且能够以最少的训练数据识别来自不同用户的不同表情。我们开发了一个基于外部注意力的对比模型,以持续地学习不同用户的表情特征,减少分布差异。涉及20名志愿者的广泛实验表明,FacER+可以准确地识别六种常见的面部表情,在不同的、与用户无关的现实生活场景中的准确率超过90%,超过领先的声学传感方法10%。FacER+为面部表情识别提供了一个强大而实用的解决方案。摘要:Expression recognition holds great promise for applications such as content recommendation and mental healthcare by accurately detecting users' emotional states. Traditional methods often rely on cameras or wearable sensors, which raise privacy concerns and add extra device burdens. In addition, existing acoustic-based methods struggle to maintain satisfactory performance when there is a distribution shift between the training dataset and the inference dataset. In this paper, we introduce FacER+, an active acoustic facial expression recognition system, which eliminates the requirement for external microphone arrays. FacER+ extracts facial expression features by analyzing the echoes of near-ultrasound signals emitted between the 3D facial contour and the earpiece speaker on a smartphone. This approach not only reduces background noise but also enables the identification of different expressions from various users with minimal training data. We develop a contrastive external attention-based model to consistently learn expression features across different users, reducing the distribution differences. Extensive experiments involving 20 volunteers, both with and without masks, demonstrate that FacER+ can accurately recognize six common facial expressions with over 90% accuracy in diverse, user-independent real-life scenarios, surpassing the performance of the leading acoustic sensing methods by 10%. FacER+ offers a robust and practical solution for facial expression recognition.


机器翻译,仅供参考