今日论文合集:cs.SD语音6篇,eess.AS音频处理7篇。本文经arXiv每日学术速递授权转载
【1】Any2Point: Empowering Any-modality Large Models for Efficient 3D Understanding
标题:Any2Point:支持任何模态大型模型以实现高效的3D理解
链接:https://arxiv.org/abs/2404.07989
作者:Yiwen Tang,Jiaming Liu,Dong Wang,Zhigang Wang,Shanghang Zhang,Bin Zhao,Xuelong Li
备注:Code and models are released at this https URL
摘要:大型基础模型最近成为人们关注的焦点,在广泛的场景中获得了卓越的性能。由于3D数据的稀缺性,已经做出了许多努力来使预先训练的Transformers从视觉适应3D域。然而,这种2D到3D的方法仍然是有限的,由于空间几何形状的潜在损失和高计算成本。更重要的是,他们的框架主要是为2D模型设计的,缺乏通用的any-to-3D范式。在本文中,我们介绍了Any 2 Point,这是一种参数高效的方法,可以为任何模态的大型模型(视觉,语言,音频)提供3D理解。给定来自任何源模态的冻结的Transformer,我们提出了3D到任何(1D或2D)虚拟投影策略,该策略将输入3D点与源模态内的原始1D或2D位置相关联。这种机制使我们能够为每个3D令牌分配与预训练模型配对的位置编码,这避免了由真实投影引起的3D几何损失,并更好地激励Transformer进行具有1D/2D位置先验的3D学习。然后,在每个Transformer块中,我们插入一个任意到3D的引导适配器模块,用于参数高效的微调。适配器结合了来自源模态的先验空间知识来指导3D标记的局部特征聚合,从而迫使任何模态Transformers进行语义适配。我们进行了大量的实验,以展示我们的方法的有效性和效率。代码和模型在https://github.com/Ivan-Tang-3D/Any2Point上发布。
摘要:Large foundation models have recently emerged as a prominent focus of interest, attaining superior performance in widespread scenarios. Due to the scarcity of 3D data, many efforts have been made to adapt pre-trained transformers from vision to 3D domains. However, such 2D-to-3D approaches are still limited, due to the potential loss of spatial geometries and high computation cost. More importantly, their frameworks are mainly designed for 2D models, lacking a general any-to-3D paradigm. In this paper, we introduce Any2Point, a parameter-efficient method to empower any-modality large models (vision, language, audio) for 3D understanding. Given a frozen transformer from any source modality, we propose a 3D-to-any (1D or 2D) virtual projection strategy that correlates the input 3D points to the original 1D or 2D positions within the source modality. This mechanism enables us to assign each 3D token with a positional encoding paired with the pre-trained model, which avoids 3D geometry loss caused by the true projection and better motivates the transformer for 3D learning with 1D/2D positional priors. Then, within each transformer block, we insert an any-to-3D guided adapter module for parameter-efficient fine-tuning. The adapter incorporates prior spatial knowledge from the source modality to guide the local feature aggregation of 3D tokens, compelling the semantic adaption of any-modality transformers. We conduct extensive experiments to showcase the effectiveness and efficiency of our method. Code and models are released at https://github.com/Ivan-Tang-3D/Any2Point.
【2】 Audio Dialogues: Dialogues dataset for audio and music understanding
标题:音频对话:音频和音乐理解的对话数据集
链接:https://arxiv.org/abs/2404.07616
作者:Arushi Goel,Zhifeng Kong,Rafael Valle,Bryan Catanzaro
备注:Demo website: this https URL
摘要:用于音频理解的现有数据集主要集中在用于以自然语言描述音频的单轮交互(即音频字幕,音频问答),从而限制了通过交互式对话理解音频。为了解决这个问题,我们引入了Audio Dialogues:一个包含163.8k个一般音频声音和音乐样本的多轮对话数据集。除了对话之外,音频对话还具有问答对,可以一起理解和比较多个输入音频。Audio Dialogues利用现有数据集的基于注释的方法和字幕注释,使用大型语言模型(LLM)生成多轮对话。我们在我们提出的数据集上评估了现有的音频增强大型语言模型,以证明音频对话的复杂性和适用性。我们用于生成数据集的代码将公开提供。详细的提示和生成的对话可以在演示网站https://audiodialogues.github.io/上找到。
摘要:Existing datasets for audio understanding primarily focus on single-turn interactions (i.e. audio captioning, audio question answering) for describing audio in natural language, thus limiting understanding audio via interactive dialogue. To address this gap, we introduce Audio Dialogues: a multi-turn dialogue dataset containing 163.8k samples for general audio sounds and music. In addition to dialogues, Audio Dialogues also has question-answer pairs to understand and compare multiple input audios together. Audio Dialogues leverages a prompting-based approach and caption annotations from existing datasets to generate multi-turn dialogues using a Large Language Model (LLM). We evaluate existing audio-augmented large language models on our proposed dataset to demonstrate the complexity and applicability of Audio Dialogues. Our code for generating the dataset will be made publicly available. Detailed prompts and generated dialogues can be found on the demo website https://audiodialogues.github.io/.
【3】 An Effective Automated Speaking Assessment Approach to Mitigating Data Scarcity and Imbalanced Distribution
标题:一种有效缓解数据稀缺和分布不均衡的自动口语评估方法
链接:https://arxiv.org/abs/2404.07575
作者:Tien-Hong Lo,Fu-An Chao,Tzu-I Wu,Yao-Ting Sung,Berlin Chen
备注:Accepted to NAACL 2023 Findings
摘要:自动口语评估(ASA)通常涉及自动语音识别(ASR)和从学习者语音的ASR转录中手工提取特征。最近,自监督学习(SSL)与传统方法相比表现出出色的性能。然而,基于SSL的ASA系统面临着至少三个与数据相关的挑战:有限的注释数据,不均匀的分布的学习者的熟练程度和不均匀的分数间隔之间的不同CEFR熟练程度。为了解决这些挑战,我们探索了两种新的建模策略的使用:基于度量的分类和损失重新加权,利用不同的基于SSL的嵌入功能。在ICNALE基准数据集上的大量实验结果表明,我们的方法可以以相当大的幅度优于现有的强基线,在CEFR预测准确度上实现了10%以上的显着提高。
摘要:Automated speaking assessment (ASA) typically involves automatic speech recognition (ASR) and hand-crafted feature extraction from the ASR transcript of a learner's speech. Recently, self-supervised learning (SSL) has shown stellar performance compared to traditional methods. However, SSL-based ASA systems are faced with at least three data-related challenges: limited annotated data, uneven distribution of learner proficiency levels and non-uniform score intervals between different CEFR proficiency levels. To address these challenges, we explore the use of two novel modeling strategies: metric-based classification and loss reweighting, leveraging distinct SSL-based embedding features. Extensive experimental results on the ICNALE benchmark dataset suggest that our approach can outperform existing strong baselines by a sizable margin, achieving a significant improvement of more than 10% in CEFR prediction accuracy.
【4】 Differentiable All-pole Filters for Time-varying Audio Systems
标题:时变音频系统的可微分全极点滤波器
链接:https://arxiv.org/abs/2404.07970
作者:Chin-Yun Yu,Christopher Mitcheltree,Alistair Carson,Stefan Bilbao,Joshua D. Reiss,György Fazekas备注:Submitted to DAFx 2024
摘要:无限脉冲响应滤波器是许多时变音频系统(如音频效果器和合成器)的重要组成部分。然而,它们的递归结构阻碍了使用自动微分对这些系统进行端到端的训练。虽然非递归滤波器近似,如频率采样和基于帧的处理已被提出并在以前的工作中被广泛使用,但它们不能准确地反映原始系统的梯度。我们通过重新表达时变全极点滤波器来通过自身反向传播梯度来缓解这一困难,因此滤波器实现不受自动微分框架的技术限制。该实现可以在包含具有极点的滤波器的任何音频系统内采用,以进行有效的梯度评估。我们展示了它的训练效率和表达能力,用于在相位器,时变减法合成器和前馈压缩器上模拟真实世界的动态音频系统。我们提供我们的代码,并在https://christhetree.github.io/all_pole_filters/的VST插件中提供经过训练的音频效果和合成模型。
摘要:Infinite impulse response filters are an essential building block of many time-varying audio systems, such as audio effects and synthesisers. However, their recursive structure impedes end-to-end training of these systems using automatic differentiation. Although non-recursive filter approximations like frequency sampling and frame-based processing have been proposed and widely used in previous works, they cannot accurately reflect the gradient of the original system. We alleviate this difficulty by re-expressing a time-varying all-pole filter to backpropagate the gradients through itself, so the filter implementation is not bound to the technical limitations of automatic differentiation frameworks. This implementation can be employed within any audio system containing filters with poles for efficient gradient evaluation. We demonstrate its training efficiency and expressive capabilities for modelling real-world dynamic audio systems on a phaser, time-varying subtractive synthesiser, and feed-forward compressor. We make our code available and provide the trained audio effect and synth models in a VST plugin at https://christhetree.github.io/all_pole_filters/.
【5】 Conformer-1: Robust ASR via Large-Scale Semisupervised Bootstrapping
标题:Conformer-1:基于大规模半监督自举的鲁棒ASR
链接:https://arxiv.org/abs/2404.07341
作者:Kevin Zhang,Luka Chkhetiani,Francis McCann Ramirez,Yash Khare,Andrea Vanzo,Michael Liang,Sergio Ramirez Martin,Gabriel Oexle,Ruben Bousbib,Taufiquzzaman Peyash,Michael Nguyen,Dillon Pulliam,Domenic Donato
摘要:本文介绍了Conformer-1,这是一种端到端的自动语音识别(ASR)模型,在57万小时的语音音频数据的广泛数据集上进行训练,其中91%来自公开来源。为了实现这一点,我们在使用强Conformer RNN-T基线模型为未标记的公共数据生成伪标签后执行Noisy Student Training。这些伪标记数据的加入导致我们的异步和实时模型的相对字错误率(WER)分别显著提高了11.5%和24.3%。此外,由于增加了这些数据,该模型对背景噪声更具鲁棒性。在这项研究中得到的结果表明,伪标记的公开可用的数据的合并是一个非常有效的策略,提高ASR的准确性和噪声鲁棒性。
摘要:This paper presents Conformer-1, an end-to-end Automatic Speech Recognition (ASR) model trained on an extensive dataset of 570k hours of speech audio data, 91% of which was acquired from publicly available sources. To achieve this, we perform Noisy Student Training after generating pseudo-labels for the unlabeled public data using a strong Conformer RNN-T baseline model. The addition of these pseudo-labeled data results in remarkable improvements in relative Word Error Rate (WER) by 11.5% and 24.3% for our asynchronous and realtime models, respectively. Additionally, the model is more robust to background noise owing to the addition of these data. The results obtained in this study demonstrate that the incorporation of pseudo-labeled publicly available data is a highly effective strategy for improving ASR accuracy and noise robustness.
【6】 Houston we have a Divergence: A Subgroup Performance Analysis of ASR Models
标题:休斯顿我们有分歧:ASR模型的子群性能分析
链接:https://arxiv.org/abs/2404.07226
作者:Alkis Koudounas,Flavio Giobergia
备注:2 pages
摘要:无畏的步骤阿波罗社区资源提供了无与伦比的机会,探索来自NASA阿波罗任务的多扬声器团队通信的潜力。这项研究的重点是发现的特点,使阿波罗录音或多或少可理解的自动语音识别(ASR)方法。我们提取,对于每个音频记录,可解释的元数据记录(信噪比,频谱平坦度,暂停的存在,句子持续时间),成绩单(说的单词数量,说话速度),或已知的先验(扬声器)。我们基于这些元数据的组合来识别音频记录的子组,并计算每个子组的性能(例如,单词错误率)和总体人群的性能差异(“差异”)。然后,我们应用不同大小的Whisper模型,在zero-shot或微调后,在仅英语或多语言数据集上进行训练。我们进行了几项分析,以(i)自动识别和描述给定模型中最有问题的子组,(ii)检查微调w.r.t.子组水平的zero-shot,(iii)了解模型大小对子组性能的影响,以及(iv)分析多语言模型是否比单语言模型对子组性能差异更敏感。这些见解增强了我们对特定于子组的性能变化的理解,为优化地空通信ASR系统的进步铺平了道路。
摘要:The Fearless Steps APOLLO Community Resource provides unparalleled opportunities to explore the potential of multi-speaker team communications from NASA Apollo missions. This study focuses on discovering the characteristics that make Apollo recordings more or less intelligible to Automatic Speech Recognition (ASR) methods. We extract, for each audio recording, interpretable metadata on recordings (signal-to-noise ratio, spectral flatness, presence of pauses, sentence duration), transcript (number of words spoken, speaking rate), or known a priori (speaker). We identify subgroups of audio recordings based on combinations of these metadata and compute each subgroup's performance (e.g., Word Error Rate) and the difference in performance (''divergence'') w.r.t the overall population. We then apply the Whisper model in different sizes, trained on English-only or multilingual datasets, in zero-shot or after fine-tuning. We conduct several analyses to (i) automatically identify and describe the most problematic subgroups for a given model, (ii) examine the impact of fine-tuning w.r.t. zero-shot at the subgroup level, (iii) understand the effect of model size on subgroup performance, and (iv) analyze if multilingual models are more sensitive than monolingual to subgroup performance disparities. The insights enhance our understanding of subgroup-specific performance variations, paving the way for advancements in optimizing ASR systems for Earth-to-space communications.
【1】 Differentiable All-pole Filters for Time-varying Audio Systems作者:Chin-Yun Yu,Christopher Mitcheltree,Alistair Carson,Stefan Bilbao,Joshua D. Reiss,György Fazekas备注:Submitted to DAFx 2024摘要:无限脉冲响应滤波器是许多时变音频系统(如音频效果器和合成器)的重要组成部分。然而,它们的递归结构阻碍了使用自动微分对这些系统进行端到端的训练。虽然非递归滤波器近似,如频率采样和基于帧的处理已被提出并在以前的工作中被广泛使用,但它们不能准确地反映原始系统的梯度。我们通过重新表达时变全极点滤波器来通过自身反向传播梯度来缓解这一困难,因此滤波器实现不受自动微分框架的技术限制。该实现可以在包含具有极点的滤波器的任何音频系统内采用,以进行有效的梯度评估。我们展示了它的训练效率和表达能力,用于在相位器,时变减法合成器和前馈压缩器上模拟真实世界的动态音频系统。我们提供我们的代码,并在https://christhetree.github.io/all_pole_filters/的VST插件中提供经过训练的音频效果和合成模型。摘要:Infinite impulse response filters are an essential building block of many time-varying audio systems, such as audio effects and synthesisers. However, their recursive structure impedes end-to-end training of these systems using automatic differentiation. Although non-recursive filter approximations like frequency sampling and frame-based processing have been proposed and widely used in previous works, they cannot accurately reflect the gradient of the original system. We alleviate this difficulty by re-expressing a time-varying all-pole filter to backpropagate the gradients through itself, so the filter implementation is not bound to the technical limitations of automatic differentiation frameworks. This implementation can be employed within any audio system containing filters with poles for efficient gradient evaluation. We demonstrate its training efficiency and expressive capabilities for modelling real-world dynamic audio systems on a phaser, time-varying subtractive synthesiser, and feed-forward compressor. We make our code available and provide the trained audio effect and synth models in a VST plugin at https://christhetree.github.io/all_pole_filters/.【2】 Conformer-1: Robust ASR via Large-Scale Semisupervised Bootstrapping标题:Conformer-1:基于大规模半监督自举的鲁棒ASR作者:Kevin Zhang,Luka Chkhetiani,Francis McCann Ramirez,Yash Khare,Andrea Vanzo,Michael Liang,Sergio Ramirez Martin,Gabriel Oexle,Ruben Bousbib,Taufiquzzaman Peyash,Michael Nguyen,Dillon Pulliam,Domenic Donato摘要:本文介绍了Conformer-1,这是一种端到端的自动语音识别(ASR)模型,在57万小时的语音音频数据的广泛数据集上进行训练,其中91%来自公开来源。为了实现这一点,我们在使用强Conformer RNN-T基线模型为未标记的公共数据生成伪标签后执行Noisy Student Training。这些伪标记数据的加入导致我们的异步和实时模型的相对字错误率(WER)分别显著提高了11.5%和24.3%。此外,由于增加了这些数据,该模型对背景噪声更具鲁棒性。在这项研究中得到的结果表明,伪标记的公开可用的数据的合并是一个非常有效的策略,提高ASR的准确性和噪声鲁棒性。摘要:This paper presents Conformer-1, an end-to-end Automatic Speech Recognition (ASR) model trained on an extensive dataset of 570k hours of speech audio data, 91% of which was acquired from publicly available sources. To achieve this, we perform Noisy Student Training after generating pseudo-labels for the unlabeled public data using a strong Conformer RNN-T baseline model. The addition of these pseudo-labeled data results in remarkable improvements in relative Word Error Rate (WER) by 11.5% and 24.3% for our asynchronous and realtime models, respectively. Additionally, the model is more robust to background noise owing to the addition of these data. The results obtained in this study demonstrate that the incorporation of pseudo-labeled publicly available data is a highly effective strategy for improving ASR accuracy and noise robustness.
【3】 Houston we have a Divergence: A Subgroup Performance Analysis of ASR Models作者:Alkis Koudounas,Flavio Giobergia摘要:无畏的步骤阿波罗社区资源提供了无与伦比的机会,探索来自NASA阿波罗任务的多扬声器团队通信的潜力。这项研究的重点是发现的特点,使阿波罗录音或多或少可理解的自动语音识别(ASR)方法。我们提取,对于每个音频记录,可解释的元数据记录(信噪比,频谱平坦度,暂停的存在,句子持续时间),成绩单(说的单词数量,说话速度),或已知的先验(扬声器)。我们基于这些元数据的组合来识别音频记录的子组,并计算每个子组的性能(例如,单词错误率)和总体人群的性能差异(“差异”)。然后,我们应用不同大小的Whisper模型,在zero-shot或微调后,在仅英语或多语言数据集上进行训练。我们进行了几项分析,以(i)自动识别和描述给定模型中最有问题的子组,(ii)检查微调w.r.t.子组水平的zero-shot,(iii)了解模型大小对子组性能的影响,以及(iv)分析多语言模型是否比单语言模型对子组性能差异更敏感。这些见解增强了我们对特定于子组的性能变化的理解,为优化地空通信ASR系统的进步铺平了道路。摘要:The Fearless Steps APOLLO Community Resource provides unparalleled opportunities to explore the potential of multi-speaker team communications from NASA Apollo missions. This study focuses on discovering the characteristics that make Apollo recordings more or less intelligible to Automatic Speech Recognition (ASR) methods. We extract, for each audio recording, interpretable metadata on recordings (signal-to-noise ratio, spectral flatness, presence of pauses, sentence duration), transcript (number of words spoken, speaking rate), or known a priori (speaker). We identify subgroups of audio recordings based on combinations of these metadata and compute each subgroup's performance (e.g., Word Error Rate) and the difference in performance (''divergence'') w.r.t the overall population. We then apply the Whisper model in different sizes, trained on English-only or multilingual datasets, in zero-shot or after fine-tuning. We conduct several analyses to (i) automatically identify and describe the most problematic subgroups for a given model, (ii) examine the impact of fine-tuning w.r.t. zero-shot at the subgroup level, (iii) understand the effect of model size on subgroup performance, and (iv) analyze if multilingual models are more sensitive than monolingual to subgroup performance disparities. The insights enhance our understanding of subgroup-specific performance variations, paving the way for advancements in optimizing ASR systems for Earth-to-space communications.【4】 Any2Point: Empowering Any-modality Large Models for Efficient 3D Understanding标题:Any2Point:支持任何模态大型模型以实现高效的3D理解作者:Yiwen Tang,Jiaming Liu,Dong Wang,Zhigang Wang,Shanghang Zhang,Bin Zhao,Xuelong Li备注:Code and models are released at this https URL摘要:大型基础模型最近成为人们关注的焦点,在广泛的场景中获得了卓越的性能。由于3D数据的稀缺性,已经做出了许多努力来使预先训练的Transformers从视觉适应3D域。然而,这种2D到3D的方法仍然是有限的,由于空间几何形状的潜在损失和高计算成本。更重要的是,他们的框架主要是为2D模型设计的,缺乏通用的any-to-3D范式。在本文中,我们介绍了Any 2 Point,这是一种参数高效的方法,可以为任何模态的大型模型(视觉,语言,音频)提供3D理解。给定来自任何源模态的冻结的Transformer,我们提出了3D到任何(1D或2D)虚拟投影策略,该策略将输入3D点与源模态内的原始1D或2D位置相关联。这种机制使我们能够为每个3D令牌分配与预训练模型配对的位置编码,这避免了由真实投影引起的3D几何损失,并更好地激励Transformer进行具有1D/2D位置先验的3D学习。然后,在每个Transformer块中,我们插入一个任意到3D的引导适配器模块,用于参数高效的微调。适配器结合了来自源模态的先验空间知识来指导3D标记的局部特征聚合,从而迫使任何模态Transformers进行语义适配。我们进行了大量的实验,以展示我们的方法的有效性和效率。代码和模型在https://github.com/Ivan-Tang-3D/Any2Point上发布。摘要:Large foundation models have recently emerged as a prominent focus of interest, attaining superior performance in widespread scenarios. Due to the scarcity of 3D data, many efforts have been made to adapt pre-trained transformers from vision to 3D domains. However, such 2D-to-3D approaches are still limited, due to the potential loss of spatial geometries and high computation cost. More importantly, their frameworks are mainly designed for 2D models, lacking a general any-to-3D paradigm. In this paper, we introduce Any2Point, a parameter-efficient method to empower any-modality large models (vision, language, audio) for 3D understanding. Given a frozen transformer from any source modality, we propose a 3D-to-any (1D or 2D) virtual projection strategy that correlates the input 3D points to the original 1D or 2D positions within the source modality. This mechanism enables us to assign each 3D token with a positional encoding paired with the pre-trained model, which avoids 3D geometry loss caused by the true projection and better motivates the transformer for 3D learning with 1D/2D positional priors. Then, within each transformer block, we insert an any-to-3D guided adapter module for parameter-efficient fine-tuning. The adapter incorporates prior spatial knowledge from the source modality to guide the local feature aggregation of 3D tokens, compelling the semantic adaption of any-modality transformers. We conduct extensive experiments to showcase the effectiveness and efficiency of our method. Code and models are released at https://github.com/Ivan-Tang-3D/Any2Point.【5】 Audio Dialogues: Dialogues dataset for audio and music understanding作者:Arushi Goel,Zhifeng Kong,Rafael Valle,Bryan Catanzaro备注:Demo website: this https URL摘要:用于音频理解的现有数据集主要集中在用于以自然语言描述音频的单轮交互(即音频字幕,音频问答),从而限制了通过交互式对话理解音频。为了解决这个问题,我们引入了Audio Dialogues:一个包含163.8k个一般音频声音和音乐样本的多轮对话数据集。除了对话之外,音频对话还具有问答对,可以一起理解和比较多个输入音频。Audio Dialogues利用现有数据集的基于注释的方法和字幕注释,使用大型语言模型(LLM)生成多轮对话。我们在我们提出的数据集上评估了现有的音频增强大型语言模型,以证明音频对话的复杂性和适用性。我们用于生成数据集的代码将公开提供。详细的提示和生成的对话可以在演示网站https://audiodialogues.github.io/上找到。摘要:Existing datasets for audio understanding primarily focus on single-turn interactions (i.e. audio captioning, audio question answering) for describing audio in natural language, thus limiting understanding audio via interactive dialogue. To address this gap, we introduce Audio Dialogues: a multi-turn dialogue dataset containing 163.8k samples for general audio sounds and music. In addition to dialogues, Audio Dialogues also has question-answer pairs to understand and compare multiple input audios together. Audio Dialogues leverages a prompting-based approach and caption annotations from existing datasets to generate multi-turn dialogues using a Large Language Model (LLM). We evaluate existing audio-augmented large language models on our proposed dataset to demonstrate the complexity and applicability of Audio Dialogues. Our code for generating the dataset will be made publicly available. Detailed prompts and generated dialogues can be found on the demo website https://audiodialogues.github.io/.
【6】 An Effective Automated Speaking Assessment Approach to Mitigating Data Scarcity and Imbalanced Distribution标题:一种有效缓解数据稀缺和分布不均衡的自动口语评估方法作者:Tien-Hong Lo,Fu-An Chao,Tzu-I Wu,Yao-Ting Sung,Berlin Chen备注:Accepted to NAACL 2023 Findings摘要:自动口语评估(ASA)通常涉及自动语音识别(ASR)和从学习者语音的ASR转录中手工提取特征。最近,自监督学习(SSL)与传统方法相比表现出出色的性能。然而,基于SSL的ASA系统面临着至少三个与数据相关的挑战:有限的注释数据,不均匀的分布的学习者的熟练程度和不均匀的分数间隔之间的不同CEFR熟练程度。为了解决这些挑战,我们探索了两种新的建模策略的使用:基于度量的分类和损失重新加权,利用不同的基于SSL的嵌入功能。在ICNALE基准数据集上的大量实验结果表明,我们的方法可以以相当大的幅度优于现有的强基线,在CEFR预测准确度上实现了10%以上的显着提高。摘要:Automated speaking assessment (ASA) typically involves automatic speech recognition (ASR) and hand-crafted feature extraction from the ASR transcript of a learner's speech. Recently, self-supervised learning (SSL) has shown stellar performance compared to traditional methods. However, SSL-based ASA systems are faced with at least three data-related challenges: limited annotated data, uneven distribution of learner proficiency levels and non-uniform score intervals between different CEFR proficiency levels. To address these challenges, we explore the use of two novel modeling strategies: metric-based classification and loss reweighting, leveraging distinct SSL-based embedding features. Extensive experimental results on the ICNALE benchmark dataset suggest that our approach can outperform existing strong baselines by a sizable margin, achieving a significant improvement of more than 10% in CEFR prediction accuracy.
【7】 PEAVS: Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores标题:PEAVS:基于观众意见评分的视听同步性知觉评价作者:Lucas Goncalves,Prashant Mathur,Chandrashekhar Lavania,Metehan Cekic,Marcello Federico,Kyu J. Han摘要:视听生成建模的最新进展受到深度学习和数据丰富的基准测试的推动。然而,增长不仅仅归因于模型和基准。普遍接受的评价指标在推动该领域的发展方面也发挥着重要作用。虽然有许多指标可用于单独评估音频和视频内容,但缺乏为“野外”视频提供视听同步的定量和可解释的度量的指标。为了解决这一差距,我们首先创建了一个大规模的人类注释数据集(100+小时),代表了视听内容中的九种类型的同步错误以及人类如何感知它们。然后,我们开发了一个PEAVS(视听同步的感知评估)评分,这是一个新颖的自动度量标准,具有5点量表,用于评估视听同步的质量。我们使用新生成的数据集验证PEAVS,与人类标签相比,在集合水平和剪辑水平上的Pearson相关性分别为0.79和0.54。在我们的实验中,我们观察到一个相对增益50%以上的自然扩展的Fr\'echet为基础的指标的视听同步,确认PEAVS的功效,客观地模拟主观感知的视听同步视频“在野外”。摘要:Recent advancements in audio-visual generative modeling have been propelled by progress in deep learning and the availability of data-rich benchmarks. However, the growth is not attributed solely to models and benchmarks. Universally accepted evaluation metrics also play an important role in advancing the field. While there are many metrics available to evaluate audio and visual content separately, there is a lack of metrics that offer a quantitative and interpretable measure of audio-visual synchronization for videos "in the wild". To address this gap, we first created a large scale human annotated dataset (100+ hrs) representing nine types of synchronization errors in audio-visual content and how human perceive them. We then developed a PEAVS (Perceptual Evaluation of Audio-Visual Synchrony) score, a novel automatic metric with a 5-point scale that evaluates the quality of audio-visual synchronization. We validate PEAVS using a newly generated dataset, achieving a Pearson correlation of 0.79 at the set level and 0.54 at the clip level when compared to human labels. In our experiments, we observe a relative gain 50% over a natural extension of Fr\'echet based metrics for Audio-Visual synchrony, confirming PEAVS efficacy in objectively modeling subjective perceptions of audio-visual synchronization for videos "in the wild".