今日论文合集:cs.SD语音5篇,eess.AS音频处理9篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
标题:OpenBEATs:一个完全开源的通用音频编码器
链接:https://arxiv.org/abs/2507.14129

作者:Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi, Satoru Fukayama, Hye-jin Shim, Soham Deshmukh, Shinji Watanabe
摘要:掩蔽标记预测已经成为跨语言,视觉和语音的强大预训练目标,提供了通过单个预训练任务统一这些不同模式的可能性。然而,它在一般音频理解方面的应用仍然没有得到充分的探索,BEAT是唯一值得注意的例子。由于缺乏开源的预训练代码,BEAT的修改有限。此外,BEAT仅在AudioSet上接受训练,限制了其更广泛的下游适用性。为了解决这些差距,我们提出了OpenBEAT,这是一个开源框架,通过多域音频预训练扩展BEAT。我们对六种类型的任务,二十五个数据集和三个音频域进行了全面的评估,包括音频推理任务,如音频问答,蕴涵和字幕。OpenBEAT在六个生物声学数据集、两个环境声音数据集和五个推理数据集上实现了最先进的性能,比参数大小为四分之一的超过十亿个参数的模型表现更好。这些结果证明了多域数据集和掩码令牌预测任务学习通用音频表示的有效性。为了促进进一步的研究和可重复性,我们在https://shikhar-s.github.io/OpenBEATs上发布了所有预训练和评估代码、预训练和微调的检查点以及训练日志。
摘要:Masked token prediction has emerged as a powerful pre-training objective across language, vision, and speech, offering the potential to unify these diverse modalities through a single pre-training task. However, its application for general audio understanding remains underexplored, with BEATs being the only notable example. BEATs has seen limited modifications due to the absence of open-source pre-training code. Furthermore, BEATs was trained only on AudioSet, restricting its broader downstream applicability. To address these gaps, we present OpenBEATs, an open-source framework that extends BEATs via multi-domain audio pre-training. We conduct comprehensive evaluations across six types of tasks, twenty five datasets, and three audio domains, including audio reasoning tasks such as audio question answering, entailment, and captioning. OpenBEATs achieves state-of-the-art performance on six bioacoustics datasets, two environmental sound datasets and five reasoning datasets, performing better than models exceeding a billion parameters at one-fourth their parameter size. These results demonstrate the effectiveness of multi-domain datasets and masked token prediction task to learn general-purpose audio representations. To promote further research and reproducibility, we release all pre-training and evaluation code, pretrained and fine-tuned checkpoints, and training logs at https://shikhar-s.github.io/OpenBEATs


【2】Controlling the Parameterized Multi-channel Wiener Filter using a tiny neural network
标题:使用微型神经网络控制参数化多通道维纳过滤器
链接:https://arxiv.org/abs/2507.13863

作者:Eric Grinstein, Ashutosh Pandey, Cole Li, Shanmukha Srinivas, Juan Azcarreta, Jacob Donley, Sanha Lee, Ali Aroudi, Cagdas Bilen
备注:Accepted to WASPAA 2025
摘要:噪声抑制和语音失真是设计多通道语音增强算法时需要平衡的两个重要方面。虽然神经网络模型已经实现了最先进的噪声抑制,但它们的非线性操作通常会引入高语音失真。相反,经典的信号处理算法,如参数化多通道维纳滤波器(PMWF)波束形成器提供了明确的机制,用于控制抑制/失真的权衡。在这项工作中,我们提出了NeuralPMWF,一个系统,其中PMWF完全控制使用低延迟,低计算神经网络,从而在一个低复杂度的系统提供高降噪和低语音失真。实验结果表明,我们提出的方法相比,使用类似的计算资源的几个竞争性的基线显着更好的感知和客观的语音增强的结果。
摘要:Noise suppression and speech distortion are two important aspects to be balanced when designing multi-channel Speech Enhancement (SE) algorithms. Although neural network models have achieved state-of-the-art noise suppression, their non-linear operations often introduce high speech distortion. Conversely, classical signal processing algorithms such as the Parameterized Multi-channel Wiener Filter ( PMWF) beamformer offer explicit mechanisms for controlling the suppression/distortion trade-off. In this work, we present NeuralPMWF, a system where the PMWF is entirely controlled using a low-latency, low-compute neural network, resulting in a low-complexity system offering high noise reduction and low speech distortion. Experimental results show that our proposed approach results in significantly better perceptual and objective speech enhancement in comparison to several competitive baselines using similar computational resources.


【3】Temporal Adaptation of Pre-trained Foundation Models for Music Structure Analysis
标题:音乐结构分析预训练基础模型的时间适应
链接:https://arxiv.org/abs/2507.13572

作者:Yixiao Zhang, Haonan Chen, Ju-Chiang Wang, Jitong Chen
备注:Accepted to WASPAA 2025. Project Page: this https URL
摘要:基于音频的音乐结构分析(MSA)是音乐信息检索中的一项重要任务,由于音乐形式的复杂性和多变性,仍然具有挑战性。最近的进展突出了微调预训练的音乐基础模型用于MSA任务的潜力。然而,这些模型通常是用高时间特征分辨率和短音频窗口训练的,这限制了它们的效率,并在应用于长格式音频时引入了偏差。本文提出了一种时间适应的方法微调音乐基础模型量身定制的MSA。我们的方法通过结合两个关键策略:(1)音频窗口扩展和(2)低分辨率自适应,可以在单次向前传递中有效分析全长歌曲。在Harmonix Set和RWC-Pop数据集上的实验表明,我们的方法显着提高了边界检测和结构函数预测,同时保持了相当的内存使用和推理速度。
摘要:Audio-based music structure analysis (MSA) is an essential task in Music Information Retrieval that remains challenging due to the complexity and variability of musical form. Recent advances highlight the potential of fine-tuning pre-trained music foundation models for MSA tasks. However, these models are typically trained with high temporal feature resolution and short audio windows, which limits their efficiency and introduces bias when applied to long-form audio. This paper presents a temporal adaptation approach for fine-tuning music foundation models tailored to MSA. Our method enables efficient analysis of full-length songs in a single forward pass by incorporating two key strategies: (1) audio window extension and (2) low-resolution adaptation. Experiments on the Harmonix Set and RWC-Pop datasets show that our method significantly improves both boundary detection and structural function prediction, while maintaining comparable memory usage and inference speed.


【4】A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models
标题:解决俄语语音生成模型中语音和韵律挑战的以数据为中心的框架
链接:https://arxiv.org/abs/2507.13563

作者:Kirill Borodin, Nikita Vasiliev, Vasiliy Kudryavtsev, Maxim Maslov, Mikhail Gorodnichev, Oleg Rogov, Grach Mkrtchian
备注:The work is still in progress
摘要:俄语语音合成提出了独特的挑战,包括元音减少,辅音清音,可变重音模式,同形异义词歧义和不自然的语调。本文介绍了Balalaika,这是一个新的数据集,包含超过2,000小时的工作室质量的俄语语音,具有全面的文本注释,包括标点符号和重音标记。实验结果表明,在Balalaika上训练的模型在语音合成和增强任务中的性能明显优于在现有数据集上训练的模型。我们详细介绍了数据集构建管道,注释方法和比较评估的结果。
摘要:Russian speech synthesis presents distinctive challenges, including vowel reduction, consonant devoicing, variable stress patterns, homograph ambiguity, and unnatural intonation. This paper introduces Balalaika, a novel dataset comprising more than 2,000 hours of studio-quality Russian speech with comprehensive textual annotations, including punctuation and stress markings. Experimental results show that models trained on Balalaika significantly outperform those trained on existing datasets in both speech synthesis and enhancement tasks. We detail the dataset construction pipeline, annotation methodology, and results of comparative evaluations.


【5】Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
标题:统一语音评分量表:语音质量评估和连续语音情感识别的比较学习框架
链接:https://arxiv.org/abs/2507.13626

作者:Cheng-Hung Hu, Yusuke Yasud, Akifumi Yoshimoto, Tomoki Toda
备注:Accepted to Interspeech 2025
摘要:语音质量评估和连续语音情感识别是语音技术中的两个关键任务,它们都依赖于收听者的评分。然而,由于个别听众的因素,这些评级本身就存在偏差。之前的方法引入了平均听众评分量表,并对训练集中的所有听众评分量表进行了建模。然而,平均听众的方法是容易从平均有序数据失真,导致潜在的偏见。此外,学习多个听众评分尺度,同时仅基于平均听众尺度进行推断限制了有效性。相比之下,我们的方法侧重于建模一个统一的听众评分尺度,使用比较分数,正确地捕捉话语之间的评分关系。实验结果表明,该方法有效地提高了SQA和CSER任务的预测性能,证明了其有效性和鲁棒性。
摘要:Speech Quality Assessment (SQA) and Continuous Speech Emotion Recognition (CSER) are two key tasks in speech technology, both relying on listener ratings. However, these ratings are inherently biased due to individual listener factors. Previous approaches have introduced a mean listener scoring scale and modeled all listener scoring scales in the training set. However, the mean listener approach is prone to distortion from averaging ordinal data, leading to potential biases. Moreover, learning multiple listener scoring scales while inferring based only on the mean listener scale limits effectiveness. In contrast, our method focuses on modeling a unified listener scoring scale, using comparison scores to correctly capture the scoring relationships between utterances. Experimental results show that our method effectively improves prediction performance in both SQA and CSER tasks, proving its effectiveness and robustness.



eess.AS音频处理


【1】TGIF: Talker Group-Informed Familiarization of Target Speaker Extraction
标题:TGIF:说话者群体熟悉的目标说话者提取
链接:https://arxiv.org/abs/2507.14044

作者:Tsun-An Hsieh, Minje Kim
摘要:最先进的目标说话人提取(TSE)系统通常被设计为推广到任何给定的混合环境,需要一个具有足够大容量的模型作为通才。个性化语音增强可以是适应单用户场景的专门解决方案,但它忽略了在仅涉及少量说话者的情况下对定制的实际需求,例如,一个特定家庭的TSE。我们解决这个差距与建议的概念,谈话者组知情的熟悉(TGIF)的TSE,其中TSE系统专门在一个特定的用户群,这是具有挑战性的,由于固有的缺乏一个干净的语音目标。为此,我们采用了一种知识蒸馏方法,其中特定于组的学生模型从大型教师模型生成的伪干净目标中学习。这使学生模型能够有效地从特定的说话者组中提取目标说话者,同时保持计算效率。实验结果表明,我们的方法优于基线通用模型,通过适应一个给定的扬声器组的独特的语音特征。我们新提出的TGIF概念强调了为多样化和现实世界的应用开发专业解决方案的潜力,例如家庭拥有的设备上的设备TSE。
摘要:State-of-the-art target speaker extraction (TSE) systems are typically designed to generalize to any given mixing environment, necessitating a model with a large enough capacity as a generalist. Personalized speech enhancement could be a specialized solution that adapts to single-user scenarios, but it overlooks the practical need for customization in cases where only a small number of talkers are involved, e.g., TSE for a specific family. We address this gap with the proposed concept, talker group-informed familiarization (TGIF) of TSE, where the TSE system specializes in a particular group of users, which is challenging due to the inherent absence of a clean speech target. To this end, we employ a knowledge distillation approach, where a group-specific student model learns from the pseudo-clean targets generated by a large teacher model. This tailors the student model to effectively extract the target speaker from the particular talker group while maintaining computational efficiency. Experimental results demonstrate that our approach outperforms the baseline generic models by adapting to the unique speech characteristics of a given speaker group. Our newly proposed TGIF concept underscores the potential of developing specialized solutions for diverse and real-world applications, such as on-device TSE on a family-owned device.


【2】Unifying Listener Scoring Scales: Comparison Learning Framework for Speech Quality Assessment and Continuous Speech Emotion Recognition
标题:统一语音评分量表:语音质量评估和连续语音情感识别的比较学习框架
链接:https://arxiv.org/abs/2507.13626

作者:Cheng-Hung Hu, Yusuke Yasud, Akifumi Yoshimoto, Tomoki Toda
备注:Accepted to Interspeech 2025
摘要:语音质量评估和连续语音情感识别是语音技术中的两个关键任务,它们都依赖于收听者的评分。然而,由于个别听众的因素,这些评级本身就存在偏差。之前的方法引入了平均听众评分量表,并对训练集中的所有听众评分量表进行了建模。然而,平均听众的方法是容易从平均有序数据失真,导致潜在的偏见。此外,学习多个听众评分尺度,同时仅基于平均听众尺度进行推断限制了有效性。相比之下,我们的方法侧重于建模一个统一的听众评分尺度,使用比较分数,正确地捕捉话语之间的评分关系。实验结果表明,该方法有效地提高了SQA和CSER任务的预测性能,证明了其有效性和鲁棒性。
摘要:Speech Quality Assessment (SQA) and Continuous Speech Emotion Recognition (CSER) are two key tasks in speech technology, both relying on listener ratings. However, these ratings are inherently biased due to individual listener factors. Previous approaches have introduced a mean listener scoring scale and modeled all listener scoring scales in the training set. However, the mean listener approach is prone to distortion from averaging ordinal data, leading to potential biases. Moreover, learning multiple listener scoring scales while inferring based only on the mean listener scale limits effectiveness. In contrast, our method focuses on modeling a unified listener scoring scale, using comparison scores to correctly capture the scoring relationships between utterances. Experimental results show that our method effectively improves prediction performance in both SQA and CSER tasks, proving its effectiveness and robustness.


【3】OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
标题:OpenBEATs:一个完全开源的通用音频编码器
链接:https://arxiv.org/abs/2507.14129

作者:Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi, Satoru Fukayama, Hye-jin Shim, Soham Deshmukh, Shinji Watanabe
摘要:掩蔽标记预测已经成为跨语言,视觉和语音的强大预训练目标,提供了通过单个预训练任务统一这些不同模式的可能性。然而,它在一般音频理解方面的应用仍然没有得到充分的探索,BEAT是唯一值得注意的例子。由于缺乏开源的预训练代码,BEAT的修改有限。此外,BEAT仅在AudioSet上接受训练,限制了其更广泛的下游适用性。为了解决这些差距,我们提出了OpenBEAT,这是一个开源框架,通过多域音频预训练扩展BEAT。我们对六种类型的任务,二十五个数据集和三个音频域进行了全面的评估,包括音频推理任务,如音频问答,蕴涵和字幕。OpenBEAT在六个生物声学数据集、两个环境声音数据集和五个推理数据集上实现了最先进的性能,比参数大小为四分之一的超过十亿个参数的模型表现更好。这些结果证明了多域数据集和掩码令牌预测任务学习通用音频表示的有效性。为了促进进一步的研究和可重复性,我们在https://shikhar-s.github.io/OpenBEATs上发布了所有预训练和评估代码、预训练和微调的检查点以及训练日志。
摘要:Masked token prediction has emerged as a powerful pre-training objective across language, vision, and speech, offering the potential to unify these diverse modalities through a single pre-training task. However, its application for general audio understanding remains underexplored, with BEATs being the only notable example. BEATs has seen limited modifications due to the absence of open-source pre-training code. Furthermore, BEATs was trained only on AudioSet, restricting its broader downstream applicability. To address these gaps, we present OpenBEATs, an open-source framework that extends BEATs via multi-domain audio pre-training. We conduct comprehensive evaluations across six types of tasks, twenty five datasets, and three audio domains, including audio reasoning tasks such as audio question answering, entailment, and captioning. OpenBEATs achieves state-of-the-art performance on six bioacoustics datasets, two environmental sound datasets and five reasoning datasets, performing better than models exceeding a billion parameters at one-fourth their parameter size. These results demonstrate the effectiveness of multi-domain datasets and masked token prediction task to learn general-purpose audio representations. To promote further research and reproducibility, we release all pre-training and evaluation code, pretrained and fine-tuned checkpoints, and training logs at https://shikhar-s.github.io/OpenBEATs


【4】Open Automatic Speech Recognition Models for Classical and Modern Standard Arabic
标题:古典和现代标准阿拉伯语的开放自动语音识别模型
链接:https://arxiv.org/abs/2507.13977

作者:Lilit Grigoryan, Nikolay Karpov, Enas Albasiri, Vitaly Lavrukhin, Boris Ginsburg
备注:Accepted to ICASSP 2025
摘要:尽管阿拉伯语是使用最广泛的语言之一,但由于语言的复杂性,阿拉伯语自动语音识别(ASR)系统的开发面临着重大挑战,并且只有有限数量的公共阿拉伯语ASR模型存在。虽然大部分的焦点都集中在现代标准阿拉伯语(MSA)上,但对语言内部的变化却很少关注。本文介绍了阿拉伯语语音和文本处理的通用方法,旨在解决语言的独特挑战。使用这种方法,我们训练两个新的模型的基础上FastConformer架构:一个专为MSA和其他,第一个统一的公共模型MSA和古典阿拉伯语(CA)。MSA模型在相关数据集上设置了一个具有最先进(SOTA)性能的新基准,而统一模型在CA的变音符号上实现了SOTA准确性,同时保持了MSA的强大性能。为了提高可重复性,我们开源了模型及其训练配方。
摘要:Despite Arabic being one of the most widely spoken languages, the development of Arabic Automatic Speech Recognition (ASR) systems faces significant challenges due to the language's complexity, and only a limited number of public Arabic ASR models exist. While much of the focus has been on Modern Standard Arabic (MSA), there is considerably less attention given to the variations within the language. This paper introduces a universal methodology for Arabic speech and text processing designed to address unique challenges of the language. Using this methodology, we train two novel models based on the FastConformer architecture: one designed specifically for MSA and the other, the first unified public model for both MSA and Classical Arabic (CA). The MSA model sets a new benchmark with state-of-the-art (SOTA) performance on related datasets, while the unified model achieves SOTA accuracy with diacritics for CA while maintaining strong performance for MSA. To promote reproducibility, we open-source the models and their training recipes.


【5】Optimizing ASR for Catalan-Spanish Code-Switching: A Comparative Analysis of Methodologies
标题:加泰罗尼亚-西班牙语代码转换的ASB优化:方法比较分析
链接:https://arxiv.org/abs/2507.13875

作者:Carlos Mena, Pol Serra, Jacobo Romero, Abir Messaoudi, Jose Giraldo, Carme Armentano-Oller, Rodolfo Zevallos, Ivan Meza, Javier Hernando
备注:Accepted at Interspeech 2025
摘要:语码转换(CS)是两种或多种语言的交替使用,由于训练数据的稀缺和语言的相似性,它对自动语音识别(ASR)提出了挑战。缺乏专用的CS数据集限制了ASR性能,因为大多数模型依赖于无法反映真实世界CS模式的单语或混合语言语料库。这个问题是至关重要的,在多语社会中,CS发生在非正式和正式的环境。一个重要的例子是加泰罗尼亚语-西班牙语CS,广泛用于媒体和议会演讲。在这项工作中,我们通过探索三种策略来改进加泰罗尼亚语-西班牙语CS的ASR:(1)生成合成CS数据,(2)连接单语音频,以及(3)利用真实CS数据和语言标记。我们从加泰罗尼亚语语料库中提取CS数据,并微调OpenAI的Whisper模型,使其可用于Hugging Face。结果表明,结合适量的合成CS数据与主导语言令牌产生最佳的转录性能。
摘要:Code-switching (CS), the alternating use of two or more languages, challenges automatic speech recognition (ASR) due to scarce training data and linguistic similarities. The lack of dedicated CS datasets limits ASR performance, as most models rely on monolingual or mixed-language corpora that fail to reflect real-world CS patterns. This issue is critical in multilingual societies where CS occurs in informal and formal settings. A key example is Catalan-Spanish CS, widely used in media and parliamentary speeches. In this work, we improve ASR for Catalan-Spanish CS by exploring three strategies: (1) generating synthetic CS data, (2) concatenating monolingual audio, and (3) leveraging real CS data with language tokens. We extract CS data from Catalan speech corpora and fine-tune OpenAI's Whisper models, making them available on Hugging Face. Results show that combining a modest amount of synthetic CS data with the dominant language token yields the best transcription performance.


【6】Controlling the Parameterized Multi-channel Wiener Filter using a tiny neural network
标题:使用微型神经网络控制参数化多通道维纳过滤器
链接:https://arxiv.org/abs/2507.13863

作者:Eric Grinstein, Ashutosh Pandey, Cole Li, Shanmukha Srinivas, Juan Azcarreta, Jacob Donley, Sanha Lee, Ali Aroudi, Cagdas Bilen
备注:Accepted to WASPAA 2025
摘要:噪声抑制和语音失真是设计多通道语音增强算法时需要平衡的两个重要方面。虽然神经网络模型已经实现了最先进的噪声抑制,但它们的非线性操作通常会引入高语音失真。相反,经典的信号处理算法,如参数化多通道维纳滤波器(PMWF)波束形成器提供了明确的机制,用于控制抑制/失真的权衡。在这项工作中,我们提出了NeuralPMWF,一个系统,其中PMWF完全控制使用低延迟,低计算神经网络,从而在一个低复杂度的系统提供高降噪和低语音失真。实验结果表明,我们提出的方法相比,使用类似的计算资源的几个竞争性的基线显着更好的感知和客观的语音增强的结果。
摘要:Noise suppression and speech distortion are two important aspects to be balanced when designing multi-channel Speech Enhancement (SE) algorithms. Although neural network models have achieved state-of-the-art noise suppression, their non-linear operations often introduce high speech distortion. Conversely, classical signal processing algorithms such as the Parameterized Multi-channel Wiener Filter ( PMWF) beamformer offer explicit mechanisms for controlling the suppression/distortion trade-off. In this work, we present NeuralPMWF, a system where the PMWF is entirely controlled using a low-latency, low-compute neural network, resulting in a low-complexity system offering high noise reduction and low speech distortion. Experimental results show that our proposed approach results in significantly better perceptual and objective speech enhancement in comparison to several competitive baselines using similar computational resources.


【7】Temporal Adaptation of Pre-trained Foundation Models for Music Structure Analysis
标题:音乐结构分析预训练基础模型的时间适应
链接:https://arxiv.org/abs/2507.13572

作者:Yixiao Zhang, Haonan Chen, Ju-Chiang Wang, Jitong Chen
备注:Accepted to WASPAA 2025. Project Page: this https URL
摘要:基于音频的音乐结构分析(MSA)是音乐信息检索中的一项重要任务,由于音乐形式的复杂性和多变性,仍然具有挑战性。最近的进展突出了微调预训练的音乐基础模型用于MSA任务的潜力。然而,这些模型通常是用高时间特征分辨率和短音频窗口训练的,这限制了它们的效率,并在应用于长格式音频时引入了偏差。本文提出了一种时间适应的方法微调音乐基础模型量身定制的MSA。我们的方法通过结合两个关键策略:(1)音频窗口扩展和(2)低分辨率自适应,可以在单次向前传递中有效分析全长歌曲。在Harmonix Set和RWC-Pop数据集上的实验表明,我们的方法显着提高了边界检测和结构函数预测,同时保持了相当的内存使用和推理速度。
摘要:Audio-based music structure analysis (MSA) is an essential task in Music Information Retrieval that remains challenging due to the complexity and variability of musical form. Recent advances highlight the potential of fine-tuning pre-trained music foundation models for MSA tasks. However, these models are typically trained with high temporal feature resolution and short audio windows, which limits their efficiency and introduces bias when applied to long-form audio. This paper presents a temporal adaptation approach for fine-tuning music foundation models tailored to MSA. Our method enables efficient analysis of full-length songs in a single forward pass by incorporating two key strategies: (1) audio window extension and (2) low-resolution adaptation. Experiments on the Harmonix Set and RWC-Pop datasets show that our method significantly improves both boundary detection and structural function prediction, while maintaining comparable memory usage and inference speed.


【8】A Data-Centric Framework for Addressing Phonetic and Prosodic Challenges in Russian Speech Generative Models
标题:解决俄语语音生成模型中语音和韵律挑战的以数据为中心的框架
链接:https://arxiv.org/abs/2507.13563

作者:Kirill Borodin, Nikita Vasiliev, Vasiliy Kudryavtsev, Maxim Maslov, Mikhail Gorodnichev, Oleg Rogov, Grach Mkrtchian
备注:The work is still in progress
摘要:俄语语音合成提出了独特的挑战,包括元音减少,辅音清音,可变重音模式,同形异义词歧义和不自然的语调。本文介绍了Balalaika,这是一个新的数据集,包含超过2,000小时的工作室质量的俄语语音,具有全面的文本注释,包括标点符号和重音标记。实验结果表明,在Balalaika上训练的模型在语音合成和增强任务中的性能明显优于在现有数据集上训练的模型。我们详细介绍了数据集构建管道,注释方法和比较评估的结果。
摘要:Russian speech synthesis presents distinctive challenges, including vowel reduction, consonant devoicing, variable stress patterns, homograph ambiguity, and unnatural intonation. This paper introduces Balalaika, a novel dataset comprising more than 2,000 hours of studio-quality Russian speech with comprehensive textual annotations, including punctuation and stress markings. Experimental results show that models trained on Balalaika significantly outperform those trained on existing datasets in both speech synthesis and enhancement tasks. We detail the dataset construction pipeline, annotation methodology, and results of comparative evaluations.


【9】Feature-based analysis of oral narratives from Afrikaans and isiXhosa children
标题:对南非荷兰语和伊斯科萨语儿童口头叙述的基于情境的分析
链接:https://arxiv.org/abs/2507.13164

作者:ratt, Annelien Smith, Retief Louw, Daleen Klop, Febe de Wet, Herman Kamper
备注:SLaTE 2025 in Nijmegen, Netherlands
摘要:口头叙述技能是后来识字发展的强有力的预测因素。本研究探讨了口头叙述的特点,从儿童谁被确定为需要干预的专家。使用简单的机器学习方法,我们分析了四岁和五岁的南非荷兰语和西科萨语儿童的故事。与以前的研究一致,我们确定词汇多样性(独特的词)和基于长度的功能(平均话语长度)作为典型发展的指标,但发音率等功能证明信息较少。尽管跨语言的变化,部分的语音模式,使用特定的动词和助动词与目标导向的讲故事与需要干预的可能性降低。我们对两种语言学上不同的语言的分析揭示了语言特异性和共同的叙事能力预测因子,并对多语言背景下的早期评估产生了影响。
摘要:Oral narrative skills are strong predictors of later literacy development. This study examines the features of oral narratives from children who were identified by experts as requiring intervention. Using simple machine learning methods, we analyse recorded stories from four- and five-year-old Afrikaans- and isiXhosa-speaking children. Consistent with prior research, we identify lexical diversity (unique words) and length-based features (mean utterance length) as indicators of typical development, but features like articulation rate prove less informative. Despite cross-linguistic variation in part-of-speech patterns, the use of specific verbs and auxiliaries associated with goal-directed storytelling is correlated with a reduced likelihood of requiring intervention. Our analysis of two linguistically distinct languages reveals both language-specific and shared predictors of narrative proficiency, with implications for early assessment in multilingual contexts.


机器翻译由腾讯交互翻译提供,仅供参考