微信公众号:arXiv_Daily
cs.SD语音
【1】Whilter: A Whisper-based Data Filter for "In-the-Wild" Speech Corpora Using Utterance-level Multi-Task Classification
链接:https://arxiv.org/abs/2507.21642
备注:Accepted for Interspeech 2025
摘要:近年来,大规模的野生语音数据集变得越来越普遍,这是因为人们对模型越来越感兴趣,这些模型可以从未标记的数据中学习有用的特征,用于语音识别或合成等任务。这些数据集通常包含不受欢迎的特征,例如多个扬声器,非目标语言和音乐,这可能会影响模型学习。Whilter模型被提出作为一个多任务的解决方案,以确定这些不需要的样本。Whilter使用Whisper编码器和基于注意力的分类器来同时解决五个不同的分类问题。此外,还为两个流行的野外语料库的子集发布了一个带注释的数据集。Whilter在五个子任务中的三个子任务中获得了85%以上的F1分数和6.5%至7.8%的相等错误率,在语音特定类上的表现优于最先进的BEAT分类器,与单任务替代方案的组合相比,处理时间显着减少。
摘要:Large-scale in-the-wild speech datasets have become more prevalent in recent years due to increased interest in models that can learn useful features from unlabelled data for tasks such as speech recognition or synthesis. These datasets often contain undesirable features, such as multiple speakers, non-target languages, and music, which may impact model learning. The Whilter model is proposed as a multitask solution to identify these undesirable samples. Whilter uses a Whisper encoder with an attention-based classifier to solve five diverse classification problems at once. In addition, an annotated dataset is published for a subset of two popular in-the-wild corpora. Whilter achieves F1 scores above 85% and equal error rates of 6.5% to 7.8% for three of five subtasks, outperforming a state-of-the-art BEATs classifier on speech-specific classes, with a notable decrease in processing time compared to a combination of single-task alternatives.
【2】Hierarchical Graph Neural Network for Compressed Speech Steganalysis
标题:用于压缩语音隐写分析的分层图神经网络
链接:https://arxiv.org/abs/2507.21591
摘要:基于深度学习(DL)的隐写分析方法通常在计算复杂性和跨不同数据集的泛化方面面临挑战。将图神经网络(GNN)应用到隐写分析方案中,可以利用关系数据来提高检测精度和适应性。本文介绍了第一个应用程序的图形神经网络(GNN),特别是GraphSAGE架构,隐写分析压缩语音IP(VoIP)语音流。该方法涉及从VoIP流中直接构建图,并使用GraphSAGE来捕获分层隐写分析信息,包括细粒度细节和高级模式,从而实现高检测准确度。实验结果表明,该方法在发现VoIP信号中基于量化索引调制(QIM)的隐写模式方面表现良好。即使对于短的0.5秒样本,它也可以实现超过98%的检测准确度,并且在具有低嵌入率的挑战性条件下实现95.17%的准确度,比最佳性能的最先进方法提高了2.8%。此外,该模型表现出卓越的效率,对于0.5秒的样本,平均检测时间低至0.016秒,提高了0.003秒。这使得它能够有效地进行在线隐写分析任务,在低嵌入率的短样本约束下提供检测精度和效率之间的卓越平衡。
摘要:Steganalysis methods based on deep learning (DL) often struggle with computational complexity and challenges in generalizing across different datasets. Incorporating a graph neural network (GNN) into steganalysis schemes enables the leveraging of relational data for improved detection accuracy and adaptability. This paper presents the first application of a Graph Neural Network (GNN), specifically the GraphSAGE architecture, for steganalysis of compressed voice over IP (VoIP) speech streams. The method involves straightforward graph construction from VoIP streams and employs GraphSAGE to capture hierarchical steganalysis information, including both fine grained details and high level patterns, thereby achieving high detection accuracy. Experimental results demonstrate that the developed approach performs well in uncovering quantization index modulation (QIM)-based steganographic patterns in VoIP signals. It achieves detection accuracy exceeding 98 percent even for short 0.5 second samples, and 95.17 percent accuracy under challenging conditions with low embedding rates, representing an improvement of 2.8 percent over the best performing state of the art methods. Furthermore, the model exhibits superior efficiency, with an average detection time as low as 0.016 seconds for 0.5-second samples an improvement of 0.003 seconds. This makes it efficient for online steganalysis tasks, providing a superior balance between detection accuracy and efficiency under the constraint of short samples with low embedding rates.
【3】Model-free Speculative Decoding for Transformer-based ASR with Token Map Drafting
标题:基于transformer的ASR无模型推测解码与标记映射
链接:https://arxiv.org/abs/2507.21522
备注:Accepted at EUSIPCO 2025
摘要:基于Transformer架构的端到端自动语音识别(ASR)系统(例如Whisper)提供高转录准确性和鲁棒性。然而,它们的自回归解码在计算上是昂贵的,因此限制了在基于CPU和资源受限的设备上的部署。推测解码(SD)通过使用较小的草稿模型来提出候选令牌,然后由主模型验证来缓解这个问题。然而,这种方法对于缺乏GPU等硬件加速器的设备来说是不切实际的。为了解决这个问题,我们提出了一种无模型的SD技术,它消除了对单独的草稿模型的需要。相反,我们利用从特定领域的训练数据中获得的预先计算的n-gram令牌映射,以最小的开销实现高效的推测解码。我们的方法显着加速ASR推理结构化,低困惑域,而不牺牲转录的准确性。实验结果表明,在CI-AVSR数据集上的解码速度为1.27\times $,在我们的内部数据集上的解码速度为1.37\times $,而不会降低识别精度。此外,我们的方法实现了10\%$的解码速度的绝对改善超过蒸馏规格的基线运行在CPU上,突出其有效性的设备上的ASR应用程序。
摘要:End-to-end automatic speech recognition (ASR) systems based on transformer architectures, such as Whisper, offer high transcription accuracy and robustness. However, their autoregressive decoding is computationally expensive, hence limiting deployment on CPU-based and resource-constrained devices. Speculative decoding (SD) mitigates this issue by using a smaller draft model to propose candidate tokens, which are then verified by the main model. However, this approach is impractical for devices lacking hardware accelerators like GPUs. To address this, we propose \emph{Token Map Drafting}, a model-free SD technique that eliminates the need for a separate draft model. Instead, we leverage a precomputed n-gram token map derived from domain-specific training data, enabling efficient speculative decoding with minimal overhead. Our method significantly accelerates ASR inference in structured, low-perplexity domains without sacrificing transcription accuracy. Experimental results demonstrate decoding speed-ups of $1.27\times$ on the CI-AVSR dataset and $1.37\times$ on our internal dataset without degrading recognition accuracy. Additionally, our approach achieves a $10\%$ absolute improvement in decoding speed over the Distill-spec baseline running on CPU, highlighting its effectiveness for on-device ASR applications.
【4】SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
标题:SpeechFake:采用尖端生成方法的大规模多语言语音Deepfake数据集
链接:https://arxiv.org/abs/2507.21463
备注:Published in ACL 2025. Dataset available at: this https URL
摘要:随着语音生成技术的进步,通过deepfake音频滥用的风险已成为一个紧迫的问题,这凸显了对强大检测系统的迫切需求。然而,许多现有的语音deepfake数据集在规模和多样性方面都是有限的,这使得训练能够很好地推广到看不见的deepfake的模型具有挑战性。为了解决这些差距,我们引入了SpeechFake,这是一个专门为语音深度伪造检测而设计的大规模数据集。SpeechFake包含超过300万个deepfake样本,总计超过3,000小时的音频,使用40种不同的语音合成工具生成。该数据集涵盖了广泛的生成技术,包括文本到语音,语音转换和神经声码器,结合了最新的尖端方法。它还提供多语言支持,涵盖46种语言。在本文中,我们提供了数据集的创建,组成和统计的详细概述。我们还通过在SpeechFake上训练检测模型来展示基线结果,在自己的测试集和各种看不见的测试集上都表现出了强大的性能。此外,我们进行实验,严格探索如何生成方法,语言多样性和扬声器的变化影响检测性能。我们相信SpeechFake将成为推进语音深度伪造检测和开发更强大模型的宝贵资源,以不断发展生成技术。
摘要:As speech generation technology advances, the risk of misuse through deepfake audio has become a pressing concern, which underscores the critical need for robust detection systems. However, many existing speech deepfake datasets are limited in scale and diversity, making it challenging to train models that can generalize well to unseen deepfakes. To address these gaps, we introduce SpeechFake, a large-scale dataset designed specifically for speech deepfake detection. SpeechFake includes over 3 million deepfake samples, totaling more than 3,000 hours of audio, generated using 40 different speech synthesis tools. The dataset encompasses a wide range of generation techniques, including text-to-speech, voice conversion, and neural vocoder, incorporating the latest cutting-edge methods. It also provides multilingual support, spanning 46 languages. In this paper, we offer a detailed overview of the dataset's creation, composition, and statistics. We also present baseline results by training detection models on SpeechFake, demonstrating strong performance on both its own test sets and various unseen test sets. Additionally, we conduct experiments to rigorously explore how generation methods, language diversity, and speaker variation affect detection performance. We believe SpeechFake will be a valuable resource for advancing speech deepfake detection and developing more robust models for evolving generation techniques.
【5】Relationship between objective and subjective perceptual measures of speech in individuals with head and neck cancer
标题:头部和颈部癌患者言语的客观和主观知觉指标之间的关系
链接:https://arxiv.org/abs/2507.21426
备注:5 pages, 1 figure, 1 table. Accepted at Interspeech 2025
摘要:有意义的言语评估在临床语音学和治疗监测中至关重要。本研究在一个大型头颈癌(HNC)数据集中检查了感知语音评估和客观声学测量之间的联系。受过训练的听众提供了可懂度,清晰度,语音质量,发声,语速,鼻音和背景噪音的语音评级。主观可懂度、清晰度和语音质量之间存在很强的相关性,这可能是由于我们说话者群体中语音症状的共同根本原因。清晰度和语速的客观测量与主观测量一致。我们的研究结果表明,一个单一的清晰度措施可能是足够的临床监测扬声器治疗HNC使用伴随放化疗。
摘要:Meaningful speech assessment is vital in clinical phonetics and therapy monitoring. This study examined the link between perceptual speech assessments and objective acoustic measures in a large head and neck cancer (HNC) dataset. Trained listeners provided ratings of intelligibility, articulation, voice quality, phonation, speech rate, nasality, and background noise on speech. Strong correlations were found between subjective intelligibility, articulation, and voice quality, likely due to a shared underlying cause of speech symptoms in our speaker population. Objective measures of intelligibility and speech rate aligned with their subjective counterpart. Our results suggest that a single intelligibility measure may be sufficient for the clinical monitoring of speakers treated for HNC using concomitant chemoradiation.
【6】Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion
标题:Sync-TVA:一个基于跨模态融合的多模态情感识别框架
链接:https://arxiv.org/abs/2507.21395
摘要:多模态情感识别(MER)对于实现感知和响应人类情感的情感智能系统至关重要。然而,现有的方法受到有限的跨模态的相互作用和跨模态的不平衡的贡献。为了解决这些问题,我们提出了Sync-TVA,一个端到端的图形注意力框架,具有特定模态的动态增强和结构化的跨模态融合。我们的设计采用了一个动态增强模块,为每一个模态和构造异构的跨模态图,跨文本,音频和视觉功能的语义关系模型。交叉注意融合机制进一步对齐多模态线索,以实现强大的情感推理。MELD和IEMOCAP上的实验表明,与最先进的模型相比,在准确性和加权F1分数方面都有了一致的改进,特别是在类不平衡的条件下。
摘要:Multimodal emotion recognition (MER) is crucial for enabling emotionally intelligent systems that perceive and respond to human emotions. However, existing methods suffer from limited cross-modal interaction and imbalanced contributions across modalities. To address these issues, we propose Sync-TVA, an end-to-end graph-attention framework featuring modality-specific dynamic enhancement and structured cross-modal fusion. Our design incorporates a dynamic enhancement module for each modality and constructs heterogeneous cross-modal graphs to model semantic relations across text, audio, and visual features. A cross-attention fusion mechanism further aligns multimodal cues for robust emotion inference. Experiments on MELD and IEMOCAP demonstrate consistent improvements over state-of-the-art models in both accuracy and weighted F1 score, especially under class-imbalanced conditions.
【7】A Deep Learning Automatic Speech Recognition Model for Shona Language
标题:Shona语言深度学习自动语音识别模型
链接:https://arxiv.org/abs/2507.21331
备注:None
摘要:这项研究提出了一种基于深度学习的自动语音识别系统的开发,用于Shona,这是一种低资源语言,其特点是独特的音调和语法复杂性。该研究旨在解决有限的训练数据,缺乏标记数据以及Shona语音中存在的复杂音调细微差别所带来的挑战,其目标是与传统统计模型相比,识别准确性得到显着提高。该研究首先探讨了使用深度学习为Shona开发准确的ASR系统的可行性。其次,它研究了设计和实现Shona语音识别深度学习架构所涉及的具体挑战,并提出了缓解这些挑战的策略。最后,它在准确性方面比较了基于深度学习的模型与现有统计模型的性能。开发的ASR系统采用了一种混合体系结构,包括用于声学建模的卷积神经网络和用于语言建模的长短期记忆网络。为了克服数据不足的问题,采用了数据扩充技术和迁移学习。注意力机制也被纳入,以适应绍纳语的音调性质。由此产生的ASR系统取得了令人印象深刻的结果,单词错误率为29%,音素错误率为12%,整体准确率为74%。这些指标表明,深度学习有潜力提高Shona等资源不足语言的ASR准确性。这项研究有助于推进ASR技术在Shona等资源不足的语言中的应用,最终促进全世界Shona语使用者的可访问性和沟通。
摘要:This study presented the development of a deep learning-based Automatic Speech Recognition system for Shona, a low-resource language characterized by unique tonal and grammatical complexities. The research aimed to address the challenges posed by limited training data, lack of labelled data, and the intricate tonal nuances present in Shona speech, with the objective of achieving significant improvements in recognition accuracy compared to traditional statistical models. The research first explored the feasibility of using deep learning to develop an accurate ASR system for Shona. Second, it investigated the specific challenges involved in designing and implementing deep learning architectures for Shona speech recognition and proposed strategies to mitigate these challenges. Lastly, it compared the performance of the deep learning-based model with existing statistical models in terms of accuracy. The developed ASR system utilized a hybrid architecture consisting of a Convolutional Neural Network for acoustic modelling and a Long Short-Term Memory network for language modelling. To overcome the scarcity of data, data augmentation techniques and transfer learning were employed. Attention mechanisms were also incorporated to accommodate the tonal nature of Shona speech. The resulting ASR system achieved impressive results, with a Word Error Rate of 29%, Phoneme Error Rate of 12%, and an overall accuracy of 74%. These metrics indicated the potential of deep learning to enhance ASR accuracy for under-resourced languages like Shona. This study contributed to the advancement of ASR technology for under-resourced languages like Shona, ultimately fostering improved accessibility and communication for Shona speakers worldwide.
【8】Combolutional Neural Networks
标题:组合神经网络
链接:https://arxiv.org/abs/2507.21202
备注:4 pages, 3 figures, accepted to WASPAA 2025
摘要:选择适当的归纳偏差是机器学习模型设计的重要步骤,尤其是在处理音频时,即使是很短的剪辑也可能包含数百万个样本。为此,我们提出了combolutional层:一个学习延迟IIR梳状滤波器和融合包络检测器,提取谐波特征的时域。我们展示了combolutional层在三个信息检索任务上的功效,评估了其相对于其他音频前端的计算成本,并提供了有效的训练实现。我们发现,在精确的谐波分析很重要的音频任务中,组合层是卷积层的有效替代,例如,钢琴转录、说话人分类和键检测。此外,与现有前端相比,组合层还有其他几个关键优势,即:低参数计数,高效的CPU推理,严格的实值计算和改进的可解释性。
摘要:Selecting appropriate inductive biases is an essential step in the design of machine learning models, especially when working with audio, where even short clips may contain millions of samples. To this end, we propose the combolutional layer: a learned-delay IIR comb filter and fused envelope detector, which extracts harmonic features in the time domain. We demonstrate the efficacy of the combolutional layer on three information retrieval tasks, evaluate its computational cost relative to other audio frontends, and provide efficient implementations for training. We find that the combolutional layer is an effective replacement for convolutional layers in audio tasks where precise harmonic analysis is important, e.g., piano transcription, speaker classification, and key detection. Additionally, the combolutional layer has several other key benefits over existing frontends, namely: low parameter count, efficient CPU inference, strictly real-valued computations, and improved interpretability.
【9】TTS-1 Technical Report
标题:TTC-1技术报告
链接:https://arxiv.org/abs/2507.21138
摘要:我们介绍Inworld TTS-1,一组两个基于transformer的自回归文本到语音(TTS)模型。我们最大的型号TTS-1-Max具有8.8B参数,专为满足苛刻应用的最高质量和表现力而设计。TTS-1是我们最高效的模型,具有1.6B参数,专为实时语音合成和设备上的用例而构建。通过扩展训练时间计算并应用语音语言模型(SpeechLM)组件的预训练,微调和RL对齐的顺序过程,这两个模型在各种基准上都实现了最先进的性能,表现出纯粹依赖于说话者语音的上下文学习的卓越质量。Inworld TTS-1和TTS-1-Max可以生成高分辨率的48 kHz语音,具有低延迟,并支持11种语言,通过音频标记进行精细的情感控制和非语言发声。我们还在MIT许可证下开源了我们的训练和建模代码。
摘要:We introduce Inworld TTS-1, a set of two Transformer-based autoregressive text-to-speech (TTS) models. Our largest model, TTS-1-Max, has 8.8B parameters and is designed for utmost quality and expressiveness in demanding applications. TTS-1 is our most efficient model, with 1.6B parameters, built for real-time speech synthesis and on-device use cases. By scaling train-time compute and applying a sequential process of pre-training, fine-tuning, and RL-alignment of the speech-language model (SpeechLM) component, both models achieve state-of-the-art performance on a variety of benchmarks, demonstrating exceptional quality relying purely on in-context learning of the speaker's voice. Inworld TTS-1 and TTS-1-Max can generate high-resolution 48 kHz speech with low latency, and support 11 languages with fine-grained emotional control and non-verbal vocalizations through audio markups. We additionally open-source our training and modeling code under an MIT license.
【1】Real-Time Audio-Visual Speech Enhancement Using Pre-trained Visual Representations
标题:基于预训练视觉表征的实时音视频语音增强
链接:https://arxiv.org/abs/2507.21448
备注:Accepted into Interspeech 2025
摘要:纯音频环境中的语音增强仍然具有挑战性,特别是在存在干扰扬声器的情况下。本文提出了一种简单而有效的实时音视频语音增强系统RAVEN,它隔离和增强屏幕上的目标说话人,同时抑制干扰说话人和背景噪声。我们研究了从视听语音识别(AVSR)和主动说话人检测(ASD)中学习到的视觉嵌入如何在不同的SNR条件和干扰说话人数量下对AVSE做出贡献。我们的研究结果表明,从AVSR和ASD模型的级联嵌入提供了最大的改善,在低信噪比,多扬声器环境,而AVSR嵌入单独执行最好的只有噪音的情况下。此外,我们还开发了一个在计算机CPU上运行的实时流媒体系统,并提供了视频演示和代码库。据我们所知,这是实时AVSE系统的第一个开源实现。
摘要:Speech enhancement in audio-only settings remains challenging, particularly in the presence of interfering speakers. This paper presents a simple yet effective real-time audio-visual speech enhancement (AVSE) system, RAVEN, which isolates and enhances the on-screen target speaker while suppressing interfering speakers and background noise. We investigate how visual embeddings learned from audio-visual speech recognition (AVSR) and active speaker detection (ASD) contribute to AVSE across different SNR conditions and numbers of interfering speakers. Our results show concatenating embeddings from AVSR and ASD models provides the greatest improvement in low-SNR, multi-speaker environments, while AVSR embeddings alone perform best in noise-only scenarios. In addition, we develop a real-time streaming system that operates on a computer CPU and we provide a video demonstration and code repository. To our knowledge, this is the first open-source implementation of a real-time AVSE system.
【2】SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods
标题:SpeechFake:采用尖端生成方法的大规模多语言语音Deepfake数据集
链接:https://arxiv.org/abs/2507.21463
备注:Published in ACL 2025. Dataset available at: this https URL
摘要:随着语音生成技术的进步,通过deepfake音频滥用的风险已成为一个紧迫的问题,这凸显了对强大检测系统的迫切需求。然而,许多现有的语音deepfake数据集在规模和多样性方面都是有限的,这使得训练能够很好地推广到看不见的deepfake的模型具有挑战性。为了解决这些差距,我们引入了SpeechFake,这是一个专门用于语音深度伪造检测的大规模数据集。SpeechFake包含超过300万个deepfake样本,总计超过3,000小时的音频,使用40种不同的语音合成工具生成。该数据集涵盖了广泛的生成技术,包括文本到语音,语音转换和神经声码器,结合了最新的尖端方法。它还提供多语言支持,涵盖46种语言。在本文中,我们提供了数据集的创建,组成和统计的详细概述。我们还通过在SpeechFake上训练检测模型来展示基线结果,在自己的测试集和各种看不见的测试集上都表现出了强大的性能。此外,我们进行实验,严格探索如何生成方法,语言多样性和扬声器的变化影响检测性能。我们相信SpeechFake将成为推进语音深度伪造检测和开发更强大模型的宝贵资源,以不断发展生成技术。
摘要:As speech generation technology advances, the risk of misuse through deepfake audio has become a pressing concern, which underscores the critical need for robust detection systems. However, many existing speech deepfake datasets are limited in scale and diversity, making it challenging to train models that can generalize well to unseen deepfakes. To address these gaps, we introduce SpeechFake, a large-scale dataset designed specifically for speech deepfake detection. SpeechFake includes over 3 million deepfake samples, totaling more than 3,000 hours of audio, generated using 40 different speech synthesis tools. The dataset encompasses a wide range of generation techniques, including text-to-speech, voice conversion, and neural vocoder, incorporating the latest cutting-edge methods. It also provides multilingual support, spanning 46 languages. In this paper, we offer a detailed overview of the dataset's creation, composition, and statistics. We also present baseline results by training detection models on SpeechFake, demonstrating strong performance on both its own test sets and various unseen test sets. Additionally, we conduct experiments to rigorously explore how generation methods, language diversity, and speaker variation affect detection performance. We believe SpeechFake will be a valuable resource for advancing speech deepfake detection and developing more robust models for evolving generation techniques.
【3】Sound Source Localization for Human-Robot Interaction in Outdoor Environments
标题:户外环境中人机交互的光源定位
链接:https://arxiv.org/abs/2507.21431
摘要:本文提出了一种声源定位策略,它依赖于一个麦克风阵列嵌入在一个无人驾驶地面车辆和一个异步近距离说话的麦克风附近的操作员。信号粗对准策略与时域声学回声消除算法相结合,估计时频理想比掩模,以将目标语音与干扰和环境噪声隔离。这允许选择性声源定位,并为机器人提供来自主动操作员的声音到达方向,从而在嘈杂的场景中实现丰富的交互。实验结果表明,在信噪比为1dB时,平均角度误差为4 °,5 °以内的定位精度为95%,明显优于现有的定位方法。
摘要:This paper presents a sound source localization strategy that relies on a microphone array embedded in an unmanned ground vehicle and an asynchronous close-talking microphone near the operator. A signal coarse alignment strategy is combined with a time-domain acoustic echo cancellation algorithm to estimate a time-frequency ideal ratio mask to isolate the target speech from interferences and environmental noise. This allows selective sound source localization, and provides the robot with the direction of arrival of sound from the active operator, which enables rich interaction in noisy scenarios. Results demonstrate an average angle error of 4 degrees and an accuracy within 5 degrees of 95\% at a signal-to-noise ratio of 1dB, which is significantly superior to the state-of-the-art localization methods.
【4】Relationship between objective and subjective perceptual measures of speech in individuals with head and neck cancer
标题:头部和颈部癌患者言语的客观和主观知觉指标之间的关系
链接:https://arxiv.org/abs/2507.21426
备注:5 pages, 1 figure, 1 table. Accepted at Interspeech 2025
摘要:有意义的言语评估在临床语音学和治疗监测中至关重要。本研究在一个大型头颈癌(HNC)数据集中检查了感知语音评估和客观声学测量之间的联系。受过训练的听众提供了可懂度,清晰度,语音质量,发声,语速,鼻音和背景噪音的语音评级。主观可懂度、清晰度和语音质量之间存在很强的相关性,这可能是由于我们说话者群体中语音症状的共同根本原因。清晰度和语速的客观测量与主观测量一致。我们的研究结果表明,一个单一的清晰度措施可能是足够的临床监测扬声器治疗HNC使用伴随放化疗。
摘要:Meaningful speech assessment is vital in clinical phonetics and therapy monitoring. This study examined the link between perceptual speech assessments and objective acoustic measures in a large head and neck cancer (HNC) dataset. Trained listeners provided ratings of intelligibility, articulation, voice quality, phonation, speech rate, nasality, and background noise on speech. Strong correlations were found between subjective intelligibility, articulation, and voice quality, likely due to a shared underlying cause of speech symptoms in our speaker population. Objective measures of intelligibility and speech rate aligned with their subjective counterpart. Our results suggest that a single intelligibility measure may be sufficient for the clinical monitoring of speakers treated for HNC using concomitant chemoradiation.
【5】A Deep Learning Automatic Speech Recognition Model for Shona Language
标题:Shona语言深度学习自动语音识别模型
链接:https://arxiv.org/abs/2507.21331
备注:None
摘要:这项研究提出了一种基于深度学习的自动语音识别系统的开发,用于Shona,这是一种低资源语言,其特点是独特的音调和语法复杂性。该研究旨在解决有限的训练数据,缺乏标记数据以及Shona语音中存在的复杂音调细微差别所带来的挑战,其目标是与传统统计模型相比,识别准确性得到显着提高。该研究首先探讨了使用深度学习为Shona开发准确的ASR系统的可行性。其次,它研究了设计和实现Shona语音识别深度学习架构所涉及的具体挑战,并提出了缓解这些挑战的策略。最后,它在准确性方面比较了基于深度学习的模型与现有统计模型的性能。开发的ASR系统采用了一种混合体系结构,包括用于声学建模的卷积神经网络和用于语言建模的长短期记忆网络。为了克服数据不足的问题,采用了数据扩充技术和迁移学习。注意力机制也被纳入,以适应绍纳语的音调性质。由此产生的ASR系统取得了令人印象深刻的结果,单词错误率为29%,音素错误率为12%,整体准确率为74%。这些指标表明,深度学习有潜力提高Shona等资源不足语言的ASR准确性。这项研究有助于推进ASR技术在Shona等资源不足的语言中的应用,最终促进全世界Shona语使用者的可访问性和沟通。
摘要:This study presented the development of a deep learning-based Automatic Speech Recognition system for Shona, a low-resource language characterized by unique tonal and grammatical complexities. The research aimed to address the challenges posed by limited training data, lack of labelled data, and the intricate tonal nuances present in Shona speech, with the objective of achieving significant improvements in recognition accuracy compared to traditional statistical models. The research first explored the feasibility of using deep learning to develop an accurate ASR system for Shona. Second, it investigated the specific challenges involved in designing and implementing deep learning architectures for Shona speech recognition and proposed strategies to mitigate these challenges. Lastly, it compared the performance of the deep learning-based model with existing statistical models in terms of accuracy. The developed ASR system utilized a hybrid architecture consisting of a Convolutional Neural Network for acoustic modelling and a Long Short-Term Memory network for language modelling. To overcome the scarcity of data, data augmentation techniques and transfer learning were employed. Attention mechanisms were also incorporated to accommodate the tonal nature of Shona speech. The resulting ASR system achieved impressive results, with a Word Error Rate of 29%, Phoneme Error Rate of 12%, and an overall accuracy of 74%. These metrics indicated the potential of deep learning to enhance ASR accuracy for under-resourced languages like Shona. This study contributed to the advancement of ASR technology for under-resourced languages like Shona, ultimately fostering improved accessibility and communication for Shona speakers worldwide.
【6】TTS-1 Technical Report
标题:TTC-1技术报告
链接:https://arxiv.org/abs/2507.21138
摘要:我们介绍Inworld TTS-1,一组两个基于transformer的自回归文本到语音(TTS)模型。我们最大的型号TTS-1-Max具有8.8B参数,专为满足苛刻应用的最高质量和表现力而设计。TTS-1是我们最高效的模型,具有1.6B参数,专为实时语音合成和设备上的用例而构建。通过扩展训练时间计算并应用语音语言模型(SpeechLM)组件的预训练,微调和RL对齐的顺序过程,这两个模型在各种基准上都实现了最先进的性能,表现出纯粹依赖于说话者语音的上下文学习的卓越质量。Inworld TTS-1和TTS-1-Max能够以低延迟生成高分辨率48 kHz语音,并通过音频标记支持11种语言的精细情绪控制和非语言发声。我们还在MIT许可证下开源了我们的训练和建模代码。
摘要:We introduce Inworld TTS-1, a set of two Transformer-based autoregressive text-to-speech (TTS) models. Our largest model, TTS-1-Max, has 8.8B parameters and is designed for utmost quality and expressiveness in demanding applications. TTS-1 is our most efficient model, with 1.6B parameters, built for real-time speech synthesis and on-device use cases. By scaling train-time compute and applying a sequential process of pre-training, fine-tuning, and RL-alignment of the speech-language model (SpeechLM) component, both models achieve state-of-the-art performance on a variety of benchmarks, demonstrating exceptional quality relying purely on in-context learning of the speaker's voice. Inworld TTS-1 and TTS-1-Max can generate high-resolution 48 kHz speech with low latency, and support 11 languages with fine-grained emotional control and non-verbal vocalizations through audio markups. We additionally open-source our training and modeling code under an MIT license.
【7】Hyperbolic Embeddings for Order-Aware Classification of Audio Effect Chains
标题:音效链的订单感知分类的双曲嵌入
链接:https://arxiv.org/abs/2507.20624
备注:7 pages, 3 figures, accepted for the 28th International Conference on Digital Audio Effects (DAFx25)
摘要:音频效果(AFX)是音乐制作中必不可少的工具,经常用于塑造音色和动态。链中AFX的顺序在确定最终声音中起着至关重要的作用,特别是当非线性(例如,失真)或时变(例如,合唱)处理器。尽管它的重要性,大多数AFX相关的研究主要集中在估计效果类型和它们的参数从湿信号。为了解决这个差距,我们制定AFX链识别的任务,联合估计AFX类型和他们的顺序从湿信号。我们提出了一种基于神经网络的方法,将湿信号嵌入到双曲空间中,并对其AFX链进行分类。双曲空间由于其指数扩展特性,可以比欧氏空间更有效地表示树结构数据。由于AFX链可以表示为树,AFX作为节点和边编码效果顺序,因此双曲空间非常适合建模有序AFX组合的指数增长和非交换性质,其中效果顺序的变化可能导致不同的最终声音。使用吉他声音的实验表明,与适当的曲率,所提出的方法优于其欧几里德对应。基于AFX类型和链长的进一步分析突出了所提出的方法在捕获AFX阶数方面的有效性。
摘要:Audio effects (AFXs) are essential tools in music production, frequently applied in chains to shape timbre and dynamics. The order of AFXs in a chain plays a crucial role in determining the final sound, particularly when non-linear (e.g., distortion) or time-variant (e.g., chorus) processors are involved. Despite its importance, most AFX-related studies have primarily focused on estimating effect types and their parameters from a wet signal. To address this gap, we formulate AFX chain recognition as the task of jointly estimating AFX types and their order from a wet signal. We propose a neural-network-based method that embeds wet signals into a hyperbolic space and classifies their AFX chains. Hyperbolic space can represent tree-structured data more efficiently than Euclidean space due to its exponential expansion property. Since AFX chains can be represented as trees, with AFXs as nodes and edges encoding effect order, hyperbolic space is well-suited for modeling the exponentially growing and non-commutative nature of ordered AFX combinations, where changes in effect order can result in different final sounds. Experiments using guitar sounds demonstrate that, with an appropriate curvature, the proposed method outperforms its Euclidean counterpart. Further analysis based on AFX type and chain length highlights the effectiveness of the proposed method in capturing AFX order.
机器翻译由腾讯交互翻译提供,仅供参考
