今日论文合集:cs.SD语音7篇,eess.AS音频处理11篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】BATON: Aligning Text-to-Audio Model with Human Preference Feedback

标题:指挥棒:使文本到音频模型与人类偏好反馈保持一致

链接:http://arxiv.org/pdf/2402.00744v1

作者:Huan Liao,Haonan Han,Kai Yang,Tianjiao Du,Rui Yang,Zunnan Xu,Qinmei Xu,Jingquan Liu,Jiasheng Lu,Xiu Li

摘要:随着人工智能生成内容(AIGC)的发展,文本到音频模型受到广泛关注。然而,由于自然语言固有的信息密度和有限的模型理解能力,这些模型生成符合人类偏好的音频是具有挑战性的。为了缓解这个问题,我们制定了指挥棒,一个框架,旨在提高生成的音频和文本提示使用人类偏好反馈之间的对齐。我们的BATON包括三个关键阶段:首先,我们策划了一个包含提示和相应生成的音频的数据集,然后根据人类反馈对其进行注释。其次,我们引入了一个奖励模型,使用构造的数据集,它可以模仿人类的偏好,通过分配奖励输入文本音频对。最后,我们采用奖励模型来微调现成的文本到音频模型。实验结果表明,我们的BATON可以显着提高原始文本到音频模型的生成质量,涉及音频的完整性,时间关系,并与人类的偏好对齐。

摘要:With the development of AI-Generated Content (AIGC), text-to-audio models are gaining widespread attention. However, it is challenging for these models to generate audio aligned with human preference due to the inherent information density of natural language and limited model understanding ability. To alleviate this issue, we formulate the BATON, a framework designed to enhance the alignment between generated audio and text prompt using human preference feedback. Our BATON comprises three key stages: Firstly, we curated a dataset containing both prompts and the corresponding generated audio, which was then annotated based on human feedback. Secondly, we introduced a reward model using the constructed dataset, which can mimic human preference by assigning rewards to input text-audio pairs. Finally, we employed the reward model to fine-tune an off-the-shelf text-to-audio model. The experiment results demonstrate that our BATON can significantly improve the generation quality of the original text-to-audio models, concerning audio integrity, temporal relationship, and alignment with human preference.


【2】 Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
标题:你能去掉具有自监督语音特征的说话人识别的下游模型吗?
链接:http://arxiv.org/pdf/2402.00340v1
作者:Zakaria Aldeneh,Takuya Higuchi,Jee-weon Jung,Skyler Seto,Tatiana Likhomanenko,Stephen Shum,Ahmed Hussen Abdelaziz,Shinji Watanabe,Barry-John Theobald
摘要:在说话人确认模型中,自监督特征通常用于代替滤波器组。然而,这些模型最初被设计为摄取滤波器组作为输入,因此,在自监督特征之上训练它们假设两种特征类型需要相同的任务学习量。在这项工作中,我们观察到预训练的自监督语音特征固有地包括下游说话人验证任务所需的信息,因此,我们可以在不牺牲性能的情况下简化下游模型。为此,我们重新审视了使用自监督特征的说话人验证下游模型的设计。我们表明,我们可以简化模型,使用少97.51%的参数,同时实现29.93%的平均性能提高SUPERB。因此,我们证明了简化的下游模型比基线模型更有效-它仅用60%的训练数据就实现了更好的性能。
摘要:Self-supervised features are typically used in place of filter-banks in speaker verification models. However, these models were originally designed to ingest filter-banks as inputs, and thus, training them on top of self-supervised features assumes that both feature types require the same amount of learning for the task. In this work, we observe that pre-trained self-supervised speech features inherently include information required for downstream speaker verification task, and therefore, we can simplify the downstream model without sacrificing performance. To this end, we revisit the design of the downstream model for speaker verification using self-supervised features. We show that we can simplify the model to use 97.51% fewer parameters while achieving a 29.93% average improvement in performance on SUPERB. Consequently, we show that the simplified downstream model is more data efficient compared to baseline--it achieves better performance with only 60% of the training data.

【3】 Exploring the limits of decoder-only models trained on public speech recognition corpora
标题:探索在公共语音识别语料库上训练的仅限解码器的模型的局限性
链接:http://arxiv.org/pdf/2402.00235v1
作者:Ankit Gupta,George Saon,Brian Kingsbury
摘要:工业规模语音识别(ASR)模型的出现,如Whisper和USM,分别在100万小时的弱标记和1200万小时的仅音频专有数据上进行训练,导致对大规模公共ASR语料库和竞争性开源管道的需求更加强烈。与上述模型不同,大型语言模型通常基于Transformer解码器,目前还不清楚仅基于公共数据训练的解码器模型是否可以提供有竞争力的性能。在这项工作中,我们调查的因素,如选择的训练数据集和建模组件,以获得最佳的性能,单独使用公共英语ASR语料库。我们的Decoder-Only Transformer for ASR(DOTA)模型在几乎所有英语ASR基准测试中的表现都全面优于Whisper(OWSM)的编码器-解码器开源复制,并在15个测试集中的7个测试中优于Whisper large-v3。我们在许可证下发布我们的代码库和模型检查点。
摘要:The emergence of industrial-scale speech recognition (ASR) models such as Whisper and USM, trained on 1M hours of weakly labelled and 12M hours of audio only proprietary data respectively, has led to a stronger need for large scale public ASR corpora and competitive open source pipelines. Unlike the said models, large language models are typically based on Transformer decoders, and it remains unclear if decoder-only models trained on public data alone can deliver competitive performance. In this work, we investigate factors such as choice of training datasets and modeling components necessary for obtaining the best performance using public English ASR corpora alone. Our Decoder-Only Transformer for ASR (DOTA) model comprehensively outperforms the encoder-decoder open source replication of Whisper (OWSM) on nearly all English ASR benchmarks and outperforms Whisper large-v3 on 7 out of 15 test sets. We release our codebase and model checkpoints under permissive license.

【4】 An Analysis of the Variance of Diffusion-based Speech Enhancement
标题:基于扩散的语音增强的方差分析
链接:http://arxiv.org/pdf/2402.00811v1
作者:Bunlong Lay,Timo Gerkmann
备注:5 pages, 3 figures, 1 table
摘要:扩散模型被证明是生成式语音增强的强大模型。在最近的SGMSE+方法中,训练涉及扩散过程的随机微分方程,逐渐将高斯和环境噪声添加到干净的语音信号中。语音增强性能的变化取决于随机微分方程的选择,该随机微分方程在添加环境和高斯噪声时控制沿扩散过程的均值和方差的演变。在这项工作中,我们强调,方差的规模是语音增强性能的主要参数,并表明它控制噪声衰减和语音失真之间的权衡。更具体地说,我们表明,更大的方差增加了噪声衰减,并允许减少计算足迹,因为需要更少的函数评估来生成估计。
摘要:Diffusion models proved to be powerful models for generative speech enhancement. In recent SGMSE+ approaches, training involves a stochastic differential equation for the diffusion process, adding both Gaussian and environmental noise to the clean speech signal gradually. The speech enhancement performance varies depending on the choice of the stochastic differential equation that controls the evolution of the mean and the variance along the diffusion processes when adding environmental and Gaussian noise. In this work, we highlight that the scale of the variance is a dominant parameter for speech enhancement performance and show that it controls the tradeoff between noise attenuation and speech distortions. More concretely, we show that a larger variance increases the noise attenuation and allows for reducing the computational footprint, as fewer function evaluations for generating the estimate are required.

【5】 Frame-Wise Breath Detection with Self-Training: An Exploration of Enhancing Breath Naturalness in Text-to-Speech
标题:基于自训练的帧级呼吸检测:增强文语转换中呼吸自然度的探索
链接:http://arxiv.org/pdf/2402.00288v1
作者:Dong Yang,Tomoki Koriyama,Yuki Saito
摘要:开发可以合成自然呼吸的文本到语音(TTS)系统对于类人语音代理是必不可少的,但需要在训练数据中对呼吸位置进行大量手动注释。为此,我们提出了一个自我训练的方法来训练呼吸检测模型,可以自动检测呼吸的位置在讲话中。我们的方法使用大型语音语料库训练模型,包括:1)利用基于规则的方法对有限的呼吸声进行注释,以及2)通过基于模型预测的伪标记对这些注释进行迭代增强。我们的检测模型采用具有下采样 上采样层的Conformer块,从而实现准确的逐帧呼吸检测。我们研究其有效性,在多扬声器TTS使用文本转录检测呼吸标记。结果表明,使用我们提出的模型的呼吸检测和呼吸标记插入合成呼吸包含的语音比基线模型更自然。
摘要:Developing Text-to-Speech (TTS) systems that can synthesize natural breath is essential for human-like voice agents but requires extensive manual annotation of breath positions in training data. To this end, we propose a self-training method for training a breath detection model that can automatically detect breath positions in speech. Our method trains the model using a large speech corpus and involves: 1) annotation of limited breath sounds utilizing a rule-based approach, and 2) iterative augmentation of these annotations through pseudo-labeling based on the model's predictions. Our detection model employs Conformer blocks with down- up-sampling layers, enabling accurate frame-wise breath detection. We investigate its effectiveness in multi-speaker TTS using text transcripts with detected breath marks. The results indicate that using our proposed model for breath detection and breath mark insertion synthesizes breath-contained speech more naturally than a baseline model.

【6】 PAM: Prompting Audio-Language Models for Audio Quality Assessment
标题:PAM:促进音频质量评估的音频语言模型
链接:http://arxiv.org/pdf/2402.00282v1
作者:Soham Deshmukh,Dareen Alharthi,Benjamin Elizalde,Hannes Gamper,Mahmoud Al Ismail,Rita Singh,Bhiksha Raj,Huaming Wang
摘要:虽然音频质量是各种音频处理任务(包括生成建模)的关键性能指标,但其客观测量仍然是一个挑战。音频语言模型(ALM)是在音频文本对上进行预训练的,这些音频文本对可能包含有关音频质量、伪像存在或噪声的信息。给定音频输入和与质量相关的文本提示,ALM可以用于计算两者之间的相似性分数。在这里,我们利用这种能力,并介绍PAM,一个无参考指标,用于评估不同的音频处理任务的音频质量。与其他“无参考”指标相反,PAM不需要在参考数据集上计算嵌入,也不需要在一组昂贵的人类听力分数上训练特定于任务的模型。我们广泛评估了PAM在四个任务上的可靠性:文本到音频(TTA),文本到音乐生成(TTM),文本到语音(TTS)和深度噪声抑制(DNS)。我们进行多个消融研究与控制失真,在野外设置,并提示选择。我们的评估表明,PAM与现有的指标和人类听力分数相关。这些结果证明了ALM用于计算通用音频质量度量的潜力。
摘要:While audio quality is a key performance metric for various audio processing tasks, including generative modeling, its objective measurement remains a challenge. Audio-Language Models (ALMs) are pre-trained on audio-text pairs that may contain information about audio quality, the presence of artifacts, or noise. Given an audio input and a text prompt related to quality, an ALM can be used to calculate a similarity score between the two. Here, we exploit this capability and introduce PAM, a no-reference metric for assessing audio quality for different audio processing tasks. Contrary to other "reference-free" metrics, PAM does not require computing embeddings on a reference dataset nor training a task-specific model on a costly set of human listening scores. We extensively evaluate the reliability of PAM against established metrics and human listening scores on four tasks: text-to-audio (TTA), text-to-music generation (TTM), text-to-speech (TTS), and deep noise suppression (DNS). We perform multiple ablation studies with controlled distortions, in-the-wild setups, and prompt choices. Our evaluation shows that PAM correlates well with existing metrics and human listening scores. These results demonstrate the potential of ALMs for computing a general-purpose audio quality metric.

【7】 Online speaker diarization of meetings guided by speech separation
标题:通过语音分离指导的在线会议发言人日记
链接:http://arxiv.org/pdf/2402.00067v1
作者:Elio Gruttadauria,Mathieu Fontaine,Slim Essid
Journal-ref:IEEE International Conference on Acoustics, Speech, and Signal Processing, Apr 2024, Seoul (Korea), South Korea
摘要:重叠语音对于说话人日记系统是出了名的问题。因此,最近提出了使用语音分离来改善它们的性能。虽然有前途,语音分离模型与现实数据的斗争,因为它们是在具有固定数量的扬声器的模拟混合物上训练的。在这项工作中,我们引入了一个新的语音分离指导的日记计划适合于在线扬声器日记的长会议录音与可变数量的扬声器,目前在AMI语料库。我们设想ConvTasNet和DPRNN作为分离网络的替代方案,具有两个或三个输出源。为了获得说话人日志化结果,对每个估计的源应用语音活动检测。最终的模型是端到端微调后,首先适应分离的实际数据使用AMI。该系统在短段上运行,并且通过使用说话人嵌入和增量聚类拼接本地预测来执行推理。结果表明,我们的系统改进了AMI耳机混合的最新技术,不使用任何oracle信息,并在充分评估(无衣领,包括重叠的语音)。最后,我们展示了我们的系统的强度,特别是在重叠的语音部分。
摘要:Overlapped speech is notoriously problematic for speaker diarization systems. Consequently, the use of speech separation has recently been proposed to improve their performance. Although promising, speech separation models struggle with realistic data because they are trained on simulated mixtures with a fixed number of speakers. In this work, we introduce a new speech separation-guided diarization scheme suitable for the online speaker diarization of long meeting recordings with a variable number of speakers, as present in the AMI corpus. We envisage ConvTasNet and DPRNN as alternatives for the separation networks, with two or three output sources. To obtain the speaker diarization result, voice activity detection is applied on each estimated source. The final model is fine-tuned end-to-end, after first adapting the separation to real data using AMI. The system operates on short segments, and inference is performed by stitching the local predictions using speaker embeddings and incremental clustering. The results show that our system improves the state-of-the-art on the AMI headset mix, using no oracle information and under full evaluation (no collar and including overlapped speech). Finally, we show the strength of our system particularly on overlapped speech sections.

eess.AS音频处理
【1】 Deep Room Impulse Response Completion
标题:深室脉冲响应完成
链接:http://arxiv.org/pdf/2402.00859v1
作者:Jackie Lin,Georg Götz,Sebastian J. Schlecht
备注:The following article has been submitted to the EURASIP Journal on Audio, Speech, and Music Processing
摘要:在虚拟现实(VR)和视频游戏中渲染沉浸式空间音频需要快速准确地生成房间脉冲响应(RIR),以重现听觉环境。然而,用于模拟或测量长RIR的传统方法要么是计算密集型的,要么是受到低信噪比的挑战。这项研究是由直接的声音和早期反射封装足够的信息,房间的几何形状和吸收特性的见解。在此前提下,我们提出了一个新的任务,称为“RIR完成”,旨在合成后期混响的早期部分(50毫秒)的响应。为此,我们介绍了DECOR,房间脉冲响应的深度指数完成,这是一种深度神经网络,结构为自动编码器,旨在预测滤波后噪声序列的多指数衰减包络。DECOR输出的可解释性有助于其与各种渲染技术的集成。所提出的方法进行了比较,对一个适应国家的最先进的网络,和可比的性能显示出可喜的成果,支持RIR完成任务的可行性。RIR完成可以广泛适用于增强RIR生成任务,其中需要快速后期混响近似。
摘要:Rendering immersive spatial audio in virtual reality (VR) and video games demands a fast and accurate generation of room impulse responses (RIRs) to recreate auditory environments plausibly. However, the conventional methods for simulating or measuring long RIRs are either computationally intensive or challenged by low signal-to-noise ratios. This study is propelled by the insight that direct sound and early reflections encapsulate sufficient information about room geometry and absorption characteristics. Building upon this premise, we propose a novel task termed "RIR completion," aimed at synthesizing the late reverberation given only the early portion (50 ms) of the response. To this end, we introduce DECOR, Deep Exponential Completion Of Room impulse responses, a deep neural network structured as an autoencoder designed to predict multi-exponential decay envelopes of filtered noise sequences. The interpretability of DECOR's output facilitates its integration with diverse rendering techniques. The proposed method is compared against an adapted state-of-the-art network, and comparable performance shows promising results supporting the feasibility of the RIR completion task. The RIR completion can be widely adapted to enhance RIR generation tasks where fast late reverberation approximation is required.

【2】 Efficient Fine-tuning of Audio Spectrogram Transformers via Soft Mixture of Adapters
标题:通过适配器的软混合实现音频频谱转换器的高效微调
链接:http://arxiv.org/pdf/2402.00828v1
作者:Umberto Cappellazzo,Daniele Falavigna,Alessio Brutti
备注:The code will be released ad: url{this https URL}
摘要:混合专家(MoE)架构最近开始蓬勃发展,由于他们的能力,规模模型的能力,同时保持可负担的计算成本。此外,它们可以应用于Transformers和状态空间模型,目前在许多领域的最先进的模型。虽然MoE主要是针对预训练阶段进行研究,但其在参数有效的迁移学习环境中的使用尚未得到充分探索。为了缩小这一差距,本文试图揭开使用MoE的音频频谱图Transformers的音频和语音下游任务的参数有效的微调。具体来说,我们提出了软混合适配器(软MoA)。它利用适配器作为专家,并利用最近的软MoE方法,它依赖于输入令牌和专家之间的软分配,以保持计算时间有限。跨4个基准测试的广泛实验表明,软MoA优于单适配器方法,并与密集MoA对应物相当。最后,我们对Soft-MoA的关键元素进行了消融研究,例如,Soft-MoA通过更多的专家实现了更好的扩展,并确保所有专家都有助于输出令牌的计算,从而免除了专家不平衡问题。
摘要:Mixture of Experts (MoE) architectures have recently started burgeoning due to their ability to scale model's capacity while maintaining the computational cost affordable. Furthermore, they can be applied to both Transformers and State Space Models, the current state-of-the-art models in numerous fields. While MoE has been mostly investigated for the pre-training stage, its use in parameter-efficient transfer learning settings is under-explored. To narrow this gap, this paper attempts to demystify the use of MoE for parameter-efficient fine-tuning of Audio Spectrogram Transformers to audio and speech downstream tasks. Specifically, we propose Soft Mixture of Adapters (Soft-MoA). It exploits adapters as the experts and, leveraging the recent Soft MoE method, it relies on a soft assignment between the input tokens and experts to keep the computational time limited. Extensive experiments across 4 benchmarks demonstrate that Soft-MoA outperforms the single adapter method and performs on par with the dense MoA counterpart. We finally present ablation studies on key elements of Soft-MoA, showing for example that Soft-MoA achieves better scaling with more experts, as well as ensuring that all experts contribute to the computation of the output tokens, thus dispensing with the expert imbalance issue.

【3】USDnet: Unsupervised Speech Dereverberation via Neural Forward Filtering
标题:USDnet:基于神经前向滤波的无监督语音去混响
链接:http://arxiv.org/pdf/2402.00820v1
作者:Zhong-Qiu Wang
备注:in submission
摘要:在具有单个扬声器的混响条件下,每个远场麦克风在不同位置处记录相同扬声器信号的混响版本。在麦克风多于扬声器的超定条件下,每个记录的混合信号可以被用作约束以将解决方案缩小到目标无回声语音,从而减少混响。基于这一认识,我们提出了USDnet,这是一种用于无监督语音去混响(USD)的新型深度神经网络(DNN)方法。在每个训练步骤中,我们首先将输入混合输入到USDnet以产生目标语音的估计,然后对DNN估计进行线性滤波以近似多麦克风混合,以便在每个麦克风处满足约束,从而正则化DNN估计以近似目标无回声语音。可以基于混合和DNN估计经由神经前向滤波算法(诸如前向卷积预测)来估计线性滤波器。我们表明,这种新的方法可以促进单源混响语音的无监督去混响。
摘要:In reverberant conditions with a single speaker, each far-field microphone records a reverberant version of the same speaker signal at a different location. In over-determined conditions, where there are more microphones than speakers, each recorded mixture signal can be leveraged as a constraint to narrow down the solutions to target anechoic speech and thereby reduce reverberation. Equipped with this insight, we propose USDnet, a novel deep neural network (DNN) approach for unsupervised speech dereverberation (USD). At each training step, we first feed an input mixture to USDnet to produce an estimate for target speech, and then linearly filter the DNN estimate to approximate the multi-microphone mixture so that the constraint can be satisfied at each microphone, thereby regularizing the DNN estimate to approximate target anechoic speech. The linear filter can be estimated based on the mixture and DNN estimate via neural forward filtering algorithms such as forward convolutive prediction. We show that this novel methodology can promote unsupervised dereverberation of single-source reverberant speech.

【4】 An Analysis of the Variance of Diffusion-based Speech Enhancement
标题:基于扩散的语音增强的方差分析
链接:http://arxiv.org/pdf/2402.00811v1
作者:Bunlong Lay,Timo Gerkmann
备注:5 pages, 3 figures, 1 table
摘要:扩散模型被证明是生成式语音增强的强大模型。在最近的SGMSE+方法中,训练涉及扩散过程的随机微分方程,逐渐将高斯和环境噪声添加到干净的语音信号中。语音增强性能的变化取决于随机微分方程的选择,该随机微分方程在添加环境和高斯噪声时控制沿扩散过程的均值和方差的演变。在这项工作中,我们强调,方差的规模是语音增强性能的主要参数,并表明它控制噪声衰减和语音失真之间的权衡。更具体地说,我们表明,更大的方差增加了噪声衰减,并允许减少计算足迹,因为需要更少的函数评估来生成估计。
摘要:Diffusion models proved to be powerful models for generative speech enhancement. In recent SGMSE+ approaches, training involves a stochastic differential equation for the diffusion process, adding both Gaussian and environmental noise to the clean speech signal gradually. The speech enhancement performance varies depending on the choice of the stochastic differential equation that controls the evolution of the mean and the variance along the diffusion processes when adding environmental and Gaussian noise. In this work, we highlight that the scale of the variance is a dominant parameter for speech enhancement performance and show that it controls the tradeoff between noise attenuation and speech distortions. More concretely, we show that a larger variance increases the noise attenuation and allows for reducing the computational footprint, as fewer function evaluations for generating the estimate are required.

【5】 Real-time Stereo Speech Enhancement with Spatial-Cue Preservation based on Dual-Path Structure
标题:基于双路径结构的保留空间线索的实时立体语音增强
链接:http://arxiv.org/pdf/2402.00337v1
作者:Masahito Togami,Jean-Marc Valin,Karim Helwani,Ritwik Giri,Umut Isik,Michael M. Goodwin
备注:Accepted for ICASSP 2024, 5 pages
摘要:我们介绍了一个实时的,多通道语音增强算法,保持立体声录音,包括两个语音源的空间线索。认识到每个源具有独特的空间信息,我们的方法利用双路径结构,通过应用特定于源的公共频带增益,确保空间线索在增强过程中不受影响。该方法还无缝集成了预训练的单声道语音增强,无需对立体声输入进行再训练。从立体声混合源分离是通过空间波束形成,与每个源的导向矢量被自适应更新使用后增强输出信号。这确保了对空间信息的准确跟踪。通过合并增强源的空间图像来导出最终的立体声输出,其功效不严重依赖于波束形成的分离性能。该算法在10 ms帧上实时运行,具有40 ms的前瞻性。评估表明,它的有效性,在充分和稀疏重叠的混合增强语音和保留空间线索。
摘要:We introduce a real-time, multichannel speech enhancement algorithm which maintains the spatial cues of stereo recordings including two speech sources. Recognizing that each source has unique spatial information, our method utilizes a dual-path structure, ensuring the spatial cues remain unaffected during enhancement by applying source-specific common-band gain. This method also seamlessly integrates pretrained monaural speech enhancement, eliminating the need for retraining on stereo inputs. Source separation from stereo mixtures is achieved via spatial beamforming, with the steering vector for each source being adaptively updated using post-enhancement output signal. This ensures accurate tracking of the spatial information. The final stereo output is derived by merging the spatial images of the enhanced sources, with its efficacy not heavily reliant on the separation performance of the beamforming. The algorithm runs in real-time on 10-ms frames with a 40 ms of look-ahead. Evaluations reveal its effectiveness in enhancing speech and preserving spatial cues in both fully and sparsely overlapped mixtures.

【6】 Frame-Wise Breath Detection with Self-Training: An Exploration of Enhancing Breath Naturalness in Text-to-Speech
标题:基于自训练的帧级呼吸检测:增强文语转换中呼吸自然度的探索
链接:http://arxiv.org/pdf/2402.00288v1
作者:Dong Yang,Tomoki Koriyama,Yuki Saito
摘要:开发可以合成自然呼吸的文本到语音(TTS)系统对于类人语音代理是必不可少的,但需要在训练数据中对呼吸位置进行大量手动注释。为此,我们提出了一个自我训练的方法来训练呼吸检测模型,可以自动检测呼吸的位置在讲话中。我们的方法使用大型语音语料库训练模型,包括:1)利用基于规则的方法对有限的呼吸声进行注释,以及2)通过基于模型预测的伪标记对这些注释进行迭代增强。我们的检测模型采用具有下采样 上采样层的Conformer块,从而实现准确的逐帧呼吸检测。我们研究其有效性,在多扬声器TTS使用文本转录检测呼吸标记。结果表明,使用我们提出的模型的呼吸检测和呼吸标记插入合成呼吸包含的语音比基线模型更自然。
摘要:Developing Text-to-Speech (TTS) systems that can synthesize natural breath is essential for human-like voice agents but requires extensive manual annotation of breath positions in training data. To this end, we propose a self-training method for training a breath detection model that can automatically detect breath positions in speech. Our method trains the model using a large speech corpus and involves: 1) annotation of limited breath sounds utilizing a rule-based approach, and 2) iterative augmentation of these annotations through pseudo-labeling based on the model's predictions. Our detection model employs Conformer blocks with down- up-sampling layers, enabling accurate frame-wise breath detection. We investigate its effectiveness in multi-speaker TTS using text transcripts with detected breath marks. The results indicate that using our proposed model for breath detection and breath mark insertion synthesizes breath-contained speech more naturally than a baseline model.

【7】 PAM: Prompting Audio-Language Models for Audio Quality Assessment
标题:PAM:促进音频质量评估的音频语言模型
链接:http://arxiv.org/pdf/2402.00282v1
作者:Soham Deshmukh,Dareen Alharthi,Benjamin Elizalde,Hannes Gamper,Mahmoud Al Ismail,Rita Singh,Bhiksha Raj,Huaming Wang
摘要:虽然音频质量是各种音频处理任务(包括生成建模)的关键性能指标,但其客观测量仍然是一个挑战。音频语言模型(ALM)是在音频文本对上进行预训练的,这些音频文本对可能包含有关音频质量、伪像存在或噪声的信息。给定音频输入和与质量相关的文本提示,ALM可以用于计算两者之间的相似性分数。在这里,我们利用这种能力,并介绍PAM,一个无参考指标,用于评估不同的音频处理任务的音频质量。与其他“无参考”指标相反,PAM不需要在参考数据集上计算嵌入,也不需要在一组昂贵的人类听力分数上训练特定于任务的模型。我们广泛评估了PAM在四个任务上的可靠性:文本到音频(TTA),文本到音乐生成(TTM),文本到语音(TTS)和深度噪声抑制(DNS)。我们进行多个消融研究与控制失真,在野外设置,并提示选择。我们的评估表明,PAM与现有的指标和人类听力分数相关。这些结果证明了ALM用于计算通用音频质量度量的潜力。
摘要:While audio quality is a key performance metric for various audio processing tasks, including generative modeling, its objective measurement remains a challenge. Audio-Language Models (ALMs) are pre-trained on audio-text pairs that may contain information about audio quality, the presence of artifacts, or noise. Given an audio input and a text prompt related to quality, an ALM can be used to calculate a similarity score between the two. Here, we exploit this capability and introduce PAM, a no-reference metric for assessing audio quality for different audio processing tasks. Contrary to other "reference-free" metrics, PAM does not require computing embeddings on a reference dataset nor training a task-specific model on a costly set of human listening scores. We extensively evaluate the reliability of PAM against established metrics and human listening scores on four tasks: text-to-audio (TTA), text-to-music generation (TTM), text-to-speech (TTS), and deep noise suppression (DNS). We perform multiple ablation studies with controlled distortions, in-the-wild setups, and prompt choices. Our evaluation shows that PAM correlates well with existing metrics and human listening scores. These results demonstrate the potential of ALMs for computing a general-purpose audio quality metric.

【8】Online speaker diarization of meetings guided by speech separation
标题:由语音分离指导的会议的在线发言者二分化
链接:http://arxiv.org/pdf/2402.00067v1
作者:Elio Gruttadauria,Mathieu Fontaine,Slim Essid
Journal-ref:IEEE International Conference on Acoustics, Speech, and Signal Processing, Apr 2024, Seoul (Korea), South Korea
摘要:重叠语音对于说话人日记系统是出了名的问题。因此,最近提出了使用语音分离来改善它们的性能。虽然有前途,语音分离模型与现实数据的斗争,因为它们是在具有固定数量的扬声器的模拟混合物上训练的。在这项工作中,我们引入了一个新的语音分离指导的日记计划适合于在线扬声器日记的长会议录音与可变数量的扬声器,目前在AMI语料库。我们设想ConvTasNet和DPRNN作为分离网络的替代方案,具有两个或三个输出源。为了获得说话人日志化结果,对每个估计的源应用语音活动检测。最终的模型是端到端微调后,首先适应分离的实际数据使用AMI。该系统在短段上运行,并且通过使用说话人嵌入和增量聚类拼接本地预测来执行推理。结果表明,我们的系统改进了AMI耳机混合的最新技术,不使用任何oracle信息,并在充分评估(无衣领,包括重叠的语音)。最后,我们展示了我们的系统的强度,特别是在重叠的语音部分。
摘要:Overlapped speech is notoriously problematic for speaker diarization systems. Consequently, the use of speech separation has recently been proposed to improve their performance. Although promising, speech separation models struggle with realistic data because they are trained on simulated mixtures with a fixed number of speakers. In this work, we introduce a new speech separation-guided diarization scheme suitable for the online speaker diarization of long meeting recordings with a variable number of speakers, as present in the AMI corpus. We envisage ConvTasNet and DPRNN as alternatives for the separation networks, with two or three output sources. To obtain the speaker diarization result, voice activity detection is applied on each estimated source. The final model is fine-tuned end-to-end, after first adapting the separation to real data using AMI. The system operates on short segments, and inference is performed by stitching the local predictions using speaker embeddings and incremental clustering. The results show that our system improves the state-of-the-art on the AMI headset mix, using no oracle information and under full evaluation (no collar and including overlapped speech). Finally, we show the strength of our system particularly on overlapped speech sections.


【9】BATON: Aligning Text-to-Audio Model with Human Preference Feedback

标题:指挥棒:使文本到音频模型与人类偏好反馈保持一致

链接:http://arxiv.org/pdf/2402.00744v1

作者:Huan Liao,Haonan Han,Kai Yang,Tianjiao Du,Rui Yang,Zunnan Xu,Qinmei Xu,Jingquan Liu,Jiasheng Lu,Xiu Li

摘要:随着人工智能生成内容(AIGC)的发展,文本到音频模型受到广泛关注。然而,由于自然语言固有的信息密度和有限的模型理解能力,这些模型生成符合人类偏好的音频是具有挑战性的。为了缓解这个问题,我们制定了指挥棒,一个框架,旨在提高生成的音频和文本提示使用人类偏好反馈之间的对齐。我们的BATON包括三个关键阶段:首先,我们策划了一个包含提示和相应生成的音频的数据集,然后根据人类反馈对其进行注释。其次,我们引入了一个奖励模型,使用构造的数据集,它可以模仿人类的偏好,通过分配奖励输入文本音频对。最后,我们采用奖励模型来微调现成的文本到音频模型。实验结果表明,我们的BATON可以显着提高原始文本到音频模型的生成质量,涉及音频的完整性,时间关系,并与人类的偏好对齐。

摘要:With the development of AI-Generated Content (AIGC), text-to-audio models are gaining widespread attention. However, it is challenging for these models to generate audio aligned with human preference due to the inherent information density of natural language and limited model understanding ability. To alleviate this issue, we formulate the BATON, a framework designed to enhance the alignment between generated audio and text prompt using human preference feedback. Our BATON comprises three key stages: Firstly, we curated a dataset containing both prompts and the corresponding generated audio, which was then annotated based on human feedback. Secondly, we introduced a reward model using the constructed dataset, which can mimic human preference by assigning rewards to input text-audio pairs. Finally, we employed the reward model to fine-tune an off-the-shelf text-to-audio model. The experiment results demonstrate that our BATON can significantly improve the generation quality of the original text-to-audio models, concerning audio integrity, temporal relationship, and alignment with human preference.

【10】 Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
标题:你能去掉具有自监督语音特征的说话人识别的下游模型吗?
链接:http://arxiv.org/pdf/2402.00340v1
作者:Zakaria Aldeneh,Takuya Higuchi,Jee-weon Jung,Skyler Seto,Tatiana Likhomanenko,Stephen Shum,Ahmed Hussen Abdelaziz,Shinji Watanabe,Barry-John Theobald
摘要:在说话人确认模型中,自监督特征通常用于代替滤波器组。然而,这些模型最初被设计为摄取滤波器组作为输入,因此,在自监督特征之上训练它们假设两种特征类型需要相同的任务学习量。在这项工作中,我们观察到预训练的自监督语音特征固有地包括下游说话人验证任务所需的信息,因此,我们可以在不牺牲性能的情况下简化下游模型。为此,我们重新审视了使用自监督特征的说话人验证下游模型的设计。我们表明,我们可以简化模型,使用少97.51%的参数,同时实现29.93%的平均性能提高SUPERB。因此,我们证明了简化的下游模型比基线模型更有效-它仅用60%的训练数据就实现了更好的性能。
摘要:Self-supervised features are typically used in place of filter-banks in speaker verification models. However, these models were originally designed to ingest filter-banks as inputs, and thus, training them on top of self-supervised features assumes that both feature types require the same amount of learning for the task. In this work, we observe that pre-trained self-supervised speech features inherently include information required for downstream speaker verification task, and therefore, we can simplify the downstream model without sacrificing performance. To this end, we revisit the design of the downstream model for speaker verification using self-supervised features. We show that we can simplify the model to use 97.51% fewer parameters while achieving a 29.93% average improvement in performance on SUPERB. Consequently, we show that the simplified downstream model is more data efficient compared to baseline--it achieves better performance with only 60% of the training data.

【11】 Exploring the limits of decoder-only models trained on public speech recognition corpora
标题:探索在公共语音识别语料库上训练的仅限解码器的模型的局限性
链接:http://arxiv.org/pdf/2402.00235v1
作者:Ankit Gupta,George Saon,Brian Kingsbury
摘要:工业规模语音识别(ASR)模型的出现,如Whisper和USM,分别在100万小时的弱标记和1200万小时的仅音频专有数据上进行训练,导致对大规模公共ASR语料库和竞争性开源管道的需求更加强烈。与上述模型不同,大型语言模型通常基于Transformer解码器,目前还不清楚仅基于公共数据训练的解码器模型是否可以提供有竞争力的性能。在这项工作中,我们调查的因素,如选择的训练数据集和建模组件,以获得最佳的性能,单独使用公共英语ASR语料库。我们的Decoder-Only Transformer for ASR(DOTA)模型在几乎所有英语ASR基准测试中的表现都全面优于Whisper(OWSM)的编码器-解码器开源复制,并在15个测试集中的7个测试中优于Whisper large-v3。我们在许可证下发布我们的代码库和模型检查点。
摘要:The emergence of industrial-scale speech recognition (ASR) models such as Whisper and USM, trained on 1M hours of weakly labelled and 12M hours of audio only proprietary data respectively, has led to a stronger need for large scale public ASR corpora and competitive open source pipelines. Unlike the said models, large language models are typically based on Transformer decoders, and it remains unclear if decoder-only models trained on public data alone can deliver competitive performance. In this work, we investigate factors such as choice of training datasets and modeling components necessary for obtaining the best performance using public English ASR corpora alone. Our Decoder-Only Transformer for ASR (DOTA) model comprehensively outperforms the encoder-decoder open source replication of Whisper (OWSM) on nearly all English ASR benchmarks and outperforms Whisper large-v3 on 7 out of 15 test sets. We release our codebase and model checkpoints under permissive license.

机器翻译由腾讯交互翻译提供,仅供参考