微信公众号:arXiv_Daily
cs.SD语音
【1】The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
标题:ICASP 2026 HumDial挑战:LLM时代的类人口语对话系统基准
链接:https://arxiv.org/abs/2601.05564
备注:Official summary paper for the ICASSP 2026 HumDial Challenge
摘要:在大型语言模型(LLM),特别是Audio-LLM和Omni-models的快速发展的推动下,口语对话系统已经发生了显着的变化,逐步缩小了人机交互和人机交互之间的差距。实现真正的“类人”沟通需要双重能力:感知用户情绪状态并与之产生共鸣的情商,以及导航动态、自然对话流的强大交互机制,如实时话轮转换。因此,我们在ICASSP 2026上发起了第一个类人语音对话系统挑战赛(HumDial),以衡量这些双重功能。该计划以来自真实人类对话的大量数据集为基础,在两个方面建立了一个公平的评估平台:(1)情商,旨在长期的情感理解和移情生成;(2)全双工互动,系统地评估“边说边说”条件下的实时决策。本文总结了数据集,轨道配置和最终结果。
摘要:Driven by the rapid advancement of Large Language Models (LLMs), particularly Audio-LLMs and Omni-models, spoken dialogue systems have evolved significantly, progressively narrowing the gap between human-machine and human-human interactions. Achieving truly ``human-like'' communication necessitates a dual capability: emotional intelligence to perceive and resonate with users' emotional states, and robust interaction mechanisms to navigate the dynamic, natural flow of conversation, such as real-time turn-taking. Therefore, we launched the first Human-like Spoken Dialogue Systems Challenge (HumDial) at ICASSP 2026 to benchmark these dual capabilities. Anchored by a sizable dataset derived from authentic human conversations, this initiative establishes a fair evaluation platform across two tracks: (1) Emotional Intelligence, targeting long-term emotion understanding and empathetic generation; and (2) Full-Duplex Interaction, systematically evaluating real-time decision-making under `` listening-while-speaking'' conditions. This paper summarizes the dataset, track configurations, and the final results.
【2】SPAM: Style Prompt Adherence Metric for Prompt-based TTS
标题:SPAM:基于预算的TTC的风格提示遵守性指标
链接:https://arxiv.org/abs/2601.05554
摘要:基于文本的文语转换(TTS)旨在生成符合文本提示中提供的细粒度风格提示的语音。然而,大多数先前的作品依赖于既不可信也不忠实的措施来评估及时遵守。也就是说,他们无法确保评估是否基于提示,是否与人类相似。因此,我们提出了一个新的自动度量,风格提示遵守度量,它明确地满足可扩展性和忠实性。受CLAP的启发,我们的方法将语音分解为声学属性,并将它们与风格提示对齐。此外,我们用监督对比损失训练评分器,这可以更清楚地区分不同的语义。我们从两个角度进行了两个实验。可信度实验表明,垃圾邮件与平均意见得分(MOS)达到了很强的相关性。此外,忠实性实验表明,垃圾邮件是成功地接地到给定的风格提示,因为它可以区分不同的语义提示。我们相信,SPAM可以提供一个可行的自动化的解决方案,评估风格即时遵守的合成语音。
摘要:Prompt-based text-to-speech (TTS) aims to generate speech that adheres to fine-grained style cues provided in a text prompt. However, most prior works depend on neither plausible nor faithful measures to evaluate prompt adherence. That is, they cannot ensure whether the evaluation is grounded on the prompt and is similar to a human. Thus, we present a new automatic metric, the Style Prompt Adherence Metric, which explicitly satisfies both plausibility and faithfulness. Inspired by the CLAP, our approach factorizes speech into acoustic attributes and aligns them with the style prompt. Also, we trained the scorer with a supervised contrastive loss, which could provide a clearer distinction between different semantics. We conducted two experiments on two perspectives. The plausibility experiment showed that SPAM achieved a strong correlation with the mean opinion score (MOS). Also, the faithfulness experiment demonstrated that SPAM is successfully grounded to the given style prompt, as it can discriminate different semantics of the prompt. We believe that SPAM can provide a viable automatic solution for evaluating style prompt adherence of synthesized speech.
【3】Closing the Modality Reasoning Gap for Speech Large Language Models
标题:缩小语音大型语言模型的情态推理差距
链接:https://arxiv.org/abs/2601.05543
摘要:虽然语音大语言模型已经取得了显着的进展,一个实质性的模态推理差距仍然存在:语音输入的推理性能明显弱于文本。这种差距可能与跨Transformer层的代表性漂移和长链推理中的行为偏差有关。为了解决这个问题,我们引入了TARS,这是一个学习框架,它通过不对称的奖励设计来调整文本条件和语音条件的轨迹。该框架采用了两个密集和互补的信号:表示对齐,它测量语音和文本条件轨迹之间的逐层隐藏状态相似性,和行为对齐,它评估生成的输出和参考文本完成之间的语义一致性。具有挑战性的推理基准,包括MMSU和OBQA的实验表明,我们的方法显着缩小了模态推理差距,并实现了最先进的性能之间的7B规模的语音LLM。
摘要:Although speech large language models have achieved notable progress, a substantial modality reasoning gap remains: their reasoning performance on speech inputs is markedly weaker than on text. This gap could be associated with representational drift across Transformer layers and behavior deviations in long-chain reasoning. To address this issue, we introduce TARS, a reinforcement-learning framework that aligns text-conditioned and speech-conditioned trajectories through an asymmetric reward design. The framework employs two dense and complementary signals: representation alignment, which measures layer-wise hidden-state similarity between speech- and text-conditioned trajectories, and behavior alignment, which evaluates semantic consistency between generated outputs and reference text completions. Experiments on challenging reasoning benchmarks, including MMSU and OBQA, show that our approach significantly narrows the modality reasoning gap and achieves state-of-the-art performance among 7B-scale Speech LLMs.
【4】CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
标题:CosyEdit:从Zero-Shot文本到语音模型中解锁端到端语音编辑能力
链接:https://arxiv.org/abs/2601.05329
摘要:自动语音编辑旨在基于文本指令修改口语内容,然而传统的级联系统遭受复杂的预处理流水线和对显式外部时间对齐的依赖。针对这些限制,我们提出了CosyEdit,这是一种通过特定于任务的微调和优化的推理程序改编自CosyVoice的端到端语音编辑模型,它将语音文本对齐内在化,同时确保编辑前后语音之间的高度一致性。通过对来自我们精心策划的GigaEdit数据集的250小时监督数据进行微调,我们的400 M参数模型实现了可靠的语音编辑性能。在RealEdit基准测试上的实验表明,CosyEdit不仅优于数十亿参数的语言模型基线,而且与最先进的级联方法的性能相当。这些结果表明,通过特定于任务的微调和推理优化,可以从zero-shot TTS模型中解锁鲁棒且高效的语音编辑功能,从而产生用于高质量语音编辑的新颖且具有成本效益的端到端解决方案。
摘要:Automatic speech editing aims to modify spoken content based on textual instructions, yet traditional cascade systems suffer from complex preprocessing pipelines and a reliance on explicit external temporal alignment. Addressing these limitations, we propose CosyEdit, an end-to-end speech editing model adapted from CosyVoice through task-specific fine-tuning and an optimized inference procedure, which internalizes speech-text alignment while ensuring high consistency between the speech before and after editing. By fine-tuning on only 250 hours of supervised data from our curated GigaEdit dataset, our 400M-parameter model achieves reliable speech editing performance. Experiments on the RealEdit benchmark indicate that CosyEdit not only outperforms several billion-parameter language model baselines but also matches the performance of state-of-the-art cascade approaches. These results demonstrate that, with task-specific fine-tuning and inference optimization, robust and efficient speech editing capabilities can be unlocked from a zero-shot TTS model, yielding a novel and cost-effective end-to-end solution for high-quality speech editing.
【5】Discriminative-Generative Target Speaker Extraction with Decoder-Only Language Models
标题:使用仅解码器语言模型的区分生成目标说话人提取
链接:https://arxiv.org/abs/2601.06006
备注:16 pages,6 figures
摘要:目标说话人提取(TSE)的目的是从混合音频记录中恢复出所需说话人的语音信号,给定一个短的注册话语。大多数现有的TSE方法是基于判别建模范式。尽管这些方法在抑制干扰说话者方面是有效的,但是这些方法通常难以产生具有高感知质量和自然度的语音。为了解决这个问题,我们首先提出了LauraTSE,一个生成TSE模型建立在自回归解码器的语言模型。然而,纯粹的生成方法可能会受到幻觉,内容漂移和有限的可控性的影响,这可能会破坏它们在复杂声学场景中的可靠性。为了克服这些挑战,我们进一步引入了一个判别生成TSE框架。在这个框架中,一个歧视性的前端鲁棒地提取目标说话人的语音,产生稳定和可控的中间表示。然后,生成后端在神经音频编解码器表示空间中操作,以重建细粒度的语音细节并增强感知质量。这种两阶段设计有效地结合了判别模型的鲁棒性和可控性与生成模型的卓越自然性和质量增强能力。此外,我们系统地研究了所提出的框架的协作训练策略,包括冻结或微调前端,纳入辅助SI-SDR损失,并探索自回归和非自回归推理机制。实验结果表明,该框架实现了更有利的权衡之间的语音质量,可懂度和说话人的一致性。
摘要:Target speaker extraction (TSE) aims to recover the speech signal of a desired speaker from a mixed audio recording, given a short enrollment utterance. Most existing TSE approaches are based on discriminative modeling paradigms. Although effective at suppressing interfering speakers, these methods often struggle to produce speech with high perceptual quality and naturalness. To address this limitation, we first propose LauraTSE, a generative TSE model built upon an auto-regressive decoder-only language model. However, purely generative approaches may suffer from hallucinations, content drift, and limited controllability, which may undermine their reliability in complex acoustic scenarios. To overcome these challenges, we further introduce a discriminative-generative TSE framework. In this framework, a discriminative front-end is employed to robustly extract the target speaker's speech, yielding stable and controllable intermediate representations. A generative back-end then operates in the neural audio codec representation space to reconstruct fine-grained speech details and enhance perceptual quality. This two-stage design effectively combines the robustness and controllability of discriminative models with the superior naturalness and quality enhancement capabilities of generative models. Moreover, we systematically investigate collaborative training strategies for the proposed framework, including freezing or fine-tuning the front-end, incorporating an auxiliary SI-SDR loss, and exploring both auto-regressive and non-auto-regressive inference mechanisms. Experimental results demonstrate that the proposed framework achieves a more favorable trade-off among speech quality, intelligibility, and speaker consistency.
【1】Discriminative-Generative Target Speaker Extraction with Decoder-Only Language Models
标题:使用仅解码器语言模型的区分生成目标说话人提取
链接:https://arxiv.org/abs/2601.06006
备注:16 pages,6 figures
摘要:目标说话人提取(TSE)的目的是从混合音频记录中恢复出所需说话人的语音信号,给定一个短的注册话语。大多数现有的TSE方法是基于判别建模范式。尽管这些方法在抑制干扰说话者方面是有效的,但是这些方法通常难以产生具有高感知质量和自然度的语音。为了解决这个问题,我们首先提出了LauraTSE,一个生成TSE模型建立在自回归解码器的语言模型。然而,纯粹的生成方法可能会受到幻觉,内容漂移和有限的可控性的影响,这可能会破坏它们在复杂声学场景中的可靠性。为了克服这些挑战,我们进一步引入了一个判别生成TSE框架。在这个框架中,一个歧视性的前端鲁棒地提取目标说话人的语音,产生稳定和可控的中间表示。然后,生成后端在神经音频编解码器表示空间中操作,以重建细粒度的语音细节并增强感知质量。这种两阶段设计有效地结合了判别模型的鲁棒性和可控性与生成模型的卓越自然性和质量增强能力。此外,我们系统地研究了所提出的框架的协作训练策略,包括冻结或微调前端,纳入辅助SI-SDR损失,并探索自回归和非自回归推理机制。实验结果表明,该框架实现了更有利的权衡之间的语音质量,可懂度和说话人的一致性。
摘要:Target speaker extraction (TSE) aims to recover the speech signal of a desired speaker from a mixed audio recording, given a short enrollment utterance. Most existing TSE approaches are based on discriminative modeling paradigms. Although effective at suppressing interfering speakers, these methods often struggle to produce speech with high perceptual quality and naturalness. To address this limitation, we first propose LauraTSE, a generative TSE model built upon an auto-regressive decoder-only language model. However, purely generative approaches may suffer from hallucinations, content drift, and limited controllability, which may undermine their reliability in complex acoustic scenarios. To overcome these challenges, we further introduce a discriminative-generative TSE framework. In this framework, a discriminative front-end is employed to robustly extract the target speaker's speech, yielding stable and controllable intermediate representations. A generative back-end then operates in the neural audio codec representation space to reconstruct fine-grained speech details and enhance perceptual quality. This two-stage design effectively combines the robustness and controllability of discriminative models with the superior naturalness and quality enhancement capabilities of generative models. Moreover, we systematically investigate collaborative training strategies for the proposed framework, including freezing or fine-tuning the front-end, incorporating an auxiliary SI-SDR loss, and exploring both auto-regressive and non-auto-regressive inference mechanisms. Experimental results demonstrate that the proposed framework achieves a more favorable trade-off among speech quality, intelligibility, and speaker consistency.
【2】The ICASSP 2026 HumDial Challenge: Benchmarking Human-like Spoken Dialogue Systems in the LLM Era
标题:ICASP 2026 HumDial挑战:LLM时代的类人口语对话系统基准
链接:https://arxiv.org/abs/2601.05564
备注:Official summary paper for the ICASSP 2026 HumDial Challenge
摘要:在大型语言模型(LLM),特别是Audio-LLM和Omni-models的快速发展的推动下,口语对话系统已经发生了显着的变化,逐步缩小了人机交互和人机交互之间的差距。实现真正的“类人”沟通需要双重能力:感知用户情绪状态并与之产生共鸣的情商,以及导航动态、自然对话流的强大交互机制,如实时话轮转换。因此,我们在ICASSP 2026上发起了第一个类人语音对话系统挑战赛(HumDial),以衡量这些双重功能。该计划以来自真实人类对话的大量数据集为基础,在两个方面建立了一个公平的评估平台:(1)情商,旨在长期的情感理解和移情生成;(2)全双工互动,系统地评估“边说边说”条件下的实时决策。本文总结了数据集,轨道配置和最终结果。
摘要:Driven by the rapid advancement of Large Language Models (LLMs), particularly Audio-LLMs and Omni-models, spoken dialogue systems have evolved significantly, progressively narrowing the gap between human-machine and human-human interactions. Achieving truly ``human-like'' communication necessitates a dual capability: emotional intelligence to perceive and resonate with users' emotional states, and robust interaction mechanisms to navigate the dynamic, natural flow of conversation, such as real-time turn-taking. Therefore, we launched the first Human-like Spoken Dialogue Systems Challenge (HumDial) at ICASSP 2026 to benchmark these dual capabilities. Anchored by a sizable dataset derived from authentic human conversations, this initiative establishes a fair evaluation platform across two tracks: (1) Emotional Intelligence, targeting long-term emotion understanding and empathetic generation; and (2) Full-Duplex Interaction, systematically evaluating real-time decision-making under `` listening-while-speaking'' conditions. This paper summarizes the dataset, track configurations, and the final results.
【3】SPAM: Style Prompt Adherence Metric for Prompt-based TTS
标题:SPAM:基于预算的TTC的风格提示遵守性指标
链接:https://arxiv.org/abs/2601.05554
摘要:基于文本的文语转换(TTS)旨在生成符合文本提示中提供的细粒度风格提示的语音。然而,大多数先前的作品依赖于既不可信也不忠实的措施来评估及时遵守。也就是说,他们无法确保评估是否基于提示,是否与人类相似。因此,我们提出了一个新的自动度量,风格提示遵守度量,它明确地满足可扩展性和忠实性。受CLAP的启发,我们的方法将语音分解为声学属性,并将它们与样式提示对齐。此外,我们用监督对比损失训练评分器,这可以更清楚地区分不同的语义。我们从两个角度进行了两个实验。可信度实验表明,垃圾邮件与平均意见得分(MOS)达到了很强的相关性。此外,忠实性实验表明,垃圾邮件是成功地接地到给定的风格提示,因为它可以区分不同的语义提示。我们相信,SPAM可以提供一个可行的自动化的解决方案,评估风格即时遵守的合成语音。
摘要:Prompt-based text-to-speech (TTS) aims to generate speech that adheres to fine-grained style cues provided in a text prompt. However, most prior works depend on neither plausible nor faithful measures to evaluate prompt adherence. That is, they cannot ensure whether the evaluation is grounded on the prompt and is similar to a human. Thus, we present a new automatic metric, the Style Prompt Adherence Metric, which explicitly satisfies both plausibility and faithfulness. Inspired by the CLAP, our approach factorizes speech into acoustic attributes and aligns them with the style prompt. Also, we trained the scorer with a supervised contrastive loss, which could provide a clearer distinction between different semantics. We conducted two experiments on two perspectives. The plausibility experiment showed that SPAM achieved a strong correlation with the mean opinion score (MOS). Also, the faithfulness experiment demonstrated that SPAM is successfully grounded to the given style prompt, as it can discriminate different semantics of the prompt. We believe that SPAM can provide a viable automatic solution for evaluating style prompt adherence of synthesized speech.
【4】Closing the Modality Reasoning Gap for Speech Large Language Models
标题:缩小语音大型语言模型的情态推理差距
链接:https://arxiv.org/abs/2601.05543
摘要:虽然语音大语言模型已经取得了显着的进展,一个实质性的模态推理差距仍然存在:语音输入的推理性能明显弱于文本。这种差距可能与跨Transformer层的代表性漂移和长链推理中的行为偏差有关。为了解决这个问题,我们引入了TARS,这是一个学习框架,它通过不对称的奖励设计来调整文本条件和语音条件的轨迹。该框架采用了两个密集和互补的信号:表示对齐,它测量语音和文本条件轨迹之间的逐层隐藏状态相似性,和行为对齐,它评估生成的输出和参考文本完成之间的语义一致性。具有挑战性的推理基准,包括MMSU和OBQA的实验表明,我们的方法显着缩小了模态推理差距,并实现了最先进的性能之间的7B规模的语音LLM。
摘要:Although speech large language models have achieved notable progress, a substantial modality reasoning gap remains: their reasoning performance on speech inputs is markedly weaker than on text. This gap could be associated with representational drift across Transformer layers and behavior deviations in long-chain reasoning. To address this issue, we introduce TARS, a reinforcement-learning framework that aligns text-conditioned and speech-conditioned trajectories through an asymmetric reward design. The framework employs two dense and complementary signals: representation alignment, which measures layer-wise hidden-state similarity between speech- and text-conditioned trajectories, and behavior alignment, which evaluates semantic consistency between generated outputs and reference text completions. Experiments on challenging reasoning benchmarks, including MMSU and OBQA, show that our approach significantly narrows the modality reasoning gap and achieves state-of-the-art performance among 7B-scale Speech LLMs.
【5】CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
标题:CosyEdit:从Zero-Shot文本到语音模型中解锁端到端语音编辑能力
链接:https://arxiv.org/abs/2601.05329
摘要:自动语音编辑旨在基于文本指令修改口语内容,然而传统的级联系统遭受复杂的预处理流水线和对显式外部时间对齐的依赖。针对这些限制,我们提出了CosyEdit,这是一种通过特定于任务的微调和优化的推理程序改编自CosyVoice的端到端语音编辑模型,它将语音文本对齐内在化,同时确保编辑前后语音之间的高度一致性。通过对来自我们精心策划的GigaEdit数据集的250小时监督数据进行微调,我们的400 M参数模型实现了可靠的语音编辑性能。在RealEdit基准测试上的实验表明,CosyEdit不仅优于数十亿参数的语言模型基线,而且与最先进的级联方法的性能相当。这些结果表明,通过特定于任务的微调和推理优化,可以从zero-shot TTS模型中解锁鲁棒且高效的语音编辑功能,从而产生用于高质量语音编辑的新颖且具有成本效益的端到端解决方案。
摘要:Automatic speech editing aims to modify spoken content based on textual instructions, yet traditional cascade systems suffer from complex preprocessing pipelines and a reliance on explicit external temporal alignment. Addressing these limitations, we propose CosyEdit, an end-to-end speech editing model adapted from CosyVoice through task-specific fine-tuning and an optimized inference procedure, which internalizes speech-text alignment while ensuring high consistency between the speech before and after editing. By fine-tuning on only 250 hours of supervised data from our curated GigaEdit dataset, our 400M-parameter model achieves reliable speech editing performance. Experiments on the RealEdit benchmark indicate that CosyEdit not only outperforms several billion-parameter language model baselines but also matches the performance of state-of-the-art cascade approaches. These results demonstrate that, with task-specific fine-tuning and inference optimization, robust and efficient speech editing capabilities can be unlocked from a zero-shot TTS model, yielding a novel and cost-effective end-to-end solution for high-quality speech editing.
机器翻译由腾讯交互翻译提供,仅供参考
