今日论文合集:cs.SD语音5篇,eess.AS音频处理6篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】DiVISe: Direct Visual-Input Speech Synthesis Preserving Speaker  Characteristics And Intelligibility

标题:DiVISe:直接视觉输入语音合成,保留说话者特征和可理解性
链接:https://arxiv.org/abs/2503.05223
作者:Yifan Liu,  Yu Fang,  Zhouhan Lin
备注:to be published in NAACL 25
摘要:视频到语音(V2S)合成,直接从无声视频输入生成语音的任务,本质上比其他语音合成任务更具挑战性,因为需要准确地重建语音内容和扬声器特性,从视觉线索单独。最近,视听预训练已经消除了V2S中对额外声学提示的需要,而之前的方法通常依赖于此来确保训练收敛。然而,即使有预训练,现有方法在实现声学可懂度和保存说话者特定特征之间的平衡方面仍然面临挑战。我们分析了这一限制,并有动机引入DiVISe(直接视觉输入语音合成),这是一种端到端的V2S模型,可以直接从视频帧中预测Mel频谱图。尽管没有采取任何声学提示,但DiVISe有效地保留了生成的音频中的扬声器特征,并在LRS2和LRS3数据集上实现了客观和主观指标的卓越性能。我们的研究结果表明,DiVISe不仅在声学可懂度方面优于现有的V2S模型,而且随着数据和模型参数的增加,其扩展效果更好。代码和重量可以在https://github.com/PussyCat0700/DiVISe上找到。
摘要:Video-to-speech (V2S) synthesis, the task of generating speech directly fromsilent video input, is inherently more challenging than other speech synthesistasks due to the need to accurately reconstruct both speech content and speakercharacteristics from visual cues alone. Recently, audio-visual pre-training haseliminated the need for additional acoustic hints in V2S, which previousmethods often relied on to ensure training convergence. However, even withpre-training, existing methods continue to face challenges in achieving abalance between acoustic intelligibility and the preservation ofspeaker-specific characteristics. We analyzed this limitation and weremotivated to introduce DiVISe (Direct Visual-Input Speech Synthesis), anend-to-end V2S model that predicts Mel-spectrograms directly from video framesalone. Despite not taking any acoustic hints, DiVISe effectively preservesspeaker characteristics in the generated audio, and achieves superiorperformance on both objective and subjective metrics across the LRS2 and LRS3datasets. Our results demonstrate that DiVISe not only outperforms existing V2Smodels in acoustic intelligibility but also scales more effectively withincreased data and model parameters. Code and weights can be found athttps://github.com/PussyCat0700/DiVISe.

【2】 UniArray: Unified Spectral-Spatial Modeling for Array-Geometry-Agnostic  Speech Separation
标题:UniArray:用于阵列-几何-不可知语音分离的统一频谱空间建模
链接:https://arxiv.org/abs/2503.05110
作者:Weiguang Chen,  Junjie Zhang,  Jielong Yang,  Eng Siong Chng,  Xionghu Zhong
备注:5 pages, Prepirnt
摘要:阵列几何无关语音分离(AGA-SS)的目的是开发一种有效的分离方法,而不管麦克风阵列的几何形状。传统的方法依赖于无置换操作,如求和或注意机制,以捕捉空间信息。然而,这些方法通常会导致高的计算成本或破坏通道内和通道间交互期间空间信息的有效使用,从而导致次优性能。为了解决这些问题,我们提出了UniArray,一种新颖的方法,放弃了传统的交织方式。UniArray由三个关键组件组成:虚拟麦克风估计(VME)模块,特征提取和融合模块,以及分层双路径分离器。VME可确保在具有不同通道数的阵列中实现稳健的性能。特征提取和融合模块利用频谱特征提取模块和空间字典学习(SDL)模块来提取和融合频率区间级特征,从而允许分离器专注于使用融合的特征。分层双路径分离器模型沿时间和频率轴具有依赖性,同时保持计算效率。实验结果表明,UniArray优于国家的最先进的方法在SI-SDRi,WB-PESQ,NB-PESQ,和STOI在可见和不可见的阵列几何形状。
摘要:Array-geometry-agnostic speech separation (AGA-SS) aims to develop aneffective separation method regardless of the microphone array geometry.Conventional methods rely on permutation-free operations, such as summation orattention mechanisms, to capture spatial information. However, these approachesoften incur high computational costs or disrupt the effective use of spatialinformation during intra- and inter-channel interactions, leading to suboptimalperformance. To address these issues, we propose UniArray, a novel approachthat abandons the conventional interleaving manner. UniArray consists of threekey components: a virtual microphone estimation (VME) module, a featureextraction and fusion module, and a hierarchical dual-path separator. The VMEensures robust performance across arrays with varying channel numbers. Thefeature extraction and fusion module leverages a spectral feature extractionmodule and a spatial dictionary learning (SDL) module to extract and fusefrequency-bin-level features, allowing the separator to focus on using thefused features. The hierarchical dual-path separator models featuredependencies along the time and frequency axes while maintaining computationalefficiency. Experimental results show that UniArray outperformsstate-of-the-art methods in SI-SDRi, WB-PESQ, NB-PESQ, and STOI across bothseen and unseen array geometries.

【3】 S2S-Arena, Evaluating Speech2Speech Protocols on Instruction Following  with Paralinguistic Information
标题:S2 S-Arena,评估带有副语言信息的教学的Speech 2 Speech协议
链接:https://arxiv.org/abs/2503.05085
作者:Feng Jiang,  Zhiyu Lin,  Fan Bu,  Yuhao Du,  Benyou Wang,  Haizhou Li
摘要:随着大语言模型(LLM)的快速发展,语音模型引起了人们的极大关注,特别是支持语音输入和输出的speech 2speech协议的最新进展。然而,现有的基准采用基于文本的自动评估器来评估这些模型的指令跟随能力,缺乏对语音理解和生成中的非语言信息的考虑。为了解决这些问题,我们引入了S2 S-Arena,这是一种新颖的竞技场风格的S2 S基准测试,可以在真实任务中的语音输入和语音输出中使用语言信息来评估跟随能力。我们设计了154个样本,融合了TTS和现场录音在四个领域的21个任务,并手动评估现有的流行的语音模型在舞台风格的方式。实验结果表明:(1)在Speech 2Speech协议中,ASR、LLM和TTS级联的语音模型除了具有GPT-4 o的优越性能外,在文语对齐后,其性能优于联合训练的模型:(2)考虑到语言信息,语音模型的知识性主要依赖于LLM骨干,其多语言支持受到语音模块的限制;(3)优秀的语音模型已经能够理解语音输入中的非语言信息,但是生成具有非语言信息的音频仍然是一个挑战。
摘要:The rapid development of large language models (LLMs) has brought significantattention to speech models, particularly recent progress in speech2speechprotocols supporting speech input and output. However, the existing benchmarksadopt automatic text-based evaluators for evaluating the instruction followingability of these models lack consideration for paralinguistic information inboth speech understanding and generation. To address these issues, we introduceS2S-Arena, a novel arena-style S2S benchmark that evaluatesinstruction-following capabilities with paralinguistic information in bothspeech-in and speech-out across real-world tasks. We design 154 samples thatfused TTS and live recordings in four domains with 21 tasks and manuallyevaluate existing popular speech models in an arena-style manner. Theexperimental results show that: (1) in addition to the superior performance ofGPT-4o, the speech model of cascaded ASR, LLM, and TTS outperforms the jointlytrained model after text-speech alignment in speech2speech protocols; (2)considering paralinguistic information, the knowledgeability of the speechmodel mainly depends on the LLM backbone, and the multilingual support of thatis limited by the speech module; (3) excellent speech models can alreadyunderstand the paralinguistic information in speech input, but generatingappropriate audio with paralinguistic information is still a challenge.

【4】 Direct Speech to Speech Translation: A Review
标题:直接语音到语音翻译:评论
链接:https://arxiv.org/abs/2503.04799
作者:Mohammad Sarim,  Saim Shakeel,  Laeeba Javed,  Jamaluddin,  Mohammad Nadeem
摘要:语音到语音翻译(S2ST)是一种变革性的技术,可以弥合全球通信差距,实现外交,旅游和国际贸易中的实时多语言互动。我们的评论研究了S2ST的演变,比较了传统的级联模型,依赖于自动语音识别(ASR),机器翻译(MT)和文本到语音(TTS)组件与新的端到端和直接语音翻译(DST)模型,绕过中间文本表示。虽然级联模型提供了模块化和优化的组件,但它们遭受错误传播,延迟增加和韵律损失。相比之下,直接S2ST模型保留说话人身份,减少延迟,并通过保留语音特征和韵律来提高翻译自然度。然而,它们仍然受到数据稀疏、高计算成本和低资源语言的泛化挑战的限制。目前的工作批判性地评估这些方法,他们的权衡,以及未来的发展方向,以提高实时多语言通信。
摘要:Speech to speech translation (S2ST) is a transformative technology thatbridges global communication gaps, enabling real time multilingual interactionsin diplomacy, tourism, and international trade. Our review examines theevolution of S2ST, comparing traditional cascade models which rely on automaticspeech recognition (ASR), machine translation (MT), and text to speech (TTS)components with newer end to end and direct speech translation (DST) modelsthat bypass intermediate text representations. While cascade models offermodularity and optimized components, they suffer from error propagation,increased latency, and loss of prosody. In contrast, direct S2ST models retainspeaker identity, reduce latency, and improve translation naturalness bypreserving vocal characteristics and prosody. However, they remain limited bydata sparsity, high computational costs, and generalization challenges forlow-resource languages. The current work critically evaluates these approaches,their tradeoffs, and future directions for improving real time multilingualcommunication.

【5】 From Voice to Safety: Language AI Powered Pilot-ATC Communication  Understanding for Airport Surface Movement Collision Risk Assessment
标题:从语音到安全:语言人工智能驱动的飞行员-ATC通信理解机场表面移动碰撞风险评估
链接:https://arxiv.org/abs/2503.04974
作者:Yutian Pang,  Andrew Paul Kendall,  Alex Porcayo,  Mariah Barsotti,  Anahita Jain,  John-Paul Clarke
摘要:这项工作将基于语言AI的语音通信理解与碰撞风险评估相结合。所提出的框架包括两个主要部分:(a)自动语音识别(ASR);(b)表面碰撞风险建模。ASR模块通过处理语音通信记录生成信息表,作为生成潜在滑行计划和计算表面移动碰撞风险的参考。对于ASR,我们基于开源视频记录和安全调查报告收集和注释我们自己的命名实体识别(NER)数据集。此外,我们参考FAA Order JO 7110.65W和FAA Order JO 7340.2N以获得在日常航空操作中使用的飞行员和空中交通管制员(ATCo)之间的通信的启发式规则和相位收缩的列表。然后,我们提出了一种新的ATC规则增强的NER方法,它集成了启发式规则的模型训练和推理阶段,导致混合规则为基础的NER模型。我们通过比较不同的设置与不同的令牌级嵌入模型,这种混合方法的有效性。对于风险建模,我们采用了美国宇航局FACET的节点-链路机场布局图,并将每个链路的飞机滑行速度建模为对数正态分布,并推导出总滑行时间分布。然后,我们提出了一个时空制定的风险概率的两个飞机移动通过潜在的碰撞节点在地面运动。我们通过模拟两个案例研究来展示我们方法的有效性,(a)2024年1月发生的Henada机场跑道碰撞事故;(b)2024年9月发生的KATL滑行道碰撞事故。我们发现,通过了解飞行员与ATC的通信记录和分析表面运动模式,该模型通过及时提供风险评估来提高机场安全性。
摘要:This work integrates language AI-based voice communication understanding withcollision risk assessment. The proposed framework consists of two major parts,(a) Automatic Speech Recognition (ASR); (b) surface collision risk modeling.ASR module generates information tables by processing voice communicationtranscripts, which serve as references for producing potential taxi plans andcalculating the surface movement collision risk. For ASR, we collect andannotate our own Named Entity Recognition (NER) dataset based on open-sourcedvideo recordings and safety investigation reports. Additionally, we refer toFAA Order JO 7110.65W and FAA Order JO 7340.2N to get the list of heuristicrules and phase contractions of communication between the pilot and the AirTraffic Controller (ATCo) used in daily aviation operations. Then, we proposethe novel ATC Rule-Enhanced NER method, which integrates the heuristic rulesinto the model training and inference stages, resulting into hybrid rule-basedNER model. We show the effectiveness of this hybrid approach by comparingdifferent setups with different token-level embedding models. For the riskmodeling, we adopt the node-link airport layout graph from NASA FACET and modelthe aircraft taxi speed at each link as a log-normal distribution and derivethe total taxi time distribution. Then, we propose a spatiotemporal formulationof the risk probability of two aircraft moving across potential collision nodesduring ground movement. We show the effectiveness of our approach by simulatingtwo case studies, (a) the Henada airport runway collision accident happened inJanuary 2024; (b) the KATL taxiway collision happened in September 2024. Weshow that, by understanding the pilot-ATC communication transcripts andanalyzing surface movement patterns, the proposed model improves airport safetyby providing risk assessment in time.

eess.AS音频处理

【1】 Musical Source Separation of Brazilian Percussion
标题:巴西打击乐的音乐来源分离
链接:https://arxiv.org/abs/2503.04995
作者:Richa Namballa,  Giovana Morais,  Magdalena Fuentes
备注:2 pages + references, 1 figure, 1 table, Extended Abstracts for the Late-Breaking Demo Session of the 25th International Society for Music Information Retrieval Conference
摘要:音乐源分离(MSS)最近在西方音乐背景下将乐器从混合物中分离出来方面取得了重大突破,但由于缺乏数据,对非西方乐器的研究仍然有限。在这个演示中,我们使用现有的巴西sama打击乐器数据集来创建人工混合物,用于训练U-Net模型来分离桑巴舞中的传统乐器surdo鼓。尽管训练数据有限,但考虑到鼓的重复模式及其特有的低音音色,该模型有效地隔离了鼓。这些结果表明,MSS系统可以成功地利用在更具有文化包容性的情况下工作,而不需要收集大量的数据。
摘要:Musical source separation (MSS) has recently seen a big breakthrough inseparating instruments from a mixture in the context of Western music, butresearch on non-Western instruments is still limited due to a lack of data. Inthis demo, we use an existing dataset of Brazilian sama percussion to createartificial mixtures for training a U-Net model to separate the surdo drum, atraditional instrument in samba. Despite limited training data, the modeleffectively isolates the surdo, given the drum's repetitive patterns and itscharacteristic low-pitched timbre. These results suggest that MSS systems canbe successfully harnessed to work in more culturally-inclusive scenarioswithout the need of collecting extensive amounts of data.

【2】 From Voice to Safety: Language AI Powered Pilot-ATC Communication  Understanding for Airport Surface Movement Collision Risk Assessment
标题:从语音到安全:语言人工智能驱动的飞行员-ATC通信理解机场表面移动碰撞风险评估
链接:https://arxiv.org/abs/2503.04974
作者:Yutian Pang,  Andrew Paul Kendall,  Alex Porcayo,  Mariah Barsotti,  Anahita Jain,  John-Paul Clarke
摘要:这项工作将基于语言AI的语音通信理解与碰撞风险评估相结合。所提出的框架包括两个主要部分:(a)自动语音识别(ASR);(b)表面碰撞风险建模。ASR模块通过处理语音通信记录生成信息表,作为生成潜在滑行计划和计算表面移动碰撞风险的参考。对于ASR,我们基于开源视频记录和安全调查报告收集和注释我们自己的命名实体识别(NER)数据集。此外,我们参考FAA Order JO 7110.65W和FAA Order JO 7340.2N以获得在日常航空操作中使用的飞行员和空中交通管制员(ATCo)之间的通信的启发式规则和相位收缩的列表。然后,我们提出了一种新的ATC规则增强的NER方法,它集成了启发式规则的模型训练和推理阶段,导致混合规则为基础的NER模型。我们通过比较不同的设置与不同的令牌级嵌入模型,这种混合方法的有效性。对于风险建模,我们采用了美国宇航局FACET的节点-链路机场布局图,并将每个链路的飞机滑行速度建模为对数正态分布,并推导出总滑行时间分布。然后,我们提出了一个时空制定的风险概率的两个飞机移动通过潜在的碰撞节点在地面运动。我们通过模拟两个案例研究来展示我们方法的有效性,(a)2024年1月发生的Henada机场跑道碰撞事故;(b)2024年9月发生的KATL滑行道碰撞事故。我们发现,通过了解飞行员与ATC的通信记录和分析表面运动模式,该模型通过及时提供风险评估来提高机场安全性。
摘要:This work integrates language AI-based voice communication understanding withcollision risk assessment. The proposed framework consists of two major parts,(a) Automatic Speech Recognition (ASR); (b) surface collision risk modeling.ASR module generates information tables by processing voice communicationtranscripts, which serve as references for producing potential taxi plans andcalculating the surface movement collision risk. For ASR, we collect andannotate our own Named Entity Recognition (NER) dataset based on open-sourcedvideo recordings and safety investigation reports. Additionally, we refer toFAA Order JO 7110.65W and FAA Order JO 7340.2N to get the list of heuristicrules and phase contractions of communication between the pilot and the AirTraffic Controller (ATCo) used in daily aviation operations. Then, we proposethe novel ATC Rule-Enhanced NER method, which integrates the heuristic rulesinto the model training and inference stages, resulting into hybrid rule-basedNER model. We show the effectiveness of this hybrid approach by comparingdifferent setups with different token-level embedding models. For the riskmodeling, we adopt the node-link airport layout graph from NASA FACET and modelthe aircraft taxi speed at each link as a log-normal distribution and derivethe total taxi time distribution. Then, we propose a spatiotemporal formulationof the risk probability of two aircraft moving across potential collision nodesduring ground movement. We show the effectiveness of our approach by simulatingtwo case studies, (a) the Henada airport runway collision accident happened inJanuary 2024; (b) the KATL taxiway collision happened in September 2024. Weshow that, by understanding the pilot-ATC communication transcripts andanalyzing surface movement patterns, the proposed model improves airport safetyby providing risk assessment in time.

【3】 DiVISe: Direct Visual-Input Speech Synthesis Preserving Speaker  Characteristics And Intelligibility
标题:DiVISe:直接视觉输入语音合成,保留说话者特征和可理解性
链接:https://arxiv.org/abs/2503.05223
作者:Yifan Liu,  Yu Fang,  Zhouhan Lin
备注:to be published in NAACL 25
摘要:视频到语音(V2S)合成,直接从无声视频输入生成语音的任务,本质上比其他语音合成任务更具挑战性,因为需要准确地重建语音内容和扬声器特性,从视觉线索单独。最近,视听预训练已经消除了V2S中对额外声学提示的需要,而之前的方法通常依赖于此来确保训练收敛。然而,即使有预训练,现有方法在实现声学可懂度和保存说话者特定特征之间的平衡方面仍然面临挑战。我们分析了这一限制,并有动机引入DiVISe(直接视觉输入语音合成),这是一种端到端的V2S模型,可以直接从视频帧中预测Mel频谱图。尽管没有采取任何声学提示,但DiVISe有效地保留了生成的音频中的扬声器特征,并在LRS2和LRS3数据集上实现了客观和主观指标的卓越性能。我们的研究结果表明,DiVISe不仅在声学可懂度方面优于现有的V2S模型,而且随着数据和模型参数的增加,其扩展效果更好。代码和重量可以在https://github.com/PussyCat0700/DiVISe上找到。
摘要:Video-to-speech (V2S) synthesis, the task of generating speech directly fromsilent video input, is inherently more challenging than other speech synthesistasks due to the need to accurately reconstruct both speech content and speakercharacteristics from visual cues alone. Recently, audio-visual pre-training haseliminated the need for additional acoustic hints in V2S, which previousmethods often relied on to ensure training convergence. However, even withpre-training, existing methods continue to face challenges in achieving abalance between acoustic intelligibility and the preservation ofspeaker-specific characteristics. We analyzed this limitation and weremotivated to introduce DiVISe (Direct Visual-Input Speech Synthesis), anend-to-end V2S model that predicts Mel-spectrograms directly from video framesalone. Despite not taking any acoustic hints, DiVISe effectively preservesspeaker characteristics in the generated audio, and achieves superiorperformance on both objective and subjective metrics across the LRS2 and LRS3datasets. Our results demonstrate that DiVISe not only outperforms existing V2Smodels in acoustic intelligibility but also scales more effectively withincreased data and model parameters. Code and weights can be found athttps://github.com/PussyCat0700/DiVISe.

【4】 UniArray: Unified Spectral-Spatial Modeling for Array-Geometry-Agnostic  Speech Separation
标题:UniArray:用于阵列-几何-不可知语音分离的统一频谱空间建模
链接:https://arxiv.org/abs/2503.05110
作者:Weiguang Chen,  Junjie Zhang,  Jielong Yang,  Eng Siong Chng,  Xionghu Zhong
备注:5 pages, Prepirnt
摘要:阵列几何无关语音分离(AGA-SS)的目的是开发一种有效的分离方法,而不管麦克风阵列的几何形状。传统的方法依赖于无置换操作,如求和或注意机制,以捕捉空间信息。然而,这些方法通常会导致高的计算成本或破坏通道内和通道间交互期间空间信息的有效使用,从而导致次优性能。为了解决这些问题,我们提出了UniArray,一种新颖的方法,放弃了传统的交织方式。UniArray由三个关键组件组成:虚拟麦克风估计(VME)模块,特征提取和融合模块,以及分层双路径分离器。VME可确保在具有不同通道数的阵列中实现稳健的性能。特征提取和融合模块利用频谱特征提取模块和空间字典学习(SDL)模块来提取和融合频率区间级特征,从而允许分离器专注于使用融合的特征。分层双路径分离器模型沿时间和频率轴具有依赖性,同时保持计算效率。实验结果表明,UniArray优于国家的最先进的方法在SI-SDRi,WB-PESQ,NB-PESQ,和STOI在可见和不可见的阵列几何形状。
摘要:Array-geometry-agnostic speech separation (AGA-SS) aims to develop aneffective separation method regardless of the microphone array geometry.Conventional methods rely on permutation-free operations, such as summation orattention mechanisms, to capture spatial information. However, these approachesoften incur high computational costs or disrupt the effective use of spatialinformation during intra- and inter-channel interactions, leading to suboptimalperformance. To address these issues, we propose UniArray, a novel approachthat abandons the conventional interleaving manner. UniArray consists of threekey components: a virtual microphone estimation (VME) module, a featureextraction and fusion module, and a hierarchical dual-path separator. The VMEensures robust performance across arrays with varying channel numbers. Thefeature extraction and fusion module leverages a spectral feature extractionmodule and a spatial dictionary learning (SDL) module to extract and fusefrequency-bin-level features, allowing the separator to focus on using thefused features. The hierarchical dual-path separator models featuredependencies along the time and frequency axes while maintaining computationalefficiency. Experimental results show that UniArray outperformsstate-of-the-art methods in SI-SDRi, WB-PESQ, NB-PESQ, and STOI across bothseen and unseen array geometries.

【5】 S2S-Arena, Evaluating Speech2Speech Protocols on Instruction Following  with Paralinguistic Information
标题:S2 S-Arena,评估带有副语言信息的教学的Speech 2 Speech协议
链接:https://arxiv.org/abs/2503.05085
作者:Feng Jiang,  Zhiyu Lin,  Fan Bu,  Yuhao Du,  Benyou Wang,  Haizhou Li
摘要:随着大语言模型(LLM)的快速发展,语音模型引起了人们的极大关注,特别是支持语音输入和输出的speech 2speech协议的最新进展。然而,现有的基准采用基于文本的自动评估器来评估这些模型的指令跟随能力,缺乏对语音理解和生成中的非语言信息的考虑。为了解决这些问题,我们引入了S2 S-Arena,这是一种新颖的竞技场风格的S2 S基准测试,可以在真实任务中的语音输入和语音输出中使用语言信息来评估跟随能力。我们设计了154个样本,融合了TTS和现场录音在四个领域的21个任务,并手动评估现有的流行的语音模型在舞台风格的方式。实验结果表明:(1)在Speech 2Speech协议中,ASR、LLM和TTS级联的语音模型除了具有GPT-4 o的优越性能外,在文语对齐后,其性能优于联合训练的模型:(2)考虑到语言信息,语音模型的知识性主要依赖于LLM骨干,其多语言支持受到语音模块的限制;(3)优秀的语音模型已经能够理解语音输入中的非语言信息,但是生成具有非语言信息的音频仍然是一个挑战。
摘要:The rapid development of large language models (LLMs) has brought significantattention to speech models, particularly recent progress in speech2speechprotocols supporting speech input and output. However, the existing benchmarksadopt automatic text-based evaluators for evaluating the instruction followingability of these models lack consideration for paralinguistic information inboth speech understanding and generation. To address these issues, we introduceS2S-Arena, a novel arena-style S2S benchmark that evaluatesinstruction-following capabilities with paralinguistic information in bothspeech-in and speech-out across real-world tasks. We design 154 samples thatfused TTS and live recordings in four domains with 21 tasks and manuallyevaluate existing popular speech models in an arena-style manner. Theexperimental results show that: (1) in addition to the superior performance ofGPT-4o, the speech model of cascaded ASR, LLM, and TTS outperforms the jointlytrained model after text-speech alignment in speech2speech protocols; (2)considering paralinguistic information, the knowledgeability of the speechmodel mainly depends on the LLM backbone, and the multilingual support of thatis limited by the speech module; (3) excellent speech models can alreadyunderstand the paralinguistic information in speech input, but generatingappropriate audio with paralinguistic information is still a challenge.

【6】 Direct Speech to Speech Translation: A Review
标题:直接语音到语音翻译:评论
链接:https://arxiv.org/abs/2503.04799
作者:Mohammad Sarim,  Saim Shakeel,  Laeeba Javed,  Jamaluddin,  Mohammad Nadeem
摘要:语音到语音翻译(S2ST)是一种变革性的技术,可以弥合全球通信差距,实现外交,旅游和国际贸易中的实时多语言互动。我们的评论研究了S2ST的演变,比较了传统的级联模型,依赖于自动语音识别(ASR),机器翻译(MT)和文本到语音(TTS)组件与新的端到端和直接语音翻译(DST)模型,绕过中间文本表示。虽然级联模型提供了模块化和优化的组件,但它们遭受错误传播,延迟增加和韵律损失。相比之下,直接S2ST模型保留说话人身份,减少延迟,并通过保留语音特征和韵律来提高翻译自然度。然而,它们仍然受到数据稀疏、高计算成本和低资源语言的泛化挑战的限制。目前的工作批判性地评估这些方法,他们的权衡,以及未来的发展方向,以提高实时多语言通信。
摘要:Speech to speech translation (S2ST) is a transformative technology thatbridges global communication gaps, enabling real time multilingual interactionsin diplomacy, tourism, and international trade. Our review examines theevolution of S2ST, comparing traditional cascade models which rely on automaticspeech recognition (ASR), machine translation (MT), and text to speech (TTS)components with newer end to end and direct speech translation (DST) modelsthat bypass intermediate text representations. While cascade models offermodularity and optimized components, they suffer from error propagation,increased latency, and loss of prosody. In contrast, direct S2ST models retainspeaker identity, reduce latency, and improve translation naturalness bypreserving vocal characteristics and prosody. However, they remain limited bydata sparsity, high computational costs, and generalization challenges forlow-resource languages. The current work critically evaluates these approaches,their tradeoffs, and future directions for improving real time multilingualcommunication.

机器翻译由腾讯交互翻译提供,仅供参考