本文经arXiv每日学术速递授权转载
微信公众号:arXiv_Daily
cs.SD语音
链接:https://arxiv.org/abs/2503.20782
备注:Project page: this https URL
摘要:在本文中,我们介绍了zero-shot音视频编辑,一种新的任务,需要转换原始的视听内容,以符合指定的文本提示,而无需额外的模型训练。为了评估这项任务,我们策划了一个基准数据集,AvED-Bench,专为zero-shot音频视频编辑。AvED-Bench包括110个视频,每个视频的持续时间为10秒,跨越VGGSound的11个类别。它提供了各种各样的提示和场景,需要听觉和视觉元素之间的精确对齐,从而实现强大的评估。我们发现现有的zero-shot音频和视频编辑方法的局限性,特别是在模态之间的同步和一致性方面,这通常会导致不一致的结果。为了解决这些挑战,我们提出了AvED,一个zero-shot交叉模态增量去噪框架,利用音频-视频交互来实现同步和连贯的编辑。AvED在AvED-Bench和最新的OAVE数据集上都展示了优异的结果,以验证其泛化能力。结果见https://genjib.github.io/project_page/AVED/index.html
摘要:In this paper, we introduce zero-shot audio-video editing, a novel task that requires transforming original audio-visual content to align with a specified textual prompt without additional model training. To evaluate this task, we curate a benchmark dataset, AvED-Bench, designed explicitly for zero-shot audio-video editing. AvED-Bench includes 110 videos, each with a 10-second duration, spanning 11 categories from VGGSound. It offers diverse prompts and scenarios that require precise alignment between auditory and visual elements, enabling robust evaluation. We identify limitations in existing zero-shot audio and video editing methods, particularly in synchronization and coherence between modalities, which often result in inconsistent outcomes. To address these challenges, we propose AvED, a zero-shot cross-modal delta denoising framework that leverages audio-video interactions to achieve synchronized and coherent edits. AvED demonstrates superior results on both AvED-Bench and the recent OAVE dataset to validate its generalization capabilities. Results are available at https://genjib.github.io/project_page/AVED/index.html
【2】 FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
标题: FireRedRTS-1 S:升级的可流传输基础文本到语音系统
链接:https://arxiv.org/abs/2503.20499
摘要:在这项工作中,我们提出了一个高质量的流媒体基础的文本到语音系统,FireRedTTS-1 S,从FireRedTTS的流媒体版本升级。FireRedTTS-1 S通过两个步骤实现流生成:文本到语义解码和语义到声学解码。在文本到语义解码中,语义感知语音标记器将语音信号转换成语义标记,语义标记可以通过语义语言模型以自回归方式从文本合成。同时,语义到声学解码模块通过超分辨率因果音频编解码器和多流声学语言模型以流式方式同时将生成的语义令牌翻译成语音信号。该设计使我们能够在zero-shot设置中产生高质量的语音音频,同时呈现具有低于150 ms的低延迟的实时生成过程。在zero-shot语音克隆的实验中,客观结果验证了FireRedTTS-1 S作为高质量的基础模型,其可懂度和说话人相似度与工业基线系统相当。此外,FireRedTTS-1 S的主观评分突出了其令人印象深刻的合成性能,达到了与地面实况录音相当的质量。这些结果验证了FireRedTTS-1 S作为一个高质量的流媒体基础TTS系统。
摘要:In this work, we propose a high-quality streaming foundation text-to-speech system, FireRedTTS-1S, upgraded from the streamable version of FireRedTTS. FireRedTTS-1S achieves streaming generation via two steps: text-to-semantic decoding and semantic-to-acoustic decoding. In text-to-semantic decoding, a semantic-aware speech tokenizer converts the speech signal into semantic tokens, which can be synthesized from the text via a semantic language model in an auto-regressive manner. Meanwhile, the semantic-to-acoustic decoding module simultaneously translates generated semantic tokens into the speech signal in a streaming way via a super-resolution causal audio codec and a multi-stream acoustic language model. This design enables us to produce high-quality speech audio in zero-shot settings while presenting a real-time generation process with low latency under 150ms. In experiments on zero-shot voice cloning, the objective results validate FireRedTTS-1S as a high-quality foundation model with comparable intelligibility and speaker similarity over industrial baseline systems. Furthermore, the subjective score of FireRedTTS-1S highlights its impressive synthesis performance, achieving comparable quality to the ground-truth recordings. These results validate FireRedTTS-1S as a high-quality streaming foundation TTS system.
【3】 Qwen2.5-Omni Technical Report
标题: Qwen 2.5-Omni技术报告
链接:https://arxiv.org/abs/2503.20215
摘要:在这份报告中,我们提出了Qwen2.5-Omni,这是一个端到端的多模态模型,旨在感知不同的模态,包括文本,图像,音频和视频,同时以流的方式生成文本和自然语音响应。为了实现多模态信息输入的流式传输,音频和视频编码器都利用逐块处理方法。为了同步视频输入的时间戳与音频,我们组织的音频和视频顺序交错的方式,并提出了一种新的位置嵌入方法,名为TMRoPE(时间对齐的多模式RoPE)。为了同时生成文本和语音,同时避免两种模式之间的干扰,我们提出了\textbf{思想者-说话者}架构。在这个框架中,Thinker作为一个大型语言模型,负责文本生成,而Talker是一个双轨自回归模型,直接利用Thinker的隐藏表示来生成音频令牌作为输出。Thinker和Talker模型都被设计为以端到端的方式进行训练和推断。为了以流式方式解码音频令牌,我们引入了限制感受野的滑动窗口DiT,旨在减少初始包延迟。Qwen2.5-Omni与类似尺寸的Qwen2.5-VL相当,并且优于Qwen 2-Audio。此外,Qwen2.5-Omni在Omni-Bench等多模态基准测试中实现了最先进的性能。值得注意的是,Qwen2.5-Omni在端到端语音指令跟踪方面的性能与其文本输入能力相当,MMLU和GSM 8 K等基准测试证明了这一点。至于语音生成,Qwen2.5-Omni的流媒体Talker在鲁棒性和自然性方面优于大多数现有的流媒体和非流媒体替代品。
摘要:In this report, we present Qwen2.5-Omni, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming manner. To enable the streaming of multimodal information inputs, both audio and visual encoders utilize a block-wise processing approach. To synchronize the timestamps of video inputs with audio, we organize the audio and video sequentially in an interleaved manner and propose a novel position embedding approach, named TMRoPE(Time-aligned Multimodal RoPE). To concurrently generate text and speech while avoiding interference between the two modalities, we propose \textbf{Thinker-Talker} architecture. In this framework, Thinker functions as a large language model tasked with text generation, while Talker is a dual-track autoregressive model that directly utilizes the hidden representations from the Thinker to produce audio tokens as output. Both the Thinker and Talker models are designed to be trained and inferred in an end-to-end manner. For decoding audio tokens in a streaming manner, we introduce a sliding-window DiT that restricts the receptive field, aiming to reduce the initial package delay. Qwen2.5-Omni is comparable with the similarly sized Qwen2.5-VL and outperforms Qwen2-Audio. Furthermore, Qwen2.5-Omni achieves state-of-the-art performance on multimodal benchmarks like Omni-Bench. Notably, Qwen2.5-Omni's performance in end-to-end speech instruction following is comparable to its capabilities with text inputs, as evidenced by benchmarks such as MMLU and GSM8K. As for speech generation, Qwen2.5-Omni's streaming Talker outperforms most existing streaming and non-streaming alternatives in robustness and naturalness.
【4】 Benchmarking Machine Learning Methods for Distributed Acoustic Sensing
标题: 分布式声学传感的机器学习方法基准测试
链接:https://arxiv.org/abs/2503.20681
摘要:分布式声学传感(DAS)技术代表了一种创新的基于光纤的传感方法,该方法通过检测光纤上的微小扰动来实现实时声学信号监测。这种传感方法具有令人信服的优势,包括广泛的测量范围,卓越的空间分辨率和广泛的动态测量频谱。 机器学习(ML)范式的集成为DAS技术带来了变革性的潜力,包括数据增强、复杂的预处理技术以及先进的声学事件分类和识别等关键领域。通过利用ML算法,DAS系统可以从传统的数据处理方法过渡到更自动化和智能的分析框架。 ML增强型DAS技术提供的计算智能促进了跨各种关键基础设施部门的前所未有的监控能力。特别值得注意的是该技术在交通基础设施,能源管理系统和自然灾害监测框架中的应用,其中数据采集的精度和智能决策机制的可靠性至关重要。 这项研究在DAS数据识别和解释的背景下批判性地审视了经典机器学习方法和最先进的深度学习模型的比较性能特征,为智能传感技术不断发展的格局提供了全面的见解。
摘要:Distributed acoustic sensing (DAS) technology represents an innovative fiber-optic-based sensing methodology that enables real-time acoustic signal monitoring through the detection of minute perturbations along optical fibers. This sensing approach offers compelling advantages, including extensive measurement ranges, exceptional spatial resolution, and an expansive dynamic measurement spectrum. The integration of machine learning (ML) paradigms presents transformative potential for DAS technology, encompassing critical domains such as data augmentation, sophisticated preprocessing techniques, and advanced acoustic event classification and recognition. By leveraging ML algorithms, DAS systems can transition from traditional data processing methodologies to more automated and intelligent analytical frameworks. The computational intelligence afforded by ML-enhanced DAS technologies facilitates unprecedented monitoring capabilities across diverse critical infrastructure sectors. Particularly noteworthy are the technology's applications in transportation infrastructure, energy management systems, and Natural disaster monitoring frameworks, where the precision of data acquisition and the reliability of intelligent decision-making mechanisms are paramount. This research critically examines the comparative performance characteristics of classical machine learning methodologies and state-of-the-art deep learning models in the context of DAS data recognition and interpretation, offering comprehensive insights into the evolving landscape of intelligent sensing technologies.
【5】 QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
标题: MIDI语音:具有自然语言推理和描述的语音质量评估数据集
链接:https://arxiv.org/abs/2503.20290
备注:23 pages, 16 figures
摘要:本文通过利用自然语言描述,探索了语音质量评估的新视角,提供了比传统数值评分方法更丰富,更细致入微的见解。自然语言反馈提供了指导性的建议和详细的评估,但现有的数据集缺乏这种方法所需的全面注释。为了弥补这一差距,我们引入了一个全面的低级别语音质量评估数据集,包括11个关键方面和详细的自然语言评论,包括推理和上下文洞察。此外,我们还提出了一种新的语音基准来评估听觉大语言模型(LLM)的低级别语音理解能力。实验结果表明,微调听觉LLM可以可靠地产生详细的描述噪声和失真,有效地识别其类型和时间特性。结果进一步突出了纳入推理的潜力,以提高质量评估的准确性和可靠性。该数据集将在https://huggingface.co/datasets/tsinghua-ee/QualiSpeech上发布。
摘要:This paper explores a novel perspective to speech quality assessment by leveraging natural language descriptions, offering richer, more nuanced insights than traditional numerical scoring methods. Natural language feedback provides instructive recommendations and detailed evaluations, yet existing datasets lack the comprehensive annotations needed for this approach. To bridge this gap, we introduce QualiSpeech, a comprehensive low-level speech quality assessment dataset encompassing 11 key aspects and detailed natural language comments that include reasoning and contextual insights. Additionally, we propose the QualiSpeech Benchmark to evaluate the low-level speech understanding capabilities of auditory large language models (LLMs). Experimental results demonstrate that finetuned auditory LLMs can reliably generate detailed descriptions of noise and distortion, effectively identifying their types and temporal characteristics. The results further highlight the potential for incorporating reasoning to enhance the accuracy and reliability of quality assessments. The dataset will be released at https://huggingface.co/datasets/tsinghua-ee/QualiSpeech.
【1】 Benchmarking Machine Learning Methods for Distributed Acoustic Sensing
标题: 分布式声学传感的机器学习方法基准测试
链接:https://arxiv.org/abs/2503.20681
摘要:分布式声学传感(DAS)技术代表了一种创新的基于光纤的传感方法,该方法通过检测光纤上的微小扰动来实现实时声学信号监测。这种传感方法具有令人信服的优势,包括广泛的测量范围,卓越的空间分辨率和广泛的动态测量频谱。 机器学习(ML)范式的集成为DAS技术带来了变革性的潜力,包括数据增强、复杂的预处理技术以及先进的声学事件分类和识别等关键领域。通过利用ML算法,DAS系统可以从传统的数据处理方法过渡到更自动化和智能的分析框架。 ML增强型DAS技术提供的计算智能促进了跨各种关键基础设施部门的前所未有的监控能力。特别值得注意的是该技术在交通基础设施,能源管理系统和自然灾害监测框架中的应用,其中数据采集的精度和智能决策机制的可靠性至关重要。 这项研究批判性地研究了DAS数据识别和解释背景下经典机器学习方法和最先进的深度学习模型的比较性能特征,为智能传感技术的不断发展提供了全面的见解。
摘要:Distributed acoustic sensing (DAS) technology represents an innovative fiber-optic-based sensing methodology that enables real-time acoustic signal monitoring through the detection of minute perturbations along optical fibers. This sensing approach offers compelling advantages, including extensive measurement ranges, exceptional spatial resolution, and an expansive dynamic measurement spectrum. The integration of machine learning (ML) paradigms presents transformative potential for DAS technology, encompassing critical domains such as data augmentation, sophisticated preprocessing techniques, and advanced acoustic event classification and recognition. By leveraging ML algorithms, DAS systems can transition from traditional data processing methodologies to more automated and intelligent analytical frameworks. The computational intelligence afforded by ML-enhanced DAS technologies facilitates unprecedented monitoring capabilities across diverse critical infrastructure sectors. Particularly noteworthy are the technology's applications in transportation infrastructure, energy management systems, and Natural disaster monitoring frameworks, where the precision of data acquisition and the reliability of intelligent decision-making mechanisms are paramount. This research critically examines the comparative performance characteristics of classical machine learning methodologies and state-of-the-art deep learning models in the context of DAS data recognition and interpretation, offering comprehensive insights into the evolving landscape of intelligent sensing technologies.
【2】 QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
标题: MIDI语音:具有自然语言推理和描述的语音质量评估数据集
链接:https://arxiv.org/abs/2503.20290
备注:23 pages, 16 figures
摘要:本文通过利用自然语言描述,探索了语音质量评估的新视角,提供了比传统数值评分方法更丰富,更细致入微的见解。自然语言反馈提供了指导性的建议和详细的评估,但现有的数据集缺乏这种方法所需的全面注释。为了弥补这一差距,我们引入了一个全面的低级别语音质量评估数据集,包括11个关键方面和详细的自然语言评论,包括推理和上下文洞察。此外,我们还提出了一种新的语音基准来评估听觉大语言模型(LLM)的低级别语音理解能力。实验结果表明,微调听觉LLM可以可靠地产生详细的描述噪声和失真,有效地识别其类型和时间特性。结果进一步突出了纳入推理的潜力,以提高质量评估的准确性和可靠性。该数据集将在https://huggingface.co/datasets/tsinghua-ee/QualiSpeech上发布。
摘要:This paper explores a novel perspective to speech quality assessment by leveraging natural language descriptions, offering richer, more nuanced insights than traditional numerical scoring methods. Natural language feedback provides instructive recommendations and detailed evaluations, yet existing datasets lack the comprehensive annotations needed for this approach. To bridge this gap, we introduce QualiSpeech, a comprehensive low-level speech quality assessment dataset encompassing 11 key aspects and detailed natural language comments that include reasoning and contextual insights. Additionally, we propose the QualiSpeech Benchmark to evaluate the low-level speech understanding capabilities of auditory large language models (LLMs). Experimental results demonstrate that finetuned auditory LLMs can reliably generate detailed descriptions of noise and distortion, effectively identifying their types and temporal characteristics. The results further highlight the potential for incorporating reasoning to enhance the accuracy and reliability of quality assessments. The dataset will be released at https://huggingface.co/datasets/tsinghua-ee/QualiSpeech.
【3】 Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising
标题: 通过跨模式Delta去噪进行Zero-Shot视听编辑
链接:https://arxiv.org/abs/2503.20782
备注:Project page: this https URL
摘要:在本文中,我们介绍了zero-shot音视频编辑,一种新的任务,需要转换原始的视听内容,以符合指定的文本提示,而无需额外的模型训练。为了评估这项任务,我们策划了一个基准数据集,AvED-Bench,专为zero-shot音频视频编辑。AvED-Bench包括110个视频,每个视频的持续时间为10秒,跨越VGGSound的11个类别。它提供了各种各样的提示和场景,需要听觉和视觉元素之间的精确对齐,从而实现强大的评估。我们发现现有的zero-shot音频和视频编辑方法的局限性,特别是在模态之间的同步和一致性方面,这通常会导致不一致的结果。为了解决这些挑战,我们提出了AvED,一个zero-shot交叉模态增量去噪框架,利用音频-视频交互来实现同步和连贯的编辑。AvED在AvED-Bench和最新的OAVE数据集上都展示了优异的结果,以验证其泛化能力。结果见https://genjib.github.io/project_page/AVED/index.html
摘要:In this paper, we introduce zero-shot audio-video editing, a novel task that requires transforming original audio-visual content to align with a specified textual prompt without additional model training. To evaluate this task, we curate a benchmark dataset, AvED-Bench, designed explicitly for zero-shot audio-video editing. AvED-Bench includes 110 videos, each with a 10-second duration, spanning 11 categories from VGGSound. It offers diverse prompts and scenarios that require precise alignment between auditory and visual elements, enabling robust evaluation. We identify limitations in existing zero-shot audio and video editing methods, particularly in synchronization and coherence between modalities, which often result in inconsistent outcomes. To address these challenges, we propose AvED, a zero-shot cross-modal delta denoising framework that leverages audio-video interactions to achieve synchronized and coherent edits. AvED demonstrates superior results on both AvED-Bench and the recent OAVE dataset to validate its generalization capabilities. Results are available at https://genjib.github.io/project_page/AVED/index.html
【4】 FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
标题: FireRedRTS-1 S:升级的可流传输基础文本到语音系统
链接:https://arxiv.org/abs/2503.20499
摘要:在这项工作中,我们提出了一个高质量的流媒体基础的文本到语音系统,FireRedTTS-1 S,从FireRedTTS的流媒体版本升级。FireRedTTS-1 S通过两个步骤实现流生成:文本到语义解码和语义到声学解码。在文本到语义解码中,语义感知语音标记器将语音信号转换成语义标记,语义标记可以通过语义语言模型以自回归方式从文本合成。同时,语义到声学解码模块通过超分辨率因果音频编解码器和多流声学语言模型以流式方式同时将生成的语义令牌翻译成语音信号。该设计使我们能够在zero-shot设置中产生高质量的语音音频,同时呈现具有低于150 ms的低延迟的实时生成过程。在zero-shot语音克隆的实验中,客观结果验证了FireRedTTS-1 S作为高质量的基础模型,其可懂度和说话人相似度与工业基线系统相当。此外,FireRedTTS-1 S的主观评分突出了其令人印象深刻的合成性能,达到了与地面实况录音相当的质量。这些结果验证了FireRedTTS-1 S作为一个高质量的流媒体基础TTS系统。
摘要:In this work, we propose a high-quality streaming foundation text-to-speech system, FireRedTTS-1S, upgraded from the streamable version of FireRedTTS. FireRedTTS-1S achieves streaming generation via two steps: text-to-semantic decoding and semantic-to-acoustic decoding. In text-to-semantic decoding, a semantic-aware speech tokenizer converts the speech signal into semantic tokens, which can be synthesized from the text via a semantic language model in an auto-regressive manner. Meanwhile, the semantic-to-acoustic decoding module simultaneously translates generated semantic tokens into the speech signal in a streaming way via a super-resolution causal audio codec and a multi-stream acoustic language model. This design enables us to produce high-quality speech audio in zero-shot settings while presenting a real-time generation process with low latency under 150ms. In experiments on zero-shot voice cloning, the objective results validate FireRedTTS-1S as a high-quality foundation model with comparable intelligibility and speaker similarity over industrial baseline systems. Furthermore, the subjective score of FireRedTTS-1S highlights its impressive synthesis performance, achieving comparable quality to the ground-truth recordings. These results validate FireRedTTS-1S as a high-quality streaming foundation TTS system.
【5】 Qwen2.5-Omni Technical Report
标题: Qwen 2.5-Omni技术报告
链接:https://arxiv.org/abs/2503.20215
摘要:在这份报告中,我们提出了Qwen2.5-Omni,这是一个端到端的多模态模型,旨在感知不同的模态,包括文本,图像,音频和视频,同时以流的方式生成文本和自然语音响应。为了实现多模态信息输入的流式传输,音频和视频编码器都利用逐块处理方法。为了同步视频输入的时间戳与音频,我们组织的音频和视频顺序交错的方式,并提出了一种新的位置嵌入方法,名为TMRoPE(时间对齐的多模式RoPE)。为了同时生成文本和语音,同时避免两种模式之间的干扰,我们提出了\textbf{思想者-说话者}架构。在这个框架中,Thinker作为一个大型语言模型,负责文本生成,而Talker是一个双轨自回归模型,直接利用Thinker的隐藏表示来生成音频令牌作为输出。Thinker和Talker模型都被设计为以端到端的方式进行训练和推断。为了以流式方式解码音频令牌,我们引入了限制感受野的滑动窗口DiT,旨在减少初始包延迟。Qwen2.5-Omni与类似尺寸的Qwen2.5-VL相当,并且优于Qwen 2-Audio。此外,Qwen2.5-Omni在Omni-Bench等多模态基准测试中实现了最先进的性能。值得注意的是,Qwen2.5-Omni在端到端语音指令跟踪方面的性能与其文本输入能力相当,MMLU和GSM 8 K等基准测试证明了这一点。至于语音生成,Qwen2.5-Omni的流媒体Talker在鲁棒性和自然性方面优于大多数现有的流媒体和非流媒体替代品。
摘要:In this report, we present Qwen2.5-Omni, an end-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming manner. To enable the streaming of multimodal information inputs, both audio and visual encoders utilize a block-wise processing approach. To synchronize the timestamps of video inputs with audio, we organize the audio and video sequentially in an interleaved manner and propose a novel position embedding approach, named TMRoPE(Time-aligned Multimodal RoPE). To concurrently generate text and speech while avoiding interference between the two modalities, we propose \textbf{Thinker-Talker} architecture. In this framework, Thinker functions as a large language model tasked with text generation, while Talker is a dual-track autoregressive model that directly utilizes the hidden representations from the Thinker to produce audio tokens as output. Both the Thinker and Talker models are designed to be trained and inferred in an end-to-end manner. For decoding audio tokens in a streaming manner, we introduce a sliding-window DiT that restricts the receptive field, aiming to reduce the initial package delay. Qwen2.5-Omni is comparable with the similarly sized Qwen2.5-VL and outperforms Qwen2-Audio. Furthermore, Qwen2.5-Omni achieves state-of-the-art performance on multimodal benchmarks like Omni-Bench. Notably, Qwen2.5-Omni's performance in end-to-end speech instruction following is comparable to its capabilities with text inputs, as evidenced by benchmarks such as MMLU and GSM8K. As for speech generation, Qwen2.5-Omni's streaming Talker outperforms most existing streaming and non-streaming alternatives in robustness and naturalness.
【6】 Dolphin: A Large-Scale Automatic Speech Recognition Model for Eastern Languages
标题: Dolphin:一种面向东方语言的大规模自动语音识别模型
链接:https://arxiv.org/abs/2503.20212
摘要:本报告介绍了Dolphin,这是一种大规模的多语言自动语音识别(ASR)模型,它扩展了Whisper架构,以支持更广泛的语言。我们的方法集成了内部专有和开源数据集,以改进和优化Dolphin的性能。该模型专门设计用于对东亚,南亚,东南亚和中东的40种东方语言实现显著的识别准确性,同时还支持22种汉语方言。实验评估表明,Dolphin在各种语言中的表现明显优于当前最先进的开源模型。为了促进可重复性和社区驱动的创新,我们正在公开我们的训练模型和推理源代码。
摘要:This report introduces Dolphin, a large-scale multilingual automatic speech recognition (ASR) model that extends the Whisper architecture to support a wider range of languages. Our approach integrates in-house proprietary and open-source datasets to refine and optimize Dolphin's performance. The model is specifically designed to achieve notable recognition accuracy for 40 Eastern languages across East Asia, South Asia, Southeast Asia, and the Middle East, while also supporting 22 Chinese dialects. Experimental evaluations show that Dolphin significantly outperforms current state-of-the-art open-source models across various languages. To promote reproducibility and community-driven innovation, we are making our trained models and inference source code publicly available.
