Here we present a conversational dataset in Mandarin Chinese, code mixed with English words and phrases.The total duration of the original dataset is about 22.54 hours, with an effective duration of about 9.57 hours. We split the dataset into two parts: the DEV set and the test set.We present only the TEST part here for open access, of which the total duration is about 10 hours. Audio files (.wav) with segments and manually annotated transcriptions are contained in the dataset.10 participants (4 males and 6 females) from whom we collected the audio data from were aged 21 - 25 years old. And in total, 42 audios were collected, corresponding to 42 annotated texts.The word correct rate of this dataset is above 99% when we test and evaluate this set.
← 返回资源分享
ASR-SECOMICSC
语音小管家 · 数据
数据Speech Recognition
提供方语音小管家
许可协议Creative Commons
发布时间2025-01-02
获取方式
https://magichub.com/datasets/chinese-english-code-mixing-conversational-speech-corpus/
