由中国科学技术大学杜俊教授,佐治亚理工学院的李锦辉教授、西北工业大学的陈景东教授、卡内基梅隆大学的Shinji Watanabe教授、西西里中部自由大学的Siniscalchi Sabato Marco教授以及代尔夫特理工大学的Odette Scharenborg教授联合举办的基于多模态信息的语音处理(MISP)国际挑战赛已经开放注册,今年关注的是家居电视场景下的多人中文聊天场景,包括音视频唤醒和音视频语音识别两个任务。MISP评测也已被接收为ICASSP 2022 Signal Processing Grand Challenge,参加两个任务排名前列的团队有机会将自己的技术方案写成论文被ICASSP 2022会议接收。欢迎大家报名参加!


MISP Challenge 2021

The challenge considers the problem of audio-visual distant multi-microphone conversational wake-up and speech recognition in everyday home environments. Both audio and video data are collected in a home TV MISP Challenge 2021 has been accepted as a Signal Processing Grand Challenge (SPGC) of ICASSP 2022!Please refer to Website for more details of ICASSP 2022 SPGC. 


Website:https://2022.ieeeicassp.org/call_for_grandchallenges.php


The challenge considers the problem of audio-visual distant multi-microphone conversational wake-up and speech recognition in everyday home environments. Both audio and video data are collected in a home TV scenario, where several people are chatting while watching TV in the living room, and they can interact with a smart speaker/TV.



Background & Task Overview

With the emergence of many speech-enable applications, the scenarios (e.g., home and meeting) are becoming increasingly challenging due to the factors of adverse acoustic environments (far-field audio, background noises, and reverberations) and conversational multi-speaker interactions with a large portion of speech overlaps. The state-of-the-art speech processing techniques based on the single audio modality encounter the performance bottlenecks, e.g., yielding the word error rate of about 40% in CHiME-6 dinner party scenario. Motivated by this, the MISP challenge aims to tackle these problems by introducing additional modality information (such as video or text), yielding better environmental and speaker robustness in realistic applications.


For the first MISP challenge, we target the home TV scenario, where several people are chatting in Chinese while watching TV in the living room and they can interact with a smart speaker/TV. As the new features, the carefully selected far-field/mid-field/near-field microphone arrays and cameras are arranged to collect both audio and video data, respectively. Also the time synchronizations among different microphone arrays and video cameras are well designed for conducting the research on the multi-modality fusion. The challenge considers the problem of distant multi-microphone conversational audio-visual wake-up and audio-visual speech recognition in everyday home environments. How to leverage on both audio and video data to improve the environmental robustness is quite interesting. The researchers from both academia and industry are warmly welcome to work on our two audio-visual tasks (with details as below) for promoting the research of speech processing using multimodal information to cross the practical threshold of realistic applications in challenging scenarios. All approaches are encouraged, whether they are emerging or established, and whether they rely on signal processing or machine learning.



Tasks

MISP Challenge 2021 features two tasks:

  1. Audio-Visual Wake Word Spotting

  2. Audio-Visual Speech Recognition with Oracle Speaker Diarization

Participants are able to submit to either one track or both.

On this web site you will find everything you need to get started, including,

  • A task overview providing details of the motivation and recording set up

  • For both Task 1 and Task 2

    1. A description of the training, development and evaluation datasets

    2. Baseline recognition and evaluation tools

    3. A detailed description of the challenge rules

    4. Instructions on how to submit your results

  • A download center with links to all tools and data packages.



Planned Schedule

  • October 20th, 2021: Registration opens

  • October 25th, 2021: Training and development set release

  • November 8th, 2021: Baseline system releases

  • December 1st - January 1st 2021: Leaderboard update for development set

  • January 1st, 2022: Evaluation set release

  • January 8th - February 3rd, 2022: Leaderboard update for evaluation set



Organizers


Jun Du

University of Science and Technology of China



Chin-Hui LEE

Georgia Institute of Technology



Jingdong Chen

Northwestern Polytechnical University



Shinji Watanabe

Carnegie Mellon University



Siniscalchi Sabato Marco

Kore University of Enna



Odette Scharenborg

Delft University of Technology


Contact Us

For additional information, 

please email us at mispchallenge@gmail.com.


Registration

Website:

https://mispchallenge.github.io/index.html