← 返回资源分享
M2H2: A Multimodal Multiparty Hindi Dataset For Humor Recognition in Conversations
论文
论文
发布时间2021-08-03
发表arXiv:2108.01260
作者:Amir Zadeh,Louis-Philippe Morency,Soujanya Poria,Navonil Majumder,Dushyant Singh Chauhan,Asif Ekbal,Pushpak Bhattacharyya,Gopendra Vikram Singh
详细介绍
Humor recognition in conversations is a challenging task that has recently gained popularity due to its importance in dialogue understanding, including in multimodal settings (i.e., text, acoustics, and visual). The few existing datasets for humor are mostly in English. However, due to the tremendous growth in multilingual content, there is a great demand to build models and systems that support multilingual information access. To this end, we propose a dataset for Multimodal Multiparty Hindi Humor (M2H2) recognition in conversations containing 6,191 utterances from 13 episodes of a very popular TV series "Shrimaan Shrimati Phir Se". Each utterance is annotated with humor/non-humor labels and encompasses acoustic, visual, and textual modalities. We propose several strong multimodal baselines and show the importance of contextual and multimodal information for humor recognition in conversations. The empirical results on M2H2 dataset demonstrate that multimodal information complements unimodal information for humor recognition. The dataset and the baselines are available at http://www.iitp.ac.in/~ai-nlp-ml/resources.html and https://github.com/declare-lab/M2H2-dataset.
代码仓库 (1)
declare-lab/M2H2-dataset官方PyTorch
