← 返回资源分享
From Masked Language Modeling to Translation: Non-English Auxiliary Tasks Improve Zero-shot Spoken Language Understanding
论文
论文
发布时间2021-05-15
发表NAACL 2021 4 · arXiv:2105.07316
作者:Mamoru Komachi,Marija Stepanović,Barbara Plank,Rob van der Goot,Ahmet Üstün,Alan Ramponi,Ibrahim Sharaf,Aizhan Imankulova,Siti Oryza Khairunnisa
详细介绍
The lack of publicly available evaluation data for low-resource languages limits progress in Spoken Language Understanding (SLU). As key tasks like intent classification and slot filling require abundant training data, it is desirable to reuse existing data in high-resource languages to develop models for low-resource scenarios. We introduce xSID, a new benchmark for cross-lingual Slot and Intent Detection in 13 languages from 6 language families, including a very low-resource dialect. To tackle the challenge, we propose a joint learning approach, with English SLU training data and non-English auxiliary tasks from raw text, syntax and translation for transfer. We study two setups which differ by type and language coverage of the pre-trained embeddings. Our results show that jointly learning the main tasks with masked language modeling is effective for slots, while machine translation transfer works best for intent classification.
代码仓库 (2)
robvanderg/xsid官方PyTorch
Kaleidophon/deep-significance官方TensorFlow
