背景描述
淘宝公开的客服对话语料库,其中包括一个训练数据集,一个验证集和一个测试集,可用于基于检索的聊天机器人开发研究。
数据说明
数据集具体的统计信息:
| Train | Val | Test | |
|---|---|---|---|
| Session-response pairs/数量 | 1m | 10k | 10k |
| Avg. positive response per session/每节回应 | 1 | 1 | 1 |
| Min turn per session/最小回应轮数 | 3 | 3 | 3 |
| Max ture per session/最大回应轮数 | 10 | 10 | 10 |
| Average turn per session /平均回应轮数 | 5.51 | 5.48 | 5.64 |
| Average Word per utterance/平均回应字数 | 7.02 | 6.99 | 7.11 |
数据来源
https://github.com/cooelf/DeepUtteranceAggregation
@inproceedings{zhang2018dua,
title = {Modeling Multi-turn Conversation with Deep Utterance Aggregation},
author = {Zhang, Zhuosheng and Li, Jiangtong and Zhu, Pengfei and Zhao, Hai},
booktitle = {Proceedings of the 27th International Conference on Computational Linguistics (COLING 2018)},
pages={3740--3752},
year = {2018}}
