StatCan对话数据集：通过与真实意图的对话检索数据表 (The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine Intents)

We introduce the StatCan Dialogue Dataset consisting of 19,379 conversation turns between agents working at Statistics Canada and online users looking for published data tables. The conversations stem from genuine intents, are held in English or French, and lead to agents retrieving one of over 5000 complex data tables. Based on this dataset, we propose two tasks: (1) automatic retrieval of relevant tables based on a on-going conversation, and (2) automatic generation of appropriate agent responses at each turn. We investigate the difficulty of each task by establishing strong baselines. Our experiments on a temporal data split reveal that all models struggle to generalize to future conversations, as we observe a significant drop in performance across both tasks when we move from the validation to the test set. In addition, we find that response generation models struggle to decide when to return a table. Considering that the tasks pose significant challenges to existing models, we encourage the community to develop models for our task, which can be directly used to help knowledge workers find relevant tables for live chat users.

翻译：我们介绍了StatCan对话数据集，它包括19379个对话轮次，涉及Statistics Canada的代理人和在线用户之间的交流，用户寻找已发布的数据表格。这些对话源自真实意图，使用英语或法语进行，并导致代理人检索其中之一，共5000多个复杂的数据表格。基于该数据集，我们提出了两个任务：（1）根据正在进行的对话自动检索相关表格，（2）在每个轮次自动生成适当的代理人响应。我们通过建立强大的基线来研究每个任务的难度。在时间数据分割的实验中，我们发现所有模型都难以推广到未来的对话中，当我们从验证集移动到测试集时，两个任务的性能都明显下降。此外，我们发现响应生成模型很难确定何时返回表格。考虑到这些任务对现有模型构成了重大挑战，我们鼓励社区开发我们的任务模型，这些模型可以直接用于帮助知识工作者为在线聊天用户找到相关表格。

相关内容

数据集

关注 88

数据集，又称为资料集、数据集合或资料集合，是一种由数据所组成的集合。
Data set（或dataset）是一个数据的集合，通常以表格形式出现。每一列代表一个特定变量。每一行都对应于某一成员的数据集的问题。它列出的价值观为每一个变量，如身高和体重的一个物体或价值的随机数。每个数值被称为数据资料。对应于行数，该数据集的数据可能包括一个或多个成员。

如何使用TensorFlow 排序构建推荐系统? How to build a recommendation system using TensorFlow Ranking?

专知会员服务

19+阅读 · 2022年3月13日

【ETH】最新《几何数据分析》2020课程，附PPT下载

专知会员服务

45+阅读 · 2020年12月18日

【KDD2020】基于知识图谱的语义融合改进会话推荐系统，Improving Conversational Recommender Systems via Knowledge Graph based Semantic Fusion

专知会员服务

90+阅读 · 2020年7月9日

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

专知会员服务

15+阅读 · 2019年10月23日