GSCLIP: 解释自然语言分布变化的框架 (GSCLIP : A Framework for Explaining Distribution Shifts in Natural Language)

Helping end users comprehend the abstract distribution shifts can greatly facilitate AI deployment. Motivated by this, we propose a novel task, dataset explanation. Given two image data sets, dataset explanation aims to automatically point out their dataset-level distribution shifts with natural language. Current techniques for monitoring distribution shifts provide inadequate information to understand datasets with the goal of improving data quality. Therefore, we introduce GSCLIP, a training-free framework to solve the dataset explanation task. In GSCLIP, we propose the selector as the first quantitative evaluation method to identify explanations that are proper to summarize dataset shifts. Furthermore, we leverage this selector to demonstrate the superiority of a generator based on language model generation. Systematic evaluation on natural data shift verifies that GSCLIP, a combined system of a hybrid generator group and an efficient selector is not only easy-to-use but also powerful for dataset explanation at scale.

翻译：帮助终端用户理解抽象分布变化可以极大地促进 AI 的部署。基于此, 我们提出一个新的任务, 数据集解释。鉴于两个图像数据集, 数据集解释旨在自动指出他们与自然语言的数据集水平分布变化。监测当前分布变化的技术为了解数据集提供了不充分的信息, 目的是提高数据质量。因此, 我们引入了 GGSCLIP, 一个无培训框架, 以解决数据集解释任务。在 GSCLIP 中, 我们提议选择器作为第一个量化评估方法, 以确定适合汇总数据集变化的解释。此外, 我们利用此选择器来显示基于语言模型生成的生成器的优势。对自然数据转移的系统评估证实, GSCLIP, 一个混合生成器组合系统和高效选择器不仅容易使用, 而且在规模上对数据集解释也非常有力。

相关内容

数据集

关注 88

数据集，又称为资料集、数据集合或资料集合，是一种由数据所组成的集合。
Data set（或dataset）是一个数据的集合，通常以表格形式出现。每一列代表一个特定变量。每一行都对应于某一成员的数据集的问题。它列出的价值观为每一个变量，如身高和体重的一个物体或价值的随机数。每个数值被称为数据资料。对应于行数，该数据集的数据可能包括一个或多个成员。

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日