用于文本摘要的印度语数据集概览 (An Overview of Indian Language Datasets used for Text Summarization)

In this paper, we survey Text Summarization (TS) datasets in Indian Languages (ILs), which are also low-resource languages (LRLs). We seek to answer one primary question: is the pool of Indian Language Text Summarization (ILTS) dataset growing or is there a resource poverty? To an-swer the primary question, we pose two sub-questions that we seek about ILTS datasets: first, what characteristics: format and domain do ILTS datasets have? Second, how different are those characteristics of ILTS datasets from high-resource languages (HRLs) particularly English. We focus on datasets reported in published ILTS research works during 2012-2022. The survey of ILTS and English datasets reveals two similarities and one contrast. The two similarities are: first, the domain of dataset commonly is news (Hermann et al., 2015). The second similarity is the format of the dataset which is both extractive and abstractive. The contrast is in how the research in dataset development has progressed. ILs face a slow speed of development and public release of datasets as compared with English. We argue that the relatively lower number of ILTS datasets is because of two reasons: first, absence of a dedicated forum for developing TS tools and resources; and second, lack of shareable standard datasets in the public domain.

翻译：在本文中,我们调查印度语言(ILS)中的文本总和数据集,这些数据集也是低资源语言(LLLs)的低资源语言(LRLs)。我们试图回答一个主要问题:印度语言文本总和数据集(ILTS)的集合正在增长,还是存在资源贫乏?对于一个更进一步的首要问题,我们提出了我们寻求ILTS数据集的两个子问题:第一,特征是什么:ILTS数据集的格式和域是什么?第二,ILTS数据集与高资源语言(HRLs),特别是英语的这些特征有何不同?我们侧重于在已出版的ILTS研究作品中报告的数据集。对 ILTS和英语数据集的调查显示两个相似之处。两个相似之处是:第一,通常的数据集领域是新闻(Hermann等人,2015年)。第二个相似之处是数据集的格式,它既是采掘的,又是抽象的。数据开发过程是如何进展的。我们关注ILTS的慢速率,而ILS的公开域数据共享是相对的论坛,因为我们开发数据缺乏了两种数据工具,而ISD的相对而言,我们开发和I标准的公开数据共享是缺乏数据工具。

相关内容

数据集

关注 88

数据集，又称为资料集、数据集合或资料集合，是一种由数据所组成的集合。
Data set（或dataset）是一个数据的集合，通常以表格形式出现。每一列代表一个特定变量。每一行都对应于某一成员的数据集的问题。它列出的价值观为每一个变量，如身高和体重的一个物体或价值的随机数。每个数值被称为数据资料。对应于行数，该数据集的数据可能包括一个或多个成员。

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日

【深度学习架构、模型和技巧集合(TensorFlow/PyTorch)】’Deep Learning Models - A collection of various deep learning architectures, models, and tips'

专知会员服务

59+阅读 · 2020年1月25日

Auto-Sizing the Transformer Network: Improving Speed, Efficiency, and Performance for Low-Resource Machine Translation

专知会员服务

49+阅读 · 2019年10月17日