In this paper, we survey Text Summarization (TS) datasets in Indian Languages (ILs), which are also low-resource languages (LRLs). We seek to answer one primary question: is the pool of Indian Language Text Summarization (ILTS) dataset growing or is there a resource poverty? To an-swer the primary question, we pose two sub-questions that we seek about ILTS datasets: first, what characteristics: format and domain do ILTS datasets have? Second, how different are those characteristics of ILTS datasets from high-resource languages (HRLs) particularly English. We focus on datasets reported in published ILTS research works during 2012-2022. The survey of ILTS and English datasets reveals two similarities and one contrast. The two similarities are: first, the domain of dataset commonly is news (Hermann et al., 2015). The second similarity is the format of the dataset which is both extractive and abstractive. The contrast is in how the research in dataset development has progressed. ILs face a slow speed of development and public release of datasets as compared with English. We argue that the relatively lower number of ILTS datasets is because of two reasons: first, absence of a dedicated forum for developing TS tools and resources; and second, lack of shareable standard datasets in the public domain.
翻译:在本文中,我们调查印度语言(ILS)中的文本总和数据集,这些数据集也是低资源语言(LLLs)的低资源语言(LRLs)。我们试图回答一个主要问题:印度语言文本总和数据集(ILTS)的集合正在增长,还是存在资源贫乏?对于一个更进一步的首要问题,我们提出了我们寻求ILTS数据集的两个子问题:第一,特征是什么:ILTS数据集的格式和域是什么?第二,ILTS数据集与高资源语言(HRLs),特别是英语的这些特征有何不同?我们侧重于在已出版的ILTS研究作品中报告的数据集。对 ILTS和英语数据集的调查显示两个相似之处。两个相似之处是:第一,通常的数据集领域是新闻(Hermann等人,2015年)。第二个相似之处是数据集的格式,它既是采掘的,又是抽象的。数据开发过程是如何进展的。我们关注ILTS的慢速率,而ILS的公开域数据共享是相对的论坛,因为我们开发数据缺乏了两种数据工具,而ISD的相对而言,我们开发和I标准的公开数据共享是缺乏数据工具。