机器与人类的改写内容比较：改写检测 (Paraphrase Detection: Human vs. Machine Content)

The growing prominence of large language models, such as GPT-4 and ChatGPT, has led to increased concerns over academic integrity due to the potential for machine-generated content and paraphrasing. Although studies have explored the detection of human- and machine-paraphrased content, the comparison between these types of content remains underexplored. In this paper, we conduct a comprehensive analysis of various datasets commonly employed for paraphrase detection tasks and evaluate an array of detection methods. Our findings highlight the strengths and limitations of different detection methods in terms of performance on individual datasets, revealing a lack of suitable machine-generated datasets that can be aligned with human expectations. Our main finding is that human-authored paraphrases exceed machine-generated ones in terms of difficulty, diversity, and similarity implying that automatically generated texts are not yet on par with human-level performance. Transformers emerged as the most effective method across datasets with TF-IDF excelling on semantically diverse corpora. Additionally, we identify four datasets as the most diverse and challenging for paraphrase detection.

翻译：人们对于大规模语言模型（如GPT-4和ChatGPT）的关注度日益增加，由于其产生机器生成的内容和改写内容的潜力，学术诚信引起了越来越多的担忧。尽管已经有一些研究探讨了人类和机器生成的改写内容检测，但是这两种类型内容的比较仍然未曾深入研究。在本文中，我们对常用的改写检测数据集进行了全面分析，并评估了各种改写检测方法。我们的研究结果显示，在表现上不同的检测方法在各个数据集上存在优缺点，其中机器生成的数据集与人类期望相距甚远。我们的主要发现是，无论在难度、多样性还是相似性方面，人类生成的改写内容都超过机器生成的内容，这说明自动生成的文本还未达到人类水平的性能。Transformer方法在不同数据集中表现最为出色，而TF-IDF在语义多样化的语料库上表现优秀。此外，我们确定了四个数据集作为改写检测中最多样化和最具挑战性的数据集。

相关内容

数据集

关注 88

数据集，又称为资料集、数据集合或资料集合，是一种由数据所组成的集合。
Data set（或dataset）是一个数据的集合，通常以表格形式出现。每一列代表一个特定变量。每一行都对应于某一成员的数据集的问题。它列出的价值观为每一个变量，如身高和体重的一个物体或价值的随机数。每个数值被称为数据资料。对应于行数，该数据集的数据可能包括一个或多个成员。

NeurlPS 2022 | 自然语言处理相关论文分类整理

专知会员服务

51+阅读 · 2022年10月2日

【Meta AI】多模态理解研究进展，Advances in multimodal understanding research at Meta AI

专知会员服务

68+阅读 · 2022年3月20日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

史上最全！358篇机器学习&自然语言处理综述论文！都这儿了

专知会员服务

129+阅读 · 2020年7月18日