探索假新闻探测精制模型的通用性 (Exploring Generalizability of Fine-Tuned Models for Fake News Detection)

The Covid-19 pandemic has caused a dramatic and parallel rise in dangerous misinformation, denoted an `infodemic' by the CDC and WHO. Misinformation tied to the Covid-19 infodemic changes continuously; this can lead to performance degradation of fine-tuned models due to concept drift. Degredation can be mitigated if models generalize well-enough to capture some cyclical aspects of drifted data. In this paper, we explore generalizability of pre-trained and fine-tuned fake news detectors across 9 fake news datasets. We show that existing models often overfit on their training dataset and have poor performance on unseen data. However, on some subsets of unseen data that overlap with training data, models have higher accuracy. Based on this observation, we also present KMeans-Proxy, a fast and effective method based on K-Means clustering for quickly identifying these overlapping subsets of unseen data. KMeans-Proxy improves generalizability on unseen fake news datasets by 0.1-0.2 f1-points across datasets. We present both our generalizability experiments as well as KMeans-Proxy to further research in tackling the fake news problem.

翻译：Covid-19大流行导致危险的错误信息急剧和平行上升,这说明疾病防治中心和世卫组织的“信息”不断发生与Covid-19进化变化相联系的错误信息;这可能导致由于概念的漂移而微调模型的性能退化;如果模型能够广泛推广,以捕捉漂流数据的某些周期性方面,那么脱色是可以减轻的。在本文中,我们探索了9个假新闻数据集中经过预先训练的和经过精细调的假冒新闻探测器的普遍适用性。我们显示,现有模型往往过分适合其培训数据集,而且无法很好地利用无法见的数据。然而,关于与培训数据重叠的一些秘密数据,模型的准确性更高。根据这一观察,我们还介绍了基于K-Means集群的快速有效方法,即快速识别这些重叠的隐形数据组。 Kmains-Proxy用0.1-0.2 f1点的无形新闻数据集改进了普通新闻数据集的通用性。我们介绍了我们的一般性实验问题,作为Kmains-Prox的进一步研究。

相关内容

MoDELS

关注 43

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

Into the Metaverse，93页ppt介绍元宇宙概念、应用、趋势

专知会员服务

49+阅读 · 2022年2月19日

【PAISS 2021 教程】概率散度与生成式模型，92页ppt

专知会员服务

34+阅读 · 2021年11月30日

【跨语言BERT模型大集合】Transfer learning is increasingly going multilingual with language-specific BERT models

专知会员服务

54+阅读 · 2020年1月30日

【深度学习表格检测、信息提取和结构化】《Table Detection, Information Extraction and Structuring using Deep Learning》by Vihar Kurama

专知会员服务

38+阅读 · 2020年1月23日