公平评估基于流动的COVID-19案例预测模型 (A fairness assessment of mobility-based COVID-19 case prediction models)

In light of the outbreak of COVID-19, analyzing and measuring human mobility has become increasingly important. A wide range of studies have explored spatiotemporal trends over time, examined associations with other variables, evaluated non-pharmacologic interventions (NPIs), and predicted or simulated COVID-19 spread using mobility data. Despite the benefits of publicly available mobility data, a key question remains unanswered: are models using mobility data performing equitably across demographic groups? We hypothesize that bias in the mobility data used to train the predictive models might lead to unfairly less accurate predictions for certain demographic groups. To test our hypothesis, we applied two mobility-based COVID infection prediction models at the county level in the United States using SafeGraph data, and correlated model performance with sociodemographic traits. Findings revealed that there is a systematic bias in models performance toward certain demographic characteristics. Specifically, the models tend to favor large, highly educated, wealthy, young, urban, and non-black-dominated counties. We hypothesize that the mobility data currently used by many predictive models tends to capture less information about older, poorer, non-white, and less educated regions, which in turn negatively impacts the accuracy of the COVID-19 prediction in these regions. Ultimately, this study points to the need of improved data collection and sampling approaches that allow for an accurate representation of the mobility patterns across demographic groups.

翻译：鉴于COVID-19的爆发,分析和衡量人员流动已变得日益重要。一系列广泛的研究已经探索了时间跨度趋势,考察了与其他变数的联系,评价了非药物干预(NPIs),预测或模拟了COVID-19使用流动数据传播情况。尽管公开的流动性数据有其好处,但一个关键问题仍然没有答案:是使用流动数据的模型在不同人口群体之间公平发挥作用?我们假设,用于培训预测模型的流动数据中的偏差可能导致对某些人口群体的不合理的准确预测。为了检验我们的假设,我们在美国县一级使用两个基于流动性的COVID感染预测模型,使用安全格拉夫数据,并用社会人口特征进行相关的模型表现。调查结果显示,模型业绩有系统偏向于某些人口特征。具体地说,模型倾向于有利于大型、教育程度高、富裕、年轻、城市和非黑人占多数的州。我们假设,许多预测模型目前使用的流动数据往往不太准确掌握关于老年、较穷、非白和低龄地区的数据。这些模型的准确性模型显示美国各州的准确性特征。结果表明,模型的准确性指标的收集方法最终需要对这些COVI的准确性区域进行。

相关内容

MoDELS

关注 0

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

100+篇《自监督学习(Self-Supervised Learning)》论文最新合集

专知会员服务

166+阅读 · 2020年3月18日

【深度学习表格检测、信息提取和结构化】《Table Detection, Information Extraction and Structuring using Deep Learning》by Vihar Kurama

专知会员服务

38+阅读 · 2020年1月23日

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

专知会员服务

15+阅读 · 2019年10月23日