Causal Inference不是一个统计问题 (Causal inference is not a statistical problem)

This paper introduces a collection of four data sets, similar to Anscombe's Quartet, that aim to highlight the challenges involved when estimating causal effects. Each of the four data sets is generated based on a distinct causal mechanism: the first involves a collider, the second involves a confounder, the third involves a mediator, and the fourth involves the induction of M-Bias by an included factor. The paper includes a mathematical summary of each data set, as well as directed acyclic graphs that depict the relationships between the variables. Despite the fact that the statistical summaries and visualizations for each data set are identical, the true causal effect differs, and estimating it correctly requires knowledge of the data-generating mechanism. These example data sets can help practitioners gain a better understanding of the assumptions underlying causal inference methods and emphasize the importance of gathering more information beyond what can be obtained from statistical tools alone. The paper also includes R code for reproducing all figures and provides access to the data sets themselves through an R package named quartet.

翻译：---- 本文介绍了四个数据集，类似于Anscombe's Quartet，旨在突显估计因果效应时涉及的挑战。每个数据集都基于不同的因果机制生成：第一个涉及碰撞器，第二个涉及混杂因素，第三个涉及中介因素，第四个涉及通过包括因素诱导M-Bias。本文包含每个数据集的数学摘要，以及描述变量之间关系的有向无环图(DAG)。尽管每个数据集的统计摘要和可视化相同，但真实的因果效应不同，正确估计它需要对数据生成机制有深入的了解。这些示例数据集可以帮助从业人员更好地理解因果推断方法的假设，并强调在统计工具所提供的信息之外收集更多信息的重要性。本文还提供了重现所有图形的R 代码，并通过名为quartet的R包提供访问数据集本身的方式。

相关内容

数据集

关注 88

数据集，又称为资料集、数据集合或资料集合，是一种由数据所组成的集合。
Data set（或dataset）是一个数据的集合，通常以表格形式出现。每一列代表一个特定变量。每一行都对应于某一成员的数据集的问题。它列出的价值观为每一个变量，如身高和体重的一个物体或价值的随机数。每个数值被称为数据资料。对应于行数，该数据集的数据可能包括一个或多个成员。

【2023新书】使用Python进行统计和数据可视化，554页pdf

专知会员服务

130+阅读 · 2023年1月29日

【ICDM 2022教程】图挖掘中的公平性:度量、算法和应用

专知会员服务

28+阅读 · 2022年12月26日

因果推断，Causal Inference：The Mixtape

专知会员服务

107+阅读 · 2021年8月27日

哈佛大学Hernan教授《因果推断:What If》新书，311页讲解因果效应（附下载）

专知会员服务

166+阅读 · 2021年1月7日