Machine learning for causal inference in Biostatistics.

Machine learning for causal inference in Biostatistics.
复制标题

生物统计学中因果推理的机器学习。

DOI:
10.1093/biostatistics/kxz045
复制
发表时间:
2019
期刊:
影响因子:
2.1
通讯作者:
D. Rizopoulos
D. Rizopoulos
中科院分区:
数学2区
文献类型:
--
作者:
Sherri Rose;D. Rizopoulos

文献摘要

被引文献

相似文献

一般推理问题和量化不确定性长期以来一直是统计科学的基石。虽然机器学习的进步已经渗透到许多学科,但这些过程的推理,特别是因果推理,尚未广泛普及。然而,这种情况正在迅速改变。随着不同的科学领域开始集中在机器学习上进行因果推理,我们认为现在是进行公开讨论的绝佳时机。作为《生物统计学》的编辑,我们决定组织一系列来自统计学、计算机科学、流行病学、卫生经济学、政策和法律领域的学者就该主题发表的评论。我们特意邀请了职业生涯早期或中期学者的领导者,并考虑了交叉多样性的多个维度,以在平台上增加声音的范围。专门用于因果推理的机器学习是一个较小的领域,因此这需要阅读会议议程、arXiv 论文和部门网站,并一波又一波地发送邀请,试图达到观点的平衡。并不是每个人都同意,这并不意外,因为我们有意接触我们的专业网络之外和跨学科的领域。我们分享这些经验是为了其他组织者的潜在利益,并认为这次投资是必要的。我们策划的收藏包含五件作品,我们在此简要介绍。第一篇评论讨论了在实施机器学习时理解结构性种族主义的中心地位(Robinson 等,2020)。结构性种族主义在健康应用中普遍存在,作者通过使用因果图熟练地展示了他们的论文。因果建模迫使研究人员批判性地思考他们的数据是如何生成的,并且还允许对他们的统计目标参数进行丰富的因果解释。评估和消除算法偏差是一个不断发展的研究领域,但大多数努力并不集中在健康和生物医学领域。我们鼓励学者参与这些问题并在每个项目中予以考虑。在创建要部署的工具或提出政策建议时应该需要这样做。在 Subbaswamy 和 Saria(2020)中,作者讨论了普遍性的关键主题。给定算法缺乏通用性可能是由于许多因素造成的,包括训练场景和应用该算法的系统之间的条件变化。需要认真对待普遍性
General inference problems and quantifying uncertainty have long been the cornerstone of statistical science. While machine learning advances have permeated many disciplines, inference for these procedures, and in particular, causal inference, has not been widespread. However, this is rapidly changing. As different scientific fields begin to converge on machine learning for causal inference, we thought now would be an excellent time to have a public discussion. In our roles as editors of Biostatistics, we decided to organize a series of commentaries on the topic from scholars with expertise in statistics, computer science, epidemiology, health economics, policy, and law. We intentionally invited leaders who are early or mid-career scholars and considered multiple dimensions of intersectional diversity to increase the range of voices given a platform. Machine learning specifically for causal inference is a smaller area, thus this involved reading conference programs, arXiv papers, and department websites along with sending invitations in waves in an attempt to achieve a balance of perspectives. Not everyone said yes, which is not unexpected given we were deliberately reaching outside our professional networks and across disciplines. We share these experiences for the potential benefit of other organizers and to argue that this time investment is necessary. The collection we curated contains five pieces that we briefly introduce here. The first commentary discusses the centrality of understanding structural racism when implementing machine learning (Robinson and others, 2020). Structural racism is pervasive in health applications, and the authors expertly present their thesis through the use of causal graphs. Causal modeling forces researchers to think critically about how their data were generated and also allows an enriched causal interpretation of their statistical target parameter. Assessing and eliminating algorithmic bias is a growing area of research, but most efforts do not focus on health and biomedicine. We encourage scholars to engage in these issues and give them consideration in each project. This should be required when creating tools meant to be deployed or making policy recommendations. In Subbaswamy and Saria (2020), the authors address the key topic of generalizability. A lack of generalizability for a given algorithm can be due to many factors, including shifts in conditions between the training scenario and the system where it was applied. A serious treatment of generalizability is needed