Text and Causal Inference: A Review of Using Text to Remove Confounding from Causal Estimates

Text and Causal Inference: A Review of Using Text to Remove Confounding from Causal Estimates
复制标题

DOI:
10.18653/v1/2020.acl-main.474
复制
发表时间:
2020-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Katherine A. Keith;David D. Jensen;Brendan T. O'Connor
Katherine A. Keith;David D. Jensen;Brendan T. O'Connor
中科院分区:
其他
文献类型:
--
作者:
Katherine A. Keith;David D. Jensen;Brendan T. O'Connor

文献摘要

被引文献

相似文献

计算社会科学的许多应用旨在从非实验数据中推断出因果结论。此类观测数据通常包含混杂因素、影响潜在原因和潜在影响的变量。未测量或潜在的混杂因素可能会使因果估计产生偏差,这激发了人们对从观察到的文本中测量潜在混杂因素的兴趣。例如,一个人的整个社交媒体帖子历史或一篇新闻文章的内容可以提供对多个混杂因素的丰富测量。然而,这个问题的方法和应用分散在不同的社区,评估实践也不一致。这篇评论是第一次收集和分类这些例子,并为数据处理和评估决策提供指南。尽管人们越来越关注使用文本来调整混淆,但仍然存在许多悬而未决的问题,我们在本文中强调了这一点。
Many applications of computational social science aim to infer causal conclusions from non-experimental data. Such observational data often contains confounders, variables that influence both potential causes and potential effects. Unmeasured or latent confounders can bias causal estimates, and this has motivated interest in measuring potential confounders from observed text. For example, an individual’s entire history of social media posts or the content of a news article could provide a rich measurement of multiple confounders.Yet, methods and applications for this problem are scattered across different communities and evaluation practices are inconsistent.This review is the first to gather and categorize these examples and provide a guide to data-processing and evaluation decisions. Despite increased attention on adjusting for confounding using text, there are still many open problems, which we highlight in this paper.