课题基金 / 基金详情

AF: Small: Algorithms and Information Theory for Causal Inference

AF: Small: Algorithms and Information Theory for Causal Inference
AF:小:因果推理的算法和信息论
批准号:
1618795
负责人:
Leonard Schulman
金额:
$45.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-08-01 至 2020-07-31

项目摘要

项目成果

Leonard Schulman的其他基金

相似基金

相关文献

中文摘要
翻译
这个项目首先关注因果推理的算法和信息论方面。除了一些纯粹为获取知识而收集的科学数据外,大多数数据都是为了潜在的干预目的而收集的:这适用于医药、公共卫生、环境法规、市场研究、歧视的法律补救以及许多其他领域。如果不了解变量之间的因果关系,决策者就无法利用数据中发现的相关性和其他结构特征。从历史上看,因果关系已经通过对照实验从相关性中分离出来。然而,有几个很好的理由,人们必须经常应付被动观察:道德原因;治理约束;系统的独特性以及无法重新运行历史。没有实验,我们就没有科学方法的主要武器。然而,有一类特殊的系统,在这些系统中,纯粹从被动观察统计量就可以进行因果推理。对于属于这一类的系统,必须能够在物理基础上建立某些可观察变量在统计上独立于其他某些变量,条件是第三组是固定的;它的形式是“半马尔可夫图形模型”。我们知道哪些半马尔可夫模型属于这一类,这取决于完美统计的假设。从这个起点出发,在这些想法对实践产生最大可能的影响之前,仍然存在重大的理论挑战。需要解决的一些挑战包括:(1)PI将旨在量化因果识别的稳定性(条件数)如何依赖于各种不确定性来源(统计误差;数值误差;模型误差)以及作为图形模型结构的函数。目的是了解从现有数据中得出的推断是合理的,并影响研究设计,以便收集到最具影响力的数据。特别是对于前一个目标,PI寻求一种有效的算法来计算给定半马尔可夫模型在特定观察统计量下的条件数。对于最后一个目标,PI寻求一种有效的算法来计算给定半马尔可夫模型的最坏情况数。(2)现有的因果识别算法,应用于与模型不一致的数据(由于统计误差,这是不可避免的,通常也是由于模型误差),会产生与模型不一致的推断。该项目将有助于了解,如果投影到模型可能会提高稳定性。(3)使用现有方法的障碍之一是它们需要的样本量在图形模型的大小中呈指数级增长。该项目旨在确定何时可以仅使用可观察变量的小子集的边际分布来推断因果关系;这将减少样本量,并可能改善条件数量。(4)在大多数半马尔可夫模型中,因果关系是不可识别的。然而,这留下了确定(或给出一个非平凡的外界)因果效应的可行区间的可能性。目前还没有有效的算法来解决这个问题,我们希望提供一个。这样的算法可以用来显示干预是有利的,尽管效果不能完全识别。(5)该项目旨在将因果推理算法提升到时间序列,并研究在此设置中通常使用的不同技术(Granger causality和Massey’s directed information)之间的联系。该项目的第二个重点包括理论计算机科学的更广泛的研究。特别是,研究算法和机器学习中使用的“增强”或“乘法权重”方法之间的联系,以及它们在生态系统动力学(“弱选择”)和经济市场(“补偿”)中产生的选择或自利变体。与研究工作密不可分的是,PI将在这些和相关的计算理论领域培养学生和博士后。
英文摘要
This project is concerned, firstly, with algorithmic and information-theoretic aspects of Causal Inference. With the exception of some scientific data that is gathered purely for knowledge, most data is gathered for the purpose of potential intervention: this holds for medicine, public health, environmental regulations, market research, legal remedies for discrimination, and in many other domains. A decision-maker cannot take advantage of correlations and other structural characterizations that are discovered in data without knowing about causal relationships between variables. Historically, causality has been teased apart from correlation through controlled experiments. However there are several good reasons that one must often make do with passive observation: ethical reasons; governance constraints; and uniqueness of the system and the inability to re-run history. Absent experiments, we are without the principal arsenal of the scientific method.Yet there is a special class of systems in which it is possible to perform causality inference purely from passive observation of the statistics. For a system to fall in this class one must be able to establish on physical grounds that certain observable variables are statistically independent of certain others, conditional on a third set being held fixed; the formalism for this is ``semi-Markovian graphical models". It is known which semi-Markovian models fall in this class, subject to the assumption of perfect statistics. From this starting point there remain significant theoretical challenges before these ideas can have the greatest possible impact on practice. Some of the challenges to be addressed include:(1) The PI will aim to quantify how the stability (condition number) of causal identification depends on the various sources of uncertainty (statistical error; numerical error; model error) and as a function of the structure of the graphical model. The purpose is both to understand what inference is justifiable from existing data, and to impact study design so that data with the greatest leverage is collected. For the former objective, in particular, the PI seeks an efficient algorithm to compute the condition number of a given semi-Markovian model at the specific observed statistics. For the last objective the PI seeks an efficient algorithm to compute the worst-case condition number of a given semi-Markovian model.(2) Existing causal identification algorithms, applied to data inconsistent with the model (which is unavoidable due to statistical error, and normally also due to model error), will yield an inference inconsistent with the model. The project will help to understand if projection onto the model may improve stability.(3) One of the obstacles to use of existing methods is that they require sample size exponential in the size of the graphical model. The project aims to determine when it is possible to infer causality using only the marginal distributions over small subsets of the observable variables; this will reduce sample size and likely improve condition number.(4) In the majority of semi-Markovian models, causality is not identifiable. This leaves open however the possibility of determining (or giving a nontrivial outer bound for) the feasible interval of causal effects. No effective algorithm is currently known for this problem, and we wish to provide one. Such an algorithm could be used to show that an intervention is favorable despite the effect not being fully identifiable.(5) The project aims to lift the causal-inference algorithm to time series, as well as study the connections with the distinct techniques (Granger causality and Massey's directed information) normally used in this setting.Secondary emphases of the project include broader research in theoretical computer science. In particular, studying connections between ``boosting" or ``multiplicative weights" methods used in algorithms and machine learning, and their variants which arise out of selection or self-interest in the system dynamics of ecosystems (``weak selection") and economic marketplaces (``tatonnement").Inseparably from the research effort, the PI will train students and postdocs in these and related areas of the theory of computation.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Edge Expansion and Spectral Gap of Nonnegative Matrices
非负矩阵的边扩展和谱间隙
DOI: 10.1137/1.9781611975994.73
发表时间: 2020
期刊: Proceedings of the 2020 ACM-SIAM Symposium on Discrete Algorithms
影响因子: --
作者: [Mehta, Jenish C., Schulman, Leonard J.]
通讯作者: Schulman, Leonard J.
NSF-BSF: AF: Small: Algorithmic and Information-Theoretic Challenges in Causal Inference
  • 批准号:
    2321079
  • 项目类别:
    Standard Grant
  • 资助金额:
    $61.6万
  • 财政年份:
    2023
  • 负责人:
    Leonard Schulman
  • 依托单位:
NSF-BSF: AF: Small: Identifying Functional Structure in Data
  • 批准号:
    1909972
  • 项目类别:
    Standard Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2019
  • 负责人:
    Leonard Schulman
  • 依托单位:
AF: Small: Algorithms for Inference
  • 批准号:
    1319745
  • 项目类别:
    Standard Grant
  • 资助金额:
    $47.39万
  • 财政年份:
    2013
  • 负责人:
    Leonard Schulman
  • 依托单位:
AF: EAGER: Algorithms in Linear Algebra and Optimization
  • 批准号:
    1038578
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2011
  • 负责人:
    Leonard Schulman
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: