课题基金 / 基金详情

AF: Small: Algorithms and Information Theory for Causal Inference

AF: Small: Algorithms and Information Theory for Causal Inference
AF:小:因果推理的算法和信息论
批准号:
1618795
负责人:
Leonard Schulman
金额:
$45.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-08-01 至 2020-07-31

项目摘要

项目成果

Leonard Schulman的其他基金

相似基金

相关文献

中文摘要
翻译
这个项目首先涉及因果推理的算法和信息论方面。除了一些纯粹为了知识而收集的科学数据外,大多数数据都是为了潜在干预的目的而收集的:这适用于医学、公共卫生、环境法规、市场研究、歧视的法律补救办法以及许多其他领域。在不知道变量之间的因果关系的情况下,决策者无法利用在数据中发现的相关性和其他结构特征。从历史上看,因果关系是通过受控实验从相关性中分离出来的。然而,人们必须经常凑合着使用被动观察的几个充分理由:道德原因;治理制约;以及制度的独特性和无法重演历史。没有实验,我们就没有科学方法的主要武器库。然而,有一类特殊的系统,在其中可以纯粹从对统计的被动观察来进行因果关系推断。要使一个系统归入这一类,就必须能够建立在物理基础上某些可观测变量在统计上独立于某些其他变量,条件是第三个集合保持不变;这方面的形式是‘’半马尔可夫图形模型‘’。根据完美统计的假设,已知哪些半马尔可夫模型属于这一类。从这一起点出发,在这些想法能够对实践产生最大可能的影响之前,仍然存在重大的理论挑战。需要解决的一些挑战包括:(1)PI将致力于量化因果识别的稳定性(条件数)如何取决于各种不确定源(统计误差、数值误差、模型误差)以及作为图形模型结构的函数。其目的既是为了了解从现有数据中得出什么推论是合理的,也是为了影响研究设计,以便以最大的杠杆收集数据。对于前一个目标,PI特别寻求一种有效的算法来计算给定的半马尔可夫模型在特定观测统计量下的条件数。对于最后一个目标,PI寻求一个有效的算法来计算给定的半马尔可夫模型的最坏情况条件数。(2)现有的因果识别算法应用于与模型不一致的数据(这是由于统计误差而不可避免的,通常也是由于模型误差),将产生与模型不一致的推断。该项目将有助于了解投影到模型上是否可以提高稳定性。(3)使用现有方法的障碍之一是,它们要求样本大小与图形模型的大小呈指数关系。该项目的目的是确定何时可以仅使用可观测变量的小子集上的边缘分布来推断因果关系;这将减少样本量,并可能改善条件数。(4)在大多数半马尔科夫模型中,因果关系是不可识别的。然而,这留下了确定(或给出一个非平凡的外界)因果效应的可行区间的可能性。对于这个问题,目前还没有有效的算法,我们希望提供一个。这样的算法可以用来表明干预是有利的,尽管影响不是完全可识别的。(5)该项目旨在将因果推理算法提升到时间序列,并研究与该背景中通常使用的不同技术(格兰杰因果关系和梅西定向信息)的联系。该项目的次要重点包括更广泛的理论计算机科学研究。特别是,研究在算法和机器学习中使用的“提升”或“乘法加权”方法与它们在生态系统动力学(“弱选择”)和经济市场(“发展”)的系统动力学(“弱选择”)中的选择或自身利益所产生的变体之间的联系。
英文摘要
This project is concerned, firstly, with algorithmic and information-theoretic aspects of Causal Inference. With the exception of some scientific data that is gathered purely for knowledge, most data is gathered for the purpose of potential intervention: this holds for medicine, public health, environmental regulations, market research, legal remedies for discrimination, and in many other domains. A decision-maker cannot take advantage of correlations and other structural characterizations that are discovered in data without knowing about causal relationships between variables. Historically, causality has been teased apart from correlation through controlled experiments. However there are several good reasons that one must often make do with passive observation: ethical reasons; governance constraints; and uniqueness of the system and the inability to re-run history. Absent experiments, we are without the principal arsenal of the scientific method.Yet there is a special class of systems in which it is possible to perform causality inference purely from passive observation of the statistics. For a system to fall in this class one must be able to establish on physical grounds that certain observable variables are statistically independent of certain others, conditional on a third set being held fixed; the formalism for this is ``semi-Markovian graphical models". It is known which semi-Markovian models fall in this class, subject to the assumption of perfect statistics. From this starting point there remain significant theoretical challenges before these ideas can have the greatest possible impact on practice. Some of the challenges to be addressed include:(1) The PI will aim to quantify how the stability (condition number) of causal identification depends on the various sources of uncertainty (statistical error; numerical error; model error) and as a function of the structure of the graphical model. The purpose is both to understand what inference is justifiable from existing data, and to impact study design so that data with the greatest leverage is collected. For the former objective, in particular, the PI seeks an efficient algorithm to compute the condition number of a given semi-Markovian model at the specific observed statistics. For the last objective the PI seeks an efficient algorithm to compute the worst-case condition number of a given semi-Markovian model.(2) Existing causal identification algorithms, applied to data inconsistent with the model (which is unavoidable due to statistical error, and normally also due to model error), will yield an inference inconsistent with the model. The project will help to understand if projection onto the model may improve stability.(3) One of the obstacles to use of existing methods is that they require sample size exponential in the size of the graphical model. The project aims to determine when it is possible to infer causality using only the marginal distributions over small subsets of the observable variables; this will reduce sample size and likely improve condition number.(4) In the majority of semi-Markovian models, causality is not identifiable. This leaves open however the possibility of determining (or giving a nontrivial outer bound for) the feasible interval of causal effects. No effective algorithm is currently known for this problem, and we wish to provide one. Such an algorithm could be used to show that an intervention is favorable despite the effect not being fully identifiable.(5) The project aims to lift the causal-inference algorithm to time series, as well as study the connections with the distinct techniques (Granger causality and Massey's directed information) normally used in this setting.Secondary emphases of the project include broader research in theoretical computer science. In particular, studying connections between ``boosting" or ``multiplicative weights" methods used in algorithms and machine learning, and their variants which arise out of selection or self-interest in the system dynamics of ecosystems (``weak selection") and economic marketplaces (``tatonnement").Inseparably from the research effort, the PI will train students and postdocs in these and related areas of the theory of computation.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Edge Expansion and Spectral Gap of Nonnegative Matrices
非负矩阵的边扩展和谱间隙
DOI: 10.1137/1.9781611975994.73
发表时间: 2020
期刊: Proceedings of the 2020 ACM-SIAM Symposium on Discrete Algorithms
影响因子: --
作者: [Mehta, Jenish C., Schulman, Leonard J.]
通讯作者: Schulman, Leonard J.
NSF-BSF: AF: Small: Algorithmic and Information-Theoretic Challenges in Causal Inference
  • 批准号:
    2321079
  • 项目类别:
    Standard Grant
  • 资助金额:
    $61.6万
  • 财政年份:
    2023
  • 负责人:
    Leonard Schulman
  • 依托单位:
NSF-BSF: AF: Small: Identifying Functional Structure in Data
  • 批准号:
    1909972
  • 项目类别:
    Standard Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2019
  • 负责人:
    Leonard Schulman
  • 依托单位:
AF: Small: Algorithms for Inference
  • 批准号:
    1319745
  • 项目类别:
    Standard Grant
  • 资助金额:
    $47.39万
  • 财政年份:
    2013
  • 负责人:
    Leonard Schulman
  • 依托单位:
AF: EAGER: Algorithms in Linear Algebra and Optimization
  • 批准号:
    1038578
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2011
  • 负责人:
    Leonard Schulman
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: