CRII: SCH: Towards robustness to data disparities: a framework for efficient and reliable data-driven decision-making tools for all
CRII: SCH: Towards robustness to data disparities: a framework for efficient and reliable data-driven decision-making tools for all
批准号:
2153083
负责人:
Maggie Makar
金额:
$17.5万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-03-15 至 2025-02-28
中文摘要
该奖项全部或部分由《2021年美国救援计划法案》(公法117-2)资助。机器学习(ML)最有前途的应用之一是它能够指导个性化决策,特别是在医疗保健领域。预测ML模型可以帮助临床医生识别高危不良后果患者,使他们能够做出明智的预防措施决定。因果ML模型可以帮助临床医生和患者了解干预措施的影响,使他们能够对治疗方案做出更明智的决定。重要的是,预测和因果ML方法的可靠性取决于用于开发它们的数据的质量。不幸的是,数据质量往往反映了在获得医疗保健和医疗保健质量方面的系统性不平等,从而导致数据差异。护理质量不平等的例子包括:黑人病人不太可能被转诊到专科医生那里,或者妇女的疼痛不太可能被认真对待,这两者都会导致诊断和治疗的潜在延误。这意味着从人口的特定亚群中收集的数据更容易丢失。在获得医疗服务方面,数据显示,黑人和西班牙裔群体更有可能没有医疗保险,也不太可能有一个经常去的地方接受医疗服务。这导致观测数据(如通常用于开发ML模型的电子健康记录)中人口亚组的代表性不足。在本提案中,我们将开发和理论上分析鲁棒的机器学习方法(预测性和因果性),以改善数据差异的影响。拟议中的研究有两个主要方面。第一个重点是开发用于诊断的预测工具,这些工具对由于少数群体代表性不足而导致的不准确性具有鲁棒性。我们将开发模型训练方法,阻止模型学习反映数据偏差的模式,而不是真正的因果机制。我们将从理论上分析我们的模型的鲁棒性和效率。第二个重点是发展对数据缺失和测量误差可靠的干预措施因果效应的估计方法。虽然大多数现有工作都试图估计干预的因果效应,但本项目将研究反映所收集数据质量不确定性的因果估计的区间或界限的估计。当使用有限的数据进行训练时,我们从理论上分析了边界的可信度和严密性。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This award is funded in whole or in part under the American Rescue Plan Act of 2021 (Public Law 117-2).One of the most promising applications of machine learning (ML), is its ability to guide personalized decision making, especially in the context of healthcare. Predictive ML models can help clinicians identify patients at high risk of adverse outcomes, enabling them to make informed decisions about preventative measures. Causal ML models can help clinicians and patients understand the effects of interventions enabling them to make more informed decisions about the treatment options. Importantly, the reliability of predictive and causal ML methods depends on the quality of data used to develop them. Unfortunately, data quality often reflects systemic inequalities in both access to and quality of healthcare leading to data disparities. Examples of unequal quality of care include settings in which Black patients are less likely to receive referrals to specialists or in which women’s pain is less likely to be taken seriously, both leading to potential delays in diagnosis and treatment. This means that data collected from specific subgroups of the population are more prone to missingness. In terms of access, data reveal that Black and Hispanic groups are more likely to be uninsured and less likely to have a usual place to go to for medical care. This results in the underrepresentation of subgroups of the population in observational data such as electronic health records typically used to develop ML models. In this proposal, we will develop and theoretically analyze robust ML methods (both predictive and causal) that ameliorate the effects of data disparities. The proposed research has two main prongs. The first prong focuses on developing prediction tools for diagnosis that are robust to inaccuracies due to underrepresentation of minorities. We will develop model training methods that discourage the models from learning patterns that are reflective of data biases rather than true causal mechanisms. We will theoretically analyze the robustness and efficiency of our models. The second prong focuses on developing methods for estimation of causal effects of interventions that are robust to data missingness and measurement error. While most existing work attempts to estimate the causal effect of an intervention, this project will study the estimation of intervals or bounds on the causal estimates which reflect the uncertainty in the quality of the collected data. We theoretically analyze the credibility and tightness of our bounds when trained using limited data.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(7)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
--
发表时间:
2022
期刊:
Advances in neural information processing systems
影响因子:
--
作者:
[Zheng, Jiayun, Makar, Maggie]
通讯作者:
Makar, Maggie
DOI:
--
发表时间:
2023
期刊:
Proceedings of Machine Learning Research
影响因子:
--
作者:
[Xu, Jiaai, Mihalcea, Rada, Frank, Elena, Sen, Srijan, Makar, Maggie]
通讯作者:
Makar, Maggie
Conditional differential measurement error: partial identifiability and estimation
条件微分测量误差:部分可识别性和估计
DOI:
--
发表时间:
2022
期刊:
NeurIPS workshop on causal machine learning for real world impact
影响因子:
--
作者:
[Huang, Pengrun, Makar, Maggie]
通讯作者:
Makar, Maggie
Learning Concept Credible Models for Mitigating Shortcuts.
学习概念减少捷径的可靠模型。
DOI:
--
发表时间:
2022
期刊:
Advances in neural information processing systems
影响因子:
--
作者:
[Wang,Jiaxuan, Jabbour,Sarah, Makar,Maggie, Sjoding,Michael, Wiens,Jenna]
通讯作者:
Wiens,Jenna
DOI:
10.48550/arxiv.2209.09423
发表时间:
2022-09
期刊:
Trans. Mach. Learn. Res.
影响因子:
--
作者:
[Maggie Makar;A. D'Amour]
通讯作者:
Maggie Makar;A. D'Amour
共 6 条
CAREER: From Fragile to Fortified: Harnessing Causal Reasoning for Trustworthy Machine Learning with Unreliable Data
-
批准号:2337529
-
项目类别:Continuing Grant
-
资助金额:$60.0万
-
财政年份:2024
-
负责人:Maggie Makar
-
依托单位:
国内基金
海外基金
登录
查看更多内容
基于生物类芬顿的LA/Sch@BB耦合系统去除水产养殖尾水中抗生素的效果与机制研究
-
批准号:42377063
-
项目类别:面上项目
-
资助金额:49万元
-
批准年份:2023
-
负责人:王电站
-
依托单位:
具有低聚合收缩和生态防龋双功能的埃洛石纳米管@SCH-79797改性复合树脂的研究
-
批准号:82170950
-
项目类别:面上项目
-
资助金额:52万元
-
批准年份:2021
-
负责人:潘乙怀
-
依托单位:
一类稳态Schödinger-Poisson-Slater方程标准化解的研究
-
批准号:11501137
-
项目类别:青年科学基金项目
-
资助金额:18.0万元
-
批准年份:2015
-
负责人:罗庭健
-
依托单位:
锥中修改的Poisson-Sch积分在无穷远点处的渐近行为及其应用
-
批准号:U1304102
-
项目类别:联合基金项目
-
资助金额:30.0万元
-
批准年份:2013
-
负责人:乔蕾
-
依托单位:
酵母中Sch9蛋白激酶信号途径调控衰老的分子机理
-
批准号:30671181
-
项目类别:面上项目
-
资助金额:24.0万元
-
批准年份:2006
-
负责人:刘科
-
依托单位: