课题基金 / 基金详情

EAGER: Towards a Computational Infrastructure for Analysis of Sensitive Data

EAGER: Towards a Computational Infrastructure for Analysis of Sensitive Data
EAGER:建立用于分析敏感数据的计算基础设施
批准号:
1551843
负责人:
Vasant Honavar
金额:
$23.16万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2019-08-31

项目摘要

项目成果

Vasant Honavar的其他基金

相似基金

相关文献

中文摘要
翻译
在许多应用中,例如,在卫生、教育等领域,由于缺乏以不违反适用的数据访问和使用政策的方式分析敏感数据的实用框架,我们充分发挥大数据潜力以改善决策和成果的能力目前受到限制。这对具有分析专业知识的研究人员参与开发和评估用于分析此类数据的先进方法,评估替代方法的性能或确保结果的可重复性构成了重大障碍。在这种背景下,这个高风险和潜在的高影响力的研究项目旨在探索一个框架和一个软件基础设施,用于敏感数据的数据访问和使用政策(DAUP)合规分析。 该项目旨在为敏感数据的数据访问和使用政策(DAUP)合规分析开发一个新的框架。该框架将支持(i)从用户或项目特定DAUP允许的数据存储中查询和检索信息。这种信息可以包括数据存储的模式、指定变量的元数据以及它们的域和范围等; (ii)执行系统或用户提供的算法实现,用于从数据存储中的数据构建预测或因果模型或可视化;(iii)评估所得模型对基准数据或用户提供的数据的预测性能;(iv)以网络服务器的形式部署经验证的模型,这些网络服务器提供对用户提交的数据或对用户提交的数据的结果的预测或可视化,针对数据存储的已定义查询;以及(v)可重用分析工作流的发布。这个探索性的项目旨在测试框架的可行性,使用预测和因果建模的数据从一个在线健康社区作为测试案例。 该项目的一个主要成果是开放源码软件基础设施,以促进敏感数据的分析和可视化。这项研究将:(i)填补了从敏感数据进行预测建模的基础设施的主要空白;(ii)显著降低了具有深入分析专业知识的研究人员进入领域的门槛(例如,卫生、教育),涉及敏感数据;(iii)通过促进算法的严格比较,提高这些领域中预测和因果建模的最新技术评估的准确性;以及(iv)促进敏感数据的可重复分析。这项研究将(i)产生一个原型开源软件基础设施,以支持敏感数据的分析和可视化;(ii)加速涉及敏感数据的领域的数据驱动的进步,例如,通过广泛吸引人才参与开发更好的算法,促进健康和教育;以及(iii)通过围绕特定敏感数据集组织的黑客马拉松和竞赛,支持将此类应用程序的实践经验纳入数据科学教育。
英文摘要
In many applications, e.g., health, education, our ability to realize the full potential of big data to improve decisions and outcomes is currently limited, by the lack of practical frameworks for analysis of sensitive data in a manner that does not violate applicable data access and use policies. This constitutes a significant barrier to the engagement of researchers with expertise in analytics in developing and evaluating advanced methods for analysis of such data, assessing the performance of alternative approaches, or ensuring the reproducibility of results. Against this background, this high-risk and potentially high-impact research project aims to explore a framework and a software infrastructure for data access and use policy (DAUP) compliant analysis of sensitive data. The project aims to develop a novel framework for data access and use policy (DAUP) compliant analysis of sensitive data. The framework will support (i) Querying and retrieval of information from the data store that are permitted by the user or project specific DAUP. Such information could include the schema of the data store, metadata that specify the variables, and their domains and ranges, etc.; (ii) Execution of system or user-supplied implementations of algorithms for construct predictive or causal models or visualizations from the data in the data store; (iii) Evaluation of the predictive performance of the resulting models on benchmark data or user-provided data; (iv) Deployment of the validated models in the form of web servers that provide predictions or visualizations over user-submitted data or over results of user-defined queries against the data store; and (v) Publication of reusable analytics workflows. This exploratory project seeks to test the feasibility of the framework using predictive and causal modeling of data from an online health community as a test case. A major outcome of this project is the open source software infrastructure for facilitating analysis and visualization of sensitive data. This research will: (i) fill a major gap in infrastructure for predictive modeling from sensitive data; (ii) significantly lower the barrier to the entry of researchers with deep expertise in analytics to domains (e.g., health, education) that involve sensitive data; (iii) improve the accuracy of assessment of the state-of-the-art in predictive and causal modeling in such domains by facilitating rigorous comparison of algorithms; and (iv) facilitate, reproducible analysis of sensitive data. This research will (i) yield a prototype open source software infrastructure to support analysis and visualization of sensitive data; (ii) Accelerate data-driven advances in domains that involve sensitive data e.g., health, education through broad engagement of talent in developing better algorithms; and (iii) Support incorporation of hands-on experience with such applications into Data Sciences education through hackathons and competitions organized around specific sensitive data sets.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: RI: III: SHF: Small: Multi-Stakeholder Decision Making: Qualitative Preference Languages, Interactive Reasoning, and Explanation
III: Small: Predictive Modeling from High-Dimensional, Sparsely and Irregularly Sampled, Longitudinal Data
AI Institute: Planning: Institute for AI-Enabled Materials Discovery, Design, and Synthesis
EAGER: Interpreting Black-Box Predictive Models Through Causal Attribution
海外基金