课题基金 / 基金详情

项目摘要

项目成果

Satya Sanket Sahoo的其他基金

相似基金

相关文献

中文摘要
翻译
 描述(由申请人提供):数据出处是确保数据质量,科学再现性和跟踪数据谱系的关键,因为它经历了转换,用于“数据驱动”的研究范式。生物医学研究和临床护理领域中新兴的“大数据”资源突出了开发可扩展和高性能来源分析引擎的多个计算挑战。这些计算挑战包括从不同来源(多样性)生成的出处信息之间的语义异质性,缺乏可扩展的出处分析算法,这些算法可以跟上以快速生成的大量数据的步伐。使用W3C推荐的新PROV表示标准(Web技术的标准机构)以及分布式云计算技术,我们建议开发一个高度可扩展的数据源不可知出处引擎。为了解决在PROV表示模型上开发此来源引擎所需的适当来源分析操作的缺乏,我们将遵循三个阶段的方法:(1)我们将首先开发一个新的代数图框架,用于分析符合PROV标准的起源图,(二)在第二阶段中,我们将利用从起源分析操作的系统特征中获得的洞察力来定义分布式算法,(3)在最后一步,我们将实现出处引擎,它将支持三个基本的出处功能:(a)科学再现性,(B)数据质量保证,和(c)信任计算。由此产生的出处引擎将有可能改变出处在生物医学“大数据”探索和分析技术中的使用,如国家睡眠研究资源等越来越多的数据存储库,以加速疾病机制的数据驱动研究。
英文摘要
 DESCRIPTION (provided by applicant): Data provenance is key to ensuring data quality, scientific reproducibility, and tracing the lineage of data as it undergoes transformation for use n the "data-driven" research paradigm. The emerging "Big Data" resources in biomedical research and clinical care domains have highlighted multiple computational challenges to develop a scalable and high performance provenance analysis engine. These computational challenges include semantic heterogeneity across provenance information generated from disparate sources (variety), lack of scalable provenance analytical algorithms that can keep pace with large volume of data generated at a rapid velocity. Using the new PROV representation standard recommended the W3C, which is the standard body for Web technologies, together with distributed cloud computing technologies we propose to develop a highly scalable data source agnostic provenance engine. To address the lack of appropriate provenance analytical operations required to develop this provenance engine over the PROV representation model, we will follow a three-phase approach: (1) we will first develop a new algebraic graph framework for analyzing provenance graphs conforming to the PROV standard, (2) in the second phase we will use the insights from the systematic characterization of provenance analysis operations to define distributed algorithms for implementation over cloud computing technologies, and (3) in the final step, we will implement the provenance engine that will support three fundamental provenance functions of (a) scientific reproducibility, (b) data quality assurance, and (c) trust computation. The resulting provenance engine will potentially transform the use of provenance in biomedical "Big Data" exploration and analysis techniques in the increasing number of data repositories such as the National Sleep Research Resource for accelerating data-driven research in disease mechanisms.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A PROV standard-based data source agnostic provenance engine for Big Data analytics (Supplement)
  • 批准号:
    9243808
  • 项目类别:
  • 资助金额:
    $15.81万
  • 财政年份:
    2015
  • 负责人:
    Satya Sanket Sahoo
  • 依托单位:
A PROV standard-based data source agnostic provenance engine for Big Data analytics
  • 批准号:
    9275507
  • 项目类别:
  • 资助金额:
    $28.98万
  • 财政年份:
    2015
  • 负责人:
    Satya Sanket Sahoo
  • 依托单位:
海外基金