A PROV standard-based data source agnostic provenance engine for Big Data analytics (Supplement)
A PROV standard-based data source agnostic provenance engine for Big Data analytics (Supplement)
批准号:
9243808
负责人:
Satya Sanket Sahoo
金额:
$15.81万
依托单位国家:
美国
项目类别:
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-06-01 至 2017-05-31
关键词:
AccelerationAcuteAddressAdoptionAlgorithmsBedsBig DataBiomedical ResearchClinicalCloud ComputingComplexComputer softwareCross-Sectional StudiesDataData AnalyticsData ProvenanceData QualityData SetData SourcesDatabasesDevelopmentDiseaseEnsureExtensible Markup LanguageGenerationsGoalsGraphHealthHeterogeneityInformaticsInternetLinkMetadataModelingPerformancePhasePrincipal InvestigatorRecording of previous eventsReportingReproducibilityResearchResearch PersonnelResourcesScienceSemanticsSleepSourceStandardizationSystemTechniquesTechnologyTestingTimeTrustUnited States National Institutes of HealthWorkbasebig biomedical dataclinical carecluster computingcohortgraph theoryhealth information technologyheuristicsinsightmiddlewareoperationprogramsquality assurancereconstructionrelational databaserepositoryresearch studysuccessworking group
中文摘要
点击翻译按钮获取中文摘要
英文摘要
DESCRIPTION (provided by applicant): Data provenance is key to ensuring data quality, scientific reproducibility, and tracing the lineage of data as it undergoes transformation for use n the "data-driven" research paradigm. The emerging "Big Data" resources in biomedical research and clinical care domains have highlighted multiple computational challenges to develop a scalable and high performance provenance analysis engine. These computational challenges include semantic heterogeneity across provenance information generated from disparate sources (variety), lack of scalable provenance analytical algorithms that can keep pace with large volume of data generated at a rapid velocity. Using the new PROV representation standard recommended the W3C, which is the standard body for Web technologies, together with distributed cloud computing technologies we propose to develop a highly scalable data source agnostic provenance engine. To address the lack of appropriate provenance analytical operations required to develop this provenance engine over the PROV representation model, we will follow a three-phase approach: (1) we will first develop a new algebraic graph framework for analyzing provenance graphs conforming to the PROV standard, (2) in the second phase we will use the insights from the systematic characterization of provenance analysis operations to define distributed algorithms for implementation over cloud computing technologies, and (3) in the final step, we will implement the provenance engine that will support three fundamental provenance functions of (a) scientific reproducibility, (b) data quality assurance, and (c) trust computation. The resulting provenance engine will potentially transform the use of provenance in biomedical "Big Data" exploration and analysis techniques in the increasing number of data repositories such as the National Sleep Research Resource for accelerating data-driven research in disease mechanisms.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A PROV standard-based data source agnostic provenance engine for Big Data analytics
-
批准号:8875904
-
项目类别:
-
资助金额:$30.44万
-
财政年份:2015
-
负责人:Satya Sanket Sahoo
-
依托单位:
A PROV standard-based data source agnostic provenance engine for Big Data analytics
-
批准号:9275507
-
项目类别:
-
资助金额:$28.98万
-
财政年份:2015
-
负责人:Satya Sanket Sahoo
-
依托单位:
海外基金