A PROV standard-based data source agnostic provenance engine for Big Data analytics
A PROV standard-based data source agnostic provenance engine for Big Data analytics
批准号:
9275507
负责人:
Satya Sanket Sahoo
金额:
$28.98万
依托单位国家:
美国
项目类别:
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-06-01 至 2019-05-31
关键词:
AccelerationAcuteAddressAdoptionAlgebraAlgorithmsBedsBig DataBiomedical ResearchClinicalCloud ComputingComplexComputer softwareCross-Sectional StudiesDataData AnalyticsData ProvenanceData QualityData SetData SourcesDatabasesDevelopmentDiseaseEnsureExtensible Markup LanguageGenerationsGoalsGraphHeterogeneityInformaticsInternetIntuitionLinkMetadataModalityModelingPerformancePhasePrincipal InvestigatorRecording of previous eventsReportingReproducibilityResearchResearch PersonnelResourcesScienceSemanticsSleepSourceStandardizationSystemTechniquesTechnologyTestingTimeTrustUnited States National Institutes of HealthWorkbasebig biomedical dataclinical carecluster computingcohortdata resourcegraph theoryhealth information technologyheuristicsinsightmiddlewareoperationprogramspublic health relevancequality assurancereconstructionrelational databaserepositoryresearch studysuccessworking group
中文摘要
描述(由申请人提供):数据来源是确保数据质量、科学再现性和追踪数据谱系的关键,因为数据正在经历转换,以用于“数据驱动”的研究范式。生物医学研究和临床护理领域不断涌现的大数据资源突显了开发可扩展和高性能来源分析引擎的多重计算挑战。这些计算挑战包括从不同来源(多样性)生成的来源信息之间的语义异构性,缺乏可扩展的来源分析算法来跟上快速生成的大量数据的步伐。使用W3C推荐的新的PROV表示标准,结合分布式云计算技术,我们提出了开发一个高度可扩展的数据源不可知来源引擎。为了解决在PROV表示模型上开发这个起源引擎所需的适当的起源分析操作的不足,我们将遵循三个阶段的方法:(1)我们将首先开发一个新的代数图框架来分析符合PROV标准的起源图,(2)在第二阶段,我们将使用来自起源分析操作的系统表征的见解来定义用于在云计算技术上实现的分布式算法,以及(3)在最后一步,我们将实现起源引擎,该引擎将支持(A)科学可重复性、(B)数据质量保证和(C)信任计算的三个基本起源功能。由此产生的起源引擎可能会改变起源在生物医学“大数据”探索和分析技术中的使用,并在越来越多的数据储存库中使用,例如国家睡眠研究资源,以加速疾病机制中的数据驱动研究。
英文摘要
DESCRIPTION (provided by applicant): Data provenance is key to ensuring data quality, scientific reproducibility, and tracing the lineage of data as it undergoes transformation for use n the "data-driven" research paradigm. The emerging "Big Data" resources in biomedical research and clinical care domains have highlighted multiple computational challenges to develop a scalable and high performance provenance analysis engine. These computational challenges include semantic heterogeneity across provenance information generated from disparate sources (variety), lack of scalable provenance analytical algorithms that can keep pace with large volume of data generated at a rapid velocity. Using the new PROV representation standard recommended the W3C, which is the standard body for Web technologies, together with distributed cloud computing technologies we propose to develop a highly scalable data source agnostic provenance engine. To address the lack of appropriate provenance analytical operations required to develop this provenance engine over the PROV representation model, we will follow a three-phase approach: (1) we will first develop a new algebraic graph framework for analyzing provenance graphs conforming to the PROV standard, (2) in the second phase we will use the insights from the systematic characterization of provenance analysis operations to define distributed algorithms for implementation over cloud computing technologies, and (3) in the final step, we will implement the provenance engine that will support three fundamental provenance functions of (a) scientific reproducibility, (b) data quality assurance, and (c) trust computation. The resulting provenance engine will potentially transform the use of provenance in biomedical "Big Data" exploration and analysis techniques in the increasing number of data repositories such as the National Sleep Research Resource for accelerating data-driven research in disease mechanisms.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
ProvCaRe Semantic Provenance Knowledgebase: Evaluating Scientific Reproducibility of Research Studies.
ProvCaRe 语义起源知识库:评估研究的科学再现性。
DOI:
--
发表时间:
2017
期刊:
AMIA ... Annual Symposium proceedings. AMIA Symposium
影响因子:
--
作者:
[Valdez,Joshua, Kim,Matthew, Rueschman,Michael, Socrates,Vimig, Redline,Susan, Sahoo,SatyaS]
通讯作者:
Sahoo,SatyaS
An Ontology-Enabled Natural Language Processing Pipeline for Provenance Metadata Extraction from Biomedical Text (Short Paper).
用于从生物医学文本中提取来源元数据的本体支持的自然语言处理管道(短论文)。
DOI:
10.1007/978-3-319-48472-3_43
发表时间:
2016
期刊:
On the move to meaningful Internet systems ... : CoopIS, DOA, and ODBASE : Confederated International Conferences, CoopIS, DOA, and ODBASE ... proceedings. OTM Confederated International Conferences
影响因子:
--
作者:
[Valdez,Joshua, Rueschman,Michael, Kim,Matthew, Redline,Susan, Sahoo,SatyaS]
通讯作者:
Sahoo,SatyaS
DOI:
10.1016/j.ijmedinf.2018.10.009
发表时间:
2019-01-01
期刊:
INTERNATIONAL JOURNAL OF MEDICAL INFORMATICS
影响因子:
4.9
作者:
[Sahoo, Satya S., Valdez, Joshua, Redline, Susan]
通讯作者:
Redline, Susan
Computing Functional Brain Connectivity in Neurological Disorders: Efficient Processing and Retrieval of Electrophysiological Signal Data
计算神经疾病中的功能性大脑连接:电生理信号数据的有效处理和检索
DOI:
--
发表时间:
2019
期刊:
AMIA Jt Summits Transl Sci Proc.
影响因子:
--
作者:
[Gershon, A, Devulapalli, P, Zonjy, B, Ghosh, K, Tatsuoka, C, Sahoo, SS]
通讯作者:
Sahoo, SS
Semantic Provenance Graph for Reproducibility of Biomedical Research Studies: Generating and Analyzing Graph Structures from Published Literature.
用于生物医学研究再现性的语义起源图:从已发表的文献中生成和分析图结构。
DOI:
10.3233/shti190237
发表时间:
2019
期刊:
Studies in health technology and informatics
影响因子:
--
作者:
[Sahoo,SatyaS, Valdez,Joshua, Rueschman,Michael, Kim,Matthew]
通讯作者:
Kim,Matthew
共 9 条
A PROV standard-based data source agnostic provenance engine for Big Data analytics
-
批准号:8875904
-
项目类别:
-
资助金额:$30.44万
-
财政年份:2015
-
负责人:Satya Sanket Sahoo
-
依托单位:
A PROV standard-based data source agnostic provenance engine for Big Data analytics (Supplement)
-
批准号:9243808
-
项目类别:
-
资助金额:$15.81万
-
财政年份:2015
-
负责人:Satya Sanket Sahoo
-
依托单位:
海外基金