Approximating and Reasoning about Data Provenance
Approximating and Reasoning about Data Provenance
批准号:
9243763
负责人:
Zachary Ives
金额:
$15.36万
依托单位国家:
美国
项目类别:
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-15 至 2018-05-31
关键词:
AdoptionAffectAreaBig DataCodeCommunitiesComplexCosts and BenefitsDataData AnalysesData ProvenanceData QualityEnsureEvaluationGoalsGraphHealthHumanInstitutionLaboratoriesLanguageLeadMedicineModelingOutcomePerceptionProceduresProcessProviderPythonsReproducibilityResearchResearch InfrastructureRunningSamplingScienceSiteSourceSpecific qualifier valueStagingStandardizationSystemTechniquesTechnologyTestingTimeUncertaintyWeightbasecomputerized data processingfunctional genomicsinsightinstrumentnext generation sequencingoperationreconstructionresearch studysuccesstool
中文摘要
点击翻译按钮获取中文摘要
英文摘要
DESCRIPTION (provided by applicant): In many Big Data applications today, such as Next-Generation Sequencing, data processing pipelines are highly complex, span multiple institutions, and include many human and computational steps. The pipelines evolve over time and vary across institutions, so it is difficult to track and reason about the processing pipelines
to ensure consistency and correctness of results. Provenance-enabled scientific workflow systems promise to aid here - yet such workflow systems are often avoided due to perceptions of inflexibility, lack of good provenance analytics tools, and emphasis on supporting the data consumer rather than producer. We propose to better incentivize the adoption of workflow and other provenance tracking tools: (1) Instead of requiring a single workflow system across the entire pipeline, which can be inflexible, we allow for integration across multiple autonomous systems (provenance- enabled workflow systems, provenance tracking systems for languages like Python and R, etc.), and even across steps performed without any provenance tracking at all. (2) We develop provenance reasoning capabilities specifically useful to the data provider, such as provenance analytics across time, sites, and users; finding the code modules that best explain why two results are different; regression testing to determine whether a code change would affect prior results; and reconstructing missing provenance for steps that were not captured. These capabilities are expected to lead to wider tracking of data provenance, and ultimately to more consistent, reproducible, and reliable science. We will validate this hypothesis
through the evaluation of our technologies within a Next-Generation Sequencing pipeline run by one of the PIs with collaborators at other institutions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Approximating and Reasoning about Data Provenance
-
批准号:8876037
-
项目类别:
-
资助金额:$45.66万
-
财政年份:2015
-
负责人:Zachary Ives
-
依托单位:
TRAINING PROGRAM IN BIOMEDICAL IMAGING AND INFORMATIONAL SCIENCES
-
批准号:10641331
-
项目类别:
-
资助金额:$5.28万
-
财政年份:2009
-
负责人:Zachary Ives
-
依托单位:
TRAINING PROGRAM IN BIOMEDICAL IMAGING AND INFORMATIONAL SCIENCES
-
批准号:10263929
-
项目类别:
-
资助金额:$29.87万
-
财政年份:2009
-
负责人:Zachary Ives
-
依托单位:
海外基金