A modular data analysis ecosystem using portable encapsulated projects
A modular data analysis ecosystem using portable encapsulated projects
批准号:
9751344
负责人:
Nathan Sheffield
金额:
$39.32万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-08-01 至 2023-07-31
关键词:
AdoptedBioinformaticsBiomedical ResearchComplexDataData AnalysesData CollectionData SetEcosystemEncapsulatedEnvironmentGoalsHumanIndividualInheritedInstitutionKnowledgeLinkManualsProcessProviderPythonsResearchResearch Project GrantsRunningSeriesStandardizationStructureSurfaceSystemTechniquesbasebioinformatics toolcluster computingcomputerized data processingdata managementdata sharinginnovationinsightinterestnext generationnovelnovel strategiesportabilitytooltool development
中文摘要
项目摘要
概述
随着可用数据量的增加,处理数据变得更具挑战性。
表面:它是从数据到分析的映射。不幸的是,太多时候,这需要一个独特的结构,为每一个组合
数据集和分析。这使得在一个数据集上运行多个不同的分析或插入多个分析变得困难
不同的数据集到一个分析,因为每个连接结构必须手动定义。
为了减轻将数据链接到工具的挑战,本提案提出了便携式封装项目的概念
(PEP)以及一系列读取和处理这些项目的工具。从本质上讲,PEP格式旨在标准化
数据收集的描述,使数据提供者和数据使用者能够通过共同接口进行通信
一个标准的格式。实际上,这意味着使用这种格式描述其项目的个人将立即继承
既有更大的分析便携性,也有更大的外部补充数据访问权。此链接围绕
简单、标准、可扩展的项目定义。
伴随着这一点,该提案开发了Python和R包,以提供一个低障碍的模块化框架,
入口,可以轻松构建以PEP格式为中心的强大管道和其他工具。该系统提供了一个
组织数据密集型生物医学研究项目的新方法。
意义和创新
该建议位于数据管理和生物信息学工具开发的接口。虽然努力很大,
已经分别致力于其中的每一个,但在将两者联系起来的层面上关注较少。这项建议
将在生物信息学的数据和工具之间建立一个标准化的接口,提供格式和工具的实际进展
来促进这种互动。这项工作以一种新颖的方式处理计算项目,并建立概念和工具
可以彻底改变生物信息学研究。目标不是开发新的工具,而是让现有的工具变得更容易
适用于现有数据。
在计算研究中,大量的精力花在数据清理上:准备数据进行分析。通过促进
从数据到工具的连接,这将鼓励用新的分析技术重新分析现有数据,
的发现它还将使分析新数据与现有数据更容易,从而增加两者的价值。它将
有助于可重用性、大规模分析、便携式计算环境和数据共享。
人们对跨科学领域的数据共享和可访问性越来越感兴趣,这项提案将促进这一点。
早期的版本已经在四个不同的研究机构被用于本地计算和集群计算,
随着该项目的成熟,它将围绕一个共同的数据描述,把各种研究环境结合起来。这将使
更容易在用户、研究小组和机构之间共享数据和工具。
1
英文摘要
Project summary
Overview
As the amount of available data increases, it becomes more challenging to process it. Data processing is simple on the
surface: it is a mapping from data to analysis. Unfortunately, too often, this requires a unique structure for each combination
of dataset and analysis. This makes it difficult to do things like run several different analyses on one dataset, or plug several
different datasets to one analysis, because each connection structure must be defined manually.
To alleviate this challenge of linking data to tools, this proposal develops the concept of Portable Encapsulated Projects
(PEP) and a series of tools that read and process such projects. Essentially, the PEP format aims to standardize the
description of data collections, enabling both data providers and data users to communicate through the common interface
of a standard format. Practically, this means individuals who describe their projects using this format will immediately inherit
both greater portability for analysis as well as greater access to external complementary data. This link operates around a
simple, standard, extensible definition of a project.
Accompanying this, this proposal develops Python and R packages to provide a modular framework with a low barrier to
entry that makes it easy to build robust pipelines and other tools centered around the PEP format. This system presents a
new approach to organizing data-intensive biomedical research projects.
Significance and innovation
This proposal sits at the interface of data management and bioinformatics tool development. While significant effort is
already dedicated to each of these individually, there has been less focus at the level of connecting the two. This proposal
will build a standardized interface between data and tools in bioinformatics, providing practical advances in formats and tools
to facilitate this interaction. This effort approaches computational projects in a novel way, and builds both concepts and tools
that can revolutionize bioinformatics research. The goal is not to develop new tools, but to make existing tools more easily
applied to existing data.
In computational research, a huge amount of effort is spent in data cleanup: preparing data for analysis. By facilitating the
connection from data to tools, this will encourage re-analysis of existing data with novel analysis techniques, leading to new
discovery. It will also make it easier to analyze new data in tandem with existing data, increasing the value of both. It will
contribute to reusability, larger-scale analysis, portable computing environments, and data sharing.
There is increasing interest in data sharing and accessibility across scientific domains, and this proposal will facilitate this.
Early versions are already adopted for both local compute and cluster computing at four different research institutions, and
as the project matures, it will unite various research environments around a common data description. This will make it
easier to share data and tools across users, research groups, and institutions.
1
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Novel methods for large-scale genomic interval comparison
-
批准号:10678947
-
项目类别:
-
资助金额:$38.4万
-
财政年份:2022
-
负责人:Nathan Sheffield
-
依托单位:
Novel methods for large-scale genomic interval comparison
-
批准号:10842040
-
项目类别:
-
资助金额:$31.47万
-
财政年份:2022
-
负责人:Nathan Sheffield
-
依托单位:
A modular data analysis ecosystem using portable encapsulated projects
-
批准号:10468680
-
项目类别:
-
资助金额:$39.38万
-
财政年份:2018
-
负责人:Nathan Sheffield
-
依托单位:
A modular data analysis ecosystem using portable encapsulated projects
-
批准号:10019399
-
项目类别:
-
资助金额:$39.38万
-
财政年份:2018
-
负责人:Nathan Sheffield
-
依托单位:
A modular data analysis ecosystem using portable encapsulated projects
-
批准号:10224819
-
项目类别:
-
资助金额:$39.38万
-
财政年份:2018
-
负责人:Nathan Sheffield
-
依托单位:
海外基金