Comprehensive data infrastructure for plant bioinformatics

Comprehensive data infrastructure for plant bioinformatics
复制标题

植物生物信息学综合数据基础设施

DOI:
10.1109/clusterwksp.2010.5613093
复制
发表时间:
2010
期刊:
2010 IEEE International Conference On Cluster Computing Workshops and Posters (CLUSTER WORKSHOPS)
影响因子:
--
通讯作者:
C. Noutsos
C. Noutsos
中科院分区:
--
文献类型:
--
作者:
C. Jordan;D. Stanzione;D. Ware;Jerry Lu;C. Noutsos

文献摘要

被引文献

相似文献

IFactory Collaborative是一项由国家科学基金会资助的为期5年的努力,旨在发展网络基础设施,以应对植物科学中的一系列重大挑战。这些重大挑战中的第二个是从基因型到表型的项目,该项目寻求以基于网络的发现环境的形式提供工具,以了解从DNA到成熟植物的发育过程。解决这一挑战需要集成多种数据类型,这些数据类型可能以多种格式存储,具有不同级别的标准化。提供再现性需要详细的信息,记录数据的实验来源,以及在数据进入iFactory环境后应用于数据的计算转换。处理高通量测序和其他生物信息学数据的实验来源所涉及的大量数据,需要一个强大的基础设施来存储和重新使用大型数据对象。我们描述了目前计划为基因型到表型发现环境开发的工作流程,必须在环境中导入和操作的数据类型和格式,并描述了为在发现环境中表达和交换数据而开发的数据模型,以及为捕获实验源和数字转换描述而定义的来源模型。本课程介绍了与参考数据库交互的能力,不仅侧重于从此类数据源检索数据的能力,还侧重于使用iPLANT发现环境进一步填充这些重要资源的能力。还介绍了未来的活动以及它们将给iPlant Collaborative的数据基础设施带来的挑战。
The iPlant Collaborative is a 5-year, National Science Foundation-funded effort to develop cyberinfrastructure to address a series of grand challenges in plant science. The second of these grand challenges is the Genotype-to-Phenotype project, which seeks to provide tools, in the form of a web-based Discovery Environment, for understanding the developmental process from DNA to a full-grown plant. Addressing this challenge requires the integration of multiple data types that may be stored in multiple formats, with varying levels of standardization. Providing for reproducibility requires that detailed information documenting the experimental provenance of data, and the computational transformations applied to data once it is brought into the iPlant environment. Handling the large quantities of data involved in high-throughput sequencing and other experimental sources of bioinformatics data requires a robust infrastructure for storing and reusing large data objects. We describe the currently planned workflows to be developed for the Genotype-to-Phenotype discovery environment, the data types and formats that must be imported and manipulated within the environment, and we describe the data model that has been developed to express and exchange data within the Discovery Environment, along with the provenance model defined for capturing experimental source and digital transformation descriptions. Capabilities for interaction with reference databases are addressed, focusing not just on the ability to retrieve data from such data sources, but on the ability to use the iPlant Discovery Environment to further populate these important resources. Future activities and the challenges they will present to the data infrastructure of the iPlant Collaborative are also described.