Data management pipeline for plant phenotyping in a multisite project

Data management pipeline for plant phenotyping in a multisite project
复制标题

DOI:
10.1071/fp12009
复制
发表时间:
2012-01-01
影响因子:
3
通讯作者:
Koehl, Karin I.
Koehl, Karin I.
中科院分区:
生物学4区
文献类型:
--
作者:
Billiau, Kenny;Sprenger, Heike;Koehl, Karin I.

文献摘要

被引文献

相似文献

在植物育种中,不同的人员必须在规定的时间跨度内在多个田间地点精确、一致且快速地表征植物。为了进行有意义的数据评估和统计分析,需要标准化的数据存储。数据访问必须长期提供,并且不受组织障碍的影响,且不会危及数据完整性或知识产权。我们讨论了相关的技术挑战,并展示了在识别马铃薯耐旱标记项目的数据管理管道中举例说明的适当解决方案。该项目涉及来自学术界和育种公司的11个小组、11个站点和4个分析平台。我们的数据仓库概念结合了数据库和文件服务器中的中央数据存储,并将针对特定数据类型的现有和专用数据库解决方案与新的、特定于项目的数据库集成。事实证明,严格使用受控词汇和应用网络访问技术对于不同机构之间以及数据管理概念和基础设施之间的成功数据交换至关重要。通过展示我们的数据管理系统并提供软件,我们的目标是支持相关的表型项目。
In plant breeding, plants have to be characterised precisely, consistently and rapidly by different people at several field sites within defined time spans. For a meaningful data evaluation and statistical analysis, standardised data storage is required. Data access must be provided on a long-term basis and be independent of organisational barriers without endangering data integrity or intellectual property rights. We discuss the associated technical challenges and demonstrate adequate solutions exemplified in a data management pipeline for a project to identify markers for drought tolerance in potato. This project involves 11 groups from academia and breeding companies, 11 sites and four analytical platforms. Our data warehouse concept combines central data storage in databases and a file server and integrates existing and specialised database solutions for particular data types with new, project-specific databases. The strict use of controlled vocabularies and the application of web-access technologies proved vital to the successful data exchange between diverse institutes and data management concepts and infrastructures. By presenting our data management system and making the software available, we aim to support related phenotyping projects.