Arabidopsis bioinformatics resources: The current state, challenges, and priorities for the future

Arabidopsis bioinformatics resources: The current state, challenges, and priorities for the future
复制标题

DOI:
10.1002/pld3.109
复制
发表时间:
2019-01-01
期刊:
影响因子:
3
通讯作者:
Wurtele, Eve
Wurtele, Eve
中科院分区:
生物学3区
文献类型:
--
作者:
Doherty, Colleen;Friesner, Joanna;Wurtele, Eve

文献摘要

被引文献

相似文献

有效的研究,教育和推广工作的拟南芥社区,以及其他科学界依赖于拟南芥资源,非常依赖于容易获得和公共共享的资源。这些资源包括参考基因组序列数据和不断增加的各种数据集和数据类型。TAIR(拟南芥信息资源)和Araport(最初命名为拟南芥信息门户)是社区信息资源,为全球30,000多名研究人员提供工具,数据和应用程序,这些研究人员在工作中使用拟南芥作为主要研究系统或来自拟南芥的数据。在Araport成立四年后,IAIC举办了另一次研讨会,评估拟南芥信息学的现状,并为未来的研究和发展制定路线。研讨会重点讨论了若干挑战,包括需要可靠和最新的注释、社区定义的数据和元数据共同标准,以及便于使用和方便用户的数据整合和可视化储存库/工具/方法。设想的解决方案包括:(a)一个集中的注释权威机构来合并来自新组的注释,建立一致的命名方案,定期和频繁地分发这种格式,并鼓励和强制采用。(b)数据和元数据格式的标准,这是必不可少的,但在不同基因型之间进行比较时以及在标准不太确定的领域(例如,表型组学、代谢组学)。需要制定社区制定的准则。(c)可搜索的中央存储库,用于分析和可视化工具。改进的版本控制和用户访问将使工具更易于访问。研讨会与会者提议建立一个“一站式”网站,即一个拟南芥“超级门户”,将各种工具、数据资源、方案标准和每一数据类型的最佳做法说明联系起来。这必须得到社区的支持和参与,以鼓励采用。
Effective research, education, and outreach efforts by the Arabidopsis thaliana community, as well as other scientific communities that depend on Arabidopsis resources, depend vitally on easily available and publicly-shared resources. These resources include reference genome sequence data and an ever-increasing number of diverse data sets and data types. TAIR (The Arabidopsis Information Resource) and Araport (originally named the Arabidopsis Information Portal) are community informatics resources that provide tools, data, and applications to the more than 30,000 researchers worldwide that use in their work either Arabidopsis as a primary system of study or data derived from Arabidopsis. Four years after Araport's establishment, the IAIC held another workshop to evaluate the current status of Arabidopsis Informatics and chart a course for future research and development. The workshop focused on several challenges, including the need for reliable and current annotation, community-defined common standards for data and metadata, and accessible and user-friendly repositories/tools/methods for data integration and visualization. Solutions envisioned included (a) a centralized annotation authority to coalesce annotation from new groups, establish a consistent naming scheme, distribute this format regularly and frequently, and encourage and enforce its adoption. (b) Standards for data and metadata formats, which are essential, but challenging when comparing across diverse genotypes and in areas with less-established standards (e.g., phenomics, metabolomics). Community-established guidelines need to be developed. (c) A searchable, central repository for analysis and visualization tools. Improved versioning and user access would make tools more accessible. Workshop participants proposed a "one-stop shop" website, an Arabidopsis "Super-Portal" to link tools, data resources, programmatic standards, and best practice descriptions for each data type. This must have community buy-in and participation in its establishment and development to encourage adoption.