Moving the mountain: analysis of the effort required to transform comparative anatomy into computable anatomy.

Moving the mountain: analysis of the effort required to transform comparative anatomy into computable anatomy.
复制标题

移动山:分析将比较解剖结构转化为可计算解剖结构所需的努力。

DOI:
10.1093/database/bav040
复制
发表时间:
2015
期刊:
Database : the journal of biological databases and curation
影响因子:
--
通讯作者:
Mabee P
Mabee P
中科院分区:
其他
文献类型:
--
作者:
Dahdul W;Dececchi TA;Ibrahim N;Lapp H;Mabee P

文献摘要

参考文献

被引文献

相似文献

生物体的各种表型已经描述了几个世纪,虽然它们可以被数字化,但它们并不容易以可计算的形式提供。Phenoscape项目使用了100多项形态学研究,证明了通过用群落本体术语注释特征,可以在新物种解剖学和可能构成它们基础的基因之间建立联系。但是,考虑到遗留文献的复杂性,如何才能使这些基本上未被利用的描述性数据财富适合大规模计算?为了确定瓶颈,我们量化了在表型策展的主要方面所涉及的时间,因为我们从脊椎动物系统发生学文献中注释字符。这涉及将由本体术语组成的完全可计算的逻辑表达式附加到逐个分类单元矩阵中的描述。工作流程包括:(i)数据准备,(ii)表型注释,(iii)本体开发和(iv)策展团队讨论和软件开发反馈。我们的研究结果表明,完成这项工作需要两个博士后,一个首席数据管理员和学生的团队两年的时间。手动数据准备需要近13%的工作量。尤其是这一部分可以通过更好的社区数据实践来大幅减少,例如将完全填充的矩阵存放在公共存储库中。表型注释需要花费40%的精力。我们正在努力通过自然语言处理工具使其更有效。然而,本体论开发(40%)仍然是一项高度手动的任务,需要领域(解剖学)的专业知识和使用专门的软件。数据准备和本体开发所需的大量开销导致注释速率较低,约为每小时两个字符,而当活动仅限于字符注释时,则为每小时14个字符。要释放大量形态学描述的潜力,需要更好的工具来有效地处理自然语言,并需要更好的社区实践来实现数字形态学。数据库URL:http://kb.phenoscape.org
The diverse phenotypes of living organisms have been described for centuries, and though they may be digitized, they are not readily available in a computable form. Using over 100 morphological studies, the Phenoscape project has demonstrated that by annotating characters with community ontology terms, links between novel species anatomy and the genes that may underlie them can be made. But given the enormity of the legacy literature, how can this largely unexploited wealth of descriptive data be rendered amenable to large-scale computation? To identify the bottlenecks, we quantified the time involved in the major aspects of phenotype curation as we annotated characters from the vertebrate phylogenetic systematics literature. This involves attaching fully computable logical expressions consisting of ontology terms to the descriptions in character-by-taxon matrices. The workflow consists of: (i) data preparation, (ii) phenotype annotation, (iii) ontology development and (iv) curation team discussions and software development feedback. Our results showed that the completion of this work required two person-years by a team of two post-docs, a lead data curator, and students. Manual data preparation required close to 13% of the effort. This part in particular could be reduced substantially with better community data practices, such as depositing fully populated matrices in public repositories. Phenotype annotation required ∼40% of the effort. We are working to make this more efficient with Natural Language Processing tools. Ontology development (40%), however, remains a highly manual task requiring domain (anatomical) expertise and use of specialized software. The large overhead required for data preparation and ontology development contributed to a low annotation rate of approximately two characters per hour, compared with 14 characters per hour when activity was restricted to character annotation. Unlocking the potential of the vast stores of morphological descriptions requires better tools for efficiently processing natural language, and better community practices towards a born-digital morphology. Database URL: http://kb.phenoscape.org
DOI: 10.1186/s12859-015-0488-1
发表时间: 2015-02-15
期刊: BMC bioinformatics
影响因子: 3
作者:
Huang F;Macklin JA;Cui H;Cole HA;Endara L
通讯作者: Endara L
DOI: 10.1186/1471-2105-10-228
发表时间: 2009-07-21
期刊: BMC bioinformatics
影响因子: 3
作者:
Van Auken K;Jaffery J;Chan J;Müller HM;Sternberg PW
通讯作者: Sternberg PW
DOI: 10.1186/2041-1480-5-34
发表时间: 2014
影响因子: 1.9
作者:
Dahdul WM;Cui H;Mabee PM;Mungall CJ;Osumi-Sutherland D;Walls RL;Haendel MA
通讯作者: Haendel MA
DOI: 10.4202/app.2010.0101
发表时间: 2012-03-01
影响因子: 1.8
作者:
Skutschas, Pavel P.;Gubin, Yuri M.
通讯作者: Gubin, Yuri M.
DOI: 10.1371/journal.pone.0018657
发表时间: 2011
期刊: PloS one
影响因子: 3.7
作者:
Piwowar HA
通讯作者: Piwowar HA