Crowdsourcing biocuration: The Community Assessment of Community Annotation with Ontologies (CACAO).

Crowdsourcing biocuration: The Community Assessment of Community Annotation with Ontologies (CACAO).
复制标题

DOI:
10.1371/journal.pcbi.1009463
复制
发表时间:
2021-10
影响因子:
4.3
通讯作者:
Hu JC
Hu JC
中科院分区:
生物学2区
文献类型:
--
作者:
Ramsey J;McIntosh B;Renfro D;Aleksander SA;LaBonte S;Ross C;Zweifel AE;Liles N;Farrar S;Gill JJ;Erill I;Ades S;Berardini TZ;Bennett JA;Brady S;Britton R;Carbon S;Caruso SM;Clements D;Dalia R;Defelice M;Doyle EL;Friedberg I;Gurney SMR;Hughes L;Johnson A;Kowalski JM;Li D;Lovering RC;Mans TL;McCarthy F;Moore SD;Murphy R;Paustian TD;Perdue S;Peterson CN;Prüß BM;Saha MS;Sheehy RR;Tansey JT;Temple L;Thorman AW;Trevino S;Vollmer AC;Walbot V;Willey J;Siegele DA;Hu JC

文献摘要

参考文献

被引文献

相似文献

从原始文献中收集的关于基因功能的实验数据对研究科学家理解生物学具有巨大价值。使用基因本体论(GO),由专家手工策展提供了一个重要的资源,研究基因的功能,特别是在模式生物。科学文献的空前扩展和预测蛋白质的验证增加了数据价值和跟上步伐的挑战。捕获基于文献的功能注释受到生物制造者处理大量且快速增长的科学文献的能力的限制。在面向社区的wiki框架GO注释称为基因本体正常使用跟踪系统(GONUTS),我们描述了一种方法来扩大biocuration通过众包与本科生。这使国际数据库中高质量注释的数量成倍增加,丰富了我们对正常基因功能文献的覆盖范围,并将该领域推向新的方向。从一个由经验丰富的生物学家评判的校际比赛,社区评估社区注释与本体(CACAO),我们已经贡献了近5,000个基于文献的注释。这些注释中的许多是目前在GO中没有很好代表的生物体。在10年的历史中,我们的社区贡献者已经推动了传统上不被专业生物学家覆盖的本体的变化。CACAO的原则是依靠社区成员参与并塑造GO中生物管理的未来,这是一个用于促进科学事业的强大且可扩展的模型。它还为本科生提供了一个独特的和丰富的介绍,主要文献的批判性阅读和市场技能的收购。主要的科学文献以人类可读的格式对公共资助的基因功能科学研究的结果进行了编目。从这些研究中捕获的信息以广泛采用的机器可读标准格式以基因本体论(GO)注释的形式出现,这些注释涉及生命所有领域的基因功能。基于直接来自科学文献的推论的手动注释,包括用于做出此类推论的证据,通过改善整个生物科学的数据可访问性并允许进化相关生物之间的新见解,代表了最佳投资回报。为了补充专业的策展,我们的社区本体注释社区评估(CACAO)项目使社区注释者能够注释科学文献,在这种情况下是本科生,这导致了数千个独特的,经过验证的条目对公共资源的贡献。重要的是,这里描述的由非专家发起的注释通常涉及专家通常不涉及的主题。这些注释现在被世界各地的科学家用于他们的研究工作。
Experimental data about gene functions curated from the primary literature have enormous value for research scientists in understanding biology. Using the Gene Ontology (GO), manual curation by experts has provided an important resource for studying gene function, especially within model organisms. Unprecedented expansion of the scientific literature and validation of the predicted proteins have increased both data value and the challenges of keeping pace. Capturing literature-based functional annotations is limited by the ability of biocurators to handle the massive and rapidly growing scientific literature. Within the community-oriented wiki framework for GO annotation called the Gene Ontology Normal Usage Tracking System (GONUTS), we describe an approach to expand biocuration through crowdsourcing with undergraduates. This multiplies the number of high-quality annotations in international databases, enriches our coverage of the literature on normal gene function, and pushes the field in new directions. From an intercollegiate competition judged by experienced biocurators, Community Assessment of Community Annotation with Ontologies (CACAO), we have contributed nearly 5,000 literature-based annotations. Many of those annotations are to organisms not currently well-represented within GO. Over a 10-year history, our community contributors have spurred changes to the ontology not traditionally covered by professional biocurators. The CACAO principle of relying on community members to participate in and shape the future of biocuration in GO is a powerful and scalable model used to promote the scientific enterprise. It also provides undergraduate students with a unique and enriching introduction to critical reading of primary literature and acquisition of marketable skills. The primary scientific literature catalogs the results from publicly funded scientific research about gene function in human-readable format. Information captured from those studies in a widely adopted, machine-readable standard format comes in the form of Gene Ontology (GO) annotations about gene functions from all domains of life. Manual annotations based on inferences directly from the scientific literature, including the evidence used to make such inferences, represent the best return on investment by improving data accessibility across the biological sciences and allowing novel insights between evolutionarily related organisms. To supplement professional curation, our Community Assessment of Community Annotation with Ontologies (CACAO) project enabled annotation of the scientific literature by community annotators, in this case undergraduates, which resulted in the contribution of thousands of unique, validated entries to public resources. Importantly, the annotations described here initiated by nonexperts often deal with topics not typically covered by the experts. These annotations are now being used by scientists worldwide in their research efforts.
DOI: 10.1093/nar/gkaa1113
发表时间: 2021-01-08
影响因子: 14.9
作者:
Gene Ontology Consortium
通讯作者: Gene Ontology Consortium
DOI: 10.1093/nar/gky1055
发表时间: 2019-01-08
影响因子: 14.9
作者:
The Gene Ontology Consortium
通讯作者: The Gene Ontology Consortium
Biopython:用于计算分子生物学和生物信息学的免费 Python 工具。
DOI: 10.1093/bioinformatics/btp163
发表时间: 2009-06-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Cock PJ;Antao T;Chang JT;Chapman BA;Cox CJ;Dalke A;Friedberg I;Hamelryck T;Kauff F;Wilczynski B;de Hoon MJ
通讯作者: de Hoon MJ
DOI: 10.1371/journal.pbio.2002846
发表时间: 2018-04-01
期刊: PLOS BIOLOGY
影响因子: 9.8
作者:
Ammari, Mais;Aryamontri, Andrew Chatr;Wood, Valerie
通讯作者: Wood, Valerie
DOI: 10.1093/database/bas045
发表时间: 2012
期刊: Database : the journal of biological databases and curation
影响因子: --
作者:
Drabkin HJ;Blake JA;Mouse Genome Informatics Database
通讯作者: Mouse Genome Informatics Database