Repositories for Taxonomic Data: Where We Are and What is Missing.

Repositories for Taxonomic Data: Where We Are and What is Missing.
复制标题

DOI:
10.1093/sysbio/syaa026
复制
发表时间:
2020-11-01
期刊:
影响因子:
6.5
通讯作者:
Vences M
Vences M
中科院分区:
生物学1区
文献类型:
--
作者:
Miralles A;Bruy T;Wolcott K;Scherz MD;Begerow D;Beszteri B;Bonkowski M;Felden J;Gemeinholzer B;Glaw F;Glöckner FO;Hawlitschek O;Kostadinov I;Nattkemper TW;Printzen C;Renz J;Rybalka N;Stadler M;Weibulat T;Wilke T;Renner SS;Vences M

文献摘要

参考文献

被引文献

相似文献

自然历史收藏正在引领标本数字化(图像、元数据、DNA条形码)的大规模成功项目,从而将分类学转变为大数据科学。然而,在保护和随后调动每年在命名15 000至20 000个物种的过程中产生的大量原始数据方面,几乎没有作出任何努力。从阿尔法分类学家的角度来看,我们提供了一个审查的属性和多样性的分类数据,评估其数量和使用,并建立优化数据存储库的标准。我们调查了2002年,2010年和2018年代表性期刊上的4113项α分类学研究,发现分子数据在物种诊断和描述中的使用越来越多,但相对有限。2018年,在专业分类学期刊上发表的2661篇论文中,分子数据广泛应用于真菌学(94%),经常应用于脊椎动物(53%),但很少应用于植物学(15%)和昆虫学(10%)。图像在所有分类群的分类研究中发挥着重要作用,在所调查的论文中,超过80%的论文使用了照片,58%的论文使用了图画。组学(高通量)方法或3D文档的使用仍然很少。改进元条形码一致性读数、基因组和转录组组装以及化学和代谢组学数据的存档策略,可以帮助动员大量高通量数据用于α-分类学。由于长期-理想情况下永久-数据存储对分类学特别重要,如果其信息内容足以用于分类学研究,则通过存储要求较低的格式减少能源足迹是优先事项。虽然分类分配是大多数生物学科的准事实,但它们仍然是关于α-分类学个体进化相关性的假设。出于这个原因,改进分类数据的重用,包括基于机器学习的物种识别和划界管道,需要一种网络标本方法-通过唯一的标本标识符链接数据,从而使它们可查找,可访问,可互操作和可重复使用的分类研究。这既带来了使现有数据中心基础设施适应以数据中心为中心的概念的定性挑战,也带来了托管和连接阿尔法分类研究每年产生的估计200万张图像以及来自数字化活动的数百万张图像的定量挑战。在全球30,000 - 40,000名分类学家中,许多人被认为是非专业人员,因此捕获数据进行在线存储和重用需要低复杂性的提交工作流程和免费的存储库使用。专家分类学家是主要的利益相关者,能够确定和规范化的学科的需求,他们的专业知识是需要实施设想的虚拟收集的网络标本。[Big数据;网络标本;新物种;组学;资料库;标本标识符;分类学;分类学数据。]
Natural history collections are leading successful large-scale projects of specimen digitization (images, metadata, DNA barcodes), thereby transforming taxonomy into a big data science. Yet, little effort has been directed towards safeguarding and subsequently mobilizing the considerable amount of original data generated during the process of naming 15,000–20,000 species every year. From the perspective of alpha-taxonomists, we provide a review of the properties and diversity of taxonomic data, assess their volume and use, and establish criteria for optimizing data repositories. We surveyed 4113 alpha-taxonomic studies in representative journals for 2002, 2010, and 2018, and found an increasing yet comparatively limited use of molecular data in species diagnosis and description. In 2018, of the 2661 papers published in specialized taxonomic journals, molecular data were widely used in mycology (94%), regularly in vertebrates (53%), but rarely in botany (15%) and entomology (10%). Images play an important role in taxonomic research on all taxa, with photographs used in >80% and drawings in 58% of the surveyed papers. The use of omics (high-throughput) approaches or 3D documentation is still rare. Improved archiving strategies for metabarcoding consensus reads, genome and transcriptome assemblies, and chemical and metabolomic data could help to mobilize the wealth of high-throughput data for alpha-taxonomy. Because long-term—ideally perpetual—data storage is of particular importance for taxonomy, energy footprint reduction via less storage-demanding formats is a priority if their information content suffices for the purpose of taxonomic studies. Whereas taxonomic assignments are quasifacts for most biological disciplines, they remain hypotheses pertaining to evolutionary relatedness of individuals for alpha-taxonomy. For this reason, an improved reuse of taxonomic data, including machine-learning-based species identification and delimitation pipelines, requires a cyberspecimen approach—linking data via unique specimen identifiers, and thereby making them findable, accessible, interoperable, and reusable for taxonomic research. This poses both qualitative challenges to adapt the existing infrastructure of data centers to a specimen-centered concept and quantitative challenges to host and connect an estimated 2 million images produced per year by alpha-taxonomic studies, plus many millions of images from digitization campaigns. Of the 30,000–40,000 taxonomists globally, many are thought to be nonprofessionals, and capturing the data for online storage and reuse therefore requires low-complexity submission workflows and cost-free repository use. Expert taxonomists are the main stakeholders able to identify and formalize the needs of the discipline; their expertise is needed to implement the envisioned virtual collections of cyberspecimens. [Big data; cyberspecimen; new species; omics; repositories; specimen identifier; taxonomy; taxonomic data.]
DOI: 10.1093/nar/gkr1178
发表时间: 2012-01
影响因子: 14.9
作者:
Federhen S
通讯作者: Federhen S
DOI: 10.1371/journal.pbio.2002231
发表时间: 2017-08
期刊: PLoS biology
影响因子: 9.8
作者:
Bik HM
通讯作者: Bik HM
DOI: 10.3897/zookeys.209.3571
发表时间: 2012
期刊: ZooKeys
影响因子: 1.3
作者:
Dietrich C;Hart J;Raila D;Ravaioli U;Sobh N;Sobh O;Taylor C
通讯作者: Taylor C
DOI: 10.1080/10635150701701083
发表时间: 2007-12-01
期刊: SYSTEMATIC BIOLOGY
影响因子: 6.5
作者:
De Queiroz, Kevin
通讯作者: De Queiroz, Kevin
DOI: 10.11646/zootaxa.4196.3.9
发表时间: 2016-11-23
期刊: ZOOTAXA
影响因子: 0.9
作者:
Ceriaco, Luis M. P.;Gutierrez, Eliecer E.;Zug, George
通讯作者: Zug, George