A Comprehensive and Standards-Aware Common Data Model (CDM) for Taxonomic Research

A Comprehensive and Standards-Aware Common Data Model (CDM) for Taxonomic Research
复制标题

用于分类学研究的全面且具有标准意识的通用数据模型 (CDM)

DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
K. Luther
K. Luther
中科院分区:
--
文献类型:
--
作者:
Andreas Müller;W. Berendsohn;A. Kohlbecker;A. Güntsch;Patrick Plitzner;K. Luther

文献摘要

被引文献

相似文献

编辑公共数据模型(CDM) (FUB, BGBM 2008)是网络分类学编辑平台的核心(FUB, BGBM 2011, Ciardelli et al. 2009)。它建立在可追溯到20世纪90年代的建模工作的基础上,旨在将与分类学领域相关的现有标准(但通常是为数据交换而设计的)与现代分类学工具的要求结合起来。它以统一建模语言(UML) (Booch et al. 2005)建模,提供了一个面向对象的信息域视图,由专家分类学家管理,独立于使用的操作系统和数据库管理系统(DBMS)实现。在过去的十年中,该模型被用于不同重点的各种国家和国际研究项目,并逐渐发展成为各种分类项目的共同基础,如植物群、动物群和清单(参见FUB, BGBM 2016,不同项目创建并公开提供的一些数据门户)。CDM严格地以分类学专家群体的需求为导向。在需求复杂的地方,它试图合理地反映它们,而不是通过(过度)简化引入歧义或减少功能。在可能简化的地方,它试图保持或变得简单。模型层面的简化由‡‡‡‡©m<e:1> ller A et al实现。这是一篇在知识共享署名许可(CC BY 4.0)条款下发布的开放获取文章,该许可允许在任何媒体上不受限制地使用、分发和复制,前提是要注明原作者和来源。通过约束实现业务规则,而不是通过类型化和子类化。用户界面级别的简化是通过大量定制选项实现的。作为各种应用程序类型和用例的通用模型,它可以被用户和开发人员适应和扩展。它结合使用静态和动态类型,既可以有效地处理复杂但定义良好的数据域,如分类分类和命名法,也可以处理定义较差的灵活域,如事实数据和描述性数据。此外,它还允许创建30多种用户定义的词汇表,如分类等级、命名状态、名称对名称关系、地理区域、存在状态等。通过使几乎所有数据的来源都可以详细引用,并提供数据谱系以追溯数据的根源,将重点放在良好的科学实践上。也很容易同时反映多种观点,例如不同的分类概念(Berendsohn 1995, Berendsohn & al.,本会议)或从不同区域植物群或动物群中获得的几种描述性处理。CDM试图全面覆盖分类学领域命名法、分类学(包括概念)、分类单元分布数据、各种描述性数据,包括涉及分类单元和/或标本的形态学数据、各种图像和多媒体数据,以及涵盖标本和标本衍生物直至DNA样本和序列的复杂系统(Kilian et al. 2015)。Stöver and m<e:1> ller 2015),这反映了分类学研究过程中知识积累的复杂性。在EDIT平台的上下文中,已经基于CDM和提供API和基于CDM的web服务接口的库开发了几个应用程序(参见Kohlbecker & al.和g<s:1> ntsch & al.,本会议)。在某些领域,CDM仍在发展,尽管基本结构已经存在,但应用程序开发的问题反馈到建模决策中。然而,“没有捷径”的建模方法在过去曾多次延迟了应用程序的开发,但现在它得到了回报:平台可以快速适应不同项目和分类专家不断变化的需求。
The EDIT Common Data Model (CDM) (FUB, BGBM 2008) is the centrepiece of the EDIT Platform for Cybertaxonomy (FUB, BGBM 2011, Ciardelli et al. 2009). Building on modelling efforts reaching back to the 1990ies, it aims to combine existing standards relevant to the taxonomic domain (but often designed for data exchange) with requirements of modern taxonomic tools. Modelled in the Unified Modelling Language (UML) (Booch et al. 2005), it offers an object oriented view on the information domain managed by expert taxonomists that is implemented independent of the used operating system and database management system (DBMS). Being used in various national and international research projects with diverse foci over the past decade, the model evolved and became the common base of a variety of taxonomic projects, such as floras, faunas and checklists (see FUB, BGBM 2016 for a number of data portals created and made publicly available by different projects). The CDM is strictly oriented towards the needs of the taxonomic experts community. Where requirements are complex it tries to reflect them reasonably rather than introducing ambiguity or reduced functionality via (over-)simplification. Where simplification is possible it tries to stay or become simple. Simplification on the model level is achieved by ‡ ‡ ‡ ‡ ‡ ‡ © Müller A et al. This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. implementing business rules via constraints rather than via typification and subclassing. Simplification on the user interface level is achieved by numerous options for customisation. Being used as a generic model for a variety of application types and use cases, it is adaptable and extendable by users and developers. It uses a combination of static and dynamic typification to allow both efficient handling of complex but well-defined data domains such as taxonomic classifications and nomenclature as well as less well-defined flexible domains like factual and descriptive data. Additionally it allows the creation of more than 30 types of user defined vocabularies such as those for taxonomic rank, nomenclatural status, name-to-name relationships, geographic area, presence status, etc. A strong focus is set on good scientific praxis by making the source of almost all data citable in detail and offering data lineage to trace data back to its roots. It is also easy to reflect multiple opinions in parallel, e.g. differing taxonomic concepts (Berendsohn 1995, Berendsohn & al., this session) or several descriptive treatments obtained from different regional floras or faunas. The CDM attempts to comprehensively cover the data used in the taxonomic domain nomenclature, taxonomy (including concepts), taxon distribution data, descriptive data of all kinds, including morphological data referring to taxa and/or specimens, images and multimedia data of various kinds, and a complex system covering specimens and specimen derivatives down to DNA samples and sequences (Kilian et al. 2015, Stöver and Müller 2015) that mirrors the complexity of knowledge accumulation in the taxonomic research process. In the context of the EDIT Platform, several applications have been developed based on the CDM and the library that provides the API and web Service interfaces based on the CDM (see Kohlbecker & al. and Güntsch & al., this session). In some areas the CDM is still evolving although the basic structures are present, questions of application development feed back into modelling decisions. However, a "no-shortcuts" approach to modelling has variously delayed application development in the past, but it now pays off: the Platform can rapidly adapt to changing requirements from different projects and taxonomic specialists.