A Chado case study: an ontology-based modular schema for representing genome-associated biological information

A Chado case study: an ontology-based modular schema for representing genome-associated biological information
复制标题

DOI:
10.1093/bioinformatics/btm189
复制
发表时间:
2007-07-01
期刊:
影响因子:
5.8
通讯作者:
Emmert, David B.
Emmert, David B.
中科院分区:
生物学3区
文献类型:
--
作者:
Mungall, Christopher J.;Emmert, David B.

文献摘要

被引文献

相似文献

动机:几年前,FlyBase着手设计一个新的数据库模式来存储果蝇数据。它将完全整合基因组序列和注释数据与文献中的书目,遗传,表型和分子数据,这些数据代表了对这一主要动物模型系统前100年研究的精华。在开发这种新的集成模式时,FlyBase还承诺确保其设计是通用的,可扩展的,并且可以作为开源,以便它可以用作任何模式生物数据库的核心模式,从而避免冗余的软件开发和潜在的互操作性。我们的问题是,我们是否可以创建一个关系数据库模式,将成功reuse.Results:查多是一个关系数据库模式,现在被用来管理生物知识的各种生物体,从人类到病原体,特别是类的信息,直接或间接地可以与基因组序列或主要的RNA和蛋白质产品编码的基因组。符合此模式的生物数据库可以相互操作,并与通用模型生物数据库(GMOD)工具包中的应用软件进行互操作。Chado的独特之处在于它的设计是由本体驱动的。本体(或受控词汇表)的使用在模式中无处不在,因为它们被用作实体类型化的一种手段。Chado模式被划分为集成的子模式(模块),每个子模式封装了不同的生物学领域,并且每个子模式使用适当的本体中的表示来描述。为了说明这种方法,我们在这里描述用于描述基因组序列的Chado模块。
Motivation: A few years ago, FlyBase undertook to design a new database schema to store Drosophila data. It would fully integrate genomic sequence and annotation data with bibliographic, genetic, phenotypic and molecular data from the literature representing a distillation of the first 100 years of research on this major animal model system. In developing this new integrated schema, FlyBase also made a commitment to ensure that its design was generic, extensible and available as open source, so that it could be employed as the core schema of any model organism data repository, thereby avoiding redundant software development and potentially increasing interoperability. Our question was whether we could create a relational database schema that would be successfully reused.Results: Chado is a relational database schema now being used to manage biological knowledge for a wide variety of organisms, from human to pathogens, especially the classes of information that directly or indirectly can be associated with genome sequences or the primary RNA and protein products encoded by a genome. Biological databases that conform to this schema can interoperate with one another, and with application software from the Generic Model Organism Database (GMOD) toolkit. Chado is distinctive because its design is driven by ontologies. The use of ontologies ( or controlled vocabularies) is ubiquitous across the schema, as they are used as a means of typing entities. The Chado schema is partitioned into integrated subschemas ( modules), each encapsulating a different biological domain, and each described using representations in appropriate ontologies. To illustrate this methodology, we describe here the Chado modules used for describing genomic sequences.