Practical application of ontologies to annotate and analyse large scale raw mouse phenotype data.

Practical application of ontologies to annotate and analyse large scale raw mouse phenotype data.
复制标题

本体论的实际应用在注释和分析大规模的原始小鼠表型数据。

DOI:
10.1186/1471-2105-10-s5-s2
复制
发表时间:
2009-05-06
期刊:
影响因子:
3
通讯作者:
Mallon AM
Mallon AM
中科院分区:
生物学4区
文献类型:
--
作者:
Beck T;Morgan H;Blake A;Wells S;Hancock JM;Mallon AM

文献摘要

被引文献

相似文献

大规模的国际项目正在进行中,以产生敲除小鼠突变体的集合,并随后进行高通量表型评估,由于表型数据的复杂性和规模,给计算研究人员带来了新的挑战。表型可以使用本体以两种不同的方法来描述。传统上,单个表型特征要么使用源自物种特异性专用表型本体的单个复合术语来定义,要么使用来自一系列不同本体的概念通过组合注释来定义表型特征作为具有相关品质(EQ)的实体。两种方法都有其优点,其中包括允许使用社区标准术语的专用方法,以及促进跨物种表型陈述比较的组合方法。以前数据库更喜欢一种方法而不是另一种方法。 EUMODIC 项目将生成大量的小鼠表型数据,这些数据是由于执行一组标准操作程序 (SOP) 而生成的,并将实施两种本体方法来捕获生成的表型数据。对于所有 SOP,都进行了四层注释: SOP 的高级描述,广泛定义 SOP 生成的数据类型;使用EQ模型的单个参数注释;对每只小鼠生成的定性数据进行注释;以及统计分析后突变系的注释。表型偏差的定性评估是在数据输入时使用子 PATO 质量作为参数质量进行的。为了方便更熟悉描述表型的单一复合术语的科学家进行数据查询,利用哺乳动物表型 (MP) 本体和 EQ PATO 模型之间的映射来允许通过 MP 术语进行查询。注释良好且可比较的表型数据库可以通过使用本体论派生的可比较表型语句来实现,并且已经通过 OBO 兼容的 EQ 注释来实现。我们描述的实现还让科学家通过 PATO 质量评估定性表型以及使用社区接受的复合 MP 术语查询数据库的能力与本体无缝合作。这项工作代表了第一次使用组合和单一专用方法来注释表型数据集。
Large-scale international projects are underway to generate collections of knockout mouse mutants and subsequently to perform high throughput phenotype assessments, raising new challenges for computational researchers due to the complexity and scale of the phenotype data. Phenotypes can be described using ontologies in two differing methodologies. Traditionally an individual phenotypic character has either been defined using a single compound term, originating from a species-specific dedicated phenotype ontology, or alternatively by a combinatorial annotation, using concepts from a range of disparate ontologies, to define a phenotypic character as an entity with an associated quality (EQ). Both methods have their merits, which include the dedicated approach allowing use of community standard terminology, and the combinatorial approach facilitating cross-species phenotypic statement comparisons. Previously databases have favoured one approach over another. The EUMODIC project will generate large amounts of mouse phenotype data, generated as a result of the execution of a set of Standard Operating Procedures (SOPs) and will implement both ontological approaches to capture the phenotype data generated. For all SOPs a four-tier annotation is made: a high-level description of the SOP, to broadly define the type of data generated by the SOP; individual parameter annotation using the EQ model; annotation of the qualitative data generated for each mouse; and the annotation of mutant lines after statistical analysis. The qualitative assessments of phenodeviance are made at the point of data entry, using child PATO qualities to the parameter quality. To facilitate data querying by scientists more familiar with single compound terms to describe phenotypes, the mappings between the Mammalian Phenotype (MP) ontology and the EQ PATO model are exploited to allow querying via MP terms. Well-annotated and comparable phenotype databases can be achieved through the use of ontologically derived comparable phenotypic statements and have been implemented here by means of OBO compatible EQ annotations. The implementation we describe also sees scientists working seamlessly with ontologies through the assessment of qualitative phenotypes in terms of PATO qualities and the ability to query the database using community-accepted compound MP terms. This work represents the first time the combinatorial and single-dedicated approaches have both been implemented to annotate a phenotypic dataset.