Structural and functional-annotation of an equine whole genome oligoarray

Structural and functional-annotation of an equine whole genome oligoarray
复制标题

DOI:
10.1186/1471-2105-10-s11-s8
复制
发表时间:
2009-01-01
期刊:
影响因子:
3
通讯作者:
McCarthy, Fiona M.
McCarthy, Fiona M.
中科院分区:
生物学4区
文献类型:
--
作者:
Bright, Lauren A.;Burgess, Shane C.;McCarthy, Fiona M.

文献摘要

被引文献

相似文献

背景资料:马的基因组测序,使马的研究人员能够使用高通量功能基因组学平台,如微阵列;下一代基因表达和蛋白质组测序。然而,对于研究人员来说,要从这些功能基因组学数据集中获得价值,他们必须能够以生物学相关的方式对这些数据进行建模;要做到这一点,需要对马的基因组进行更全面的注释。有两种相互关联的基因组注释类型:结构和功能。结构注释是对基因组元件(如基因、启动子和调控元件)进行描绘和划分。功能注释是指将功能分配给结构元素。基因本体论(GO)是事实上的标准功能注释,并经常被用作建模和假设检验的基础,大型功能基因组dataset.Results:马全基因组寡核苷酸(EWGO)阵列与21,351元素在德克萨斯州A&M大学开发。使用马基因组的约7 x组装和注释序列设计该70聚体寡核苷酸阵列,以成为可用于表达马序列的最全面的阵列之一。为了帮助研究人员确定来自该阵列的数据的生物学意义,我们通过将元素映射到多个数据库登录来对其进行结构注释,这些数据库包括UniProtKB,NRPD(非冗余蛋白质数据库)和UniGene。我们接下来提供了该阵列上代表的基因转录物的GO功能注释。总的来说,我们GO注释了14,531个基因产物(EWGO阵列上代表的基因产物的68.1%),其中注释57,912个。在我们添加GO注释之前和之后,计算该阵列的GAQ(GO注释质量)分数。额外的注释将平均GAQ评分提高了16倍。这些数据可在AgBase http://www.agbase.msstate.edu/.Conclusion上公开获得:提供有关链接到阵列上所代表的基因产物的公共数据库的额外信息,使用户在使用基因表达建模和假设检验计算工具时具有更大的灵活性。此外,由于不同的数据库提供不同类型的信息,用户可以访问多种数据源。此外,我们的GO注释支持大多数基因表达分析工具的功能建模,并使马研究人员能够以生物学相关的方式对大量差异表达的转录本进行建模。
Background: The horse genome is sequenced, allowing equine researchers to use high-throughput functional genomics platforms such as microarrays; next-generation sequencing for gene expression and proteomics. However, for researchers to derive value from these functional genomics datasets, they must be able to model this data in biologically relevant ways; to do so requires that the equine genome be more fully annotated. There are two interrelated types of genomic annotation: structural and functional. Structural annotation is delineating and demarcating the genomic elements (such as genes, promoters, and regulatory elements). Functional annotation is assigning function to structural elements. The Gene Ontology (GO) is the de facto standard for functional annotation, and is routinely used as a basis for modelling and hypothesis testing, large functional genomics datasets.Results: An Equine Whole Genome Oligonucleotide (EWGO) array with 21,351 elements was developed at Texas A&M University. This 70-mer oligoarray was designed using the approximately 7x assembled and annotated sequence of the equine genome to be one of the most comprehensive arrays available for expressed equine sequences. To assist researchers in determining the biological meaning of data derived from this array, we have structurally annotated it by mapping the elements to multiple database accessions, including UniProtKB, Entrez Gene, NRPD (Non-Redundant Protein Database) and UniGene. We next provided GO functional annotations for the gene transcripts represented on this array. Overall, we GO annotated 14,531 gene products (68.1% of the gene products represented on the EWGO array) with 57,912 annotations. GAQ (GO Annotation Quality) scores were calculated for this array both before and after we added GO annotation. The additional annotations improved the meanGAQ score 16-fold. This data is publicly available at AgBase http://www.agbase.msstate.edu/.Conclusion: Providing additional information about the public databases which link to the gene products represented on the array allows users more flexibility when using gene expression modelling and hypothesis-testing computational tools. Moreover, since different databases provide different types of information, users have access to multiple data sources. In addition, our GO annotation underpins functional modelling for most gene expression analysis tools and enables equine researchers to model large lists of differentially expressed transcripts in biologically relevant ways.