Correlation between gene expression and GO semantic similarity

Correlation between gene expression and GO semantic similarity
复制标题

DOI:
10.1109/tcbb.2005.50
复制
发表时间:
2005-10-01
影响因子:
4.5
通讯作者:
Rubio, A
Rubio, A
中科院分区:
工程技术3区
文献类型:
--
作者:
Sevilla, JL;Segura, V;Rubio, A

文献摘要

被引文献

相似文献

本研究分析了基因表达、基因功能和基因注释之间的关系。最近的许多研究都隐含地基于这样的假设,即生物学和功能相关的基因产物将在其表达谱以及其基因本体(GO)注释中保持这种相似性。我们分析如何准确地证明这一假设是使用真实的公开可用的数据。我们的目标还包括验证GO注释的语义相似性度量。我们使用Pearson相关系数及其绝对值作为基因产物表达谱之间相似性的度量。我们探索了一些语义相似性度量(Resnik,Jiang和Lin),并计算使用GO注释的基因产物之间的相似性。最后,我们计算相关系数来比较基因表达相似性与GO语义相似性。我们的研究结果表明,Resnik相似性度量优于其他,似乎更适合用于基因本体。我们还推断,在GO注释和基因表达的三个GO本体的语义相似性之间似乎有相关性。我们表明,这种相关性是可以忽略不计的一定的语义相似性值,然后,对于更高的相似性值,关系的趋势变得几乎线性。这些结果可用于增强聚类算法提供的知识,并在生物信息学工具的发展,寻找和表征基因产物。
This research analyzes some aspects of the relationship between gene expression, gene function, and gene annotation. Many recent studies are implicitly based on the assumption that gene products that are biologically and functionally related would maintain this similarity both in their expression profiles as well as in their Gene Ontology (GO) annotation. We analyze how accurate this assumption proves to be using real publicly available data. We also aim to validate a measure of semantic similarity for GO annotation. We use the Pearson correlation coefficient and its absolute value as a measure of similarity between expression profiles of gene products. We explore a number of semantic similarity measures (Resnik, Jiang, and Lin) and compute the similarity between gene products annotated using the GO. Finally, we compute correlation coefficients to compare gene expression similarity against GO semantic similarity. Our results suggest that the Resnik similarity measure outperforms the others and seems better suited for use in Gene Ontology. We also deduce that there seems to be correlation between semantic similarity in the GO annotation and gene expression for the three GO ontologies. We show that this correlation is negligible up to a certain semantic similarity value; then, for higher similarity values, the relationship trend becomes almost linear. These results can be used to augment the knowledge provided by clustering algorithms and in the development of bioinformatic tools for finding and characterizing gene products.