Measures of co-expression for improved function prediction of long non-coding RNAs

Measures of co-expression for improved function prediction of long non-coding RNAs
复制标题

DOI:
10.1186/s12859-018-2546-y
复制
发表时间:
2018-12-19
期刊:
影响因子:
3
通讯作者:
Drablos, Finn
Drablos, Finn
中科院分区:
生物学4区
文献类型:
--
作者:
Ehsani, Rezvan;Drablos, Finn

文献摘要

被引文献

相似文献

背景在GENCODE项目中已经发现了大约16,000个人类长非编码RNA(LncRNA)基因。然而,它们中的大多数的功能还有待于发现。通过在已经注释的基因中识别与lncRNAs共表达的显著丰富的注释术语,可以预测lncRNAs和其他新基因的功能。然而,这些方法对用于估计共表达水平的方法是敏感的。结果我们使用19个正常人体组织的实验表达数据,测试并比较了两个著名的统计度量(Pearson和Spearman)和两个几何度量(Sobolev和Fisher)以识别共表达的基因。我们还使用了一种基于语义相似性的基准方法来评估这些方法使用一组标注良好的蛋白质编码基因来预测标注术语的能力。结论这项工作表明,几何度量,特别是与统计度量相结合,将比传统方法更有效地预测标注术语。对选定的lncRNA的测试证实,如果有一组可靠的表达数据,就有可能预测这些基因的功能。用于这项调查的软件是免费提供的。
BackgroundAlmost 16,000 human long non-coding RNA (lncRNA) genes have been identified in the GENCODE project. However, the function of most of them remains to be discovered. The function of lncRNAs and other novel genes can be predicted by identifying significantly enriched annotation terms in already annotated genes that are co-expressed with the lncRNAs. However, such approaches are sensitive to the methods that are used to estimate the level of co-expression.ResultsWe have tested and compared two well-known statistical metrics (Pearson and Spearman) and two geometrical metrics (Sobolev and Fisher) for identification of the co-expressed genes, using experimental expression data across 19 normal human tissues. We have also used a benchmarking approach based on semantic similarity to evaluate how well these methods are able to predict annotation terms, using a well-annotated set of protein-coding genes.ConclusionThis work shows that geometrical metrics, in particular in combination with the statistical metrics, will predict annotation terms more efficiently than traditional approaches. Tests on selected lncRNAs confirm that it is possible to predict the function of these genes given a reliable set of expression data. The software used for this investigation is freely available.