Computational selection of transcriptomics experiments improves Guilt-by-Association analyses.

Computational selection of transcriptomics experiments improves Guilt-by-Association analyses.
复制标题

DOI:
10.1371/journal.pone.0039681
复制
发表时间:
2012
期刊:
影响因子:
3.7
通讯作者:
Paccanaro A
Paccanaro A
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Bhat P;Yang H;Bögre L;Devoto A;Paccanaro A

文献摘要

参考文献

被引文献

相似文献

关联罪 (GBA) 原理,根据该原理,具有相似表达谱的基因在功能上相关联,广泛应用于使用大量异质转录组数据集进行功能分析。然而,使用如此大的集合可能会妨碍对表达条件特异性的基因进行 GBA 功能分析。在这些情况下,应该使用较小的条件相关实验集,但仅根据文献知识从大量集合中识别此类功能相关的实验是一项不切实际的任务。我们首先从数学和生物学的角度分析为什么在 GBA 功能分析中只应使用特定条件的实验。我们能够证明这种现象与功能分类方案和所分析的生物体无关。然后,我们提出了一种半监督算法,可以从大量转录组学实验中选择功能相关的实验。我们的算法能够选择与给定 GO 术语、MIPS FunCat 术语甚至 KEGG 路径相关的实验。我们在酵母和拟南芥的大型数据集上广泛测试了我们的算法。我们证明:使用选定的实验,感兴趣的功能类别中基因之间的相关性在统计上有显着的改善;所选实验改进了基于 GBA 的基因功能预测;所选实验的有效性随着注释特异性的增加而增加;我们的算法可以成功应用于基于GBA的路径重建。重要的是,算法选择的实验集反映了有关实验的现有文献知识。 [该算法的MATLAB实现以及本文使用的所有数据可以从论文网站下载:http://www.paccanarolab.org/papers/CorrGene/]。
The Guilt-by-Association (GBA) principle, according to which genes with similar expression profiles are functionally associated, is widely applied for functional analyses using large heterogeneous collections of transcriptomics data. However, the use of such large collections could hamper GBA functional analysis for genes whose expression is condition specific. In these cases a smaller set of condition related experiments should instead be used, but identifying such functionally relevant experiments from large collections based on literature knowledge alone is an impractical task. We begin this paper by analyzing, both from a mathematical and a biological point of view, why only condition specific experiments should be used in GBA functional analysis. We are able to show that this phenomenon is independent of the functional categorization scheme and of the organisms being analyzed. We then present a semi-supervised algorithm that can select functionally relevant experiments from large collections of transcriptomics experiments. Our algorithm is able to select experiments relevant to a given GO term, MIPS FunCat term or even KEGG pathways. We extensively test our algorithm on large dataset collections for yeast and Arabidopsis. We demonstrate that: using the selected experiments there is a statistically significant improvement in correlation between genes in the functional category of interest; the selected experiments improve GBA-based gene function prediction; the effectiveness of the selected experiments increases with annotation specificity; our algorithm can be successfully applied to GBA-based pathway reconstruction. Importantly, the set of experiments selected by the algorithm reflects the existing literature knowledge about the experiments. [A MATLAB implementation of the algorithm and all the data used in this paper can be downloaded from the paper website: http://www.paccanarolab.org/papers/CorrGene/].
DOI: 10.1186/gb-2009-10-12-r139
发表时间: 2009
期刊: Genome biology
影响因子: 12.3
作者:
Adler P;Kolde R;Kull M;Tkachenko A;Peterson H;Reimand J;Vilo J
通讯作者: Vilo J
DOI: 10.1093/nar/gkm815
发表时间: 2008-01
影响因子: 14.9
作者:
Faith JJ;Driscoll ME;Fusaro VA;Cosgrove EJ;Hayete B;Juhn FS;Schneider SJ;Gardner TS
通讯作者: Gardner TS
DOI: 10.1186/gb-2004-5-10-r80
发表时间: 2004
期刊: Genome biology
影响因子: 12.3
作者:
Gentleman RC;Carey VJ;Bates DM;Bolstad B;Dettling M;Dudoit S;Ellis B;Gautier L;Ge Y;Gentry J;Hornik K;Hothorn T;Huber W;Iacus S;Irizarry R;Leisch F;Li C;Maechler M;Rossini AJ;Sawitzki G;Smith C;Smyth G;Tierney L;Yang JY;Zhang J
通讯作者: Zhang J
系统的调查揭示了基因共表达网络中“逐罪”的一般适用性。
DOI: 10.1186/1471-2105-6-227
发表时间: 2005-09-14
期刊: BMC bioinformatics
影响因子: 3
作者:
Wolfe CJ;Kohane IS;Butte AJ
通讯作者: Butte AJ
DOI: 10.1109/tcbb.2004.2
发表时间: 2004-01-01
影响因子: 4.5
作者:
Madeira, SC;Oliveira, AL
通讯作者: Oliveira, AL