Finding consistent disease subnetworks across microarray datasets.

Finding consistent disease subnetworks across microarray datasets.
复制标题

DOI:
10.1186/1471-2105-12-s13-s15
复制
发表时间:
2011
期刊:
影响因子:
3
通讯作者:
Wong L
Wong L
中科院分区:
生物学4区
文献类型:
--
作者:
Soh D;Dong D;Guo Y;Wong L

文献摘要

被引文献

相似文献

虽然当代微阵列分析方法是研究单个微阵列数据集的优秀工具,但它们倾向于从同一疾病的不同数据集产生不同的结果。我们的目标是通过引入一种技术(SNet)来解决这种再现性问题。SNet通过识别有意义的通路的特定连接部分来提供微阵列数据集的定量和描述性分析。我们将通路中的这些部分称为“子网络”。我们在几种疾病的独立数据集上测试了SNet,包括儿童ALL,DMD和肺癌。对于这些疾病中的每一种,我们获得了由不同实验室在不同平台上产生的两个独立的微阵列数据集。在每种情况下,我们的技术始终产生几乎相同的列表显着的非平凡的子网络从两个独立的微阵列数据集。这些重要的子网络的基因水平的一致性在51.18%到93.01%之间。相比之下,当使用GSEA、t检验和SAM分析相同的微阵列数据集时,GSEA的百分比下降2.38%至28.90%,t检验的百分比下降49.60%至73.01%,SAM的百分比下降49.96%至81.25%。此外,使用这些现有方法选择的基因没有形成实质性大小的子网络。因此,更有可能的是,我们的技术选择的子网络可以为研究人员提供更多的描述性信息,这些信息是关于实际受疾病影响的通路部分的。这些结果清楚地表明,与其他流行的方法(GSEA,t-test和SAM)相比,我们的技术生成了重要的子网络和基因,这些子网络和基因在数据集之间更加一致和可重复。我们生成的子网络的大小表明它们通常更具生物学意义(不太可能是虚假的)。此外,我们选择了两个样本子网络,并验证他们与生物文献中的参考。这表明我们的算法能够生成描述性的生物学结论。
While contemporary methods of microarray analysis are excellent tools for studying individual microarray datasets, they have a tendency to produce different results from different datasets of the same disease. We aim to solve this reproducibility problem by introducing a technique (SNet). SNet provides both quantitative and descriptive analysis of microarray datasets by identifying specific connected portions of pathways that are significant. We term such portions within pathways as “subnetworks”. We tested SNet on independent datasets of several diseases, including childhood ALL, DMD and lung cancer. For each of these diseases, we obtained two independent microarray datasets produced by distinct labs on distinct platforms. In each case, our technique consistently produced almost the same list of significant nontrivial subnetworks from two independent sets of microarray data. The gene-level agreement of these significant subnetworks was between 51.18% to 93.01%. In contrast, when the same pairs of microarray datasets were analysed using GSEA, t-test and SAM, this percentage fell between 2.38% to 28.90% for GSEA, 49.60% tp 73.01% for t-test, and 49.96% to 81.25% for SAM. Furthermore, the genes selected using these existing methods did not form subnetworks of substantial size. Thus it is more probable that the subnetworks selected by our technique can provide the researcher with more descriptive information on the portions of the pathway actually affected by the disease. These results clearly demonstrate that our technique generates significant subnetworks and genes that are more consistent and reproducible across datasets compared to the other popular methods available (GSEA, t-test and SAM). The large size of subnetworks which we generate indicates that they are generally more biologically significant (less likely to be spurious). In addition, we have chosen two sample subnetworks and validated them with references from biological literature. This shows that our algorithm is capable of generating descriptive biologically conclusions.