An additional k-means clustering step improves the biological features of WGCNA gene co-expression networks.

An additional k-means clustering step improves the biological features of WGCNA gene co-expression networks.
复制标题

DOI:
10.1186/s12918-017-0420-6
复制
发表时间:
2017-04-12
影响因子:
--
通讯作者:
Weale ME
Weale ME
中科院分区:
生物2区
文献类型:
--
作者:
Botía JA;Vandrovcova J;Forabosco P;Guelfi S;D'Sa K;United Kingdom Brain Expression Consortium;Hardy J;Lewis CM;Ryten M;Weale ME

文献摘要

被引文献

相似文献

加权基因共表达网络分析(WGCNA)是一个广泛使用的R软件包,用于生成基因共表达网络(GCN)。WGCNA生成GCN和基因簇(模块)的衍生分区。我们提出k均值聚类作为传统WGCNA的附加处理步骤,我们已经在R包km 2gcn中实现了这一步骤(k-means to gene co-expression network,https://github.com/juanbot/km2gcn)。我们在从UKBEC数据(10种不同的人脑组织)创建的网络上,在从GTEx数据(42种人体组织,包括13种脑组织)创建的网络上,以及在从GTEx数据导出的模拟网络上评估了我们的方法。我们观察到显著改善的模块特性,包括:(1)很少或零错位基因;(2)增加交替组织中可复制簇的计数(平均x3.1);(3)改进基因本体论术语的富集(见于48/52例GCN)(4)细胞类型富集信号改善(见于21/23的脑GCN);和(5)根据一系列相似性指数,模拟数据中的更准确分区。从我们的调查得到的结果表明,我们的k-means方法,作为标准WGCNA的辅助,结果在更好的网络分区。这些改进的分区使下游分析更富有成效,因为基因模块更具生物学意义。本文的在线版本(doi:10.1186/s12918-017-0420-6)包含补充材料,可供授权用户使用。
Weighted Gene Co-expression Network Analysis (WGCNA) is a widely used R software package for the generation of gene co-expression networks (GCN). WGCNA generates both a GCN and a derived partitioning of clusters of genes (modules). We propose k-means clustering as an additional processing step to conventional WGCNA, which we have implemented in the R package km2gcn (k-means to gene co-expression network, https://github.com/juanbot/km2gcn). We assessed our method on networks created from UKBEC data (10 different human brain tissues), on networks created from GTEx data (42 human tissues, including 13 brain tissues), and on simulated networks derived from GTEx data. We observed substantially improved module properties, including: (1) few or zero misplaced genes; (2) increased counts of replicable clusters in alternate tissues (x3.1 on average); (3) improved enrichment of Gene Ontology terms (seen in 48/52 GCNs) (4) improved cell type enrichment signals (seen in 21/23 brain GCNs); and (5) more accurate partitions in simulated data according to a range of similarity indices. The results obtained from our investigations indicate that our k-means method, applied as an adjunct to standard WGCNA, results in better network partitions. These improved partitions enable more fruitful downstream analyses, as gene modules are more biologically meaningful. The online version of this article (doi:10.1186/s12918-017-0420-6) contains supplementary material, which is available to authorized users.