Maximizing capture of gene co-expression relationships through pre-clustering of input expression samples: an Arabidopsis case study.

Maximizing capture of gene co-expression relationships through pre-clustering of input expression samples: an Arabidopsis case study.
复制标题

DOI:
10.1186/1752-0509-7-44
复制
发表时间:
2013-06-05
影响因子:
--
通讯作者:
Smith MC
Smith MC
中科院分区:
生物2区
文献类型:
--
作者:
Feltus FA;Ficklin SP;Gibson SM;Smith MC

文献摘要

参考文献

被引文献

相似文献

在基因组学中,通过发现表达数据集中基因之间显着的成对相关性,构建了高度相关的基因相互作用(共表达)网络。然后挖掘这些网络以阐明多基因水平的生物功能。在某些情况下,可以根据输入样本构建网络,这些样本测量各种不同条件下的基因表达,例如不同的基因型、环境、疾病状态和组织。当从公共存储库获取大量样本时,通常难以将样本关联到特定于条件的组,并且组合来自不同条件的样本会对网络规模产生负面影响。通常应用固定的显着性阈值,这也限制了最终网络的大小。因此,我们建议对输入表达样本进行预聚类,以近似特定条件的样本分组以及每组的单独网络构建,作为动态显着性阈值的方法。净效应是增加敏感性,从而最大化最终共表达网络纲要中的总共表达关系。对 7,105 个公开可用的 ATH1 Affymetrix 微阵列样本进行 k 均值划分后,构建了总共 86 个拟南芥共表达网络。我们将每个预排序网络称为基因交互层(GIL)。随机矩阵理论 (RMT) 是一种无监督阈值方法,用于独立地对 86 个网络中的每个网络进行阈值处理,有效地为网络提供动态(非全局)阈值。所有 GIL 的总体基因计数达到 19,588 个基因(94.7% 测量的基因覆盖率)和 558,022 个独特的共表达关系。相比之下,未对输入样本进行预排序的网络构建仅产生 3,297 个基因 (15.9%) 和 129,134 个关系。在全球网络中。在这里,我们表明,微阵列样本的预聚类有助于近似特定条件的网络,并允许使用无监督方法进行动态阈值处理。由于 RMT 确保仅保留高度显着的相互作用,因此 GIL 纲要包含 558,022 个独特的高质量拟南芥共表达关系,涵盖 ATH1 阵列上几乎所有可测量的基因。对于拟南芥来说,这些网络代表了迄今为止最大的重要基因共表达关系的概要,并且是探索这种焦点模型植物的复杂途径、多基因和多效性关系的一种手段。可以在 sysbio.genome.clemson.edu 上探索该网络。最后,该方法适用于任何生物体的任何大型表达谱集合,并且最适合需要独立于知识的网络构建方法的情况。
In genomics, highly relevant gene interaction (co-expression) networks have been constructed by finding significant pair-wise correlations between genes in expression datasets. These networks are then mined to elucidate biological function at the polygenic level. In some cases networks may be constructed from input samples that measure gene expression under a variety of different conditions, such as for different genotypes, environments, disease states and tissues. When large sets of samples are obtained from public repositories it is often unmanageable to associate samples into condition-specific groups, and combining samples from various conditions has a negative effect on network size. A fixed significance threshold is often applied also limiting the size of the final network. Therefore, we propose pre-clustering of input expression samples to approximate condition-specific grouping of samples and individual network construction of each group as a means for dynamic significance thresholding. The net effect is increase sensitivity thus maximizing the total co-expression relationships in the final co-expression network compendium. A total of 86 Arabidopsis thaliana co-expression networks were constructed after k-means partitioning of 7,105 publicly available ATH1 Affymetrix microarray samples. We term each pre-sorted network a Gene Interaction Layer (GIL). Random Matrix Theory (RMT), an un-supervised thresholding method, was used to threshold each of the 86 networks independently, effectively providing a dynamic (non-global) threshold for the network. The overall gene count across all GILs reached 19,588 genes (94.7% measured gene coverage) and 558,022 unique co-expression relationships. In comparison, network construction without pre-sorting of input samples yielded only 3,297 genes (15.9%) and 129,134 relationships. in the global network. Here we show that pre-clustering of microarray samples helps approximate condition-specific networks and allows for dynamic thresholding using un-supervised methods. Because RMT ensures only highly significant interactions are kept, the GIL compendium consists of 558,022 unique high quality A. thaliana co-expression relationships across almost all of the measurable genes on the ATH1 array. For A. thaliana, these networks represent the largest compendium to date of significant gene co-expression relationships, and are a means to explore complex pathway, polygenic, and pleiotropic relationships for this focal model plant. The networks can be explored at sysbio.genome.clemson.edu. Finally, this method is applicable to any large expression profile collection for any organism and is best suited where a knowledge-independent network construction method is desired.
DOI: 10.1371/journal.pbio.0040109
发表时间: 2006-04
期刊: PLoS biology
影响因子: 9.8
作者:
Conant GC;Wolfe KH
通讯作者: Wolfe KH
DOI: 10.1371/journal.pone.0055871
发表时间: 2013
期刊: PloS one
影响因子: 3.7
作者:
Gibson SM;Ficklin SP;Isaacson S;Luo F;Feltus FA;Smith MC
通讯作者: Smith MC
DOI: 10.1186/gb-2004-5-10-r80
发表时间: 2004
期刊: Genome biology
影响因子: 12.3
作者:
Gentleman RC;Carey VJ;Bates DM;Bolstad B;Dettling M;Dudoit S;Ellis B;Gautier L;Ge Y;Gentry J;Hornik K;Hothorn T;Huber W;Iacus S;Irizarry R;Leisch F;Li C;Maechler M;Rossini AJ;Sawitzki G;Smith C;Smyth G;Tierney L;Yang JY;Zhang J
通讯作者: Zhang J
DOI: 10.1186/1748-7188-1-24
发表时间: 2006-01-01
影响因子: 1
作者:
Hwang, Woochang;Cho, Young-Rae;Ramanathan, Murali
通讯作者: Ramanathan, Murali
DOI: 10.1093/nar/gkm908
发表时间: 2008-01
影响因子: 14.9
作者:
Avraham, Shulamit;Tung, Chih-Wei;Ilic, Katica;Jaiswal, Pankaj;Kellogg, Elizabeth A.;McCouch, Susan;Pujar, Anuradha;Reiser, Leonore;Rhee, Seung Y.;Sachs, Martin M.;Schaeffer, Mary;Stein, Lincoln;Stevens, Peter;Vincent, Leszek;Zapata, Felipe;Ware, Doreen
通讯作者: Ware, Doreen