Maximizing capture of gene co-expression relationships through pre-clustering of input expression samples: an Arabidopsis case study.
Maximizing capture of gene co-expression relationships through pre-clustering of input expression samples: an Arabidopsis case study.
复制标题
DOI:
10.1186/1752-0509-7-44
复制
发表时间:
2013-06-05
影响因子:
--
通讯作者:
Smith MC
中科院分区:
文献类型:
--
作者:
Feltus FA;Ficklin SP;Gibson SM;Smith MC
In genomics, highly relevant gene interaction (co-expression) networks have been constructed by finding significant pair-wise correlations between genes in expression datasets. These networks are then mined to elucidate biological function at the polygenic level. In some cases networks may be constructed from input samples that measure gene expression under a variety of different conditions, such as for different genotypes, environments, disease states and tissues. When large sets of samples are obtained from public repositories it is often unmanageable to associate samples into condition-specific groups, and combining samples from various conditions has a negative effect on network size. A fixed significance threshold is often applied also limiting the size of the final network. Therefore, we propose pre-clustering of input expression samples to approximate condition-specific grouping of samples and individual network construction of each group as a means for dynamic significance thresholding. The net effect is increase sensitivity thus maximizing the total co-expression relationships in the final co-expression network compendium. A total of 86 Arabidopsis thaliana co-expression networks were constructed after k-means partitioning of 7,105 publicly available ATH1 Affymetrix microarray samples. We term each pre-sorted network a Gene Interaction Layer (GIL). Random Matrix Theory (RMT), an un-supervised thresholding method, was used to threshold each of the 86 networks independently, effectively providing a dynamic (non-global) threshold for the network. The overall gene count across all GILs reached 19,588 genes (94.7% measured gene coverage) and 558,022 unique co-expression relationships. In comparison, network construction without pre-sorting of input samples yielded only 3,297 genes (15.9%) and 129,134 relationships. in the global network. Here we show that pre-clustering of microarray samples helps approximate condition-specific networks and allows for dynamic thresholding using un-supervised methods. Because RMT ensures only highly significant interactions are kept, the GIL compendium consists of 558,022 unique high quality A. thaliana co-expression relationships across almost all of the measurable genes on the ATH1 array. For A. thaliana, these networks represent the largest compendium to date of significant gene co-expression relationships, and are a means to explore complex pathway, polygenic, and pleiotropic relationships for this focal model plant. The networks can be explored at sysbio.genome.clemson.edu. Finally, this method is applicable to any large expression profile collection for any organism and is best suited where a knowledge-independent network construction method is desired.
登录
查看更多内容
影响因子:
9.8
作者:
Conant GC;Wolfe KH
通讯作者:
Wolfe KH
影响因子:
3.7
作者:
Gibson SM;Ficklin SP;Isaacson S;Luo F;Feltus FA;Smith MC
通讯作者:
Smith MC
影响因子:
12.3
作者:
Gentleman RC;Carey VJ;Bates DM;Bolstad B;Dettling M;Dudoit S;Ellis B;Gautier L;Ge Y;Gentry J;Hornik K;Hothorn T;Huber W;Iacus S;Irizarry R;Leisch F;Li C;Maechler M;Rossini AJ;Sawitzki G;Smith C;Smyth G;Tierney L;Yang JY;Zhang J
通讯作者:
Zhang J
影响因子:
1
作者:
Hwang, Woochang;Cho, Young-Rae;Ramanathan, Murali
通讯作者:
Ramanathan, Murali
影响因子:
14.9
作者:
Avraham, Shulamit;Tung, Chih-Wei;Ilic, Katica;Jaiswal, Pankaj;Kellogg, Elizabeth A.;McCouch, Susan;Pujar, Anuradha;Reiser, Leonore;Rhee, Seung Y.;Sachs, Martin M.;Schaeffer, Mary;Stein, Lincoln;Stevens, Peter;Vincent, Leszek;Zapata, Felipe;Ware, Doreen
通讯作者:
Ware, Doreen