Maximal extraction of biological information from genetic interaction data.
Maximal extraction of biological information from genetic interaction data.
复制标题
DOI:
10.1371/journal.pcbi.1000347
复制
发表时间:
2009-04
影响因子:
4.3
通讯作者:
Galitski T
中科院分区:
文献类型:
--
作者:
Carter GW;Galas DJ;Galitski T
Extraction of all the biological information inherent in large-scale genetic interaction datasets remains a significant challenge for systems biology. The core problem is essentially that of classification of the relationships among phenotypes of mutant strains into biologically informative “rules” of gene interaction. Geneticists have determined such classifications based on insights from biological examples, but it is not clear that there is a systematic, unsupervised way to extract this information. In this paper we describe such a method that depends on maximizing a previously described context-dependent information measure to obtain maximally informative biological networks. We have successfully validated this method on two examples from yeast by demonstrating that more biological information is obtained when analysis is guided by this information measure. The context-dependent information measure is a function only of phenotype data and a set of interaction rules, involving no prior biological knowledge. Analysis of the resulting networks reveals that the most biologically informative networks are those with the greatest context-dependent information scores. We propose that these high-complexity networks reveal genetic architecture at a modular level, in contrast to classical genetic interaction rules that order genes in pathways. We suggest that our analysis represents a powerful, data-driven, and general approach to genetic interaction analysis, with particular potential in the study of mammalian systems in which interactions are complex and gene annotation data are sparse. Targeted genetic perturbation is a powerful tool for inferring gene function in model organisms. Functional relationships between genes can be inferred by observing the effects of multiple genetic perturbations in a single strain. The study of these relationships, generally referred to as genetic interactions, is a classic technique for ordering genes in pathways, thereby revealing genetic organization and gene-to-gene information flow. Genetic interaction screens are now being carried out in high-throughput experiments involving tens or hundreds of genes. These data sets have the potential to reveal genetic organization on a large scale, and require computational techniques that best reveal this organization. In this paper, we use a complexity metric based in information theory to determine the maximally informative network given a set of genetic interaction data. We find that networks with high complexity scores yield the most biological information in terms of (i) specific associations between genes and biological functions, and (ii) mapping modules of co-functional genes. This information-based approach is an automated, unsupervised classification of the biological rules underlying observed genetic interactions. It might have particular potential in genetic studies in which interactions are complex and prior gene annotation data are sparse.
登录
查看更多内容
影响因子:
64.8
作者:
Steinmetz, LM;Sinha, H;Davis, RW
通讯作者:
Davis, RW
影响因子:
30.8
作者:
St Onge, Robert P.;Mani, Ramamurthy;Giaever, Guri
通讯作者:
Giaever, Guri
影响因子:
56.9
作者:
Milo, R;Shen-Orr, S;Alon, U
通讯作者:
Alon, U
DOI:
10.1007/pl00013817
发表时间:
1999-12-01
期刊:
MOLECULAR AND GENERAL GENETICS
影响因子:
--
作者:
Entian, KD;Schuster, T;Hinnen, A
通讯作者:
Hinnen, A
影响因子:
--
作者:
Zhang LV;King OD;Wong SL;Goldberg DS;Tong AH;Lesage G;Andrews B;Bussey H;Boone C;Roth FP
通讯作者:
Roth FP