Maximal extraction of biological information from genetic interaction data.

Maximal extraction of biological information from genetic interaction data.
复制标题

DOI:
10.1371/journal.pcbi.1000347
复制
发表时间:
2009-04
影响因子:
4.3
通讯作者:
Galitski T
Galitski T
中科院分区:
生物学2区
文献类型:
--
作者:
Carter GW;Galas DJ;Galitski T

文献摘要

参考文献

被引文献

相似文献

提取大规模遗传相互作用数据集中固有的所有生物信息仍然是系统生物学面临的重大挑战。核心问题基本上是将突变菌株的表型之间的关系归类为具有生物学信息的基因相互作用的“规则”。遗传学家已经根据生物学例子的洞察力确定了这样的分类,但尚不清楚是否有一种系统的、无监督的方法来提取这些信息。在这篇文章中,我们描述了这样一种方法,它依赖于最大化先前描述的上下文相关信息度量来获得最大信息量的生物网络。我们已经在酵母的两个例子上成功地验证了这种方法,证明了当用这种信息测量来指导分析时,可以获得更多的生物信息。上下文相关的信息度量仅是表型数据和一组交互规则的函数,不涉及先前的生物学知识。对所得网络的分析表明,生物信息量最大的网络是那些具有最大上下文相关信息分数的网络。我们认为,这些高复杂性的网络揭示了模块化水平的遗传结构,而不是经典的基因相互作用规则,即在路径中对基因进行排序。我们认为,我们的分析代表了一种强大的、数据驱动的和通用的遗传交互分析方法,在研究交互作用复杂而基因注释数据稀疏的哺乳动物系统方面具有特别的潜力。靶向遗传扰动是推断模式生物基因功能的有力工具。基因之间的功能关系可以通过观察单一菌株中多个遗传扰动的影响来推断。对这些关系的研究,通常被称为遗传相互作用,是对途径中的基因进行排序的经典技术,从而揭示遗传组织和基因之间的信息流。基因相互作用筛选现在正在涉及数十或数百个基因的高通量实验中进行。这些数据集有可能揭示大规模的遗传组织,并需要最好地揭示这种组织的计算技术。在本文中,我们使用基于信息论的复杂性度量来确定给定一组遗传交互数据的最大信息量网络。我们发现,复杂性分数高的网络在(I)基因和生物功能之间的特定关联以及(Ii)协同功能基因的映射模块方面产生了最多的生物信息。这种基于信息的方法是对观察到的遗传相互作用背后的生物规则进行自动、无监督的分类。它可能在基因研究中具有特别的潜力,因为在这些研究中,相互作用是复杂的,之前的基因注释数据是稀疏的。
Extraction of all the biological information inherent in large-scale genetic interaction datasets remains a significant challenge for systems biology. The core problem is essentially that of classification of the relationships among phenotypes of mutant strains into biologically informative “rules” of gene interaction. Geneticists have determined such classifications based on insights from biological examples, but it is not clear that there is a systematic, unsupervised way to extract this information. In this paper we describe such a method that depends on maximizing a previously described context-dependent information measure to obtain maximally informative biological networks. We have successfully validated this method on two examples from yeast by demonstrating that more biological information is obtained when analysis is guided by this information measure. The context-dependent information measure is a function only of phenotype data and a set of interaction rules, involving no prior biological knowledge. Analysis of the resulting networks reveals that the most biologically informative networks are those with the greatest context-dependent information scores. We propose that these high-complexity networks reveal genetic architecture at a modular level, in contrast to classical genetic interaction rules that order genes in pathways. We suggest that our analysis represents a powerful, data-driven, and general approach to genetic interaction analysis, with particular potential in the study of mammalian systems in which interactions are complex and gene annotation data are sparse. Targeted genetic perturbation is a powerful tool for inferring gene function in model organisms. Functional relationships between genes can be inferred by observing the effects of multiple genetic perturbations in a single strain. The study of these relationships, generally referred to as genetic interactions, is a classic technique for ordering genes in pathways, thereby revealing genetic organization and gene-to-gene information flow. Genetic interaction screens are now being carried out in high-throughput experiments involving tens or hundreds of genes. These data sets have the potential to reveal genetic organization on a large scale, and require computational techniques that best reveal this organization. In this paper, we use a complexity metric based in information theory to determine the maximally informative network given a set of genetic interaction data. We find that networks with high complexity scores yield the most biological information in terms of (i) specific associations between genes and biological functions, and (ii) mapping modules of co-functional genes. This information-based approach is an automated, unsupervised classification of the biological rules underlying observed genetic interactions. It might have particular potential in genetic studies in which interactions are complex and prior gene annotation data are sparse.
DOI: 10.1038/416326a
发表时间: 2002-03-21
期刊: NATURE
影响因子: 64.8
作者:
Steinmetz, LM;Sinha, H;Davis, RW
通讯作者: Davis, RW
DOI: 10.1038/ng1948
发表时间: 2007-02-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
St Onge, Robert P.;Mani, Ramamurthy;Giaever, Guri
通讯作者: Giaever, Guri
DOI: 10.1126/science.298.5594.824
发表时间: 2002-10-25
期刊: SCIENCE
影响因子: 56.9
作者:
Milo, R;Shen-Orr, S;Alon, U
通讯作者: Alon, U
DOI: 10.1007/pl00013817
发表时间: 1999-12-01
期刊: MOLECULAR AND GENERAL GENETICS
影响因子: --
作者:
Entian, KD;Schuster, T;Hinnen, A
通讯作者: Hinnen, A
DOI: 10.1186/jbiol23
发表时间: 2005
期刊: Journal of biology
影响因子: --
作者:
Zhang LV;King OD;Wong SL;Goldberg DS;Tong AH;Lesage G;Andrews B;Bussey H;Boone C;Roth FP
通讯作者: Roth FP