Learning a prior on regulatory potential from eQTL data.

Learning a prior on regulatory potential from eQTL data.
复制标题

DOI:
10.1371/journal.pgen.1000358
复制
发表时间:
2009-01
期刊:
影响因子:
4.5
通讯作者:
Koller D
Koller D
中科院分区:
生物学2区
文献类型:
--
作者:
Lee SI;Dudley AM;Drubin D;Silver PA;Krogan NJ;Pe'er D;Koller D

文献摘要

参考文献

被引文献

相似文献

全基因组RNA表达数据提供了生物体生物状态的详细视图;因此,测量遗传多样个体间表达差异的数据集(eQTL数据)可能为复杂性状的遗传学提供重要见解。然而,利用来自相对较少个体的数据,很难从大量可能性中区分出真正的因果多态性。在具有显著连锁不平衡的群体中,这个问题尤其具有挑战性,在这些群体中,性状常常与包含许多基因的大染色体区域相关联。在此,我们提出一种新方法,Lirnet,它自动学习每个序列多态性的调控潜能,估计其对基因表达产生显著影响的可能性。这种调控潜能是根据“调控特征”来定义的——包括基因的功能以及遗传多态性的保守性、类型和位置——这些特征对任何生物体都是可用的。不同特征对调控潜能的影响程度是自动学习的,这使得Lirnet很容易应用于不同的数据集、生物体和特征集。我们将Lirnet应用于人类HapMap eQTL数据集和一个酵母eQTL数据集,并提供统计和生物学结果,证明Lirnet比其他近期方法能产生明显更好的调控程序。我们在酵母数据中证明,Lirnet能够正确地指出在一个大的连锁染色体区域内的特定因果序列变异。在一个例子中,Lirnet揭示了Puf3(一种序列特异性RNA结合蛋白)和P - 小体(调节翻译和RNA稳定性的细胞质结构)之间一种新的、经实验验证的联系,以及特定的致病多态性,即Mkt1中的一个单核苷酸多态性(SNP),它诱导了该通路中的变异。 遗传多样个体的基因表达数据(eQTL数据)为遗传变异对细胞通路的影响提供了独特视角。然而,多重假设的负担,加上连锁不平衡的挑战,使得正确识别因果多态性变得困难。研究人员传统上应用启发式方法在合理的假设中进行选择,倾向于更保守的、导致显著氨基酸变化的或位于其功能与目标相关的基因中的多态性。但是我们如何知道要给不同的调控特征赋予多少权重呢?我们描述了Lirnet,它从eQTL数据中学习如何对调控特征进行加权并推断序列变异的调控潜能。Lirnet在学习调控网络的同时评估这些权重,找到能使网络更具预测性的权重。我们表明Lirnet构建了高精度的调控程序,并证明了它正确识别致病多态性的能力。Lirnet可以灵活使用任何调控特征,包括任何已测序生物体可用的序列特征,并以特定于数据集的方式自动学习它们的权重。这一特征使其在哺乳动物系统中尤其具有优势,在哺乳动物系统中,简单模式生物中使用的许多形式的先验知识是不完整的或不可用的。
Genome-wide RNA expression data provide a detailed view of an organism's biological state; hence, a dataset measuring expression variation between genetically diverse individuals (eQTL data) may provide important insights into the genetics of complex traits. However, with data from a relatively small number of individuals, it is difficult to distinguish true causal polymorphisms from the large number of possibilities. The problem is particularly challenging in populations with significant linkage disequilibrium, where traits are often linked to large chromosomal regions containing many genes. Here, we present a novel method, Lirnet, that automatically learns a regulatory potential for each sequence polymorphism, estimating how likely it is to have a significant effect on gene expression. This regulatory potential is defined in terms of “regulatory features”—including the function of the gene and the conservation, type, and position of genetic polymorphisms—that are available for any organism. The extent to which the different features influence the regulatory potential is learned automatically, making Lirnet readily applicable to different datasets, organisms, and feature sets. We apply Lirnet both to the human HapMap eQTL dataset and to a yeast eQTL dataset and provide statistical and biological results demonstrating that Lirnet produces significantly better regulatory programs than other recent approaches. We demonstrate in the yeast data that Lirnet can correctly suggest a specific causal sequence variation within a large, linked chromosomal region. In one example, Lirnet uncovered a novel, experimentally validated connection between Puf3—a sequence-specific RNA binding protein—and P-bodies—cytoplasmic structures that regulate translation and RNA stability—as well as the particular causative polymorphism, a SNP in Mkt1, that induces the variation in the pathway. Gene expression data of genetically diverse individuals (eQTL data) provide a unique perspective on the effect of genetic variation on cellular pathways. However, the burden of multiple hypotheses, combined with the challenges of linkage disequilibrium, makes it difficult to correctly identify causal polymorphisms. Researchers traditionally apply heuristics for selecting among plausible hypotheses, favoring polymorphisms that are more conserved, that lead to significant amino acid change, or that reside in genes whose function is related to that of the targets. But how do we know how much weight to attribute to different regulatory features? We describe Lirnet, which learns from eQTL data how to weight regulatory features and induce a regulatory potential for sequence variations. Lirnet assesses these weights simultaneously to learning a regulatory network, finding weights that lead to a more predictive network. We show that Lirnet constructs high-accuracy regulatory programs and demonstrate its ability to correctly identify causative polymorphisms. Lirnet can flexibly use any regulatory features, including sequence features that are available for any sequenced organism, and automatically learn their weights in a dataset-specific way. This feature makes it especially advantageous for mammalian systems, where many forms of prior knowledge used in simple model organisms are incomplete or unavailable.
DOI: 10.1038/nature06258
发表时间: 2007-10-18
期刊: NATURE
影响因子: 64.8
作者:
Frazer, Kelly A.;Ballinger, Dennis G.;Cox, David R.;Hinds, David A.;Stuve, Laura L.;Gibbs, Richard A.;Belmont, John W.;Boudreau, Andrew;Hardenbol, Paul;Leal, Suzanne M.;Pasternak, Shiran;Wheeler, David A.;Willis, Thomas D.;Yu, Fuli;Yang, Huanming;Zeng, Changqing;Gao, Yang;Hu, Haoran;Hu, Weitao;Li, Chaohua;Lin, Wei;Liu, Siqi;Pan, Hao;Tang, Xiaoli;Wang, Jian;Wang, Wei;Yu, Jun;Zhang, Bo;Zhang, Qingrun;Zhao, Hongbin;Zhao, Hui;Zhou, Jun;Gabriel, Stacey B.;Barry, Rachel;Blumenstiel, Brendan;Camargo, Amy;Defelice, Matthew;Faggart, Maura;Goyette, Mary;Gupta, Supriya;Moore, Jamie;Nguyen, Huy;Onofrio, Robert C.;Parkin, Melissa;Roy, Jessica;Stahl, Erich;Winchester, Ellen;Ziaugra, Liuda;Altshuler, David;Shen, Yan;Yao, Zhijian;Huang, Wei;Chu, Xun;He, Yungang;Jin, Li;Liu, Yangfan;Shen, Yayun;Sun, Weiwei;Wang, Haifeng;Wang, Yi;Wang, Ying;Xiong, Xiaoyan;Xu, Liang;Waye, Mary M. Y.;Tsui, Stephen K. W.;Wong, J. Tze-Fei;Galver, Luana M.;Fan, Jian-Bing;Gunderson, Kevin;Murray, Sarah S.;Oliphant, Arnold R.;Chee, Mark S.;Montpetit, Alexandre;Chagnon, Fanny;Ferretti, Vincent;Leboeuf, Martin;Olivier, Jean-Franccois;Phillips, Michael S.;Roumy, Stephanie;Sallee, Clementine;Verner, Andrei;Hudson, Thomas J.;Kwok, Pui-Yan;Cai, Dongmei;Koboldt, Daniel C.;Miller, Raymond D.;Pawlikowska, Ludmila;Taillon-Miller, Patricia;Xiao, Ming;Tsui, Lap-Chee;Mak, William;Song, You Qiang;Tam, Paul K. H.;Nakamura, Yusuke;Kawaguchi, Takahisa;Kitamoto, Takuya;Morizono, Takashi;Nagashima, Atsushi;Ohnishi, Yozo;Sekine, Akihiro;Tanaka, Toshihiro;Tsunoda, Tatsuhiko;Deloukas, Panos;Bird, Christine P.;Delgado, Marcos;Dermitzakis, Emmanouil T.;Gwilliam, Rhian;Hunt, Sarah;Morrison, Jonathan;Powell, Don;Stranger, Barbara E.;Whittaker, Pamela;Bentley, David R.;Daly, Mark J.;de Bakker, Paul I. W.;Barrett, Jeff;Chretien, Yves R.;Maller, Julian;McCarroll, Steve;Patterson, Nick;Pe'er, Itsik;Price, Alkes;Purcell, Shaun;Richter, Daniel J.;Sabeti, Pardis;Saxena, Richa;Schaffner, Stephen F.;Sham, Pak C.;Varilly, Patrick;Altshuler, David;Stein, Lincoln D.;Krishnan, Lalitha;Smith, Albert Vernon;Tello-Ruiz, Marcela K.;Thorisson, Gudmundur A.;Chakravarti, Aravinda;Chen, Peter E.;Cutler, David J.;Kashuk, Carl S.;Lin, Shin;Abecasis, Goncalo R.;Guan, Weihua;Li, Yun;Munro, Heather M.;Qin, Zhaohui Steve;Thomas, Daryl J.;McVean, Gilean;Auton, Adam;Bottolo, Leonardo;Cardin, Niall;Eyheramendy, Susana;Freeman, Colin;Marchini, Jonathan;Myers, Simon;Spencer, Chris;Stephens, Matthew;Donnelly, Peter;Cardon, Lon R.;Clarke, Geraldine;Evans, David M.;Morris, Andrew P.;Weir, Bruce S.;Tsunoda, Tatsuhiko;Johnson, Todd A.;Mullikin, James C.;Sherry, Stephen T.;Feolo, Michael;Skol, Andrew
通讯作者: Skol, Andrew
DOI: 10.1534/genetics.105.041103
发表时间: 2005-06-01
期刊: GENETICS
影响因子: 3.3
作者:
Bing, N;Hoeschele, I
通讯作者: Hoeschele, I
DOI: 10.1073/pnas.0605140103
发表时间: 2006-08-08
影响因子: 11.1
作者:
Chua, Gordon;Morris, Quaid D.;Hughes, Timothy R.
通讯作者: Hughes, Timothy R.
DOI: 10.1038/nature05649
发表时间: 2007-04-12
期刊: NATURE
影响因子: 64.8
作者:
Collins, Sean R.;Miller, Kyle M.;Krogan, Nevan J.
通讯作者: Krogan, Nevan J.
DOI: 10.1214/009053604000000067
发表时间: 2004-04-01
影响因子: 4.5
作者:
Efron, B;Hastie, T;Tibshirani, R
通讯作者: Tibshirani, R