Efficient and accurate causal inference with hidden confounders from genome-transcriptome variation data

Efficient and accurate causal inference with hidden confounders from genome-transcriptome variation data
复制标题

利用基因组转录组变异数据中隐藏的混杂因素进行高效、准确的因果推断

DOI:
--
复制
发表时间:
2016
期刊:
bioRxiv
影响因子:
--
通讯作者:
T. Michoel
T. Michoel
中科院分区:
--
文献类型:
--
作者:
Lingfei Wang;T. Michoel

文献摘要

参考文献

被引文献

相似文献

利用全基因组测序和转录组分析将基因表达定位为数量性状,可以发现遗传变异的功能后果。我们开发了一种新的方法和超快速软件Findr,使用顺式调控DNA变异作为因果锚,可以在基因表达性状之间进行高精度的因果推断,通过考虑隐藏的混杂因素和弱调控来改进现有方法。在DREAM5系统遗传学挑战和预测人类淋巴母细胞中的microRNA和转录因子靶标方面,Findr的表现优于现有方法,同时速度快近100万倍。Findr可在https://github.com/lingfeiwang/findr.上公开获得作者总结了解个体之间的遗传变异如何决定可观察到的特征或疾病风险的变异是遗传学的核心目标之一。众所周知,遗传变异经常影响基因调控的DNA元件,并直接导致邻近基因表达的变化。这种效应又通过复杂的途径和基因相互作用网络向下传递到其他基因,最终决定细胞在不断变化的环境中如何运作。理论上,当在大量个体中同时测量遗传变异和基因表达水平时,可以使用与随机对照试验中使用的统计模型类似的统计模型来推断基因之间的因果关系。我们开发了一种新的方法和超高速软件Findr,与现有方法不同,它在预测特定基因对之间的因果关系时考虑了复杂但未知的网络环境。与现有的方法相比,Findr的预测与已知基因网络的重叠程度明显更高,使用的都是模拟数据和真实数据。Findr的速度也快了近一百万倍,因此是同类软件中唯一可以处理现代数据集的软件,在现代数据集中,数以万计的基因的表达水平在数百到数千个个体中同时进行测量。
Mapping gene expression as a quantitative trait using whole genome-sequencing and transcriptome analysis allows to discover the functional consequences of genetic variation. We developed a novel method and ultra-fast software Findr for higly accurate causal inference between gene expression traits using cis-regulatory DNA variations as causal anchors, which improves current methods by taking into account hidden confounders and weak regulations. Findr outperformed existing methods on the DREAM5 Systems Genetics challenge and on the prediction of microRNA and transcription factor targets in human lymphoblastoid cells, while being nearly a million times faster. Findr is publicly available at https://github.com/lingfeiwang/findr. Author summary Understanding how genetic variation between individuals determines variation in observable traits or disease risk is one of the core aims of genetics. It is known that genetic variation often affects gene regulatory DNA elements and directly causes variation in expression of nearby genes. This effect in turn cascades down to other genes via the complex pathways and gene interaction networks that ultimately govern how cells operate in an ever changing environment. In theory, when genetic variation and gene expression levels are measured simultaneously in a large number of individuals, the causal effects of genes on each other can be inferred using statistical models similar to those used in randomized controlled trials. We developed a novel method and ultra-fast software Findr which, unlike existing methods, takes into account the complex but unknown network context when predicting causality between specific gene pairs. Findr’s predictions have a significantly higher overlap with known gene networks compared to existing methods, using both simulated and real data. Findr is also nearly a million times faster, and hence the only software in its class that can handle modern datasets where the expression levels of ten-thousands of genes are simultaneously measured in hundreds to thousands of individuals.
DOI: 10.1038/nature12531
发表时间: 2013-09-26
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
DOI: 10.1093/ije/dyp334
发表时间: 2010-04-01
影响因子: 7.7
作者:
Cole, Stephen R.;Platt, Robert W.;Poole, Charles
通讯作者: Poole, Charles
DOI: 10.1016/j.tig.2010.09.002
发表时间: 2010-12
期刊: TRENDS IN GENETICS
影响因子: 11.4
作者:
Li, Yang;Tesson, Bruno M.;Churchill, Gary A.;Jansen, Ritsert C.
通讯作者: Jansen, Ritsert C.