Efficient and accurate causal inference with hidden confounders from genome-transcriptome variation data
Efficient and accurate causal inference with hidden confounders from genome-transcriptome variation data
复制标题
利用基因组转录组变异数据中隐藏的混杂因素进行高效、准确的因果推断
DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
T. Michoel
中科院分区:
文献类型:
--
作者:
Lingfei Wang;T. Michoel
Mapping gene expression as a quantitative trait using whole genome-sequencing and transcriptome analysis allows to discover the functional consequences of genetic variation. We developed a novel method and ultra-fast software Findr for higly accurate causal inference between gene expression traits using cis-regulatory DNA variations as causal anchors, which improves current methods by taking into account hidden confounders and weak regulations. Findr outperformed existing methods on the DREAM5 Systems Genetics challenge and on the prediction of microRNA and transcription factor targets in human lymphoblastoid cells, while being nearly a million times faster. Findr is publicly available at https://github.com/lingfeiwang/findr. Author summary Understanding how genetic variation between individuals determines variation in observable traits or disease risk is one of the core aims of genetics. It is known that genetic variation often affects gene regulatory DNA elements and directly causes variation in expression of nearby genes. This effect in turn cascades down to other genes via the complex pathways and gene interaction networks that ultimately govern how cells operate in an ever changing environment. In theory, when genetic variation and gene expression levels are measured simultaneously in a large number of individuals, the causal effects of genes on each other can be inferred using statistical models similar to those used in randomized controlled trials. We developed a novel method and ultra-fast software Findr which, unlike existing methods, takes into account the complex but unknown network context when predicting causality between specific gene pairs. Findr’s predictions have a significantly higher overlap with known gene networks compared to existing methods, using both simulated and real data. Findr is also nearly a million times faster, and hence the only software in its class that can handle modern datasets where the expression levels of ten-thousands of genes are simultaneously measured in hundreds to thousands of individuals.
影响因子:
64.8
作者:
通讯作者:
--
影响因子:
7.7
作者:
Cole, Stephen R.;Platt, Robert W.;Poole, Charles
通讯作者:
Poole, Charles
影响因子:
11.4
作者:
Li, Yang;Tesson, Bruno M.;Churchill, Gary A.;Jansen, Ritsert C.
通讯作者:
Jansen, Ritsert C.