Risk Prediction Modeling of Sequencing Data Using a Forward Random Field Method.

Risk Prediction Modeling of Sequencing Data Using a Forward Random Field Method.
复制标题

使用前向随机场方法对测序数据进行风险预测建模。

DOI:
10.1038/srep21120
复制
发表时间:
2016-02-19
期刊:
影响因子:
4.6
通讯作者:
Lu Q
Lu Q
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Wen Y;He Z;Li M;Lu Q

文献摘要

相似文献

随着高通量测序技术的进步,研究常见和罕见变异在疾病风险预测中的作用是可行的。虽然这项新技术在改善疾病预测方面有很大的希望,但大量的数据和罕见变异的低频率对风险预测建模提出了巨大的分析挑战。在本文中,我们开发了一个前向随机场方法(FRF)的风险预测建模使用测序数据。在FRF中,受试者的表型被视为由受试者的基因型形成的遗传空间上的随机场的随机实现,并且个体的表型可以由具有相似基因型的相邻受试者预测。FRF方法允许模型中有多个相似性度量和候选基因,并自适应地选择最佳相似性度量和疾病相关基因以反映潜在的疾病模型。它还避免了对罕见变异阈值的规定,并允许不同方向和程度的遗传效应。通过模拟,我们证明了FRF方法在各种疾病模型下比常用的基于支持向量机的方法获得更高或相当的准确性。我们进一步说明了FRF方法与应用程序的测序数据从达拉斯心脏研究。
With the advance in high-throughput sequencing technology, it is feasible to investigate the role of common and rare variants in disease risk prediction. While the new technology holds great promise to improve disease prediction, the massive amount of data and low frequency of rare variants pose great analytical challenges on risk prediction modeling. In this paper, we develop a forward random field method (FRF) for risk prediction modeling using sequencing data. In FRF, subjects’ phenotypes are treated as stochastic realizations of a random field on a genetic space formed by subjects’ genotypes, and an individual’s phenotype can be predicted by adjacent subjects with similar genotypes. The FRF method allows for multiple similarity measures and candidate genes in the model, and adaptively chooses the optimal similarity measure and disease-associated genes to reflect the underlying disease model. It also avoids the specification of the threshold of rare variants and allows for different directions and magnitudes of genetic effects. Through simulations, we demonstrate the FRF method attains higher or comparable accuracy over commonly used support vector machine based methods under various disease models. We further illustrate the FRF method with an application to the sequencing data obtained from the Dallas Heart Study.