Mapping complex traits using Random Forests.

Mapping complex traits using Random Forests.
复制标题

使用随机森林绘制复杂性状。

DOI:
10.1186/1471-2156-4-s1-s64
复制
发表时间:
2003-12-31
期刊:
影响因子:
2.9
通讯作者:
Van Eerdewegh, P
Van Eerdewegh, P
中科院分区:
生物学3区
文献类型:
--
作者:
Bureau, A;Dupuis, J;Hayward, B;Falls, K;Van Eerdewegh, P

文献摘要

被引文献

相似文献

随机森林是一种基于自举数据样本上生长的树木的预测技术,结合随机选择的解释变量来定义每个节点的最佳分割。在定量结果的情况下,树预测器采用数值。我们将随机森林应用于遗传分析研讨会13模拟数据集的第一个复制,将兄弟姐妹对作为我们的分析单位,并将选择的基因座上的血统识别(IBD)作为我们的解释变量。有了真实模型的知识,我们对三种表型进行了两组分析:HDL,甘油三酯和葡萄糖。目标是从多变量的角度来研究复杂特征的映射。第一组分析模拟了候选基因方法,预测因子中真实基因的比例很高,而第二组分析代表了使用微卫星标记的基因组扫描分析。随机森林能够确定一些影响表型的主要基因,如基线HDL和甘油三酯,但未能确定调节基线葡萄糖水平的主要基因。
Random Forest is a prediction technique based on growing trees on bootstrap samples of data, in conjunction with a random selection of explanatory variables to define the best split at each node. In the case of a quantitative outcome, the tree predictor takes on a numerical value. We applied Random Forest to the first replicate of the Genetic Analysis Workshop 13 simulated data set, with the sibling pairs as our units of analysis and identity by descent (IBD) at selected loci as our explanatory variables. With the knowledge of the true model, we performed two sets of analyses on three phenotypes: HDL, triglycerides, and glucose. The goal was to approach the mapping of complex traits from a multivariate perspective. The first set of analyses mimics a candidate gene approach with a high proportion of true genes among the predictors while the second set represents a genome scan analysis using microsatellite markers. Random Forest was able to identify a few of the major genes influencing the phenotypes, such as baseline HDL and triglycerides, but failed to identify the major genes regulating baseline glucose levels.