Practical considerations for dividing data into subsets prior to PPL analysis

Practical considerations for dividing data into subsets prior to PPL analysis
复制标题

DOI:
10.1159/000143405
复制
发表时间:
2008-01-01
期刊:
影响因子:
1.8
通讯作者:
Vieland, V. J.
Vieland, V. J.
中科院分区:
生物学4区
文献类型:
--
作者:
Govil, M.;Vieland, V. J.

文献摘要

被引文献

相似文献

目的:PPL是一类用于人类复杂性状遗传作图的统计量,它利用贝叶斯序贯更新来积累证据,以支持或反对潜在异质数据(子)集之间的连锁。在这里,我们系统地探索替代子集方法在PPL计算中的相对有效性。方法:基于一个正在进行的研究中的家系,我们模拟了三个家系集合(同胞对;2-3代;64代)的基因分型。对于每个系谱集,在不同的异质性水平下产生了100个重复(在没有连锁的情况下产生了1000个重复)。在每个重复中,跨随机定义的子集(RAND2、RAND4)、按真(真)连锁状态、具有真实(真)分类、按单个系谱(PED)或不带任何子集(无)执行更新。结果:与NONE、RAND2、RAND4或PED相比,在“联动”下,REAL产生更大的PPL。在‘无连锁’的情况下,RAND2、RAND4和PED的收益率几乎为零。结论:我们研究了不同的子集策略对PPL抽样行为的影响。我们的结果强调了寻找能够帮助描绘更多同质数据子集的变量的效用,并表明,一旦找到这样的变量,在链接的位置存在明显的异质性的情况下,顺序更新可以非常有益,而不会在未链接的位置发生膨胀。版权所有(C)2008 S.Karger AG,巴塞尔。
Objective: The PPL, a class of statistics for complex trait genetic mapping in humans, utilizes Bayesian sequential updating to accumulate evidence for or against linkage across potentially heterogeneous data (sub) sets. Here, we systematically explore the relative efficacy of alternative subsetting approaches for purposes of PPL calculation. Methods: We simulated genotypes for three pedigree sets (sib pairs; 2-3 generations; 6 4 generations) based on families from an ongoing study. For each pedigree set, 100 replicates were generated under different levels of heterogeneity (1000 under 'no linkage'). Within each replicate, updating was performed across subsets defined randomly (RAND2, RAND4), by true (TRUE) linkage status, with a realistic (REAL) classification, by individual pedigree (PED), or without any subsetting (NONE). Results: Under 'linkage', REAL yields larger PPLs compared to NONE, RAND2, RAND4, or PED. Under 'no linkage', RAND2, RAND4 and PED yield PPLs close to NONE. Conclusions: We have examined the impact of different subsetting strategies on the sampling behavior of the PPL. Our results underscore the utility of finding variables that can help delineate more homogeneous data subsets and demonstrate that, once such variables are found, sequential updating can be highly beneficial in the presence of appreciable heterogeneity at a linked locus, without inflation at an unlinked locus. Copyright (C) 2008 S. Karger AG, Basel.