Approximate likelihood methods for estimating local recombination rates

Approximate likelihood methods for estimating local recombination rates
复制标题

DOI:
10.1111/1467-9868.00355
复制
发表时间:
2002-01-01
影响因子:
5.8
通讯作者:
Donnelly, P
Donnelly, P
中科院分区:
数学1区
文献类型:
--
作者:
Fearnhead, P;Donnelly, P

文献摘要

被引文献

相似文献

目前,人们对了解整个人类基因组中重组率在短期内的变化方式非常感兴趣。除了固有的兴趣之外,对这种局部变异的理解对于许多旨在阐明常见疾病或人类历史的遗传基础的研究的合理设计和分析至关重要。基于谱系的标准方法不具备解决此问题所需的精细分辨率。相比之下,来自群体中不相关染色体的脱氧核糖核酸序列样本携带相关信息,但从此类数据进行推断极具挑战性。尽管最近人们对开发用于根据此类数据估计局部重组率的完全似然推理方法产生了很大兴趣,但它们目前对于现代实验技术生成的数据集大小并不实用。我们介绍并研究了两种近似似然方法。第一个是边际可能性,忽略了一些数据。仔细选择要忽略的内容可以节省大量计算量,并且几乎不会丢失相关信息。对于较大的序列,我们引入了“复合”似然,它通过忽略某些长期依赖性来近似感兴趣的模型。非正式的渐近分析和模拟研究表明,基于复合似然的推理是可行的并且表现良好。我们结合这两种方法来重新分析脂蛋白脂肪酶基因的数据,结果严重质疑了这些数据的一些早期研究的结论。
There is currently great interest in understanding the way in which recombination rates vary, over short scales, across the human genome. Aside from inherent interest, an understanding of this local variation is essential for the sensible design and analysis of many studies aimed at elucidating the genetic basis of common diseases or of human population histories. Standard pedigree-based approaches do not have the fine scale resolution that is needed to address this issue. In contrast, samples of deoxyribonucleic acid sequences from unrelated chromosomes in the population carry relevant information, but inference from such data is extremely challenging. Although there has been much recent interest in the development of full likelihood inference methods for estimating local recombination rates from such data, they are not currently practicable for data sets of the size being generated by modern experimental techniques. We introduce and study two approximate likelihood methods. The first, a marginal likelihood, ignores some of the data. A careful choice of what to ignore results in substantial computational savings with virtually no loss of relevant information. For larger sequences, we introduce a 'composite' likelihood, which approximates the model of interest by ignoring certain long-range dependences. An informal asymptotic analysis and a simulation study suggest that inference based on the composite likelihood is practicable and performs well. We combine both methods to reanalyse data from the lipoprotein lipase gene, and the results seriously question conclusions from some earlier studies of these data.