Likelihood-based inference on haplotype effects in genetic association studies

Likelihood-based inference on haplotype effects in genetic association studies
复制标题

DOI:
10.1198/016214505000000808
复制
发表时间:
2006-03-01
影响因子:
3.7
通讯作者:
Zeng, D
Zeng, D
中科院分区:
数学1区
文献类型:
--
作者:
Lin, DY;Zeng, D

文献摘要

被引文献

相似文献

单倍型是单个染色体上核苷酸的特定序列。单倍型和疾病表型之间的群体关联为复杂人类疾病的遗传基础提供了重要信息。标准的基因分型技术无法区分个体的两条同源染色体,因此只能直接观察到未分期的基因型(即两条同源单倍型的组合)。基于非阶段性基因型数据的单倍型-表型关联的统计推断提出了一个有趣的数据缺失问题,特别是当采样取决于疾病状态时。本文的目的是为这个问题提供一个系统而严谨的处理方法。所有常用的研究设计。包括横断面。考虑了病例对照和队列研究。表型可以是一种疾病指标,一种数量性状。或者是一个可能被删减的致病时间变量。单倍型对表型的影响是通过灵活的回归模型来表述的。它可以适应各种遗传机制和基因与环境的相互作用。构建可能涉及高维参数的适当可能性。证明了参数的可辨识性和极大似然估计的相合性、渐近正态性和有效性。开发了高效可靠的数值算法。仿真研究表明,基于似然的程序在实际设置中表现良好。提供了在芬兰-美国NIDDM遗传研究调查中的应用。讨论了需要进一步发展的领域。
A haplotype is a specific sequence of nucleotides on a single chromosome. The population associations between haplotypes and disease phenotypes provide critical information about the genetic basis of complex human diseases. Standard genotyping techniques cannot distinguish the two homologous chromosomes of an individual, so only the unphased genotype (i.e., the combination of the two homologous haplotypes) is directly observable. Statistical inference about haplotype-phenotype associations based on unphased genotype data presents an intriguing missing-data problem, especially when the sampling depends on the disease status. The objective of this article is to provide a systematic and rigorous treatment of this problem. All commonly used study designs. including cross-sectional. case-control, and cohort studies, are considered. The phenotype can be a disease indicator, a quantitative trait. or a potentially censored time-to-disease variable. The effects of haplotypes on the phenotype are formulated through flexible regression models. which can accommodate various genetic mechanisms and gene-environment interactions. Appropriate likelihoods are constructed that may involve high-dimensional parameters. The identifiability of the parameters and the consistency, asymptotic normality, and efficiency of the maximum likelihood estimators are established. Efficient and reliable numerical algorithms are developed. Simulation studies show that the likelihood-based procedures perform well in practical settings. An application to the Finland-United States Investigation of NIDDM Genetics Study is provided. Areas in need of further development are discussed.