Statistical resolution of ambiguous HLA typing data.

Statistical resolution of ambiguous HLA typing data.
复制标题

DOI:
10.1371/journal.pcbi.1000016
复制
发表时间:
2008-02-29
影响因子:
4.3
通讯作者:
Heckerman D
Heckerman D
中科院分区:
生物学2区
文献类型:
--
作者:
Listgarten J;Brumme Z;Kadie C;Xiaojiang G;Walker B;Carrington M;Goulder P;Heckerman D

文献摘要

参考文献

被引文献

相似文献

高分辨率HLA分型在免疫学的许多领域中起着核心作用,例如在识别疾病的免疫遗传风险因素中,在研究病原体的基因组如何响应免疫选择压力而进化中,以及在疫苗设计中,其中HLA限制性表位的识别可用于指导疫苗免疫原的选择。也许最直接的应用之一是直接的医疗决策,涉及干细胞移植供体与无关受体的匹配。然而,高分辨率HLA分型由于其高成本或无法重新分型历史数据而经常不可用。在本文中,我们介绍和评估一种方法,统计,在硅细化模糊和/或低分辨率HLA数据。我们的方法,这需要一个独立的,高分辨率的训练数据集从同一人口的数据进行细化,使用HLA单倍型的连锁不平衡,以及四位数的等位基因频率数据,以概率细化HLA分型。我们的方法的核心是使用单倍型推断。我们引入新的方法,这方面的改进后,期望最大化(EM)为基础的方法目前使用的HLA社区。我们的改进是通过使用一个简约的参数化单倍型分布和平滑的最大似然(ML)的解决方案。这些改进使得可以以计算效率更高且更稳定的方式将细化扩展到更多数量的等位基因和基因座。我们还展示了如何增强我们的方法,以纳入种族信息(HLA等位基因分布根据种族/种族以及地理区域的差异很大),并证明了这种实验的潜在效用。基于我们的方法的工具可以在http://microsoft.com/science上免费获得。人类适应性免疫应答的核心是训练杀伤机制,其中特化的免疫细胞被敏化以识别来自外源的小肽(例如,艾滋病毒或细菌)。在这种致敏之后,这些免疫细胞然后被激活以杀死展示这种相同肽(并且含有这种相同外源肽)的其他细胞。然而,为了使致敏和杀死发生,外来肽必须与感染者的另一种专门的免疫分子-HLA分子“配对”。肽与这些HLA分子相互作用的方式决定了是否以及如何产生免疫应答。有一个巨大的库,这样的HLA分子,几乎没有两个人有相同的设置。此外,一个人的HLA类型可以决定他们对疾病的易感性,或移植的成功,例如。然而,获得高质量的HLA数据的患者往往是困难的,因为巨大的成本和专业实验室的要求,或者因为数据是历史的,不能用现代方法重新分型。因此,我们引入了一个统计模型,它可以利用现有的高质量的HLA数据,从低质量的数据推断出高质量的HLA数据。
High-resolution HLA typing plays a central role in many areas of immunology, such as in identifying immunogenetic risk factors for disease, in studying how the genomes of pathogens evolve in response to immune selection pressures, and also in vaccine design, where identification of HLA-restricted epitopes may be used to guide the selection of vaccine immunogens. Perhaps one of the most immediate applications is in direct medical decisions concerning the matching of stem cell transplant donors to unrelated recipients. However, high-resolution HLA typing is frequently unavailable due to its high cost or the inability to re-type historical data. In this paper, we introduce and evaluate a method for statistical, in silico refinement of ambiguous and/or low-resolution HLA data. Our method, which requires an independent, high-resolution training data set drawn from the same population as the data to be refined, uses linkage disequilibrium in HLA haplotypes as well as four-digit allele frequency data to probabilistically refine HLA typings. Central to our approach is the use of haplotype inference. We introduce new methodology to this area, improving upon the Expectation-Maximization (EM)-based approaches currently used within the HLA community. Our improvements are achieved by using a parsimonious parameterization for haplotype distributions and by smoothing the maximum likelihood (ML) solution. These improvements make it possible to scale the refinement to a larger number of alleles and loci in a more computationally efficient and stable manner. We also show how to augment our method in order to incorporate ethnicity information (as HLA allele distributions vary widely according to race/ethnicity as well as geographic area), and demonstrate the potential utility of this experimentally. A tool based on our approach is freely available for research purposes at http://microsoft.com/science. At the core of the human adaptive immune response is the train-to-kill mechanism in which specialized immune cells are sensitized to recognize small peptides from foreign sources (e.g., from HIV or bacteria). Following this sensitization, these immune cells are then activated to kill other cells which display this same peptide (and which contain this same foreign peptide). However, in order for sensitization and killing to occur, the foreign peptide must be “paired up” with one of the infected person's other specialized immune molecules—an HLA molecule. The way in which peptides interact with these HLA molecules defines if and how an immune response will be generated. There is a huge repertoire of such HLA molecules, with almost no two people having the same set. Furthermore, a person's HLA type can determine their susceptibility to disease, or the success of a transplant, for example. However, obtaining high quality HLA data for patients is often difficult because of the great cost and specialized laboratories required, or because the data are historical and cannot be retyped with modern methods. Therefore, we introduce a statistical model which can make use of existing high-quality HLA data, to infer higher-quality HLA data from lower-quality data.
DOI: 10.1371/journal.ppat.0030094
发表时间: 2007-07
期刊: PLoS pathogens
影响因子: 6.7
作者:
Brumme ZL;Brumme CJ;Heckerman D;Korber BT;Daniels M;Carlson J;Kadie C;Bhattacharya T;Chui C;Szinger J;Mo T;Hogg RS;Montaner JS;Frahm N;Brander C;Walker BD;Harrigan PR
通讯作者: Harrigan PR
DOI: 10.1056/nejm200105313442203
发表时间: 2001-05-31
影响因子: 158.5
作者:
Gao, XJ;Nelson, GW;Carrington, M
通讯作者: Carrington, M
DOI: 10.1093/bioinformatics/btl324
发表时间: 2007-01-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Hertz, Tomer;Yanover, Chen
通讯作者: Yanover, Chen
DOI: 10.1126/science.1131528
发表时间: 2007-03-16
期刊: SCIENCE
影响因子: 56.9
作者:
Bhattacharya, Tanmoy;Daniels, Marcus;Korber, Bette
通讯作者: Korber, Bette
DOI: 10.1371/journal.pmed.0040287
发表时间: 2007-09
期刊: PLoS medicine
影响因子: 15.8
作者:
Ellison GT;Smart A;Tutton R;Outram SM;Ashcroft R;Martin P
通讯作者: Martin P