Whole genome association mapping by incompatibilities and local perfect phylogenies

Whole genome association mapping by incompatibilities and local perfect phylogenies
复制标题

DOI:
10.1186/1471-2105-7-454
复制
发表时间:
2006-10-16
期刊:
影响因子:
3
通讯作者:
Schierup, Mikkel H.
Schierup, Mikkel H.
中科院分区:
生物学4区
文献类型:
--
作者:
Mailund, Thomas;Besenbacher, Soren;Schierup, Mikkel H.

文献摘要

被引文献

相似文献

背景资料:随着目前的技术,大量的数据可以廉价和有效地产生关联研究,并防止数据分析成为研究的瓶颈,快速和有效的分析方法,规模等数据集sizes.Results:我们提出了一个快速的方法,准确定位的致病变异在高密度的病例-对照关联映射实验与大量的病例和对照。该方法在由与单个系统发育树兼容的每个标记周围的最大区域定义的“完美”系统发育树中搜索病例染色体的显著聚类。这个完美的系统发育树被视为用于确定疾病状态的决策树,并根据其作为决策树的准确性进行评分。这样做的理由是,与随机树相比,疾病影响突变附近的完美遗传学应该提供更多关于受影响/未受影响分类的信息。如果相容性区域包含很少的标记,由于e。G.大的标记间距,该算法可以允许包括不亲和标记,以便在估计它们的同源性之前扩大区域。可以分析单倍型数据和定相基因型数据。该方法的功率和效率进行了研究1)模拟基因型数据下的不同模型的疾病确定2)人工数据集创建的HapMap资源,和3)数据集用于测试的其他方法,以便与这些进行比较。我们的方法在最简单的情况下具有与单一标记关联(SMA)相同的准确性,即单一致病突变和恒定重组率。然而,当涉及到更复杂的突变异质性和更复杂的单倍型结构时,例如在HapMap数据中发现的,我们的方法优于SMA以及其他快速的数据挖掘方法,例如HapMiner和单倍型模式挖掘(HPM),尽管速度明显更快。对于未定相的基因型数据,估计相位的初始步骤仅略微降低该方法的功效。还发现该方法可以准确定位经验数据集中已知的易感性变体-囊性纤维化的Delta F508突变-其中易感性变体是已知的-并找到CYP 2D 6基因与药物代谢不良之间关联的显著信号,尽管对于该数据集,最高关联得分距离CYP 2D 6基因约60 kb。我们的方法已在BLOCKASSOCiation软件中实现。使用微阵列,可以在不到两个CPU小时的时间内分析1000例病例和1000例对照中300万个SNP的全基因组芯片调查。
Background: With current technology, vast amounts of data can be cheaply and efficiently produced in association studies, and to prevent data analysis to become the bottleneck of studies, fast and efficient analysis methods that scale to such data set sizes must be developed.Results: We present a fast method for accurate localisation of disease causing variants in high density case-control association mapping experiments with large numbers of cases and controls. The method searches for significant clustering of case chromosomes in the "perfect" phylogenetic tree defined by the largest region around each marker that is compatible with a single phylogenetic tree. This perfect phylogenetic tree is treated as a decision tree for determining disease status, and scored by its accuracy as a decision tree. The rationale for this is that the perfect phylogeny near a disease affecting mutation should provide more information about the affected/unaffected classification than random trees. If regions of compatibility contain few markers, due to e. g. large marker spacing, the algorithm can allow the inclusion of incompatibility markers in order to enlarge the regions prior to estimating their phylogeny. Haplotype data and phased genotype data can be analysed. The power and efficiency of the method is investigated on 1) simulated genotype data under different models of disease determination 2) artificial data sets created from the HapMap ressource, and 3) data sets used for testing of other methods in order to compare with these. Our method has the same accuracy as single marker association (SMA) in the simplest case of a single disease causing mutation and a constant recombination rate. However, when it comes to more complex scenarios of mutation heterogeneity and more complex haplotype structure such as found in the HapMap data our method outperforms SMA as well as other fast, data mining approaches such as HapMiner and Haplotype Pattern Mining (HPM) despite being significantly faster. For unphased genotype data, an initial step of estimating the phase only slightly decreases the power of the method. The method was also found to accurately localise the known susceptibility variants in an empirical data set - the Delta F508 mutation for cystic fibrosis - where the susceptibility variant is already known - and to find significant signals for association between the CYP2D6 gene and poor drug metabolism, although for this dataset the highest association score is about 60 kb from the CYP2D6 gene.Conclusion: Our method has been implemented in the Blossoc (BLOck aSSOCiation) software. Using Blossoc, genome wide chip-based surveys of 3 million SNPs in 1000 cases and 1000 controls can be analysed in less than two CPU hours.