Appropriate data cleaning methods for genome-wide association study

Appropriate data cleaning methods for genome-wide association study
复制标题

DOI:
10.1007/s10038-008-0322-y
复制
发表时间:
2008-10-01
影响因子:
3.5
通讯作者:
Tokunaga, Katsushi
Tokunaga, Katsushi
中科院分区:
生物学3区
文献类型:
--
作者:
Miyagawa, Taku;Nishida, Nao;Tokunaga, Katsushi

文献摘要

被引文献

相似文献

全基因组关联研究(GWAS)利用大量的单核苷酸多态性(SNP)已成功地应用于确定常见疾病的遗传变异。然而,使用新的阵列技术的基因分型通常与可能不利地影响GWAS分析的假结果相关。因此,数据清理在排除假基因分型结果方面至关重要。在这项研究中,我们调查了389个无关的健康日本样本的适当清洁所需的标准,使用基因芯片人类图谱500 K阵列集GWAS分析。将样本随机分为两组,并比较各组中单个SNP的等位基因频率,作为准病例对照研究。然后,通过四个参数(SNP调用率、使用具有Mahalanobis基因型调用算法的贝叶斯鲁棒线性模型获得的置信度得分、Hardy-Weinberg平衡和次要等位基因频率)过滤观察结果,并评估与零假设的偏差。我们发现,使用这四个参数可以实现适当的数据清洗。我们的研究结果提供了一个途径,从GWAS获得适当的数据。
Genome-wide association studies (GWAS) using a large number of single nucleotide polymorphisms (SNPs) have successfully been applied to identify genetic variants of common diseases. However, genotyping using the new array technologies is often associated with spurious results that could unfavorably affect analyses of GWAS. Consequently, data cleaning is of paramount importance in excluding spurious genotyping results. In this study, we investigated the criteria required for the appropriate cleaning of 389 unrelated healthy Japanese samples analyzed using the GeneChip Human Mapping 500K Array Set for GWAS. The samples were randomly subdivided into two groups, and the allele frequencies in the groups were compared for individual SNPs as a quasi-case-control study. Then, observed results were filtered by four parameters (SNP call rate, confidence score obtained using the Bayesian Robust Linear Model with Mahalanobis genotype-calling algorithm, Hardy-Weinberg equilibrium, and minor allele frequency) and assessed for deviation from the null hypothesis. We found that appropriate data cleaning could be achieved using these four parameters. Our findings offer an avenue for obtaining appropriate data from GWAS.