Assessing batch effects of genotype calling algorithm BRLMM for the Affymetrix GeneChip Human Mapping 500 K array set using 270 HapMap samples.

Assessing batch effects of genotype calling algorithm BRLMM for the Affymetrix GeneChip Human Mapping 500 K array set using 270 HapMap samples.
复制标题

评估使用270个HAPMAP样品的Affymetrix Genechip人类映射的基因型调用算法BRLMM的批处理效应。

DOI:
10.1186/1471-2105-9-s9-s17
复制
发表时间:
2008-08-12
期刊:
影响因子:
3
通讯作者:
Tong W
Tong W
中科院分区:
生物学4区
文献类型:
--
作者:
Hong H;Su Z;Ge W;Shi L;Perkins R;Fang H;Xu J;Chen JJ;Han T;Kaput J;Fuscoe JC;Tong W

文献摘要

被引文献

相似文献

全基因组关联研究旨在确定整个人类基因组中与疾病状态和药物反应等表型特征相关的遗传变异(通常是单核苷酸多态[SNPs])。高精确度和可重复性的基因调用是最重要的,因为调用算法引入的错误可能会导致基因和表型之间的错误关联膨胀。目前用于GWAS的大多数基因分型算法都是基于多个阵列的。由于数以百亿兆字节(GB)的原始数据是从GWAS中产生的,样本通常被分成包含整个数据集的子集的批次,用于基因分型。实现了高呼叫率和准确率。然而,批次大小(即一起分析的芯片数量)和批次组成(即批次中芯片的选择)对呼叫率和准确性的影响以及这些影响对已确定的显著关联的SNP的影响尚未被调查。本文利用Affymetrix Human Maping500K数组分析了270个HapMap样本的原始数据,分析了批次大小和批次组成对BRLMM算法的影响。使用Affymetrix Human map 500K阵列集合询问的270个HapMap样本的数据,使用BRLMM算法对三种不同的批次大小和三种不同的批次成分进行基因分型。通过关联分析,将调用结果与相应的显著SNPs列表进行对比分析,发现批次大小和组成都影响了基因型调用结果,并且显著关联SNPs。批次大小和批次组成效应在低呼叫率的样本和SNPs上比高呼叫率的样本和SNPs更严重,而在杂合型呼叫上比纯合子基因型呼叫更严重。批次大小和组成影响使用BRLMM的GWA型调用结果。批次大小的差异越大,影响越大。批次中样品的同质性越高,基因分型的一致性就越高。这种不一致会传播到下游关联分析中确定的显著关联的SNP列表。因此,应使用统一的大批次大小进行基因分型。此外,均质度较高的样品应放入同一批次。
Genome-wide association studies (GWAS) aim to identify genetic variants (usually single nucleotide polymorphisms [SNPs]) across the entire human genome that are associated with phenotypic traits such as disease status and drug response. Highly accurate and reproducible genotype calling are paramount since errors introduced by calling algorithms can lead to inflation of false associations between genotype and phenotype. Most genotype calling algorithms currently used for GWAS are based on multiple arrays. Because hundreds of gigabytes (GB) of raw data are generated from a GWAS, the samples are typically partitioned into batches containing subsets of the entire dataset for genotype calling. High call rates and accuracies have been achieved. However, the effects of batch size (i.e., number of chips analyzed together) and of batch composition (i.e., the choice of chips in a batch) on call rate and accuracy as well as the propagation of the effects into significantly associated SNPs identified have not been investigated. In this paper, we analyzed both the batch size and batch composition for effects on the genotype calling algorithm BRLMM using raw data of 270 HapMap samples analyzed with the Affymetrix Human Mapping 500 K array set. Using data from 270 HapMap samples interrogated with the Affymetrix Human Mapping 500 K array set, three different batch sizes and three different batch compositions were used for genotyping using the BRLMM algorithm. Comparative analysis of the calling results and the corresponding lists of significant SNPs identified through association analysis revealed that both batch size and composition affected genotype calling results and significantly associated SNPs. Batch size and batch composition effects were more severe on samples and SNPs with lower call rates than ones with higher call rates, and on heterozygous genotype calls compared to homozygous genotype calls. Batch size and composition affect the genotype calling results in GWAS using BRLMM. The larger the differences in batch sizes, the larger the effect. The more homogenous the samples in the batches, the more consistent the genotype calls. The inconsistency propagates to the lists of significantly associated SNPs identified in downstream association analysis. Thus, uniform and large batch sizes should be used to make genotype calls for GWAS. In addition, samples of high homogeneity should be placed into the same batch.