Cleaning Genotype Data from Diversity Outbred Mice

Cleaning Genotype Data from Diversity Outbred Mice
复制标题

DOI:
10.1534/g3.119.400165
复制
发表时间:
2019-05-01
影响因子:
2.6
通讯作者:
Churchill, Gary A.
Churchill, Gary A.
中科院分区:
生物学3区
文献类型:
--
作者:
Broman, Karl W.;Gatti, Daniel M.;Churchill, Gary A.

文献摘要

被引文献

相似文献

数据清理是大多数统计分析中重要的第一步,包括绘制导致数量性状变异的遗传位点图谱。在这里,我们使用来自 291 只多样性远交 (DO) 小鼠的 MegaMUGA 阵列数据,说明了多亲群体(源自两个以上创始品系的实验杂交)的质量控制和基于阵列的基因分型数据清理的方法。我们的方法采用数据可视化,可以揭示个体小鼠水平或个体 SNP 标记的问题。我们发现每只小鼠缺失基因型的比例是样本质量的有效指标。我们使用 X 和 Y 染色体上的 SNP 的微阵列探针强度来确认每只小鼠的性别,并使用成对小鼠之间匹配的 SNP 基因型的比例来检测样本重复。我们使用隐马尔可夫模型 (HMM) 重建每个小鼠基因组的创始人单倍型镶嵌,以估计交叉数量并识别潜在的基因分型错误。为了评估标记质量,我们发现缺失数据和基因分型错误率是最有效的诊断。我们还检查了 SNP 基因型频率,其中标记根据创始菌株中的次要等位基因频率进行分组。对于具有高表观错误率的标记,等位基因特异性探针强度的散点图可以揭示错误基因型识别的根本原因。包含或排除低质量样本的决定可能会对给定研究的绘图结果产生重大影响。我们发现低质量标记对特定研究的影响通常很小,但报告有问题的标记可以提高基因分型阵列在许多研究中的实用性。
Data cleaning is an important first step in most statistical analyses, including efforts to map the genetic loci that contribute to variation in quantitative traits. Here we illustrate approaches to quality control and cleaning of array-based genotyping data for multiparent populations (experimental crosses derived from more than two founder strains), using MegaMUGA array data from a set of 291 Diversity Outbred (DO) mice. Our approach employs data visualizations that can reveal problems at the level of individual mice or with individual SNP markers. We find that the proportion of missing genotypes for each mouse is an effective indicator of sample quality. We use microarray probe intensities for SNPs on the X and Y chromosomes to confirm the sex of each mouse, and we use the proportion of matching SNP genotypes between pairs of mice to detect sample duplicates. We use a hidden Markov model (HMM) reconstruction of the founder haplotype mosaic across each mouse genome to estimate the number of crossovers and to identify potential genotyping errors. To evaluate marker quality, we find that missing data and genotyping error rates are the most effective diagnostics. We also examine the SNP genotype frequencies with markers grouped according to their minor allele frequency in the founder strains. For markers with high apparent error rates, a scatterplot of the allele-specific probe intensities can reveal the underlying cause of incorrect genotype calls. The decision to include or exclude low-quality samples can have a significant impact on the mapping results for a given study. We find that the impact of low-quality markers on a given study is often minimal, but reporting problematic markers can improve the utility of the genotyping array across many studies.