Using Mendelian inheritance errors as quality control criteria in whole genome sequencing data set.

Using Mendelian inheritance errors as quality control criteria in whole genome sequencing data set.
复制标题

DOI:
10.1186/1753-6561-8-s1-s21
复制
发表时间:
2014-01-01
期刊:
影响因子:
--
通讯作者:
Martin, Lisa J
Martin, Lisa J
中科院分区:
其他
文献类型:
--
作者:
Pilipenko, Valentina V;He, Hua;Martin, Lisa J

文献摘要

被引文献

相似文献

尽管人们普遍认识到全基因组测序的技术和分析复杂性,但数据清理和质量控制的最佳实践尚未定义。基于家族的数据可用于指导非基于家族的数据中特定质量控制指标的标准化。鉴于突变率较低,孟德尔遗传错误很可能是由于错误的基因型识别造成的。因此,我们的目标是确定决定孟德尔遗传错误的特征。为了实现这一目标,我们使用了来自遗传分析研讨会 18 的基于 3 号染色体全基因组测序家族的数据。孟德尔遗传错误作为 GAW18 数据集的一部分提供。此外,对于二元变体,我们使用 PLINK 计算了孟德尔遗传错误。根据我们的分析,非二元单核苷酸变异体固有地存在大量孟德尔遗传错误。此外,在二元变体中,孟德尔遗传错误不是随机分布的。事实上,我们确定了 3 个孟德尔遗传错误峰,这些峰富含重复元素。然而,通过包含测序文件中的单个过滤器可以减少这些峰值。总之,我们证明了错误的测序调用在基因组中非随机分布,质量控制指标可以显着减少孟德尔遗传错误的数量。适当的质量控制将允许遗传数据的最佳利用,以实现全基因组测序的全部潜力。
Although the technical and analytic complexity of whole genome sequencing is generally appreciated, best practices for data cleaning and quality control have not been defined. Family based data can be used to guide the standardization of specific quality control metrics in nonfamily based data. Given the low mutation rate, Mendelian inheritance errors are likely as a result of erroneous genotype calls. Thus, our goal was to identify the characteristics that determine Mendelian inheritance errors. To accomplish this, we used chromosome 3 whole genome sequencing family based data from the Genetic Analysis Workshop 18. Mendelian inheritance errors were provided as part of the GAW18 data set. Additionally, for binary variants we calculated Mendelian inheritance errors using PLINK. Based on our analysis, nonbinary single-nucleotide variants have an inherently high number of Mendelian inheritance errors. Furthermore, in binary variants, Mendelian inheritance errors are not randomly distributed. Indeed, we identified 3 Mendelian inheritance error peaks that were enriched with repetitive elements. However, these peaks can be lessened with the inclusion of a single filter from the sequencing file. In summary, we demonstrated that erroneous sequencing calls are nonrandomly distributed across the genome and quality control metrics can dramatically reduce the number of mendelian inheritance errors. Appropriate quality control will allow optimal use of genetic data to realize the full potential of whole genome sequencing.