The struggle to find reliable results in exome sequencing data: filtering out Mendelian errors

The struggle to find reliable results in exome sequencing data: filtering out Mendelian errors
复制标题

DOI:
10.3389/fgene.2014.00016
复制
发表时间:
2014-02-12
影响因子:
3.7
通讯作者:
Kaufman, Kenneth M.
Kaufman, Kenneth M.
中科院分区:
生物学3区
文献类型:
--
作者:
Patel, Zubin H.;Kottyan, Leah C.;Kaufman, Kenneth M.

文献摘要

被引文献

相似文献

下一代测序研究以相对成本和时间有效的方式生成大量遗传数据,并为识别导致疾病表型的候选致病变体提供了前所未有的机会。这些研究的一个挑战是通过当前技术产生测序伪像。为了识别和表征区分假阳性变体与真实变体的特性,我们使用从三个来源(血液,口腔细胞和唾液)分离的DNA对一个孩子和父母(一个三人组)进行测序。三人策略使我们能够识别先证者中不可能从父母遗传的变异(孟德尔错误),并且很可能表明测序伪像。质量控制测量进行了检查,发现三个测量,以确定最大数量的孟德尔错误。这些包括读取深度、基因型质量评分和交替等位基因比率。过滤这些测量的变体去除了近似95%的孟德尔误差,同时保留了80%的调用变体。这些过滤器独立使用。不同来源的同一样品经过滤后的一致率为99.99%,而过滤前的一致率为87%。这种高一致性表明,不同来源的DNA可以用于三重研究,而不会影响鉴定致病多态性的能力。为了促进下一代测序数据的分析,我们开发了辛辛那提测序信息分析套件(CASSI)来存储测序文件、元数据(例如,相关性信息)、文件版本化、数据过滤、变体注释,并鉴定遵循从头、罕见隐性纯合或复合杂合遗传模型的候选致病多态性。我们的结论是,数据清洗过程提高了变异体的信噪比,并有助于识别候选致病多态性。
Next Generation Sequencing studies generate a large quantity of genetic data in a relatively cost and time efficient manner and provide an unprecedented opportunity to identify candidate causative variants that lead to disease phenotypes. A challenge to these studies is the generation of sequencing artifacts by current technologies. To identify and characterize the properties that distinguish false positive variants from true variants, we sequenced a child and both parents (one trio) using DNA isolated from three sources (blood, buccal cells, and saliva). The trio strategy allowed us to identify variants in the proband that could not have been inherited from the parents (Mendelian errors) and would most likely indicate sequencing artifacts. Quality control measurements were examined and three measurements were found to identify the greatest number of Mendelian errors. These included read depth, genotype quality score, and alternate allele ratio. Filtering the variants on these measurements removed similar to 95% of the Mendelian errors while retaining 80% of the called variants. These filters were applied independently. After filtering, the concordance between identical samples isolated from different sources was 99.99% as compared to 87% before filtering. This high concordance suggests that different sources of DNA can be used in trio studies without affecting the ability to identify causative polymorphisms. To facilitate analysis of next generation sequencing data, we developed the Cincinnati Analytical Suite for Sequencing Informatics (CASSI) to store sequencing files, metadata (eg. relatedness information), file versioning, data filtering, variant annotation, and identify candidate causative polymorphisms that follow either de novo, rare recessive homozygous or compound heterozygous inheritance models. We conclude the data cleaning process improves the signal to noise ratio in terms of variants and facilitates the identification of candidate disease causative polymorphisms.