Who's Who? Detecting and Resolving Sample Anomalies in Human DNA Sequencing Studies with Peddy

Who's Who? Detecting and Resolving Sample Anomalies in Human DNA Sequencing Studies with Peddy
复制标题

DOI:
10.1016/j.ajhg.2017.01.017
复制
发表时间:
2017-03-02
影响因子:
9.8
通讯作者:
Quinlan, Aaron R.
Quinlan, Aaron R.
中科院分区:
生物学1区
文献类型:
--
作者:
Pedersen, Brent S.;Quinlan, Aaron R.

文献摘要

被引文献

相似文献

如果来自队列的DNA样本被错误标记、交换或污染,或者如果其中包括非预期的个人,那么在人类DNA测序研究中发现基因的可能性会大大降低。不幸的是,这种错误的可能性很大,因为DNA样本在测序过程中经常被几个协议、实验室或科学家操纵。我们已经开发了一个名为Peddy的软件包,通过交互可视化和报告将所述的性别、亲属关系和祖先与从全基因组(WGS)或全外显子组(WES)测序得出的个体基因类型进行比较,来识别和促进此类错误的补救。佩迪使用机器学习模型预测样本的祖先,该模型针对1000基因组计划参考小组中不同祖先的个人进行了训练。Peddy促进了自动和交互,对样本交换、糟糕的测序质量和其他指示样本问题的可视检测,如果没有检测到这些问题,将阻碍发现。
The potential for genetic discovery in human DNA sequencing studies is greatly diminished if DNA samples from a cohort are mislabeled, swapped, or contaminated or if they include unintended individuals. Unfortunately, the potential for such errors is significant since DNA samples are often manipulated by several protocols, labs, or scientists in the process-of-sequencing. We have developed a software package, peddy, to identify and facilitate the remediation of such errors via interactive visualizations and reports comparing the stated sex, relatedness, and ancestry to what is inferred from the individual genotypes derived from whole-genome (WGS) or whole-exome (WES) sequencing. Peddy predicts a sample's ancestry using a machine learning model trained on individuals of diverse ancestries from the 1000 Genomes Project reference panel. Peddy facilitates both automated and interactive, visual detection of sample swaps, poor sequencing quality, and other indicators of sample problems that, if left undetected, would inhibit discovery.