Impact of SNP microarray analysis of compromised DNA on kinship classification success in the context of investigative genetic genealogy

Impact of SNP microarray analysis of compromised DNA on kinship classification success in the context of investigative genetic genealogy
复制标题

DOI:
10.1016/j.fsigen.2021.102625
复制
发表时间:
2021-11-06
影响因子:
3.1
通讯作者:
Kayser, Manfred
Kayser, Manfred
中科院分区:
医学2区
文献类型:
--
作者:
de Vries, Jard H.;Kling, Daniel;Kayser, Manfred

文献摘要

被引文献

相似文献

利用微阵列技术生成的单核苷酸多态性(SNP)数据已被用于通过从识别包括在可访问的基因组数据库中的未知肇事者的亲属获得的调查线索来解决谋杀案件,这种方法被称为调查遗传谱系(IGG)。然而,开发SNP微阵列用于相对高的输入DNA数量和质量,而通常可从犯罪现场染色剂获得的DNA具有低DNA数量和质量,并且从受损DNA获得的SNP微阵列数据在很大程度上缺失。通过将Illumina Global Screening Array(GSA)应用于264个数量和质量发生系统性改变的DNA样本,我们经验性地测试了受损DNA的SNP微阵列分析对亲缘关系分类成功的影响,与IGG相关。根据专家推荐的输入DNA质量和数量的参考数据,估计受损DNA样本的基因型准确度,并模拟不同程度的亲属数据。虽然输入DNA量从200 ng逐步减少到6.25 pg导致SNP调用率降低和基因分型错误增加,但亲属分类成功率并没有降低到兄弟姐妹和第一堂兄弟姐妹的250 pg,第二堂兄弟姐妹的1 ng,而在25 pg及以下亲属分类成功率为零。通过增加DNA片段化而逐步降低输入DNA质量导致基因分型准确性和亲缘关系分类成功率降低,在平均DNA片段大小为150个碱基对时,亲缘关系分类成功率降至零。在模拟案例和骨骼样本中结合降低的DNA数量和质量进一步突出了可能性和局限性。总的来说,GSA分析取得了最大的亲属分类成功,从800到200倍的输入DNA数量比专家推荐的低,虽然DNA质量也起着关键作用,而受损的DNA产生假阴性亲属分类,而不是假阳性。
Single nucleotide polymorphism (SNP) data generated with microarray technologies have been used to solve murder cases via investigative leads obtained from identifying relatives of the unknown perpetrator included in accessible genomic databases, an approach referred to as investigative genetic genealogy (IGG). However, SNP microarrays were developed for relatively high input DNA quantity and quality, while DNA typically obtainable from crime scene stains is of low DNA quantity and quality, and SNP microarray data obtained from compromised DNA are largely missing. By applying the Illumina Global Screening Array (GSA) to 264 DNA samples with systematically altered quantity and quality, we empirically tested the impact of SNP microarray analysis of compromised DNA on kinship classification success, as relevant in IGG. Reference data from manufacturerrecommended input DNA quality and quantity were used to estimate genotype accuracy in the compromised DNA samples and for simulating data of different degree relatives. Although stepwise decrease of input DNA amount from 200 ng to 6.25 pg led to decreased SNP call rates and increased genotyping errors, kinship classification success did not decrease down to 250 pg for siblings and 1st cousins, 1 ng for 2nd cousins, while at 25 pg and below kinship classification success was zero. Stepwise decrease of input DNA quality via increased DNA fragmentation resulted in the decrease of genotyping accuracy as well as kinship classification success, which went down to zero at the average DNA fragment size of 150 base pairs. Combining decreased DNA quantity and quality in mock casework and skeletal samples further highlighted possibilities and limitations. Overall, GSA analysis achieved maximal kinship classification success from 800 to 200 times lower input DNA quantities than manufacturer-recommended, although DNA quality plays a key role too, while compromised DNA produced false negative kinship classifications rather than false positive ones.