Evaluation of microhaplotype panels for complex kinship analysis using massively parallel sequencing

Evaluation of microhaplotype panels for complex kinship analysis using massively parallel sequencing
复制标题

DOI:
10.1016/j.fsigen.2023.102887
复制
发表时间:
2023-05-18
影响因子:
3.1
通讯作者:
Liang,Weibo
Liang,Weibo
中科院分区:
医学2区
文献类型:
--
作者:
Xue,Jiaming;Tan,Mengyu;Liang,Weibo

文献摘要

相似文献

近年来,微单倍型(microhaplotypes,MHs)已成为法医遗传学领域的研究热点。传统的MH仅包含在短片段内紧密连接的SNP。在这里,我们扩大了一般MH的概念,包括短的InDel。复杂亲属关系鉴定在灾害受害人身份识别和刑事侦查中具有重要作用。对于远亲(例如,第三级),需要许多遗传标记来增强亲属关系测试的能力。本研究以千人基因组计划中的中国南方汉族人群为研究对象,在全基因组范围内筛选由220 bp以内的两个或两个以上变异(InDel或SNP)组成的新的MH标记。成功开发了基于NGS的67重MH面板(面板B),并对124个无关个体样品进行测序以获得群体遗传数据,包括等位基因和等位基因频率。据我们所知,在67个遗传标记中,有65个MH是新发现的,其中32个MH的有效等位基因数(Ae)值大于5.0。群体平均Ae为5.34,平均杂合度为0.7352。接下来,收集来自先前研究的53个MH作为图A(平均Aeof 7.43),并且通过组合图A和B形成具有87个MH(平均Aeof 7.02)的图C。我们调查了这三个面板在亲属关系分析中的效用(父母-子女,全兄弟姐妹,第二度,第三度,第四度和第五度亲属),面板C表现出比其他两个面板更好的性能。图C能够在真实的系谱数据中将父母-子女、全同胞和二级亲属二人组与无关对照分开,在模拟的二级二人组中具有0.11%的小的假检验水平(FTL)。对于更远的关系,FTL要高得多:三级为8.99%,四级为35.46%,五级为61.55%。当一个精心挑选的额外的亲戚是已知的,这可能会提高测试力量的远亲关系分析。Q家系(2-5和2-7)和W家系(3-18和3-19)的两对双生子在所有检测的MH中具有相同的基因型,导致将叔侄二人组归类为亲子二人组的错误结论。此外,C组在亲子鉴定中表现出很大的排除近亲(二级和三级亲属)的能力。在18,246个真实的和10,000个模拟的无关对中,没有一个被误解为在log 10(LR)截止值为4的2度内的亲属。本文所提出的面板可以为复杂亲属关系的分析提供补充力量。
In recent years, microhaplotypes (MHs) have become a research hotspot within the field of forensic genetics. Traditional MHs contain only SNPs that are closely linked within short fragments. Herein, we broaden the concept of general MHs to include short InDels. Complex kinship identification plays an important role in disaster victim identification and criminal investigations. For distant relatives (e.g., 3rd-degree), many genetic markers are required to enhance power of kinship testing. We performed genome-wide screening for new MH markers composed of two or more variants (InDel or SNP) within 220 bp based on the Chinese Southern Han from the 1000 Genomes Project. An NGS-based 67plex MH panel (Panel B) was successfully developed, and 124 unrelated individual samples were sequenced to obtain population genetic data, including alleles and allele frequencies. Of the 67 genetic markers, 65 MHs were, as far as we know, newly discovered, and 32 MHs had effective number of allele (Ae) values greater than 5.0. The average Aeand heterozygosity of the panel were 5.34 and 0.7352, respectively. Next, 53 MHs from a previous study were collected as Panel A (average Aeof 7.43), and Panel C with 87 MHs (average Aeof 7.02) was formed by combining Panels A and B. We investigated the utility of these three panels in kinship analysis (parent-child, full siblings, 2nd-degree, 3rd-degree, 4th-degree, and 5th-degree relatives), with Panel C exhibiting better performance than the two other panels. Panel C was able to separate parent-child, full-sibling, and 2nd-degree relative duos from unrelated controls in real pedigree data, with a small false testing level (FTL) of 0.11% in simulated 2nd-degree duos. For more distant relationships, the FTL was much higher: 8.99% for 3rd-degree, 35.46% for 4th-degree, and 61.55% for 5th-degree. When a carefully chosen extra relative was known, this may enhance the testing power for distant kinship analysis. Two twins from the Q family (2–5 and 2–7) and W family (3–18 and 3–19) shared the same genotypes in all tested MHs, which led to the incorrect conclusion that an uncle-nephew duo was classified as a parent-child duo. In addition, Panel C showed great capacity for excluding close relatives (2nd-degree and 3rd-degree relatives) during paternity tests. Among 18,246 real and 10,000 simulated unrelated pairs, none were misinterpreted as a relative within 2nd-degree at a log10(LR) cutoff of 4. The panels presented herein could provide supplementary power for the analysis of complex kinship.