Evaluating the utility of identity-by-descent segment numbers for relatedness inference via information theory and classification.

Evaluating the utility of identity-by-descent segment numbers for relatedness inference via information theory and classification.
复制标题

DOI:
10.1093/g3journal/jkac072
复制
发表时间:
2022-05-30
期刊:
G3 (Bethesda, Md.)
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

相似文献

尽管在遗传研究中对亲属进行分类的方法已经发展了几十年,但两两亲缘关系方法的召回率只有在一级到三级亲缘关系中才超过90%。表现最好的方法是利用血统识别片段,通常只使用亲属系数,而其他方法,包括最近共同祖先的估计(ERSA),使用亲属共享的片段数量。为了量化在亲缘关系推断中使用片段数的潜力,我们利用信息论措施来分析来自模拟亲属的精确(即由模拟器产生)血统识别片段。在一系列设置中,我们发现亲缘度与亲缘系数和段数组成的元组之间的互信息平均比亲缘度与亲缘系数之间的互信息大4.6%。通过构建贝叶斯分类器来使用不同的特征集预测第一到六度关系,我们进一步评估了下降身份分段数的效用。当用精确的片段进行训练和测试时,片段数的包含将第二到第六度亲属的召回率提高了0.28%到3%。然而,当使用推断片段时,召回率每度提高不到1.8%,这表明由于血统识别检测准确性的限制。最后,我们将包含段数的贝叶斯分类器与ERSA和IBIS进行了比较,发现了可比较的召回率,贝叶斯分类器和ERSA在不同程度上都略微优于对方。总体而言,本研究表明,根据血统识别的片段数可以改善相关性推断,但目前基于SNP阵列的检测方法的错误在实践中会产生抑制信号。
Despite decades of methods development for classifying relatives in genetic studies, pairwise relatedness methods’ recalls are above 90% only for first through third-degree relatives. The top-performing approaches, which leverage identity-by-descent segments, often use only kinship coefficients, while others, including estimation of recent shared ancestry (ERSA), use the number of segments relatives share. To quantify the potential for using segment numbers in relatedness inference, we leveraged information theory measures to analyze exact (i.e. produced by a simulator) identity-by-descent segments from simulated relatives. Over a range of settings, we found that the mutual information between the relatives’ degree of relatedness and a tuple of their kinship coefficient and segment number is on average 4.6% larger than between the degree and the kinship coefficient alone. We further evaluated identity-by-descent segment number utility by building a Bayes classifier to predict first through sixth-degree relationships using different feature sets. When trained and tested with exact segments, the inclusion of segment numbers improves the recall by between 0.28% and 3% for second through sixth-degree relatives. However, the recalls improve by less than 1.8% per degree when using inferred segments, suggesting limitations due to identity-by-descent detection accuracy. Last, we compared our Bayes classifier that includes segment numbers with both ERSA and IBIS and found comparable recalls, with the Bayes classifier and ERSA slightly outperforming each other across different degrees. Overall, this study shows that identity-by-descent segment numbers can improve relatedness inference, but errors from current SNP array-based detection methods yield dampened signals in practice.
DOI: 10.1371/journal.pone.0087357
发表时间: 2014
期刊: PloS one
影响因子: 3.7
作者:
Ross BC
通讯作者: Ross BC
DOI: 10.1038/s41586-018-0579-z
发表时间: 2018-10
期刊: Nature
影响因子: 64.8
作者:
Bycroft C;Freeman C;Petkova D;Band G;Elliott LT;Sharp K;Motyer A;Vukcevic D;Delaneau O;O'Connell J;Cortes A;Welsh S;Young A;Effingham M;McVean G;Leslie S;Allen N;Donnelly P;Marchini J
通讯作者: Marchini J
DOI: 10.1093/molbev/msaa328
发表时间: 2021-05-04
影响因子: 10.7
作者:
Freyman WA;McManus KF;Shringarpure SS;Jewett EM;Bryc K;23 and Me Research Team;Auton A
通讯作者: Auton A
DOI: 10.1186/s13742-015-0047-8
发表时间: 2015
期刊: GigaScience
影响因子: 9.2
作者:
Chang CC;Chow CC;Tellier LC;Vattikuti S;Purcell SM;Lee JJ
通讯作者: Lee JJ
DOI: 10.1016/j.eswa.2014.04.019
发表时间: 2014-10-15
影响因子: 8.5
作者:
Hoque, N.;Bhattacharyya, D. K.;Kalita, J. K.
通讯作者: Kalita, J. K.