Automated analysis of immunosequencing datasets reveals novel immunoglobulin D genes across diverse species

Automated analysis of immunosequencing datasets reveals novel immunoglobulin D genes across diverse species
复制标题

DOI:
10.1371/journal.pcbi.1007837
复制
发表时间:
2020-04-01
影响因子:
4.3
通讯作者:
Safonova, Yana
Safonova, Yana
中科院分区:
生物学2区
文献类型:
--
作者:
Bhardwaj, Vinnu;Franceschetti, Massimo;Safonova, Yana

文献摘要

被引文献

相似文献

作者摘要抗体可与多种抗原特异性结合,是适应性免疫系统的关键组成部分。免疫测序已成为一种生成数百万读数的首选方法,这些读数对抗体库进行采样,并提供监测对疾病和疫苗接种的免疫反应的见解。以前的免疫基因组学研究大多数依赖于免疫球蛋白位点中的参考种系基因,而不是特定患者的种系基因。这种方法是有缺陷的,因为已知种系基因组不完整(特别是对于非欧洲人类和非人类物种)并且包含由测序和注释错误导致的等位基因。从免疫测序数据中从头推断多样性 (D) 基因的问题一直悬而未决,直到 2019 年 IgScout 算法被开发出来。我们通过开发用于 D 基因重建的概率 MINING-D 算法来解决 IgScout 的局限性,并推断标准数据库中不存在的跨多个物种的多个 D 基因。免疫球蛋白基因是通过 V(D)J 重组形成的,该重组将变量 (V)、多样性 (D) 结合起来,并加入多样性 (D) (J)种系基因。由于种系基因的变异与各种疾病有关,因此个性化免疫基因组学的重点是寻找不同患者的种系基因的等位基因。尽管 V 和 J 基因的重建是一个经过充分研究的问题,但重建 D 基因的更具挑战性的任务仍然悬而未决,直到 2019 年 IgScout 算法被开发出来。在这项工作中,我们通过开发用于 D 基因重建的概率 MINING-D 算法来解决 IgScout 的局限性,将其应用于来自多个物种的数百个免疫测序数据集,并通过分析不同的全基因组测序数据集和对杂合 V 进行单倍型分析来验证新推断的 D 基因基因。
Author summaryAntibodies provide specific binding to an enormous range of antigens and represent a key component of the adaptive immune system. Immunosequencing has emerged as a method of choice for generating millions of reads that sample antibody repertoires and provides insights into monitoring immune response to disease and vaccination. Most of the previous immunogenomics studies rely on the reference germline genes in the immunoglobulin locus rather than the germline genes in a specific patient. This approach is deficient since the set of known germline genes is incomplete (particularly for non-European humans and non-human species) and contains alleles that resulted from sequencing and annotation errors. The problem of de novo inference of diversity (D) genes from immunosequencing data remained open until the IgScout algorithm was developed in 2019. We address limitations of IgScout by developing a probabilistic MINING-D algorithm for D gene reconstruction and infer multiple D genes across multiple species that are not present in standard databases.Immunoglobulin genes are formed through V(D)J recombination, which joins the variable (V), diversity (D), and joining (J) germline genes. Since variations in germline genes have been linked to various diseases, personalized immunogenomics focuses on finding alleles of germline genes across various patients. Although reconstruction of V and J genes is a well-studied problem, the more challenging task of reconstructing D genes remained open until the IgScout algorithm was developed in 2019. In this work, we address limitations of IgScout by developing a probabilistic MINING-D algorithm for D gene reconstruction, apply it to hundreds of immunosequencing datasets from multiple species, and validate the newly inferred D genes by analyzing diverse whole genome sequencing datasets and haplotyping heterozygous V genes.