Calling pangenes from plant genome alignments confirms presence-absence variation

Calling pangenes from plant genome alignments confirms presence-absence variation
复制标题

从植物基因组比对中调用泛基因证实了存在与不存在的变异

DOI:
10.1101/2023.01.03.520531
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Contreras-Moreira B
Contreras-Moreira B
中科院分区:
--
文献类型:
--
作者:
Contreras-Moreira B

文献摘要

参考文献

相似文献

随着新品种基因组的频繁发表,农作物中一致的基因注释变得越来越困难。最近测序的种质的基因集与参考种质的基因标识符不同,并且由于技术进步可能具有更高的质量。由于这些原因,需要定义泛基因,它代表基因模型的所有已知同线性直系同源物,并且可以链接回原始来源。泛基因集有效地总结了我们目前对作物编码潜力的理解,可用于为新品种的基因模型注释提供信息。在这里,我们提出了一种方法(get_pangenes)来识别和分析不偏向参考注释的泛基因。该方法涉及计算全基因组比对(WGA),用于估计基因模型重叠。在对拟南芥、水稻、小麦和大麦数据集进行基准测试后,我们发现两种不同的 WGA 算法(minimap2 和 GSAlign)产生相似的泛基因集。我们的结果表明,泛基因概括了已知的基于系统发育的直系学,同时在水稻中添加了额外的核心基因模型。更重要的是,get_pangenes 还可以产生与其他品种中注释的基因模型重叠的基因组片段 (gDNA) 簇。通过提升 CDS 序列,gDNA 簇可以帮助完善个体的基因模型,并确认或拒绝观察到的基因存在-不存在变异。文档和源代码可在 https://github.com/Ensembl/plant-scripts/tree/master/pangenes 获取。核心思想全基因组比对捕获重叠的基因模型和基因组片段。泛基因代表来自不同基因集的同源共线基因模型。Lift-over 可用于细化基因模型并确认基因存在-不存在变异。
Consistent gene annotation in crops is becoming harder as genomes for new cultivars are frequently published. Gene sets from recently sequenced accessions have different gene identifiers to those on the reference accession, and might be of higher quality due to technical advances. For these reasons there is a need to define pangenes, which represent all known syntenic orthologues for a gene model and can be linked back to the original sources. A pangene set effectively summarizes our current understanding of the coding potential of a crop and can be used to inform gene model annotation in new cultivars. Here we present an approach (get_pangenes) to identify and analyze pangenes that is not biased towards the reference annotation. The method involves computing Whole Genome Alignments (WGA), which are used to estimate gene model overlaps. After a benchmark onArabidopsis, rice, wheat and barley datasets, we find that two different WGA algorithms (minimap2 and GSAlign) produce similar pangene sets. Our results show that pangenes recapitulate known phylogeny-based orthologies while adding extra core gene models in rice. More importantly, get_pangenes can also produce clusters of genome segments (gDNA) that overlap with gene models annotated in other cultivars. By lifting-over CDS sequences, gDNA clusters can help refine gene models across individuals and confirm or reject observed gene Presence-Absence Variation. Documentation and source code are available at https://github.com/Ensembl/plant-scripts/tree/master/pangenes.Core ideasWhole Genome Alignments capture overlapping gene models and genome segments.A pangene represents homologous collinear gene models from different gene sets.Lift-over can be used to refine gene models and to confirm gene Presence-Absence Variation.
DOI: 10.1007/978-1-0716-2067-0_2
发表时间: 2022
期刊: Methods in molecular biology (Clifton, N.J.)
影响因子: --
作者:
Contreras-Moreira B;Naamati G;Rosello M;Allen JE;Hunt SE;Muffato M;Gall A;Flicek P
通讯作者: Flicek P
DOI: 10.1093/bioinformatics/btac308
发表时间: 2022-06-27
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
通讯作者: --
DOI: 10.1016/j.cub.2022.04.085
发表时间: 2022-06-20
期刊: CURRENT BIOLOGY
影响因子: 9.2
作者:
Weisman, Caroline M.;Murray, Andrew W.;Eddy, Sean R.
通讯作者: Eddy, Sean R.