Systematic analysis of dark and camouflaged genes reveals disease-relevant genes hiding in plain sight

Systematic analysis of dark and camouflaged genes reveals disease-relevant genes hiding in plain sight
复制标题

DOI:
10.1186/s13059-019-1707-2
复制
发表时间:
2019-05-20
期刊:
影响因子:
12.3
通讯作者:
Fryer, John D.
Fryer, John D.
中科院分区:
生物学1区
文献类型:
--
作者:
Ebbert, Mark T. W.;Jensen, Tanner D.;Fryer, John D.

文献摘要

被引文献

相似文献

背景人类基因组包含暗基因区域,使用标准的短读测序技术无法充分组装或比对,从而阻止研究人员识别这些基因区域内可能与人类疾病相关的突变。在这里,我们识别有少量可映射读数的区域,我们将其称为深度暗区域,以及其他具有模糊对齐的区域,称为伪装区域。结果基于标准的全基因组Illumina测序数据,我们从6054个基因体中识别出对人类健康、发育和生殖至关重要的36,794个暗区。在这些基因体中,8.7%是完全暗的,35.2%是5%的暗。我们确定了748个基因的蛋白质编码外显子中存在的暗区。10x基因组、PacBio和牛津纳米孔技术公司的连锁阅读或长阅读测序技术将暗蛋白质编码区分别减少到约50.5%、35.6%和9.6%。我们提出了一种算法来解析大多数伪装区域,并将其应用到阿尔茨海默病测序项目中。我们挽救了在阿尔茨海默病顶级基因CR1中罕见的十核苷酸移码缺失,在疾病病例中发现,但在对照组中未发现。结论尽管由于样本量不足,我们无法正式评估CR1移码突变与阿尔茨海默病的关联,但我们认为它值得在更大的队列中进行研究。仍然有数千个潜在的重要基因组区域被短读测序忽视,这些区域在很大程度上是通过长读技术解决的。
BackgroundThe human genome contains dark gene regions that cannot be adequately assembled or aligned using standard short-read sequencing technologies, preventing researchers from identifying mutations within these gene regions that may be relevant to human disease. Here, we identify regions with few mappable reads that we call dark by depth, and others that have ambiguous alignment, called camouflaged. We assess how well long-read or linked-read technologies resolve these regions.ResultsBased on standard whole-genome Illumina sequencing data, we identify 36,794 dark regions in 6054 gene bodies from pathways important to human health, development, and reproduction. Of these gene bodies, 8.7% are completely dark and 35.2% are 5% dark. We identify dark regions that are present in protein-coding exons across 748 genes. Linked-read or long-read sequencing technologies from 10x Genomics, PacBio, and Oxford Nanopore Technologies reduce dark protein-coding regions to approximately 50.5%, 35.6%, and 9.6%, respectively. We present an algorithm to resolve most camouflaged regions and apply it to the Alzheimer's Disease Sequencing Project. We rescue a rare ten-nucleotide frameshift deletion in CR1, a top Alzheimer's disease gene, found in disease cases but not in controls.ConclusionsWhile we could not formally assess the association of the CR1 frameshift mutation with Alzheimer's disease due to insufficient sample-size, we believe it merits investigating in a larger cohort. There remain thousands of potentially important genomic regions overlooked by short-read sequencing that are largely resolved by long-read technologies.