De novo pattern discovery enables robust assessment of functional consequences of non-coding variants.

De novo pattern discovery enables robust assessment of functional consequences of non-coding variants.
复制标题

从头模式发现能够对非编码变体的功能后果进行稳健评估。

DOI:
10.1093/bioinformatics/bty826
复制
发表时间:
2019
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Li,Bingshan
Li,Bingshan
中科院分区:
--
文献类型:
--
作者:
Yang,Hai;Chen,Rui;Wang,Quan;Wei,Qiang;Ji,Ying;Zheng,Guangze;Zhong,Xue;Cox,NancyJ;Li,Bingshan

文献摘要

相似文献

考虑到基因组区域的复杂性,优先考虑非编码变体的功能效应仍然是一个挑战。虽然已经提出了几个框架来评估非编码变体的功能,但其中大多数使用“黑盒”方法,将任务简化为致病性/良性分类问题,这忽略了变体的独特调控机制,并导致不太理想的性能。在这项研究中,我们开发了DVAR,一个无监督的框架,利用各种生化和进化的证据来区分基因调控类别的变体,并评估其综合功能的影响,同时.ResultsDVAR performedde novopattern发现在高维数据,并确定了五个监管集群的非编码变体。利用对多功能模式的新见解,它可以测量变体的类间和类内功能含义,以实现准确的优先级排序。与其他两类学习方法相比,它在识别临床显著变异、精细定位的GWAS变异、eQTL和表达调节变异方面表现出更好的性能。此外,它对通过基因组编辑(如CRISPR-Cas9)验证的疾病因果变异具有上级性能,这可以为整个基因组的基因组编辑技术提供预选策略。最后,在BioVU和UK Biobank(两个与完整电子健康记录相关的大规模DNA生物库)中进行评估,DVAR证明了其在优先考虑与医学表型相关的非编码变体方面的有效性。可用性和实施C++和Python源代码,整个基因组的预先计算的DVAR簇标签和DVAR分数可在https://www.vumc.org/cgg/dvar.Supplementary信息中获得补充数据可在Bioinformaticsonline。
MotivationGiven the complexity of genome regions, prioritize the functional effects of non-coding variants remains a challenge. Although several frameworks have been proposed for the evaluation of the functionality of non-coding variants, most of them used ‘black boxes’ methods that simplify the task as the pathogenicity/benign classification problem, which ignores the distinct regulatory mechanisms of variants and leads to less desirable performance. In this study, we developed DVAR, an unsupervised framework that leverage various biochemical and evolutionary evidence to distinguish the gene regulatory categories of variants and assess their comprehensive functional impact simultaneously.ResultsDVAR performedde novopattern discovery in high-dimensional data and identified five regulatory clusters of non-coding variants. Leveraging the new insights into the multiple functional patterns, it measures both the between-class and the within-class functional implication of the variants to achieve accurate prioritization. Compared to other two-class learning methods, it showed improved performance in identification of clinically significant variants, fine-mapped GWAS variants, eQTLs and expression-modulating variants. Moreover, it has superior performance on disease causal variants verified by genome-editing (like CRISPR-Cas9), which could provide a pre-selection strategy for genome-editing technologies across the whole genome. Finally, evaluated in BioVU and UK Biobank, two large-scale DNA biobanks linked to complete electronic health records, DVAR demonstrated its effectiveness in prioritizing non-coding variants associated with medical phenotypes.Availability and implementationThe C++ and Python source codes, the pre-computed DVAR-cluster labels and DVAR-scores across the whole genome are available at https://www.vumc.org/cgg/dvar.Supplementary informationSupplementary data are available atBioinformaticsonline.