Evolutionarily informed deep learning methods for predicting relative transcript abundance from DNA sequence

Evolutionarily informed deep learning methods for predicting relative transcript abundance from DNA sequence
复制标题

DOI:
10.1073/pnas.1814551116
复制
发表时间:
2019-03-19
影响因子:
11.1
通讯作者:
Wang, Hai
Wang, Hai
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Washburn, Jacob D.;Mejia-Guerra, Maria Katherine;Wang, Hai

文献摘要

被引文献

相似文献

深度学习方法已经使许多领域的预测发生了革命性的变化,并在分子生物学和遗传学中显示出同样的潜力。然而,以目前的形式应用这些方法会忽略生物系统内的进化依赖关系,并可能导致假阳性和虚假结论。我们开发了两种方法来解释机器学习模型中的进化相关性:(I)基因家族引导的分裂和(Ii)直系同源对比。第一种方法通过约束模型训练和测试集以包括不同的基因家族来解释进化。第二种方法使用进化知情的同源基因之间的比较,在训练过程中控制和利用进化差异。这两种方法是在mRNA表达水平预测的背景下进行探索和验证的,其ROC曲线下面积(AuROC)值在0.75到0.94之间。模型重量检查显示了生物学上可解释的模式,导致假设3‘UTR对于微调mRNA丰度水平更重要,而5’UTR对于大规模变化更重要。
Deep learning methodologies have revolutionized prediction in many fields and show potential to do the same in molecular biology and genetics. However, applying these methods in their current forms ignores evolutionary dependencies within biological systems and can result in false positives and spurious conclusions. We developed two approaches that account for evolutionary relatedness in machine learning models: (i) gene-family-guided splitting and (ii) ortholog contrasts. The first approach accounts for evolution by constraining model training and testing sets to include different gene families. The second approach uses evolutionarily informed comparisons between orthologous genes to both control for and leverage evolutionary divergence during the training process. The two approaches were explored and validated within the context of mRNA expression level prediction and have the area under the ROC curve (auROC) values ranging from 0.75 to 0.94. Model weight inspections showed biologically interpretable patterns, resulting in the hypothesis that the 3' UTR is more important for fine-tuning mRNA abundance levels while the 5' UTR is more important for large-scale changes.