Identification of the expressome by machine learning on omics data

Identification of the expressome by machine learning on omics data
复制标题

DOI:
10.1073/pnas.1813645116
复制
发表时间:
2019-09-03
影响因子:
11.1
通讯作者:
Briggs, Steven P.
Briggs, Steven P.
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Sartor, Ryan C.;Noshay, Jaclyn;Briggs, Steven P.

文献摘要

被引文献

相似文献

由于全基因组复制产生的冗余或转座元件捕获和移动基因片段而产生许多假基因,植物基因组的准确注释仍然很复杂。基于转录组和蛋白质组训练数据的全基因组表观遗传标记的机器学习,可用于通过将所有假定的蛋白质编码基因分类为组成型沉默或能够表达来改进注释。表达的基因被细分为能够表达 mRNA 和蛋白质或仅表达 RNA,并且 CG 基因体甲基化仅与前一亚类相关。玉米自交系 B73 的参考基因组中已注释了 60,000 多个蛋白质编码基因。这些基因中大约三分之二被转录并被指定为过滤基因集(FGS)。我们训练有素的随机森林算法对基因进行分类是准确的,并且仅依赖于基因体内的组蛋白修饰或 DNA 甲基化模式;启动子甲基化并不重要。已知其他近交系转录显着不同的基因组,表明 FGS 是 B73 特异的。我们准确地对来自近交系特异性 DNA 甲基化模式的其他自交系中的转录基因组进行了分类。这种方法凸显了使用染色质信息来改进功能基因注释的潜力。
Accurate annotation of plant genomes remains complex due to the presence of many pseudogenes arising from whole-genome duplication-generated redundancy or the capture and movement of gene fragments by transposable elements. Machine learning on genome-wide epigenetic marks, informed by transcriptomic and proteomic training data, could be used to improve annotations through classification of all putative protein-coding genes as either constitutively silent or able to be expressed. Expressed genes were subclassified as able to express both mRNAs and proteins or only RNAs, and CG gene body methylation was associated only with the former subclass. More than 60,000 protein-coding genes have been annotated in the reference genome of maize inbred B73. About two-thirds of these genes are transcribed and are designated the filtered gene set (FGS). Classification of genes by our trained random forest algorithm was accurate and relied only on histone modifications or DNA methylation patterns within the gene body; promoter methylation was unimportant. Other inbred lines are known to transcribe significantly different sets of genes, indicating that the FGS is specific to B73. We accurately classified the sets of transcribed genes in additional inbred lines, arising from inbred-specific DNA methylation patterns. This approach highlights the potential of using chromatin information to improve annotations of functional genes.