A computational method for prediction of matrix proteins in endogenous retroviruses.

A computational method for prediction of matrix proteins in endogenous retroviruses.
复制标题

预测内源逆转录病毒基质蛋白的计算方法

DOI:
10.1371/journal.pone.0176909
复制
发表时间:
2017
期刊:
影响因子:
3.7
通讯作者:
Zhang X
Zhang X
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Ma Y;Liu R;Lv H;Han J;Zhong D;Zhang X

文献摘要

被引文献

相似文献

人类内源性逆转录病毒(HERV)编码活性逆转录病毒蛋白,可能与癌症和其他疾病的进展有关。基质蛋白(MA)位于逆转录病毒组特异性抗原基因(GAG)中,在大多数哺乳动物逆转录病毒中与病毒包膜糖蛋白相关,可能参与病毒颗粒的组装、运输和萌发。然而,到目前为止,ERV中注释MA的数量仍然处于较低水平。到目前为止,还没有提出一种计算方法来预测GAS中MAS的准确起始和结束坐标。本文提出了一种识别ERV中MA的计算方法。设计了一种分治技术,并将其应用到传统的预测模型中,以获得更好的处理不同长度的基因序列的结果。对起始点和终止点分别进行预测,然后按间隔组合。应用并比较了三种不同的算法:加权支持向量机、加权极限学习机和随机森林算法。由随机森林模型生成的5次交叉验证的起始位点和终止位点的G-−均值(敏感性和特异性的几何平均值)分别为0.9869和0.9755,是所用算法中最高的。我们的预测模型结合了RF和WSVM算法,以获得最佳的预测结果。在收集到的完整mA的ERV序列中,98.4%的序列(总共125个)可以被该模型准确地预测。用该模型对118个家系的94,671个HERV序列进行了扫描,在人类染色体上预测了104个新的MAs。文中还分析了假设MA的分布和模型参数的优化。我们的预测方法也被扩展到其他逆转录病毒,并获得了令人满意的结果。
Human endogenous retroviruses (HERVs) encode active retroviral proteins, which may be involved in the progression of cancer and other diseases. Matrix protein (MA), in group-specific antigen genes (gag) of retroviruses, is associated with the virus envelope glycoproteins in most mammalian retroviruses and may be involved in virus particle assembly, transport and budding. However, the amount of annotated MAs in ERVs is still at a low level so far. No computational method to predict the exact start and end coordinates of MAs in gags has been proposed yet. In this paper, a computational method to identify MAs in ERVs is proposed. A divide and conquer technique was designed and applied to the conventional prediction model to acquire better results when dealing with gene sequences with various lengths. Initiation sites and termination sites were predicted separately and then combined according to their intervals. Three different algorithms were applied and compared: weighted support vector machine (WSVM), weighted extreme learning machine (WELM) and random forest (RF). G − mean (geometric mean of sensitivity and specificity) values of initiation sites and termination sites under 5-fold cross validation generated by random forest models are 0.9869 and 0.9755 respectively, highest among the algorithms applied. Our prediction models combine RF & WSVM algorithms to achieve the best prediction results. 98.4% of all the collected ERV sequences with complete MAs (125 in total) could be predicted exactly correct by the models. 94,671 HERV sequences from 118 families were scanned by the model, 104 new putative MAs were predicted in human chromosomes. Distributions of the putative MAs and optimizations of model parameters were also analyzed. The usage of our predicting method was also expanded to other retroviruses and satisfying results were acquired.