Continuous Distributed Representation of Biological Sequences for Deep Proteomics and Genomics.

Continuous Distributed Representation of Biological Sequences for Deep Proteomics and Genomics.
复制标题

DOI:
10.1371/journal.pone.0141287
复制
发表时间:
2015
期刊:
影响因子:
3.7
通讯作者:
Mofrad MR
Mofrad MR
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Asgari E;Mofrad MR

文献摘要

被引文献

相似文献

介绍了一种新的生物序列表示和特征提取方法。生物载体(BioVec)是指与蛋白质(氨基酸序列)的蛋白质载体(ProtVec)和基因序列的基因载体(GeneVec)一起统称为生物序列的生物载体,这种表示法可广泛应用于蛋白质组学和基因组学的深度学习中。在本文中,我们将重点介绍可用于广泛的生物信息学研究的蛋白质载体,如家族分类、蛋白质可视化、结构预测、无序蛋白质识别和蛋白质相互作用预测。在该方法中,我们采用人工神经网络的方法,用单个稠密的n维向量来表示蛋白质序列。为了对该方法进行评估,我们将其应用于从Swiss-Prot获得的属于7,027个蛋白质家族的324,018个蛋白质序列的分类,获得了93%±0.06%的平均家族分类准确率,优于现有的家族分类方法。此外,我们使用ProtVec表示法来预测结构蛋白中的无序蛋白。使用了两个无序序列数据库:DisProt数据库以及以富含苯丙氨酸-甘氨酸重复序列(FG-NUP)的核孔蛋白无序区域为特征的数据库。使用支持向量机分类器,FG-Nup序列与蛋白质数据库中的结构化蛋白质序列的区分准确率为99.8%,非结构化DisProt序列与结构化DisProt序列的区分准确率为100.0。这些结果表明,只需向该模型提供各种蛋白质的序列数据,就可以确定关于蛋白质结构的准确信息。重要的是,这个模型只需要训练一次,然后就可以应用于提取关于感兴趣蛋白质的全面信息集。此外,这种表示可以被认为是深度学习在生物信息学中的各种应用的预训练。相关数据可在生活语言处理网站:http://llp.berkeley.edu和哈佛数据中心:http://dx.doi.org/10.7910/DVN/JMFHTN.获得
We introduce a new representation and feature extraction method for biological sequences. Named bio-vectors (BioVec) to refer to biological sequences in general with protein-vectors (ProtVec) for proteins (amino-acid sequences) and gene-vectors (GeneVec) for gene sequences, this representation can be widely used in applications of deep learning in proteomics and genomics. In the present paper, we focus on protein-vectors that can be utilized in a wide array of bioinformatics investigations such as family classification, protein visualization, structure prediction, disordered protein identification, and protein-protein interaction prediction. In this method, we adopt artificial neural network approaches and represent a protein sequence with a single dense n-dimensional vector. To evaluate this method, we apply it in classification of 324,018 protein sequences obtained from Swiss-Prot belonging to 7,027 protein families, where an average family classification accuracy of 93%±0.06% is obtained, outperforming existing family classification methods. In addition, we use ProtVec representation to predict disordered proteins from structured proteins. Two databases of disordered sequences are used: the DisProt database as well as a database featuring the disordered regions of nucleoporins rich with phenylalanine-glycine repeats (FG-Nups). Using support vector machine classifiers, FG-Nup sequences are distinguished from structured protein sequences found in Protein Data Bank (PDB) with a 99.8% accuracy, and unstructured DisProt sequences are differentiated from structured DisProt sequences with 100.0% accuracy. These results indicate that by only providing sequence data for various proteins into this model, accurate information about protein structure can be determined. Importantly, this model needs to be trained only once and can then be applied to extract a comprehensive set of information regarding proteins of interest. Moreover, this representation can be considered as pre-training for various applications of deep learning in bioinformatics. The related data is available at Life Language Processing Website: http://llp.berkeley.edu and Harvard Dataverse: http://dx.doi.org/10.7910/DVN/JMFHTN.