From Unsupervised Multi-Instance Learning to Identification of Near-Native Protein Structures
From Unsupervised Multi-Instance Learning to Identification of Near-Native Protein Structures
复制标题
DOI:
10.29007/pjcf
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
F. Alam;Amarda Shehu
中科院分区:
文献类型:
--
作者:
F. Alam;Amarda Shehu
A major challenge in computational biology regards recognizing one or more biologically-active/native tertiary protein structures among thousands of physically-realistic structures generated via template-free protein structure prediction algorithms. Clustering structures based on structural similarity remains a popular approach. However, clustering organizes structures into groups and does not directly provide a mechanism to select individual structures for prediction. In this paper, we provide a few algorithms for this selection problem. We approach the problem under unsupervised multi-instance learning and address it in three stages, first organizing structures into bags, identifying relevant bags, and then drawing individual structures/instances from these bags. We present both non-parametric and parametric algorithms for drawing individual instances. In the latter, parameters are trained over training data and evaluated over testing data via rigorous metrics.