From Unsupervised Multi-Instance Learning to Identification of Near-Native Protein Structures

From Unsupervised Multi-Instance Learning to Identification of Near-Native Protein Structures
复制标题

DOI:
10.29007/pjcf
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
F. Alam;Amarda Shehu
F. Alam;Amarda Shehu
中科院分区:
其他
文献类型:
--
作者:
F. Alam;Amarda Shehu

文献摘要

相似文献

计算生物学中的一个主要挑战是在通过无模板蛋白质结构预测算法生成的数千个物理现实结构中识别一个或多个生物活性/天然三级蛋白质结构。基于结构相似性的结构聚类仍然是一种流行的方法。然而,聚类将结构组织成组,并且不直接提供选择用于预测的单个结构的机制。在本文中,我们提供了一些算法,这个选择问题。我们在无监督的多实例学习下处理这个问题,并分三个阶段解决它,首先将结构组织成包,识别相关的包,然后从这些包中提取单个结构/实例。我们提出了非参数化和参数化算法绘制个别实例。在后者中,参数通过训练数据进行训练,并通过严格的指标在测试数据上进行评估。
A major challenge in computational biology regards recognizing one or more biologically-active/native tertiary protein structures among thousands of physically-realistic structures generated via template-free protein structure prediction algorithms. Clustering structures based on structural similarity remains a popular approach. However, clustering organizes structures into groups and does not directly provide a mechanism to select individual structures for prediction. In this paper, we provide a few algorithms for this selection problem. We approach the problem under unsupervised multi-instance learning and address it in three stages, first organizing structures into bags, identifying relevant bags, and then drawing individual structures/instances from these bags. We present both non-parametric and parametric algorithms for drawing individual instances. In the latter, parameters are trained over training data and evaluated over testing data via rigorous metrics.