Adaptive Stochastic Optimization to Improve Protein Conformation Sampling

Adaptive Stochastic Optimization to Improve Protein Conformation Sampling
复制标题

DOI:
10.1109/tcbb.2021.3134103
复制
发表时间:
2021-12
期刊:
IEEE/ACM Transactions on Computational Biology and Bioinformatics
影响因子:
--
通讯作者:
Ahmed Bin Zaman;Toki Tahmid Inan;K. De Jong;Amarda Shehu
Ahmed Bin Zaman;Toki Tahmid Inan;K. De Jong;Amarda Shehu
中科院分区:
其他
文献类型:
--
作者:
Ahmed Bin Zaman;Toki Tahmid Inan;K. De Jong;Amarda Shehu

文献摘要

相似文献

我们很早就知道,确定蛋白质的结构是理解蛋白质功能的关键。计算方法在很大程度上解决了这个问题的狭隘公式,试图从氨基酸序列中计算出一个天然结构。现在,AlphaFold2被证明能够揭示许多蛋白质的高质量天然结构。然而,多年来,研究人员一直主张拓宽我们的视野,以解释自然结构的多样性。我们现在知道,许多蛋白质分子在不同的结构之间切换,以调节与细胞内分子伙伴的相互作用。从新开始阐明这种结构是异常困难的,因为它需要探索可能非常大的结构空间,以寻找竞争的、接近最佳的结构。在这里,我们报告了一种新的随机优化方法,该方法能够从已知的氨基酸序列中揭示给定蛋白质非常不同的结构。该方法利用进化搜索技术并调整其对搜索空间的探索,以在存在计算预算的情况下在探索和开发之间取得平衡。除了展示这种方法在识别多个本地结构方面的实用性外,我们还提供了一个基准数据集,供研究人员继续研究这个问题。
We have long known that characterizing protein structures structure is key to understanding protein function. Computational approaches have largely addressed a narrow formulation of the problem, seeking to compute one native structure from an amino-acid sequence. Now AlphaFold2 is shown to be able to reveal a high-quality native structure for many proteins. However, researchers over the years have argued for broadening our view to account for the multiplicity of native structures. We now know that many protein molecules switch between different structures to regulate interactions with molecular partners in the cell. Elucidating such structures de novo is exceptionally difficult, as it requires exploration of possibly a very large structure space in search of competing, near-optimal structures. Here we report on a novel stochastic optimization method capable of revealing very different structures for a given protein from knowledge of its amino-acid sequence. The method leverages evolutionary search techniques and adapts its exploration of the search space to balance between exploration and exploitation in the presence of a computational budget. In addition to demonstrating the utility of this method for identifying multiple native structures, we additionally provide a benchmark dataset for researchers to continue work on this problem.