The impact of AI-based modeling on the accuracy of protein assembly prediction: Insights from CASP15

The impact of AI-based modeling on the accuracy of protein assembly prediction: Insights from CASP15
复制标题

DOI:
10.1002/prot.26598
复制
发表时间:
2023-10-20
影响因子:
2.9
通讯作者:
Karaca,Ezgi
Karaca,Ezgi
中科院分区:
生物学4区
文献类型:
--
作者:
Ozden,Burcu;Kryshtafovych,Andriy;Karaca,Ezgi

文献摘要

被引文献

相似文献

在CASP15中,87个预测者提交了41个装配目标的约11000个模型。该社区在整体折叠和界面接触预测方面表现出色,实现了令人印象深刻的90%的成功率(与CASP14的31%相比)。这一非凡的成就很大程度上归功于DeepMind将AF2 - multitimer方法整合到定制的预测管道中。为了评估参与方法的附加价值,我们将社区模型与基线AF2‐multitimer预测器进行了比较。在超过1/3的病例中,社区模型优于基线预测器。这种性能提高的主要原因是使用了定制的多序列比对,优化的AF2 - Multimer采样,以及手工组装AF2 - Multimer构建的亚复合物。最好的三组,依次是郑,冯卡瓦斯和沃尔纳。Zheng和Venclovas的成功率为73.2%(41例),而Wallner的成功率为69.4%(36例)。尽管如此,在预测具有弱进化信号的结构方面仍然存在挑战,例如纳米体-抗原、抗体-抗原和病毒复合物。可以预见的是,由于大型复合体的高内存计算需求,其建模仍然具有挑战性。除了装配类外,我们还评估了三级结构预测目标中域间界面建模的准确性。分析了具有17个独特界面的7个目标的模型。最好的预测方法达到了76.5%的成功率,其中UM - TBM组处于领先地位。在域间类别中,我们观察到当给定的域对的进化信号较弱或结构较大时,预测器面临挑战,就像在装配类别的情况下一样。总的来说,CASP15在接口建模方面有了前所未有的改进,反映了CASP14中看到的AI革命。
In CASP15, 87 predictors submitted around 11 000 models on 41 assembly targets. The community demonstrated exceptional performance in overall fold and interface contact predictions, achieving an impressive success rate of 90% (compared to 31% in CASP14). This remarkable accomplishment is largely due to the incorporation of DeepMind's AF2‐Multimer approach into custom‐built prediction pipelines. To evaluate the added value of participating methods, we compared the community models to the baseline AF2‐Multimer predictor. In over 1/3 of cases, the community models were superior to the baseline predictor. The main reasons for this improved performance were the use of custom‐built multiple sequence alignments, optimized AF2‐Multimer sampling, and the manual assembly of AF2‐Multimer‐built subcomplexes. The best three groups, in order, are Zheng, Venclovas, and Wallner. Zheng and Venclovas reached a 73.2% success rate over all (41) cases, while Wallner attained 69.4% success rate over 36 cases. Nonetheless, challenges remain in predicting structures with weak evolutionary signals, such as nanobody–antigen, antibody–antigen, and viral complexes. Expectedly, modeling large complexes also remains challenging due to their high memory compute demands. In addition to the assembly category, we assessed the accuracy of modeling interdomain interfaces in the tertiary structure prediction targets. Models on seven targets featuring 17 unique interfaces were analyzed. Best predictors achieved a 76.5% success rate, with the UM‐TBM group being the leader. In the interdomain category, we observed that the predictors faced challenges, as in the case of the assembly category, when the evolutionary signal for a given domain pair was weak or the structure was large. Overall, CASP15 witnessed unprecedented improvement in interface modeling, reflecting the AI revolution seen in CASP14.