Machine Learning at the Interface of Polymer Science and Biology: How Far Can We Go?

Machine Learning at the Interface of Polymer Science and Biology: How Far Can We Go?
复制标题

DOI:
10.1021/acs.biomac.1c01436
复制
发表时间:
2022-02-08
期刊:
影响因子:
6.2
通讯作者:
Percec, Simona
Percec, Simona
中科院分区:
化学2区
文献类型:
--
作者:
Gianti, Eleonora;Percec, Simona

文献摘要

被引文献

相似文献

该观点概述了使用机器学习(ML)(一种数据驱动的方法)来解决生物大分子设计、合成、加工和表征中的关键问题的最新进展和未来方向。这些任务的实现需要在广阔而复杂的化学和生物空间中航行,难以以合理的速度完成。使用现代算法和超级计算机,量子物理学方法能够检查包含数百个相互作用物种的系统,并确定在相空间的特定区域中找到它们的概率,从而预测它们的性质。同样,在高性能计算的支持下,化学和生物分子模拟的现代方法最终产生了规模不断扩大和内在高度复杂的数据集。因此,使用ML从这些领域提取相关信息对于促进我们对化学和生物分子系统的理解至关重要。ML方法的核心是统计算法,通过评估给定数据集的一部分,识别、学习和操纵管理整个数据集的底层规则。组装质量模型来表示数据,然后预测和消除误差源是ML中的关键步骤。除了不断增长的机器学习工具基础设施来解决复杂问题外,越来越多与我们对生物大分子基本性质的理解相关的方面都暴露在机器学习中。这些领域,包括那些位于聚合物科学和生物学界面的领域(即,结构确定,从头设计,折叠和动力学),努力采用并利用ML领域方法提供的变革力量,这显然具有加速生物大分子领域研究的潜力。
This Perspective outlines recent progress and future directions for using machine learning (ML), a data-driven method, to address critical questions in the design, synthesis, processing, and characterization of biomacromolecules. The achievement of these tasks requires the navigation of vast and complex chemical and biological spaces, difficult to accomplish with reasonable speed. Using modern algorithms and supercomputers, quantum physics methods are able to examine systems containing a few hundred interacting species and determine the probability of finding them in a particular region of phase space, thereby anticipating their properties. Likewise, modern approaches in chemistry and biomolecular simulation, supported by high performance computing, have culminated in producing data sets of escalating size and intrinsically high complexity. Hence, using ML to extract relevant information from these fields is of paramount importance to advance our understanding of chemical and biomolecular systems. At the heart of ML approaches lie statistical algorithms, which by evaluating a portion of a given data set, identify, learn, and manipulate the underlying rules that govern the whole data set. The assembly of a quality model to represent the data followed by the predictions and elimination of error sources are the key steps in ML. In addition to a growing infrastructure of ML tools to address complex problems, an increasing number of aspects related to our understanding of the fundamental properties of biomacromolecules are exposed to ML. These fields, including those residing at the interface of polymer science and biology (i.e., structure determination, de novo design, folding, and dynamics), strive to adopt and take advantage of the transformative power offered by approaches in the ML domain, which clearly has the potential of accelerating research in the field of biomacromolecules.