Quantifying Pairwise Similarity for Complex Polymers

Quantifying Pairwise Similarity for Complex Polymers
复制标题

DOI:
10.1021/acs.macromol.3c00761
复制
发表时间:
2023-09-06
期刊:
影响因子:
5.5
通讯作者:
Olsen,Bradley D.
Olsen,Bradley D.
中科院分区:
化学1区
文献类型:
--
作者:
Shi,Jiale;Rebello,Nathan J.;Olsen,Bradley D.

文献摘要

相似文献

定义化学实体之间的相似性是聚合物信息学中的一项基本任务,可以进行排序、聚类和分类。尽管聚合物的成对化学相似性很重要,但它仍然是一个悬而未决的问题。在这里,基于典型的BigSMILEs生成的聚合物的随机图形表示,设计了具有明确主链的聚合物的相似性函数。随机图表示分为三个部分:重复单元、端基和聚合物拓扑。利用推土机的距离计算重复单元和端组的相似度,利用图形编辑距离计算拓扑的相似度。这三个值可以线性或非线性组合,以产生聚合物的总体成对化学相似性得分,该得分与专家用户的化学直觉基本一致,并且可以根据不同化学特征对给定相似性问题的相对重要性进行调整。这种方法为定量计算聚合物的成对化学相似分数提供了可靠的解决方案,并代表着朝着建立聚合物数据的搜索引擎和定量设计工具迈出的重要一步。
Defining the similarity between chemical entities is an essential task in polymer informatics, enabling ranking, clustering, and classification. Despite its importance, the pairwise chemical similarity of polymers remains an open problem. Here, a similarity function for polymers with well-defined backbones is designed based on polymers’ stochastic graph representations generated from canonical BigSMILES, a structurally based line notation for describing macromolecules. The stochastic graph representations are separated into three parts: repeat units, end groups, and polymer topology. The earth mover’s distance is utilized to calculate the similarity of the repeat units and end groups, while the graph edit distance is used to calculate the similarity of the topology. These three values can be linearly or nonlinearly combined to yield an overall pairwise chemical similarity score for polymers that is largely consistent with the chemical intuition of expert users and is adjustable based on the relative importance of different chemical features for a given similarity problem. This method gives a reliable solution to quantitatively calculate the pairwise chemical similarity score for polymers and represents a vital step toward building search engines and quantitative design tools for polymer data.