Reply to Skinnider and Magarvey: Rates of novel natural product discovery remain high.

Reply to Skinnider and Magarvey: Rates of novel natural product discovery remain high.
复制标题

回复 Skinnider 和 Magarvey:新型天然产品的发现率仍然很高。

DOI:
10.1073/pnas.1711139114
复制
发表时间:
2017
影响因子:
11.1
通讯作者:
Linington,RogerG
Linington,RogerG
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Pye,CameronR;Bertin,MatthewJ;Lokey,RScott;Gerwick,WilliamH;Linington,RogerG

文献摘要

相似文献

令人鼓舞的是,我们最近的文章研究了天然产物(NP)的发现率和结构多样性的趋势(1)正在这个迷人的领域(2)引起讨论。然而,我们希望纠正Skinnider和Magarvey(3)评论中提出的几个误解。Skinnider和Magarvey(3)的信错误地总结了我们工作的关键结论。这封信指出:“他们的分析表明,结构独特的NP发现的速度正在下降。我们的研究得出了完全相反的结论:“粗略地回顾这些数据可能表明,天然产物领域不再发现新的化学实体。[然而,]同样重要的是评估具有低相似性分数的分子的分布。总的来说,这项分析表明,天然产物中新分子结构的发现率自这一领域的起源以来一直在增加,尽管已发表的天然产物数量不断增加,但仍保持着显著的增长率。(一). Skinnider和Magarvey(3)对我们(1)的分析提出了两个主要的关注点。首先,他们认为观察到的相似性趋势仅仅是由于样本量的增加,而不是与NP多样性相关的问题。事实上,正如伯克及其同事(2)所指出的那样,正是这种数值的上升使我们能够得出结论,即迄今为止发现的NP化学空间是有界的;如果可用的NP化学空间远远大于已经观察到的空间,那么这条曲线(见参考文献1中的图1B)将包含几乎一致的低值。Skinnider和Magarvey的分析(3),用ZINC数据库中的化合物替换NP的子集(4),产生了低得多的相似性值,表明:(i)NP不像ZINC化合物,(ii)如果删除一些NP,那么描述NP空间边界的能力就会下降。这一发现进一步得到了我们对源亚类的分析的支持(参考文献1中的图2D),这表明来自一个源亚类的化合物与所有其他海洋化合物具有较低的结构相似性,无论有多少化合物被添加到数据集中。这些图反驳了文库大小单独对观察到的趋势负责的建议。其次,Skinnider和Magarvey(3)质疑我们的结论,即近年来发表的更多化合物是已知支架的衍生物,而不是过去几十年的真实情况。为了支持他们的立场,Skinnider和Magarvey(3)将实际相似性趋势与随机数据集进行了比较。由于实际原因,这种方法从根本上是有缺陷的,因为它忽略了一个事实,即许多NP手稿在一篇文章中报告了多个家庭成员。以2015年为例,我们的数据集包含来自484篇论文的1,576种化合物。其中,49%的化合物在同一文章中具有至少一种其他化合物,Tanimoto评分> 0.9。我们的原始分析排除了文章内比较,因为它只比较了化合物与前几年发现的化合物。通过将化合物随机化,Skinnider和Magarvey(3)在家族成员之间引入了大量的衍生关系,排除了实际数据集和随机数据集之间的直接比较。因此,我们必须得出这样的结论:Skinnider和Magarvey(3)的信中提出的分析并不能准确地反映NP研究的现状。相反,为了呼应我们最初的结论,“天然产品的未来确实非常光明”(1)。
It is encouraging that our recent article examining trends in discovery rates and structural diversity for natural products (NP)(1) is generating discussion in this fascinating area (2). However, we wish to correct several misconceptions presented in the comments from Skinnider and Magarvey (3). Skinnider and Magarvey’s (3) letter incorrectly summarizes the key conclusion of our work. The letter states that “[t] heir analysis suggests that the pace of structurally unique NP discovery is decreasing.” Our study makes precisely the opposite conclusion:“A cursory review of these data might suggest that the field of natural products is no longer discovering novel chemical entities...[However,] it is also important to evaluate the distribution of molecules with low similarity scores... Overall, this analysis indicates that the discovery rate of new molecular architectures among natural products has increased since the origins of this field and has remained at a significant rate despite the everincreasing number of published natural products...”(1). Skinnider and Magarvey (3) raise two main concerns about our (1) analyses. First, they suggest that the observed trends in similarity are due solely to increasing sample size, and do not inform questions related to NP diversity. In fact, as noted by Burke and coworkers (2), it is precisely this rise in values that allows us to conclude that the NP chemical space discovered to date is bounded; if available NP chemical space were vastly greater than what has been observed, this curve (see figure 1B in ref. 1) would contain almost uniformly low values. The analysis by Skinnider and Magarvey (3), replacing a subset of NPs with compounds from the ZINC database (4), yields much lower similarity values, demonstrating that:(i) NPs are not like ZINC compounds and (ii) if one removes some of the NPs, then the ability to describe the boundary of NP space goes down. This finding is further supported by our analysis of source subclasses (figure 2D in ref. 1), which demonstrates that compounds from one source subclass bear low structural similarities to all other marine compounds, regardless of how many compounds are added to the dataset. These plots counter the suggestion that library size alone is responsible for the observed trends. Second, Skinnider and Magarvey (3) question our (1) conclusion that more compounds published in recent years are derivatives of known scaffolds than was true in previous decades. To support their position, Skinnider and Magarvey (3) compare actual similarity trends against a randomized dataset. For practical reasons this approach is fundamentally flawed because it ignores the fact that many NP manuscripts report multiple family members in a single article. Taking 2015 as an example, our dataset contains 1,576 compounds derived from 484 papers. Of these, 49% of compounds possess at least one other compound in the same article with a Tanimoto score> 0.9. Our original analysis excludes intra-article comparisons because it only compares compounds to those found in previous years. By randomizing compounds across bins, Skinnider and Magarvey (3) have introduced a large number of derivative relationships between family members, precluding direct comparison between the actual and randomized datasets. Therefore, we must conclude that the analyses presented in Skinnider and Magarvey’s (3) letter do not accurately reflect the current situation for NP research. Instead, to echo our original conclusion,“the future for natural products is very bright indeed”(1).