Estimating Error Rates in Bioactivity Databases

Estimating Error Rates in Bioactivity Databases
复制标题

DOI:
10.1021/ci400099q
复制
发表时间:
2013-10-01
影响因子:
5.6
通讯作者:
Franke, Lutz
Franke, Lutz
中科院分区:
化学2区
文献类型:
--
作者:
Tiikkainen, Pekka;Bellis, Louisa;Franke, Lutz

文献摘要

被引文献

相似文献

生物活性数据库通常用于药物发现,以查找并使用预测工具来预测小分子的潜在靶标。这些数据库通常是从专利和科学文章中手动精选的。除了源文档中的错误外,人为因素还可能在提取过程中导致错误。这些错误可能会导致早期药物发现过程中的错误决定。在目前的工作中,我们比较了来自三个大型数据库(ChEMBL、Liceptor和袋熊)的生物活性数据,这些数据库已经从相同的源文件中整理了数据。因此,我们能够报告单个活动参数和单个生物活性数据库的错误率估计。小分子结构的估计错误率最大,其次是靶标、活性值和活性类型。这一订单也反映在供应商特定错误率估计中。这些结果也有助于识别要重新计算的数据点。我们希望这些结果将使科学家们更广泛地认识到生物活性数据中错误的频率和类型。
Bioactivity databases are routinely used in drug discovery to look-up and, using prediction tools, to predict potential targets for small molecules. These databases are typically manually curated from patents and scientific articles. Apart from errors in the source document, the human factor can cause errors during the extraction process. These errors can lead to wrong decisions in the early drug discovery 0 process. In the current work, we have compared bioactivity data from three large databases (ChEMBL, Liceptor, and WOMBAT) who have curated data from the same source documents. As a result, we are able to report error rate estimates for individual activity parameters and individual bioactivity databases. Small molecule structures have the greatest estimated error rate followed by target, activity value, and activity type. This order is also reflected in supplier-specific error rate estimates. The results are also useful in identifying data points for recuration. We hope the results will lead to a more widespread awareness among scientists on the frequencies and types of errors in bioactivity data.