A document classifier for medicinal chemistry publications trained on the ChEMBL corpus.

A document classifier for medicinal chemistry publications trained on the ChEMBL corpus.
复制标题

DOI:
10.1186/s13321-014-0040-8
复制
发表时间:
2014-12
影响因子:
8.6
通讯作者:
Overington JP
Overington JP
中科院分区:
化学2区
文献类型:
--
作者:
Papadatos G;van Westen GJ;Croset S;Santos R;Trubian S;Overington JP

文献摘要

参考文献

被引文献

相似文献

科学出版物数量的大幅增加推动了对半自动和全自动文本挖掘方法的需求,以便协助对个别科学家进行分类,并将更大规模的数据提取和整理到公共数据库中。在这里,我们介绍了一个文档分类器,它能够成功地区分类似ChEMBL的出版物(即与小分子药物发现有关并可能包含定量生物活性数据的出版物)和不是此类出版物的出版物。药物化学文献集的规模史无前例,再加上人工管理和映射到化学和生物学的优势,使ChEMBL语料库成为文本挖掘的独特资源。该方法已分别作为管道试点(8.5版)和Knime(2.9版)的数据协议/工作流实现。工作流程和模型均可在以下网址免费获得:ftp://ftp.ebi.ac.uk/pub/databases/chembl/text-mining.可以容易地对这些进行修改,以包括附加的关键字约束,以进一步聚焦搜索。大规模机器学习文档分类对于这一特定的应用程序被证明是非常健壮和灵活的,如四个不同的基于文本挖掘的用例所示。这些模型在两个数据工作流程平台上都很容易获得,我们相信这将允许大多数科学界将它们应用于他们自己的数据。本文的在线版本(doi:10.1186/s13321-0140040-8)包含补充材料,授权用户可以使用。
The large increase in the number of scientific publications has fuelled a need for semi- and fully automated text mining approaches in order to assist in the triage process, both for individual scientists and also for larger-scale data extraction and curation into public databases. Here, we introduce a document classifier, which is able to successfully distinguish between publications that are `ChEMBL-like’ (i.e. related to small molecule drug discovery and likely to contain quantitative bioactivity data) and those that are not. The unprecedented size of the medicinal chemistry literature collection, coupled with the advantage of manual curation and mapping to chemistry and biology make the ChEMBL corpus a unique resource for text mining. The method has been implemented as a data protocol/workflow for both Pipeline Pilot (version 8.5) and KNIME (version 2.9) respectively. Both workflows and models are freely available at: ftp://ftp.ebi.ac.uk/pub/databases/chembl/text-mining. These can be readily modified to include additional keyword constraints to further focus searches. Large-scale machine learning document classification was shown to be very robust and flexible for this particular application, as illustrated in four distinct text-mining-based use cases. The models are readily available on two data workflow platforms, which we believe will allow the majority of the scientific community to apply them to their own data. The online version of this article (doi:10.1186/s13321-014-0040-8) contains supplementary material, which is available to authorized users.
DOI: 10.1093/bioinformatics/btm557
发表时间: 2008-01-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Rebholz-Schuhmann, Dietrich;Arregui, Miguel;Jimeno, Antonio
通讯作者: Jimeno, Antonio
DOI: 10.1371/journal.pone.0058201
发表时间: 2013
期刊: PloS one
影响因子: 3.7
作者:
Davis AP;Wiegers TC;Johnson RJ;Lay JM;Lennon-Hopkins K;Saraceni-Richards C;Sciaky D;Murphy CG;Mattingly CJ
通讯作者: Mattingly CJ
DOI: 10.1093/bioinformatics/bts183
发表时间: 2012-06-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Rocktaschel, Tim;Weidlich, Michael;Leser, Ulf
通讯作者: Leser, Ulf
DOI: 10.1371/journal.pbio.0030065
发表时间: 2005-02
期刊: PLoS biology
影响因子: 9.8
作者:
Rebholz-Schuhmann D;Kirsch H;Couto F
通讯作者: Couto F
DOI: 10.3163/1536-5050.100.2.007
发表时间: 2012-04-01
影响因子: 2
作者:
Brown, Heather L.
通讯作者: Brown, Heather L.