Using workflows to explore and optimise named entity recognition for chemistry.

Using workflows to explore and optimise named entity recognition for chemistry.
复制标题

DOI:
10.1371/journal.pone.0020181
复制
发表时间:
2011
期刊:
影响因子:
3.7
通讯作者:
Ananiadou S
Ananiadou S
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Kolluru B;Hawizy L;Murray-Rust P;Tsujii J;Ananiadou S

文献摘要

参考文献

被引文献

相似文献

化学文本挖掘工具应该具有互操作性和适应性,而不考虑系统级的实现、安装甚至编程问题。我们的目标是通过可重新配置的工作流从底层实现中抽象出这些工具的功能,以自动识别化学名称。为了实现这一点,我们重构了一个已经建立的命名实体识别器OSCAR(在化学领域),并研究了每个组件对网络性能的影响。我们使用可互操作的文本挖掘框架U-Compare从OSCAR开发了两个可重构的工作流。这些工作流程可以使用U-Compare图形用户界面的拖放机制进行更改。这些工作流还提供了一个平台来研究文本挖掘组件之间的关系,例如标记化和命名实体识别(使用最大熵马尔可夫模型(MEMM)和基于模式识别的分类器)。结果表明,特别是对于化学来说,消除由标记化技术产生的噪声在命名实体识别(NER)准确性方面的性能略好于其他技术。较差的标记化转化为对分类器组件的较差输入,进而导致类型I或类型II错误的增加,从而降低整体性能。在Sciborg语料库上,基于工作流的系统使用了新的标记器,同时保留了相同的MEMM组件,将F-Score从82.35%提高到84.44%。在PubMed语料库上,它的F值为84.84%,而OSCAR的F值为84.23%。
Chemistry text mining tools should be interoperable and adaptable regardless of system-level implementation, installation or even programming issues. We aim to abstract the functionality of these tools from the underlying implementation via reconfigurable workflows for automatically identifying chemical names. To achieve this, we refactored an established named entity recogniser (in the chemistry domain), OSCAR and studied the impact of each component on the net performance. We developed two reconfigurable workflows from OSCAR using an interoperable text mining framework, U-Compare. These workflows can be altered using the drag-&-drop mechanism of the graphical user interface of U-Compare. These workflows also provide a platform to study the relationship between text mining components such as tokenisation and named entity recognition (using maximum entropy Markov model (MEMM) and pattern recognition based classifiers). Results indicate that, for chemistry in particular, eliminating noise generated by tokenisation techniques lead to a slightly better performance than others, in terms of named entity recognition (NER) accuracy. Poor tokenisation translates into poorer input to the classifier components which in turn leads to an increase in Type I or Type II errors, thus, lowering the overall performance. On the Sciborg corpus, the workflow based system, which uses a new tokeniser whilst retaining the same MEMM component, increases the F-score from 82.35% to 84.44%. On the PubMed corpus, it recorded an F-score of 84.84% as against 84.23% by OSCAR.
DOI: 10.1002/cpe.994
发表时间: 2006-08-25
影响因子: 2
作者:
Ludascher, Bertram;Altintas, Ilkay;Zhao, Yang
通讯作者: Zhao, Yang
DOI: 10.1093/bioinformatics/btp535
发表时间: 2009-11-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Hettne, Kristina M.;Stierum, Rob H.;Kors, Jan A.
通讯作者: Kors, Jan A.
DOI: 10.1021/ci990052b
发表时间: 1999-11-01
期刊: JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES
影响因子: --
作者:
Murray-Rust, P;Rzepa, HS
通讯作者: Rzepa, HS
DOI: 10.1093/bioinformatics/btp289
发表时间: 2009-08-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Kano Y;Baumgartner WA Jr;McCrohon L;Ananiadou S;Cohen KB;Hunter L;Tsujii J
通讯作者: Tsujii J
DOI: 10.1021/ci800332w
发表时间: 2009-02-01
影响因子: 5.6
作者:
Jiao, Dazhi;Wild, David J.
通讯作者: Wild, David J.