Comparison of character-level and part of speech features for name recognition in biomedical texts

Comparison of character-level and part of speech features for name recognition in biomedical texts
复制标题

DOI:
10.1016/j.jbi.2004.08.008
复制
发表时间:
2004-12-01
影响因子:
4.5
通讯作者:
Takeuchi, K
Takeuchi, K
中科院分区:
医学3区
文献类型:
--
作者:
Collier, N;Takeuchi, K

文献摘要

被引文献

相似文献

现在从分子生物学实验中可以获得的大量数据导致报告结果的爆炸式增长,其中大多数只能以非结构化文本格式获得。因此,人们对文本挖掘任务非常感兴趣,以帮助事实提取、文档筛选、引文分析以及与大型基因和基因产物数据库的链接。特别是,对命名实体 (NE) 任务作为所有这些任务的核心技术进行了深入研究,这是由 GENIA v3.02 语料库等大容量训练集的可用性驱动的。尽管训练集如此之大,生物学 NE 的准确性已被证明始终远低于新闻领域的高水平表现,新闻领域的 F 分数通常高于 90,可以被认为接近人类表现。我们认为,至关重要的是,对影响模型性能的因素进行更严格的分析,以发现潜在的局限性以及我们未来的研究方向应该是什么。我们在本文中的研究报告了两种广泛使用的特征类型(词性 (POS) 标签和字符级正字法特征)的变化,并比较了这些变化如何影响性能。我们的实验基于经过验证的最先进模型、支持向量机,使用 100 个带注释的 MEDLINE 摘要的高质量子集。实验表明,性能最好的特征是正交特征,F 得分为 72.6。尽管在 GENIA v3.02p POS 语料库上进行域内训练的 Brill 标注器在所有 POS 标注器中提供了最佳的整体性能,F 得分为 68.6,但这仍然明显低于拼写特征。结合起来,这两种功能类型似乎会相互干扰,并导致性能略有下降,F 分数为 72.3。 (C) 2004 Elsevier Inc. 保留所有权利。
The immense volume of data which is now available from experiments in molecular biology has led to an explosion in reported results most of which are available only in unstructured text format. For this reason there has been great interest in the task of text mining to aid in fact extraction, document screening, citation analysis, and linkage with large gene and gene-product databases. In particular there has been an intensive investigation into the named entity (NE) task as a core technology in all of these tasks which has been driven by the availability of high volume training sets such as the GENIA v3.02 corpus. Despite such large training sets accuracy for biology NE has proven to be consistently far below the high levels of performance in the news domain where F scores above 90 are commonly reported which can be considered near to human performance. We argue that it is crucial that more rigorous analysis of the factors that contribute to the model's performance be applied to discover where the underlying limitations are and what our future research direction should be. Our investigation in this paper reports on variations of two widely used feature types, part of speech (POS) tags and character-level orthographic features, and makes a comparison of how these variations influence performance. We base our experiments on a proven state-of-the-art model, support vector machines using a high quality subset of 100 annotated MEDLINE abstracts. Experiments reveal that the best performing features are orthographic features with F score of 72.6. Although the Brill tagger trained in-domain on the GENIA v3.02p POS corpus gives the best overall performance of any POS tagger, at an F score of 68.6, this is still significantly below the orthographic features. In combination these two features types appear to interfere with each other and degrade performance slightly to an F score of 72.3. (C) 2004 Elsevier Inc. All rights reserved.