Porting a lexicalized-grammar parser to the biomedical domain

Porting a lexicalized-grammar parser to the biomedical domain
复制标题

DOI:
10.1016/j.jbi.2008.12.004
复制
发表时间:
2009-10-01
影响因子:
4.5
通讯作者:
Clark, Stephen
Clark, Stephen
中科院分区:
医学3区
文献类型:
--
作者:
Rimell, Laura;Clark, Stephen

文献摘要

被引文献

相似文献

本文将一种最先进的、基于语言的统计解析器引入生物医学文本挖掘领域,并提出了一种方法,使其适用于只需要有限资源进行数据标注的生物医学领域。解析器最初是使用Penn Treebank开发的,因此调整为报纸文本。我们的方法利用词汇化的语法形式主义,组合范畴语法(CCG),在比完全句法派生更低的表示级别上训练解析器。CCG解析器使用三级表示:第一级由词性(POS)标签组成;第二级由更细粒度的CCG词法类别组成;以及第三级,由CCG派生组成。我们发现,简单地对生物医学数据上的词性标记器进行重新训练可以大大提高句法分析的性能,并且在表示的中间词汇类别级别使用标注数据可以进一步提高句法分析的准确率。我们描述了评估解析器所涉及的过程,并获得了与报纸文本报告的相同范围内的生物医学数据的准确性,而高于我们评估的生物医学资源的先前报告的准确性。我们的结论是,将报纸解析器移植到生物医学领域,至少对于使用词汇化语法的解析器来说,可能并不像最初想象的那么困难。(C)2008 Elsevier Inc.保留所有权利。
This paper introduces a state-of-the-art, linguistically motivated statistical parser to the biomedical text mining community, and proposes a method of adapting it to the biomedical domain requiring only limited resources for data annotation. The parser was originally developed using the Penn Treebank and is therefore tuned to newspaper text. Our approach takes advantage of a lexicalized grammar formalism, Combinatory Categorial Grammar (CCG), to train the parser at a lower level of representation than full syntactic derivations. The CCG parser uses three levels of representation: a first level consisting of part-ofspeech (POS) tags; a second level consisting of more fine-grained CCG lexical categories: and a third, hierarchical level consisting of CCG; derivations. We find that simply retraining the POS tagger on biomedical data leads to a large improvement in parsing performance, and that using annotated data at the intermediate lexical category level of representation improves parsing accuracy further. We describe the procedure involved in evaluating the parser, and obtain accuracies for biomedical data in the same range as those reported for newspaper text, and higher than those previously reported for the biomedical resource on which we evaluate. Our conclusion is that porting newspaper parsers to the biomedical domain, at least for parsers which use lexicalized grammars, may not be as difficult as first thought. (C) 2008 Elsevier Inc. All rights reserved.