Subdomain adaptation of a POS tagger with a small corpus

Subdomain adaptation of a POS tagger with a small corpus
复制标题

DOI:
10.3115/1567619.1567651
复制
发表时间:
2006-06
期刊:
--
影响因子:
--
通讯作者:
Yuka Tateisi;Yoshimasa Tsuruoka;Junichi Tsujii
Yuka Tateisi;Yoshimasa Tsuruoka;Junichi Tsujii
中科院分区:
其他
文献类型:
--
作者:
Yuka Tateisi;Yoshimasa Tsuruoka;Junichi Tsujii

文献摘要

被引文献

相似文献

对于生物医学研究摘要领域,有两个大型语料库可用,即 GENIA (Kim et al 2003) 和 Penn BioIE (Kulik et al 2004)。两者基本上都属于人类领域,并且在这些语料库上训练的系统在应用于处理其他物种的摘要时的性能尚不清楚。在基于机器学习的系统中,通过在目标域中添加语料库来重新训练模型已经取得了有希望的结果(例如 Tsuruoka et al 2005、Lease et al 2005)。在本文中,我们比较了针对 GENIA 和 Penn BioIE 语料库训练的词性标注器适应果蝇(果蝇)域的两种方法。
For the domain of biomedical research abstracts, two large corpora, namely GENIA (Kim et al 2003) and Penn BioIE (Kulik et al 2004) are available. Both are basically in human domain and the performance of systems trained on these corpora when they are applied to abstracts dealing with other species is unknown. In machine-learning-based systems, re-training the model with addition of corpora in the target domain has achieved promising results (e.g. Tsuruoka et al 2005, Lease et al 2005). In this paper, we compare two methods for adaptation of POS taggers trained for GENIA and Penn BioIE corpora to Drosophila melanogaster (fruit fly) domain.