Domain-specific language models and lexicons for tagging

Domain-specific language models and lexicons for tagging
复制标题

DOI:
10.1016/j.jbi.2005.02.009
复制
发表时间:
2005-12-01
影响因子:
4.5
通讯作者:
Chute, CG
Chute, CG
中科院分区:
医学3区
文献类型:
--
作者:
Coden, AR;Pakhomov, SV;Chute, CG

文献摘要

被引文献

相似文献

准确可靠的词性标记对于许多自然语言处理 (NLP) 任务非常有用,这些任务构成了基于 NLP 的信息检索和数据挖掘方法的基础。一般来说,需要大型注释语料库才能达到所需的词性标注准确性。我们表明,大型带注释的通用英语语料库不足以构建足以标记医学领域文档的词性标注器模型。然而,在大型通用英语语料库中添加一个相当小的特定领域语料库,可以将性能从我们研究中的 87% 提高到 92% 以上。我们还建议了一些特征来量化训练语料库和测试数据之间的相似性。这些结果为创建适当的语料库来构建词性标注器模型提供了指导,该模型以相对较小的成本在新领域上提供了令人满意的准确性结果。 (c) 2005 Elsevier Inc. 保留所有权利。
Accurate and reliable part-of-speech tagging is useful for many Natural Language Processing (NLP) tasks that form the foundation of NLP-based approaches to information retrieval and data mining. In general, large annotated corpora are necessary to achieve desired part-of-speech tagger accuracy. We show that a large annotated general-English corpus is not sufficient for building a part-of-speech tagger model adequate for tagging documents from the medical domain. However, adding a quite small domain-specific corpus to a large general-English one boosts performance to over 92% accuracy from 87% in our studies. We also suggest a number of characteristics to quantify the similarities between a training corpus and the test data. These results give guidance for creating an appropriate corpus for building a part-of-speech tagger model that gives satisfactory accuracy results on a new domain at a relatively small cost. (c) 2005 Elsevier Inc. All rights reserved.