Adapting a Lexicalized-Grammar Parser to Contrasting Domains

Adapting a Lexicalized-Grammar Parser to Contrasting Domains
复制标题

DOI:
10.3115/1613715.1613775
复制
发表时间:
2008-10
期刊:
--
影响因子:
--
通讯作者:
Laura Rimell;S. Clark
Laura Rimell;S. Clark
中科院分区:
其他
文献类型:
--
作者:
Laura Rimell;S. Clark

文献摘要

被引文献

相似文献

大多数最先进的广覆盖解析器都是在报纸文本上训练的,在其他领域的准确性会下降,这使得解析器的适应成为一个紧迫的问题。在本文中,我们证明了CCG解析器可以适用于两个新的领域,生物医学文本和QA系统的问题,通过仅在词法类别级别使用手动注释的训练数据。这种方法实现了与报纸数据相当的解析器精度,而不需要在新域中使用带注释的解析树。我们发现,在词汇类别级别上的再训练对问题的性能提高比对生物医学文本的性能提高更大,并分析了这两个数据集,以调查为什么不同的领域在解析器适应方面可能表现不同。
Most state-of-the-art wide-coverage parsers are trained on newspaper text and suffer a loss of accuracy in other domains, making parser adaptation a pressing issue. In this paper we demonstrate that a CCG parser can be adapted to two new domains, biomedical text and questions for a QA system, by using manually-annotated training data at the pos and lexical category levels only. This approach achieves parser accuracy comparable to that on newspaper data without the need for annotated parse trees in the new domain. We find that retraining at the lexical category level yields a larger performance increase for questions than for biomedical text and analyze the two datasets to investigate why different domains might behave differently for parser adaptation.