Improving Unsupervised Dependency Parsing with Knowledge from Query Logs

Improving Unsupervised Dependency Parsing with Knowledge from Query Logs
复制标题

DOI:
10.1145/2903720
复制
发表时间:
2016-06
期刊:
ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP)
影响因子:
--
通讯作者:
Xiuming Qiao;Hailong Cao;T. Zhao
Xiuming Qiao;Hailong Cao;T. Zhao
中科院分区:
其他
文献类型:
--
作者:
Xiuming Qiao;Hailong Cao;T. Zhao

文献摘要

相似文献

无监督依存句法分析由于不需要有监督和半监督依存句法分析所需的树库等昂贵的标注,近年来得到了越来越多的应用。然而,它的准确性仍然远远低于监督依赖分析器,部分原因是它们的分析模型不足以捕捉文本背后的语言现象。通过从文本中挖掘知识并将其纳入模型,可以提高无监督依赖分析的性能。在这篇文章中,语法知识是从查询日志中获得的,以帮助估计更好的概率依赖模型与效价。该方法与语言无关,通过利用搜狗和百度查询日志中的附加依赖关系,在Penn中文树库上获得了4.1%的无标注准确率的提高.实验结果表明,该模型在使用AOL查询日志的CoNLL 2007英语上取得了8.07%的改进。我们相信查询日志是许多自然语言处理(NLP)任务的有用语法知识来源。
Unsupervised dependency parsing becomes more and more popular in recent years because it does not need expensive annotations, such as treebanks, which are required for supervised and semi-supervised dependency parsing. However, its accuracy is still far below that of supervised dependency parsers, partly due to the fact that their parsing model is insufficient to capture linguistic phenomena underlying texts. The performance for unsupervised dependency parsing can be improved by mining knowledge from the texts and by incorporating it into the model. In this article, syntactic knowledge is acquired from query logs to help estimate better probabilities in dependency models with valence. The proposed method is language independent and obtains an improvement of 4.1% unlabeled accuracy on the Penn Chinese Treebank by utilizing additional dependency relations from the Sogou query logs and Baidu query logs. Morever, experiments show that the proposed model achieves improvements of 8.07% on CoNLL 2007 English using the AOL query logs. We believe query logs are useful sources of syntactic knowledge for many natural language processing (NLP) tasks.