An Empirical Investigation of Word Representations for Parsing the Web

An Empirical Investigation of Word Representations for Parsing the Web
复制标题

DOI:
--
复制
发表时间:
2013
期刊:
--
影响因子:
--
通讯作者:
Sorami Hisamoto;Kevin Duh;Yuji Matsumoto
Sorami Hisamoto;Kevin Duh;Yuji Matsumoto
中科院分区:
其他
文献类型:
--
作者:
Sorami Hisamoto;Kevin Duh;Yuji Matsumoto

文献摘要

相似文献

解析Web文本对于自然语言处理中的许多应用(例如机器翻译、信息检索和情感分析)逐渐变得重要。当前的句法分析一直集中在规范数据,如新闻专线。在华尔街日报数据集等标准基准上进行评估时,当前最先进的解析器的准确率远高于90%。然而,当它们被应用到新的领域(如Web数据)时,准确率急剧下降,仅超过80%。为了在许多依赖解析的应用程序中取得进展,我们需要能够处理此类文本的健壮的解析器。最近流行的一种方法是使用无监督的单词表示作为额外的特征。Koo等人。[1]已经表明,无监督聚类特征可以有效地改善依赖解析。Turian等人。[2]研究了组块和命名实体识别任务的聚类和无监督词嵌入特征。无监督词嵌入是表示词的密集、低维和实值向量,通常由神经语言模型诱导。他们已经表明,这些单词表示功能导致性能的改善。这些词表示是由无监督的方法诱导的,因此它们适用于新的领域,如Web,它有大量的未标记数据,但很少有标记数据。在本文中,我们调查的无监督词表示功能的依赖分析与Web文本的效果。我们考虑两种不同的词表示,即布朗聚类和词嵌入诱导神经语言模型。据我们所知,这是第一个系统地研究这些词表示的依赖解析Web文本的任务。
Parsing web text is progressively becoming important for many applications in natural language processing, such as machine translation, information retrieval, and sentiment analysis. Current syntactic parsing has been focused on canonical data such as newswires. When evaluated on standard benchmarks such as Wall Street Journal data set, current state-of-the-art parsers achieve accuracies well above 90%. However the accuracy drops dramatically when they are applied to new domains such as web data, barely over 80%. In order to make progress in many applications that rely on parsing, we need robust parsers that can handle such texts. One approach that is becoming popular recently is to use unsupervised word representations as extra features. Koo et al. [1] has shown that unsupervised clustering features are effective to improve dependency parsing. Turian et al. [2] examined clustering and unsupervised word embedding features on chunking and named entity recognition tasks. Unsupervised word embeddings are dense, low-dimensional and real-value vectors representing words, often induced by neural language models. They have shown that these word representation features lead to improvement in the performances. These word representations are induced by unsupervised methods, thus they are good for new domains such as the web, which has enormous amount of unlabeled data but little labeled data. In this paper we investigate the effect of unsupervised word representation features on dependency parsing with web texts. We consider two different kinds of word representations, namely Brown clustering and word embeddings induced from a neural language model. To the best of our knowledge, this is the first work that systematically examines these word representations on the task of dependency parsing on web text.