Word Representations: A Simple and General Method for Semi-Supervised Learning

Word Representations: A Simple and General Method for Semi-Supervised Learning
复制标题

DOI:
--
复制
发表时间:
2010-07
期刊:
--
影响因子:
--
通讯作者:
Joseph P. Turian;Lev-Arie Ratinov;Yoshua Bengio
Joseph P. Turian;Lev-Arie Ratinov;Yoshua Bengio
中科院分区:
其他
文献类型:
--
作者:
Joseph P. Turian;Lev-Arie Ratinov;Yoshua Bengio

文献摘要

被引文献

相似文献

如果我们采用现有的有监督 NLP 系统,提高准确性的一种简单而通用的方法是使用无监督词表示作为额外的词特征。我们评估了 NER 和分块上的 Brown 聚类、Collobert 和 Weston (2008) 嵌入以及 HLBL (Mnih & Hinton, 2009) 单词嵌入。我们使用接近最先进的监督基线,并发现这三个单词表示中的每一个都提高了这些基线的准确性。我们通过组合不同的单词表示找到了进一步的改进。您可以在此处下载我们的单词功能以及我们的代码,以便在现有 NLP 系统中使用现成的功能:http://metaoptimize.com/projects/wordreprs/
If we take an existing supervised NLP system, a simple and general way to improve accuracy is to use unsupervised word representations as extra word features. We evaluate Brown clusters, Collobert and Weston (2008) embeddings, and HLBL (Mnih & Hinton, 2009) embeddings of words on both NER and chunking. We use near state-of-the-art supervised baselines, and find that each of the three word representations improves the accuracy of these baselines. We find further improvements by combining different word representations. You can download our word features, for off-the-shelf use in existing NLP systems, as well as our code, here: http://metaoptimize.com/projects/wordreprs/