Programming word embeddings in Snap!

Programming word embeddings in Snap!
复制标题

在 Snap 中编程词嵌入!

DOI:
--
复制
发表时间:
--
期刊:
影响因子:
--
通讯作者:
Ming Gao
Ming Gao
中科院分区:
--
文献类型:
--
作者:
Ken Kahn;Yu Lu;N. Winters;Ming Gao

文献摘要

被引文献

相似文献

词嵌入是自然语言处理中的一种技术,它将词嵌入到高维空间中。它们被用于情感分析、实体检测、推荐系统、释义、文本摘要、问答、翻译以及历史和地理语言学。我们描述一个Snap!包含15种语言的20,000个单词嵌入的库。使用一个报告任何已知单词的300个数字列表的块,可以创建搜索相似单词的程序,找到其他单词的平均值,探索文化偏见,并解决单词类比问题。这些程序可以在单一语言中工作,也可以依靠不同语言的词嵌入空间的对齐来执行粗略的翻译。要用词嵌入进行计算,需要执行向量算术。这可以通过提供矢量算术块来实现。更高级的用户可以利用Snap!支持高阶函数使用列表映射块来执行向量操作。
Word embeddings is a technique in natural language processing whereby words are ​ embedded in a high-dimensional space. They are used in sentiment analysis, entity detection, recommender systems, paraphrasing, text summarisation, question answering, translation, and historical and geographic linguistics. We describe a Snap! library that contains 20,000 word embeddings in 15 languages. Using a block that reports a list of 300 numbers for any of the known words, one can create programs that search for similar words, find words that are the average of other words, explore cultural biases, and solve word analogy problems. These programs can work in a single language or rely upon the alignment of the word embedding spaces of different languages to perform rough translations. To compute with word embeddings one needs to perform vector arithmetic. This can be accomplished by providing vector arithmetic blocks. More advanced users can instead take advantage of Snap!’s support of higher-order functions to use list mapping blocks to perform the vector operations.