Explorations of morphological structure in distributional space

Explorations of morphological structure in distributional space
复制标题

DOI:
10.1075/ml.00021.baa
复制
发表时间:
2023-11
期刊:
The Mental Lexicon
影响因子:
--
通讯作者:
Harald Baayen;Dunstan Brown;Yu-Ying Chuang
Harald Baayen;Dunstan Brown;Yu-Ying Chuang
中科院分区:
其他
文献类型:
--
作者:
Harald Baayen;Dunstan Brown;Yu-Ying Chuang

文献摘要

相似文献

本期特刊汇集了五个研究成果,它们是两个研究项目之间密切互动的成果:由英国艺术与人文研究委员会资助的“盛宴与饥荒”项目,以及由欧洲研究委员会资助的WIDE项目。“盛宴与饥荒”项目在形态范式中解决了过剩和缺陷问题。WIDE项目研究了一种心理词汇和形态处理模型,其中形式和意义由高维数字向量表示。将这两个项目结合在一起的是对探索分布语义对理解形态学的有用性的共同兴趣。分布语义学是人工智能、心理学和计算语义学交叉的一个研究领域,它通过从大型语料库中计算出实数的高维向量来表示单词的含义。有许多方法可以获得这样的向量,通常被称为“嵌入”或“语义向量”。潜在语义分析(Landauer and Dumais, 1997)方法首先计算单词在文档中出现的频率,得到一个单词按文档频率表。意思相似或语义相关的词往往出现在同一份文件中。因此,具有单词文档频率的向量提供了该单词的语义指纹。第二步,对单词-文档频率表进行降维技术(奇异值分解),得到由n个潜在维组成的单词矩阵。n的典型值是300。简而言之,LSA利用全局统计信息,了解单词如何在涵盖广泛主题的文档中出现。其他各种方法使用滑动窗口技术,跟踪其他单词在目标单词的直接上下文中出现的频率(例如,HAL Burgess和Lund (1998);HiDEx, Shaoul and Westbury (2010);word2vec, Mikolov et al.(2013), FastText, Bojanowski et al.(2017))。这些方法基于单词的局部统计,而不是全局统计。FastText嵌入可用于各种语言,见https://Interactive,图可从https://doi获得。org/10.1075/ml。00021. 英国机场管理局。数字https://doi。org/10.1075/ml。00021. 在线出版:2023年9月12日The Mental Lexicon ISSN 1871-1340| E-ISSN 1871-1375在CC BY 4.0许可下提供。©2023约翰·本杰明出版公司
This special issue brings together five studies that are the fruit of intense interactions between two research projects: The ‘Feast and Famine’project funded by the UK’s Arts and Humanities Research Council, and the WIDE project funded by the European Research Council. The Feast and Famine project addresses overabundance and defectiveness in morphological paradigms. The WIDE project worked on a model of the mental lexicon and morphological processing in which form and meaning are represented by high-dimensional numeric vectors. What brought the two projects together is a shared interest in exploring the usefulness of distributional semantics for understanding morphology. Distributional semantics, a research area at the intersection of artificial intelligence, psychology, and computational semantics, represents words’ meanings by means of high-dimensional vectors of real numbers calculated from large corpora. There are many ways in which such vectors, often referred to as ‘embeddings’, or ‘semantic vectors’, can be obtained. The latent semantic analysis (Landauer and Dumais, 1997) method first calculates how often words occur in documents, resulting in a word by document frequency table. Words that are similar in meaning or that are semantically related tend to occur in the same documents. As a consequence, the vector with a word’s document frequencies provides a semantic fingerprint of that word. As a second step, the word-document frequency table is subjected to a dimension reduction technique (singular value decomposition), resulting in a matrix of words by n latent dimensions. A typical value for n is 300. In short, LSA makes use of global statistics of how words cooccur across documents that cover a wide range of topics. Various other methods use a sliding window technique that keeps track of the frequencies with which other words occur in the immediate context of a target word (eg, HAL Burgess and Lund (1998); HiDEx, Shaoul and Westbury (2010); word2vec, Mikolov et al.(2013), and FastText, Bojanowski et al.(2017)). These methods build on the local statistics of words, rather than on their global statistics. FastText embeddings are available for a wide range of languages at https://Interactive figure available from https://doi. org/10.1075/ml. 00021. baa. figures https://doi. org/10.1075/ml. 00021. baa| Published online: 12 September 2023 The Mental Lexicon ISSN 1871-1340| E-ISSN 1871-1375 Available under the CC BY 4.0 license.© 2023 John Benjamins Publishing Company