Explorations of morphological structure in distributional space
Explorations of morphological structure in distributional space
复制标题
DOI:
10.1075/ml.00021.baa
复制
发表时间:
2023-11
期刊:
影响因子:
--
通讯作者:
Harald Baayen;Dunstan Brown;Yu-Ying Chuang
中科院分区:
文献类型:
--
作者:
Harald Baayen;Dunstan Brown;Yu-Ying Chuang
This special issue brings together five studies that are the fruit of intense interactions between two research projects: The ‘Feast and Famine’project funded by the UK’s Arts and Humanities Research Council, and the WIDE project funded by the European Research Council. The Feast and Famine project addresses overabundance and defectiveness in morphological paradigms. The WIDE project worked on a model of the mental lexicon and morphological processing in which form and meaning are represented by high-dimensional numeric vectors. What brought the two projects together is a shared interest in exploring the usefulness of distributional semantics for understanding morphology. Distributional semantics, a research area at the intersection of artificial intelligence, psychology, and computational semantics, represents words’ meanings by means of high-dimensional vectors of real numbers calculated from large corpora. There are many ways in which such vectors, often referred to as ‘embeddings’, or ‘semantic vectors’, can be obtained. The latent semantic analysis (Landauer and Dumais, 1997) method first calculates how often words occur in documents, resulting in a word by document frequency table. Words that are similar in meaning or that are semantically related tend to occur in the same documents. As a consequence, the vector with a word’s document frequencies provides a semantic fingerprint of that word. As a second step, the word-document frequency table is subjected to a dimension reduction technique (singular value decomposition), resulting in a matrix of words by n latent dimensions. A typical value for n is 300. In short, LSA makes use of global statistics of how words cooccur across documents that cover a wide range of topics. Various other methods use a sliding window technique that keeps track of the frequencies with which other words occur in the immediate context of a target word (eg, HAL Burgess and Lund (1998); HiDEx, Shaoul and Westbury (2010); word2vec, Mikolov et al.(2013), and FastText, Bojanowski et al.(2017)). These methods build on the local statistics of words, rather than on their global statistics. FastText embeddings are available for a wide range of languages at https://Interactive figure available from https://doi. org/10.1075/ml. 00021. baa. figures https://doi. org/10.1075/ml. 00021. baa| Published online: 12 September 2023 The Mental Lexicon ISSN 1871-1340| E-ISSN 1871-1375 Available under the CC BY 4.0 license.© 2023 John Benjamins Publishing Company