Random indexing of text samples for latent semantic analysis

Random indexing of text samples for latent semantic analysis
复制标题

DOI:
--
复制
发表时间:
2000
期刊:
--
影响因子:
--
通讯作者:
P. Kanerva;Jan Kristoferson;Anders Holst
P. Kanerva;Jan Kristoferson;Anders Holst
中科院分区:
其他
文献类型:
--
作者:
P. Kanerva;Jan Kristoferson;Anders Holst

文献摘要

被引文献

相似文献

用于潜在语义分析的文本样本随机索引Pentti Kanerva Jan Kristoferson Anders Holst kanerva@sics.se, janke@sics.se, aho@sics.se RWCP理论基础SICS实验室瑞典计算机科学研究所,Box 1263, SE-16429 Kista, Sweden潜在语义分析是一种计算向量|的方法,它有几个随机放置的;1和高维语义向量,或上下文向量,1,其余的0(例如,1和1各4个,或者来自它们的共现统计的单词。由Landauer & Dumais(1997)的研究,每1800个实验中有8个非0,而不是每30000个实验中有1个非0,这涵盖了一个词汇量(如上所述)。因此,我们将在30,000个上下文(60,000个30,000个文本样本)中积累相同的60,000个单词的数据集(由分隔成60,000个1,800个单词的逐上下文矩阵的唯一字母字符串,而不是单词空间字符)。或150字左右的文件)。我们的方法已经用不同的数据进行了验证,第一个数据收集到由79,000个共现矩阵组成的60,000个30,000个单词/上下文的1000万单词/ TASA语料库中,每一行代表一个单词词汇表(当单词在第8行之后被截断时,每列代表一个文本样本,以便每个字符)37,600个文本样本。该数据被计算机输入,给出给定单词在给定单词中出现的频率,模拟成一个79000个单词按上下文排列的矩阵,即文本样本。对频率进行归一化处理,将归一化后的归一化矩阵用奇异值变换成归一化矩阵;1 0 1。未归一化的1800维分解(SVD)将其原始的30,000个文档上下文向量减少,使托福考试维度的35{44%正确率降低到潜在数量要少得多,而归一化的维度正确率为48{51%,其中300维被证明是最优的。因此,单词响应Landauer & Dumais的36%的常态-由300维语义向量表示。在SVD之前化了30,000维向量,对于一个dier-所有这一切的重点是向量捕获语料库(见上文)。我们的词根据上下文矩阵的意思。Landauer和Dumais用a证明了它可以进一步转换,例如用SVD作为同义词测试托福(用于\ test Of English作为LSA),除了矩阵要小得多。数学上,是30,000或37,600维(外语)。对于每个测试词,四个备选索引向量是正交的,而给出了1800个维度,并要求参赛者找出最同义的那个。随机选择横向的,只是几乎正交的。它们似乎会产生25%的正确率。然而,当语义工作良好时,除了测试词的语义向量比语义类脑向量多,而且受四个备选词的文本语义向量数量影响较小之外,它将大多数组(1800维索引向量在64%的情况下可以覆盖很宽的范围)与正确的备选词高度相关。文本样本数量不等)。然而,当同样的测试是基于30,000个向量,也在狭窄的上下文窗口中索引单词时,维度向量在SVD之前,结果几乎没有得到62.70%的正确率,并得出结论,随机输入也很好:只有36%的正确率。作者的结论是,索引值得更充分地研究和理解。致谢。这项研究在某种程度上得到了SVD信息重组的支持——日本国际贸易和工业部回应了人类的心理。在Real World Computing Partnership下,我们研究了高维随机分布(MITI) (RWCP)的TASA语料库和80个托福表征,作为测试项目程序的类脑表征模型。由Pro- information提供给我们(Kanerva, 1994; Kanerva & Sj - odin, 1999)。托马斯·兰道尔教授,科罗拉多大学。在这张海报中,我们报告了使用这种表示来降低原始单词-上下文矩阵的维数。该方法可以解释Kanerva, P.(1994)。通过查看多个级别的60000 - 30000个频率概念矩阵来进行编码的飞溅代码。在马里纳罗先生和p.g.上面。假设每个文本样本都由Morasso(编辑),ICANN '94, Proc. international 'l Conference的30,000位向量表示,其中单个1标记人工神经网络(Sorrento, Italy),卷1,样本在所有样本列表中的位置,并将其称为样本的pp. 226{229}。伦敦:斯普林格出版社。索引向量(即Kanerva, P.和Sj - odin, G.(1999)的索引向量的第n位)。随机模式:文本样本为1|,表示为一元或低计算。程序。2000真实世界计算系统)。然后是频率的单词-上下文矩阵(TR-99-002报告,第271{276页)。筑波-可以通过以下程序获得:每次城市,日本:真实世界计算伙伴关系。由于单词w出现在第n个文本样本中,因此将第n个索引向量添加到单词w的行中。兰道尔,T. K.和杜马,S. T.(1997)。我们使用相同的程序来积累单词——柏拉图的问题:根据上下文的潜在语义分析矩阵,除了索引向量是获取、归纳和表示的理论——不是统一的。文本样本的索引向量是知识的一小部分。心理评论104(2):211{通过比较|我们使用了1800维指数
Random Indexing of Text Samples for Latent Semantic Analysis Pentti Kanerva Jan Kristoferson Anders Holst kanerva@sics.se, janke@sics.se, aho@sics.se RWCP Theoretical Foundation SICS Laboratory Swedish Institute of Computer Science, Box 1263, SE-16429 Kista, Sweden Latent Semantic Analysis is a method of computing vectors|and it has several randomly placed ; 1s and high-dimensional semantic vectors, or context vectors, 1s, with the rest 0s (e.g., four each of ; 1 and 1, or for words from their co-occurrence statistics. An exper- eight non-0s in 1,800, instead of one non-0 in 30,000 iment by Landauer & Dumais (1997) covers a vocabu- as above). Thus, we would accumulate the same data lary of 60,000 words (unique letter strings delimited by into a 60,000 1,800 words-by-contexts matrix instead word-space characters) in 30,000 contexts (text samples of 60,000 30,000. or \documents of about 150 words each). The data are Our method has been veried with dierent data, a rst collected into a 60,000 30,000 words-by-contexts ten-million-word \TASA corpus consisting of a 79,000- co-occurrence matrix, with each row representing a word word vocabulary (when words are truncated after the 8th and each column representing a text sample so that each character) in 37,600 text samples. The data were accu- entry gives the frequency of a given word in a given mulated into a 79,000 1,800 words-by-contexts matrix, text sample. The frequencies are normalized, and the which was normalized by thresholding into a matrix of normalized matrix is transformed with Singular-Value ; 1s, 0s, and 1s. The unnormalized 1,800-dimensional Decomposition (SVD) reducing its original 30,000 doc- context vectors gave 35{44% correct in the TOEFL test ument dimensions into a much smaller number of latent and the normalized ones gave 48{51% correct, which cor- dimensions, 300 proving to be optimal. Thus words are respond to Landauer & Dumais' 36% for their normal- represented by 300-dimensional semantic vectors. ized 30,000-dimensional vectors before SVD, for a dier- The point in all of this is that the vectors capture ent corpus (see above). Our words-by-contexts matrix meaning. Landauer and Dumais demonstrate it with a can be transformed further, for example with SVD as in synonym test called TOEFL (for \Test Of English as a LSA, except that the matrix is much smaller. Mathematically, the 30,000- or 37,600-dimensional in- Foreign Language ). For each test word, four alterna- dex vectors are orthogonal, whereas the 1,800-dimen- tives are given, and the \contestant is asked to nd the one that's the most synonymous. Choosing at random sional ones are only nearly orthogonal. They seem to would yield 25% correct. However, when the seman- work just as well, in addition to which they are more tic vector for the test word is compared to the seman- \brainlike and less aected by the number of text sam- tic vectors for the four alternatives, it correlates most ples (1,800-dimensional index vectors can cover a wide- highly with the correct alternative in 64% of the cases. ranging number of text samples). We have used such However, when the same test is based on the 30,000- vectors also to index words in narrow context windows, dimensional vectors before SVD, the result is not nearly getting 62{70% correct, and conclude that random in- as good: only 36% correct. The authors conclude that dexing deserves to be studied and understood more fully. Acknowledgments. This research is supported by the reorganization of information by SVD somehow cor- Japan's Ministry of International Trade and Industry responds to human psychology. under the Real World Computing Partnership We have studied high-dimensional random distributed (MITI) (RWCP) The TASA corpus and 80 TOEFL representations, as models of brainlike representation of test items program. were made available to us by courtesy of Pro- information (Kanerva, 1994; Kanerva & Sjodin, 1999). fessor Thomas Landauer, University of Colorado. In this poster we report on the use of such a repre- sentation to reduce the dimensionality of the original words-by-contexts matrix. The method can be explained Kanerva, P. (1994). References The Spatter Code for encoding by looking at the 60,000 30,000 matrix of frequencies concepts at many levels. In M. Marinaro and P. G. above. Assume that each text sample is represented by a Morasso (eds.), ICANN '94, Proc. Int'l Conference 30,000-bit vector with a single 1 marking the place of the on Articial Neural Networks (Sorrento, Italy), vol. 1, sample in a list of all samples, and call it the sample's pp. 226{229. London: Springer-Verlag. index vector (i.e., the n th bit of the index vector for the Kanerva, P., and Sjodin, G. (1999). Stochastic Pattern n th text sample is 1|the representation is unitary or lo- Computing. Proc. 2000 Real World Computing Sym- cal). Then the words-by-contexts matrix of frequencies bosium (Report TR-99-002, pp. 271{276). Tsukuba- can be gotten by the following procedure: every time city, Japan: Real World Computing Partnership. that the word w occurs in the n th text sample, the n th index vector is added to the row for the word w . Landauer, T. K., and Dumais, S. T. (1997). A solution We use the same procedure for accumulating a words- to Plato's problem: The Latent Semantic Analysis by-contexts matrix, except that the index vectors are theory of the acquisition, induction, and representa- not unitary. A text-sample's index vector is \small tion of knowledge. Psychological Review 104 (2):211{ by comparison|we have used 1,800-dimensional index