An introduction to latent semantic analysis

An introduction to latent semantic analysis
复制标题

DOI:
10.1080/01638539809545028
复制
发表时间:
1998-01-01
影响因子:
2.2
通讯作者:
Laham, D
Laham, D
中科院分区:
心理学3区
文献类型:
--
作者:
Landauer, TK;Foltz, PW;Laham, D

文献摘要

被引文献

相似文献

潜在语义分析(Latent Semantic Analysis,LSA)是一种通过对大量文本进行统计计算来提取和表示词的上下文使用意义的理论和方法(Landauer & Dumais,1997)。其基本思想是,给定单词出现和不出现的所有单词上下文的集合提供了一组相互约束,这在很大程度上决定了单词和单词集合彼此含义的相似性。LSA对人类知识的充分反映已经以各种方式建立。例如,它的分数与人类在标准词汇和主题测试中的分数重叠;它模仿人类的单词排序和类别判断;它模拟单词-单词和单词-单词词汇启动数据;并且,正如本期以下3篇文章所报道的那样,它准确地估计了段落连贯性,个别学生段落的可学习性以及文章中包含的知识的质量和数量。
Latent Semantic Analysis (LSA) is a theory and me:hod for extracting and representing the contextual-usage meaning of words by statistical computations applied to a large corpus of text (Landauer & Dumais, 1997). The underlying idea is that the aggregate of all the word contexts in which a given word does and does not appear provides a set of mutual constraints that largely determines the similarity of meaning of words and sets of words to each other. The adequacy of LSA's reflection of human knowledge has been established in a variety of ways. For example, its scores overlap those of humans on standard vocabulary and subject matter tests; it mimics human word sorting and category judgments; it simulates word-word and passage-word lexical priming data; and, as reported in 3 following articles in this issue, it accurately estimates passage coherence, learnability of passages by individual students, and the quality and quantity of knowledge contained in an essay.