Latent Semantic Analysis: five methodological recommendations
Latent Semantic Analysis: five methodological recommendations
复制标题
DOI:
10.1057/ejis.2010.61
复制
发表时间:
2012-01-01
影响因子:
9.5
通讯作者:
Prybutok, Victor R.
中科院分区:
文献类型:
--
作者:
Evangelopoulos, Nicholas;Zhang, Xiaoni;Prybutok, Victor R.
The recent influx in generation, storage, and availability of textual data presents researchers with the challenge of developing suitable methods for their analysis. Latent Semantic Analysis (LSA), a member of a family of methodological approaches that offers an opportunity to address this gap by describing the semantic content in textual data as a set of vectors, was pioneered by researchers in psychology, information retrieval, and bibliometrics. LSA involves a matrix operation called singular value decomposition, an extension of principal component analysis. LSA generates latent semantic dimensions that are either interpreted, if the researcher's primary interest lies with the understanding of the thematic structure in the textual data, or used for purposes of clustering, categorization, and predictive modeling, if the interest lies with the conversion of raw text into numerical data, as a precursor to subsequent analysis. This paper reviews five methodological issues that need to be addressed by the researcher who will embark on LSA. We examine the dilemmas, present the choices, and discuss the considerations under which good methodological decisions are made. We illustrate these issues with the help of four small studies, involving the analysis of abstracts for papers published in the European Journal of Information Systems. European Journal of Information Systems (2012) 21, 70-86. doi:10.1057/ejis.2010.61; published online 21 December 2010