LDA v. LSA: A Comparison of Two Computational Text Analysis Tools for the Functional Categorization of Patents
LDA v. LSA: A Comparison of Two Computational Text Analysis Tools for the Functional Categorization of Patents
复制标题
DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
Tonči Cvitanić;Bumsoo Lee;H. Song;Katherine K. Fu;D. Rosen
中科院分区:
文献类型:
--
作者:
Tonči Cvitanić;Bumsoo Lee;H. Song;Katherine K. Fu;D. Rosen
. One means to support for design-by-analogy (DbA) in practice involves giving designers efficient access to source analogies as inspiration to solve problems. The patent database has been used for many DbA support efforts, as it is a pre-existing repository of catalogued technology. Latent Semantic Analysis (LSA) has been shown to be an effective computational text processing method for extracting meaningful similarities between patents for useful functional exploration during DbA. However, this has only been shown to be useful at a small-scale (100 patents). Considering the vastness of the patent database and realistic exploration at a large-scale, it is important to consider how these computational analyses change with orders of magnitude more data. We present analysis of 1,000 random mechanical patents, comparing the ability of LSA to Latent Dirichlet Allocation (LDA) to categorize patents into meaningful groups. Resulting implications for large(r) scale data mining of patents for DbA support are detailed.