LDA v. LSA: A Comparison of Two Computational Text Analysis Tools for the Functional Categorization of Patents

LDA v. LSA: A Comparison of Two Computational Text Analysis Tools for the Functional Categorization of Patents
复制标题

DOI:
--
复制
发表时间:
2016
期刊:
--
影响因子:
--
通讯作者:
Tonči Cvitanić;Bumsoo Lee;H. Song;Katherine K. Fu;D. Rosen
Tonči Cvitanić;Bumsoo Lee;H. Song;Katherine K. Fu;D. Rosen
中科院分区:
其他
文献类型:
--
作者:
Tonči Cvitanić;Bumsoo Lee;H. Song;Katherine K. Fu;D. Rosen

文献摘要

相似文献

.在实践中支持类比设计(DbA)的一种方法是让设计人员有效地访问源类比作为解决问题的灵感。专利数据库已用于许多DbA支持工作,因为它是一个预先存在的编目技术库。潜在语义分析(LSA)已被证明是一种有效的计算文本处理方法,用于提取专利之间有意义的相似性,以便在DbA期间进行有用的功能探索。然而,这仅在小规模(100项专利)上显示出有用。考虑到专利数据库的巨大性和大规模的现实探索,重要的是要考虑这些计算分析如何随着数量级的数据而变化。我们分析了1,000个随机的机械专利,比较了LSA和潜在狄利克雷分配(LDA)将专利分类为有意义的组的能力。所产生的影响,大(r)规模数据挖掘的专利DbA的支持进行了详细说明。
. One means to support for design-by-analogy (DbA) in practice involves giving designers efficient access to source analogies as inspiration to solve problems. The patent database has been used for many DbA support efforts, as it is a pre-existing repository of catalogued technology. Latent Semantic Analysis (LSA) has been shown to be an effective computational text processing method for extracting meaningful similarities between patents for useful functional exploration during DbA. However, this has only been shown to be useful at a small-scale (100 patents). Considering the vastness of the patent database and realistic exploration at a large-scale, it is important to consider how these computational analyses change with orders of magnitude more data. We present analysis of 1,000 random mechanical patents, comparing the ability of LSA to Latent Dirichlet Allocation (LDA) to categorize patents into meaningful groups. Resulting implications for large(r) scale data mining of patents for DbA support are detailed.