Patentopia: A multi-stage patent extraction platform with disambiguation for certain semantic challenges

Patentopia: A multi-stage patent extraction platform with disambiguation for certain semantic challenges
复制标题

DOI:
10.1109/bigdata55660.2022.10020918
复制
发表时间:
2022-12
期刊:
2022 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
A. Belz;Alexandra Graddy-Reed;Fnu Shweta;Aleksandar Giga;Shivesh Meenakshi Murali
A. Belz;Alexandra Graddy-Reed;Fnu Shweta;Aleksandar Giga;Shivesh Meenakshi Murali
中科院分区:
其他
文献类型:
--
作者:
A. Belz;Alexandra Graddy-Reed;Fnu Shweta;Aleksandar Giga;Shivesh Meenakshi Murali

文献摘要

相似文献

书目名称消歧是一个重大的语义挑战,但重要的知识资产的社会科学研究的关键。在这里,我们以多种方式为创新研究做出贡献。我们显示了一个显着的同义词问题,作者的名字,并讨论如何预处理启发式步骤标准化的名称变体的帮助,但同音异义词与中文名称生成的是特别难以解决和体现在相关的位置列表。在这里,我们确定了一个新的现象“专名丰富”,频繁使用的某些单词在公司名称的语义原因,可以混淆消歧聚类算法。我们用Patentopia来说明这些问题,Patentopia是我们定制的平台,可以访问美国专利商标局数据库的PatentsView门户网站,并可供免费学术使用。这个多阶段系统使用与PatentsView聚类过程相一致的分析方法,并报告元数据以进一步帮助分析。作为高度相关的用例,我们说明了系统性能与来自两个重要的公共创新计划,I-Corps和小企业创新研究(SBIR)的数据,我们关闭与文献计量学分析当前的专利数据的影响。
Bibliographic name disambiguation is an major semantic challenge, but critical to social sciences studies of important intellectual assets. Here we contribute to innovation research in several ways. We show a significant synonym problem in author names and discuss how a pre-processing heuristic step standardizing name variants helps, but homonyms generated with Chinese names are particularly difficult to resolve and manifest in an associated location list. Here we identify a new phenomenon of "onomastic profusion," the frequent use of certain words in firm names for semantic reasons that can confound disambiguation clustering algorithms. We illustrate these concerns with Patentopia, our customized platform accessing the PatentsView portal for the United States Patent and Trademark Office database and available for free academic use. This multi-stage system uses heuristics in concert with the PatentsView clustering process and reports meta-data to further assist analysis. As highly relevant use cases, we illustrate system performance with data derived from two important public innovation programs, I-Corps and Small Business Innovation Research (SBIR), and we close with implications for bibliometric analysis of current patent data.