Large-scale text analysis using generative language models: A case study in discovering public value expressions in AI patents

Large-scale text analysis using generative language models: A case study in discovering public value expressions in AI patents
复制标题

DOI:
10.1162/qss_a_00285
复制
发表时间:
2024-03-01
影响因子:
6.4
通讯作者:
Shapira,Philip
Shapira,Philip
中科院分区:
其他
文献类型:
--
作者:
Pelaez,Sergio;Verma,Gaurav;Shapira,Philip

文献摘要

被引文献

相似文献

我们提出了一种使用生成语言模型(GPT-4)来生成大规模文本分析的标签和基本原理的新颖方法。该方法用于发现专利中的公共价值表达。使用美国专利商标局 (USPTO) 的 154,934 份美国人工智能专利文件的文本(540 万个句子),我们设计了一个半自动化、人工监督的框架,用于识别和标记这些句子中的公共价值表达。开发了 GPT-4 提示,其中包括文本分类的定义、指南、示例和基本原理。我们使用 BLEU 分数和主题建模评估 GPT-4 生成的标签和基本原理,发现它们准确、多样化且忠实。 GPT-4 从我们的框架中实现了对公共价值表达的高级识别,它还用它来发现看不见的公共价值表达。 GPT 生成的标签用于训练基于 BERT 的分类器并在整个数据库上预测句子,在 3 类(0.85)和 2 类分类(0.91)任务中获得高 F1 分数。我们讨论了我们的方法对复杂和抽象概念进行大规模文本分析的影响。通过仔细的框架设计和交互式人类监督,我们建议生成语言模型可以在生成标签和基本原理方面提供重要帮助。
We put forward a novel approach using a generative language model (GPT-4) to produce labels and rationales for large-scale text analysis. The approach is used to discover public value expressions in patents. Using text (5.4 million sentences) for 154,934 US AI patent documents from the United States Patent and Trademark Office (USPTO), we design a semi-automated, human-supervised framework for identifying and labeling public value expressions in these sentences. A GPT-4 prompt is developed that includes definitions, guidelines, examples, and rationales for text classification. We evaluate the labels and rationales produced by GPT-4 using BLEU scores and topic modeling, finding that they are accurate, diverse, and faithful. GPT-4 achieved an advanced recognition of public value expressions from our framework, which it also uses to discover unseen public value expressions. The GPT-produced labels are used to train BERT-based classifiers and predict sentences on the entire database, achieving high F1 scores for the 3-class (0.85) and 2-class classification (0.91) tasks. We discuss the implications of our approach for conducting large-scale text analyses with complex and abstract concepts. With careful framework design and interactive human oversight, we suggest that generative language models can offer significant assistance in producing labels and rationales.