AnaDE1.0: A Novel Data Set for Benchmarking Analogy Detection and Extraction

AnaDE1.0: A Novel Data Set for Benchmarking Analogy Detection and Extraction
复制标题

DOI:
--
复制
发表时间:
2024
影响因子:
--
通讯作者:
B. Bhavya;Shradha Sehgal;Jinjun Xiong;ChengXiang Zhai
B. Bhavya;Shradha Sehgal;Jinjun Xiong;ChengXiang Zhai
中科院分区:
计算机科学3区
文献类型:
--
作者:
B. Bhavya;Shradha Sehgal;Jinjun Xiong;ChengXiang Zhai

文献摘要

相似文献

在两个概念之间进行比较的文本类比通常用于解释复杂的想法,创造性写作和科学发现。在本文中,我们提出并研究了一个新的任务,称为类比检测和提取(AnaDE),其中包括三个协同子任务:1)检测包含类比的文档,2)提取组成类比的文本片段,以及3)识别(源和目标)比较的概念。为了便于研究这个新任务,我们通过抓取Metamia.com创建了一个基准数据集,并调查了所有子任务上最先进模型的性能,以建立这个新任务的第一代基准结果。我们发现,Longformer模型在所有三个子任务上都取得了最佳性能,证明了其处理长文本的有效性。此外,在我们的数据集上微调的较小模型比未微调的ChatGPT表现更好,这表明任务难度很高。总体而言,该模型在文档检测上实现了高性能,这表明它可以用于开发类似于类比搜索引擎的应用程序。此外,在片段和概念提取任务上还有很大的改进空间。
Textual analogies that make comparisons between two concepts are often used for explaining complex ideas, creative writing, and scientific discovery. In this paper, we propose and study a new task, called Analogy Detection and Extraction (AnaDE), which includes three synergistic sub-tasks: 1) detecting documents containing analogies, 2) extracting text segments that make up the analogy, and 3) identifying the (source and target) concepts being compared. To facilitate the study of this new task, we create a benchmark dataset by scraping Metamia.com and investigate the performances of state-of-the-art models on all sub-tasks to establish the first-generation benchmark results for this new task. We find that the Longformer model achieves the best performance on all the three sub-tasks demonstrating its effectiveness for handling long texts. Moreover, smaller models fine-tuned on our dataset perform better than non-finetuned ChatGPT, suggesting high task difficulty. Overall, the models achieve a high performance on documents detection suggesting that it could be used to develop applications like analogy search engines. Further, there is a large room for improvement on the segment and concept extraction tasks.