Thinking about GPT-3 In-Context Learning for Biomedical IE? Think Again

Thinking about GPT-3 In-Context Learning for Biomedical IE? Think Again
复制标题

DOI:
10.48550/arxiv.2203.08410
复制
发表时间:
2022-03
期刊:
--
影响因子:
--
通讯作者:
Bernal Jimenez Gutierrez;Nikolas McNeal;Clay Washington;You Chen;Lang Li;Huan Sun;Yu Su
Bernal Jimenez Gutierrez;Nikolas McNeal;Clay Washington;You Chen;Lang Li;Huan Sun;Yu Su
中科院分区:
其他
文献类型:
--
作者:
Bernal Jimenez Gutierrez;Nikolas McNeal;Clay Washington;You Chen;Lang Li;Huan Sun;Yu Su

文献摘要

被引文献

相似文献

大型预训练语言模型(PLM)(例如GPT-3)的强烈射击中文化学习能力对应用程序域(例如生物医学)高度吸引力,例如生物医学,这些域具有高度和多样化的语言技术需求,但也具有很高的数据注释成本。在本文中,我们介绍了第一项系统和全面的研究,用于比较GPT-3中文本学习的几次表现,并在两个高度代表性的生物医学信息提取任务上进行微调较小(即Bert尺寸)PLMS,命名为实体识别和关系提取。我们遵循真正的几弹性设置,以避免高估模型在大型验证集上的模型选择中的少量射击性能。我们还通过已知技术(例如上下文校准和动态内部的示例检索)来优化GPT-3的性能。但是,我们的结果表明,与简单地调整较小的PLM相比,GPT-3的表现仍然显着不足。此外,当更多培训数据可用时,GPT-3中文学习还会产生较小的准确性。我们的深入分析进一步揭示了可能对信息提取任务有害的信息中的学习设置。考虑到GPT-3实验的高昂成本,我们希望我们的研究为生物医学研究人员和从业人员提供指导,向更有前途的方向(例如微调小PLM)提供指导。
The strong few-shot in-context learning capability of large pre-trained language models (PLMs) such as GPT-3 is highly appealing for application domains such as biomedicine, which feature high and diverse demands of language technologies but also high data annotation costs. In this paper, we present the first systematic and comprehensive study to compare the few-shot performance of GPT-3 in-context learning with fine-tuning smaller (i.e., BERT-sized) PLMs on two highly representative biomedical information extraction tasks, named entity recognition and relation extraction. We follow the true few-shot setting to avoid overestimating models' few-shot performance by model selection over a large validation set. We also optimize GPT-3's performance with known techniques such as contextual calibration and dynamic in-context example retrieval. However, our results show that GPT-3 still significantly underperforms compared to simply fine-tuning a smaller PLM. In addition, GPT-3 in-context learning also yields smaller gains in accuracy when more training data becomes available. Our in-depth analyses further reveal issues of the in-context learning setting that may be detrimental to information extraction tasks in general. Given the high cost of experimenting with GPT-3, we hope our study provides guidance for biomedical researchers and practitioners towards more promising directions such as fine-tuning small PLMs.