Large Language Models are Zero-Shot Clinical Information Extractors

Large Language Models are Zero-Shot Clinical Information Extractors
复制标题

大型语言模型是零样本临床信息提取器

DOI:
--
复制
发表时间:
2022
期刊:
arXiv.org
影响因子:
--
通讯作者:
David A. Sontag
David A. Sontag
中科院分区:
--
文献类型:
--
作者:
Monica Agrawal;S. Hegselmann;Hunter Lang;Yoon Kim;David A. Sontag

文献摘要

参考文献

被引文献

相似文献

我们表明,大型语言模型,如GPT-3[1],尽管没有专门针对临床领域进行训练,但在从临床文本中提取零采样信息方面表现良好。我们提出了几个例子,展示了如何使用这些模型作为工具来完成(i)概念消歧,(ii)证据提取,(iii)共同参考解决,(iv)概念提取,所有这些都是在临床文本上。良好性能的关键是使用简单的特定于任务的程序,这些程序从语言模型输出映射到任务的标签空间。我们将这些程序称为解析器,它是定义输出令牌和离散标签空间[2]之间映射的语言分析器的泛化。在我们的示例中,我们展示了好的解析器共享公共组件(例如,确保语言模型输出忠实地匹配输入数据的“安全检查”),并且跨任务的公共模式使解析器轻量级且易于创建。为了更好地评估这些系统,我们还引入了两个新的数据集,用于基准测试零射击临床信息提取,这些数据集基于CASI数据集[3]的手动重新标记,并为新任务添加标签。在我们研究的临床提取任务中,GPT-3 +解析器系统明显优于现有的零次和少次基线。
We show that large language models , such as GPT-3 [1], perform well at zero-shot information extraction from clinical text despite not being trained specifically for the clinical domain. We present several examples showing how to use these models as tools for the diverse tasks of (i) concept disambiguation, (ii) evidence extraction, (iii) coreference resolution, and (iv) concept extraction, all on clinical text . The key to good performance is the use of simple task-specific programs that map from the language model outputs to the label space of the task. We refer to these programs as resolvers , a generalization of the verbalizer which defines a mapping between output tokens and a discrete label space [2]. We show in our examples that good resolvers share common components (e.g., “safety checks” that ensure the language model outputs faithfully match the input data), and that the common patterns across tasks make resolvers lightweight and easy to create. To better evaluate these systems, we also introduce two new datasets for benchmarking zero-shot clinical information extraction based on manual relabeling of the CASI dataset [3] with labels for new tasks. On the clinical extraction tasks we studied, the GPT-3 + resolver systems significantly outperform existing zero-and few-shot baselines.
DOI: 10.18653/v1/2020.emnlp-main.685
发表时间: 2020-10
期刊: ArXiv
影响因子: --
作者:
Shubham Toshniwal;Sam Wiseman;Allyson Ettinger;Karen Livescu;Kevin Gimpel
通讯作者: Shubham Toshniwal;Sam Wiseman;Allyson Ettinger;Karen Livescu;Kevin Gimpel
DOI: 10.48550/arxiv.2203.08410
发表时间: 2022-03
期刊: --
影响因子: --
作者:
Bernal Jimenez Gutierrez;Nikolas McNeal;Clay Washington;You Chen;Lang Li;Huan Sun;Yu Su
通讯作者: Bernal Jimenez Gutierrez;Nikolas McNeal;Clay Washington;You Chen;Lang Li;Huan Sun;Yu Su
DOI: --
发表时间: 2019
期刊: AMIA Joint Summits on Translational Science proceedings. AMIA Joint Summits on Translational Science
影响因子: --
作者:
Yuqi Si;Kirk Roberts
通讯作者: Yuqi Si;Kirk Roberts
DOI: --
发表时间: 2014
期刊: J. Am. Medical Informatics Assoc.
影响因子: --
作者:
Rachel Chasin;Anna Rumshisky;Özlem Uzuner;Peter Szolovits
通讯作者: Rachel Chasin;Anna Rumshisky;Özlem Uzuner;Peter Szolovits
理解护理笔记中的缩写:死亡率预测的案例研究。
DOI: --
发表时间: 2019
期刊: AMIA Joint Summits on Translational Science proceedings. AMIA Joint Summits on Translational Science
影响因子: --
作者:
Nakayama,JasmineY;Hertzberg,Vicki;Ho,JoyceC
通讯作者: Ho,JoyceC