Exploring the Value of Pre-trained Language Models for Clinical Named Entity Recognition

Exploring the Value of Pre-trained Language Models for Clinical Named Entity Recognition
复制标题

DOI:
10.1109/bigdata59044.2023.10386154
复制
发表时间:
2022-10
期刊:
2023 IEEE International Conference on Big Data (BigData)
影响因子:
--
通讯作者:
Yuping Wu;Lifeng Han;Valerio Antonini;G. Nenadic
Yuping Wu;Lifeng Han;Valerio Antonini;G. Nenadic
中科院分区:
其他
文献类型:
--
作者:
Yuping Wu;Lifeng Han;Valerio Antonini;G. Nenadic

文献摘要

相似文献

在自然语言处理(NLP)领域中,将预训练语言模型(PLMs)从一般或特定领域的数据微调到资源有限的特定任务的实践已经得到了广泛的应用。在这项工作中,我们重新审视了这一假设,并对临床NLP进行了调查,特别是对药物及其相关属性的命名实体识别(NER)。我们比较了从零开始训练的Transformer模型和微调的基于BERT的大型语言模型(llm),即BERT、BioBERT和ClinicalBERT。此外,我们研究了额外的条件随机场(CRF)层对这些模型的影响,以鼓励上下文学习。我们使用n2c2-2018共享任务数据进行模型开发和评估。实验结果表明:1)CRF层改进了所有语言模型;2)参考使用宏观平均F1分数进行的biostrict span level评价,尽管微调后的llm获得了0.83+的分数,但从头训练的TransformerCRF模型获得了0.78+的分数,其性能相当,成本更低,例如训练参数减少了39.80%;3)采用F1评分加权平均评价BIO-strict跨度水平时,ClinicalBERT-CRF、BERT-CRF和TransformerCRF得分差异较小,分别为97.59%、97.44%和96.84%。4)通过下采样进行有效的训练以获得更好的数据分布,进一步降低了训练成本和对数据的需求,同时保持了相似的分数——即与使用完整数据集相比,大约低0.02分。这个TRANSFORMERCRF项目托管在https://github.com/HECTA-UoM/TransformerCRF
The practice of fine-tuning Pre-trained Language Models (PLMs) from general or domain-specific data to a specific task with limited resources, has gained popularity within the field of natural language processing (NLP). In this work, we re-visit this assumption and carry out an investigation in clinical NLP, specifically Named Entity Recognition (NER) on drugs and their related attributes. We compare Transformer models that are trained from scratch to fine-tuned BERT-based Large Language Models (LLMs) namely BERT, BioBERT, and ClinicalBERT. Furthermore, we examine the impact of an additional Conditional Random Field (CRF) layer on such models to encourage contextual learning. We use n2c2-2018 shared task data for model development and evaluations. The experimental outcomes show that 1) CRF layers improved all language models; 2) referring to BIO-strict span level evaluation using macro-average F1 score, although the fine-tuned LLMs achieved 0.83+ scores, the TransformerCRF model trained from scratch achieved 0.78+, demonstrating comparable performances with much lower cost, e.g. with 39.80% less training parameters; 3) referring to BIO-strict span-level evaluation using weighted-average F1 score, ClinicalBERT-CRF, BERT-CRF, and TransformerCRF exhibited lower score differences, with 97.59%/97.44%/96.84% respectively. 4) applying efficient training by down-sampling for better data distribution further reduced the training cost and need for data, while maintaining similar scores -i.e. around 0.02 points lower compared to using the full dataset. This This TRANSFORMERCRF project is hosted at https://github.com/HECTA-UoM/TransformerCRF