An Ensemble Learning Strategy for Eligibility Criteria Text Classification for Clinical Trial Recruitment: Algorithm Development and Validation

An Ensemble Learning Strategy for Eligibility Criteria Text Classification for Clinical Trial Recruitment: Algorithm Development and Validation
复制标题

DOI:
10.2196/17832
复制
发表时间:
2020-07-01
影响因子:
3.2
通讯作者:
Qu, Yingying
Qu, Yingying
中科院分区:
医学3区
文献类型:
--
作者:
Zeng, Kun;Pan, Zhiwei;Qu, Yingying

文献摘要

被引文献

相似文献

背景:资格标准是筛选合适的临床试验参与者的主要策略。利用自然语言处理技术,通过数字筛选自动分析临床试验资格标准,可以提高招募效率并降低促进临床研究的成本。目的:我们的目标是创建一个自然语言处理模型来自动分类临床试验资格标准。方法:我们提出了一种基于集成学习的短文本资格标准分类器,其中集成了一组预训练模型。预训练模型包括用于训练和分类的最先进的深度学习方法,包括 Transformers 的双向编码器表示 (BERT)、XLNet 和鲁棒优化的 BERT 预训练方法 (RoBERTa)。集成模型的分类结果被组合起来作为新特征,用于训练用于资格标准分类的光梯度提升机(LightGBM)模型。结果:我们提出的方法在国际会议共享任务的标准数据集上获得了 0.846 的准确率、0.803 的精确度和 0.817 的召回率。宏观F1值为0.807,在共享任务上优于最先进的基线方法。结论:我们设计了一个基于多模型集成学习的临床试验短文本分类标准筛选模型。通过实验,我们得出的结论是,与单个模型相比,模型集成的性能得到了显着提高。焦点损失的引入可以减少类别不平衡的影响,从而获得更好的性能。
Background: Eligibility criteria are the main strategy for screening appropriate participants for clinical trials. Automatic analysis of clinical trial eligibility criteria by digital screening, leveraging natural language processing techniques, can improve recruitment efficiency and reduce the costs involved in promoting clinical research.Objective: We aimed to create a natural language processing model to automatically classify clinical trial eligibility criteria.Methods: We proposed a classifier for short text eligibility criteria based on ensemble learning, where a set of pretrained models was integrated. The pretrained models included state-of-the-art deep learning methods for training and classification, including Bidirectional Encoder Representations from Transformers (BERT), XLNet, and A Robustly Optimized BERT Pretraining Approach (RoBERTa). The classification results by the integrated models were combined as new features for training a Light Gradient Boosting Machine (LightGBM) model for eligibility criteria classification.Results: Our proposed method obtained an accuracy of 0.846, a precision of 0.803, and a recall of 0.817 on a standard data set from a shared task of an international conference. The macro F1 value was 0.807, outperforming the state-of-the-art baseline methods on the shared task.Conclusions: We designed a model for screening short text classification criteria for clinical trials based on multimodel ensemble learning. Through experiments, we concluded that performance was improved significantly with a model ensemble compared to a single model. The introduction of focal loss could reduce the impact of class imbalance to achieve better performance.