Automated classification of clinical trial eligibility criteria text based on ensemble learning and metric learning.

Automated classification of clinical trial eligibility criteria text based on ensemble learning and metric learning.
复制标题

基于集成学习和度量学习的临床试验资格标准文本自动分类

DOI:
10.1186/s12911-021-01492-z
复制
发表时间:
2021-07-30
影响因子:
3.5
通讯作者:
Hao T
Hao T
中科院分区:
医学3区
文献类型:
--
作者:
Zeng K;Xu Y;Lin G;Liang L;Hao T

文献摘要

参考文献

被引文献

相似文献

资格标准是筛选临床试验目标参与者的主要策略。利用机器学习方法对临床试验资格标准文本进行自动分类,提高了招募效率,降低了临床研究成本。然而,现有方法由于资格标准文本数据的复杂性和不平衡性,导致分类性能较差。方法提出一种基于集成学习的度量学习模型,用于资格标准的分类。该模型集成了一组预训练模型,包括来自变形器的双向编码器表示(BERT)、鲁棒优化的BERT预训练方法(RoBERTa)、XLNet、作为鉴别器而不是生成器的文本编码器预训练(ELECTRA)和通过知识集成的增强表示(ERNIE)。焦损用作损失函数来解决数据不平衡问题。利用度量学习训练每个基模型的嵌入,进行特征区分。采用软投票对集成模型进行最终分类。数据集来自第五届中国卫生信息处理大会标准评价任务3,包含44个类别38341个合格标准文本。结果该方法的准确率为0.8497,精密度为0.8229,召回率为0.8216。宏观f1得分为0.8169,比最先进的基线方法平均提高0.84%。此外,通过标准t检验,性能改进的p值为2.152e-07,表明我们的模型实现了显著的改进。结论提出了一种基于多模型集成学习和度量学习的临床试验合格标准文本分类模型。实验表明,集成模型显著提高了分类性能。此外,度量学习能够改善词嵌入表示,焦点丢失减少了数据不平衡对模型性能的影响。
BackgroundEligibility criteria are the primary strategy for screening the target participants of a clinical trial. Automated classification of clinical trial eligibility criteria text by using machine learning methods improves recruitment efficiency to reduce the cost of clinical research. However, existing methods suffer from poor classification performance due to the complexity and imbalance of eligibility criteria text data.MethodsAn ensemble learning-based model with metric learning is proposed for eligibility criteria classification. The model integrates a set of pre-trained models including Bidirectional Encoder Representations from Transformers (BERT), A Robustly Optimized BERT Pretraining Approach (RoBERTa), XLNet, Pre-training Text Encoders as Discriminators Rather Than Generators (ELECTRA), and Enhanced Representation through Knowledge Integration (ERNIE). Focal Loss is used as a loss function to address the data imbalance problem. Metric learning is employed to train the embedding of each base model for feature distinguish. Soft Voting is applied to achieve final classification of the ensemble model. The dataset is from the standard evaluation task 3 of 5th China Health Information Processing Conference containing 38,341 eligibility criteria text in 44 categories.ResultsOur ensemble method had an accuracy of 0.8497, a precision of 0.8229, and a recall of 0.8216 on the dataset. The macro F1-score was 0.8169, outperforming state-of-the-art baseline methods by 0.84% improvement on average. In addition, the performance improvement had a p-value of 2.152e-07 with a standard t-test, indicating that our model achieved a significant improvement.ConclusionsA model for classifying eligibility criteria text of clinical trials based on multi-model ensemble learning and metric learning was proposed. The experiments demonstrated that the classification performance was improved by our ensemble model significantly. In addition, metric learning was able to improve word embedding representation and the focal loss reduced the impact of data imbalance to model performance.
DOI: 10.1016/j.jbi.2014.01.009
发表时间: 2014-12
影响因子: 4.5
作者:
Hao, Tianyong;Rusanov, Alexander;Boland, Mary Regina;Weng, Chunhua
通讯作者: Weng, Chunhua
DOI: 10.1200/jco.2017.74.4144
发表时间: 2017-11-20
影响因子: 45.3
作者:
Gore, Lia;Ivy, S. Percy;Reaman, Gregory
通讯作者: Reaman, Gregory
DOI: 10.1093/jamia/ocx019
发表时间: 2017-11-01
影响因子: 6.4
作者:
Kang, Tian;Zhang, Shaodian;Weng, Chunhua
通讯作者: Weng, Chunhua
DOI: 10.1200/jop.2012.000646
发表时间: 2012-11-01
影响因子: --
作者:
Penberthy, Lynne T.;Dahman, Bassam A.;DeShazo, Jonathan P.
通讯作者: DeShazo, Jonathan P.
DOI: 10.1093/bib/bbv024
发表时间: 2016-01-01
影响因子: 9.5
作者:
Huang, Chung-Chi;Lu, Zhiyong
通讯作者: Lu, Zhiyong