CCS Explorer: Relevance Prediction, Extractive Summarization, and Named Entity Recognition from Clinical Cohort Studies

CCS Explorer: Relevance Prediction, Extractive Summarization, and Named Entity Recognition from Clinical Cohort Studies
复制标题

DOI:
10.1109/bigdata55660.2022.10020807
复制
发表时间:
2022-11
期刊:
2022 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Irfan Al-Hussaini;D. An;Albert Lee;Sarah Bi;Cassie S. Mitchell
Irfan Al-Hussaini;D. An;Albert Lee;Sarah Bi;Cassie S. Mitchell
中科院分区:
其他
文献类型:
--
作者:
Irfan Al-Hussaini;D. An;Albert Lee;Sarah Bi;Cassie S. Mitchell

文献摘要

相似文献

临床队列研究(CCS),如随机临床试验,是记录临床研究的重要来源。理想情况下,临床专家检查这些文章进行探索性分析,从药物发现评估现有药物在应对新兴疾病方面的疗效到新开发药物的首次测试。然而,PubMed上每天都有100多篇关于COVID-19等单一流行疾病的文章发表。因此,医生可能需要几天时间才能找到文章并提取相关信息。我们能否开发一个系统来更快地筛选这些文章,并记录每一篇文章的关键内容?在这项工作中,我们提出了CCS资源管理器,一个端到端的系统,用于句子的相关性预测,提取摘要,患者,结果和干预实体检测CCS。CCS Explorer打包在一个基于Web的图形用户界面中,用户可以提供任何疾病名称。CCS Explorer然后根据后端自动生成的查询结果,从PubMed上的文章中提取并汇总所有相关信息。对于每项任务,CCS Explorer都会基于具有额外层的transformer微调预训练的语言表示模型。这些模型使用两个公开的数据集进行评估。CCS Explorer在使用BioBERT进行句子相关性预测时获得了80.2%的召回率、0.843的AUC-ROC和88.3%的准确率,并在使用PubMedBERT进行患者、干预、结果检测(PIO)时获得了77.8%的平均Micro F1分数。因此,CCS Explorer可以可靠地提取相关信息来汇总文章,节省约660倍的时间。
Clinical Cohort Studies (CCS), such as randomized clinical trials, are a great source of documented clinical research. Ideally, a clinical expert inspects these articles for exploratory analysis ranging from drug discovery for evaluating the efficacy of existing drugs in tackling emerging diseases to the first test of newly developed drugs. However, more than 100 articles are published daily on a single prevalent disease like COVID-19 in PubMed. As a result, it can take days for a physician to find articles and extract relevant information. Can we develop a system to sift through these articles faster and document the crucial takeaways from each of these articles? In this work, we propose CCS Explorer, an end-to-end system for relevance prediction of sentences, extractive summarization, and patient, outcome, and intervention entity detection from CCS. CCS Explorer is packaged in a web-based graphical user interface where the user can provide any disease name. CCS Explorer then extracts and aggregates all relevant information from articles on PubMed based on the results of an automatically generated query produced on the back-end. For each task, CCS Explorer fine-tunes pre-trained language representation models based on transformers with additional layers. The models are evaluated using two publicly available datasets. CCS Explorer obtains a recall of 80.2%, AUC-ROC of 0.843, and an accuracy of 88.3% on sentence relevance prediction using BioBERT and achieves an average Micro F1-Score of 77.8% on Patient, Intervention, Outcome detection (PIO) using PubMedBERT. Thus, CCS Explorer can reliably extract relevant information to summarize articles, saving time by ~660×.