ClinicalRisk: A New Therapy-related Clinical Trial Dataset for Predicting Trial Status and Failure Reasons.

ClinicalRisk: A New Therapy-related Clinical Trial Dataset for Predicting Trial Status and Failure Reasons.
复制标题

ClinicalRisk:用于预测试验状态和失败原因的新治疗相关临床试验数据集。

DOI:
10.1145/3583780.3615113
复制
发表时间:
2023
期刊:
Proceedings of the ... ACM International Conference on Information & Knowledge Management. ACM International Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
Ma,Fenglong
Ma,Fenglong
中科院分区:
--
文献类型:
--
作者:
Luo,Junyu;Qiao,Zhi;Glass,Lucas;Xiao,Cao;Ma,Fenglong

文献摘要

相似文献

临床试验旨在研究新的测试并评估其对人类健康结果的影响,这是一个巨大的市场规模。然而,进行临床试验既昂贵又耗时,而且往往以无结果告终。如果我们能够开发一个有效的模型来自动评估临床试验的状态并找出可能的失败原因,这将给临床实践带来革命性的变化。然而,由于缺乏基准数据集,开发这样一个模型是具有挑战性的。为了应对这些挑战,在本文中,我们首先通过从ClinicalTrials.gov中提取公开可用的临床试验报告来构建新的数据集。每个报告的关联状态被视为状态标签。为了分析失败原因,领域专家帮助我们根据与其相关的描述来手动注释每个失败的报告。更重要的是,我们检查了这项任务的几个最新的文本分类基线,并发现临床试验方案的独特格式在影响预测准确性方面起着至关重要的作用,证明了特殊设计的临床试验分类模型的必要性。
Clinical trials aim to study new tests and evaluate their effects on human health outcomes, which has a huge market size. However, carrying out clinical trials is expensive and time-consuming and often ends in no results. It will revolutionize clinical practice if we can develop an effective model to automatically estimate the status of a clinical trial and find out possible failure reasons. However, it is challenging to develop such a model because of the lack of a benchmark dataset. To address these challenges, in this paper, we first build a new dataset by extracting the publicly available clinical trial reports from ClinicalTrials.gov. The associated status of each report is treated as the status label. To analyze the failure reasons, domain experts help us manually annotate each failed report based on the description associated with it. More importantly, we examine several state-of-the-art text classification baselines on this task and find out that the unique format of the clinical trial protocols plays an essential role in affecting prediction accuracy, demonstrating the need for specially designed clinical trial classification models.