Clustering clinical trials with similar eligibility criteria features.

Clustering clinical trials with similar eligibility criteria features.
复制标题

DOI:
10.1016/j.jbi.2014.01.009
复制
发表时间:
2014-12
影响因子:
4.5
通讯作者:
Weng, Chunhua
Weng, Chunhua
中科院分区:
医学3区
文献类型:
--
作者:
Hao, Tianyong;Rusanov, Alexander;Boland, Mary Regina;Weng, Chunhua

文献摘要

参考文献

被引文献

相似文献

自动识别和聚类具有相似资格特征的临床试验。使用公共存储库ClinicalTrials.gov作为数据源,我们从所有临床试验的合格标准文本中提取语义特征,并构建了一个试验特征矩阵。我们根据其合格性特征计算了所有临床试验的成对相似性。对于所有试验,通过每次选择一个试验作为中心,我们确定了与中心试验的相似性大于或等于预定义阈值的试验,并构建了基于中心的聚类。然后,我们确定了独特的审判集与独特的审判成员组成中心为基础的集群忽视其结构信息。从ClinicalTrials.gov上的145,745个临床试验中,我们提取了5,508,491个语义特征。其中,459,936个是唯一的,160,951个至少由一对试验共享。使用Amazon Mechanical Turk(MTurk)众包聚类评估,我们确定了最佳相似度阈值0.9。使用这个阈值,我们生成了8,806个基于中心的聚类。MTurk对聚类样本的评价结果为平均得分4.331±0.796(1-5分)(5分表示“强烈同意聚类中的试验相似”)。我们提供了一种自动化的方法来聚类具有相似资格特征的临床试验。这种方法可以是潜在的有用的研究知识重用模式在临床试验合格标准的设计和改善临床试验招募。我们还提供了一个有效的众包方法来评估信息干预。
To automatically identify and cluster clinical trials with similar eligibility features. Using the public repository ClinicalTrials.gov as the data source, we extracted semantic features from the eligibility criteria text of all clinical trials and constructed a trial-feature matrix. We calculated the pairwise similarities for all clinical trials based on their eligibility features. For all trials, by selecting one trial as the center each time, we identified trials whose similarities to the central trial were greater than or equal to a predefined threshold and constructed center-based clusters. Then we identified unique trial sets with distinctive trial membership compositions from center-based clusters by disregarding their structural information. From the 145,745 clinical trials on ClinicalTrials.gov, we extracted 5,508,491 semantic features. Of these, 459,936 were unique and 160,951 were shared by at least one pair of trials. Crowdsourcing the cluster evaluation using Amazon Mechanical Turk (MTurk), we identified the optimal similarity threshold, 0.9. Using this threshold, we generated 8,806 center-based clusters. Evaluation of a sample of the clusters by MTurk resulted in a mean score 4.331±0.796 on a scale of 1–5 (5 indicating “strongly agree that the trials in the cluster are similar”). We contribute an automated approach to clustering clinical trials with similar eligibility features. This approach can be potentially useful for investigating knowledge reuse patterns in clinical trial eligibility criteria designs and for improving clinical trial recruitment. We also contribute an effective crowdsourcing method for evaluating informatics interventions.
DOI: 10.1016/j.ins.2003.03.011
发表时间: 2004-02-15
影响因子: 8.1
作者:
Hirano, S;Sun, XG;Tsumoto, S
通讯作者: Tsumoto, S
DOI: 10.1093/rheumatology/kes261
发表时间: 2013-02-01
期刊: RHEUMATOLOGY
影响因子: 5.5
作者:
Li, Philip Hei;Wong, Wilfred Hing Sang;Lau, Yu-Lung
通讯作者: Lau, Yu-Lung
DOI: 10.1016/j.jbi.2006.06.004
发表时间: 2007-06-01
影响因子: 4.5
作者:
Pedersen, Ted;Pakhomov, Serguei V. S.;Chute, Christopher G.
通讯作者: Chute, Christopher G.
DOI: 10.3758/s13428-011-0124-6
发表时间: 2012-03-01
影响因子: 5.4
作者:
Mason, Winter;Suri, Siddharth
通讯作者: Suri, Siddharth
DOI: 10.1186/1472-6947-12-s1-s3
发表时间: 2012-04-30
影响因子: 3.5
作者:
Korkontzelos I;Mu T;Ananiadou S
通讯作者: Ananiadou S