Feasibility of feature-based indexing, clustering, and search of clinical trials. A case study of breast cancer trials from ClinicalTrials.gov.

Feasibility of feature-based indexing, clustering, and search of clinical trials. A case study of breast cancer trials from ClinicalTrials.gov.
复制标题

DOI:
10.3414/me12-01-0092
复制
发表时间:
2013
影响因子:
1.7
通讯作者:
Weng C
Weng C
中科院分区:
医学4区
文献类型:
--
作者:
Boland MR;Miotto R;Gao J;Weng C

文献摘要

被引文献

相似文献

当标准疗法失败时,临床试验为耐药疾病或绝症患者提供实验性治疗机会。临床试验还可以为否则可能无法获得此类护理的个人提供免费治疗和教育。为了找到相关的临床试验,患者经常在网上搜索;然而,由于大量的试验和减少试验搜索空间的无效索引方法,它们经常遇到重大障碍。本研究探讨了基于特征的索引、聚类和临床试验搜索的可行性,并告知设计自动化这些过程。我们将80个随机选择的III期乳腺癌临床试验分解成一个合格特征向量,并将其组织成一个层次结构。我们根据它们的资格特征相似性对试验进行聚类。在模拟搜索过程中,使用手动选择的特征来生成特定的资格问题,以迭代过滤试验。我们提取了1437个不同的资格特征,并在20多个试验中对37个频繁出现的特征进行了特征提取,获得了0.73的评分间一致性。使用所有1437个特征,我们将80个试验分为6个组,其中包括按患者特征特征招募相似患者的试验,按疾病特征特征招募5个组,按混合特征招募2个组。大多数特征被映射到一个或多个统一医学语言系统(UMLS)概念,展示了命名实体识别在与UMLS进行映射之前用于自动特征提取的效用。开发基于特征的临床试验索引和聚类方法来识别具有相似目标人群的试验,提高试验检索效率是可行的。
When standard therapies fail, clinical trials provide experimental treatment opportunities for patients with drug-resistant illnesses or terminal diseases. Clinical Trials can also provide free treatment and education for individuals who otherwise may not have access to such care. To find relevant clinical trials, patients often search online; however, they often encounter a significant barrier due to the large number of trials and in-effective indexing methods for reducing the trial search space. This study explores the feasibility of feature-based indexing, clustering, and search of clinical trials and informs designs to automate these processes. We decomposed 80 randomly selected stage III breast cancer clinical trials into a vector of eligibility features, which were organized into a hierarchy. We clustered trials based on their eligibility feature similarities. In a simulated search process, manually selected features were used to generate specific eligibility questions to filter trials iteratively. We extracted 1,437 distinct eligibility features and achieved an inter-rater agreement of 0.73 for feature extraction for 37 frequent features occurring in more than 20 trials. Using all the 1,437 features we stratified the 80 trials into six clusters containing trials recruiting similar patients by patient-characteristic features, five clusters by disease-characteristic features, and two clusters by mixed features. Most of the features were mapped to one or more Unified Medical Language System (UMLS) concepts, demonstrating the utility of named entity recognition prior to mapping with the UMLS for automatic feature extraction. It is feasible to develop feature-based indexing and clustering methods for clinical trials to identify trials with similar target populations and to improve trial search efficiency.