Criteria2Query: a natural language interface to clinical databases for cohort definition

Criteria2Query: a natural language interface to clinical databases for cohort definition
复制标题

DOI:
10.1093/jamia/ocy178
复制
发表时间:
2019-04-01
影响因子:
6.4
通讯作者:
Weng, Chunhua
Weng, Chunhua
中科院分区:
管理学2区
文献类型:
--
作者:
Yuan, Chi;Ryan, Patrick B.;Weng, Chunhua

文献摘要

被引文献

相似文献

客观队列定义是进行临床研究的瓶颈,取决于领域专家的主观决策。数据驱动的队列定义很有吸引力,但需要大量的术语和临床数据模型知识。 Criteria2Query 是一种自然语言界面,可促进使用临床数据库进行队列定义和执行的人机协作。材料和方法 Criteria2Query 使用混合信息提取管道,结合机器学习和基于规则的方法来系统地解析资格标准文本,首先将其转换为结构化标准表示,然后转换为可共享且可执行的临床数据查询,表示为符合 OMOP 通用数据模型的 SQL 查询。用户可以在 ATLAS Web 应用程序中交互式地查看、优化和执行查询。为了测试有效性,我们评估了 ClinicalTrials.gov 中不同疾病领域的 125 项标准和 52 项用户输入的标准。我们针对 2 位领域专家评估了 F1 分数和准确性,并计算了全自动查询制定的平均计算时间。我们进行了一项评估可用性的匿名调查。结果 Criteria2Query 在实体识别和关系提取方面分别获得了 0.795 和 0.805 F1 分数。否定检测、逻辑检测、实体标准化和属性标准化的准确度分别为 0.984、0.864、0.514 和 0.793。全自动查询制定需要 1.22 秒/标准。超过 80%(13 人中的 11 人以上)的用户将在未来的队列定义任务中使用 Criteria2Query。 结论 我们为临床数据库提供了一种新颖的自然语言界面。它是开源的,支持完全自动化和交互模式,由研究人员以最少的人力进行自主数据驱动的队列定义。我们展示了其有前途的用户友好性和可用性。
Objective Cohort definition is a bottleneck for conducting clinical research and depends on subjective decisions by domain experts. Data-driven cohort definition is appealing but requires substantial knowledge of terminologies and clinical data models. Criteria2Query is a natural language interface that facilitates human-computer collaboration for cohort definition and execution using clinical databases.Materials and Methods Criteria2Query uses a hybrid information extraction pipeline combining machine learning and rule-based methods to systematically parse eligibility criteria text, transforms it first into a structured criteria representation and next into sharable and executable clinical data queries represented as SQL queries conforming to the OMOP Common Data Model. Users can interactively review, refine, and execute queries in the ATLAS web application. To test effectiveness, we evaluated 125 criteria across different disease domains from ClinicalTrials.gov and 52 user-entered criteria. We evaluated F1 score and accuracy against 2 domain experts and calculated the average computation time for fully automated query formulation. We conducted an anonymous survey evaluating usability.Results Criteria2Query achieved 0.795 and 0.805 F1 score for entity recognition and relation extraction, respectively. Accuracies for negation detection, logic detection, entity normalization, and attribute normalization were 0.984, 0.864, 0.514 and 0.793, respectively. Fully automatic query formulation took 1.22 seconds/criterion. More than 80% (11+ of 13) of users would use Criteria2Query in their future cohort definition tasks.Conclusions We contribute a novel natural language interface to clinical databases. It is open source and supports fully automated and interactive modes for autonomous data-driven cohort definition by researchers with minimal human effort. We demonstrate its promising user friendliness and usability.