Chia, a large annotated corpus of clinical trial eligibility criteria

Chia, a large annotated corpus of clinical trial eligibility criteria
复制标题

DOI:
10.1038/s41597-020-00620-0
复制
发表时间:
2020-08-27
期刊:
影响因子:
9.8
通讯作者:
Weng, Chunhua
Weng, Chunhua
中科院分区:
综合性期刊2区
文献类型:
--
作者:
Kury, Fabricio;Butler, Alex;Weng, Chunhua

文献摘要

被引文献

相似文献

我们推出了 Chia,这是一个新型的、带有注释的大型患者资格标准语料库,从在 ClinicalTrials.gov 注册的 1,000 项介入性 IV 期临床试验中提取。该数据集包括 12,409 个带注释的资格标准,由 15 种实体类型的 41,487 个独特实体和 12 种关系类型的 25,017 个关系表示。每个条件都表示为有向无环图,可以轻松地将其转换为布尔逻辑以形成数据库查询。 Chia 可以作为共享基准来开发和测试未来的机器学习、基于规则或混合方法,以从自由文本临床试验资格标准中提取信息。
We present Chia, a novel, large annotated corpus of patient eligibility criteria extracted from 1,000 interventional, Phase IV clinical trials registered in ClinicalTrials.gov. This dataset includes 12,409 annotated eligibility criteria, represented by 41,487 distinctive entities of 15 entity types and 25,017 relationships of 12 relationship types. Each criterion is represented as a directed acyclic graph, which can be easily transformed into Boolean logic to form a database query. Chia can serve as a shared benchmark to develop and test future machine learning, rule-based, or hybrid methods for information extraction from free-text clinical trial eligibility criteria.