Concept selection for phenotypes and diseases using learn to rank

Concept selection for phenotypes and diseases using learn to rank
复制标题

DOI:
10.1186/s13326-015-0019-z
复制
发表时间:
2015-06-01
影响因子:
1.9
通讯作者:
Groza, Tudor
Groza, Tudor
中科院分区:
工程技术4区
文献类型:
--
作者:
Collier, Nigel;Oellrich, Anika;Groza, Tudor

文献摘要

被引文献

相似文献

背景:表型是根据给定证据确定疾病存在的基础。然而,这些证据中的大部分仍然被锁在文本中——科学文章、临床试验报告和电子病历(EPR)——作者使用人类语言的全部表现力来报告他们的观察结果。结果:在本文中,我们利用现成工具的组合来提取机器可理解的表型表示和其他有关疾病诊断和治疗的相关概念。这些都是针对一个金标准EPR集合进行测试的,该集合已经用统一医学语言系统(UMLS)概念标识符进行了注释:ShARE/CLEF 2013语料库,用于疾病检测。我们将四个管道作为独立系统进行评估,然后尝试使用几种学习排序(LTR)方法来优化基于语义类型的性能——三种成对方法和一种列表方法。我们观察到,虽然整体Apache ctake倾向于在强召回(R = 0.57)上优于其他独立系统,但精度较低(P = 0.09),导致低至中等的F1测量(F1 = 0.16)。此外,在障碍的不同语义类型中,系统性能有很大的差异。例如,概念Findings (T033)似乎对所有系统都非常具有挑战性。LTR内的联合系统大大提高了F1 (F1 = 0.24),特别是对于疾病或综合征(T047)和解剖异常(T190)。虽然召回率显著提高,但精度仍然是一个挑战(P = 0.15, R = 0.59)。
Background: Phenotypes form the basis for determining the existence of a disease against the given evidence. Much of this evidence though remains locked away in text - scientific articles, clinical trial reports and electronic patient records (EPR) - where authors use the full expressivity of human language to report their observations.Results: In this paper we exploit a combination of off-the-shelf tools for extracting a machine understandable representation of phenotypes and other related concepts that concern the diagnosis and treatment of diseases. These are tested against a gold standard EPR collection that has been annotated with Unified Medical Language System (UMLS) concept identifiers: the ShARE/CLEF 2013 corpus for disorder detection. We evaluate four pipelines as stand-alone systems and then attempt to optimise semantic-type based performance using several learn-to-rank (LTR) approaches - three pairwise and one listwise. We observed that whilst overall Apache cTAKES tended to outperform other stand-alone systems on a strong recall (R = 0.57), precision was low (P = 0.09) leading to low-to-moderate F1 measure (F1 = 0.16). Moreover, there is substantial variation in system performance across semantic types for disorders. For example, the concept Findings (T033) seemed to be very challenging for all systems. Combining systems within LTR improved F1 substantially (F1 = 0.24) particularly for Disease or syndrome (T047) and Anatomical abnormality (T190). Whilst recall is improved markedly, precision remains a challenge (P = 0.15, R = 0.59).