A text mining approach to the prediction of disease status from clinical discharge summaries.

A text mining approach to the prediction of disease status from clinical discharge summaries.
复制标题

DOI:
10.1197/jamia.m3096
复制
发表时间:
2009-07
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
通讯作者:
Hui Yang;Irena Spasic;J. Keane;G. Nenadic
Hui Yang;Irena Spasic;J. Keane;G. Nenadic
中科院分区:
其他
文献类型:
--
作者:
Hui Yang;Irena Spasic;J. Keane;G. Nenadic

文献摘要

被引文献

相似文献

目的作者提出了一个系统开发的挑战在自然语言处理的临床数据i2 b2肥胖的挑战,其目的是自动识别肥胖的状态和15个相关的合并症的患者使用他们的临床出院摘要。挑战包括两个任务,文本和直觉。文本的任务是确定明确的参考疾病,而直观的任务集中在预测的疾病状态时,证据没有明确断言。设计:作者收集了一组资源,从词汇和语义上描述疾病及其相关症状,治疗方法等。这些特征在混合文本挖掘方法中进行了探索,该方法结合了字典查找,基于规则和机器学习方法。这些方法应用于一组507个以前看不见的出院摘要,并对手动准备的金标准进行了评估的预测。参赛队伍的整体排名主要基于宏观平均F-测量。结果:实施的方法实现了文本任务的宏观平均F-测量值为81%(这是挑战中达到的最高值),直观任务的宏观平均F-测量值为63%(在28个团队中排名第7,最高为66%)。微平均F-测量显示文本的平均准确度为97%,直观注释的平均准确度为96%。结论:所实现的性能与人类注释者之间的一致性一致,表明文本挖掘在从临床出院摘要准确有效地预测疾病状态方面的潜力。
OBJECTIVE The authors present a system developed for the Challenge in Natural Language Processing for Clinical Data-the i2b2 obesity challenge, whose aim was to automatically identify the status of obesity and 15 related co-morbidities in patients using their clinical discharge summaries. The challenge consisted of two tasks, textual and intuitive. The textual task was to identify explicit references to the diseases, whereas the intuitive task focused on the prediction of the disease status when the evidence was not explicitly asserted. DESIGN The authors assembled a set of resources to lexically and semantically profile the diseases and their associated symptoms, treatments, etc. These features were explored in a hybrid text mining approach, which combined dictionary look-up, rule-based, and machine-learning methods. MEASUREMENTS The methods were applied on a set of 507 previously unseen discharge summaries, and the predictions were evaluated against a manually prepared gold standard. The overall ranking of the participating teams was primarily based on the macro-averaged F-measure. RESULTS The implemented method achieved the macro-averaged F-measure of 81% for the textual task (which was the highest achieved in the challenge) and 63% for the intuitive task (ranked 7(th) out of 28 teams-the highest was 66%). The micro-averaged F-measure showed an average accuracy of 97% for textual and 96% for intuitive annotations. CONCLUSIONS The performance achieved was in line with the agreement between human annotators, indicating the potential of text mining for accurate and efficient prediction of disease statuses from clinical discharge summaries.