Hierarchical Annotation for Building A Suite of Clinical Natural Language Processing Tasks: Progress Note Understanding
Hierarchical Annotation for Building A Suite of Clinical Natural Language Processing Tasks: Progress Note Understanding
复制标题
DOI:
10.48550/arxiv.2204.03035
复制
发表时间:
2022-04
期刊:
影响因子:
--
通讯作者:
Yanjun Gao;Dmitriy Dligach;Timothy Miller;S. Tesch;Ryan Laffin;M. Churpek;M. Afshar
中科院分区:
文献类型:
--
作者:
Yanjun Gao;Dmitriy Dligach;Timothy Miller;S. Tesch;Ryan Laffin;M. Churpek;M. Afshar
Applying methods in natural language processing on electronic health records (EHR) data has attracted rising interests. Existing corpus and annotation focus on modeling textual features and relation prediction. However, there are a paucity of annotated corpus built to model clinical diagnostic thinking, a processing involving text understanding, domain knowledge abstraction and reasoning. In this work, we introduce a hierarchical annotation schema with three stages to address clinical text understanding, clinical reasoning and summarization. We create an annotated corpus based on a large collection of publicly available daily progress notes, a type of EHR that is time-sensitive, problem-oriented, and well-documented by the format of Subjective, Objective, Assessment and Plan (SOAP). We also define a new suite of tasks, Progress Note Understanding, with three tasks utilizing the three annotation stages. This new suite aims at training and evaluating future NLP models for clinical text understanding, clinical knowledge representation, inference and summarization.