Family member information extraction via neural sequence labeling models with different tag schemes

Family member information extraction via neural sequence labeling models with different tag schemes
复制标题

DOI:
10.1186/s12911-019-0996-4
复制
发表时间:
2019-12-27
影响因子:
3.5
通讯作者:
Dai, Hong-Jie
Dai, Hong-Jie
中科院分区:
医学3区
文献类型:
--
作者:
Dai, Hong-Jie

文献摘要

被引文献

相似文献

背景:在非结构化电子病历(EHRs)中描述的家族史信息(FHI)是患者护理和科学研究的宝贵信息来源。由于FHI通常以自由文本的形式进行描述,因此整个FHI提取过程包括切片分割、家族成员和临床观察的提取以及发现提取的成员与其观察之间的关系等多个步骤。提取步骤包括识别FHI概念及其属性,如家庭成员概念的家庭方面属性。方法:本研究重点关注提取步骤,并将其表述为序列标记问题。我们采用神经序列标记模型和不同的标记方案来区分家庭成员和他们的观察结果。针对不同的标记方案,对识别出来的实体进行聚合和处理,以确定所需的属性。结果:我们通过评估标签方案在BioCreative/OHNLP挑战赛2018发布的数据集上的性能,研究了标签方案中编码所需属性的有效性。观察发现,所提出的侧方案结合所开发的特征和神经网络架构,在测试集上的整体f1得分为0.849,在FHI实体识别子任务中排名第二。结论:通过与条件随机场模型的性能比较,所开发的基于神经网络的模型的性能明显优于条件随机场模型。然而,我们的误差分析揭示了当前方法的两个具有挑战性的问题。一是一些属性需要跨句推理。另一个是,目前的模型无法区分描述患者家庭成员的叙述和指定患者家庭成员的亲属的叙述。
Background: Family history information (FHI) described in unstructured electronic health records (EHRs) is a valuable information source for patient care and scientific researches. Since FHI is usually described in the format of free text, the entire process of FHI extraction consists of various steps including section segmentation, family member and clinical observation extraction, and relation discovery between the extracted members and their observations. The extraction step involves the recognition of FHI concepts along with their properties such as the family side attribute of the family member concept.Methods: This study focuses on the extraction step and formulates it as a sequence labeling problem. We employed a neural sequence labeling model along with different tag schemes to distinguish family members and their observations. Corresponding to different tag schemes, the identified entities were aggregated and processed by different algorithms to determine the required properties.Results: We studied the effectiveness of encoding required properties in the tag schemes by evaluating their performance on the dataset released by the BioCreative/OHNLP challenge 2018. It was observed that the proposed side scheme along with the developed features and neural network architecture can achieve an overall F1-score of 0.849 on the test set, which ranked second in the FHI entity recognition subtask.Conclusions: By comparing with the performance of conditional random fields models, the developed neural network-based models performed significantly better. However, our error analysis revealed two challenging issues of the current approach. One is that some properties required cross-sentence inferences. The other is that the current model is not able to distinguish between the narratives describing the family members of the patient and those specifying the relatives of the patient's family members.