Risk prediction using natural language processing of electronic mental health records in an inpatient forensic psychiatry setting

Risk prediction using natural language processing of electronic mental health records in an inpatient forensic psychiatry setting
复制标题

DOI:
10.1016/j.jbi.2018.08.007
复制
发表时间:
2018-10-01
影响因子:
4.5
通讯作者:
Scanlan, Joel
Scanlan, Joel
中科院分区:
医学3区
文献类型:
--
作者:
Duy Van Le;Montgomery, James;Scanlan, Joel

文献摘要

被引文献

相似文献

目的:自我和他人伤害风险评估工具广泛应用于住院司法精神病学中。一种潜在的替代或补充风险预测方法是使用自然语言处理(NLP)对电子健康记录(EHR)中的病例记录进行自动分析。这项探索性研究评估了法医EHR数据集中单词的存在或不存在以及频率,比较了四个参考词典。使用七种机器学习算法和不同时间段的EHR分析进行了探索。哪一个字典和哪一个时间段是最能预测风险评估分数的有效instruments.Materials和方法:EHR数据集包括来自塔斯马尼亚Wilfred Lopes中心的去识别法医住院笔记。数据包括非结构化自由文本病例记录条目和三个风险评估量表的系列评级:历史临床风险管理-20(HCR-20)、风险和可治疗性短期评估(START)和。情境攻击的动态评估(DASA)。四个NLP字典单词列表被选中:6865心理健康症状词从统一医学语言系统(UMLS),455 DSM-IV诊断从UMLS库,6790英语积极和消极的情绪词,和1837高频词从当代美国英语语料库(COCA)。使用七种机器学习方法Bagging,J 48,Jrip,Logistic Model Trees(LMT),Logistic Regression,Linear Regression和Support Vector Machine(SVM)来识别最佳预测风险评估分数的词典和算法的组合。结果:使用情感词典和LMT和SVM算法在DASA数据集上实现了最准确的预测。NLP与NLP字典和机器学习结合使用,根据EHR内容预测HCR-20,START和DASA的风险评级。需要进一步的研究来确定NLP方法在预测实际自我伤害、伤害他人或自我毁灭的终点方面的效用。
Objective: Instruments rating risk of harm to self and others are widely used in inpatient forensic psychiatry settings. A potential alternate or supplementary means of risk prediction is from the automated analysis of case notes in Electronic Health Records (EHRs) using Natural Language Processing (NLP). This exploratory study rated presence or absence and frequency of words in a forensic EHR dataset, comparing four reference dictionaries. Seven machine learning algorithms and different time periods of EHR analysis were used to probe. which dictionary and which time period were most predictive of risk assessment scores on validated instruments.Materials and methods: The EHR dataset comprised de-identified forensic inpatient notes from the Wilfred Lopes Centre in Tasmania. The data comprised unstructured free-text case note entries and serial ratings of three risk assessment scales: Historical Clinical Risk Management-20 (HCR-20), Short-Term Assessment of Risk and Treatability (START) and. Dynamic Appraisal of Situational Aggression (DASA). Four NLP dictionary word lists were selected: 6865 mental health symptom words from the Unified Medical Language System (UMLS), 455 DSM-IV diagnoses from UMLS repository, 6790 English positive and negative sentiment words, and 1837 high frequency words from the Corpus of Contemporary American English (COCA). Seven machine learning methods Bagging, J48, Jrip, Logistic Model Trees (LMT), Logistic Regression, Linear Regression and Support Vector Machine (SVM) were used to identify the combination of dictionaries and algorithms that best predicted risk assessment scores.Results: The most accurate prediction was attained on the DASA dataset using the sentiment dictionary and the LMT and SVM algorithms.Conclusions: NLP, used in conjunction with NLP dictionaries and machine learning, predicted risk ratings on the HCR-20, START, and DASA, based on EHR content. Further research is required to ascertain the utility of NLP approaches in predicting endpoints of actual self-harm, harm to others or victimisation.