Using weak supervision and deep learning to classify clinical notes for identification of current suicidal ideation.

Using weak supervision and deep learning to classify clinical notes for identification of current suicidal ideation.
复制标题

DOI:
10.1016/j.jpsychires.2021.01.052
复制
发表时间:
2021-04
影响因子:
4.8
通讯作者:
Pathak J
Pathak J
中科院分区:
医学2区
文献类型:
--
作者:
Cusick M;Adekkanattu P;Campion TR Jr;Sholle ET;Myers A;Banerjee S;Alexopoulos G;Wang Y;Pathak J

文献摘要

参考文献

被引文献

相似文献

心理健康问题,如自杀想法,经常被提供者记录在临床笔记中,而不是结构化的编码数据。在这项研究中,我们评估了弱监督的方法检测“当前”的自杀意念,从非结构化的临床笔记在电子健康记录(EHR)系统。弱监督机器学习方法利用不完美的标签进行训练,减轻了创建大型手动注释数据集的负担。在确定了600名有自杀意念风险的患者队列后,我们使用基于规则的自然语言处理方法(NLP)来标记训练和验证注释(n= 17,978)。使用这个大型的临床笔记语料库,我们训练了几个统计机器学习模型-逻辑分类器,支持向量机(SVM),朴素贝叶斯分类器-和一个深度学习模型,即文本分类卷积神经网络(CNN),在手动审查的测试集(n=837)上进行评估。CNN模型的表现优于所有其他方法,在具有“当前”自杀意念的文档上实现了94%的总体准确率和0.82的F1分数。该算法正确识别了另外42次遭遇和9例指示自杀意念但缺少结构化诊断代码的患者。当应用于5,000份临床笔记的随机子集时,该算法将0.46%(n=23)归类为“当前”自杀意念,其中87%通过人工审查确实具有指示性。实施这种方法进行大规模的文件筛选可能会发挥重要作用,有针对性的自杀预防干预措施的床旁临床信息系统,并改善研究的途径,从构思尝试。
Mental health concerns, such as suicidal thoughts, are frequently documented by providers in clinical notes, as opposed to structured coded data. In this study, we evaluated weakly supervised methods for detecting “current” suicidal ideation from unstructured clinical notes in electronic health record (EHR) systems. Weakly supervised machine learning methods leverage imperfect labels for training, alleviating the burden of creating a large manually annotated dataset. After identifying a cohort of 600 patients at risk for suicidal ideation, we used a rule-based natural language processing approach (NLP) approach to label the training and validation notes (n=17,978). Using this large corpus of clinical notes, we trained several statistical machine learning models—logistic classifier, support vector machines (SVM), Naive Bayes classifier—and one deep learning model, namely a text classification convolutional neural network (CNN), to be evaluated on a manually-reviewed test set (n=837). The CNN model outperformed all other methods, achieving an overall accuracy of 94% and a F1-score of 0.82 on documents with “current” suicidal ideation. This algorithm correctly identified an additional 42 encounters and 9 patients indicative of suicidal ideation but missing a structured diagnosis code. When applied to a random subset of 5,000 clinical notes, the algorithm classified 0.46% (n=23) for “current” suicidal ideation, of which 87% were truly indicative via manual review. Implementation of this approach for large-scale document screening may play an important role in point-of-care clinical information systems for targeted suicide prevention interventions and improve research on the pathways from ideation to attempt.
DOI: 10.1371/journal.pone.0085733
发表时间: 2014
期刊: PloS one
影响因子: 3.7
作者:
Poulin C;Shiner B;Thompson P;Vepstas L;Young-Xu Y;Goertzel B;Watts B;Flashman L;McAllister T
通讯作者: McAllister T
DOI: 10.1038/s41598-018-25773-2
发表时间: 2018-05-09
期刊: Scientific reports
影响因子: 4.6
作者:
Fernandes AC;Dutta R;Velupillai S;Sanyal J;Stewart R;Chandran D
通讯作者: Chandran D
DOI: 10.1037//0022-006x.68.3.371
发表时间: 2000-06-01
影响因子: 5.9
作者:
Brown, GK;Beck, AT;Grisham, JR
通讯作者: Grisham, JR
DOI: 10.1007/s11606-014-2767-3
发表时间: 2014-06-01
影响因子: 5.7
作者:
Ahmedani, Brian K.;Simon, Gregory E.;Solberg, Leif I.
通讯作者: Solberg, Leif I.
DOI: 10.1111/sltb.12657
发表时间: 2020-12
影响因子: 3.2
作者:
Brown LA;Boudreaux ED;Arias SA;Miller IW;May AM;Camargo CA Jr;Bryan CJ;Armey MF
通讯作者: Armey MF