Automatically annotating topics in transcripts of patient-provider interactions via machine learning.

Automatically annotating topics in transcripts of patient-provider interactions via machine learning.
复制标题

DOI:
10.1177/0272989x13514777
复制
发表时间:
2014-05
期刊:
Medical decision making : an international journal of the Society for Medical Decision Making
影响因子:
--
通讯作者:
Trikalinos TA
Trikalinos TA
中科院分区:
其他
文献类型:
--
作者:
Wallace BC;Laws MB;Small K;Wilson IB;Trikalinos TA

文献摘要

相似文献

带注释的患者-提供者接触可以为临床沟通提供重要的见解,最终建议如何改进以实现更好的健康结果。但是,用Roter或通用医疗交互分析系统(GMIAS)代码注释门诊记录是昂贵的,限制了此类分析的范围。我们建议通过机器学习自动注释患者与主题代码的交互记录。我们使用一个条件随机场(CRF)模型的话语主题概率。该模型考虑了会话的顺序结构和话语的组成。我们通过对360次门诊就诊(超过230,000次话语)的GMIAS注释的成绩单进行10倍交叉验证来评估预测性能。然后,我们使用自动化代替手动注释,重新分析了一项随机试验的116次额外访问,该试验使用GMIAS评估旨在改善抗逆转录病毒(ARV)依从性的干预措施的有效性。对于6个主题代码,CRF与人类注释者相比的平均成对Kappa值为0.49(范围:0.47,0.53),平均总体准确度为0.64(范围:0.62,0.66)。关于RCT再分析,使用自动注释的结果与使用手动注释获得的结果一致。根据手动注释,没有干预和有干预的抗逆转录病毒相关话语的中位数分别为49.5和76(配对符号检验p=0.07)。使用自动注释,相应的数字为39与55(p=0.04)。虽然适度准确,但预测的注释远非完美。会话主题是中间结果;它们的效用仍在研究中。这种对自动主题推理的尝试表明,机器学习方法可以将包括患者-提供者交互的话语以合理的准确性分类为临床相关主题。
Annotated patient-provider encounters can provide important insights into clinical communication, ultimately suggesting how it might be improved to effect better health outcomes. But annotating outpatient transcripts with Roter or General Medical Interaction Analysis System (GMIAS) codes is expensive, limiting the scope of such analyses. We propose automatically annotating transcripts of patient-provider interactions with topic codes via machine learning. We use a conditional random field (CRF) to model utterance topic probabilities. The model accounts for the sequential structure of conversations and the words comprising utterances. We assess predictive performance via 10- fold cross-validation over GMIAS-annotated transcripts of 360 outpatient visits (over 230,000 utterances). We then used automated in place of manual annotations to reproduce an analysis of 116 additional visits from a randomized trial that used GMIAS to assess the efficacy of an intervention aimed at improving communication around antiretroviral (ARV) adherence. With respect to six topic codes, the CRF achieved a mean pairwise kappa compared with human annotators of 0.49 (range: 0.47, 0.53) and a mean overall accuracy of 0.64 (range: 0.62, 0.66). With respect to the RCT re-analysis, results using automated annotations agreed with those obtained using manual ones. According to the manual annotations, the median number of ARV-related utterances without and with the intervention was 49.5 versus 76, respectively (paired sign test p=0.07). Using automated annotations, the respective numbers were 39 versus 55 (p=0.04). While moderately accurate, the predicted annotations are far from perfect. Conversational topics are intermediate outcomes; their utility is still being researched. This foray into automated topic inference suggests that machine learning methods can classify utterances comprising patient-provider interactions into clinically relevant topics with reasonable accuracy.