Classifying unstructured electronic consult messages to understand primary care physician specialty information needs.

Classifying unstructured electronic consult messages to understand primary care physician specialty information needs.
复制标题

对非结构化电子咨询消息进行分类,以了解初级保健医生的专业信息需求。

DOI:
10.1093/jamia/ocac092
复制
发表时间:
2022
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
通讯作者:
Miller,TimothyA
Miller,TimothyA
中科院分区:
--
文献类型:
--
作者:
Ding,Xiyu;Barnett,Michael;Mehrotra,Ateev;Tuot,DelphineS;Bitterman,DanielleS;Miller,TimothyA

文献摘要

相似文献

目的电子会诊(EConsulting)内容反映了组织内有关推荐临床医生需求的重要信息,但提取起来具有挑战性。这项工作的目标是开发机器学习模型,用于根据问题类型和问题内容对EConsult问题进行分类。这项工作的另一个目标是调查在有限的专家时间资源下解决这项任务的能力。材料和方法我们的数据来源是旧金山健康网络咨询系统,从2008-2017年间有超过70万个已识别的问题,来自胃肠病学、泌尿学和神经学专业。我们基于Transformers的双向编码器表示开发分类器,试验多任务学习,以了解何时可以在分类器之间共享信息。我们绘制学习曲线,以了解何时我们可能能够减少所需的人类标记量。结果多任务学习仅在神经病学-泌尿科对中显示出好处,因为它们在问题类型的分布上有很大的相似之处。继续对新领域的模型进行预培训是非常有效的。在神经病学-泌尿学对中,仅用10%的泌尿学训练数据获得接近峰值的性能,给出所有神经病学数据。讨论跨分类器类型共享信息没有什么好处,而跨专业共享分类器组件可能会有所帮助,如果它们在程序性和认知性患者护理的平衡方面相似。结论我们可以使用足够的标记数据准确地对EConsult内容进行分类,但只有在特殊情况下才适用减少标记工作量的方法。未来的工作应该探索新的学习范式,以进一步减少标记工作。
ObjectiveElectronic consultation (eConsult) content reflects important information about referring clinician needs across an organization, but is challenging to extract. The objective of this work was to develop machine learning models for classifying eConsult questions for question type and question content. Another objective of this work was to investigate the ability to solve this task with constrained expert time resources.Materials and MethodsOur data source is the San Francisco Health Network eConsult system, with over 700 000 deidentified questions from the years 2008–2017, from gastroenterology, urology, and neurology specialties. We develop classifiers based on Bidirectional Encoder Representations from Transformers, experimenting with multitask learning to learn when information can be shared across classifiers. We produce learning curves to understand when we may be able to reduce the amount of human labeling required.ResultsMultitask learning shows benefits only in the neurology–urology pair where they shared substantial similarities in the distribution of question types. Continued pretraining of models in new domains is highly effective. In the neurology–urology pair, near-peak performance is achieved with only 10% of the urology training data given all of the neurology data.DiscussionSharing information across classifier types shows little benefit, whereas sharing classifier components across specialties can help if they are similar in the balance of procedural versus cognitive patient care.ConclusionWe can accurately classify eConsult content with enough labeled data, but only in special cases do methods for reducing labeling effort apply. Future work should explore new learning paradigms to further reduce labeling effort.