Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum

Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum
复制标题

DOI:
10.1001/jamainternmed.2023.1838
复制
发表时间:
2023-04-28
影响因子:
39
通讯作者:
Smith, Davey M.
Smith, Davey M.
中科院分区:
医学1区
文献类型:
--
作者:
Ayers, John W.;Poliak, Adam;Smith, Davey M.

文献摘要

被引文献

相似文献

重要性虚拟医疗保健的快速扩张导致了患者信息的激增,同时也伴随着医疗保健专业人员的更多工作和倦怠。人工智能(AI)助手可以通过起草可以由临床医生审查的回答来帮助创建对患者问题的答案。目的评估2022年11月发布的AI聊天机器人助手(ChatGPT)对患者问题提供高质量和同情的回答的能力。设计,设置和参与者在这项横断面研究中,来自公共社交媒体论坛(Reddit的r/Askeleton)的公共和不可识别的问题数据库被用于从2022年10月随机抽取195个交流,其中一名经过验证的医生回答了一个公共问题。聊天机器人的回答是通过在2022年12月22日和23日将原始问题输入到新的会话中(在会话中没有先前的问题)来生成的。最初的问题沿着匿名和随机排序的医生和聊天机器人的回答由一组有执照的医疗保健专业人员一式三份进行评估。评估者选择“哪种反应更好”,并判断“提供的信息质量”(非常差、差、可接受、好或非常好)和“提供的同理心或床边态度”(不同情、轻微同情、中度同情、同情和非常同情)。结果在195个问题和回答中,在585个评价中,评价者更喜欢聊天机器人回答而不是医生回答,占78.6%(95%CI,75.0%-81.8%)。平均(IQR)医生响应明显短于聊天机器人响应(52 [17-62]字vs 211 [168-245]字; t = 25.4; P
IMPORTANCE The rapid expansion of virtual health care has caused a surge in patient messages concomitant with more work and burnout among health care professionals. Artificial intelligence (AI) assistants could potentially aid in creating answers to patient questions by drafting responses that could be reviewed by clinicians.OBJECTIVE To evaluate the ability of an AI chatbot assistant (ChatGPT), released in November 2022, to provide quality and empathetic responses to patient questions.DESIGN, SETTING, AND PARTICIPANTS In this cross-sectional study, a public and nonidentifiable database of questions from a public social media forum (Reddit's r/AskDocs) was used to randomly draw 195 exchanges from October 2022 where a verified physician responded to a public question. Chatbot responses were generated by entering the original question into a fresh session (without prior questions having been asked in the session) on December 22 and 23, 2022. The original question along with anonymized and randomly ordered physician and chatbot responses were evaluated in triplicate by a team of licensed health care professionals. Evaluators chose "which response was better" and judged both "the quality of information provided" (very poor, poor, acceptable, good, or very good) and "the empathy or bedside manner provided" (not empathetic, slightly empathetic, moderately empathetic, empathetic, and very empathetic). Mean outcomes were ordered on a 1 to 5 scale and compared between chatbot and physicians.RESULTS Of the 195 questions and responses, evaluators preferred chatbot responses to physician responses in 78.6%(95% CI, 75.0%-81.8%) of the 585 evaluations. Mean (IQR) physician responses were significantly shorter than chatbot responses (52 [17-62] words vs 211 [168-245] words; t = 25.4; P