How Well Do Artificial Intelligence Chatbots Respond to the Top Search Queries About Urological Malignancies?

How Well Do Artificial Intelligence Chatbots Respond to the Top Search Queries About Urological Malignancies?
复制标题

DOI:
10.1016/j.eururo.2023.07.004
复制
发表时间:
2023-12-15
期刊:
影响因子:
23.4
通讯作者:
Kabarriti, Abdo E.
Kabarriti, Abdo E.
中科院分区:
医学1区
文献类型:
--
作者:
Musheyev, David;Pan, Alexander;Kabarriti, Abdo E.

文献摘要

被引文献

相似文献

人工智能(AI)聊天机器人正在成为一种流行的信息来源,但它们提供的关于泌尿系统恶性肿瘤信息质量的数据有限。我们的目标是从四个AI聊天机器人ChatGPT、Perplexity、Chat Sonic和Microsoft Bing AI中表征信息质量并检测有关前列腺癌、膀胱癌、肾癌和睾丸癌的错误信息。我们使用了2021年1月至2023年1月期间谷歌趋势中与前列腺癌、膀胱癌、肾癌和睾丸癌相关的前五个搜索查询,并将它们输入到AI聊天机器人中。使用已公布的工具对回答的质量、可理解性、可操作性、错误信息和可读性进行评估。AI聊天机器人的回答具有中等到高的信息质量(中位数辨别分数4分,范围2-5),并且没有错误信息。可理解性中等(可打印材料患者教育材料评估工具[PEMAT-P]可理解性中位数66.7%,范围44.4-90.9%),可操作性中等至差(PEMAT-P可操作性中位数40%,范围0-40%)。答案是以相当难的阅读水平写成的。人工智能聊天机器人生成的信息通常是准确的,质量中等到高,以响应与泌尿系统恶性肿瘤相关的顶级搜索查询,但响应缺乏明确的、可操作的说明,并超过了为消费者健康信息推荐的阅读水平。患者简介:人工智能聊天机器人产生的信息通常是准确的,质量中等,以响应流行的谷歌搜索有关泌尿系癌症的信息。然而,他们的回答相当难读,也相当难理解,而且缺乏明确的用户操作说明。(C)2023年欧洲泌尿外科协会。爱思唯尔出版,版权所有。
Artificial intelligence (AI) chatbots are becoming a popular source of information but there are limited data on the quality of information on urological malignancies that they provide. Our objective was to characterize the quality of information and detect misinformation about prostate, bladder, kidney, and testicular cancers from four AI chatbots: ChatGPT, Perplexity, Chat Sonic, and Microsoft Bing AI. We used the top five search queries related to prostate, bladder, kidney, and testicular cancers according to Google Trends from January 2021 to January 2023 and input them into the AI chatbots. Responses were evaluated for quality, understandability, actionability, misinformation, and readability using published instruments. AI chatbot responses had moderate to high information quality (median DISCERN score 4 out of 5, range 2-5) and lacked misinformation. Understandability was moderate (median Patient Education Material Assessment Tool for Printable Materials [PEMAT-P] understandability 66.7%, range 44.4-90.9%) and actionability was moderate to poor (median PEMAT-P actionability 40%, range 0-40%). The responses were written at a fairly difficult reading level. AI chat-bots produce information that is generally accurate and of moderate to high quality in response to the top urological malignancy-related search queries, but the responses lack clear, actionable instructions and exceed the reading level recommended for consumer health information. Patient summary: Artificial intelligence chatbots produce information that is generally accurate and of moderately high quality in response to popular Google searches about urological cancers. However, their responses are fairly difficult to read, are moderately hard to understand, and lack clear instructions for users to act on.(c) 2023 European Association of Urology. Published by Elsevier B.V. All rights reserved.