Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers.

Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers.
复制标题

DOI:
10.1038/s41746-023-00819-6
复制
发表时间:
2023-04-26
影响因子:
15.2
通讯作者:
Pearson, Alexander T.
Pearson, Alexander T.
中科院分区:
医学1区
文献类型:
--
作者:
Gao, Catherine A.;Howard, Frederick M.;Markov, Nikolay S.;Dyer, Emma C.;Ramesh, Siddhi;Luo, Yuan;Pearson, Alexander T.

文献摘要

参考文献

被引文献

相似文献

像ChatGPT这样的大型语言模型可以产生越来越逼真的文本,但在科学写作中使用这些模型的准确性和完整性方面的信息是未知的。我们从五种高影响因子医学期刊中收集了第五篇研究摘要,并要求ChatGPT根据它们的标题和期刊生成研究摘要。使用AI输出检测器“GPT-2输出检测器”检测到大多数生成的摘要,与原始摘要的中位数0.02% [IQR 0.02%,0.09%]相比,中位数[四分位距]的%“假”分数(更高意味着更可能生成)为99.98%[12.73%,99.98%]。AI输出检测器的AUROC为0.94。当通过剽窃检测器网站和iThenticate运行时,生成的摘要得分低于原始摘要(得分越高意味着找到更多匹配的文本)。当给出原始摘要和一般摘要的混合时,盲态人类评审员正确地识别出68%的生成摘要是由ChatGPT生成的,但错误地识别出14%的原始摘要是由ChatGPT生成的。审稿人指出,这是令人惊讶的难以区分两者,虽然他们怀疑产生的摘要是更简洁,更公式化。ChatGPT编写可信的科学摘要,尽管使用的是完全生成的数据。根据特定的指南,AI输出检测器可以作为编辑工具,帮助维护科学标准。使用大型语言模型来帮助科学写作的道德和可接受的界限仍在讨论中,不同的期刊和会议正在采取不同的政策。
Large language models such as ChatGPT can produce increasingly realistic text, with unknown information on the accuracy and integrity of using these models in scientific writing. We gathered fifth research abstracts from five high-impact factor medical journals and asked ChatGPT to generate research abstracts based on their titles and journals. Most generated abstracts were detected using an AI output detector, ‘GPT-2 Output Detector’, with % ‘fake’ scores (higher meaning more likely to be generated) of median [interquartile range] of 99.98% ‘fake’ [12.73%, 99.98%] compared with median 0.02% [IQR 0.02%, 0.09%] for the original abstracts. The AUROC of the AI output detector was 0.94. Generated abstracts scored lower than original abstracts when run through a plagiarism detector website and iThenticate (higher scores meaning more matching text found). When given a mixture of original and general abstracts, blinded human reviewers correctly identified 68% of generated abstracts as being generated by ChatGPT, but incorrectly identified 14% of original abstracts as being generated. Reviewers indicated that it was surprisingly difficult to differentiate between the two, though abstracts they suspected were generated were vaguer and more formulaic. ChatGPT writes believable scientific abstracts, though with completely generated data. Depending on publisher-specific guidelines, AI output detectors may serve as an editorial tool to help maintain scientific standards. The boundaries of ethical and acceptable use of large language models to help scientific writing are still being discussed, and different journals and conferences are adopting varying policies.
DOI: 10.1038/s41746-021-00464-x
发表时间: 2021-06-03
影响因子: 15.2
作者:
Korngiebel DM;Mooney SD
通讯作者: Mooney SD
DOI: 10.1371/journal.pdig.0000198
发表时间: 2023-02
期刊: PLOS digital health
影响因子: --
作者:
通讯作者: --
DOI: 10.3389/fpsyg.2020.513474
发表时间: 2020
影响因子: 3.8
作者:
Bishop JM
通讯作者: Bishop JM
DOI: 10.1038/s41467-021-24698-1
发表时间: 2021-07-20
影响因子: 16.6
作者:
Howard FM;Dolezal J;Kochanny S;Schulte J;Chen H;Heij L;Huo D;Nanda R;Olopade OI;Kather JN;Cipriani N;Grossman RL;Pearson AT
通讯作者: Pearson AT
DOI: 10.1038/s41592-019-0686-2
发表时间: 2020-02-03
期刊: NATURE METHODS
影响因子: 48
作者:
Virtanen, Pauli;Gommers, Ralf;van Mulbregt, Paul
通讯作者: van Mulbregt, Paul