Using Generative Artificial Intelligence to Classify Primary Progressive Aphasia from Connected Speech.

Using Generative Artificial Intelligence to Classify Primary Progressive Aphasia from Connected Speech.
复制标题

使用生成人工智能对互联言语中的原发性进行性失语症进行分类。

DOI:
10.1101/2023.12.22.23300470
复制
发表时间:
2023
期刊:
medRxiv : the preprint server for health sciences
影响因子:
--
通讯作者:
Wolff,Phillip
Wolff,Phillip
中科院分区:
--
文献类型:
--
作者:
Rezaii,Neguine;Quimby,Megan;Wong,Bonnie;Hochberg,Daisy;Brickhouse,Michael;Touroutoglou,Alexandra;Dickerson,BradfordC;Wolff,Phillip

文献摘要

相似文献

神经退行性痴呆综合征,如原发性进行性失语症(PPA),传统上部分基于语言和非语言认知特征进行诊断。关于PPA是否最好细分为三个变体以及关于PPA变体分类的最独特的语言特征的争论仍在继续。在这项研究中,我们利用人工智能(AI)和自然语言处理(NLP)的能力,首先对78名PPA患者的简洁、连接的语音样本进行无监督分类。大型语言模型识别出三个不同的PPA集群,与独立的临床诊断有88.5%的一致性。三个数据驱动聚类的皮质萎缩模式与临床诊断标准中的定位相对应。然后,我们使用NLP来识别最能分离这三种PPA变体的语言特征。17个特征对于这一目的是最有价值的,包括观察到将动词分为高频和低频类型显着提高分类准确性。使用这些来自简短连接语音样本分析的语言特征,我们开发了一个分类器,在预测PPA亚型和健康对照方面达到了97.9%的准确率。我们的研究结果为完善早期痴呆诊断提供了关键见解,加深了我们对这些神经退行性表型和语言处理神经生物学特征的理解,并提高了诊断评估的准确性。
Neurodegenerative dementia syndromes, such as Primary Progressive Aphasias (PPA), have traditionally been diagnosed based in part on verbal and nonverbal cognitive profiles. Debate continues about whether PPA is best subdivided into three variants and also regarding the most distinctive linguistic features for classifying PPA variants. In this study, we harnessed the capabilities of artificial intelligence (AI) and natural language processing (NLP) to first perform unsupervised classification of concise, connected speech samples from 78 PPA patients. Large Language Models discerned three distinct PPA clusters, with 88.5% agreement with independent clinical diagnoses. Patterns of cortical atrophy of three data-driven clusters corresponded to the localization in the clinical diagnostic criteria. We then used NLP to identify linguistic features that best dissociate the three PPA variants. Seventeen features emerged as most valuable for this purpose, including the observation that separating verbs into high and low-frequency types significantly improves classification accuracy. Using these linguistic features derived from the analysis of brief connected speech samples, we developed a classifier that achieved 97.9% accuracy in predicting PPA subtypes and healthy controls. Our findings provide pivotal insights for refining early-stage dementia diagnosis, deepening our understanding of the characteristics of these neurodegenerative phenotypes and the neurobiology of language processing, and enhancing diagnostic evaluation accuracy.