Linguistic features and automatic classifiers for identifying mild cognitive impairment and dementia

Linguistic features and automatic classifiers for identifying mild cognitive impairment and dementia
复制标题

DOI:
10.1016/j.csl.2020.101113
复制
发表时间:
2021-01-01
影响因子:
4.3
通讯作者:
Tamburini, Fabio
Tamburini, Fabio
中科院分区:
计算机科学3区
文献类型:
--
作者:
Calza, Laura;Gagliardi, Gloria;Tamburini, Fabio

文献摘要

被引文献

相似文献

2018年,全球有近5000万人患有痴呆症,这一数字每20年将翻一番。现有药物治疗的有效性仅限于症状控制,并且没有一种药物能够预防、逆转或关闭导致痴呆的神经退行性过程;因此,为了开发和测试新药并支持临床和家庭环境的管理,及时检测“疾病特征”是一个关键问题。最近的研究表明,语言改变可能是病理学的最早迹象之一,比其他神经认知缺陷变得明显要早几年。传统的测试无法识别这些轻微但明显的变化,而自然语言处理(NLP)技术的口语分析可以生态地和廉价地识别潜在患者的微小语言修改。这项跨学科的研究旨在量化和描述由于认知下降而导致的语言特征的改变,并建立一个自动化系统,用于早期诊断和筛查目的。为此,我们招募了96名参与者:48名健康对照组和48名受损受试者。在后者中,32人被诊断患有轻度认知障碍,16人患有早期痴呆症(艾德)。每个受试者都进行了简短的神经心理学筛查,并通过三个启发任务收集了半自发语音产品的样本。记录的会话进行了正字法转录,PoS标记和解析,建立了两个不同的语料库:在第一个语料库中,我们保留了自动注释,而在第二个语料库中,为了删除所有错误,我们手动更正了转录本。一个多维参数计算的数据进行,考虑到一组87个声学,节奏,形态句法和词汇的功能,以及一些可读性指标和人口统计信息。在这些准备步骤之后,训练一些自动分类器以采用两种不同的算法(支持向量(SVC)和随机森林分类器(RFC))来区分健康对照与MCI受试者。我们的系统能够区分对照组和表现出高F1分数(约75%)的MCI受试者,因此它似乎是识别痴呆症临床前阶段的一种有前途的方法。(C)2020爱思唯尔有限公司保留所有权利。
Almost 50 million people are living with dementia in 2018 worldwide, and the number will double every 20 years. The effectiveness of existing pharmacologic treatments for the disease is limited to symptoms control, and none of them are able to prevent, reverse or turn off the neurodegenerative process that leads to dementia; therefore, a prompt detection of the "disease signature" is a key problem, in order to develop and test new drugs and to support the management of clinical and domestic context. Recent studies showed that linguistic alterations may be one of the earliest signs of the pathology, years before other neurocognitive deficits become evident. Traditional tests fail to identify these slight but noticeable changes; whereas, the analysis of spoken language productions by Natural Language Processing (NLP) techniques can ecologically and inexpensively identify minor language modifications in potential patients.This interdisciplinary study aims at quantifying and describing alterations of linguistic features due to cognitive decline and build an automatic system for early diagnosis and screening purpose. To this aim, we enrolled 96 participants: 48 healthy controls and 48 impaired subjects. Of the latter, 32 was diagnosed with Mild Cognitive Impairment and 16 with early Dementia (eD). Each subject underwent a brief neuropsychological screening, and samples of semi-spontaneous speech productions was collected by means of three elicitation tasks. Recorded sessions were orthographically transcribed, PoS tagged and parsed building two different corpora: in the first we kept the automatic annotations, while in the second the transcripts were manually corrected in order to remove all mistakes. A multidimensional parameter computation was performed on the data, taking into consideration a set of 87 acoustical, rhythmical, morpho-syntactic and lexical feature as well as some readability indexes and demographic information. After these preparatory steps, some automatic classifiers were trained to distinguish healthy controls from MCI subjects employing two different algorithms, Support Vector (SVC) and Random Forest Classifiers (RFC). Our system was able to distinguish between controls and MCI subjects exhibiting high F1 scores, around 75%, thus it seems to be a promising approach for the identification of preclinical stages of dementia. (C) 2020 Elsevier Ltd. All rights reserved.