Algorithmic Classification of Five Characteristic Types of Paraphasias

Algorithmic Classification of Five Characteristic Types of Paraphasias
复制标题

DOI:
10.1044/2016_ajslp-15-0147
复制
发表时间:
2016-12-01
影响因子:
2.6
通讯作者:
Bedrick, Steven
Bedrick, Steven
中科院分区:
医学2区
文献类型:
--
作者:
Fergadiotis, Gerasimos;Gorman, Kyle;Bedrick, Steven

文献摘要

被引文献

相似文献

目的:本研究旨在评估一系列用于对失语错误(形式错误、语义错误、混合错误、新学错误和不相关错误)进行自动分类的算法。方法:我们分析了 Moss 失语症心理语言学项目数据库(Mirman 等人,2010)中的 7,111 例失语症,并评估了 3 个自动化工具的分类准确性。首先,我们使用 SUBTLEXus 数据库(Brysbaert & New,2009)中的频率规范来区分非单词错误和真实单词产生。然后,我们实现了语音相似性算法来识别语音相关的真实单词错误。最后,我们评估了基于 word2vec 的语义相似性标准的性能(Mikolov、Yih 和 Zweig,2013)。结果:总体而言,算法分类复制了人类对失语症主要类别的评分,具有较高的准确性。该工具基于 SUBTLEXus 频率规范,词汇判断准确率超过 97%。语音相似性标准的准确度约为 91%,语义分类器的总体分类准确度范围为 86% 至 90%。 结论:总体而言,结果凸显了自然语言处理领域的工具在开发适合收集用于研究和临床目的的高质量测量数据的高度可靠、经济高效的诊断工具方面的潜力。
Purpose: This study was intended to evaluate a series of algorithms developed to perform automatic classification of paraphasic errors (formal, semantic, mixed, neologistic, and unrelated errors).Method: We analyzed 7,111 paraphasias from the Moss Aphasia Psycholinguistics Project Database (Mirman et al., 2010) and evaluated the classification accuracy of 3 automated tools. First, we used frequency norms from the SUBTLEXus database (Brysbaert & New, 2009) to differentiate nonword errors and real-word productions. Then we implemented a phonological-similarity algorithm to identify phonologically related real-word errors. Last, we assessed the performance of a semantic-similarity criterion that was based on word2vec (Mikolov, Yih, & Zweig, 2013).Results: Overall, the algorithmic classification replicated human scoring for the major categories of paraphasias studied with high accuracy. The tool that was based on the SUBTLEXus frequency norms was more than 97% accurate in making lexicality judgments. The phonological-similarity criterion was approximately 91% accurate, and the overall classification accuracy of the semantic classifier ranged from 86% to 90%.Conclusion: Overall, the results highlight the potential of tools from the field of natural language processing for the development of highly reliable, cost-effective diagnostic tools suitable for collecting high-quality measurement data for research and clinical purposes.