Automation of Language Sample Analysis.

Automation of Language Sample Analysis.
复制标题

DOI:
10.1044/2023_jslhr-22-00642
复制
发表时间:
2023-07-12
影响因子:
2.6
通讯作者:
Lanzi, Alyssa
Lanzi, Alyssa
中科院分区:
医学2区
文献类型:
--
作者:
Liu, Houjun;MacWhinney, Brian;Fromm, Davida;Lanzi, Alyssa

文献摘要

相似文献

语言样本分析(LSA)广泛应用的一个主要障碍是转录非常耗时。减少所需时间和精力的方法有助于促进LSA在临床实践和研究中的应用。本文描述了一个称为Batchalign的自动化管道,它获取原始音频并以“人类对话分析代码”(CHAT)转录格式创建完整的转录本,并完成了话语和单词级别的时间对齐和形态句法分析。管道只需要主要的人工干预来进行最终检查。它将一系列现有工具与其他新颖的重新格式化过程结合在一起。管道中的步骤是(a)自动语音识别,(b)话语标记化,(c)自动更正,(d)说话者ID分配,(e)强制对齐,(f)用户调整,以及(g)自动形态句法和分析。在对患有语言障碍的成年人的录音进行研究时,获得了六个主要结果:(a)单词错误率在对照组的2.4%和患者的3.4%之间,(b)话语分词准确率在无语言障碍的说话者报告的水平,(c)对照参与者的单词分词准确率为93%,语言障碍参与者的单词分词准确率为83%,(d)基于单词分词的单词分词准确率很高,(e)对CHAT格式的坚持是完全准确的,(f)人类转录时间减少了75%。该管道极大地缩短了数据收集和数据分析之间的时间间隔,并提供了优于通常由人类转录器生成的输出。
A major barrier to the wider use of language sample analysis (LSA) is the fact that transcription is very time intensive. Methods that can reduce the required time and effort could help in promoting the use of LSA for clinical practice and research. This article describes an automated pipeline, called Batchalign, that takes raw audio and creates full transcripts in Codes for the Human Analysis of Talk (CHAT) transcription format, complete with utterance- and word-level time alignments and morphosyntactic analysis. The pipeline only requires major human intervention for final checking. It combines a series of existing tools with additional novel reformatting processes. The steps in the pipeline are (a) automatic speech recognition, (b) utterance tokenization, (c) automatic corrections, (d) speaker ID assignment, (e) forced alignment, (f) user adjustments, and (g) automatic morphosyntactic and profiling analyses. For work with recordings from adults with language disorders, six major results were obtained: (a) The word error rate was between 2.4% for controls and 3.4% for patients, (b) utterance tokenization accuracy was at the level reported for speakers without language disorders, (c) word-level diarization accuracy was at 93% for control participants and 83% for participants with language disorders, (d) utterance-level diarization accuracy based on word-level diarization was high, (e) adherence to CHAT format was fully accurate, and (f) human transcriber time was reduced by up to 75%. The pipeline dramatically shortens the time gap between data collection and data analysis and provides an output superior to that typically generated by human transcribers.