Look who's talking: A comparison of automated and human-generated speaker tags in naturalistic day-long recordings

Look who's talking: A comparison of automated and human-generated speaker tags in naturalistic day-long recordings
复制标题

DOI:
10.3758/s13428-019-01265-7
复制
发表时间:
2020-04-01
影响因子:
5.4
通讯作者:
Bergelson, Elika
Bergelson, Elika
中科院分区:
心理学2区
文献类型:
--
作者:
Bulgarelli, Federica;Bergelson, Elika

文献摘要

被引文献

相似文献

LENA系统彻底改变了对语言习得的研究,提供了一种可穿戴设备来收集儿童环境的全天记录,以及一组使用专有算法处理,识别和分类语音的自动输出。该输出包括关于输入源的信息(例如,成年男性,电子产品)。虽然这个系统已经在各种设置中进行了测试,但在这里,我们更深入地验证了LENA自动日记的准确性和可靠性,即,标签是谁在说话具体来说,我们将LENA的输出与手动生成的说话者标签的黄金标准集进行了比较,这些标签来自一个包含88天录音的数据集,这些录音来自44名6个月和7个月大的婴儿,其中包括57,983个话语。我们比较了原始Lena技术报告中一系列分类的准确性,以及一组按话语类型检查分类准确性的分析(例如,歌唱,歌唱)。与之前的验证一致,我们发现人类和LENA生成的扬声器标签之间的总体高度一致性,特别是对于成人语音,识别儿童,重叠,噪声和电子语音的性能较差(所有测量的准确性范围:0-92%)。我们讨论了几个明显的好处,使用这个自动化系统的基础上,我们观察到的错误模式,潜在的警告,最后与使用LENA生成的扬声器标签的研究的影响。
The LENA system has revolutionized research on language acquisition, providing both a wearable device to collect day-long recordings of children's environments, and a set of automated outputs that process, identify, and classify speech using proprietary algorithms. This output includes information about input sources (e.g., adult male, electronics). While this system has been tested across a variety of settings, here we delve deeper into validating the accuracy and reliability of LENA's automated diarization, i.e., tags of who is talking. Specifically, we compare LENA's output with a gold standard set of manually generated talker tags from a dataset of 88 day-long recordings, taken from 44 infants at 6 and 7 months, which includes 57,983 utterances. We compare accuracy across a range of classifications from the original Lena Technical Report, alongside a set of analyses examining classification accuracy by utterance type (e.g., declarative, singing). Consistent with previous validations, we find overall high agreement between the human and LENA-generated speaker tags for adult speech in particular, with poorer performance identifying child, overlap, noise, and electronic speech (accuracy range across all measures: 0-92%). We discuss several clear benefits of using this automated system alongside potential caveats based on the error patterns we observe, concluding with implications for research using LENA-generated speaker tags.