Describing Vocalizations in Young Children: A Big Data Approach Through Citizen Science Annotation.

Describing Vocalizations in Young Children: A Big Data Approach Through Citizen Science Annotation.
复制标题

描述幼儿的发声:通过公民科学注释的大数据方法。

DOI:
10.1044/2021_jslhr-20-00661
复制
发表时间:
2021
期刊:
Journal of speech, language, and hearing research : JSLHR
影响因子:
--
通讯作者:
Cristia,Alejandrina
Cristia,Alejandrina
中科院分区:
--
文献类型:
--
作者:
Semenzin,Chiara;Hamrick,Lisa;Seidl,Amanda;Kelleher,BridgetteL;Cristia,Alejandrina

文献摘要

被引文献

相似文献

通过可穿戴设备记录幼儿的发声是一种很有前途的评估语言发展的方法。然而,准确和快速地注释这些文件仍然具有挑战性。与公民科学家合作的在线众包可能是一个可行的解决方案。在这篇文章中,我们评估了公民科学家的注释与实验室中收集的幼儿录音的注释一致的程度。10名低风险对照儿童和10名诊断为Angelman综合征的儿童,这是一种以严重语言障碍为特征的神经遗传综合征。语音样本由实验室中训练有素的注释者以及Zooniverse上的公民科学家进行注释。所有的注释者都为每个样本分配了五个标签之一:Canonical,Noncanonical,Crying,Laughing和Junk。这允许推导两个儿童级的发声指标:语言的比例和典型的proportion.ResultsAt段水平,Zooniverse分类有适度的精度和召回。更重要的是,来自Zooniverse注释的语言比例和规范比例与来自实验室annotation.ConclusionsAnnotations通过公民科学平台获得的高度相关,可以帮助我们克服注释一整天的语音记录的过程所带来的挑战。特别是当在复合或衍生度量中使用时,这样的注释可以用于调查语言延迟的早期标记。
PurposeRecording young children's vocalizations through wearables is a promising method to assess language development. However, accurately and rapidly annotating these files remains challenging. Online crowdsourcing with the collaboration of citizen scientists could be a feasible solution. In this article, we assess the extent to which citizen scientists' annotations align with those gathered in the lab for recordings collected from young children.MethodSegments identified by Language ENvironment Analysis as produced by the key child were extracted from one daylong recording for each of 20 participants: 10 low-risk control children and 10 children diagnosed with Angelman syndrome, a neurogenetic syndrome characterized by severe language impairments. Speech samples were annotated by trained annotators in the laboratory as well as by citizen scientists on Zooniverse. All annotators assigned one of five labels to each sample: Canonical, Noncanonical, Crying, Laughing, and Junk. This allowed the derivation of two child-level vocalization metrics: the Linguistic Proportion and the Canonical Proportion.ResultsAt the segment level, Zooniverse classifications had moderate precision and recall. More importantly, the Linguistic Proportion and the Canonical Proportion derived from Zooniverse annotations were highly correlated with those derived from laboratory annotations.ConclusionsAnnotations obtained through a citizen science platform can help us overcome challenges posed by the process of annotating daylong speech recordings. Particularly when used in composites or derived metrics, such annotations can be used to investigate early markers of language delays.