Describing Vocalizations in Young Children: A Big Data Approach Through Citizen Science Annotation.
Describing Vocalizations in Young Children: A Big Data Approach Through Citizen Science Annotation.
复制标题
描述幼儿的发声:通过公民科学注释的大数据方法。
DOI:
10.1044/2021_jslhr-20-00661
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Cristia,Alejandrina
中科院分区:
文献类型:
--
作者:
Semenzin,Chiara;Hamrick,Lisa;Seidl,Amanda;Kelleher,BridgetteL;Cristia,Alejandrina
PurposeRecording young children's vocalizations through wearables is a promising method to assess language development. However, accurately and rapidly annotating these files remains challenging. Online crowdsourcing with the collaboration of citizen scientists could be a feasible solution. In this article, we assess the extent to which citizen scientists' annotations align with those gathered in the lab for recordings collected from young children.MethodSegments identified by Language ENvironment Analysis as produced by the key child were extracted from one daylong recording for each of 20 participants: 10 low-risk control children and 10 children diagnosed with Angelman syndrome, a neurogenetic syndrome characterized by severe language impairments. Speech samples were annotated by trained annotators in the laboratory as well as by citizen scientists on Zooniverse. All annotators assigned one of five labels to each sample: Canonical, Noncanonical, Crying, Laughing, and Junk. This allowed the derivation of two child-level vocalization metrics: the Linguistic Proportion and the Canonical Proportion.ResultsAt the segment level, Zooniverse classifications had moderate precision and recall. More importantly, the Linguistic Proportion and the Canonical Proportion derived from Zooniverse annotations were highly correlated with those derived from laboratory annotations.ConclusionsAnnotations obtained through a citizen science platform can help us overcome challenges posed by the process of annotating daylong speech recordings. Particularly when used in composites or derived metrics, such annotations can be used to investigate early markers of language delays.