Harnessing multimodal data to enhance machine learning of children’s vocalizations
Harnessing multimodal data to enhance machine learning of children’s vocalizations
批准号:
10411575
负责人:
DANIEL S MESSINGER
金额:
$20.0万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-02-01 至 2026-01-31
关键词:
Administrative SupplementAdultAffectAfrican AmericanAfrican CaribbeanAgeAlgorithmsArchivesAsiansBenchmarkingCellsCharacteristicsChildChild DevelopmentChild LanguageChild SupportCochlear implant procedureCodeComplexComputer ModelsComputer Vision SystemsComputer softwareConsentDataData CollectionData SetDatabasesDetectionDevelopmentDimensionsEnvironmentEquilibriumExhibitsExposure toFaceFeedbackFemaleFosteringFrequenciesFundingGardenalGaussian modelGoalsHearingHispanicsHourHumanIndividualLanguageLanguage DevelopmentLearningLeftLegal patentLifeLinguisticsLinkLocationMachine LearningManuscriptsMeasurementMeasuresMetadataModelingModernizationMovementNursery SchoolsOutputParentsParticipantPatternPersonsPlayPositioning AttributePostdoctoral FellowPredictive FactorPrivacyProbabilityProceduresProcessProductionPublicationsPythonsRadialRadioRandomizedReportingResearchResearch PersonnelResearch Project GrantsResourcesSamplingSeriesSex DistributionShoulderSignal TransductionSocial DevelopmentSourceSpeechSpeech SoundStatistical ModelsStreamSystemTensorFlowTestingTimeTimeLineTrainingUnited States National Institutes of HealthWalkersWeightautomated analysisbasecomputerized data processingcontextual factorsconvolutional neural networkcostdata analysis pipelinedata de-identificationdata integrationdata pipelinedeafnessdeep learningdeep neural networkdemographicsdenoisingdesigndyadic interactioneducational atmosphereexperiencefeedingfunctional outcomeshearing impairmentimprovedinteractive feedbackinterestlight weightmachine learning algorithmmalemarkov modelmetermulti-ethnicmultimodal datamultimodalitymultiple datasetsopen dataopen sourceparent grantpeerprediction algorithmprototyperadio frequencyrecurrent neural networkrecursive neural networkrepositoryresponsesensorsignal processingsocialsocioeconomicssoundspeech processingsupervised learningsupplemental instructionteachertoolundergraduate studentvocalization
中文摘要
项目摘要
本行政补充建议实施多模式数据管道,以支持
机器学习在复杂自然环境中的儿童语言生产。补充剂
建立在父R 01(DC 018542)的基础上,该父R 01收集客观的纵向数据以捕获声音
儿童听力损失(HL)的影响。即使有人工耳蜗植入,HL是一个改变生活的
社会成本高。纳入HL儿童和正常听力(TH)同龄人,
学前班教室是一个国家标准,但目前还不清楚如何早期声乐互动有助于
HL儿童及其TH同龄人的语言发展。父R 01采用
儿童位置和方向的计算模型,以指示儿童何时处于社会接触中
和他们的同学和老师一起。实现R 01广泛目标的另一项战略是,
识别儿童产生音位复杂的发声的互动环境,
交互式语音是机器学习。机器学习算法可以确定上下文,
预测儿童发声和声音互动的个体和互动因素。然而,在这方面,
父R 01没有提出机器学习,也没有以设计用于
促进机器学习。为了促进课堂上的机器学习,
需要一个过程来确定说话人身份,这是可操作的可能性,每个
发声是由特定的孩子或老师说出的。我们将整合每个目标的音频处理
孩子和老师的第一人称录音,并处理他们的互动伙伴的录音。
同伴录音的影响将取决于他们的物理距离和相对方位
到目标这将产生每个发声的加权说话者识别分数。对于25%的
样本,算法得分将与由受过训练的编码器提供的说话者识别进行比较,
量化系统间可靠性。处理的数据集将包括7,160小时的多模式记录,
教室里的儿童和教师运动与连续记录的儿童和教师运动同步,
教师专用(第一人称)录音。去识别输出数据将表征发声
关于算法计算的说话者识别概率,编码器识别的说话者
身份(样本的25%)、音素复杂性和音频特征(例如,基频),
以及教室中所有个人的位置和相对方向,
人口统计学(包括HL的特征)。在补充过程中,输出数据,
Python处理代码和处理管道的元数据描述将在
专门的分发门户,包括Github、Kaggle和UCI仓库。记录将
通过NIH资助的数据库(如Databrary和Homebank)发布给经过认证的研究人员。
英文摘要
Project Summary
This Administrative Supplement proposes implementation of a multimodal data pipeline to support
machine learning of child language production in complex naturalistic environments. The Supplement
builds on the parent R01 (DC018542) that gathers objective, longitudinal data to capture the vocal
interactions of children with hearing loss (HL). Even with cochlear implantation, HL is a life-altering
condition with high social costs. Inclusion of children with HL and typically hearing (TH) peers in
preschool classrooms is a national standard, but it is not clear how early vocal interaction contributes to
the language development of children with HL and their TH peers. The parent R01 employs
computational models of child location and orientation to indicate when children are in social contact
with their peers and teachers. An additional strategy for pursuing the broad goals of the R01—
identifying interactive contexts in which children produce phonemically complex vocalizations and
interactive speech—is machine learning. Machine learning algorithms can determine the contextual,
individual, and interactive factors that predict children’s vocalizations and vocal interactions. However,
the parent R01 does not propose machine learning, nor are data disseminated in a format designed to
facilitate machine learning. To facilitate machine learning in the classroom, a rigorous diarization
process is required to determine speaker identity, which is operationalized as the likelihood that each
vocalization was spoken by a given child or teacher. We will integrate audio processing of each target
child and teacher’s first-person audio recording with processing of their interactive partners’ recordings.
The influence of partner recordings will be determined by their physical distance and orientation relative
to the target. This will yield a weighted speaker identification score for each vocalization. For 25% of the
sample, the algorithmic score will be compared to speaker identification provided by trained coders to
quantify intersystem reliability. Processed datasets will include 7,160 hours of multimodal recordings of
child and teacher movement in classrooms synchronized with continuously recorded, child- and
teacher-specific (first-person) audio recordings. De-identified output data will characterize vocalizations
with respect to algorithmically computed speaker identification probabilities, coder-identified speaker
identity (25% of sample), phonemic complexity and audio characteristics (e.g., fundamental frequency),
as well as the position and relative orientation of all individuals in the classroom, and child
demographics (including characterizations of HL). Over the course of the supplement, output data,
Python processing code, and metadata descriptions of the processing pipeline will be disseminated in
dedicated distribution portals including Github, Kaggle, and the UCI repository. Recordings will be
released to certified investigators via NIH-funded repositories such as Databrary and Homebank.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Bioethical Issues Associated with Objective Behavioral Measurement of Children with Hearing Loss in Naturalistic Environments
-
批准号:10790269
-
项目类别:
-
资助金额:$29.7万
-
财政年份:2023
-
负责人:DANIEL S MESSINGER
-
依托单位:
Characterizing bilingual spoken language experiences in preschoolers with hearing loss
-
批准号:10802499
-
项目类别:
-
资助金额:$5.62万
-
财政年份:2023
-
负责人:DANIEL S MESSINGER
-
依托单位:
Language Development and Social Interaction in Children with Hearing Loss
-
批准号:10605307
-
项目类别:
-
资助金额:$29.08万
-
财政年份:2021
-
负责人:DANIEL S MESSINGER
-
依托单位:
Language Development and Social Interaction in Children with Hearing Loss
-
批准号:10335271
-
项目类别:
-
资助金额:$29.62万
-
财政年份:2021
-
负责人:DANIEL S MESSINGER
-
依托单位:
Social-Emotional Development of Infants At Risk for Autism Spectrum
-
批准号:7694276
-
项目类别:
-
资助金额:$60.66万
-
财政年份:2008
-
负责人:DANIEL S MESSINGER
-
依托单位:
Social-Emotional Development of Infants At Risk for Autism Spectrum
-
批准号:8323829
-
项目类别:
-
资助金额:$66.27万
-
财政年份:2008
-
负责人:DANIEL S MESSINGER
-
依托单位:
Social-Emotional Development of Infants At Risk for Autism Spectrum
-
批准号:8421563
-
项目类别:
-
资助金额:$3.9万
-
财政年份:2008
-
负责人:DANIEL S MESSINGER
-
依托单位:
Social-Emotional Development of Infants At Risk for Autism Spectrum
-
批准号:8141259
-
项目类别:
-
资助金额:$59.9万
-
财政年份:2008
-
负责人:DANIEL S MESSINGER
-
依托单位:
Social-Emotional Development of Infants At Risk for Autism Spectrum
-
批准号:7527975
-
项目类别:
-
资助金额:$62.08万
-
财政年份:2008
-
负责人:DANIEL S MESSINGER
-
依托单位:
Social-Emotional Development of Infants At Risk for Autism Spectrum
-
批准号:7901094
-
项目类别:
-
资助金额:$60.5万
-
财政年份:2008
-
负责人:DANIEL S MESSINGER
-
依托单位:
Naive Observers' Ratings of Behavior: A Multi-Construct Validation Study
-
批准号:7195203
-
项目类别:
-
资助金额:$22.13万
-
财政年份:2007
-
负责人:DANIEL S MESSINGER
-
依托单位:
Naive Observers' Ratings of Behavior: A Multi-Construct Validation Study
-
批准号:7352781
-
项目类别:
-
资助金额:$18.38万
-
财政年份:2007
-
负责人:DANIEL S MESSINGER
-
依托单位:
Emotion, communication, & EEG: Development & risk
-
批准号:7816917
-
项目类别:
-
资助金额:$29.52万
-
财政年份:2006
-
负责人:DANIEL S MESSINGER
-
依托单位:
Emotion, communication, & EEG: Development & risk
-
批准号:7623228
-
项目类别:
-
资助金额:$29.82万
-
财政年份:2006
-
负责人:DANIEL S MESSINGER
-
依托单位:
Emotion, communication, & EEG: Development & risk
-
批准号:7426481
-
项目类别:
-
资助金额:$29.82万
-
财政年份:2006
-
负责人:DANIEL S MESSINGER
-
依托单位:
Emotion, communication, & EEG: Development & risk
-
批准号:7144489
-
项目类别:
-
资助金额:$30.57万
-
财政年份:2006
-
负责人:DANIEL S MESSINGER
-
依托单位:
Emotion, communication, & EEG: Development & risk
-
批准号:7261921
-
项目类别:
-
资助金额:$30.01万
-
财政年份:2006
-
负责人:DANIEL S MESSINGER
-
依托单位:
A Multi-Method Investigation of Infant Emotion
-
批准号:6420956
-
项目类别:
-
资助金额:$7.55万
-
财政年份:2002
-
负责人:DANIEL S MESSINGER
-
依托单位:
A Multi-Method Investigation of Infant Emotion
-
批准号:6687602
-
项目类别:
-
资助金额:$0.72万
-
财政年份:2002
-
负责人:DANIEL S MESSINGER
-
依托单位:
A Multi-Method Investigation of Infant Emotion
-
批准号:6620716
-
项目类别:
-
资助金额:$9.77万
-
财政年份:2002
-
负责人:DANIEL S MESSINGER
-
依托单位:
海外基金