课题基金 / 基金详情

Harnessing multimodal data to enhance machine learning of children’s vocalizations

Harnessing multimodal data to enhance machine learning of children’s vocalizations
利用多模态数据增强儿童发声的机器学习
批准号:
10411575
负责人:
DANIEL S MESSINGER
金额:
$20.0万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-02-01 至 2026-01-31
关键词:
Administrative SupplementAdultAffectAfrican AmericanAfrican CaribbeanAgeAlgorithmsArchivesAsiansBenchmarkingCellsCharacteristicsChildChild DevelopmentChild LanguageChild SupportCochlear implant procedureCodeComplexComputer ModelsComputer Vision SystemsComputer softwareConsentDataData CollectionData SetDatabasesDetectionDevelopmentDimensionsEnvironmentEquilibriumExhibitsExposure toFaceFeedbackFemaleFosteringFrequenciesFundingGardenalGaussian modelGoalsHearingHispanicsHourHumanIndividualLanguageLanguage DevelopmentLearningLeftLegal patentLifeLinguisticsLinkLocationMachine LearningManuscriptsMeasurementMeasuresMetadataModelingModernizationMovementNursery SchoolsOutputParentsParticipantPatternPersonsPlayPositioning AttributePostdoctoral FellowPredictive FactorPrivacyProbabilityProceduresProcessProductionPublicationsPythonsRadialRadioRandomizedReportingResearchResearch PersonnelResearch Project GrantsResourcesSamplingSeriesSex DistributionShoulderSignal TransductionSocial DevelopmentSourceSpeechSpeech SoundStatistical ModelsStreamSystemTensorFlowTestingTimeTimeLineTrainingUnited States National Institutes of HealthWalkersWeightautomated analysisbasecomputerized data processingcontextual factorsconvolutional neural networkcostdata analysis pipelinedata de-identificationdata integrationdata pipelinedeafnessdeep learningdeep neural networkdemographicsdenoisingdesigndyadic interactioneducational atmosphereexperiencefeedingfunctional outcomeshearing impairmentimprovedinteractive feedbackinterestlight weightmachine learning algorithmmalemarkov modelmetermulti-ethnicmultimodal datamultimodalitymultiple datasetsopen dataopen sourceparent grantpeerprediction algorithmprototyperadio frequencyrecurrent neural networkrecursive neural networkrepositoryresponsesensorsignal processingsocialsocioeconomicssoundspeech processingsupervised learningsupplemental instructionteachertoolundergraduate studentvocalization

项目摘要

项目成果

DANIEL S MESSINGER的其他基金

相似基金

相关文献

中文摘要
翻译
项目摘要 本行政补充建议实施多模式数据管道,以支持 机器学习在复杂自然环境中的儿童语言生产。补充剂 建立在父R 01(DC 018542)的基础上,该父R 01收集客观的纵向数据以捕获声音 儿童听力损失(HL)的影响。即使有人工耳蜗植入,HL是一个改变生活的 社会成本高。纳入HL儿童和正常听力(TH)同龄人, 学前班教室是一个国家标准,但目前还不清楚如何早期声乐互动有助于 HL儿童及其TH同龄人的语言发展。父R 01采用 儿童位置和方向的计算模型,以指示儿童何时处于社会接触中 和他们的同学和老师一起。实现R 01广泛目标的另一项战略是, 识别儿童产生音位复杂的发声的互动环境, 交互式语音是机器学习。机器学习算法可以确定上下文, 预测儿童发声和声音互动的个体和互动因素。然而,在这方面, 父R 01没有提出机器学习,也没有以设计用于 促进机器学习。为了促进课堂上的机器学习, 需要一个过程来确定说话人身份,这是可操作的可能性,每个 发声是由特定的孩子或老师说出的。我们将整合每个目标的音频处理 孩子和老师的第一人称录音,并处理他们的互动伙伴的录音。 同伴录音的影响将取决于他们的物理距离和相对方位 到目标这将产生每个发声的加权说话者识别分数。对于25%的 样本,算法得分将与由受过训练的编码器提供的说话者识别进行比较, 量化系统间可靠性。处理的数据集将包括7,160小时的多模式记录, 教室里的儿童和教师运动与连续记录的儿童和教师运动同步, 教师专用(第一人称)录音。去识别输出数据将表征发声 关于算法计算的说话者识别概率,编码器识别的说话者 身份(样本的25%)、音素复杂性和音频特征(例如,基频), 以及教室中所有个人的位置和相对方向, 人口统计学(包括HL的特征)。在补充过程中,输出数据, Python处理代码和处理管道的元数据描述将在 专门的分发门户,包括Github、Kaggle和UCI仓库。记录将 通过NIH资助的数据库(如Databrary和Homebank)发布给经过认证的研究人员。
英文摘要
Project Summary This Administrative Supplement proposes implementation of a multimodal data pipeline to support machine learning of child language production in complex naturalistic environments. The Supplement builds on the parent R01 (DC018542) that gathers objective, longitudinal data to capture the vocal interactions of children with hearing loss (HL). Even with cochlear implantation, HL is a life-altering condition with high social costs. Inclusion of children with HL and typically hearing (TH) peers in preschool classrooms is a national standard, but it is not clear how early vocal interaction contributes to the language development of children with HL and their TH peers. The parent R01 employs computational models of child location and orientation to indicate when children are in social contact with their peers and teachers. An additional strategy for pursuing the broad goals of the R01— identifying interactive contexts in which children produce phonemically complex vocalizations and interactive speech—is machine learning. Machine learning algorithms can determine the contextual, individual, and interactive factors that predict children’s vocalizations and vocal interactions. However, the parent R01 does not propose machine learning, nor are data disseminated in a format designed to facilitate machine learning. To facilitate machine learning in the classroom, a rigorous diarization process is required to determine speaker identity, which is operationalized as the likelihood that each vocalization was spoken by a given child or teacher. We will integrate audio processing of each target child and teacher’s first-person audio recording with processing of their interactive partners’ recordings. The influence of partner recordings will be determined by their physical distance and orientation relative to the target. This will yield a weighted speaker identification score for each vocalization. For 25% of the sample, the algorithmic score will be compared to speaker identification provided by trained coders to quantify intersystem reliability. Processed datasets will include 7,160 hours of multimodal recordings of child and teacher movement in classrooms synchronized with continuously recorded, child- and teacher-specific (first-person) audio recordings. De-identified output data will characterize vocalizations with respect to algorithmically computed speaker identification probabilities, coder-identified speaker identity (25% of sample), phonemic complexity and audio characteristics (e.g., fundamental frequency), as well as the position and relative orientation of all individuals in the classroom, and child demographics (including characterizations of HL). Over the course of the supplement, output data, Python processing code, and metadata descriptions of the processing pipeline will be disseminated in dedicated distribution portals including Github, Kaggle, and the UCI repository. Recordings will be released to certified investigators via NIH-funded repositories such as Databrary and Homebank.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Bioethical Issues Associated with Objective Behavioral Measurement of Children with Hearing Loss in Naturalistic Environments
  • 批准号:
    10790269
  • 项目类别:
  • 资助金额:
    $29.7万
  • 财政年份:
    2023
  • 负责人:
    DANIEL S MESSINGER
  • 依托单位:
Characterizing bilingual spoken language experiences in preschoolers with hearing loss
  • 批准号:
    10802499
  • 项目类别:
  • 资助金额:
    $5.62万
  • 财政年份:
    2023
  • 负责人:
    DANIEL S MESSINGER
  • 依托单位:
Language Development and Social Interaction in Children with Hearing Loss
  • 批准号:
    10605307
  • 项目类别:
  • 资助金额:
    $29.08万
  • 财政年份:
    2021
  • 负责人:
    DANIEL S MESSINGER
  • 依托单位:
Language Development and Social Interaction in Children with Hearing Loss
  • 批准号:
    10335271
  • 项目类别:
  • 资助金额:
    $29.62万
  • 财政年份:
    2021
  • 负责人:
    DANIEL S MESSINGER
  • 依托单位:
海外基金