Crowd Sourcing Labels From Electronic Medical Records to Enable Biomedical Research
Crowd Sourcing Labels From Electronic Medical Records to Enable Biomedical Research
批准号:
9270528
负责人:
Daniel Fabbri
金额:
$31.6万
依托单位国家:
美国
项目类别:
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-05-06 至 2019-04-30
关键词:
Accident and Emergency departmentAddressAlgorithmsArchitectureAreaAsthmaBiomedical ResearchCharacteristicsChildhoodClinicalClinical DataCollectionComputerized Medical RecordCrowdingDataData SetData SourcesDevelopmentDisclosureEnsureEvaluationEventExtravasationFutureGoalsHealthHealth Insurance Portability and Accountability ActHuman ResourcesIncentivesInterviewLabelMachine LearningManagement AuditManualsMeasuresMechanicsMedical RecordsMedical ResearchMedical StudentsMedical centerMethodsModelingNursesOutcomePatientsPrivacyProductivityReceiver Operating CharacteristicsRelapseResearchResearch DesignResearch PersonnelResourcesRoleSecuritySpecificityStructureSupervisionSystemTimeTrainingclinical predictorscohortcomputer sciencecrowdsourcingdata sharingdesignmembermodel developmentopen sourcepublic health relevanceresearch studyresponsescale uptool
中文摘要
描述(申请人提供):有监督的机器学习是一种流行的方法,它使用标记的训练样本来预测未来的结果。不幸的是,用于生物医学研究的有监督的机器学习往往受到缺乏标签数据的限制。目前产生带标签数据的方法包括手动查看图表,这很费力,而且不能随数据创建率进行调整。该项目旨在开发一个框架,通过形成一群临床人员标签员来从电子病历中众包标签数据集。这些有标签的数据集的构建将允许进行以前不可行的新的生物医学研究。开发一个临床数据的众包平台在实践和理论上都面临着许多挑战。第一,受欢迎的公众人群
亚马逊的机械土耳其等采购平台不适合为病历贴标签,因为HIPAA会使临床数据共享存在风险。其次,人们还不能很好地理解哪些类型的临床问题是可以众包的。第三,目前尚不清楚临床人群能否快速准确地制作标签。这些挑战中的每一个都将在一个单独的目标中得到解决。作为该项目的第一个目标,该团队将评估不同的临床众包架构。该架构必须利用人群的规模,同时最大限度地减少患者信息的暴露。将考虑使用消除身份识别工具来擦除临床笔记或减少信息泄露。使用这一设计,团队将扩展流行的开源众包工具PyBossa,并将其发布给公众。作为第二个目标,该团队将研究临床预测问题的类型、结构、主题和特异性,以及这些特征如何影响标记器质量。最后,团队将评估质量和准确性
收集关于两个现有图表审查问题的临床众包数据,以确定该平台的实用性。
英文摘要
DESCRIPTION (provided by applicant): Supervised machine learning is a popular method that uses labeled training examples to predict future outcomes. Unfortunately, supervised machine learning for biomedical research is often limited by a lack of labeled data. Current methods to produce labeled data involve manual chart reviews that are laborious and do not scale with data creation rates. This project aims to develop a framework to crowd source labeled data sets from electronic medical records by forming a crowd of clinical personnel labelers. The construction of these labeled data sets will allow for new biomedical research studies that were previously infeasible to conduct. There are numerous practical and theoretical challenges of developing a crowd sourcing platform for clinical data. First, popular, public crowd
sourcing platforms such as Amazon's Mechanical Turk are not suitable for medical record labeling as HIPAA makes clinical data sharing risky. Second, the types of clinical questions that are amenable for crowd sourcing are not well understood. Third, it is unclear if the clinical crowd can produce labels quickly and accurately. Each of these challenges will be addressed in a separate Aim. As the first Aim of this project, the team will evaluate different clinical crowd sourcing architectures. The architecture must leverage the scale of the crowd, while minimizing patient information exposure. De-identification tools will be considered to scrub clinical notes t reduce information leakage. Using this design, the team will extend a popular open source crowd sourcing tool, Pybossa, and release it to the public. As the second Aim, the team will study the type, structure, topic and specificity of clinical prediction questions, and how these characteristics impact labeler quality. Lastly, the team will evaluate the quality and accuracy of
collected clinical crowd sourced data on two existing chart review problems to determine the platform's utility.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
A Crowdsourcing Framework for Medical Data Sets.
医疗数据集的众包框架。
DOI:
--
发表时间:
2018
期刊:
AMIA Joint Summits on Translational Science proceedings. AMIA Joint Summits on Translational Science
影响因子:
--
作者:
[Ye,Cheng, Coco,Joseph, Epishova,Anna, Hajaj,Chen, Bogardus,Henry, Novak,Laurie, Denny,Joshua, Vorobeychik,Yevgeniy, Lasko,Thomas, Malin,Bradley, Fabbri,Daniel]
通讯作者:
Fabbri,Daniel
Crowd Sourcing Labels From Electronic Medical Records to Enable Biomedical Research
-
批准号:9076555
-
项目类别:
-
资助金额:$31.6万
-
财政年份:2016
-
负责人:Daniel Fabbri
-
依托单位:
海外基金