Transfer Learning for Digital Curation of the EMR Clinical Narrative
Transfer Learning for Digital Curation of the EMR Clinical Narrative
批准号:
10468604
负责人:
GUERGANA K. SAVOVA
金额:
$37.61万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-08-12 至 2025-05-31
关键词:
Adverse eventApacheAppearanceAreaArtificial IntelligenceAutomated AnnotationBiomedical ResearchBrain AneurysmsClassificationClinicalClinical InvestigatorCommunitiesComputer Vision SystemsComputerized Medical RecordCoupledDataData ScienceData SetDatabasesDevelopmentDiseaseE-learningEngineeringEpigenetic ProcessEvaluationEventGeneticHealthHepatotoxicityImageInformation RetrievalIntuitionInvestigationLabelLanguageMachine LearningMedicineMethodologyMethodsMethotrexateModelingMultiple SclerosisNatural Language ProcessingNeural Network SimulationPatientsPerformancePharmaceutical PreparationsPhenotypeProceduresPublic HealthPublicationsPublishingRare DiseasesReportingResearchResourcesRheumatoid ArthritisSemanticsSeveritiesSigns and SymptomsSourceSpeechStructureSupervisionSystemTechniquesTestingTextTrainingTranslational ResearchVisionWorkautism spectrum disorderbasecomorbiditycomputing resourcesdeep learningdeep neural networkdigitalimaging Segmentationimprovedinformatics toollearning strategymachine learning methodmultitaskneural networknext generationnovelopen sourcephenotypic datarelating to nervous systemresponsesupport vector machinetransfer learningunstructured data
中文摘要
项目摘要
这项建议是对PAR 18-796的回应,目的是寻求对推进转让方法学的支持
电子病历(EMR)临床叙述数字策划的学习框架。在
在当今人工智能(AI)在生物医学中日益重要的时代,我们的建议解决了一个关键的
AI组件--健康相关文本的自动标注。自2015年以来,该技术的发展和应用
机器学习(ML)方法在大量数字化非结构化的融合的推动下得到了爆炸性的发展
数据(文本、语音、图像)、硬件和神经网络的改进或深度学习。2018年标志着
自然语言处理(NLP)的转折点,特别是通过预先训练的模型进行迁移学习
像通用语言模型精调文本分类,Allen AI的Elmo,OpenAI的Open-GPT。在……里面
2018年11月,谷歌发布了Transformers(BERT)的双向编码表示法,a
基于变压器的模型在海量通用文本数据库(共33亿字)上进行了预训练。这份出版物
报告使用BERT表示为11个NLP任务构建了分类器,这些任务的表现优于最新状态-
艺术品(SOTA),利润率很高。NLP研究界立即提出了探索这一新技术的想法
框架,但很快意识到从头开始构建Bert风格的模型是负担得起的,而且
只有少数人才能做到。因此,研究调查朝着使用这些巨大模型的方向进行
作为语言表示的资源。科学工作的重点是预先培训的模型(例如BERT)作为
提取高质量语言功能或对特定任务进行微调的源代码,即使用模型作为
检查点和使用数量小得多的特定于任务的数据进行重新培训,以生成预测,通常
在表示的顶部添加一个完全连接的层,并针对几个纪元进行培训。这位将军
自然语言处理向迁移学习的分水岭转变与几年来计算机视觉的发展相平行
再加上我们的最新工作,将一个关键的NLP研究课题带到了前沿,这个课题已经成熟,可以探索--一个
电子病历临床叙事数字策划的迁移学习框架。拟做的工作是研究
从健康相关文本中提取详细信息的新科学方法,特别是电子病历,
患者表型数据的主要来源。需要精确的表型信息来推进翻译
研究,特别是揭示遗传、表观遗传和系统变化对反应能力的影响。
这项研究总体上符合神经深度学习方法和人工智能的最新发展
预计将加强生物医学研究,并通过此促进公众的健康。
英文摘要
Project Summary
This proposal is in response to PAR 18-796 to seek support for advancing methodologies for a transfer
learning framework for the digital curation of the Electronic Medical Records (EMR) clinical narrative. In the
current era of increasing importance of Artificial Intelligence (AI) in biomedicine, our proposal tackles a critical
AI component – automated annotation of health-related text. Since 2015 the development and application of
machine learning (ML) methods has exploded propelled by the convergence of plentiful digitized unstructured
data (text, speech, images), hardware and the refinement of neural networks or deep learning. 2018 marked a
turning point in Natural Language Processing (NLP), particularly transfer learning through pre-trained models
like Universal Language Model Fine-tuning for Text Classification, Allen AI's ELMO, OpenAI's Open-GPT. In
November 2018, Google published the Bidirectional Encodings Representations from Transformers (BERT), a
transformer-based model pre-trained on massive general text databases (3.3B words total). The publication
reported using BERT representations to build classifiers for 11 NLP tasks which outperformed the state-of-the-
art (SOTA) with large margins. The NLP research community jumped to the idea of exploring this new
framework but quickly came to the realization that building BERT-style models from scratch is affordable and
feasible to only a few. Thus, research investigation proceeded in the direction of using these gigantic models
as resources for language representations. Scientific efforts focused on pre-trained models (e.g. BERT) as a
source of extracting high quality language features or fine-tuning on a specific task, i.e. using a model as a
checkpoint and re-training with much smaller amounts of task-specific data to produce predictions by typically
adding one fully-connected layer on top of the representations and training for a few epochs. This general
watershed shift in NLP to transfer learning which parallels the developments in computer vision a few years
ago coupled with our latest work brings to the forefront a critical NLP research topic ripe for exploration – a
transfer learning framework for the digital curation of the EMR clinical narrative. The proposed work is research
of novel scientific methods for extracting detailed information from health-related text especially the EMR, the
major source of phenotype data for patients. Precise phenotype information is needed to advance translational
research, particularly to unravel the effects of genetic, epigenetic, and systems changes on responsiveness.
This research is in line with the latest developments in neural deep learning approaches and AI in general and
is expected to enhance biomedical research and through that the health of the public.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Transfer Learning for Digital Curation of the EMR Clinical Narrative
-
批准号:10092340
-
项目类别:
-
资助金额:$37.61万
-
财政年份:2021
-
负责人:GUERGANA K. SAVOVA
-
依托单位:
Transfer Learning for Digital Curation of the EMR Clinical Narrative
-
批准号:10647748
-
项目类别:
-
资助金额:$37.61万
-
财政年份:2021
-
负责人:GUERGANA K. SAVOVA
-
依托单位:
Cancer Deep Phenotype Extraction from Electronic Medical Records
-
批准号:9538366
-
项目类别:
-
资助金额:$56.79万
-
财政年份:2014
-
负责人:GUERGANA K. SAVOVA
-
依托单位:
Multi-source clinical Question Answering system
-
批准号:7842799
-
项目类别:
-
资助金额:$49.75万
-
财政年份:2009
-
负责人:GUERGANA K. SAVOVA
-
依托单位:
Multi-source clinical Question Answering system
-
批准号:7936991
-
项目类别:
-
资助金额:$49.14万
-
财政年份:2009
-
负责人:GUERGANA K. SAVOVA
-
依托单位:
国内基金
海外基金
基于Apache Spark的可扩展宏基因组序列组装方法研究
-
批准号:61802246
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:邓丽
-
依托单位: