Transfer Learning for Digital Curation of the EMR Clinical Narrative
Transfer Learning for Digital Curation of the EMR Clinical Narrative
批准号:
10468604
负责人:
GUERGANA K. SAVOVA
金额:
$37.61万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-08-12 至 2025-05-31
关键词:
Adverse eventApacheAppearanceAreaArtificial IntelligenceAutomated AnnotationBiomedical ResearchBrain AneurysmsClassificationClinicalClinical InvestigatorCommunitiesComputer Vision SystemsComputerized Medical RecordCoupledDataData ScienceData SetDatabasesDevelopmentDiseaseE-learningEngineeringEpigenetic ProcessEvaluationEventGeneticHealthHepatotoxicityImageInformation RetrievalIntuitionInvestigationLabelLanguageMachine LearningMedicineMethodologyMethodsMethotrexateModelingMultiple SclerosisNatural Language ProcessingNeural Network SimulationPatientsPerformancePharmaceutical PreparationsPhenotypeProceduresPublic HealthPublicationsPublishingRare DiseasesReportingResearchResourcesRheumatoid ArthritisSemanticsSeveritiesSigns and SymptomsSourceSpeechStructureSupervisionSystemTechniquesTestingTextTrainingTranslational ResearchVisionWorkautism spectrum disorderbasecomorbiditycomputing resourcesdeep learningdeep neural networkdigitalimaging Segmentationimprovedinformatics toollearning strategymachine learning methodmultitaskneural networknext generationnovelopen sourcephenotypic datarelating to nervous systemresponsesupport vector machinetransfer learningunstructured data
中文摘要
项目摘要
本提案是对PAR 18-796的回应,旨在寻求支持,
电子病历(EMR)临床叙述数字化管理的学习框架。在
当前人工智能(AI)在生物医学中的重要性日益增加,我们的建议解决了一个关键问题,
AI组件-健康相关文本的自动注释。自2015年以来,
机器学习(ML)方法在大量数字化非结构化
数据(文本、语音、图像)、硬件和神经网络或深度学习的改进。2018年,
自然语言处理(NLP)的转折点,特别是通过预训练模型的迁移学习
比如用于文本分类的通用语言模型微调,艾伦AI的埃尔莫,OpenAI的Open-GPT。在
2018年11月,谷歌发布了变形金刚双向编码表示(BERT),
基于transformer的模型在大规模的通用文本数据库(总共33亿字)上进行了预训练。出版
报告使用BERT表示来构建11个NLP任务的分类器,这些任务的性能超过了
艺术(SOTA)与大利润率。NLP研究社区跳到探索这个新的想法,
框架,但很快就意识到,从头开始构建BERT风格的模型是负担得起的,
只对少数人可行。因此,研究调查朝着使用这些巨大模型的方向进行
作为语言表征的资源。科学工作集中在预先训练的模型(例如BERT)上,
提取高质量语言特征或对特定任务进行微调的来源,即使用模型作为
检查点和重新训练,使用更少量的特定于任务的数据来生成预测,
在表示的顶部添加一个完全连接的层,并训练几个时期。这个一般
NLP向迁移学习的分水岭转变,与计算机视觉几年的发展相似
前加上我们的最新工作带来了前沿的一个关键的NLP研究课题成熟的探索-一个
迁移学习框架,用于EMR临床叙述的数字化管理。拟议的工作是研究
从健康相关文本中提取详细信息的新科学方法,特别是EMR,
患者表型数据的主要来源。需要精确的表型信息来推进翻译
研究,特别是解开遗传,表观遗传和系统变化对反应的影响。
这项研究符合神经深度学习方法和人工智能的最新发展,
预计将加强生物医学研究,并通过公众的健康。
英文摘要
Project Summary
This proposal is in response to PAR 18-796 to seek support for advancing methodologies for a transfer
learning framework for the digital curation of the Electronic Medical Records (EMR) clinical narrative. In the
current era of increasing importance of Artificial Intelligence (AI) in biomedicine, our proposal tackles a critical
AI component – automated annotation of health-related text. Since 2015 the development and application of
machine learning (ML) methods has exploded propelled by the convergence of plentiful digitized unstructured
data (text, speech, images), hardware and the refinement of neural networks or deep learning. 2018 marked a
turning point in Natural Language Processing (NLP), particularly transfer learning through pre-trained models
like Universal Language Model Fine-tuning for Text Classification, Allen AI's ELMO, OpenAI's Open-GPT. In
November 2018, Google published the Bidirectional Encodings Representations from Transformers (BERT), a
transformer-based model pre-trained on massive general text databases (3.3B words total). The publication
reported using BERT representations to build classifiers for 11 NLP tasks which outperformed the state-of-the-
art (SOTA) with large margins. The NLP research community jumped to the idea of exploring this new
framework but quickly came to the realization that building BERT-style models from scratch is affordable and
feasible to only a few. Thus, research investigation proceeded in the direction of using these gigantic models
as resources for language representations. Scientific efforts focused on pre-trained models (e.g. BERT) as a
source of extracting high quality language features or fine-tuning on a specific task, i.e. using a model as a
checkpoint and re-training with much smaller amounts of task-specific data to produce predictions by typically
adding one fully-connected layer on top of the representations and training for a few epochs. This general
watershed shift in NLP to transfer learning which parallels the developments in computer vision a few years
ago coupled with our latest work brings to the forefront a critical NLP research topic ripe for exploration – a
transfer learning framework for the digital curation of the EMR clinical narrative. The proposed work is research
of novel scientific methods for extracting detailed information from health-related text especially the EMR, the
major source of phenotype data for patients. Precise phenotype information is needed to advance translational
research, particularly to unravel the effects of genetic, epigenetic, and systems changes on responsiveness.
This research is in line with the latest developments in neural deep learning approaches and AI in general and
is expected to enhance biomedical research and through that the health of the public.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Transfer Learning for Digital Curation of the EMR Clinical Narrative
-
批准号:10092340
-
项目类别:
-
资助金额:$37.61万
-
财政年份:2021
-
负责人:GUERGANA K. SAVOVA
-
依托单位:
Transfer Learning for Digital Curation of the EMR Clinical Narrative
-
批准号:10647748
-
项目类别:
-
资助金额:$37.61万
-
财政年份:2021
-
负责人:GUERGANA K. SAVOVA
-
依托单位:
Cancer Deep Phenotype Extraction from Electronic Medical Records
-
批准号:9538366
-
项目类别:
-
资助金额:$56.79万
-
财政年份:2014
-
负责人:GUERGANA K. SAVOVA
-
依托单位:
Multi-source clinical Question Answering system
-
批准号:7842799
-
项目类别:
-
资助金额:$49.75万
-
财政年份:2009
-
负责人:GUERGANA K. SAVOVA
-
依托单位:
Multi-source clinical Question Answering system
-
批准号:7936991
-
项目类别:
-
资助金额:$49.14万
-
财政年份:2009
-
负责人:GUERGANA K. SAVOVA
-
依托单位:
国内基金
海外基金
基于Apache Spark的可扩展宏基因组序列组装方法研究
-
批准号:61802246
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:邓丽
-
依托单位: