Transfer Learning for Digital Curation of the EMR Clinical Narrative
Transfer Learning for Digital Curation of the EMR Clinical Narrative
批准号:
10647748
负责人:
GUERGANA K. SAVOVA
金额:
$37.61万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-08-12 至 2025-05-31
关键词:
Adverse eventApacheAppearanceAreaArtificial IntelligenceAutomated AnnotationBiomedical ResearchBrain AneurysmsClassificationClinicalClinical InvestigatorCommunitiesComputer Vision SystemsComputerized Medical RecordCoupledCreativenessDataData ScienceData SetDatabasesDevelopmentDiseaseEngineeringEpigenetic ProcessEvaluationEventExplosionGeneticHealthHepatotoxicityImageInformation RetrievalIntuitionInvestigationLabelLanguageLearningMachine LearningMedicineMethodologyMethodsMethotrexateModelingMultiple SclerosisNatural Language ProcessingNeural Network SimulationPatientsPerformancePharmaceutical PreparationsPhenotypeProceduresPublic HealthPublicationsPublishingRare DiseasesReportingResearchResourcesRheumatoid ArthritisRunningSemanticsSeveritiesSourceSpeechStructureSymptomsSystemTechniquesTestingTextTimeTrainingTranslational ResearchVisionWorkautism spectrum disorderbasecomorbiditycomputing resourcesdeep learningdeep neural networkdigitalimaging Segmentationimprovedinformatics toollearning strategymachine learning methodmultitaskneuralneural networknext generationnovelopen sourcephenotypic dataresponsesupport vector machinetransfer learningunstructured data
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Project Summary
This proposal is in response to PAR 18-796 to seek support for advancing methodologies for a transfer
learning framework for the digital curation of the Electronic Medical Records (EMR) clinical narrative. In the
current era of increasing importance of Artificial Intelligence (AI) in biomedicine, our proposal tackles a critical
AI component – automated annotation of health-related text. Since 2015 the development and application of
machine learning (ML) methods has exploded propelled by the convergence of plentiful digitized unstructured
data (text, speech, images), hardware and the refinement of neural networks or deep learning. 2018 marked a
turning point in Natural Language Processing (NLP), particularly transfer learning through pre-trained models
like Universal Language Model Fine-tuning for Text Classification, Allen AI's ELMO, OpenAI's Open-GPT. In
November 2018, Google published the Bidirectional Encodings Representations from Transformers (BERT), a
transformer-based model pre-trained on massive general text databases (3.3B words total). The publication
reported using BERT representations to build classifiers for 11 NLP tasks which outperformed the state-of-the-
art (SOTA) with large margins. The NLP research community jumped to the idea of exploring this new
framework but quickly came to the realization that building BERT-style models from scratch is affordable and
feasible to only a few. Thus, research investigation proceeded in the direction of using these gigantic models
as resources for language representations. Scientific efforts focused on pre-trained models (e.g. BERT) as a
source of extracting high quality language features or fine-tuning on a specific task, i.e. using a model as a
checkpoint and re-training with much smaller amounts of task-specific data to produce predictions by typically
adding one fully-connected layer on top of the representations and training for a few epochs. This general
watershed shift in NLP to transfer learning which parallels the developments in computer vision a few years
ago coupled with our latest work brings to the forefront a critical NLP research topic ripe for exploration – a
transfer learning framework for the digital curation of the EMR clinical narrative. The proposed work is research
of novel scientific methods for extracting detailed information from health-related text especially the EMR, the
major source of phenotype data for patients. Precise phenotype information is needed to advance translational
research, particularly to unravel the effects of genetic, epigenetic, and systems changes on responsiveness.
This research is in line with the latest developments in neural deep learning approaches and AI in general and
is expected to enhance biomedical research and through that the health of the public.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
Improving the Transferability of Clinical Note Section Classification Models with BERT and Large Language Model Ensembles
使用 BERT 和大型语言模型集成提高临床记录部分分类模型的可迁移性
DOI:
10.18653/v1/2023.clinicalnlp-1.16
发表时间:
2023
期刊:
Clinical Natural Language Processing Workshop
影响因子:
--
作者:
[Weipeng Zhou, M. Afshar, Dmitriy Dligach, Yanjun Gao, Timothy Miller]
通讯作者:
Timothy Miller
DOI:
10.1038/s41746-023-00970-0
发表时间:
2024-01-11
期刊:
NPJ digital medicine
影响因子:
15.2
作者:
[]
通讯作者:
Transfer Learning for Digital Curation of the EMR Clinical Narrative
-
批准号:10092340
-
项目类别:
-
资助金额:$37.61万
-
财政年份:2021
-
负责人:GUERGANA K. SAVOVA
-
依托单位:
Transfer Learning for Digital Curation of the EMR Clinical Narrative
-
批准号:10468604
-
项目类别:
-
资助金额:$37.61万
-
财政年份:2021
-
负责人:GUERGANA K. SAVOVA
-
依托单位:
Cancer Deep Phenotype Extraction from Electronic Medical Records
-
批准号:9538366
-
项目类别:
-
资助金额:$56.79万
-
财政年份:2014
-
负责人:GUERGANA K. SAVOVA
-
依托单位:
Multi-source clinical Question Answering system
-
批准号:7842799
-
项目类别:
-
资助金额:$49.75万
-
财政年份:2009
-
负责人:GUERGANA K. SAVOVA
-
依托单位:
Multi-source clinical Question Answering system
-
批准号:7936991
-
项目类别:
-
资助金额:$49.14万
-
财政年份:2009
-
负责人:GUERGANA K. SAVOVA
-
依托单位:
国内基金
海外基金
基于Apache Spark的可扩展宏基因组序列组装方法研究
-
批准号:61802246
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:邓丽
-
依托单位: