Automated domain adaptation for clinical natural language processing
Automated domain adaptation for clinical natural language processing
批准号:
9768545
负责人:
Timothy A Miller
金额:
$38.39万
依托单位国家:
美国
项目类别:
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-09-01 至 2021-07-31
关键词:
AdultAdverse drug eventAlgorithmsApacheAreaCharacteristicsChildhoodClinicalClinical InformaticsClinical ResearchColon CarcinomaCommunitiesComputer softwareComputersConceptionsDataData SetData SourcesDetectionDimensionsEcosystemEducational process of instructingElectronic Health RecordEvaluationHumanInstitutionKnowledgeLabelLanguageLeadLearningLinguisticsMachine LearningMalignant NeoplasmsMalignant neoplasm of brainManualsMeasurementMeasuresMedicalMethodsModelingNatural Language ProcessingNetwork-basedOutputPathologyPatientsPerformancePharmaceutical PreparationsPopulationProcessPulmonary HypertensionRadiology SpecialtyResearchSoftware ToolsSourceStatistical ModelsStructureSystemTestingTextTimeLineTrainingUpdateVisionWorkbasecase findingimprovedlearning strategymalignant breast neoplasmmethod developmentnatural languageneural networknew technologynewsnovelopen sourcepoint of careside effectsocial mediasoftware systemsstatisticssupervised learningtooltumorunsupervised learning
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Project Summary
Automatic extraction of useful information from clinical texts enables new clinical research tasks
and new technologies at the point of care. The natural language processing (NLP) systems that
perform this extraction rely on supervised machine learning. The learning process uses
manually labeled datasets that are limited in size and scope, and as a result, applying NLP
systems to unseen datasets often results in severely degraded performance. Obtaining larger
and broader datasets is unlikely due to the expense of the manual labeling process and the
difficulty of sharing text data between multiple different institutions. Therefore, this project
develops unsupervised domain adaptation algorithms to adapt NLP systems to new data.
Domain adaptation describes the process of adapting a machine learning system to new data
sources. The proposed methods are unsupervised in that they do not require manual labels for
the new data.
This project has three aims. The first aim makes use of multiple existing datasets for the same
task to study the differences in domains, and uses this information to develop new domain
adaptation algorithms. Evaluation uses standard machine learning metrics, and analysis of
performance is tightly bounded by strong baselines from below and realistic upper bounds, both
based on theoretical research on machine learning generalization. The second aim develops
open source software tools to simplify the process of incorporating domain adaptation into
clinical text processing workflows. This software will have input interfaces to connect to methods
developed in Aim 1 and output interfaces to connect with Apache cTAKES, a widely used open-
source NLP tool. Aim 3 tests these methods in an end-to-end use case, adverse drug event
(ADE) extraction on a dataset of pediatric pulmonary hypertension notes. ADE extraction relies
on multiple NLP systems, so this use case is able to show how broad improvements to NLP
methods can improve downstream methods. This aim also creates new manual labels for the
dataset for an end-to-end evaluation that directly measures how improvements to the NLP
systems lead to improvement in ADE extraction.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Learning Universal Patient Representations with Hierarchical Transformers
-
批准号:10587270
-
项目类别:
-
资助金额:$64.82万
-
财政年份:2019
-
负责人:Timothy A Miller
-
依托单位:
Bone Tissue Engineering Using Mineralized Collagen-GAG Scaffolds
-
批准号:8621974
-
项目类别:
-
资助金额:$0.0万
-
财政年份:2012
-
负责人:Timothy A Miller
-
依托单位:
Bone Tissue Engineering Using Mineralized Collagen-GAG Scaffolds
-
批准号:8440695
-
项目类别:
-
资助金额:$0.0万
-
财政年份:2012
-
负责人:Timothy A Miller
-
依托单位:
海外基金