DeconDTN: Deconfounding Deep Transformer Networks for Clinical NLP
DeconDTN: Deconfounding Deep Transformer Networks for Clinical NLP
批准号:
10626888
负责人:
Trevor Cohen
金额:
$34.2万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-06-01 至 2026-02-28
关键词:
AddressArchitectureAreaAutomobile DrivingBehaviorBridge to Artificial IntelligenceCOVID-19CaringCategoriesCharacteristicsClassificationClinicalClinical ServicesCognitiveComputer softwareConfounding Factors (Epidemiology)CoupledDataData AggregationData SetData SourcesDementiaDevelopmentDiagnosisDiagnosticEnsureEquilibriumEvaluationGenetic TranscriptionGoalsHigh PrevalenceIndividualInstitutionInvestmentsLabelLanguageLearningLinguisticsLocationMedicalMethodsModelingModificationNatural Language ProcessingNatureNeural Network SimulationOutcomeOutputParticipantPatientsPerformancePhysiciansPrevalenceResearchSARS-CoV-2 positiveSamplingServicesSiteSourceSpeechSystematic BiasTestingTextTimeTrainingTranscriptUnited States National Institutes of HealthUnited States National Library of MedicineUpdateVisionWeightWorkartificial intelligence methodcoronavirus diseasedeep learningdeep learning modeldesignheterogenous datainterestlarge datasetslearning strategyloss of functionmachine learning modelnetwork modelsneuralnovelopen sourceopen source toolportabilitypredictive modelingprogramsstatistical and machine learning
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Natural Language Processing (NLP) methods have been broadly applied to clinical problems, from recognition
of clinical findings in physician notes to identification of transcribed speech samples indicating changes in
cognitive status. Deep transformer networks (DTNs) have dramatically advanced NLP accuracy. These deep
learning models have multiple hidden layers that may correspond to billions of trainable parameters, allowing
them to apply information learned from training on large unlabeled corpora to a specific task of interest. However,
their size leaves them especially vulnerable to confounding bias, induced by variables that can influence both
the predictor (text) and the outcome (e.g. an associated diagnosis) of a predictive model. Such systematic biases
are a recognized danger in the application of artificial intelligence methods to clinical problems, and are the focus
of NLM NOT-LM-19-003 which invites applications proposing methods to identify and address them. Deep
learning models in general require large amounts of training data, spurring initiatives to aggregate medical data
from across institutional siloes. This can increase data set size and enhance model portability, but leaves the
resulting models vulnerable to confounding by provenance, where models learn to recognize the origin of dataset
components and make biased predictions based on site-specific class distributions (e.g. COVID prevalence).
Such models will assign classes based on indicators of dataset provenance, rather than diagnostically
meaningful linguistic differences, and make erroneous predictions when the provenance-specific distributions at
the point of deployment differ from those in the training set. Confounding of this nature is a pervasive problem
that presents a fundamental barrier to the portability of trained models, and threatens the utility of datasets
assembled from across institutions and services. Unlike traditional statistical and machine learning models, with
deep transformer networks feature representations are distributed across parameters spread throughout the
entire network. New methods are needed to meet the challenge of identifying and mitigating the influence of
confounding variables in such models. In the proposed research we will develop a systematic approach to
Deconfounding Deep Transformer Networks (DeconDTN), embodied in an eponymous and publicly available
set of open source tools for (1) identification of provenance-related biases, (2) mitigation of these biases using
a novel set of validated methods, and (3) systematic evaluation of the resulting effects on model performance.
While DeconDTN will be generally applicable, development and evaluation will occur in the context of three use
cases involving data sets drawn from different sources: classification of speech transcripts from participants with
dementia drawn from two locations, identification of goals-of-care discussions in clinical notes drawn from
multiple studies involving a range of clinical services, and prediction of COVID-19 status in notes drawn from
different clinical units. Our driving hypothesis is that the resulting models will make more accurate predictions in
these heterogenous datasets than corresponding models without correction for confounding by provenance.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Professional to Plain Language Neural Translation: A Path Toward Actionable Health Information
-
批准号:10349319
-
项目类别:
-
资助金额:$19.04万
-
财政年份:2022
-
负责人:Trevor Cohen
-
依托单位:
Professional to Plain Language Neural Translation: A Path Toward Actionable Health Information
-
批准号:10579898
-
项目类别:
-
资助金额:$21.16万
-
财政年份:2022
-
负责人:Trevor Cohen
-
依托单位:
DeconDTN: Deconfounding Deep Transformer Networks for Clinical NLP
-
批准号:10467107
-
项目类别:
-
资助金额:$34.53万
-
财政年份:2022
-
负责人:Trevor Cohen
-
依托单位:
DeconDTN: Deconfounding Deep Transformer Networks for Clinical NLP
-
批准号:10711315
-
项目类别:
-
资助金额:$31.12万
-
财政年份:2022
-
负责人:Trevor Cohen
-
依托单位:
Computerized assessment of linguistic indicators of lucidity in Alzheimer's Disease dementia
-
批准号:10093304
-
项目类别:
-
资助金额:$44.26万
-
财政年份:2020
-
负责人:Trevor Cohen
-
依托单位:
Using Biomedical Knowledge to Identify Plausible Signals for Pharmacovigilance
-
批准号:8914098
-
项目类别:
-
资助金额:$16.0万
-
财政年份:2013
-
负责人:Trevor Cohen
-
依托单位:
Using Biomedical Knowledge to Identify Plausible Signals for Pharmacovigilance
-
批准号:8727094
-
项目类别:
-
资助金额:$30.26万
-
财政年份:2013
-
负责人:Trevor Cohen
-
依托单位:
Encoding Semantic Knowledge in Vector Space for Biomedical Information
-
批准号:8138564
-
项目类别:
-
资助金额:$18.0万
-
财政年份:2010
-
负责人:Trevor Cohen
-
依托单位:
Encoding Semantic Knowledge in Vector Space for Biomedical Information
-
批准号:7977263
-
项目类别:
-
资助金额:$22.15万
-
财政年份:2010
-
负责人:Trevor Cohen
-
依托单位:
海外基金