Interactive machine learning methods for clinical natural language processing
Interactive machine learning methods for clinical natural language processing
批准号:
9132834
负责人:
HUA XU
金额:
$46.34万
依托单位国家:
美国
项目类别:
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-05-31 至 2018-09-28
关键词:
AbbreviationsActive LearningAddressAdoptionAlgorithmsAttentionBiomedical ResearchClassificationClinicalClinical DataClinical InformaticsClinical ResearchCognitiveCommunitiesDataData SetDevelopmentDiseaseEducational workshopElectronic Health RecordFaceGoalsGrantHumanHybridsKnowledgeLabelLearningLinguisticsMachine LearningManualsMedicalMethodologyMethodsModelingNamesNatural Language ProcessingPatientsPatternPerformancePharmaceutical PreparationsPhysiciansProcessResearchResearch PersonnelResearch PriorityResourcesSamplingSourceSpecific qualifier valueStatistical MethodsStatistical ModelsSystemTechnologyTestingTextTimeUnited States National Library of Medicinebaseclinical applicationclinical phenotypecohortcomputer human interactioncomputerizedcostexperienceimprovedlearning strategymodel buildingmodel developmentnovelopen sourcereal world applicationstatisticssuccesstoolusability
中文摘要
点击翻译按钮获取中文摘要
英文摘要
DESCRIPTION (provided by applicant): Growing deployments of electronic health records (EHRs) systems have made massive clinical data available electronically. However, much of detailed clinical information of patients is embedded in narrative text and is not directly accessible for computerized clinical applications. Therefore, natural language processing (NLP) technologies, which can unlock information in narrative document, have received great attention in the medical domain. Current state-of-the-art NLP approaches often involve building probabilistic models. However, the wide adoption of statistical methods in clinical NLP faces two grand challenges: 1) the lack of large annotated clinical corpora; and 2) the lack of methodologies that can efficiently integrate linguistic and domain knowledge with statistical learning. High-performance statistical NLP methods rely on large scale and high quality annotations of clinical text, but it is time-consuming and costly to create large annotated clinica corpora as it often requires manual review by physicians. Moreover, the medical domain is knowledge intensive. To achieve optimal performance, probabilistic models need to leverage medical domain knowledge. Therefore, methods that can efficiently integrate domain and expert knowledge with machine learning processes to quickly build high-quality probabilistic models with minimum annotation cost would be highly desirable for clinical text processing.
In this study, we propose to investigate interactive machine learning (IML) methods to address the above challenges in clinical NLP. An IML system builds a classification model in an iterative process, which can actively select informative samples for annotation based on models built on previously annotated samples, thus reducing the annotation cost for model development. More importantly, an IML system also involves human inputs to the learning process (e.g., an expert can specify important features for a classification task based on domain knowledge). Thus, IML is an ideal framework for efficiently integrating rule-based (via domain experts specifying features) and statistics-based (via different learning algorithms) approaches to clinical NLP. To achieve our goal, we propose three specific aims. In Aim 1, we plan to investigate different aspects of IML for word sense disambiguation, including developing new active learning algorithms and conducting cognitive usability analysis for efficient feature annotation by users. To demonstrate the broad uses of IML, we further extend IML approaches to two other important clinical NLP classification tasks: named entity recognition and clinical phenoytping in Aim 2. Finally we propose to disseminate the IML methods and tools to the biomedical research community in Aim 3.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Leveraging Longitudinal Data and Informatics Technology to Understand the Role of Bilingualism in Cognitive Resilience, Aging and Dementia
-
批准号:10583170
-
项目类别:
-
资助金额:$141.18万
-
财政年份:2023
-
负责人:HUA XU
-
依托单位:
Detecting synergistic effects of pharmacological and non-pharmacological interventions for AD/ADRD
-
批准号:10501245
-
项目类别:
-
资助金额:$78.15万
-
财政年份:2022
-
负责人:HUA XU
-
依托单位:
Engagement and outreach to achieve a FAIR data ecosystem for the BICAN
-
批准号:10523908
-
项目类别:
-
资助金额:$97.37万
-
财政年份:2022
-
负责人:HUA XU
-
依托单位:
Interactive machine learning methods for clinical natural language processing
-
批准号:8818096
-
项目类别:
-
资助金额:$55.84万
-
财政年份:2010
-
负责人:HUA XU
-
依托单位:
Real-time Disambiguation of Abbreviations in Clinical Notes
-
批准号:8077875
-
项目类别:
-
资助金额:$37.4万
-
财政年份:2010
-
负责人:HUA XU
-
依托单位:
Real-time Disambiguation of Abbreviations in Clinical Notes
-
批准号:7866149
-
项目类别:
-
资助金额:$38.75万
-
财政年份:2010
-
负责人:HUA XU
-
依托单位:
Real-time Disambiguation of Abbreviations in Clinical Notes
-
批准号:8589822
-
项目类别:
-
资助金额:$23.79万
-
财政年份:2010
-
负责人:HUA XU
-
依托单位:
Real-time Disambiguation of Abbreviations in Clinical Notes
-
批准号:8305149
-
项目类别:
-
资助金额:$12.9万
-
财政年份:2010
-
负责人:HUA XU
-
依托单位:
An in-silico method for epidemiological studies using Electronic Medical Records
-
批准号:8110041
-
项目类别:
-
资助金额:$25.23万
-
财政年份:2009
-
负责人:HUA XU
-
依托单位:
An in-silico method for epidemiological studies using Electronic Medical Records
-
批准号:7726747
-
项目类别:
-
资助金额:$27.33万
-
财政年份:2009
-
负责人:HUA XU
-
依托单位:
An insilico method for epidemiological studies using Electonic Medical Records
-
批准号:8589201
-
项目类别:
-
资助金额:$19.58万
-
财政年份:2009
-
负责人:HUA XU
-
依托单位:
An in-silico method for epidemiological studies using Electronic Medical Records
-
批准号:8298614
-
项目类别:
-
资助金额:$5.66万
-
财政年份:2009
-
负责人:HUA XU
-
依托单位:
海外基金