课题基金 / 基金详情

Automated domain adaptation for clinical natural language processing

Automated domain adaptation for clinical natural language processing
临床自然语言处理的自动领域适应
批准号:
9768545
负责人:
Timothy A Miller
金额:
$38.39万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-09-01 至 2021-07-31

项目摘要

项目成果

Timothy A Miller的其他基金

相似基金

相关文献

中文摘要
翻译
项目摘要 从临床文本中自动提取有用信息使新的临床研究任务成为可能 和新技术在护理点。自然语言处理(NLP)系统 依靠有监督的机器学习来执行此提取。学习过程使用 手动标记大小和范围受限的数据集,因此应用NLP 系统到不可见的数据集通常会导致性能严重下降。获得更大的收益 而更广泛的数据集不太可能,因为手动标注过程和 难以在多个不同机构之间共享文本数据。因此,这个项目 开发无监督的域自适应算法以使NLP系统适应新数据。 领域适应描述了使机器学习系统适应新数据的过程 消息来源。建议的方法是无监督的,因为它们不需要手动标记 新的数据。 这个项目有三个目标。第一个目标是使用多个现有数据集来实现相同的目标 任务是研究域中的差异,并使用这些信息开发新的域 自适应算法。评估使用标准的机器学习度量,并分析 业绩受到严格的底线和现实上限的约束,两者都是如此 基于机器学习泛化的理论研究。第二个目标是发展 开放源码软件工具,可简化将领域适应合并到 临床文本处理工作流。该软件将有输入接口以连接到方法 在AIM 1中开发并输出接口,以与广泛使用的开放的. 源NLP工具。AIM 3在端到端的用例--药物不良事件中测试了这些方法 (ADE)对儿童肺动脉高压笔记的数据集进行提取。ADE提取依赖于 在多个NLP系统上,因此此用例能够展示对NLP的广泛改进 方法可以改进下游方法。这一目标还将为 端到端评估的数据集,可直接衡量NLP的改进情况 系统导致ADE提取的改进。
英文摘要
Project Summary Automatic extraction of useful information from clinical texts enables new clinical research tasks and new technologies at the point of care. The natural language processing (NLP) systems that perform this extraction rely on supervised machine learning. The learning process uses manually labeled datasets that are limited in size and scope, and as a result, applying NLP systems to unseen datasets often results in severely degraded performance. Obtaining larger and broader datasets is unlikely due to the expense of the manual labeling process and the difficulty of sharing text data between multiple different institutions. Therefore, this project develops unsupervised domain adaptation algorithms to adapt NLP systems to new data. Domain adaptation describes the process of adapting a machine learning system to new data sources. The proposed methods are unsupervised in that they do not require manual labels for the new data. This project has three aims. The first aim makes use of multiple existing datasets for the same task to study the differences in domains, and uses this information to develop new domain adaptation algorithms. Evaluation uses standard machine learning metrics, and analysis of performance is tightly bounded by strong baselines from below and realistic upper bounds, both based on theoretical research on machine learning generalization. The second aim develops open source software tools to simplify the process of incorporating domain adaptation into clinical text processing workflows. This software will have input interfaces to connect to methods developed in Aim 1 and output interfaces to connect with Apache cTAKES, a widely used open- source NLP tool. Aim 3 tests these methods in an end-to-end use case, adverse drug event (ADE) extraction on a dataset of pediatric pulmonary hypertension notes. ADE extraction relies on multiple NLP systems, so this use case is able to show how broad improvements to NLP methods can improve downstream methods. This aim also creates new manual labels for the dataset for an end-to-end evaluation that directly measures how improvements to the NLP systems lead to improvement in ADE extraction.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Learning Universal Patient Representations with Hierarchical Transformers
  • 批准号:
    10587270
  • 项目类别:
  • 资助金额:
    $64.82万
  • 财政年份:
    2019
  • 负责人:
    Timothy A Miller
  • 依托单位:
Bone Tissue Engineering Using Mineralized Collagen-GAG Scaffolds
Bone Tissue Engineering Using Mineralized Collagen-GAG Scaffolds
海外基金