课题基金 / 基金详情

POET-2: High-performance computing for advanced clinical narrative preprocessing

POET-2: High-performance computing for advanced clinical narrative preprocessing
POET-2:用于高级临床叙述预处理的高性能计算
批准号:
8326648
负责人:
JOHN F. HURDLE
金额:
$31.84万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-09-01 至 2014-08-31

项目摘要

项目成果

JOHN F. HURDLE的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供): 这个项目的重点是临床自然语言处理(CNLP),一个新兴的重要领域的信息学。从20世纪70年代语言字符串项目的医学语言处理器(纽约大学)开始,研究人员通过实证研究和构建复杂的高级cNLP软件应用程序(例如哥伦比亚的MedLEE),在cNLP方面取得了稳步进展。至少有四个专门讨论生物医学/临床NLP的科学会议。CNLP文献在过去十年中一直在增长,随着更多临床文本存储库的发布,如MIMIC II和匹兹堡大学Blu实验室语料库,这将获得势头。 然而,cNLP领域的持续成功受到这样一个现实的阻碍,即临床文本比NLP中传统研究的文本有更多的噪音,如新闻专线文章、生物医学摘要和出院摘要。本文中的噪声是由语言的可分析性特征和文本中出现的语言结构来定义的。临床文本有各种各样的笔记类型,研究最好的类型是出院总结、放射学报告和病理报告。这些笔记类型共享一个重要特征:它们是为了在医疗保健提供者之间交流护理问题而编写的,因此通常结构良好、编辑良好,并且通常是口述的。但电子健康记录中的绝大多数笔记主要是为了记录医疗问题。当然,他们也会进行沟通,但与出院总结和报告相比,他们在创作中使用的谨慎要少得多。因此,它们往往不符合语法;由简短的电报短语组成;充斥着拼写错误和速记(例如缩写);模板格式错误和空格的随意使用;并嵌入了“非散文”(例如实验室数值字符串)。所有这些噪声源都会使原本简单的NLP任务复杂化,比如标记化、句子分割以及最终的信息提取本身。 我们提出了一项系统的研究,以提高临床叙述中的信噪比,以改善cNLP。这项工作扩展了我们(在PEET项目下)的初步研究,并有以下目标: O开发和实施一套针对来自多个医疗机构的所有临床病历类型的可解析性改进工具。 O评估可解析性改进工具的经验和功能上的成功。 O设计并实施HIPAA兼容的基于ULMA的流水线cNLP框架,用于典型的高性能、多处理器计算环境。
英文摘要
DESCRIPTION (provided by applicant): This project focuses on clinical natural language processing (cNLP), a field of emerging importance in informatics. Starting with the Linguistic String Project's Medical Language Processor (New York University) in the 1970s, researchers have made steady gains in cNLP through empirical studies and by building sophisticated high-level cNLP software applications (e.g., Columbia's MedLEE). There are no fewer than four scientific conferences devoted exclusively to biomedical/clinical NLP. The cNLP literature has been growing over the past decade, and this will gain momentum as more clinical text repositories are released, such as the MIMIC II and University of Pittsburgh BLU Lab corpora. However, sustained success in the field of cNLP is hampered by the reality that clinical texts have a far more noise than do texts traditionally studied in NLP, such as newswire articles, biomedical abstracts, and discharge summaries. Noise in this context is defined by the parseability characteristics of the language and the linguistic structures that appear in text. Clinical texts come in a striking variety of note types, with the best studied types being discharge summaries, radiology reports, and pathology reports. These note types share an important feature: they are written to communicate care issues between healthcare providers and hence typically are well-composed, well-edited, and often are dictated. But the vast majority of notes in the electronic health record are written primarily to document care issues. They communicate as well, of course, but much less care is used in their creation than with discharge summaries and reports. As a result they are often ungrammatical; are composed of short, telegraphic phrases; are replete with misspellings and shorthand (e.g., abbreviations); are ill-formatted with templates and liberal use of white space; and are embedded with "non-prose" (e.g., strings of laboratory values). All of these sources of noise complicate otherwise straightforward NLP tasks like tokenization, sentence segmentation, and ultimately information extraction itself. We propose a systematic study of ways to increase the signal-to-noise ratio in clinical narratives to improve cNLP. This work extends our preliminary research (under the POET project) and has the following aims: o Develop and implement a suite of parseability improvement tools designed for all clinical note types from multiple healthcare institutions. o Evaluate the empirical and the functional success of the parseability improvement tools. o Design and implement a HIPAA-compliant UlMA-based pipeline cNLP framework for use in a typical high-performance, multi-processor computing environment.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
University of Utah Biomedical Informatics Training Grant Supplement
  • 批准号:
    9380137
  • 项目类别:
  • 资助金额:
    $0.24万
  • 财政年份:
    2016
  • 负责人:
    JOHN F. HURDLE
  • 依托单位:
POET-2: High-performance computing for advanced clinical narrative preprocessing
  • 批准号:
    8182025
  • 项目类别:
  • 资助金额:
    $32.52万
  • 财政年份:
    2011
  • 负责人:
    JOHN F. HURDLE
  • 依托单位:
POET: Consolidated, Comprehensive Clinical Text Preprocessing
  • 批准号:
    7570254
  • 项目类别:
  • 资助金额:
    $16.93万
  • 财政年份:
    2008
  • 负责人:
    JOHN F. HURDLE
  • 依托单位:
POET: Consolidated, Comprehensive Clinical Text Preprocessing
  • 批准号:
    7689273
  • 项目类别:
  • 资助金额:
    $16.66万
  • 财政年份:
    2008
  • 负责人:
    JOHN F. HURDLE
  • 依托单位:
海外基金