Automating Delirium Identification and Risk Prediction in Electronic Health Records (Supplement)
Automating Delirium Identification and Risk Prediction in Electronic Health Records (Supplement)
批准号:
10410694
负责人:
RICHARD E KENNEDY
金额:
$26.65万
依托单位国家:
美国
项目类别:
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-02-15 至 2022-12-31
中文摘要
摘要
谵妄或急性意识模糊状态影响30-40%的住院老年人,估计额外的护理费用高达70亿美元。虽然最初被认为是一种短暂性疾病,但现在认识到谵妄具有重大后果,包括死亡风险增加,功能下降和长期认知障碍。由于高达75%的病例未被提供者识别,因此迫切需要先进的方法来识别临床和研究目的的谵妄,并根据谵妄风险对患者进行分层。不幸的是,谵妄的监测是耗时的,导致很少有机构对所有老年人实施系统的筛查程序。使用定期谵妄筛查评估的研究通常是在患者子集中进行的小型单机构研究,由于对隐私的担忧、缺乏技术专长或其他担忧,这些研究不会发布其数据进行复制。不使用谵妄筛查评估的流行病学研究依赖于代理措施,例如敏感性低至3%的行政代码,而不是可以恢复高达74%的谵妄病例的图表审查。自然语言处理(NLP)和机器学习(ML)等先进方法有可能使这种图表审查过程自动化,并促进对谵妄的大规模研究,但由于缺乏用于算法开发的合适数据而受到阻碍。我们建议利用亚拉巴马大学伯明翰分校(UAB)虚拟急性老年护理(ACE)质量改进计划提供的系统性谵妄筛查,创建并发布一个去识别的谵妄数据集,以解决谵妄流行病学研究中的数据和诊断差距。我们的虚拟ACE项目在六年的时间里已经确定了超过33,000名患者的谵妄状态,提供了一组丰富的数据,本项目将从中提取。我们将测试一个假设,即我们基于迁移学习的去识别方法可以帮助注释者更快地去识别临床文本,为更大、更快、更广泛的数据集发布打开大门。为了验证完整数据集的实用性,我们将确定我们的去识别语料库的统计能力,以检测在常用的ML研究设计中有和没有谵妄的参与者之间的差异。
我们的谵妄数据集发布,包含3,000个去识别的临床笔记和相关的结构数据,将是有史以来发布的最大的文本语料库之一,也是唯一一个专门用于谵妄研究的文本语料库。拟议的数据集将可通过数据使用协议(DUA)在Physionet上下载,以促进NLP和ML方法的进一步发展,从而通过迁移学习在其他机构和其他人群中确定谵妄状态、风险因素和后遗症。我们的去识别算法和方法的发布也将促进其他领域大规模文本语料库的开发和发布。
障碍
英文摘要
ABSTRACT
Delirium, or acute confusional state, affects 30-40% of hospitalized older adults, with the added cost of care estimated to be up to $7 billion. Although originally conceptualized as a transient disorder, delirium is now recognized to have significant consequences, including increased risk of death, functional decline, and long-term cognitive impairment. As up to 75% of cases are not recognized by providers, there is an critical need for advanced methods to identify delirium for clinical and research purposes, and to stratify patients based on delirium risk. Unfortunately, surveillance of delirium is time consuming, resulting in few institutions implementing systematic screening procedures on all older adults. Research studies that use regular delirium screening assessments are typically small, single institution studies in a subset of patients that do not release their data for replication due to concern over privacy, lack of technical expertise, or other concerns. Epidemiological studies that do not use delirium screening assessments rely on proxy measures, such administrative codes with sensitivity as low as 3%, instead of chart review which can recover up to 74% of delirium cases. Advanced methods such as natural language processing (NLP) and machine learning (ML) have the potential to automate this chart review process and facilitate large-scale studies of delirium, but are hampered by lack of suitable data for algorithm development. We propose to leverage systematic delirium screening available through the University of Alabama of Birmingham (UAB) Virtual Acute Care for Elders (ACE) quality improvement program to create and release a de-identified delirium dataset to address the data and diagnosis gap in epidemiological studies of delirium. Our Virtual ACE program has determined delirium status on more than 33,000 patients across a six-year period, providing a rich set of data from which this project will draw. We will test the hypothesis that our transfer learning based deidentification method can assist annotators to more rapidly de-identify clinical text, opening the door to larger, faster, more widely available dataset releases. To validate the utility of the full dataset, we will determine the statistical power of our de-identified corpus to detect differences between participants with and without delirium in commonly used ML study designs.
Our delirium dataset release, containing 3,000 de-identified clinical notes and associated structural data, will be one of the largest text corpora ever released and the only text inclusive corpora specifically for the study of delirium. The proposed dataset will be available for download on Physionet with a Data Use Agreement (DUA) to facilitate further development of NLP and ML approaches for determining delirium status, risk factors, and sequelae at other institutions and in other populations by transfer learning. Release of our de-identification algorithms and methodology will also facilitate the development and release of large-scale text corpora in other
disorders
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
Automating Delirium Identification and Risk Prediction in Electronic Health Records
-
批准号:10341053
-
项目类别:
-
资助金额:$37.69万
-
财政年份:2019
-
负责人:RICHARD E KENNEDY
-
依托单位:
Automating Delirium Identification and Risk Prediction in Electronic Health Records
-
批准号:10091381
-
项目类别:
-
资助金额:$37.69万
-
财政年份:2019
-
负责人:RICHARD E KENNEDY
-
依托单位:
In Silico Screening of Medications for Slowing Alzheimer's Disease Progression.
-
批准号:9884696
-
项目类别:
-
资助金额:$64.6万
-
财政年份:2017
-
负责人:RICHARD E KENNEDY
-
依托单位:
Mixed Effects Modeling of Microarrays Using the S-score
-
批准号:6935669
-
项目类别:
-
资助金额:$6.59万
-
财政年份:2005
-
负责人:RICHARD E KENNEDY
-
依托单位:
Mixed Effects Modeling of Microarrays Using the S-score
-
批准号:7121993
-
项目类别:
-
资助金额:$6.59万
-
财政年份:2005
-
负责人:RICHARD E KENNEDY
-
依托单位:
Mixed Effects Modeling of Microarrays Using the S-score
-
批准号:7272023
-
项目类别:
-
资助金额:$2.78万
-
财政年份:2005
-
负责人:RICHARD E KENNEDY
-
依托单位:
海外基金