课题基金 / 基金详情

Open Health Natural Language Processing Collaboratory

Open Health Natural Language Processing Collaboratory
开放健康自然语言处理合作实验室
批准号:
10005506
负责人:
Xiaoqian Jiang
金额:
$150.08万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-09-01 至 2022-08-31

项目摘要

项目成果

Xiaoqian Jiang的其他基金

相似基金

相关文献

中文摘要
翻译
项目摘要 将电子健康记录(EHR)数据用于临床和翻译的主要障碍之一 科学是广泛使用非结构化或半结构化的临床叙述来记录临床 信息。自然语言处理(NLP)从叙事中提取结构化信息,具有 得到了极大的关注,并在使EHR二次用于临床和治疗方面发挥了关键作用 翻译研究。ACT(临床患者应计)等大规模努力证明了这一点 试验)、Emerge和PCORnet,使用EHR数据进行研究依赖于强大的数据和 信息学基础设施,允许构建临床叙述并支持临床描述的提取 下游应用程序的信息。当前成功的NLP使用案例通常需要强大的信息学知识 团队(与NLP专家一起)与临床医生合作,提供他们的领域知识并构建定制的NLP 引擎反复使用。这需要NLP专家和临床医生之间的密切合作,但在 信息学支持有限的机构。此外,的可用性、可移植性和通用性 NLP系统仍然有限,部分原因是各机构无法接触企业人力资源管理人员,以培训 系统。EHR数据的有限限制了可用于提高员工能力的培训 在临床NLP中。我们的目标是通过扩大我们现有的合作伙伴关系来应对上述挑战 多个CTSA中心关于开放健康自然语言处理(OHNLP),以共享 NLP构件(即单词、n元语法、短语、句子、概念提及、概念和文本片段) 从多个机构的真实EHR获得。我们将利用先进的隐私保护 IDASH的计算基础设施(集成数据以供分析、匿名和共享)以实现隐私- 保留数据分析模型,并将与包括观察性健康数据在内的不同社区合作 科学与信息学(OHDSI)、精确医学倡议(PMI)、PCORnet和罕见疾病临床 研究网络(RDCRN),以演示自然语言处理在翻译研究中的效用。这项CTSA创新 大奖RFA为我们提供了一个独特的机会来应对临床NLP和 通过与多个研究社区的强大合作伙伴关系和研究团队在 临床NLP,我们预计该项目的成功交付将扩大临床NLP的应用范围 在整个研究界。计划有四个目标:i)通过以下方式获得PHI抑制的NLP伪像 保留跨多个机构的分发信息,并评估访问PHI的隐私风险 被抑制的人工产物,ii)生成用于临床叙述的探索性分析的合成文本语料库 评估其在利用各种NLP挑战的NLP任务中的效用,iii)开发隐私保护计算 支持NLP的表型模型,以及iv)与不同的社区合作演示该实用程序 我们的翻译研究项目。
英文摘要
Project Summary One of the major barriers in leveraging Electronic Health Record (EHR) data for clinical and translational science is the prevalent use of unstructured or semi-structured clinical narratives for documenting clinical information. Natural Language Processing (NLP), which extracts structured information from narratives, has received great attention and has played a critical role in enabling secondary use of EHRs for clinical and translational research. As demonstrated by large scale efforts such as ACT (Accrual of patients for Clinical Trials), eMERGE, and PCORnet, using EHR data for research rests on the capabilities of a robust data and informatics infrastructure that allows the structuring of clinical narratives and supports the extraction of clinical information for downstream applications. Current successful NLP use cases often require a strong informatics team (with NLP experts) to work with clinicians to supply their domain knowledge and build customized NLP engines iteratively. This requires close collaboration between NLP experts and clinicians, not feasible at institutions with limited informatics support. Additionally, the usability, portability, and generalizability of the NLP systems are still limited, partially due to the lack of access to EHRs across institutions to train the systems. The limited availability of EHR data limits the training available to improve the workforce competence in clinical NLP. We aim to address the above challenges by extending our existing collaboration among multiple CTSA hubs on open health natural language processing (OHNLP) to share distributional information of NLP artifacts (i.e., words, n-grams, phrases, sentences, concept mentions, concepts, and text segments) acquired from real EHRs across multiple institutions. We will leverage the advanced privacy-preserving computing infrastructure of iDASH (integrating Data for Analysis, Anonymization, and SHaring) for privacy- preserving data analysis models and will partner with diverse communities including Observational Health Data Sciences and Informatics (OHDSI), Precision Medicine Initiative (PMI), PCORnet, and Rare Diseases Clinical Research Network (RDCRN) to demonstrate the utility of NLP for translational research. This CTSA innovation award RFA provides us with a unique opportunity to address the challenges faced with clinical NLP and through strong partnership with multiple research communities and leadership roles of the research team in clinical NLP, we envision that the successful delivery of this project will broaden the utilization of clinical NLP across the research community. There are four aims planned: i) obtain PHI-suppressed NLP artifacts with retained distribution information across multiple institutions and assess the privacy risk of accessing PHI- suppressed artifacts, ii) generate a synthetic text corpus for exploratory analysis of clinical narratives and assess its utility in NLP tasks leveraging various NLP challenges, iii) develop privacy-preserving computational phenotyping models empowered with NLP, and iv) partner with diverse communities to demonstrate the utility of our project for translational research.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Harmonizing multiple clinical trials for Alzheimer's disease to investigate differential responses to treatment via federated counterfactual learning
Robust privacy preserving distributed analysis platform for cancer research: addressing data bias and disparities
  • 批准号:
    10642562
  • 项目类别:
  • 资助金额:
    $41.19万
  • 财政年份:
    2023
  • 负责人:
    Xiaoqian Jiang
  • 依托单位:
iDASH Genome Privacy and Security Competition Workshop
Decentralized differentially-private methods for dynamic data release and analysis
  • 批准号:
    10740597
  • 项目类别:
  • 资助金额:
    $61.37万
  • 财政年份:
    2023
  • 负责人:
    Xiaoqian Jiang
  • 依托单位:
海外基金