National NLP Clinical Challenges (n2c2): Challenges in Natural Language Processing for Clinical Narratives
National NLP Clinical Challenges (n2c2): Challenges in Natural Language Processing for Clinical Narratives
批准号:
9759499
负责人:
Ozlem Uzuner
金额:
$2.0万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-05-15 至 2024-04-30
关键词:
Access to InformationAddressAmericanClinicClinicalCollaborationsCommunitiesCommunity DevelopmentsComplementDataData ScienceData SetDevelopmentEducational workshopElectronic Health RecordEvaluationFosteringFutureFuture GenerationsGoalsGoldGrantGrowthHandHeadHealthcareImprove AccessIndividualInformaticsInstitutionIsraelJournalsKnowledgeMeasuresMedical InformaticsMedical centerMethodologyNatural Language ProcessingOutcomePaperPeer ReviewPerformancePrivacyPublicationsPublishingRecordsResearchResearch PersonnelRestRunningSeriesSourceStructureSystemSystems DevelopmentTargeted ResearchTechnologyTimeUnited States National Institutes of HealthUniversitiesbasebiomedical informaticsclinical developmentcomputerizedfallshead-to-head comparisonhealth dataindexingmedical schoolsmeetingspractical applicationprogramssymposiumusabilityworking group
中文摘要
项目摘要和摘要
电子健康记录(EHR)的叙述包含难以自动提取的有用信息,
索引、搜索或解释。自然语言处理(NLP)技术可以提取这些信息并
将其转换为计算机化系统更容易访问的结构化格式。然而,
NLP系统的开发取决于能否获得相关数据,众所周知,EHR很难获得
因为隐私的原因。尽管最近努力去识别和发布用于研究的叙述性EHR,
这些数据仍然非常罕见。因此,临床NLP作为一个领域已经相对滞后。为了解决这个问题,
自2006年以来,我们组织了13项共同任务,并辅之以研讨会和期刊出版物。十二
在这些分担的任务中,重点是临床NLP系统的开发,其余的任务是
这些系统的可用性。我们在共享任务方面涵盖了深度和广度,准备任务
它研究了来自多个机构的各种电子病历数据的尖端NLP问题。我们的共同任务是
运行时间最长的临床NLP共享任务系列,具有不断增长的EHR数据集、任务和参与度。
我们最受欢迎的三个数据集被引用了495(2010年数据)、284(2006年去ID数据)和274(2009年数据)
时间分别代表了仅这三个数据集就产生的数百篇文章。我们的
该提案的目标是继续我们2006年在i2b2共同任务挑战(i2b2,NIH)下开始的努力
NLMU54LM008748,PI:Kohane和R13 LM011411,PI:Zuuner)来识别EHR,用金色注释-
临床NLP任务的标准注释,并将其发布给研究社区以供开发
以及临床NLP系统的面对面比较,以促进最先进的技术水平。继续
我们在国家NLP临床挑战(N2c2)下的努力基于美国卫生数据科学计划
哈佛医学院新成立的生物医学信息学系,我们的目标是建立合作伙伴关系
与社区一起通过以下几种方式增加共享任务的努力:(1)增加可用的已取消确认的EHR数据
通过合作伙伴关系为数据量和种类做出贡献,以及(2)扩大可用的
在NLP任务的深度和广度方面的黄金标准注释。鉴于这些目标和伙伴关系,我们
计划举办一系列共享任务。我们将通过研讨会补充这些共同的任务,这些研讨会将在
与美国医学信息学协会秋季研讨会和《特殊》杂志合作
这些问题可以加快最先进技术的进步,后代可以在过去的基础上再接再厉。
英文摘要
Project Summary and Abstract
Narratives of electronic health records (EHRs) contain useful information that is difficult to automatically extract,
index, search, or interpret. Natural language processing (NLP) technologies can extract this information and
convert it in to a structured format that is more readily accessible by computerized systems. However, the
development of NLP systems is contingent on access to relevant data and EHRs are notoriously difficult to obtain
because of privacy reasons. Despite the recent efforts to de-identify and release narrative EHRs for research,
these data are still very rare. As a result, clinical NLP, as a field has lagged behind. To address this problem,
since 2006, we organized thirteen shared tasks, accompanied with workshops and journal publications. Twelve
of these shared tasks have focused on the development of clinical NLP systems and the remaining one on the
usability of these systems. We have covered both depth and breadth in terms of shared tasks, preparing tasks
that study cutting-edge NLP problems on a variety of EHR data from multiple institutions. Our shared tasks are
the longest running series of clinical NLP shared tasks, with ever growing EHR data sets, tasks, and participation.
Our most popular three data sets have been cited 495 (2010 data), 284 (2006 de-id data), and 274 (2009 data)
times, respectively, representing hundreds of articles that have come out of these three data sets alone. Our
goal in this proposal is to continue the efforts we started in 2006 under i2b2 shared task challenges (i2b2, NIH
NLM U54LM008748, PI: Kohane and R13 LM011411, PI: Uzuner) to de-identify EHRs, annotate them with gold-
standard annotations for clinical NLP tasks, and release them to the research community for the development
and head-to-head comparison of clinical NLP systems, for the advancement of the state of the art. Continuing
our efforts under National NLP Clinical Challenges (n2c2) based at the Health Data Science program of the
newly established Department of Biomedical Informatics at Harvard Medical School, we aim to form partnerships
with the community to grow the shared task efforts in several ways: (1) grow the available de-identified EHR data
sets through partnerships that can contribute to the volume and variety of the data, and (2) grow the available
gold-standard annotations in terms of depth and breadth of NLP tasks. Given these aims and partnerships, we
plan to hold a series of shared tasks. We will complement these shared tasks with workshops that meet in
conjunction with the Fall Symposium of the American Medical Informatics Association and with journal special
issues so that advancement of the state of the art can be sped up and future generations can build on the past.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Joint learning methods for event and relation extraction from clinical narratives
-
批准号:10507223
-
项目类别:
-
资助金额:$42.49万
-
财政年份:2022
-
负责人:Ozlem Uzuner
-
依托单位:
National NLP Clinical Challenges (n2c2): Challenges in Natural Language Processing for Clinical Narratives
-
批准号:10670801
-
项目类别:
-
资助金额:$2.0万
-
财政年份:2019
-
负责人:Ozlem Uzuner
-
依托单位:
Leveraging Unlabeled and Pseudo Data for Clinical Information Extraction
-
批准号:9813134
-
项目类别:
-
资助金额:$41.48万
-
财政年份:2019
-
负责人:Ozlem Uzuner
-
依托单位:
National NLP Clinical Challenges (n2c2): Challenges in Natural Language Processing for Clinical Narratives
-
批准号:10393499
-
项目类别:
-
资助金额:$2.0万
-
财政年份:2019
-
负责人:Ozlem Uzuner
-
依托单位:
Challenges in Natural Language Processing in Clinical Text
-
批准号:9597333
-
项目类别:
-
资助金额:$2.0万
-
财政年份:2017
-
负责人:Ozlem Uzuner
-
依托单位:
Challenges in Natural Language Processing for Clinical Narratives
-
批准号:8722031
-
项目类别:
-
资助金额:$2.0万
-
财政年份:2012
-
负责人:Ozlem Uzuner
-
依托单位:
Challenges in Natural Language Processing for Clinical Narratives
-
批准号:8913773
-
项目类别:
-
资助金额:$1.98万
-
财政年份:2012
-
负责人:Ozlem Uzuner
-
依托单位:
Challenges in Natural Language Processing for Clinical Narratives
-
批准号:8400218
-
项目类别:
-
资助金额:$2.0万
-
财政年份:2012
-
负责人:Ozlem Uzuner
-
依托单位:
海外基金