Investigating the generalizability of natural language processing of EMR data
Investigating the generalizability of natural language processing of EMR data
批准号:
7691692
负责人:
BRIAN L HAZLEHURST
金额:
$17.78万
依托单位国家:
美国
项目类别:
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-09-30 至 2011-09-29
关键词:
AddressAdoptedAdoptionAffectAlcoholsArchitectureBehavioralCaringClassificationClinicalClinical DataCodeComplexComputer Systems DevelopmentComputerized Medical RecordCoughingCounselingDataDatabasesDevelopmentEventGoldHealthcare SystemsHuman ResourcesInformaticsInformation SystemsInformation TechnologyInstitute of Medicine (U.S.)KnowledgeLanguageLearningMeasuresMedical RecordsModificationNatural Language ProcessingPerformancePersonsPositioning AttributePrevalenceProcessProliferatingQuality of CareRecordsResearchRunningSolutionsSourceSpecificityStructureSymptomsSystemTechniquesTestingTextTimeTranslatingUnited States National Academy of SciencesUrsidae FamilyValidationVisionVocabularycare deliverycostdesignevaluation/testingexperiencehealth care qualitystemsuccesstooltool developmentweb site
中文摘要
描述(由申请人提供):
电子病历(EMR)为提高护理质量提供了令人印象深刻的机会,但挑战阻碍了这一愿景的实现。例如,易于分析的编码EMR数据通常是不完整的(由于EMR实施中普遍存在自由文本临床记录),而来自不同EMR的数据通常由于标准词汇表和系统实施的差异而不相称。虽然信息学研究已经显示了使用自然语言处理(NLP)对临床文本的特定方面进行自动编码的可行性,但将这些信息学的发展转化为大规模医疗质量评估仍然存在挑战。到目前为止,用于自动化质量评估的成功的NLP解决方案往往是特定于(A)目标问题或临床焦点、(B)EMR数据系统和(C)实施NLP解决方案的个人或团队的应用程序。在这项研究中,我们建议通过开发、评估和免费提供可推广的NLP开发工具套件来开始解决实现团队专用性的问题。这些工具将使NLP系统能够被广泛采用,以从免费文本临床记录中提取数据并对其进行编码。知识编辑工具包将帮助用户定义构成领域特定知识模块的规则、概念和术语,从而简化特定问题知识的开发,从而允许任何信息学家开发NLP应用程序。NLP应用程序验证工具包将允许根据来自任何EMR的独立编码测试记录的黄金标准对应用程序进行快速测试和评估。为了评估工具包对NLP普适性的影响,我们将让三名临床信息学家每人构建两个NLP应用程序(总共有六个不同的应用程序)。它们的应用之一将识别一系列常见的临床体征或症状(例如“持续性咳嗽”),这些体征或症状是相对独立的概念,使用简单的语言术语用于许多不同的临床目的。他们的第二个应用将评估行为咨询(例如,“酒精咨询”),它使用复杂的语言结构专门用于临床目的。我们将针对独立编码的医疗记录测试集描述和评估解决方案的准确性。我们将通过时间、构建应用程序所需的迭代次数以及使用的概念和规则的数量来量化和比较创建这些解决方案的难度,并分析创建的解决方案的内容和准确性的可变性。此外,我们将使用定性技术来评估开发工具的易用性;学习工具的难度;以及遇到的特定类型的问题、限制和错误。这样的NLP开发工具套件有可能提供简单、优雅和可靠的良好NLP解决方案,而不受临床问题领域或解决方案开发人员的影响。
英文摘要
DESCRIPTION (provided by applicant):
The electronic medical record (EMR) offers impressive opportunities for increasing care quality, but challenges stand in the way of realizing this vision. For example, coded EMR data readily available for analysis typically are incomplete (due to the prevalence of free-text clinical notes in EMR implementations), and data from different EMRs are often incommensurate due to differences in standard vocabularies and system implementations. While informatics research has shown the feasibility of automatically coding specific aspects of clinical text using Natural Language Processing (NLP), challenges remain for translating these informatics developments into large-scale care quality assessments. To date, successful NLP solutions for automated quality assessment have tended to be applications that are specific to (a) the target problem or clinical focus, (b) the EMR data system, and (c) the person or team that implements the NLP solution. In this study, we propose to begin addressing the problem of implementation team specificity by developing, evaluating, and making freely available a generalizable NLP development tool suite. The tools will enable widespread adoption of NLP systems to extract and code data from free text clinical notes. The Knowledge Editing Toolkit will simplify development of problem-specific knowledge by helping the user define the rules, concepts, and terms that constitute a domain-specific knowledge module, thus allowing any informaticist to develop an NLP application. The NLP Application Validation Toolkit will allow rapid testing and evaluation of the application against a gold standard of independently-coded test records from any EMR. To evaluate the effects of the toolkits on NLP generalizability, we will have three clinical informaticists each build two NLP applications (for a total of six distinct applications). One of their applications will identify a constellation of common clinical signs or symptoms (e.g., "persistent cough") that are relatively discrete concepts using simple language terms for many different clinical purposes. Their second application will assess behavioral counseling (e.g., "alcohol counseling"), which uses complex language constructs for dedicated clinical purposes. We will describe and evaluate the accuracy of the solutions against independently coded test sets of medical records. We will quantify and compare the difficulty of creating these solutions as measured by the time, number of iterations required to build the applications, and the number of concepts and rules employed, as well as analyze variability in content and accuracy of the solutions created. In addition, we will use qualitative techniques to assess the ease of using the development tools; the difficulty in learning the tools; and specific types of problems, limitations, and bugs encountered. Such an NLP development tool suite has the potential to allow simple, elegant, and reliably good NLP solutions regardless of the clinical problem domain or the person developing the solution.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Enhancing Clinical Effectiveness Research with Natural Language Processing of EMR
-
批准号:8032928
-
项目类别:
-
资助金额:$869.69万
-
财政年份:2010
-
负责人:BRIAN L HAZLEHURST
-
依托单位:
Investigating the generalizability of natural language processing of EMR data
-
批准号:7850343
-
项目类别:
-
资助金额:$10.0万
-
财政年份:2009
-
负责人:BRIAN L HAZLEHURST
-
依托单位:
Automating assessment of obesity care quality
-
批准号:7941068
-
项目类别:
-
资助金额:$49.61万
-
财政年份:2009
-
负责人:BRIAN L HAZLEHURST
-
依托单位:
Automating assessment of obesity care quality
-
批准号:8136907
-
项目类别:
-
资助金额:$24.79万
-
财政年份:2009
-
负责人:BRIAN L HAZLEHURST
-
依托单位:
Investigating the generalizability of natural language processing of EMR data
-
批准号:7529967
-
项目类别:
-
资助金额:$21.33万
-
财政年份:2008
-
负责人:BRIAN L HAZLEHURST
-
依托单位:
Automating Assessment of Asthma Care Quality
-
批准号:7355952
-
项目类别:
-
资助金额:$42.14万
-
财政年份:2007
-
负责人:BRIAN L HAZLEHURST
-
依托单位:
Automating Assessment of Asthma Care Quality
-
批准号:7498546
-
项目类别:
-
资助金额:$39.58万
-
财政年份:2007
-
负责人:BRIAN L HAZLEHURST
-
依托单位:
海外基金