Automated Extraction of Diagnostic Criteria From Electronic Health Records for Autism Spectrum Disorders: Development, Evaluation, and Application.

Automated Extraction of Diagnostic Criteria From Electronic Health Records for Autism Spectrum Disorders: Development, Evaluation, and Application.
复制标题

DOI:
10.2196/10497
复制
发表时间:
2018-11-07
影响因子:
7.4
通讯作者:
Kurzius-Spencer M
Kurzius-Spencer M
中科院分区:
医学2区
文献类型:
--
作者:
Leroy G;Gu Y;Pettygrove S;Galindo MK;Arora A;Kurzius-Spencer M

文献摘要

参考文献

被引文献

相似文献

电子健康档案(EHR)为信息利用带来了许多机会。其中一个用途是由疾病控制和预防中心进行的监测,以跟踪自闭症谱系障碍(ASD)的病例。该过程目前包括手动收集和审查美国11个州4岁和8岁儿童的EHR,以确定ASD标准的存在。这项工作既费时又费钱。我们的目标是从EHR中自动提取临床医生在精神疾病诊断和统计手册(DSM)中的诊断标准的证据中注意到的行为描述。以前,我们报告了整个EHR的分类ASD或没有。在这项工作中,我们专注于提取文本中不同ASD标准的个体表达。我们打算促进ASD的大规模监测工作,并支持随时间变化的分析,以及与其他相关数据的整合。我们开发了一个自然语言处理(NLP)解析器,使用104个模式和92个词典(1787个术语)提取12个DSM标准的表达式。解析器是基于规则的,可以从文本中精确提取实体。实体本身包含在EHR中,作为不同人在不同时间(临床医生,言语病理学家等)编写的诊断标准的非常不同的表达。由于数据的稀疏性,基于规则的方法最适合,直到可以为机器学习算法生成更大的数据集。我们评估了基于规则的解析器,并将其与机器学习基线(决策树)进行了比较。使用6636个句子(50个EHR)的测试集,我们发现我们的解析器实现了76%的准确率,43%的召回率(即灵敏度)和>99%的标准提取特异性。基于规则的方法的性能优于机器学习基线(60%的准确率和30%的召回率)。对于某些单独的标准,准确率高达97%,召回率为57%。由于精确度非常高,我们确信标准很少被错误地分配,我们的数字代表了它们在EHR中存在的下限。然后,我们进行了一项案例研究,并解析了4480个新的EHR,涵盖了亚利桑那州发育障碍监测计划10年的监测记录。社会标准(A1标准)显示了多年来最大的变化。交流标准(A2标准)没有区分ASD和非ASD记录。在行为和兴趣标准(A3标准)中,1(A3b)在ASD EHR中的出现频率远高于非ASD EHR。我们的研究结果表明,NLP可以支持对ASD监测和研究有用的大规模分析。今后,我们打算促进对国家数据集的详细分析和整合。
Electronic health records (EHRs) bring many opportunities for information utilization. One such use is the surveillance conducted by the Centers for Disease Control and Prevention to track cases of autism spectrum disorder (ASD). This process currently comprises manual collection and review of EHRs of 4- and 8-year old children in 11 US states for the presence of ASD criteria. The work is time-consuming and expensive. Our objective was to automatically extract from EHRs the description of behaviors noted by the clinicians in evidence of the diagnostic criteria in the Diagnostic and Statistical Manual of Mental Disorders (DSM). Previously, we reported on the classification of entire EHRs as ASD or not. In this work, we focus on the extraction of individual expressions of the different ASD criteria in the text. We intend to facilitate large-scale surveillance efforts for ASD and support analysis of changes over time as well as enable integration with other relevant data. We developed a natural language processing (NLP) parser to extract expressions of 12 DSM criteria using 104 patterns and 92 lexicons (1787 terms). The parser is rule-based to enable precise extraction of the entities from the text. The entities themselves are encompassed in the EHRs as very diverse expressions of the diagnostic criteria written by different people at different times (clinicians, speech pathologists, among others). Due to the sparsity of the data, a rule-based approach is best suited until larger datasets can be generated for machine learning algorithms. We evaluated our rule-based parser and compared it with a machine learning baseline (decision tree). Using a test set of 6636 sentences (50 EHRs), we found that our parser achieved 76% precision, 43% recall (ie, sensitivity), and >99% specificity for criterion extraction. The performance was better for the rule-based approach than for the machine learning baseline (60% precision and 30% recall). For some individual criteria, precision was as high as 97% and recall 57%. Since precision was very high, we were assured that criteria were rarely assigned incorrectly, and our numbers presented a lower bound of their presence in EHRs. We then conducted a case study and parsed 4480 new EHRs covering 10 years of surveillance records from the Arizona Developmental Disabilities Surveillance Program. The social criteria (A1 criteria) showed the biggest change over the years. The communication criteria (A2 criteria) did not distinguish the ASD from the non-ASD records. Among behaviors and interests criteria (A3 criteria), 1 (A3b) was present with much greater frequency in the ASD than in the non-ASD EHRs. Our results demonstrate that NLP can support large-scale analysis useful for ASD surveillance and research. In the future, we intend to facilitate detailed analysis and integration of national datasets.
DOI: 10.1371/journal.pcbi.1002854
发表时间: 2013
影响因子: 4.3
作者:
Cunningham H;Tablan V;Roberts A;Bontcheva K
通讯作者: Bontcheva K
DOI: 10.1016/j.jbi.2013.07.006
发表时间: 2013-10-01
影响因子: 4.5
作者:
Kwak, Myungjae;Leroy, Gondy;Harwell, Jeffrey
通讯作者: Harwell, Jeffrey
DOI: 10.1002/aur.1581
发表时间: 2016-08-01
期刊: AUTISM RESEARCH
影响因子: 4.7
作者:
Luo, Sean X.;Shinall, Jacqueline A.;Gerber, Andrew J.
通讯作者: Gerber, Andrew J.
DOI: 10.12688/f1000research.4591.1
发表时间: 2014-01-01
期刊: F1000Research
影响因子: --
作者:
Bagewadi, Shweta;Bobic, Tamara;Klinger, Roman
通讯作者: Klinger, Roman
DOI: 10.1111/jcpp.12864
发表时间: 2018-07-01
影响因子: 7.6
作者:
Arvidsson, Olof;Gillberg, Christopher;Lundstrom, Sebastian
通讯作者: Lundstrom, Sebastian