Text Mining in Electronic Medical Records Enables Quick and Efficient Identification of Pregnancy Cases Occurring After Breast Cancer

Text Mining in Electronic Medical Records Enables Quick and Efficient Identification of Pregnancy Cases Occurring After Breast Cancer
复制标题

DOI:
10.1200/cci.19.00031
复制
发表时间:
2019-10-18
影响因子:
4.2
通讯作者:
Hamy, Anne-Sophie
Hamy, Anne-Sophie
中科院分区:
其他
文献类型:
--
作者:
Labrosse, Julie;Lam, Thanh;Hamy, Anne-Sophie

文献摘要

被引文献

相似文献

目的将文本挖掘(Text Mining,TM)技术应用于乳腺癌患者的电子病历(EMR)中,以检索乳腺癌诊断后妊娠的发生情况,并与人工治疗进行比较。材料与方法训练队列(队列A)包括2005-2007年间在居里研究所接受治疗的344名年龄在40岁以下的乳腺癌患者。人工指导包括手动审查每一份EMR以恢复妊娠。TM包括首先应用关键字筛选器(“accouch*”或“enceinte”,法语中分别表示“分娩*”和“怀孕”的术语)来选择EMR的子集,然后手动检查EMR以确认怀孕。然后,我们将我们的TM算法应用于2008至2012年间收治的独立队列中的BC患者(队列B)。结果在队列A中,在344名患者中确定了36例妊娠(10.5%;2829人年的EMR)。通过人工审查确定了30个,而TM确定了35个。TM导致人工检查的百分比较低(分别为26.7%和100%),并显著缩短了时间(识别怀孕的时间:TM分别为13分钟和244分钟)。两种TM滤器中任何一种的存在都显示出极高的灵敏度(97%)和阴性预测值(100%)。在队列B中,在1226名患者中发现了67例怀孕(5.5%;7349人年的EMR)。同样,对于B组,TM免除了904例(73.7%)EMR的人工审查,并在BC之后迅速产生了67例妊娠。两组中BC后妊娠的发生率均为0.01/EMR/人年。结论TM是一种快速识别罕见事件的高效工具,有望提高医学研究的速度、效率和成本。
PURPOSE To apply text mining (TM) technology on electronic medical records (EMRs) of patients with breast cancer (BC) to retrieve the occurrence of a pregnancy after BC diagnosis and compare its performance to manual curation.MATERIALS AND METHODS The training cohort (Cohort A) comprised 344 patients with BC age ? 40 years old treated at Institut Curie between 2005 and 2007. Manual curation consisted in manually reviewing each EMR to retrieve pregnancies. TM consisted of first applying a keyword filter ("accouch*" or "enceinte," French terms for "deliver*" and "pregnant," respectively) to select a subset of EMRs, and, second, checking manually EMRs to confirm the pregnancy. Then, we applied our TM algorithm on an independent cohort of patients with BC treated between 2008 and 2012 (Cohort B).RESULTS In Cohort A, 36 pregnancies were identified among 344 patients (10.5%; 2,829 person-years of EMR). Thirty were identified by manual review versus 35 by TM. TM resulted in a lower percentage of manual checking (26.7% v 100%, respectively) and substantial time gains (time to identify a pregnancy: 13 minutes for TM v 244 minutes for manual curation, respectively). Presence of any of the two TM filters showed excellent sensitivity (97%) and negative predictive value (100%). In Cohort B, 67 pregnancies were identified among 1,226 patients (5.5%; 7,349 person-years of EMR). Similarly, for Cohort B, TM spared 904 (73.7%) EMRs from manual review and quickly generated a cohort of 67 pregnancies after BC. Incidence rate of pregnancy after BC was 0.01 pregnancy per person-year of EMR in both cohorts.CONCLUSION TM is highly efficient to quickly identify rare events and is a promising tool to improve rapidity, efficiency, and costs of medical research.