A deep learning approach for transgender and gender diverse patient identification in electronic health records.

A deep learning approach for transgender and gender diverse patient identification in electronic health records.
复制标题

电子健康记录中跨性别和性别多样化患者识别的深度学习方法。

DOI:
10.1016/j.jbi.2023.104507
复制
发表时间:
2023
影响因子:
4.5
通讯作者:
Zhou,Li
Zhou,Li
中科院分区:
医学3区
文献类型:
--
作者:
Hua,Yining;Wang,Liqin;Nguyen,Vi;Rieu-Werden,Meghan;McDowell,Alex;Bates,DavidW;Foer,Dinah;Zhou,Li

文献摘要

相似文献

尽管准确识别电子健康记录(EHR)中的性别身份对于提供公平的医疗服务至关重要,特别是对于变性人和性别多元化(TGD)人群,但由于结构化EHR领域中性别信息的不完整,这仍然是一项具有挑战性的任务。目的以TGD身份识别为例,本研究使用自然语言处理和深度学习来构建准确的患者性别身份预测模型,旨在解决从EHR数据中识别相关患者级别信息的挑战,并减少注释工作。为了从大量的临床记录中识别相关信息,我们通过专家筛选、文献回顾和通过微调的BioWordVec模型进行扩展,编制了与性别相关的关键字列表。该关键字列表用于预先筛选潜在的TGD个体,并创建用于模型训练、测试和验证的两个数据集。数据集I是一个平衡的数据集,包含临床医生确认的TGD患者和没有关键字的病例。数据集二包含带有关键字的案例。结果最终的关键词列表由109个关键词组成,其中58个关键词(53.2%)通过BioWordVec模型进行扩展。数据集I包含3150名患者(50%的TGD),而数据集II包含200名患者(90%的TGD)。在数据集I上,深度学习模型的F1得分为0.917,灵敏度为0.854,精度为0.980;在数据集II上,F1得分为0.969,灵敏度为0.967,精度为0.972。深度学习模型的性能明显优于基于规则的算法。结论首次研究表明,深度学习与自然语言处理相结合的算法能够利用EHR数据准确地识别性别身份。未来的工作应该利用和评估其他不同的数据源,以生成更具普遍性的算法。
BackgroundAlthough accurate identification of gender identity in the electronic health record (EHR) is crucial for providing equitable health care, particularly for transgender and gender diverse (TGD) populations, it remains a challenging task due to incomplete gender information in structured EHR fields.ObjectiveUsing TGD identification as a case study, this research uses NLP and deep learning to build an accurate patient gender identity predictive model, aiming to tackle the challenges of identifying relevant patient-level information from EHR data and reducing annotation work.MethodsThis study included adult patients in a large healthcare system in Boston, MA, between 4/1/2017 to 4/1/2022. To identify relevant information from massive clinical notes, we compiled a list of gender-related keywords through expert curation, literature review, and expansion via a fine-tuned BioWordVec model. This keyword list was used to pre-screen potential TGD individuals and create two datasets for model training, testing, and validation. Dataset I was a balanced dataset that contained clinician-confirmed TGD patients and cases without keywords. Dataset II contained cases with keywords. The performance of the deep learning model was compared to traditional machine learning and rule-based algorithms.ResultsThe final keyword list consists of 109 keywords, of which 58 (53.2%) were expanded by the BioWordVec model. Dataset I contained 3,150 patients (50% TGD) while Dataset II contained 200 patients (90% TGD). On Dataset I the deep learning model achieved a F1 score of 0.917, sensitivity of 0.854, and a precision of 0.980; and on Dataset II a F1 score of 0.969, sensitivity of 0.967, and precision of 0.972. The deep learning model significantly outperformed rule-based algorithms.ConclusionThis is the first study to show that deep learning-integrated NLP algorithms can accurately identify gender identity using EHR data. Future work should leverage and evaluate additional diverse data sources to generate more generalizable algorithms.