Does BERT Pretrained on Clinical Notes Reveal Sensitive Data?

Does BERT Pretrained on Clinical Notes Reveal Sensitive Data?
复制标题

DOI:
10.18653/v1/2021.naacl-main.73
复制
发表时间:
2021-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Eric P. Lehman;Sarthak Jain;Karl Pichotta;Yoav Goldberg;Byron C. Wallace
Eric P. Lehman;Sarthak Jain;Karl Pichotta;Yoav Goldberg;Byron C. Wallace
中科院分区:
其他
文献类型:
--
作者:
Eric P. Lehman;Sarthak Jain;Karl Pichotta;Yoav Goldberg;Byron C. Wallace

文献摘要

被引文献

相似文献

通过电子健康记录(EHR)预处理的大型变压器(EHR)在预测临床任务上的绩效取得了可观的增长。限制性模型(例如临床)的释放,尽管大多数努力都使用了EHR,但许多研究人员都可以使用大量敏感的他们可能会训练BERT模型的非识别EHR(或类似)。 PHI)特别是在训练有素的BERT中。在EHR的模拟物III语料库中接受了培训。 exposing_patient_data_release。
Large Transformers pretrained over clinical notes from Electronic Health Records (EHR) have afforded substantial gains in performance on predictive clinical tasks. The cost of training such models (and the necessity of data access to do so) coupled with their utility motivates parameter sharing, i.e., the release of pretrained models such as ClinicalBERT. While most efforts have used deidentified EHR, many researchers have access to large sets of sensitive, non-deidentified EHR with which they might train a BERT model (or similar). Would it be safe to release the weights of such a model if they did? In this work, we design a battery of approaches intended to recover Personal Health Information (PHI) from a trained BERT. Specifically, we attempt to recover patient names and conditions with which they are associated. We find that simple probing methods are not able to meaningfully extract sensitive information from BERT trained over the MIMIC-III corpus of EHR. However, more sophisticated “attacks” may succeed in doing so: To facilitate such research, we make our experimental setup and baseline probing models available at https://github.com/elehman16/exposing_patient_data_release.