De-identification of clinical narratives through writing complexity measures.

De-identification of clinical narratives through writing complexity measures.
复制标题

DOI:
10.1016/j.ijmedinf.2014.07.002
复制
发表时间:
2014-10
影响因子:
4.9
通讯作者:
Malin BA
Malin BA
中科院分区:
医学2区
文献类型:
--
作者:
Li M;Carrell D;Aberdeen J;Hirschman L;Malin BA

文献摘要

参考文献

被引文献

相似文献

电子健康记录包含大量的临床叙述,越来越多地被重复用于研究目的。为了大规模共享数据并尊重隐私,删除患者标识符至关重要。已经提出了基于机器学习的去识别工具;然而,模型训练通常基于随机的文档组或预先存在的文档类型指定(例如,出院总结)。这项工作调查,如果固有的功能,如写作的复杂性,可以识别文档子集,以提高去识别性能。我们应用了一种无监督聚类方法,根据写作复杂性度量对两个语料库进行分组:超过4500个不同文档类型的文档的集合(例如,出院摘要、病史和体检报告以及放射学报告)和889份出院摘要的公开可用i2 b2语料库。我们比较了在这些集群上训练的去识别模型与在随机分组的文档或VUMC文档类型上训练的模型的性能(通过召回率,精度和F-测量)。对于范德比尔特数据集,观察到在相同的风格聚类(平均F度量为0.917)上训练和测试去识别模型往往优于基于随机文档聚类的模型(平均F度量为0.881)。进一步观察到,增加从特定聚类采样的训练子集的大小可以产生改进的结果(例如,当训练集的规模从10个增加到50个时,F测度从0.743上升到0.841;当训练集的规模达到200个时,F测度达到0.901)。对于i2 b2数据集,基于复杂性度量(平均F分数0.966)对相同聚类进行训练和测试并没有显著超过随机选择的聚类(平均F分数0.965)。我们的研究结果表明,在由各种临床文档组成的环境中,在写作复杂性度量上训练的去识别模型比在随机组和许多情况下的文档类型上训练的模型更好。
Electronic health records contain a substantial quantity of clinical narrative, which is increasingly reused for research purposes. To share data on a large scale and respect privacy, it is critical to remove patient identifiers. De-identification tools based on machine learning have been proposed; however, model training is usually based on either a random group of documents or a pre-existing document type designation (e.g., discharge summary). This work investigates if inherent features, such as the writing complexity, can identify document subsets to enhance de-identification performance. We applied an unsupervised clustering method to group two corpora based on writing complexity measures: a collection of over 4500 documents of varying document types (e.g., discharge summaries, history and physical reports, and radiology reports) from Vanderbilt University Medical Center (VUMC) and the publicly available i2b2 corpus of 889 discharge summaries. We compare the performance (via recall, precision, and F-measure) of de-identification models trained on such clusters with models trained on documents grouped randomly or VUMC document type. For the Vanderbilt dataset, it was observed that training and testing de-identification models on the same stylometric cluster (with the average F-measure of 0.917) tended to outperform models based on clusters of random documents (with an average F-measure of 0.881). It was further observed that increasing the size of a training subset sampled from a specific cluster could yield improved results (e.g., for subsets from a certain stylometric cluster, the F-measure raised from 0.743 to 0.841 when training size increased from 10 to 50 documents, and the F-measure reached 0.901 when the size of the training subset reached 200 documents). For the i2b2 dataset, training and testing on the same clusters based on complexity measures (average F-score 0.966) did not significantly surpass randomly selected clusters (average F-score 0.965). Our findings illustrate that, in environments consisting of a variety of clinical documentation, de-identification models trained on writing complexity measures are better than models trained on random groups and, in many instances, document types.
DOI: 10.1037/h0057532
发表时间: 1948-06-01
影响因子: 9.9
作者:
Flesch, Rudolf
通讯作者: Flesch, Rudolf
DOI: 10.1197/jamia.m2131
发表时间: 2008-01-01
影响因子: 6.4
作者:
Johnson, Stephen B.;Bakken, Suzanne;Stetson, Peter
通讯作者: Stetson, Peter
DOI: 10.1309/e6k33gbpe5c27fyu
发表时间: 2004-02-01
影响因子: 3.5
作者:
Gupta, D;Saul, M;Gilbertson, J
通讯作者: Gilbertson, J
DOI: 10.1007/bf01830689
发表时间: 1994-04-01
期刊: COMPUTERS AND THE HUMANITIES
影响因子: --
作者:
HOLMES, DI
通讯作者: HOLMES, DI
DOI: 10.1016/0306-4573(93)90079-s
发表时间: 1993-09-01
影响因子: 8.6
作者:
LEHNER, F
通讯作者: LEHNER, F