Word embeddings trained on published case reports are lightweight, effective for clinical tasks, and free of protected health information.

Word embeddings trained on published case reports are lightweight, effective for clinical tasks, and free of protected health information.
复制标题

DOI:
10.1016/j.jbi.2021.103971
复制
发表时间:
2022-01
影响因子:
4.5
通讯作者:
Weissman GE
Weissman GE
中科院分区:
医学3区
文献类型:
--
作者:
Flamholz ZN;Crane-Droesch A;Ungar LH;Weissman GE

文献摘要

参考文献

被引文献

相似文献

量化开发临床相关词嵌入的几种策略中的性能,再现性和资源需求的权衡。我们对Pubmed Central(PMC)开放获取子集中的所有全文手稿、其中的病例报告、英文维基百科语料库、重症监护医学信息市场(MIMIC)III数据集以及宾夕法尼亚大学卫生系统(UPHS)电子健康记录中的所有注释进行了单独的嵌入训练。我们在六个临床相关任务中测试了嵌入,包括死亡率预测和去识别,并分别使用缩放Brier评分(SBS)和成功去识别的笔记比例评估了性能。UPHS的嵌入最能预测死亡率(SBS 0.30,95% CI 0.15 - 0.45),而维基百科的嵌入最差(SBS 0.12,95% CI-0.05-0.28)。维基百科嵌入最一致(78%的注释),而完整的PMC语料库嵌入最不一致(48%)去识别注释。在所有六个任务中,完整的PMC语料库表现出最一致的性能,维基百科语料库表现最不一致。语料库的规模从4900万个代币(PMC案例报告)到100亿个代币(UPHS)不等。在大多数任务中,在已发表的病例报告上训练的嵌入与在其他语料库上训练的嵌入一样少,临床语料库的表现始终优于非临床语料库。没有一个语料库在所有任务中产生严格的主导嵌入集,因此最佳训练语料库取决于预期用途。在已发表的病例报告上训练的嵌入在大多数临床任务上的表现优于在较大语料库上训练的嵌入。开放获取语料库允许训练临床相关的、有效的和可再现的嵌入。
Quantify tradeoffs in performance, reproducibility, and resource demands across several strategies for developing clinically relevant word embeddings. We trained separate embeddings on all full-text manuscripts in the Pubmed Central (PMC) Open Access subset, case reports therein, the English Wikipedia corpus, the Medical Information Mart for Intensive Care (MIMIC) III dataset, and all notes in the University of Pennsylvania Health System (UPHS) electronic health record. We tested embeddings in six clinically relevant tasks including mortality prediction and de-identification, and assessed performance using the scaled Brier score (SBS) and the proportion of notes successfully de-identified, respectively. Embeddings from UPHS notes best predicted mortality (SBS 0.30, 95% CI 0.15 to 0.45) while Wikipedia embeddings performed worst (SBS 0.12, 95% CI −0.05 to 0.28). Wikipedia embeddings most consistently (78% of notes) and the full PMC corpus embeddings least consistently (48%) de-identified notes. Across all six tasks, the full PMC corpus demonstrated the most consistent performance, and the Wikipedia corpus the least. Corpus size ranged from 49 million tokens (PMC case reports) to 10 billion (UPHS). Embeddings trained on published case reports performed as least as well as embeddings trained on other corpora in most tasks, and clinical corpora consistently outperformed non-clinical corpora. No single corpus produced a strictly dominant set of embeddings across all tasks and so the optimal training corpus depends on intended use. Embeddings trained on published case reports performed comparably on most clinical tasks to embeddings trained on larger corpora. Open access corpora allow training of clinically relevant, effective, and reproducible embeddings.
DOI: 10.1038/sdata.2016.35
发表时间: 2016-05-24
期刊: Scientific data
影响因子: 9.8
作者:
Johnson AE;Pollard TJ;Shen L;Lehman LW;Feng M;Ghassemi M;Moody B;Szolovits P;Celi LA;Mark RG
通讯作者: Mark RG
DOI: 10.1093/bioinformatics/btz682
发表时间: 2020-02-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Lee J;Yoon W;Kim S;Kim D;Kim S;So CH;Kang J
通讯作者: Kang J
用生物医学和一般领域知识库评估神经单词嵌入的语义关系。
DOI: 10.1186/s12911-018-0630-x
发表时间: 2018-07-23
影响因子: 3.5
作者:
Chen Z;He Z;Liu X;Bian J
通讯作者: Bian J
DOI: 10.1016/j.jbi.2012.07.012
发表时间: 2012-12
影响因子: 4.5
作者:
Keselman, Alla;Smith, Catherine Arnott
通讯作者: Smith, Catherine Arnott
DOI: 10.1016/j.yjbinx.2019.100057
发表时间: 2019-01-01
影响因子: 4.5
作者:
Khattak, Faiza Khan;Jeblee, Serena;Rudzicz, Frank
通讯作者: Rudzicz, Frank