A system for de-identifying medical message board text.

A system for de-identifying medical message board text.
复制标题

DOI:
10.1186/1471-2105-12-s3-s2
复制
发表时间:
2011-06-09
期刊:
影响因子:
3
通讯作者:
Holmes JH
Holmes JH
中科院分区:
生物学4区
文献类型:
--
作者:
Benton A;Hill S;Ungar L;Chung A;Leonard C;Freeman C;Holmes JH

文献摘要

被引文献

相似文献

用户在医疗留言板上发布了数百万条公共帖子,寻求有关各种医疗状况的支持和信息。事实证明,这些帖子可以用来更好地了解患者的经历和担忧。随着研究人员出于研究目的继续探索大量医学讨论板数据,保护这些在线社区成员的隐私成为需要应对的重要挑战。用于更结构化文本的现有实体识别方法是不够的,因为消息帖子带来了额外的挑战:帖子包含许多印刷错误、大量可能的名称、术语和缩写,特定于互联网帖子或特定留言板,以及提及作者的个人生活。考虑到上述挑战,本文的主要贡献是一个自动去识别留言板帖子作者身份的系统。我们在两个不同的留言板语料库上展示了我们的系统,一个是关于乳腺癌的,另一个是关于关节炎的。我们表明,我们的方法明显优于其他公开的命名实体识别和去识别系统,这些系统已针对更结构化的文本进行了调整,例如手术报告、病理报告、出院摘要或新闻专线。
There are millions of public posts to medical message boards by users seeking support and information on a wide range of medical conditions. It has been shown that these posts can be used to gain a greater understanding of patients’ experiences and concerns. As investigators continue to explore large corpora of medical discussion board data for research purposes, protecting the privacy of the members of these online communities becomes an important challenge that needs to be met. Extant entity recognition methods used for more structured text are not sufficient because message posts present additional challenges: the posts contain many typographical errors, larger variety of possible names, terms and abbreviations specific to Internet posts or a particular message board, and mentions of the authors’ personal lives. The main contribution of this paper is a system to de-identify the authors of message board posts automatically, taking into account the aforementioned challenges. We demonstrate our system on two different message board corpora, one on breast cancer and another on arthritis. We show that our approach significantly outperforms other publicly available named entity recognition and de-identification systems, which have been tuned for more structured text like operative reports, pathology reports, discharge summaries, or newswire.