NLM's Software Application to De-identify Clinical Text Documents
NLM's Software Application to De-identify Clinical Text Documents
批准号:
8158053
负责人:
Mehmet Kayaalp
金额:
$31.71万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
中文摘要
临床文本文档包含一套丰富的临床知识,对临床研究来说是无价的。不幸的是,它们在很大程度上仍然是一个未开发的资源,因为按原样传播这些数据将危及患者的隐私,并泄露受保护的健康信息。
计算型去身份识别是克服这一问题的一种方法。它涉及使用自然语言处理(NLP)工具和技术处理临床文本文档,识别文本中与患者相关的个人可识别信息(例如,姓名、地址以及电话和社会保障号码),并仅编辑这些识别符。这样,患者的隐私得到了保护,临床知识得到了保存。
在没有计算工具的情况下,去身份识别给临床医生带来了沉重的负担,但这是根据《健康保险携带和责任法案》(HIPAA)的隐私规则和1974年的《隐私法案》保护患者隐私的必要步骤。
在探索了现有的识别工具后,美国国家医学图书馆(NLM)正在开发新的软件,该软件能够高精度地识别多种临床文本文档。
软件设计使用了许多确定性和概率模式识别算法以及各种计算语言方法。我们正在使用许多用于姓名、地址和组织的大型数据集,所有这些数据集都有可能识别患者,以便从文本中找到并删除这些内容。
该应用程序接受纯文本或HL7格式的文本文档。如果文档是以HL7格式提供的,则应用程序利用嵌入在各个HL7段和字段中的与患者相关的信息,以便高精度地从文本语料库中查找和删除该信息,包括排版错误和拼写错误。
该应用软件包括一个用于可视化和标记的编辑器,称为可视化标记工具(VTT)。尽管VTT是专门为标记包含个人可识别的受保护健康信息的标识符而设计的,但它将向更大的NLP社区公开提供,以扩展词汇标记和文本注释。
我们正在开始一系列研究,以评估在大量标记的临床文档语料库上去身份识别的成功。这项研究的初步结果表明,在包含个人身份信息的大范围识别器上,计算识别方法可能达到或更好地达到99%的灵敏度和99%的特异度的水平。
英文摘要
Clinical text documents contain a rich set of clinical knowledge that is invaluable for clinical research. Unfortunately, they remain a largely untapped resource since disseminating such data as-is would jeopardize the privacy of patients and reveal protected health information.
Computational de-identification is a means to overcome this problem. It involves processing clinical text documents using natural language processing (NLP) tools and techniques, recognizing patient-related individually identifiable information (e.g., names, addresses, and telephone and social security numbers) in the text, and redacting only those identifiers. In this way, patient privacy is protected and clinical knowledge is preserved.
Without computational tools, de-identification places a heavy burden on clinicians shoulders, but it is a necessary step for protecting patient privacy as mandated by both the Privacy Rule of the Health Insurance Portability and Accountability Act (HIPAA) and the Privacy Act of 1974.
After exploring existing de-identification tools, the U.S. National Library of Medicine (NLM) is developing new software that is capable of de-identifying many kinds of clinical text documents with high accuracy.
The software design uses a number of deterministic and probabilistic pattern recognition algorithms and various computational linguistic methods. We are using many large datasets for names, addresses, and organizations, all of which have the potential to identify patients, in order to find and remove such content from the text.
The application accepts text documents in plain text or in HL7 format. If documents are provided in an HL7 format, the application makes use of patient related information embedded in various HL7 segments and fields in order to find and remove that information, including typographical errors and misspellings, from the corpus of the text with high accuracy.
The application software includes an editor for visualization and markup called the Visual Tagging Tool (VTT). Although designed specifically for tagging identifiers that contain personally identifiable protected health information, VTT will be made publicly available to the greater NLP community for expanded lexical tagging and text annotation.
We are beginning a series of studies to assess the success of de-identifying on a large corpus of tagged clinical documents. The preliminary results of this study suggest that computational de-identification methods may attain an accuracy at or better than the level of 99% sensitivity and 99% specificity across a large spectrum of identifiers containing personally identifiable information.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
NLM's Software Application to De-identify Clinical Text Documents
-
批准号:8558114
-
项目类别:
-
资助金额:$34.93万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
NLM Scrubber: NLM's Software Application to De-identify Clinical Text Documents
-
批准号:9554455
-
项目类别:
-
资助金额:$46.34万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
NLM's Software Application to De-identify Clinical Text Documents
-
批准号:8344957
-
项目类别:
-
资助金额:$33.32万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
NLM Scrubber: NLM's Software Application to De-identify Clinical Text Documents
-
批准号:10268072
-
项目类别:
-
资助金额:$84.51万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
NLM's Software Application to De-identify Clinical Text Documents
-
批准号:8943232
-
项目类别:
-
资助金额:$38.53万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
国内基金
海外基金
Molecular Interaction Reconstruction of Rheumatoid Arthritis Therapies Using Clinical Data
-
批准号:31070748
-
项目类别:面上项目
-
资助金额:34.0万元
-
批准年份:2010
-
负责人:Christine Nardini
-
依托单位: