NLM's Software Application to De-identify Clinical Text Documents
NLM's Software Application to De-identify Clinical Text Documents
批准号:
8158053
负责人:
Mehmet Kayaalp
金额:
$31.71万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
中文摘要
临床文本文档包含一组丰富的临床知识,对于临床研究来说是非常宝贵的。不幸的是,这些数据在很大程度上仍然是一种未开发的资源,因为按原样传播这些数据将危及患者的隐私,并泄露受保护的健康信息。
计算去识别是克服这个问题的一种手段。它涉及使用自然语言处理(NLP)工具和技术处理临床文本文档,识别与患者相关的个体可识别信息(例如,姓名、地址、电话号码和社会安全号码),并仅编辑这些标识符。通过这种方式,患者隐私得到保护,临床知识得到保存。
在没有计算工具的情况下,去识别化给临床医生带来了沉重的负担,但这是保护患者隐私的必要步骤,这是《健康保险流通与责任法案》(HIPAA)和1974年《隐私法》规定的。
在探索了现有的去识别工具之后,美国国家医学图书馆(NLM)正在开发新的软件,该软件能够以高精度去识别多种临床文本文档。
软件设计使用了一些确定性和概率模式识别算法和各种计算语言学方法。我们正在使用许多大型数据集的名称,地址和组织,所有这些都有可能识别患者,以便从文本中找到并删除此类内容。
应用程序接受纯文本或HL7格式的文本文档。如果文档以HL7格式提供,则应用程序利用嵌入在各种HL7段和字段中的患者相关信息,以便从文本语料库中高精度地查找和删除该信息,包括印刷错误和拼写错误。
应用软件包括一个可视化和标记的编辑器,称为可视化标记工具(VTT)。虽然VTT是专门为标记包含个人身份受保护的健康信息的标识符而设计的,但它将公开提供给更大的NLP社区,以扩展词汇标记和文本注释。
我们正在开始一系列研究,以评估在大量标记的临床文档语料库上去识别的成功率。这项研究的初步结果表明,计算去识别方法可以达到或超过99%的灵敏度和99%的特异性的水平,在一个大范围的标识符包含个人身份信息的准确性。
英文摘要
Clinical text documents contain a rich set of clinical knowledge that is invaluable for clinical research. Unfortunately, they remain a largely untapped resource since disseminating such data as-is would jeopardize the privacy of patients and reveal protected health information.
Computational de-identification is a means to overcome this problem. It involves processing clinical text documents using natural language processing (NLP) tools and techniques, recognizing patient-related individually identifiable information (e.g., names, addresses, and telephone and social security numbers) in the text, and redacting only those identifiers. In this way, patient privacy is protected and clinical knowledge is preserved.
Without computational tools, de-identification places a heavy burden on clinicians shoulders, but it is a necessary step for protecting patient privacy as mandated by both the Privacy Rule of the Health Insurance Portability and Accountability Act (HIPAA) and the Privacy Act of 1974.
After exploring existing de-identification tools, the U.S. National Library of Medicine (NLM) is developing new software that is capable of de-identifying many kinds of clinical text documents with high accuracy.
The software design uses a number of deterministic and probabilistic pattern recognition algorithms and various computational linguistic methods. We are using many large datasets for names, addresses, and organizations, all of which have the potential to identify patients, in order to find and remove such content from the text.
The application accepts text documents in plain text or in HL7 format. If documents are provided in an HL7 format, the application makes use of patient related information embedded in various HL7 segments and fields in order to find and remove that information, including typographical errors and misspellings, from the corpus of the text with high accuracy.
The application software includes an editor for visualization and markup called the Visual Tagging Tool (VTT). Although designed specifically for tagging identifiers that contain personally identifiable protected health information, VTT will be made publicly available to the greater NLP community for expanded lexical tagging and text annotation.
We are beginning a series of studies to assess the success of de-identifying on a large corpus of tagged clinical documents. The preliminary results of this study suggest that computational de-identification methods may attain an accuracy at or better than the level of 99% sensitivity and 99% specificity across a large spectrum of identifiers containing personally identifiable information.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
NLM's Software Application to De-identify Clinical Text Documents
-
批准号:8558114
-
项目类别:
-
资助金额:$34.93万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
NLM Scrubber: NLM's Software Application to De-identify Clinical Text Documents
-
批准号:9554455
-
项目类别:
-
资助金额:$46.34万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
NLM's Software Application to De-identify Clinical Text Documents
-
批准号:8344957
-
项目类别:
-
资助金额:$33.32万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
NLM Scrubber: NLM's Software Application to De-identify Clinical Text Documents
-
批准号:10268072
-
项目类别:
-
资助金额:$84.51万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
NLM's Software Application to De-identify Clinical Text Documents
-
批准号:8943232
-
项目类别:
-
资助金额:$38.53万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
国内基金
海外基金
Molecular Interaction Reconstruction of Rheumatoid Arthritis Therapies Using Clinical Data
-
批准号:31070748
-
项目类别:面上项目
-
资助金额:34.0万元
-
批准年份:2010
-
负责人:Christine Nardini
-
依托单位: