NLM Scrubber: NLM's Software Application to De-identify Clinical Text Documents
NLM Scrubber: NLM's Software Application to De-identify Clinical Text Documents
批准号:
9554455
负责人:
Mehmet Kayaalp
金额:
$46.34万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
AddressAlgorithmsAreaClinicalClinical ResearchComputational LinguisticsComputer softwareDatabasesDictionaryGoalsGuidelinesHealthInformation SystemsKnowledgeLawsMethodsNamesNational Cancer InstitutePathology ReportPatientsPattern RecognitionPerformancePersonally Identifiable InformationPoliciesPrivacyRegulationReportingResearchRiskSEER ProgramSocial Security NumberSoftware DesignSystemTelephoneTestingTextTranslational ResearchUnited States National Institutes of HealthUnited States National Library of MedicineWorkbaseimprovedpatient privacyrepositorytool
中文摘要
叙述性临床报告包含丰富的临床知识,对临床研究可能是无价的。但是,它们也可能包含个人身份信息(PII),使这些临床报告被归类为PHI,这与使用限制和隐私风险有关。计算去识别试图删除此类叙述性文本中的所有PII实例,以便生成去识别文档,这些文档将不再被归类为PHI,并且可以在限制较少且几乎没有隐私风险的情况下用于研究。计算去识别使用模式识别和计算语言学方法来识别表示PII的单词和其他字母数字标记(例如,姓名、地址、电话号码和社会安全号码),并对它们进行编辑。通过这种方式,既保护了患者隐私,又保留了临床知识。
在探索了现有的去识别工具之后,美国国家医学图书馆(NLM)开始开发一种名为NLM Scrubber的新软件应用程序,该软件能够以高准确度去识别许多类型的临床报告。软件设计基于确定性和概率模式识别以及利用人名、地址和组织的大型词典的计算语言学方法。应用程序接受纯文本或HL7格式的叙述性报告。当输入报告被格式化为HL7消息时,应用软件利用嵌入在HL7段中的患者信息在HL7消息的文本部分中查找此类信息。
2014年11月,我们发布了NLM Scrubber的第一个测试版,可从https://scrubber.nlm.nih.gov免费下载。NLM Scrubber在检测单词和其他包含口述报告中发现的PII的字母数字令牌方面表现出色。我们的重点是扩展我们的工作,以进一步提高NLM Scrubers在大量标识符和其他报告类型中的去标识性能。NLM Scrubber将用于对NIH的临床叙述性报告的整个生物医学转化研究信息系统(BTRIS)存储库以及国家癌症研究所维护的监测、流行病学和最终结果(SEER)数据库中的叙述性病理学报告进行去识别。
英文摘要
Narrative clinical reports contain a rich set of clinical knowledge that could be invaluable for clinical research. However, they may also contain personally identifiable information (PII) that make those clinical reports classified as PHI, which is associated with use restrictions and risks to privacy. Computational de-identification seeks to remove all instances of PII in such narrative text in order to produce de-identified documents, which would no longer be classified as PHI and can be used in research with fewer constraints and with almost no risk to privacy. Computational de-identification uses pattern recognition and computational linguistic methods to recognize words and other alphanumeric tokens denoting PII (e.g., names, addresses, and telephone and social security numbers) in the text, and redacts them. In this way, both patient privacy is protected and clinical knowledge is preserved.
After exploring existing de-identification tools, the U.S. National Library of Medicine (NLM) began developing a new software application called NLM Scrubber, which is capable of de-identifying many types of clinical reports with high accuracy. The software design is based on both deterministic and probabilistic pattern recognition and computational linguistic methods utilizing large dictionaries of personal names, addresses, and organizations. The application accepts narrative reports in plain text or in HL7 format. When the input reports are formatted as HL7 messages, the application software leverages patient information embedded in HL7 segments to find such information in the text portion of the HL7 message.
In November 2014, we released the first beta version of NLM Scrubber, which is freely downloadable from https://scrubber.nlm.nih.gov. NLM Scrubber performs quite well on detecting words and other alphanumeric tokens containing PII found on dictated reports. Our focus is on extending our work to further improve NLM Scrubbers de-identification performance across a large spectrum of identifiers and additional report types. NLM Scrubber will be used to de-identify the entire Biomedical Translational Research Information System (BTRIS) repository of clinical narrative reports at NIH as well as the narrative pathology reports in Surveillance, Epidemiology and End Results (SEER) database maintained by National Cancer Institute.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
NLM's Software Application to De-identify Clinical Text Documents
-
批准号:8558114
-
项目类别:
-
资助金额:$34.93万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
NLM's Software Application to De-identify Clinical Text Documents
-
批准号:8344957
-
项目类别:
-
资助金额:$33.32万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
NLM Scrubber: NLM's Software Application to De-identify Clinical Text Documents
-
批准号:10268072
-
项目类别:
-
资助金额:$84.51万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
NLM's Software Application to De-identify Clinical Text Documents
-
批准号:8158053
-
项目类别:
-
资助金额:$31.71万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
NLM's Software Application to De-identify Clinical Text Documents
-
批准号:8943232
-
项目类别:
-
资助金额:$38.53万
-
财政年份:--
-
负责人:Mehmet Kayaalp
-
依托单位:
海外基金