Radiology Text Analysis System (RadText): Architecture and Evaluation.

Radiology Text Analysis System (RadText): Architecture and Evaluation.
复制标题

DOI:
10.1109/ichi54592.2022.00050
复制
发表时间:
2022-06
期刊:
Proceedings. IEEE International Conference on Healthcare Informatics
影响因子:
--
通讯作者:
Peng, Yifan
Peng, Yifan
中科院分区:
其他
文献类型:
--
作者:
Wang, Song;Lin, Mingquan;Ding, Ying;Shih, George;Lu, Zhiyong;Peng, Yifan

文献摘要

参考文献

相似文献

分析放射学报告是一项耗时且容易出错的任务,这就提出了对高效的自动化放射学报告分析系统的需求,以减轻放射科医生的工作量,促进精确诊断。在这项工作中,我们介绍了一个高性能的开源的Python放射学文本分析系统RadText。RadText提供了一个简单易用的文本分析流水线,包括去识别、部分分割、句子分割和单词标记化、命名实体识别、语法分析和否定检测。RadText优于现有的广泛使用的工具包,具有混合文本处理模式,支持原始文本处理和本地处理,从而实现更高的准确性、更好的可用性和改善的数据隐私。RadText采用Bioc作为统一接口,并将输出标准化为与观察性医疗结果伙伴关系(OMOP)公共数据模型(CDM)兼容的结构化表示法,从而允许对跨多个不同数据源的观察性研究采取更系统的方法。我们在MIMIC-CXR数据集上评估了RadText,并为这项工作注释了五个新的疾病标签。RadText的分类准确率很高,平均准确率为0.91,平均召回率为0.94,平均F-1得分为0.92。我们还为五个新的疾病标签的测试集添加了注释,以便于未来的研究或应用。我们已经在https://github.com/bionlplab/radtext.上提供了我们的代码、文档、示例和测试集
Analyzing radiology reports is a time-consuming and error-prone task, which raises the need for an efficient automated radiology report analysis system to alleviate the workloads of radiologists and encourage precise diagnosis. In this work, we present RadText, a high-performance open-source Python radiology text analysis system. RadText offers an easy-to-use text analysis pipeline, including de-identification, section segmentation, sentence split and word tokenization, named entity recognition, parsing, and negation detection. Superior to existing widely used toolkits, RadText features a hybrid text processing schema, supports raw text processing and local processing, which enables higher accuracy, better usability and improved data privacy. RadText adopts BioC as the unified interface, and also standardizes the output into a structured representation that is compatible with Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM), which allows for a more systematic approach to observational research across multiple, disparate data sources. We evaluated RadText on the MIMIC-CXR dataset, with five new disease labels that we annotated for this work. RadText demonstrates highly accurate classification performances, with a 0.91 average precision, 0.94 average recall and 0.92 average F-1 score. We also annotated a test set for the five new disease labels to facilitate future research or applications. We have made our code, documentations, examples and the test set available at https://github.com/bionlplab/radtext.
DOI: 10.1136/jamia.2009.001560
发表时间: 2010-09-01
影响因子: 6.4
作者:
Savova, Guergana K.;Masanz, James J.;Chute, Christopher G.
通讯作者: Chute, Christopher G.
DOI: 10.1136/amiajnl-2013-001810
发表时间: 2013-11-01
影响因子: 6.4
作者:
Fan, Jung-wei;Yang, Elly W.;Huang, Yang
通讯作者: Huang, Yang
DOI: 10.1093/database/bat064
发表时间: 2013
期刊: Database : the journal of biological databases and curation
影响因子: --
作者:
Comeau DC;Islamaj Doğan R;Ciccarese P;Cohen KB;Krallinger M;Leitner F;Lu Z;Peng Y;Rinaldi F;Torii M;Valencia A;Verspoor K;Wiegers TC;Wu CH;Wilbur WJ
通讯作者: Wilbur WJ
DOI: 10.1016/j.jbi.2011.03.011
发表时间: 2011-10
影响因子: 4.5
作者:
Chapman BE;Lee S;Kang HP;Chapman WW
通讯作者: Chapman WW
DOI: 10.1373/49.4.624
发表时间: 2003-04-01
期刊: CLINICAL CHEMISTRY
影响因子: 9.3
作者:
McDonald, CJ;Huff, SM;Maloney, P
通讯作者: Maloney, P