Evaluating and reducing the effect of data corruption when applying bag of words approaches to medical records

Evaluating and reducing the effect of data corruption when applying bag of words approaches to medical records
复制标题

DOI:
10.1016/s1386-5056(02)00057-6
复制
发表时间:
2002-12-04
影响因子:
4.9
通讯作者:
Geissb端hler, A
Geissb端hler, A
中科院分区:
医学2区
文献类型:
--
作者:
Ruch, P;Baud, R;Geissb端hler, A

文献摘要

被引文献

相似文献

与期刊语料库不同,期刊语料库在出版前应该经过仔细审查,病历中的文件质量经常被拼写错误的单词和传统的图表或缩写所破坏。经过调查的域,本文着重于评估这种腐败的信息检索(IR)引擎的影响。IR系统采用经典的词袋方法,以词干为代表项,词频-逆文档频率(tf-idf)为权重模式,我们特别注意归一化因子。第一个结果表明,即使是低腐败水平(3%)影响检索效率(4-7%),而更高的腐败水平可以影响检索效率的25%。然后,我们表明,使用改进的自动拼写校正系统,应用于损坏的集合,几乎可以恢复引擎的检索效果。(C)2002爱思唯尔科学爱尔兰有限公司保留所有权利。
Unlike journal corpora, which are supposed to be carefully reviewed before being published, the quality of documents in a patient record are often corrupted by misspelled words and conventional graphies or abbreviations. After a survey of the domain, the paper focuses on evaluating the effect of such corruption on an information retrieval (IR) engine. The IR system uses a classical bag of words approach, with stems as representation items and term frequency-inverse document frequency (tf-idf) as weighting schema; we pay special attention to the normalization factor. First results shows that even low corruption levels (3%) do affect retrieval effectiveness (4-7%), whereas higher corruption levels can affect retrieval effectiveness by 25%. Then, we show that the use of an improved automatic spelling correction system, applied on the corrupted collection, can almost restore the retrieval effectiveness of the engine. (C) 2002 Elsevier Science Ireland Ltd. All rights reserved.