Concept-based document readability in domain specific information retrieval

Concept-based document readability in domain specific information retrieval
复制标题

特定领域信息检索中基于概念的文档可读性

DOI:
--
复制
发表时间:
2006
期刊:
International Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
Xue Li
Xue Li
中科院分区:
--
文献类型:
--
作者:
Xin Yan;D. Song;Xue Li

文献摘要

被引文献

相似文献

特定领域的信息检索已成为需求。不仅领域专家,而且普通非专家用户也对搜索特定领域(例如,医疗和健康)信息。然而,对于普通用户来说,一个典型的问题是搜索结果总是具有不同可读性水平的文档的混合。非专业用户可能希望在列表顶部看到可读性更高的文档。因此,搜索结果需要按照可读性的降序重新排序。领域专家手动标记大型数据库文档的可读性通常是不切实际的。可读性的计算模型需要研究。然而,传统的可读性公式是为通用文本设计的,不足以处理特定领域信息检索的技术材料。更先进的算法,如文本一致性模型是计算昂贵的重新排序大量检索到的文档。在本文中,我们提出了一个有效的和计算易处理的基于概念的文本可读性模型。除了文档的文本类型之外,我们的模型还考虑了特定领域的知识,即,文档中包含的特定于领域的概念如何影响文档的可读性。提出了三个主要的可读性公式,并将其应用于健康和医学信息检索。实验结果表明,我们提出的可读性公式导致显着的改善与用户的可读性评级的相关性比四个传统的可读性措施。
Domain specific information retrieval has become in demand. Not only domain experts, but also average non-expert users are interested in searching domain specific (e.g., medical and health) information from online resources. However, a typical problem to average users is that the search results are always a mixture of documents with different levels of readability. Non-expert users may want to see documents with higher readability on the top of the list. Consequently the search results need to be re-ranked in a descending order of readability. It is often not practical for domain experts to manually label the readability of documents for large databases. Computational models of readability needs to be investigated. However, traditional readability formulas are designed for general purpose text and insufficient to deal with technical materials for domain specific information retrieval. More advanced algorithms such as textual coherence model are computationally expensive for re-ranking a large number of retrieved documents. In this paper, we propose an effective and computationally tractable concept-based model of text readability. In addition to textual genres of a document, our model also takes into account domain specific knowledge, i.e., how the domain-specific concepts contained in the document affect the document's readability. Three major readability formulas are proposed and applied to health and medical information retrieval. Experimental results show that our proposed readability formulas lead to remarkable improvements in terms of correlation with users' readability ratings over four traditional readability measures.