课题基金 / 基金详情

TEXTUAL INFORMATION RETRIEVAL TESTING

TEXTUAL INFORMATION RETRIEVAL TESTING
文本信息检索测试
批准号:
5203618
负责人:
W J WILBUR
金额:
$0.0万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至

项目摘要

项目成果

W J WILBUR的其他基金

相似基金

相关文献

中文摘要
翻译
检索测试之所以困难,有几个原因。首先,集合 对其执行实际检索的文档太大, 指出,创建完整的人工判断测试集并不是 有可能。正因为如此,我们刚刚完成了对一项大型 在分子生物学领域的收藏。对于这个大型数据库, 对检索操作进行了模拟,并给出了相应的性能指标 与绝对衡量标准相比。这些模拟是基于 McCarn-Lewis模型方程。现已发现,这一亲属 正如我们所构建的那样,度量是绝对的敏感指标 性能,但与绝对衡量标准略有偏离 一种系统性的方法,在某些情况下可以很容易地预测到 正在测试的检索方法的可测量特性。 导致检索测试困难的第二个因素是缺乏 实际检索性能的模型。这使得很难 评估在进行测试时获得的结果的重要性。一 必须使用有关测试的对称性的特殊假设 通过不同检索方法等获得的结果。这些假设 很少感到满意。因此,我们已经研究了自举 方法:研究方法。根据这些结果,我们得出结论,引导程序 方法提供了一种可靠而实用的显著性检验方法 以获取检索性能结果。 导致检索测试困难的第三个方面是缺乏 人类相关性判断的重复性。我们采取的方法是 因此,相关性必须被理解为一种概率。这允许一个 对概率排序原理的独特解释。我们的研究 表明这样的解释提供了一种稳定的测试方法 这对性能设置了限制,并允许将人类和 机器性能在同等基础上
英文摘要
Retrieval testing is difficult for several reasons. First, the sets of documents on which actual retrieval are carried out are large to the point that the creation of complete humanly judged test sets is not possible. Because of this we have just completed a study of a large collection in the area of molecular biology. For this large database the retrieval operation is simulated and relative measures of performance are compared with absolute measures. The simulations are based on the McCarn-Lewis model equation. It has been found that the relative measure, as we have constructed it, is a sensitive indicator of absolute performance but one which deviates slightly from the absolute measure in a systematic way which can in some circumstances be predicted from easily measurable properties of the retrieval method being tested. The second factor that makes retrieval testing difficult is the lack of a model for actual retrieval performance. This makes it difficult to assess the significance of results obtained when testing is done. One must use special assumptions about the symmetry properties of test results obtained by different retrieval methods, etc. Such assumptions are seldom satisfied. We have therefore investigated the bootstrap methods. Based on these results we have concluded that the bootstrap methods provide a reliable and practical method of significance testing for retrieval performance results. The third area that makes retrieval testing difficult is the lack of reproducibility of human relevance judgements. We take the approach that relevance must therefore be understood as a probability. This allows a unique interpretation of the probability ranking principle. Our studies show that such an interpretation provides a stable approach to testing that places limits on performance and allows the comparison of human and machine performance on an equal footing
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
TEXTUAL INFORMATION RETRIEVAL TESTING
  • 批准号:
    2578621
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    W J WILBUR
  • 依托单位:
AUTOMATIC BAYESIAN METHODS IN TEXT RETRIEVAL
  • 批准号:
    2578622
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    W J WILBUR
  • 依托单位:
DYNAMIC MODELS OF PROTEIN FOLDING
  • 批准号:
    2578639
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    W J WILBUR
  • 依托单位:
A DOCUMENT PROCESSING SYSTEM
  • 批准号:
    3845112
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    W J WILBUR
  • 依托单位:
海外基金