课题基金 / 基金详情

A DOCUMENT PROCESSING SYSTEM

A DOCUMENT PROCESSING SYSTEM
文件处理系统
批准号:
3781267
负责人:
W J WILBUR
金额:
$0.0万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至

项目摘要

项目成果

W J WILBUR的其他基金

相似基金

相关文献

中文摘要
翻译
一个系统的软件已经开发的目的是找到 Medline中的相关文件。该系统由超过 用C语言编写的35个程序加上一些实用程序, 程序.该系统具有许多独特的功能: 1)它是高度模块化的,因此系统中的更改相对容易 易于执行。 2)该系统目前在ASN1格式的Medline数据上运行, 系统接口部分的变化将允许它 适用于任何由离散文本记录组成的大型数据库。 3)该系统的设计具有一定程度的安全性,防止数据丢失 由于操作系统崩溃或断电。 4)系统处理的所有数据都以永久形式存储, 倒排文件结构等。这些结构是可更新的, 当新数据变得可用时,可以将其不断地添加到系统中。 5)使用贝叶斯形式的 分析和统计数据,术语的相关性权重是 根据以前的文件比较得出。这些统计数据 随着每个新的处理周期而更新。 6)文档相关的概率由系统计算 基于使用一组文档产生的原始分数的缩放, 由人类评判员判断其亲缘关系的配对。这种规模 每次更新术语权重时重新计算, 不同的是,对于具有的文档和不具有的文档, 摘要。 对该系统最明显的故障进行分析, 用于文档相似度度量的测试集, 出去这表明,所经历的问题的很大一部分可能 是由于几个地区普遍存在不同程度的 的描述和不同程度的统一描述中, 单一文件。这些描述区域的数字表示 一般来说,在一个文件中, 定义文档内容。
英文摘要
A system of software has been developed for the purpose of finding the closely related documents in Medline. The system consists of over thirty-five programs written in the C language plus a number of utility programs. The system has a number of unique features: 1) It is highly modular so that alterations in the system are relatively simple to perform. 2) The system currently operates on Medline data in the ASN1 format but a change in the interface portion of the system would allow it to be applied to any large database consisting of discrete textual records. 3) The system is designed with a degree of security against loss of data due to operating system crashes or power outages. 4) All data processed by the system is stored in permanent form as inverted file structures, etc. These structures are updateable so that new data may be continually added to the system as it becomes available. 5) Documents are compared with each other using a Bayesian form of analysis and the statistics on which the relevance weighting of terms is based are derived from previous document comparisons. These statistics are updated with each new cycle of processing. 6) The probability that documents are related is computed by the system based on a scaling of the raw scores produced using a set of document pairs that have been judged for relatedness by human judges. This scale is recalculated each time term weights are updated and it is calculated differently for documents with as opposed to documents without abstracts. An analysis of the most glaring failures of the system, as identified on the test set used for scaling of document similarity, has been carried out. This shows that a significant part of the problems experienced may be due to the common occurrence of several areas with different levels of description and different levels of uniformity in description in a single document. The numerical representation of these descriptive areas within a document does not in general correlate with their importance in defining document content.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
TEXTUAL INFORMATION RETRIEVAL TESTING
  • 批准号:
    2578621
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    W J WILBUR
  • 依托单位:
AUTOMATIC BAYESIAN METHODS IN TEXT RETRIEVAL
  • 批准号:
    2578622
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    W J WILBUR
  • 依托单位:
DYNAMIC MODELS OF PROTEIN FOLDING
  • 批准号:
    2578639
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    W J WILBUR
  • 依托单位:
A DOCUMENT PROCESSING SYSTEM
  • 批准号:
    3845112
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    W J WILBUR
  • 依托单位:
海外基金