课题基金 / 基金详情

From Text Corpora to Text Databases: Research in Text Processing and Retrieval

From Text Corpora to Text Databases: Research in Text Processing and Retrieval
从文本语料库到文本数据库:文本处理与检索研究
批准号:
9302615
负责人:
Ralph Grishman
金额:
$20.35万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
1993
资助国家:
美国
项目状态:
已结题
起止时间:
1993-08-01 至 1997-01-31

项目摘要

项目成果

Ralph Grishman的其他基金

相似基金

相关文献

中文摘要
翻译
9302615 Strzalkowski从文本语料库到文本数据库:文本处理和检索研究这是一项为期三年的持续性奖励的第一年资助。本研究的目的是探索自然语言处理在大型、最小结构文本库自动信息检索中的潜力。这项工作包括开发更有效的索引、路由、内容近似、抽象和从文本数据创建层次领域地图的技术。语言和统计方法都被使用。主要的信任是为以下问题找到满意的解决方案:(1)为搜索目的获得数据库内容的准确和通用的表示;(2)设计能够以速度和鲁棒性匹配或超过统计系统的算法来完成这项任务。为了创建能够支持各种搜索类型的数据库内容的准确表示,创建了一个广泛的自然语言处理组件。语言处理包括随机词性标注、词典辅助词干提取、句法解析、短语提取和消歧以及数据库域基础概念的语义关联。本研究广泛地基于大型文本集的实证实验。预计将产生能够显著提高顶级全文信息检索系统预期性能水平的技术。***
英文摘要
9302615 Strzalkowski From Text Corpora to Text Databases: Research in Text Processing and Retrieval This is the first year funding of a three-year continuing award. The goal of this research is to explore the potential of natural language processing in automated information retrieval from large, minimally structured text libraries. This effort includes development of more effective techniques of indexing, routing, contents approximation, abstracting and creation of hierarchical domain maps from textual data. Both linguistic and statistical methods are used. The main trust is to find satisfactory solutions to the following problems: (1) obtaining an accurate and versatile representation of database contents for search purposes; and (2) devising algorithms that can accomplish this task with the speed and robustness to match or exceed that of statistical systems. In order to create an accurate representation of database contents that would be able to support various types of search, an extensive natural language processing component is created. Linguistic processing includes stochastic part of speech tagging, dictionary- assisted stemming, syntactic parsing, phrase extraction and disambiguation, and semantic correlation of concepts underlying the database domain. This research is based extensively on empirical experiments with large text collections. It is expected to produce technologies that will significantly improve the expected performance levels for top-of-the-line full-text information retrieval systems. ***
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
ITR: Automated Structuring of Text Information
  • 批准号:
    0081962
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.98万
  • 财政年份:
    2000
  • 负责人:
    Ralph Grishman
  • 依托单位:
A Dictionary of Nominal Complements
  • 批准号:
    9633286
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $47.37万
  • 财政年份:
    1996
  • 负责人:
    Ralph Grishman
  • 依托单位:
Collaborative Research on Knowledge Aquisition for Japanese-English Machine Translation
  • 批准号:
    9303013
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $16.09万
  • 财政年份:
    1993
  • 负责人:
    Ralph Grishman
  • 依托单位:
A Sublanguage Approach to Japanese-English Machine Translation
  • 批准号:
    8902304
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $12.02万
  • 财政年份:
    1989
  • 负责人:
    Ralph Grishman
  • 依托单位:
海外基金