课题基金 / 基金详情

Natural Language Processing Techniques To Enhance Information Access.

Natural Language Processing Techniques To Enhance Information Access.
增强信息访问的自然语言处理技术。
批准号:
8943224
负责人:
Willy Wilbur
金额:
$56.15万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:

项目摘要

项目成果

Willy Wilbur的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Recently we have been involved in several subprojects which use natural language processing techniques: 1) We have developed a machine learning algorithm for abbreviation definition identification in text which makes use of what we term naturally labeled data. Positive training examples are naturally occurring potential abbreviation-definition pairs in text. Negative training examples are generated by randomly mixing potential abbreviations with unrelated potential definitions. The machine learner is trained to distinguish between these two sets of examples. Then, the learned feature weights are used to identify the abbreviation full form. This approach does not require manually labeled training data. We evaluate the performance of our algorithm on the Ab3P, BIOADI and Medstract corpora. Our system demonstrated results that compare favourably to the existing Ab3P and BIOADI systems. We achieve an F-measure of 91.36% on Ab3P corpus, and an F-measure of 87.13% on BIOADI corpus which are superior to the results reported by Ab3P and BIOADI systems. Moreover, we outperform these systems in terms of recall, which is one of our goals. 2) We are studying paraphrases in MEDLINE abstracts. These come about because an author is describing some entity of interest and uses a phrase like "drug abuse" and then needing to describe the same entity again a sentence or two latter does not wish to use exactly the same wording again and may use a variant of the phrase such as "drug use" which in the context of "drug abuse" has substantially the same meaning. 3) An author disambiguation algorithm has been developed which relies on machine learning based on the assumption that if an author name is infrequent in the data it probably represents the same person in all documents where it is found. This gives us positive instances. Negative instances are sampled from pairs of documents that have no author in common. Such positive and negative data allows us to do machine learning on all aspects of the document other than the name in question. This allows us to learn how to weight this data for best performance in distinguishing the positive and negative instances from each other. This learning is then applied in individual name cases or spaces to determine which author document pairs represent the same author. 4) We are using results of dependency parsers and syntactic parsers to create features for improved machine learning and also to automatically find good titles for document clusters.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A Document Processing System
  • 批准号:
    8344939
  • 项目类别:
  • 资助金额:
    $7.99万
  • 财政年份:
    --
  • 负责人:
    Willy Wilbur
  • 依托单位:
Automatic Analysis and Annotation of Document Keywords in Biomedical Literature
  • 批准号:
    8344960
  • 项目类别:
  • 资助金额:
    $23.98万
  • 财政年份:
    --
  • 负责人:
    Willy Wilbur
  • 依托单位:
General and Semi-supervised Machine Learning Applied to Bioinformatics
  • 批准号:
    8558105
  • 项目类别:
  • 资助金额:
    $56.4万
  • 财政年份:
    --
  • 负责人:
    Willy Wilbur
  • 依托单位:
PubMed Query Log Analysis and Use in Access Inhancement
  • 批准号:
    7969244
  • 项目类别:
  • 资助金额:
    $77.4万
  • 财政年份:
    --
  • 负责人:
    Willy Wilbur
  • 依托单位:
海外基金