课题基金 / 基金详情

项目摘要

项目成果

LAWRENCE E HUNTER的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供): 有一个明显的社区需要一个由生物医学期刊文章全文组成的带注释的语料库。有很多理由相信,今天阻碍生物医学语言处理进展的速度限制因素是缺乏合适的、经过专业注释的数据。带注释的语料库是具有与特定文本元素相关联的意义或结构的信息的文本的集合。标注语料库在两个方面是生物医学自然语言处理研究的重要组成部分。首先,大多数当代的语言处理方法至少部分依赖于机器学习或统计模型。这类系统必须根据具有已知输出的样例集进行“训练”,因此带注释的语料库提供了对构建现代自然语言处理系统至关重要的训练数据。其次,带注释的语料库提供了评估特定文本挖掘任务的各种方法的黄金标准。由于它们在培训和测试语言处理系统方面的核心作用,带注释的语料库的设计和业务创建的质量对这种系统所能完成的工作构成了根本限制。虽然在摘要标注方面已经做了一些有价值的工作,但从文本挖掘的角度来看,摘要和全文文章之间存在着重要的区别,对全文期刊文章的标注可以忽略不计。生物学(特别是模式生物数据库管理)和文本挖掘社区的工作人员都独立地指出,如果生物医学世界能够充分利用文本挖掘,处理科学出版物全文的重要性。我们建议建立一个大型的、完全注释的语料库,由生物医学期刊文章的全文组成。此外,以前的生物医学语料库注释工作经常利用特别的本体,这限制了它们在创建它们的小组之外的用途。我们将通过注释社区共识本体,如基因本体论和UMLS,来确保社区的可接受性。由于这项任务涉及昂贵的人力,因此效率是创建语料库的关键问题。为此,我们建议建立一个团队,其中包括迄今最大的语义标注语料库的构建者,模型生物体数据库的先驱之一,以及一批经验丰富的语言学和领域专家注释员。
英文摘要
DESCRIPTION (provided by applicant): There is a demonstrated community need for an annotated corpus consisting of the full texts of biomedical journal articles. There are many reasons to believe that the rate-limiting factor impeding progress in biomedical language processing today is the lack of availability of the right kind of expertly annotated data. An annotated corpus is a collection of texts with information about the meaning or structure associated with particular textual elements. Annotated corpora are a critical component of biomedical natural language processing research in two ways. First, most contemporary approaches to language processing rely at least in part on machine learning or statistical models. Such systems must be "trained" on sets of examples with known outputs, so annotated corpora provide the training data vital to the construction of modern NLP systems. Second, annotated corpora provide the gold standard by which various approaches to particular text mining tasks are evaluated. Due to their central roles in training and testing language processing systems, the quality of the design and operational creation of annotated corpora place fundamental limits on what can be accomplished with such systems. Although there has been valuable work done on annotating abstracts, there are important differences between abstracts and full-text articles from a text mining perspective, and annotation of full-text journal articles has been negligible. Workers in both the biological (especially model organism database curation) community and the text mining community have independently pointed out the importance of processing the full text of scientific publications if the biomedical world is to be able to fully utilize text mining. We propose to build a large, fully annotated corpus consisting of full texts of biomedical journal articles. Additionally, previous biomedical corpus annotation efforts have often utilized ad hoc ontologies that have limited their utility outside of the groups that created them. We will ensure community acceptability by annotating with respect to community-consensus ontologies such as the Gene Ontology and the UMLS. Since the task involves expensive human labor, efficiency is a key issue in creating corpora. For this reason, we propose to build a team that includes the builder of the largest semantically annotated corpus to date, one of the pioneers of the model organism databases, and an already-assembled cadre of experienced linguistic and domain-expert annotators.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
High Performance Text Mining for Translator
  • 批准号:
    10334356
  • 项目类别:
  • 资助金额:
    $47.12万
  • 财政年份:
    2020
  • 负责人:
    LAWRENCE E HUNTER
  • 依托单位:
Scientific Questions: A New Target for Biomedical NLP
  • 批准号:
    10223438
  • 项目类别:
  • 资助金额:
    $45.31万
  • 财政年份:
    2020
  • 负责人:
    LAWRENCE E HUNTER
  • 依托单位:
Scientific Questions: A New Target for Biomedical NLP
  • 批准号:
    10454968
  • 项目类别:
  • 资助金额:
    $44.52万
  • 财政年份:
    2020
  • 负责人:
    LAWRENCE E HUNTER
  • 依托单位:
High Performance Text Mining for Translator
  • 批准号:
    10548337
  • 项目类别:
  • 资助金额:
    $46.61万
  • 财政年份:
    2020
  • 负责人:
    LAWRENCE E HUNTER
  • 依托单位:
海外基金