课题基金 / 基金详情

项目摘要

项目成果

LAWRENCE E HUNTER的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供): 生物医学语言处理是将计算技术应用于生物医学中人类生成的文本,是基础和应用生物医学研究的一项日益重要的使能技术。同行评议文献的指数增长和与高通量技术相关的学科界限的打破,增加了自动化工具的重要性,使科学家能够了解与其工作相关的所有已发表材料。然而,尽管经过几十年的研究,最先进的工具在信息提取和文档检索等基本语言处理任务中的性能仍然低于充分实用和广泛采用这项技术所需的水平。文本挖掘系统的开发、性能和评价在很大程度上取决于是否有适当的语料库:具有代表性的文件的集合,这些文件带有与语言处理任务有关的人类判断的注释。语料库在这项技术的发展中扮演着两个角色:第一,它们是可以公平比较替代自动化方法的“黄金标准”,第二,它们为统计和机器学习系统的培训提供数据,这些系统创建语言使用模式的经验模型。传统观点认为,语料库是感兴趣领域的中性随机样本。我们的初步工作表明,现有语料库在大小、质量、体裁和表征图式方面的限制本身就是生物医学文本处理技术近期突破的关键限制因素。因此,我们建议检验以下假设:从多个流派创建大型、高质量的生物医学语料库将导致生物医学文本挖掘系统的性能显著提高,并创建新的文本挖掘任务方法。具体目标包括建立涵盖一系列体裁的几个大型语料库并纳入丰富的知识表示;确定影响全文与摘要差异表现的因素;以及开发新的语言处理方法,特别是全文处理。由于从许多文本类型中自动提取信息的能力的改进将帮助科学家和临床医生完成跟上蓬勃发展的生物医学文献的关键任务,因此潜在的公共健康影响相当大。
英文摘要
DESCRIPTION (provided by applicant): Biomedical language processing, the application of computational techniques to human-generated texts in biomedicine, is an increasingly important enabling technology for basic and applied biomedical research. The exponential growth of the peer-reviewed literature and the breakdown of disciplinary boundaries associated with high-throughput techniques have increased the importance of automated tools for keeping scientists abreast of all of the published material relevant to their work. However, despite decades of research, the performance of state-of-the-art tools for basic language processing tasks like information extraction and document retrieval remain below the level necessary for adequate utility and widespread adoption of this technology. The development, performance and evaluation of text mining systems depend crucially on the availability of appropriate corpora: collections of representative documents that have been annotated with human judgments relevant to a language-processing task. Corpora play two roles in the development of this technology: first, they act as "gold standards" by which alternative automated methods can be fairly compared, and second, they provide data for the training of statistical and machine learning systems that create empirical models of patterns in language use. The conventional view is that corpora are neutral, random samples of the domain of interest. Our preliminary work suggests that the restrictions in size, quality, genre, and representational schema of the small number of existing corpora are themselves a critical limiting factor for near-term breakthroughs in biomedical text processing technology. Therefore, we propose to test the following hypothesis: Creation of large, high-quality, biomedical corpora from multiple genres will lead to significant improvements in the performance of biomedical text mining systems and the creation of new approaches to text mining tasks. Specific aims include constructing several large corpora covering a range of genres and incorporating a rich knowledge representation; identifying factors that affect differential performance on full text versus abstracts; and developing new methods for language processing, especially of full text. Because improvements in the ability to automatically extract information from many textual genres will assist scientists and clinicians in the crucial task of keeping up with the burgeoning biomedical literature, the potential public health impact is quite large.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
High Performance Text Mining for Translator
  • 批准号:
    10334356
  • 项目类别:
  • 资助金额:
    $47.12万
  • 财政年份:
    2020
  • 负责人:
    LAWRENCE E HUNTER
  • 依托单位:
Scientific Questions: A New Target for Biomedical NLP
  • 批准号:
    10223438
  • 项目类别:
  • 资助金额:
    $45.31万
  • 财政年份:
    2020
  • 负责人:
    LAWRENCE E HUNTER
  • 依托单位:
Scientific Questions: A New Target for Biomedical NLP
  • 批准号:
    10454968
  • 项目类别:
  • 资助金额:
    $44.52万
  • 财政年份:
    2020
  • 负责人:
    LAWRENCE E HUNTER
  • 依托单位:
High Performance Text Mining for Translator
  • 批准号:
    10548337
  • 项目类别:
  • 资助金额:
    $46.61万
  • 财政年份:
    2020
  • 负责人:
    LAWRENCE E HUNTER
  • 依托单位:
海外基金