课题基金 / 基金详情

Research on Advanced Natural Language Processing and Text Mining

Research on Advanced Natural Language Processing and Text Mining
先进自然语言处理与文本挖掘研究
批准号:
18002007
负责人:
TSUJII Junichi
金额:
$319.57万
依托单位:
依托单位国家:
日本
项目类别:
Grant-in-Aid for Specially Promoted Research
财政年份:
2006
资助国家:
日本
项目状态:
已结题
起止时间:
2006 至 2010

项目摘要

项目成果

TSUJII Junichi的其他基金

相似基金

相关文献

中文摘要
翻译
该项目的目标是将统计建模与基于结构的符号处理相结合的方法应用于更具挑战性的任务,如深层语义处理、基于知识的信息提取和上下文处理。我们在以下方面取得了显著的成果:(1)基于语言形式的高效、健壮的深度分析;(2)用于生物领域的大规模语义标注语料库(Genia语料库);(3)用于生物领域的信息提取程序(命名实体识别器和事件识别器);(4)用于以数据为中心的并行处理的工作流软件。它曾两次被用作国际共享任务竞赛的训练和测试语料库(BioNLP 09和BioNLP 11)。在(3)中开发的提取程序在这些国际共享任务比赛中成功地展示了最先进的表现。基于(1)和(4)的系统表明,该项目开发的技术对于处理真实世界的文本是实用的。我们成功地处理了整个MEDLINE(2000万摘要,20多亿个句子),并在不到一周的时间内对它们进行了语义索引。MEDLINE的处理结果已通过智能文件检索系统(MEDIE)向公众公布
英文摘要
The objective of the project was to apply the methodology of combining statistical modeling with structure-based symbolic processing, which had proven successful in sentence parsing, to more challenging tasks such as deep semantic processing, knowledge-based information extraction and contextual processing. We have achieved significant results in (1) efficient and robust deep parsing based on a linguistically sound formalism, (2) a large scale semantically annotated corpus for the biology domain (GENIA corpus), (3) information extraction programs (named entity recognizers and event recognizers) for the biology domain which combine the deep parsing in (1) and structural machine learning algorithms, and (4) Workflow software for data-centered parallel processing.The GENIA corpus in (2) has been recognized as the gold standard corpus for research of text mining for biology and has been used by many groups in the world. It was adopted as the training and test corpus for international shared task competition twice (BioNLP 09 and BioNLP 11). The extraction programs developed in (3) successfully showed the state of the art performance in these international shared task competitions. The system based on (1) and (4) showed that the technology developed by this project was practical for processing the real world text. We successfully processed the whole of MEDLINE (20 million abstracts, more than 2 billion sentences) and indexed them semantically in less than a week. The processing results of MEDLINE has been made publicly available through an intelligent document retrieval system (MEDIE)
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1186/1471-2105-9-10
发表时间: 2008-01-08
期刊: BMC bioinformatics
影响因子: 3
作者: [Kim JD, Ohta T, Tsujii J]
通讯作者: Tsujii J
DOI: 10.1093/bioinformatics/btn469
发表时间: 2008-11-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者: [Tsuruoka Y, Tsujii J, Ananiadou S]
通讯作者: Ananiadou S
DOI: --
发表时间: 2008-06
期刊:
影响因子: --
作者: [Yusuke Miyao;Rune Sætre;Kenji Sagae;Takuya Matsuzaki;Junichi Tsujii]
通讯作者: Yusuke Miyao;Rune Sætre;Kenji Sagae;Takuya Matsuzaki;Junichi Tsujii
DOI: 10.3115/1220175.1220234
发表时间: 2006-07
期刊:
影响因子: --
作者: [Daisuke Okanohara;Yusuke Miyao;Yoshimasa Tsuruoka;Junichi Tsujii]
通讯作者: Daisuke Okanohara;Yusuke Miyao;Yoshimasa Tsuruoka;Junichi Tsujii
共 160 条
    Grammar Formalism with Self-Productivity and Development of Superhigh-speed Parsers for the Grammar Formalism
    • 批准号:
      08408009
    • 项目类别:
      Grant-in-Aid for Scientific Research (A)
    • 资助金额:
      $14.14万
    • 财政年份:
      1996
    • 负责人:
      TSUJII Junichi
    • 依托单位:
    海外基金