Construction of large scale annotated corpus and its management system
Construction of large scale annotated corpus and its management system
批准号:
12480082
负责人:
TOKUNAGA Takenobu
金额:
$8.96万
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (B)
财政年份:
2000
资助国家:
日本
项目状态:
已结题
起止时间:
2000 至 2002
中文摘要
自20世纪80年代中期以来,基于大规模语言数据的自然语言处理已成为该研究领域的主流。在这类研究中,语言资源扮演着重要的角色,人们已经尝试创造各种各样的资源。本课题旨在构建一个大规模的日语语料库句法标注环境。为了实现这一目标,我们进行了以下课题的研究。2000年,我们构建了一个注释工具,支持用户以交互式方式对句子的句法结构进行注释。该工具与现有的解析器一起工作,用户可以有效地从许多解析器的输出中选择正确的语法结构。此外,该工具还具有通过建议选择顺序来引导用户的能力。按照这个顺序,用户可以有效地注释句子。2001年,我们从EDR语料库中提取了语法规则,这是日本现存最大的语料库之一。EDR语料库的缺点是缺少对语料库进行注释所基于的语法。因此,我们首先从EDR语料库中自动提取语法,并对其进行改进,使语法的歧义尽可能小。此外,我们还提出了一种新的语义知识构建框架,该框架不仅在语义分析中有重要作用,而且在句法分析中也有重要作用。从头开始构建语义知识是很困难的,因此我们采取了一种结合已有语义知识的方法。2002年,我们继续开展2001年开始的两个专题工作。除此之外,我们还构建了一个标注语料库管理系统。该系统允许用户高效地检索各种语法结构。语句中的结构存储在关系数据库系统中,为用户提供了多种检索功能。为了验证上述研究结果,我们建立了一个由大约2万个句子组成的日语语料库。这个句子集摘自EDR语料库。该语料库基于从EDR语料库中提取的语法,并在本项目中进行了改进。使用本项目开发的标注工具对语料库进行标注,并通过上述系统对生成的语料库进行管理。少
英文摘要
Since the middle of 1980's, natural language processing based on a large scale linguistic data has become a main stream in this research area. For this kind of research, linguistic resources play important role, and there has been many attempts to create various kinds of resources. This research project aims to construct an environment to create syntactically annotated Japanese corpora in large scale. To achieve this goal, we conducted the research in the following topics.In 2000, we built an annotation tool which supports a user to annotate syntactic structure on sentences in interactive way. This tool works with an existing parser and the user cab efficiently select a correct syntactic structure from a number of parser's output. In addition, the tool has an ability to navigate the user by suggesting the order of choices. Following this order, the user can efficiently annotate sentences.In 2001, we extracted grammar rules from the EDR corpus, which is one of the existing largest Japan … More ese coypus. The drawback of the EDR corpus is that the grammar based on which the corpus is annotated is missing. Thus we first extract the grammar from the EDR corpus automatically and improve it so that the ambiguities of the grammar became as small as possible.In addition, we proposed a new framework to build semantic knowledge which plays important role not only in semantic analysis but also in syntactic analysis. It is difficult to build semantic knowledge from scratch, therefore we took an approach to combine existing semantic knowledge.In 2002, we continued to work on the two topics started in 2001. In addition to this, we constructed a management system of annotated corpora. This system allows users to retrieve various kinds of syntactic structures efficiently. The structures in a sentence are stored in a relational database system, providing users versatile retrieve capability.In order to verify the results of above research, we built a Japanese corpus consisting of about 20,000 sentences. This sentence set is an excerpt from the EDR corpus. This corpus is based on the grammar extracted from the EDR corpus and improved in this project. To annotate the corpus, the annotation tool developed in this project was used, and the resultant corpus was managed by the system mentioned above. Less
期刊论文(40)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
徳永健伸, 阿辺川武: "統計情報による連体修飾節の解析"日本語学. 20・12. 20-27 (2001)
Takenobu Tokunaga、Takeshi Abekawa:“使用统计信息的主语修饰语从句分析”《日语研究》20・12(2001)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Mino, H., Hasimoto, T. Tokunaga, T and Tanaka, H.: "Disambiguation of adverbial phrase attachment by using decision tree"Annual meeting of Association of Natural Language Processing. 411-414 (2002)
Mino, H.、Hasimoto, T. Tokunaga, T 和 Tanaka, H.:“使用决策树消除状语短语附件的歧义”自然语言处理协会年会。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Sirai, K., Ueki, M., Hasimoto, T., Tokunaga, T. and Tanaka H.: "The MSLR parser : A toolkit of natural language processing"Natural Language Processing. 7, No. 5. 93-112 (2000)
Sirai, K.、Ueki, M.、Hasimoto, T.、Tokunaga, T. 和 Tanaka H.:“MSLR 解析器:自然语言处理工具包”自然语言处理。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
美野秀弥, 橋本泰一, 徳永健伸, 田中穂積: "決定リストを利用した形容動詞の修飾先の決定"言語処理学会第8回年次大会予稿集. 411-414 (2002)
Hideya Mino、Taiichi Hashimoto、Kennobu Tokunaga、Hozumi Tanaka:“使用决策列表确定形容词动词的修饰语”语言处理学会第八届年会论文集 411-414(2002 年)。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
Noro, T., Okazaki, A., Tokunaga, T. and Tanaka, H.: "A study on large Japanese grammar development"Annual meeting of Association of Natural Language Processing. 387-390 (2002)
Noro, T.、Okazaki, A.、Tokunaga, T. 和 Tanaka, H.:“大型日语语法发展研究”自然语言处理协会年会。
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[]
通讯作者:
共 16 条
Understanding and generation of referring expressions using gaze
-
批准号:21300049
-
项目类别:Grant-in-Aid for Scientific Research (B)
-
资助金额:$12.06万
-
财政年份:2009
-
负责人:TOKUNAGA Takenobu
-
依托单位:
Understanding referring expression in dialogue with embodied agents
-
批准号:19500116
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$2.91万
-
财政年份:2007
-
负责人:TOKUNAGA Takenobu
-
依托单位:
海外基金