课题基金 / 基金详情

CI-P: Planning for a Multilingual FrameNet Lexical Resource

CI-P: Planning for a Multilingual FrameNet Lexical Resource
CI-P:规划多语言 FrameNet 词汇资源
批准号:
1406048
负责人:
Collin Baker
金额:
$9.98万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-06-01 至 2016-05-31

项目摘要

项目成果

Collin Baker的其他基金

相似基金

相关文献

中文摘要
翻译
位于加州伯克利的国际计算机科学研究所的框架项目(http://framenet.icsi.berkeley.edu)一直在构建一个独特的日常英语单词词典,它将单词与语义框架连接起来,每个语义框架代表一种事件或关系,以及参与其中的人或物体的角色。例如,亲属框架包括“侄子”、“岳母”、“兄弟姐妹”等词,领导框架包括“下士”、“主教”、“首席执行官”、“校长”等名词和“领导”、“主席”、“规则”等动词。自1997年以来,FrameNet项目已经开发了1100多个这样的框架,涵盖了超过12700个单词的意思,并手工注释了近20万个例子,展示了这些语义角色如何适应句子的语法。这个数据库是免费的,每天都有人下载,供世界各地的NLP应用使用,比如问答和推理事件的因果。其他地方的项目正在为许多其他语言构建框架风格的数据库,包括西班牙语、德语、日语、中文、法语、巴西葡萄牙语、阿拉伯语等。这些其他项目在很大程度上遵循了英语框架网的例子,使用相同的语义框架和角色。他们的结论是,目标语言中大约80%的词汇单位与最初为英语定义的语义框架非常吻合。但是,目前还没有广泛的、系统的努力来整合所有这些数据库,以产生一个免费可用的统一的多语言框架语义资源。该奖项用于计划这样的努力,为将所有这些框架连接到一个多语言数据库奠定基础,这将允许许多新的NLP应用,例如基于框架的机器翻译和识别两种不同语言的新闻账户讨论同一事件。在这个计划项目中,研究人员正在计划如何根据框架名称(其他框架使用英语框架名称或其翻译)和框架、词汇单位和语义角色的相似性的定量测量来对齐单独的框架。定量度量将利用每种语言中由帧对帧关系创建的网络,不仅比较单个帧,还比较它们在网络中的邻居,它们对应的语义角色,以及跨语言中相应角色的填充符的相似性。通过这种方式,将在完全不同的语法结构(价模式)之间建立对应关系,其中这些角色跨语言实现,类似于通过投影在其他语言中创建框架所使用的方法。该计划包括面对面和虚拟会议,以定义社区需求和优先级,并收集研究人员和开发人员在校准方法、应用程序接口等方面的建议。作为该项目的结果,研究人员计划向CISE研究基础设施计划准备一份全面的提案,以实际创建多语言数据库,并继续与开发各种框架的团队以及已经在研究和实际应用中使用框架的团队进行磋商。
英文摘要
The FrameNet Project (http://framenet.icsi.berkeley.edu) at the International Computer Science Institute in Berkeley, California, has been building a unique dictionary of everyday English words, which connects words with semantic frames, each representing a type of event or relation and the roles of people or objects that participate in it. For example, the Kinship frame includes the words 'nephew', 'mother-in-law', 'sibling', etc. and the Leadership frame includes both nouns like 'corporal', 'bishop', 'CEO', and 'headmaster' and verbs like 'lead', 'preside', and 'rule'. Since 1997, the FrameNet project has developed more than 1,100 such frames covering more than 12,700 word senses, and manually annotated almost 200,000 examples showing how these semantic roles fit into the grammar of the sentences. The database is freely available and is being downloaded daily for use around the world, in NLP applications like question answering and reasoning about the causes and effects of events. Projects elsewhere are building FrameNet-style databases for many other languages, including Spanish, German, Japanese, Chinese, French, Brazilian Portuguese, Arabic, etc. These other projects have largely followed the English FrameNet example, using the same semantic frames and roles. Their conclusion has been that roughly 80% of the lexical units in the target languages fit nicely into semantic frames that were originally defined for English. But there has been no broad, systematic effort to align all these databases to produce a freely available unified multilingual frame semantic resource. This award is used to plan such an effort, to lay the groundwork for connecting all these FrameNets into one multilingual database, which will permit many new NLP applications, such as frame-based machine translation and recognizing when news accounts in two different languages are discussing the same event.During this planning project, the investigators are planning how to go about aligning the separate FrameNets, based in part on the frame names (other FrameNets either use the English frame names or translations of them) and in part on quantitative measures of the similarity of frames, lexical units, and semantic roles. The quantitative measures will exploit the networks created by frame-to-frame relations in each language, comparing not only individual frames, but also their neighbors in the network, their corresponding semantic roles, and the similarity of the fillers of corresponding roles across languages. In this way, correspondences will be established between the quite different syntactic constructions (valence patterns) in which these roles are realized across languages, similar to the methodology used in creating FrameNets in other languages by projection. The plan includes both face-to-face and virtual meetings to define community requirements and priorities and to gather researchers and developers' suggestions on alignment methods, application interfaces, etc. As a result of this project, the investigators plan to prepare a comprehensive proposal to the CISE Research Infrastructure Program to actually create the multilingual database, guided by continuing consultation both with the teams developing the various FrameNets and with those who are already using FrameNet in research and practical applications.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Berkeley FrameNet Website Migration
CI-NEW: Multilingual FrameNet: A Resource Enabling Cross-Lingual Research for the Natural Language Processing Community
FrameNet Workshop: Developing New NLP Applications
CI-P: Collaborative Research: LexLink: Aligning WordNet, FrameNet, PropBank and VerbNet
海外基金