Creating anaphorically annotated resources through semantic wikis (AnaWiki)
Creating anaphorically annotated resources through semantic wikis (AnaWiki)
批准号:
EP/F00575X/1
负责人:
Massimo Poesio
金额:
$18.26万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2007
资助国家:
英国
项目状态:
已结题
起止时间:
2007 至 --
中文摘要
在自然语言处理方面取得进展的能力--无论是开发更好的自然语言处理系统,还是开发关于人类如何处理语言的更好的理论--取决于大型标注语料库的可用性:用人类判断标注的文档集合,比如,在特定上下文中,‘bank’或‘stock’等模棱两可的词的解释是什么,或者像‘语料库’这样的回指短语的解释是什么。因此,当前用于语义信息标注的语料库不够大,并且没有收集足够多的主题的判断,这是NLP的主要障碍。然而,用目前的方法创建更大的手写标注语料库是非常昂贵和耗时的;在实践中,想要标注超过100万个单词是不可行的。在文献中已经提出了各种通过半自动标注来解决问题的技术,例如自举和主动学习;然而,它们的有效性还没有得到令人信服的证明。然而,维基百科的成功表明,另一种方法可能是可能的:利用网络人口在协作资源创建工作中做出贡献的意愿。这种意愿已经被利用来通过ESP游戏来标记图像;我们建议开发工具,使Web上的大量志愿者能够合作创建语义标注的语料库(特别是用共指信息标注的语料库)。在这方面,我们将在现有努力的基础上开发MediaWiki版本,以支持语义Web上的工作,并自行开发可靠且易于遵循的指令,以标记关于回指的语义判断。至少,这些工具将使NLP研究人员社区自己能够合作创建一个回指银行。然而,我们还将运行一个试验性的开发方法来吸引Web社区的兴趣;如果这些测试成功,我们可能能够通过Web使用协作努力的力量来创建真正大型的带注释的语料库。我们将采取的方法的一个显著特点是,我们将允许志愿者标记语义判断的差异,并对之前表达的语义判断发表评论,以识别那些意见广泛一致的判断和那些存在分歧的判断。
英文摘要
The ability to make progress in Natural Language Processing - both to develop better NLP systems and to develop better theories of how humans process language - depends on the availability of large annotated corpora: collections of documents annotated with human judgments about, say, what is the interpretation of ambiguous words such as 'bank' or 'stock' in a particular context, or what is the interpretation of anaphoric expressions like 'the corpus'. So the fact that current corpora annotated for semantic information are not large enough and do not collect the judgments of a large enough number of subjects is a major obstacle for NLP. Creating larger hand-annotated corpora with the current methods, however, is very expensive and time consuming; in practice, it is unfeasible to think of annotating more than 1M words. A variety of techniques for solving the problem by semi-automatic annotation have been proposed in the literature, such as bootstrapping and active learning; however, their usefulness has not yet been convincingly demonstrated. However, the success of Wikipedia shows that another approach might be possible: take advantage of the willingness of the Web population to contribute in collaborative resource creation efforts. This willingness has already been harnessed to tag images through the ESP game; we propose to develop tools that will make it possible for large numbers of volunteers over the Web to collaborate in the creation of semantically annotated corpora (specifically, of a corpus annotated with coreference information) . In this, we will build on existing efforts to develop versions of MediaWiki to support work on the Semantic Web, and on our own to develop reliable and easy-to-follow instructions for marking semantic judgments about anaphora. At the very least, these tools will make it possible for the community of NLP researchers themselves to collaborate in the creation of an Anaphoric Bank. We will however also run a pilot developing methods to attract the interest of the Web community at large; if these tests are successful, we may be able to use the power of collaborative effort through the Web to create really large annotated corpora. A distinctive feature of the approach we will adopt is that we will allow volunteers to mark differences in semantic judgments, and to express comments on previously expressed semantic judgments, so as to identify those judgments on which there is wide agreement and ones on which there is disagreement.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1515/itit-2017-0020
发表时间:
2018
期刊:
it - Information Technology
影响因子:
--
作者:
[Chamberlain J]
通讯作者:
Chamberlain J
DOI:
10.1145/1600150.1600156
发表时间:
2009-06
期刊:
影响因子:
--
作者:
[Jon Chamberlain;Massimo Poesio;Udo Kruschwitz]
通讯作者:
Jon Chamberlain;Massimo Poesio;Udo Kruschwitz
Aggregating crowdsourced and automatic judgments to scale up a corpus of anaphoric reference for fiction and Wikipedia texts
聚合众包和自动判断,以扩大小说和维基百科文本的照应参考语料库
DOI:
--
发表时间:
2023
期刊:
影响因子:
--
作者:
[Yu, J]
通讯作者:
Yu, J
The annotation-validation (AV) model
注释验证(AV)模型
DOI:
10.1145/2594776.2594779
发表时间:
2014
期刊:
影响因子:
--
作者:
[Chamberlain J]
通讯作者:
Chamberlain J
ANAWIKI: Creating anaphorically annotated resources through Web cooperation
ANAWIKI:通过网络合作创建照应注释资源
DOI:
--
发表时间:
2008
期刊:
Proceedings of the 6th International Conference on Language Resources and Evaluation, LREC 2008
影响因子:
--
作者:
[Poesio M.]
通讯作者:
Poesio M.
共 6 条
Annotating Reference and Coreference In Dialogue Using Conversational Agents in games
-
批准号:EP/W001632/1
-
项目类别:Research Grant
-
资助金额:$139.06万
-
财政年份:2022
-
负责人:Massimo Poesio
-
依托单位: