Understanding the annotation process: annotation for Big data
Understanding the annotation process: annotation for Big data
批准号:
AH/L010364/1
负责人:
Robert Villa
金额:
$10.29万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2014
资助国家:
英国
项目状态:
已结题
起止时间:
2014 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Data is being collected and created at the fastest rate in human history; by the far the vast majority of this is in digital format. Allied with this, what was previously "offline" information can now be digitised quickly and cheaply e.g. old manuscripts, maps etc. This vast collection of existing and new information creates new opportunities and also difficulties. For a lot of this information to be useful it must be categorised and annotated in some way, so that sense can be made of the data and also so that the correct data can be accessed more easily. It is possible to complete this categorisation by hand with human annotators, but this effort can be expensive in terms of time, money and resources. This is especially true for large data sets or for data sets that require niche expertise to annotate. With this expense in mind, many have turned to machine learning to annotate data; however machine learning approaches still require human intervention to both create training sets for algorithms and judge the output of algorithms. Thus it is inevitable that human intervention is involved at some stage of the categorisation and annotation process. In this project we aim to gain a better understanding of this annotation process so that we can provide guidelines, approaches and processes for providing the most cost effective and accurate annotations for data sets. We propose to work with the three main types of unstructured data faced in big data: text, image, and video. The first challenge is to better understand the process assessors go through when annotating and judging different types of material. This will be carried out using a mixture of qualitative and quantitative techniques, using smaller scale lab-based studies. By better understanding the process by which individuals annotate and classify material, we hope to provide insights which can be used to make the annotation process more efficient, and identify a set of initial factors which affect annotation performance, such as degree of domain expertise and time. Based on this initial work, the aim is to then investigate which of these factors most affect assessment, using large scale crowdsourcing style methods. The final challenge is related to the classification task: how should annotation be approached, to give the best results when used in machine learning? Based on this, the project aims to create a set of guidelines for the creation of annotation and relevance sets.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Video Test Collection with Graded Relevance Assessments
具有分级相关性评估的视频测试集
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[Qiying, W]
通讯作者:
Qiying, W
Augmented Test Collections: A Step in the Right Direction
增强测试集:朝着正确方向迈出的一步
DOI:
10.48550/arxiv.1501.06370
发表时间:
2015
期刊:
arXiv e-prints
影响因子:
--
作者:
[Hasler Laura]
通讯作者:
Hasler Laura
SIGIR 2014 workshop on gathering efficient assessments of relevance (GEAR)
SIGIR 2014 年收集有效相关性评估研讨会 (GEAR)
DOI:
10.1145/2600428.2600735
发表时间:
2014
期刊:
影响因子:
--
作者:
[Halvey M]
通讯作者:
Halvey M
A Comparison of Primary and Secondary Relevance Judgements for Real-Life Topics
现实生活主题的主要和次要相关性判断的比较
DOI:
--
发表时间:
期刊:
影响因子:
--
作者:
[Wakeling, S]
通讯作者:
Wakeling, S
DOI:
10.1145/2600428.2609466
发表时间:
2014
期刊:
影响因子:
--
作者:
[Halvey M]
通讯作者:
Halvey M
海外基金