课题基金 / 基金详情

Understanding the annotation process: annotation for Big data

Understanding the annotation process: annotation for Big data
了解标注过程:大数据标注
批准号:
AH/L010364/1
负责人:
Robert Villa
金额:
$10.29万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2014
资助国家:
英国
项目状态:
已结题
起止时间:
2014 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
数据正在以人类历史上最快的速度收集和创建;到目前为止,绝大多数数据都是数字格式的。与此相结合,以前的“离线”信息现在可以快速、廉价地数字化,例如旧手稿、地图等。对于许多有用的信息,必须以某种方式对其进行分类和注释,以便使数据具有意义,并且可以更容易地访问正确的数据。可以通过人工注释来完成这种分类,但这种努力在时间,金钱和资源方面可能是昂贵的。对于大型数据集或需要专业知识进行注释的数据集尤其如此。考虑到这一成本,许多人转向机器学习来注释数据;然而,机器学习方法仍然需要人为干预来创建算法的训练集并判断算法的输出。因此,在分类和注释过程的某个阶段,不可避免地涉及人为干预。在这个项目中,我们的目标是更好地了解这个注释过程,以便我们可以提供指导方针,方法和过程,为数据集提供最具成本效益和准确的注释。我们建议使用大数据中面临的三种主要类型的非结构化数据:文本,图像和视频。第一个挑战是更好地理解评估员在注释和判断不同类型的材料时所经历的过程。这将使用定性和定量技术的混合物,使用较小规模的实验室研究进行。通过更好地理解个人注释和分类材料的过程,我们希望提供可用于使注释过程更有效的见解,并确定一组影响注释性能的初始因素,例如领域专业知识和时间的程度。基于这一初步工作,目的是然后调查这些因素中哪些最影响评估,使用大规模众包式的方法。最后一个挑战与分类任务有关:应该如何处理注释,以便在机器学习中使用时提供最佳结果?在此基础上,该项目旨在创建一套用于创建注释和相关性集的指南。
英文摘要
Data is being collected and created at the fastest rate in human history; by the far the vast majority of this is in digital format. Allied with this, what was previously "offline" information can now be digitised quickly and cheaply e.g. old manuscripts, maps etc. This vast collection of existing and new information creates new opportunities and also difficulties. For a lot of this information to be useful it must be categorised and annotated in some way, so that sense can be made of the data and also so that the correct data can be accessed more easily. It is possible to complete this categorisation by hand with human annotators, but this effort can be expensive in terms of time, money and resources. This is especially true for large data sets or for data sets that require niche expertise to annotate. With this expense in mind, many have turned to machine learning to annotate data; however machine learning approaches still require human intervention to both create training sets for algorithms and judge the output of algorithms. Thus it is inevitable that human intervention is involved at some stage of the categorisation and annotation process. In this project we aim to gain a better understanding of this annotation process so that we can provide guidelines, approaches and processes for providing the most cost effective and accurate annotations for data sets. We propose to work with the three main types of unstructured data faced in big data: text, image, and video. The first challenge is to better understand the process assessors go through when annotating and judging different types of material. This will be carried out using a mixture of qualitative and quantitative techniques, using smaller scale lab-based studies. By better understanding the process by which individuals annotate and classify material, we hope to provide insights which can be used to make the annotation process more efficient, and identify a set of initial factors which affect annotation performance, such as degree of domain expertise and time. Based on this initial work, the aim is to then investigate which of these factors most affect assessment, using large scale crowdsourcing style methods. The final challenge is related to the classification task: how should annotation be approached, to give the best results when used in machine learning? Based on this, the project aims to create a set of guidelines for the creation of annotation and relevance sets.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
Video Test Collection with Graded Relevance Assessments
具有分级相关性评估的视频测试集
DOI: --
发表时间:
期刊:
影响因子: --
作者: [Qiying, W]
通讯作者: Qiying, W
Augmented Test Collections: A Step in the Right Direction
增强测试集:朝着正确方向迈出的一步
DOI: 10.48550/arxiv.1501.06370
发表时间: 2015
期刊: arXiv e-prints
影响因子: --
作者: [Hasler Laura]
通讯作者: Hasler Laura
SIGIR 2014 workshop on gathering efficient assessments of relevance (GEAR)
SIGIR 2014 年收集有效相关性评估研讨会 (GEAR)
DOI: 10.1145/2600428.2600735
发表时间: 2014
期刊:
影响因子: --
作者: [Halvey M]
通讯作者: Halvey M
Evaluating the effort involved in relevance assessments for images
评估图像相关性评估所涉及的工作
DOI: 10.1145/2600428.2609466
发表时间: 2014
期刊:
影响因子: --
作者: [Halvey M]
通讯作者: Halvey M
海外基金