Incisive Tagging: Humans-in-the-Loop in Selection and Labelling of Remote Sensing Data Sets
Incisive Tagging: Humans-in-the-Loop in Selection and Labelling of Remote Sensing Data Sets
批准号:
2440657
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
We are developing a pre-trained deep neural network to function 'under the hood' of multiple solutions to extract geospatial information from remote sensing imagery. So far, better results are achieved with larger data sets. However, we observe that much of the imagery contains little apparent unique information and are interested in developing a way to select only the pivotal examples for training. Further, we are keen to work more effectively with our in-house image interpretation experts - using their experience and specific abilities in ways that are rewarding and motivating. Within our work, image interpreters may be labelling data for training networks - and hence may be best deployed labelling the most pivotal examples. Another application is to present image interpreters with multiple examples of image clips that highly activate specific parts of the network and ask them to provide their own interpretation of the representations learned by the deep network. Our questions (not all of which may be addressed in this PhD): Can we improve sample efficiency? Even with an unsupervised target Can we actively label data? Presenting humans with the most pivotal examples for labelling Can we make labelling enjoyable? Using our experts most effectively?Aims and Intended ImpactMore efficient training and updating of machine learning models with remote sensing dataMore efficient and rewarding labelling of training examplesMore human-interpretable neural networksThis project will be grounded in investigating the interplay of human creativity, intelligence and fulfilment with efficient ML tools. The work will begin and iterate around deep and intensive understandings of the labellers' perspectives of the task; leading to prototypes and evaluations. These prototypes might involve novel gamification (e.g. [1]); visualisation techniques; or even the use of multiple modalities - e.g. from simple gestures to emotional state recognition [2] - to provide input to ML tools. The interaction design of the labelling tool will inform and be informed by algorithmic innovations within the ML tool. For instance: One approach to making the task more efficient, immersive and less onerous would be to make spotting of pivotal examples easier. So, for example, we can present samples clustered on similarity as groups of thumbnails, allowing the labeller to spot outliers faster. Another strategy might be to use active labelling - i.e. given a small labelled data set can we present larger sets to users and gain their feedback to (a) label large amounts of data quicklyand (b) resultingly make the labelling task less onerous. We might also consider how to improve sample efficiency - that is, reducing the number of samples required without reducing the efficacy of the approach. Some theoreticalmodels on sample complexity have been investigated [3]. Monte Carlo techniques would be a suggested research direction. For example, Importance Sampling has long been studied in Path Tracing and more advanced techniques such as Hamiltonian Monte Carlo or Gradient Domain [4] demonstrate orders of magnitude performance gains through a great reduction in required samples. Ensembles are used to increase robustness and stability which lend well to importance sampling. We could also examine the literature on robust statistics and M-estimators as methods for drawing samples (see [5] for a recent review). In these sorts of investigation, we will draw on the labelers' experience and insights to supplement any quantitative or theoretical evaluations of the power or limitations of the proposed approaches. The human-centered improvements discussed above could also drive machine performance improvement. Models suffer from needing long training times. There is a potential that the current problem size can be compressed so it just fits into GPU memory toimprove cache coherency during training
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
多语言环境下Social Tagging的内涵机理与应用框架研究-基于比较的视角
-
批准号:71103203
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2011
-
负责人:徐晨
-
依托单位: