EAGER: Human Computation: Integrating the Crowd and the Machine
EAGER: Human Computation: Integrating the Crowd and the Machine
批准号:
1145291
负责人:
Albert Lin
金额:
$6.6万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2011
资助国家:
美国
项目状态:
已结题
起止时间:
2011-08-01 至 2013-07-31
中文摘要
由于数字技术,今天的信息和连接比以往任何时候都更容易获得,现在可以通过招募大量人口统计数据来解决问题,以补充计算机计算的局限性。这在视觉分析的情况下尤为重要,在视觉分析中,人类的直觉仍然远远优于现有的计算机对象识别算法。虽然算法受到预标记要求的限制,但人类可以感知细微的变化和细微差别,以识别和分类意外对象。然而,这些任务的规模往往太大,一个人无法完成。将这项任务分配到一个庞大的网络上,不仅可以成功地对数据进行分类,还可以生成大量的人类量词(训练数据),从而有可能教会计算机视觉算法模仿人类感知,以区分正常和异常。这个探索性项目将把人类集体视觉感知与机器学习和物体识别结合起来,通过对6000多名志愿者提供的125万份群众输入数据进行研究,为蒙古北部的异常情况进行标记。这些数据是通过PI与国家地理数字媒体合作开发的在线平台从2010年6月至今收集的,提供了一个理想的“案例研究”环境来调查人群生成数据的性质,以及将人类输入的广泛变异性提炼成计算算法的方法。在线参与者对发现成吉思汗墓的可能性感到兴奋,他们检查了大量超高分辨率多光谱卫星图像,将松散定义的异常标记为不同类别。从大量的标签中出现的趋势代表了人类对图像所包含内容的集体观点。由PI领导的一个团队前往蒙古实地考察用户输入高度趋同的地区。由此产生的地面真实异常提供了一个独特的机会,既可以准确测量人类/自动化分析的质量,也可以研究在机器学习中使用小池绝对数据补充嘈杂的人群数据集的效果。在目前的项目中,PI将开发一个框架,用于应用和评估以下三个研究阶段,旨在研究大规模人类生成数据的性质,以整合到监督学习算法中。共识聚类——基于相邻标签的数量和一致性以及创建这些标签的个体能力的标签评估机制。无监督的“合并”标签方法也将适用于扩展的异常,如道路和河流。特征向量提取——检测异常所需的特征类型(如颜色、亮度、边缘和梯度、尺度、方向等)和邻域范围(如局部、广泛和全局)都是先验未知的。因此,目标是确定足够多样的特征,以捕获图像中的所有相关线索。机器学习-将根据上述第2阶段的结果确定给定类别像素组的代表和排除的主导特征。更广泛的影响:在这项探索性研究中,PI将为从人群资源中提取新的机器/人类合作机会奠定基础。理解人类和计算机智能之间的联系将对许多科学分支产生深远的影响。因此,通过将基于项目的分布式分析工具的众包迁移到连接人类集体感知和机器学习的门户,在此努力中开发的概念可能最终证明具有变革性。
英文摘要
Because both information and connectivity are more available today than ever before thanks to digital technologies, questions can now be addressed by enlisting massive human demographics to supplement the limitations of computer computation. This is especially relevant in the case of visual analytics, where human intuition remains far superior to existing computer object recognition algorithms. While algorithms are limited by pre-labeling requirements, humans can perceive subtle variations and nuances to identify and classify unexpected objects. These tasks, however, are often too massive in scale for a single human to accomplish. Distributing this task over a massive network not only succeeds in categorizing data, but generates massive quantities of human quantifiers (training data) to potentially teach computer vision algorithms to mimic human perception in order to distinguish the normal from the abnormal.This exploratory project will combine collective human visual perception with machine learning and object recognition, through a study of 1.25 million crowd-sourced inputs provided by over 6,000 volunteers labeling satellite imagery in a search for anomalies in northern Mongolia. These data, collected from June 2010 to the present via an online platform developed by the PI in collaboration with National Geographic Digital Media, afford an ideal "case study" environment to investigate the nature of crowd generated data and methods that distill the wide variability of human input into computational algorithms. The online participants, excited by the potential of discovering the tomb of Genghis Khan, examined massive amounts of ultra-high resolution multispectral satellite imagery to label loosely defined anomalies into various categories. Trends that emerged from the massive volume of labels represent a collective human perspective on what the images contain. A team led by the PI traveled to Mongolia to ground-truth areas of high user input convergence. The resulting ground-truthed anomalies provide a unique opportunity to both accurately measure the quality of human/automated analysis and to investigate the effect of supplementing noisy crowd-sourced data sets with small pools of absolute data in machine learning. In the current project the PI will develop a framework for applying and evaluating the following three research phases designed to study the nature of large scale human generated data for integration into supervised learning algorithms:1. Consensus Clustering - Tag evaluation mechanisms based upon the volume and consistency of neighboring tags and the ability of the individuals creating those tags. Unsupervised methods for "merging" labels will also be applied for extended anomalies such as roads and rivers.2. Feature Vector Extraction - Both the type of features (e.g., color, luminance, edges and gradients, scale, orientation, etc.) and the extent of the neighborhoods (e.g., local, wide and global) required to detect anomalies are unknown a priori. Thus, the aim is to determine sufficiently diverse features to capture all relevant cues within the image.3. Machine Learning - Dominant features representative of, and excluded from, pixel groups of given categories will be determined from the results of Phase 2 above.Broader Impacts: In this exploratory study the PI will lay the foundation for extracting new machine/human collaborative opportunities from the resource of the crowd. Understanding the bonds between human and computer intelligence will have a profound impact on many branches of science. Thus, concepts developed in this effort may ultimately prove transformative by affording migration of crowd-sourcing from a project-based tool for distributed analytics into a portal bridging collective human perception and machine learning.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
HCC: Small: Examining the Super User versus the Crowd in Human-Centered Computation
-
批准号:1219138
-
项目类别:Continuing Grant
-
资助金额:$49.73万
-
财政年份:2012
-
负责人:Albert Lin
-
依托单位:
Research Initiation: Response of Full Scale Thin Concrete Shells to Transient Vibration
-
批准号:8503993
-
项目类别:Standard Grant
-
资助金额:$6.8万
-
财政年份:1985
-
负责人:Albert Lin
-
依托单位:
国内基金
海外基金
靶向Human ZAG蛋白的降糖小分子化合物筛选以及疗效观察
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:胡文静
-
依托单位:
HBV S-Human ESPL1融合基因在慢性乙型肝炎发病进程中的分子机制研究
-
批准号:81960115
-
项目类别:地区科学基金项目
-
资助金额:34.0万元
-
批准年份:2019
-
负责人:江建宁
-
依托单位:
基于自适应表面肌电模型的下肢康复机器人“Human-in-Loop”控制研究
-
批准号:61005070
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2010
-
负责人:李庆玲
-
依托单位: