Integrating Crowdsourcing and Active Learning for Classification of Work-Life Events from Tweets

Integrating Crowdsourcing and Active Learning for Classification of Work-Life Events from Tweets
复制标题

集成众包和主动学习,对推文中的工作生活事件进行分类

DOI:
10.1007/978-3-030-55789-8_30
复制
发表时间:
2020
期刊:
Trends in Artificial Intelligence Theory and Applications. Artificial Intelligence Practices
影响因子:
--
通讯作者:
Bian, Jiang
Bian, Jiang
中科院分区:
--
文献类型:
--
作者:
Zhao, Yunpeng;Prosperi, Mattia;Lyu, Tianchen;Guo, Yi;Zhou, Le;Bian, Jiang

文献摘要

相似文献

社交媒体,尤其是Twitter,正越来越多地用于预测分析研究。在社交媒体研究中,自然语言处理(NLP)技术与基于专家的手动和定性分析结合使用。然而,社交媒体数据是非结构化的,必须经过复杂的操作才能用于研究。手动注释是多个专家评分员必须就每个项目达成共识的最耗费资源和时间的过程,但对于创建用于训练基于NLP的机器学习分类器的黄金标准数据集至关重要。为了减轻手动注释的负担,同时保持其可靠性,我们设计了一个结合主动学习策略的众包管道。我们通过一个案例研究证明了它的有效性,该案例研究从个人推文中识别失业事件。我们使用Amazon Mechanical Turk平台从互联网上招募注释员,并设计了一些质量控制措施来确保注释的准确性。我们评估了4种不同的主动学习策略(即,最小置信度、熵、投票熵和Kullback-Leibler散度)。主动学习策略的目的是减少所需的推文的数量,以达到预期的自动分类性能。结果表明,众包是有用的,以创建高质量的注释和积极的学习有助于减少所需的推文的数量,虽然有测试的策略之间没有实质性的差异。
Social media, especially Twitter, is being increasingly used for research with predictive analytics. In social media studies, natural language processing (NLP) techniques are used in conjunction with expert-based, manual and qualitative analyses. However, social media data are unstructured and must undergo complex manipulation for research use. The manual annotation is the most resource and time-consuming process that multiple expert raters have to reach consensus on every item, but is essential to create gold-standard datasets for training NLP-based machine learning classifiers. To reduce the burden of the manual annotation, yet maintaining its reliability, we devised a crowdsourcing pipeline combined with active learning strategies. We demonstrated its effectiveness through a case study that identifies job loss events from individual tweets. We used Amazon Mechanical Turk platform to recruit annotators from the Internet and designed a number of quality control measures to assure annotation accuracy. We evaluated 4 different active learning strategies (i.e., least confident, entropy, vote entropy, and Kullback-Leibler divergence). The active learning strategies aim at reducing the number of tweets needed to reach a desired performance of automated classification. Results show that crowdsourcing is useful to create high-quality annotations and active learning helps in reducing the number of required tweets, although there was no substantial difference among the strategies tested.