Large-scale high-precision topic modeling on twitter

Large-scale high-precision topic modeling on twitter
复制标题

DOI:
10.1145/2623330.2623336
复制
发表时间:
2014-08
期刊:
Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining
影响因子:
--
通讯作者:
Shuang-Hong Yang;A. Kolcz;A. Schlaikjer;Pankaj Gupta
Shuang-Hong Yang;A. Kolcz;A. Schlaikjer;Pankaj Gupta
中科院分区:
其他
文献类型:
--
作者:
Shuang-Hong Yang;A. Kolcz;A. Schlaikjer;Pankaj Gupta

文献摘要

被引文献

相似文献

我们感兴趣的是在真实的时间内组织一个连续的稀疏和嘈杂的文本流,称为“推文”,进入一个本体的数百个主题,可测量的和严格的高精度。这种推断是在一个完整的Twitter数据流上进行的,其统计分布随着时间的推移而迅速演变。在工业环境中实施,有可能影响到真实的用户并为他们所见,因此有必要克服一系列实际挑战。我们提出了一系列有助于部署系统的主题建模技术。这些包括非主题推文检测,自动标记数据采集,人工计算评估,诊断和纠正学习,最重要的是,高精度主题推理。后者代表了一种新颖的推文文本分类两阶段训练算法和一种将文本与其他信息源相结合的闭环推理机制。由此产生的系统达到93%的精度在实质性的整体覆盖。
We are interested in organizing a continuous stream of sparse and noisy texts, known as "tweets", in real time into an ontology of hundreds of topics with measurable and stringently high precision. This inference is performed over a full-scale stream of Twitter data, whose statistical distribution evolves rapidly over time. The implementation in an industrial setting with the potential of affecting and being visible to real users made it necessary to overcome a host of practical challenges. We present a spectrum of topic modeling techniques that contribute to a deployed system. These include non-topical tweet detection, automatic labeled data acquisition, evaluation with human computation, diagnostic and corrective learning and, most importantly, high-precision topic inference. The latter represents a novel two-stage training algorithm for tweet text classification and a close-loop inference mechanism for combining texts with additional sources of information. The resulting system achieves 93% precision at substantial overall coverage.