A Framework for Collecting YouTube Meta-Data

A Framework for Collecting YouTube Meta-Data
复制标题

收集 YouTube 元数据的框架

DOI:
10.1016/j.procs.2017.08.347
复制
发表时间:
2017
影响因子:
0.9
通讯作者:
Zifeng Tian
Zifeng Tian
中科院分区:
医学4区
文献类型:
--
作者:
Haroon Malik;Zifeng Tian

文献摘要

被引文献

相似文献

YouTube是目前最流行和最成功的视频分享网站。YouTube上的视频已经成为数据的宝藏,可以用于从STEM教育到医学科学的各个研究领域。然而,到目前为止,所有以前的研究,研究和分析都是在非常小的YouTube视频数据上进行的。到目前为止,还没有任何机制可以系统地、持续地收集、处理和存储YouTube丰富的数据集。在本文中,我们提出了一种方法来填补差距,即,系统地、持续地挖掘和存储YouTube数据。该方法有两个模块,视频发现和视频元数据收集。我们的方法是强大的,高效的和可扩展的。在两个月的时间里,使用我们的方法,我们发现了16,000,000个视频,并挖掘了超过42,000个视频的完整元数据。
YouTube is currently the most popular and successful video sharing website. The videos on YouTube have become a treasure of data, which can be used in various fields of research ranging from STEM education to Medical science. However, all the previous research, studies, and analysis so far, are only conducted on very small volume of YouTube video data. To date, no mechanism exists to systematically and continuously collect, process and store the rich set of YouTube data. In this paper, we present a methodology to fill the gap, i.e., systematically and continuously mine and store the YouTube data. The methodology has two modules, a video discovery and a video meta-data collection. Our methodology is robust, efficient and scalable. Over the period of two months, using our methodology, we discovered 16,000,000 videos and mined the complete meta-data of more than 42,000 videos.