A German Twitter Snapshot

A German Twitter Snapshot
复制标题

德国推特快照

DOI:
--
复制
发表时间:
2014
期刊:
International Conference on Language Resources and Evaluation
影响因子:
--
通讯作者:
Tatjana Scheffler
Tatjana Scheffler
中科院分区:
--
文献类型:
--
作者:
Tatjana Scheffler

文献摘要

被引文献

相似文献

我们提出了一个新的语料库的德国推文。由于Twitter上的德语消息数量相对较少,因此可以在一段时间内收集德国Twitter消息的几乎完整快照。在本文中,我们提出了我们的收集方法,产生了2400万推文语料库,代表了绝大多数的德国推文发送在2013年4月。此外,我们分析了这一代表性的数据集,并描述了德国twitterverse。虽然德国Twitter数据在时间分布方面与其他Twitter数据相似,但德国Twitter用户更不愿意在他们的推文中分享地理位置信息。最后,语料库收集方法允许在Twitter数据的话语现象的研究,结构化的讨论线程。
We present a new corpus of German tweets. Due to the relatively small number of German messages on Twitter, it is possible to collect a virtually complete snapshot of German twitter messages over a period of time. In this paper, we present our collection method which produced a 24 million tweet corpus, representing a large majority of all German tweets sent in April, 2013. Further, we analyze this representative data set and characterize the German twitterverse. While German Twitter data is similar to other Twitter data in terms of its temporal distribution, German Twitter users are much more reluctant to share geolocation information with their tweets. Finally, the corpus collection method allows for a study of discourse phenomena in the Twitter data, structured into discussion threads.