Automatically Profiling the Author of an Anonymous Text

Automatically Profiling the Author of an Anonymous Text
复制标题

DOI:
10.1145/1461928.1461959
复制
发表时间:
2009-02-01
影响因子:
22.7
通讯作者:
Schler, Jonathan
Schler, Jonathan
中科院分区:
计算机科学3区
文献类型:
--
作者:
Argamon, Shlomo;Koppel, Moshe;Schler, Jonathan

文献摘要

被引文献

相似文献

想象一下,你得到了一篇作者未知的重要文章,你希望通过分析这篇文章,尽可能多地了解这位作者(人口统计、个性、文化背景等)。作者身份分析问题在当前的全球信息环境中变得越来越重要——在取证、安全和商业环境中应用广泛。例如,当需要考虑的具体嫌疑人太少(或太多)时,作者身份分析可以帮助警察识别犯罪者的特征。类似地,大公司可能有兴趣知道什么样的人喜欢或不喜欢他们的产品,基于对博客和在线产品评论的分析。因此,我们要问的问题是:仅仅通过分析文本本身,我们能在多大程度上辨别出文本的作者?事实证明,在不同程度的准确性下,我们确实可以得出很多结论。与Li、Zheng和chen最近讨论的作者归属问题(从给定的候选集合中确定文本的作者)不同,作者身份分析不是从已知候选作者的一组写作样本开始的。相反,我们利用社会语言学的观察,不同群体的人用一种特定的体裁和一种特定的语言说话或写作,使用这种语言是不同的。也就是说,他们在使用某些单词或句法结构的频率上有所不同(例如,除了发音或语调的变化之外)。我们在这里考虑的特定概况维度是作者性别,1年龄,8土著地区
ImagIne that you have been gIven an Important text of unknown authorship, and wish to know as much as possible about the unknown author (demographics, personality, cultural background, among others), just by analyzing the given text. This authorship profiling problem is of growing importance in the current global information environment–applications abound in forensics, security, and commercial settings. For example, authorship profiling can help police identify characteristics of the perpetrator of a crime when there are too few (or too many) specific suspects to consider. Similarly, large corporations may be interested in knowing what types of people like or dislike their products, based on analysis of blogs and online product reviews. The question we therefore ask is: How much can we discern about the author of a text simply by analyzing the text itself? It turns out that, with varying degrees of accuracy, we can say a great deal indeed.Unlike the problem of authorship attribution (determining the author of a text from a given candidate set) discussed recently in these pages by Li, Zheng, and Chen9 authorship profiling does not begin with a set of writing samples from known candidate authors. Instead, we exploit the sociolinguistic observation that different groups of people speaking or writing in a particular genre and in a particular language use that language differently. 2 That is, they vary in how often they use certain words or syntactic constructions (in addition to variation in pronunciation or intonation, for example). The particular profile dimensions we consider here are author gender, 1 age, 8 native lan-