E-mail address categorization based on semantics of surnames

E-mail address categorization based on semantics of surnames
复制标题

基于姓氏语义的电子邮件地址分类

DOI:
--
复制
发表时间:
2013
期刊:
IEEE Symposium on Computational Intelligence and Data Mining
影响因子:
--
通讯作者:
M. Rajarajan
M. Rajarajan
中科院分区:
--
文献类型:
--
作者:
Suresh Veluru;Y. Rahulamathavan;P. Viswanath;P. Longley;M. Rajarajan

文献摘要

被引文献

相似文献

姓氏(姓氏)分析在地理学中被用来了解人口的起源、迁移、身份、社会规范和文化习俗。其中一些被认为是经过几代人进化而来的。姓氏具有良好的统计特性,可用于提取姓名数据集中的信息,例如自动检测姓名中的种族或社区群体。一个电子邮件地址,通常包含姓氏作为子字符串。这个容器可以是完全的,也可以是部分的。基于姓氏语义的电子邮件地址分类是本文的研究目标。这可以分两个阶段实现。第一阶段处理姓氏表示和聚类。本文提出了一个向量空间模型,其中进行了潜在语义分析。聚类的方法称为平均链接法。在第二阶段,将电子邮件归类为属于类别之一(在第一阶段发现)。为此,需要子字符串匹配,这是通过使用后缀树数据结构以一种有效的方式完成的。我们对印度和英国500个最常见的姓氏进行了实验评估。此外,我们将具有这些姓氏的电子邮件地址分类为子字符串。
Surname (family name) analysis is used in geography to understand population origins, migration, identity, social norms and cultural customs. Some of these are supposedly evolved over generations. Surnames exhibit good statistical properties that can be used to extract information in names data set such as automatic detection of ethnic or community groups in names. An e-mail address, often contains surname as a substring. This containment may be full or partial. An e-mail address categorization based on semantics of surnames is the objective of this paper. This is achieved in two phases. First phase deals with surname representation and clustering. Here, a vector space model is proposed where latent semantic analysis is performed. Clustering is done using the method called average-linkage method. In the second phase, an email is categorized as belonging to one of the categories (discovered in first phase). For this, substring matching is required, which is done in an efficient way by using suffix tree data structure. We perform experimental evaluation for the 500 most frequently occurring surnames in India and United Kingdom. Also, we categorize the e-mail addresses that have these surnames as substrings.