N-Gram Approach for Gender Prediction

N-Gram Approach for Gender Prediction
复制标题

用于性别预测的 N-Gram 方法

DOI:
10.1109/iacc.2017.0176
复制
发表时间:
2017
期刊:
2017 IEEE 7th International Advance Computing Conference (IACC)
影响因子:
--
通讯作者:
P. Vijayapal Reddy
P. Vijayapal Reddy
中科院分区:
--
文献类型:
--
作者:
T. Raghunadha Reddy;B. V. Vardhan;P. Vijayapal Reddy

文献摘要

参考文献

被引文献

相似文献

通过博客、Twitter推文、评论、社交媒体网络和其他信息内容,互联网随着大量信息的增长而增长。互联网上的大多数文本都是非结构化和匿名的。作者分析是一种文本分类技术,通过分析作者的文本,预测作者的性别、年龄、国家、母语和教育背景等特征。研究人员提出了不同类型的特征,如词汇特征、内容特征、结构特征和句法特征,以识别作者的写作风格。作者分析中的大多数现有方法使用特征组合来表示用于分类的文档向量。本文提出了一种结合词频n图和最频繁项计算文档权重的新模型。这些文档权重被用来表示用于分类的文档向量。本实验在评论域上进行了作者性别预测,与现有的作者分析方法相比,取得了令人满意的结果。
The Internet was growing with huge amount of information, through Blogs, Twitter tweets, Reviews, social media network and with other information content. Most of the text in the internet was unstructured and anonymous. Author Profiling is a text classification technique that is used to predict the profiling characteristics of the authors like gender, age, country, native language and educational background by analyzing their texts. Researchers proposed different types of features such as lexical, content based, structural and syntactic features to identify the writing styles of the authors. Most of the existing approaches in Author Profiling used the combination of features to represent a document vector for classification. In this paper, a new model was proposed in which document weights were calculated with combination of POS N-grams and most frequent terms. These document weights were used to represent the document vectors for classification. This experiment was carried out on the reviews domain to predict the gender of the authors and the achieved results were promising when compared with the existing approaches in Author Profiling.
DOI: 10.1002/asi.v60:3
发表时间: 2009-03
影响因子: 3.5
作者:
N. Shibata;Y. Kajikawa;Y. Takeda;K. Matsushima
通讯作者: N. Shibata;Y. Kajikawa;Y. Takeda;K. Matsushima