Gender Classification with Data Independent Features in Multiple Languages
Gender Classification with Data Independent Features in Multiple Languages
复制标题
多种语言的具有数据独立特征的性别分类
DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
Katie Cohen
中科院分区:
文献类型:
--
作者:
T. Isbister;Lisa Kaati;Katie Cohen
Gender classification is a well-researched problem, and state-of-the-art implementations achieve an accuracy of over 85%. However, most previous work has focused on gender classification of texts written in the English language, and in many cases, the results cannot be transferred to different datasets since the features used to train the machine learning models are dependent on the data. In this work, we investigate the possibilities to classify the gender of an author on five different languages: English, Swedish, French, Spanish, and Russian. We use features of the word counting program Linguistic Inquiry and Word Count (LIWC) with the benefit that these features are independent of the dataset. Our results show that by using machine learning with features from LIWC, we can obtain an accuracy of 79% and 73% depending on the language. We also, show some interesting differences between the uses of certain categories among the genders in different languages.