Gender Classification with Data Independent Features in Multiple Languages

Gender Classification with Data Independent Features in Multiple Languages
复制标题

多种语言的具有数据独立特征的性别分类

DOI:
--
复制
发表时间:
2017
期刊:
European Intelligence and Security Informatics Conference
影响因子:
--
通讯作者:
Katie Cohen
Katie Cohen
中科院分区:
--
文献类型:
--
作者:
T. Isbister;Lisa Kaati;Katie Cohen

文献摘要

被引文献

相似文献

性别分类是一个经过深入研究的问题,最先进的实现准确率超过 85%。然而,之前的大多数工作都集中在用英语编写的文本的性别分类上,并且在许多情况下,结果无法转移到不同的数据集,因为用于训练机器学习模型的特征依赖于数据。在这项工作中,我们研究了对五种不同语言的作者性别进行分类的可能性:英语、瑞典语、法语、西班牙语和俄语。我们使用字数统计程序语言查询和字数统计 (LIWC) 的功能,优点是这些功能独立于数据集。我们的结果表明,通过使用具有 LIWC 特征的机器学习,我们可以获得 79% 和 73% 的准确率,具体取决于语言。我们还展示了不同语言中性别之间某些类别的使用之间的一些有趣差异。
Gender classification is a well-researched problem, and state-of-the-art implementations achieve an accuracy of over 85%. However, most previous work has focused on gender classification of texts written in the English language, and in many cases, the results cannot be transferred to different datasets since the features used to train the machine learning models are dependent on the data. In this work, we investigate the possibilities to classify the gender of an author on five different languages: English, Swedish, French, Spanish, and Russian. We use features of the word counting program Linguistic Inquiry and Word Count (LIWC) with the benefit that these features are independent of the dataset. Our results show that by using machine learning with features from LIWC, we can obtain an accuracy of 79% and 73% depending on the language. We also, show some interesting differences between the uses of certain categories among the genders in different languages.