Automatic Classification of Wikipedia Articles by Using Convolutional Neural Network

Automatic Classification of Wikipedia Articles by Using Convolutional Neural Network
复制标题

DOI:
--
复制
发表时间:
2017
期刊:
International Journal of Service and Knowledge Management
影响因子:
--
通讯作者:
Keita Tsuji
Keita Tsuji
中科院分区:
其他
文献类型:
--
作者:
Keita Tsuji

文献摘要

相似文献

维基百科已成为大学生的重要信息来源。据报道,学生们倾向于从谷歌开始搜索,甚至在大学图书馆也能找到维基百科的文章。最近的研究结果表明,真正搜索和阅读书籍的学生相对较少。在此背景下,我们现在正在开发一个系统,根据维基百科用户在图书馆阅读的文章推荐书籍(该系统将添加到图书馆台式电脑的网络浏览器中)。这样的系统旨在鼓励学生阅读图书馆书籍作为更可靠的信息来源,而不是依赖维基百科文章。日本十进制分类 (NDC) 类别被发现是一种有效的图书推荐机器学习方法。因此,如果可以将 NDC 类别分配给维基百科文章,它们可能会被用作图书推荐的有效工具。因此,我们开发了一种利用卷积神经网络(CNN)自动为维基百科文章分配NDC类别的方法,这是深度学习的代表性方法之一。我们发现NDC分配顶级(即主类)和二级(即主类和分区的组合)的准确率分别达到87.7%和74.7%。这些结果是通过使用维基百科文章的标题和类别作为 CNN 的输入来实现的,而标题、类别和主要文本等其他组合获得的准确性相对较差。
Wikipedia has emerged as an important source of information for university students. It has been reported that the students tend to start their search with Google that leads to Wikipedia articles even in university libraries. Recent research findings indicate that relatively few students actually search and read books. Within this context, we are now developing a system that recommends books based on the articles Wikipedia users read in libraries (the system will be added to web browsers of libraries’ desktop PCs). Such a system aims to encourage students to read library books as a more reliable source of information rather than relying on Wikipedia articles. Nippon Decimal Classification (NDC) categories are found to be an effective machine learning method for book recommendation. Therefore, if NDC categories could be assigned to Wikipedia articles, they might be used as an effective tool for book recommendation. Accordingly, we developed a method to automatically assign NDC categories to Wikipedia articles by using convolutional neural network (CNN), which is one of the representative methods of deep learning. We found that the accuracy of assigning top-level (i.e. Main Class) and second-level (i.e. combination of Main Class and Division) of NDC reached 87.7% and 74.7%, respectively. These results were achieved by using titles and categories of Wikipedia articles as input to CNN, while the accuracies obtained by other combinations such as titles, categories, and main texts were relatively poor.