Investigation on University Websites for Semi-automated Syllabus Crawling

Investigation on University Websites for Semi-automated Syllabus Crawling
复制标题

大学网站半自动教学大纲抓取的调查

DOI:
10.1109/fie43999.2019.9028479
复制
发表时间:
2019
期刊:
IEEE Frontiers in Education Conference (FIE)
影响因子:
--
通讯作者:
Yamaguchi Kazunori
Yamaguchi Kazunori
中科院分区:
--
文献类型:
--
作者:
Sekiya Takayuki;Matsuda Yoshitatsu;Yamaguchi Kazunori

文献摘要

相似文献

本文介绍了对近百所大学的网站进行调查的结果,以实现大规模课程教学大纲的半自动抓取。教学大纲提供大学课程的基本信息.对于学生来说,教学大纲是学习一门课程最重要的文件之一,因为学生可以通过教学大纲掌握课程所涵盖的主题。对于学院来说,一套教学大纲有助于理解大学提供的课程。因此,教学大纲是分析教育活动的重要线索之一。已有的研究报道,来自50多所大学的计算机科学(CS)课程的大量课程大纲可以揭示CS课程中有趣的结构。此外,它们还有助于开发一种支持学生和教师的工具。然而,课程大纲是手工收集的。因此,很难大量增加教学大纲的数量。同时也很难更新收集到的教学大纲。大量课程大纲的半自动抓取需要进一步分析。本文调查了2018年泰晤士高等教育世界大学排名(THE2018)中排名前100位的大学网站上的课程大纲页面,其中计算机科学为主题。这些网页是通过简单的爬行和手动选择收集的。由此,我们可以在大约60所大学的网站上找到教学大纲页面。接着,我们发现教学大纲页面的结构可以分为三种类型:链接型、整体型和数据库型。一个链接类型包括一个目录页面,其中有多个教学大纲页面的超链接,每个页面对应一个教学大纲。一个完整的类型包括一个完整的页面,其中包括许多教学大纲在它。一个数据库类型包括一个入口页的数据库,从其中所有的教学大纲可以搜索。链接类型、整体类型和数据库类型的数量分别为31、17和12。此外,我们发现,只有少数关键页(即目录页,整个网页,入口页)可以在一定程度上自动发现的线性支持向量机。特别是目录页面和整个页面可以被相当准确地找到。这些结果预计将是有用的,使半自动抓取的教学大纲从网站。
This Research Paper presents investigation results on the websites of about 100 universities for enabling the semiautomatic crawling of massive course syllabi. A syllabus gives fundamental information about a course in a university. For students, a syllabus is one of the most important documents to take a course, because the students can grasp the topics covered by the course through the syllabus. For faculties, a set of syllabi is useful to understand the curriculum offered by the university. Thus, a syllabus is one of the most important clues in the analysis of the educational activities. Some previous works reported that the massive course syllabi of computer science (CS) curricula from about 50 universities can disclose the interesting structures in the CS curricula. In addition, they were useful to develop a tool for supporting students and faculties. However, the course syllabi were collected manually. Therefore, it was difficult to increase the number of syllabi largely. It was also difficult to keep the collected syllabi up-to-date. The semi-automatic crawling of massive course syllabi is needed for further analysis. In this paper, we investigate the pages of the course syllabi at the websites of the top 100 universities in Times Higher Education World University Rankings 2018 (THE2018) with computer science as subject. The pages were collected by a simple crawling and selected manually. From this, we could find the syllabus pages at the websites of about 60 universities. Then, we discovered that the structures of the syllabus pages can be categorized into three types: Link Type, Whole Type, and Database Type. A Link Type consists of a directory page with the hyperlinks to many syllabus pages, where each page corresponds to one syllabus. A Whole Type consists of a whole page which includes many syllabi in it. A Database Type consists of an entrance page to the database from which all the syllabi can be searched. The numbers of Link Type, Whole Type, and Database Type were 31, 17, and 12, respectively. Furthermore, we found that only a few key pages (namely, the directory pages, the whole pages, and the entrance pages) can be discovered automatically to a certain degree by the linear support vector machine. Especially, the directory pages and the whole pages could be found quite accurately. These results are expected to be useful to enable the semi-automatic crawling of the syllabi from websites.