Exploring a corpus of scientific texts using data mining

Exploring a corpus of scientific texts using data mining
复制标题

使用数据挖掘探索科学文本语料库

DOI:
--
复制
发表时间:
2010
期刊:
影响因子:
--
通讯作者:
Péter Fankhauser
Péter Fankhauser
中科院分区:
--
文献类型:
--
作者:
E. Teich;Péter Fankhauser

文献摘要

被引文献

相似文献

我们报告的项目调查的基础上,从九个学科的期刊文章语料库的英语科技文本的语言特性。该项目的目标是深入了解在计算机科学和其他一些学科(例如,生物信息学、计算语言学、计算工程)。我们在本文中关注的问题是:(a)它所代表的元语域的语料库的特征如何?(B)就它们所例示的更具体的语域而言,子语料库的差异/相似程度如何?我们分析语料库使用几种数据挖掘技术,包括功能排名,聚类和分类,看看如何选择语言特征的子语料库组。结果表明,我们的语料库是很好的区分在元寄存器的科技写作方面,也,我们发现有趣的独特功能的子语料库作为指标的语域多样化。除了展示我们的分析结果,我们还将反思和评估数据挖掘在语料库探索和分析任务中的应用。
We report on a project investigating the linguistic properties of English scientific texts on the basis of a corpus of journal articles from nine academic disciplines. The goal of the project is to gain insights on registers emerging at the boundaries of computer science and some other discipline (e.g., bioinformatics, computational linguistics, computational engineering). The questions we focus on in this paper are (a) how characteristic is the corpus of the meta-register it represents, and (b) how different/similar are the subcorpora in terms of the more specific registers they instantiate? We analyze the corpus using several data-mining techniques, including feature ranking, clustering, and classification, to see how the subcorpora group in terms of selected linguistic features. The results show that our corpus is well distinguished in terms of the meta-register of scientific writing; also, we find interesting distinctive features for the subcorpora as indicators of register diversification. Apart from presenting the results of our analyses, we will also reflect upon and assess the use of data mining for the tasks of corpus exploration and analysis.