'What is this corpus about?': using topic modelling to explore a specialised corpus

'What is this corpus about?': using topic modelling to explore a specialised corpus
复制标题

DOI:
10.3366/cor.2017.0118
复制
发表时间:
2017-08-01
期刊:
影响因子:
0.5
通讯作者:
Vajn, Dominik
Vajn, Dominik
中科院分区:
其他
文献类型:
--
作者:
Murakami, Akira;Thompson, Paul;Vajn, Dominik

文献摘要

被引文献

相似文献

本文介绍了主题建模,这是一种自动识别给定语料库中“主题”的机器学习技术。该论文阐述了它在学术英语语料库探索中的用途。它首先对主题建模的底层机制进行了直观的解释,并描述了构建模型的过程,包括模型构建过程中涉及的决策。然后本文探讨了该模型。主题模型中的主题的特征是一组同时出现的单词,我们将证明这些主题为我们带来了对语料库性质的丰富见解。作为示范任务,本文识别了论文不同部分的突出主题,调查了期刊的时间变化,并揭示了期刊中不同类型的论文。本文进一步将主题建模与语料库语言学中两种更传统的技术——语义标注和关键词分析进行了比较,并强调了主题建模的优势。我们相信主题建模在语料库的初始探索中特别有用。
This paper introduces topic modelling, a machine learning technique that automatically identifies 'topics' in a given corpus. The paper illustrates its use in the exploration of a corpus of academic English. It first offers the intuitive explanation of the underlying mechanism of topic modelling and describes the procedure for building a model, including the decisions involved in the model-building process. The paper then explores the model. A topic in topic models is characterised by a set of co-occurring words, and we will demonstrate that such topics bring us rich insights into the nature of a corpus. As exemplary tasks, this paper identifies the prominent topics in different parts of papers, investigates the chronological change of a journal, and reveals different types of papers in the journal. The paper further compares topic modelling to two more traditional techniques in corpus linguistics, semantic annotation and keywords analysis, and highlights the strengths of topic modelling. We believe that topic modelling is particularly useful in the initial exploration of a corpus.