Improving the Identification of the Discourse Function of News Article Paragraphs

Improving the Identification of the Discourse Function of News Article Paragraphs
复制标题

DOI:
10.18653/v1/2020.nuse-1.3
复制
发表时间:
2020-07
期刊:
--
影响因子:
--
通讯作者:
Deya Banisakher;W. V. Yarlott;Mohammed Aldawsari;N. Rishe;Mark A. Finlayson
Deya Banisakher;W. V. Yarlott;Mohammed Aldawsari;N. Rishe;Mark A. Finlayson
中科院分区:
其他
文献类型:
--
作者:
Deya Banisakher;W. V. Yarlott;Mohammed Aldawsari;N. Rishe;Mark A. Finlayson

文献摘要

相似文献

识别文档的语篇结构是理解书面语篇的一项重要任务。在以前工作的基础上,我们展示了一种改进的方法来自动识别新闻文章中段落的话语功能。我们首先从货车Dijk(1988)提出的新闻语篇层次理论入手,该理论提出了段落在新闻文章中的功能。这种话语信息是介于短语或句子大小的话语片段和文档体裁之间的中间层次,表征了各个段落如何传达关于文章故事情节中事件的信息。具体而言,该理论将叙述事件与(1)整体故事情节(如主要事件,背景或后果)以及(2)评论(如口头反应和评价)之间的关系进行分类。我们训练和测试了一个具有新特征的线性链条件随机场(CRF)来模拟货车Dijk的标签,并将其与以前工作中提出的几种机器学习模型进行了比较。我们的模型显著优于所有基线和先前的方法,平均达到0.71 F1分数,比先前表现最好的支持向量机模型提高了31.5%。
Identifying the discourse structure of documents is an important task in understanding written text. Building on prior work, we demonstrate an improved approach to automatically identifying the discourse function of paragraphs in news articles. We start with the hierarchical theory of news discourse developed by van Dijk (1988) which proposes how paragraphs function within news articles. This discourse information is a level intermediate between phrase- or sentence-sized discourse segments and document genre, characterizing how individual paragraphs convey information about the events in the storyline of the article. Specifically, the theory categorizes the relationships between narrated events and (1) the overall storyline (such as Main Events, Background, or Consequences) as well as (2) commentary (such as Verbal Reactions and Evaluations). We trained and tested a linear chain conditional random field (CRF) with new features to model van Dijk’s labels and compared it against several machine learning models presented in previous work. Our model significantly outperformed all baselines and prior approaches, achieving an average of 0.71 F1 score which represents a 31.5% improvement over the previously best-performing support vector machine model.