Heroes, Villains, and the In-Between: A Natural Language Processing Approach to Fairy Tales

Heroes, Villains, and the In-Between: A Natural Language Processing Approach to Fairy Tales
复制标题

英雄、恶棍和中间人:童话故事的自然语言处理方法

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Ruby Alling Ostrow
Ruby Alling Ostrow
中科院分区:
--
文献类型:
--
作者:
Ruby Alling Ostrow

文献摘要

被引文献

相似文献

虽然在过去的几十年里,自然语言处理(NLP)技术已经取得了很大的进步,但对于将NLP用于小说类型的研究却明显缺乏。这个项目试图通过考虑使用NLP技术来总结欧洲童话来解决这个问题。由于其原型人物和相对简单的故事情节,这一小说亚类型是一个适当的研究起点。我的方法是提取文本中的主要人物,沿着以修饰形容词和人物参与的言语行为为形式的关键描述符。通过这种方法,我建议我们如何通过跟踪字符与某些语言事件的概率关联来将字符解析为Proppian原型。这种分类模式反过来又使童话的更广泛的分类成为可能。该模型的整体F1分数为0.77,各个部分的F1分数分别为0.89、0.75和0.66,用于字符检索、形容词提取和动词提取。这个项目还可以进一步扩展,为进一步自动化人物分类和最终故事本身奠定关键基础。
While great strides have been made with natural language processing (NLP) techniques in the last few decades, there has been a notable lack of research into utilizing NLP for the genre of fiction. This project seeks to address this gap by considering the use of NLP techniques for the summarization of European fairy tales. This subgenre of fiction is an appropriate starting point for investigation due to its archetypal characters and relatively simple story arcs. My approach is to extract the main characters of texts, along with key descriptors in the form of modifying adjectives and verbal actions the characters take part in. Through this method, I suggest how we may parse characters into Proppian archetypes by tracking their probabilistic association with certain linguistic occurrences. This classification schema in turn makes possible the broader classification of fairy tales into types. The model has an overall F1 score of 0.77, the individual parts having F1 scores of 0.89, 0.75, and 0.66 for character retrieval, adjective extraction, and verb extraction, respectively. This project may also be extended further, laying key groundwork for further automatization of categorization of characters and ultimately stories themselves.