FanfictionNLP: A Text Processing Pipeline for Fanfiction

FanfictionNLP: A Text Processing Pipeline for Fanfiction
复制标题

DOI:
10.18653/v1/2021.nuse-1.2
复制
发表时间:
2021-06
期刊:
Proceedings of the Third Workshop on Narrative Understanding
影响因子:
--
通讯作者:
Michael Miller Yoder;Sopan Khosla;Qinlan Shen;Aakanksha Naik;Huiming Jin;Hariharan Muralidharan;C. Rosé
Michael Miller Yoder;Sopan Khosla;Qinlan Shen;Aakanksha Naik;Huiming Jin;Hariharan Muralidharan;C. Rosé
中科院分区:
其他
文献类型:
--
作者:
Michael Miller Yoder;Sopan Khosla;Qinlan Shen;Aakanksha Naik;Huiming Jin;Hariharan Muralidharan;C. Rosé

文献摘要

被引文献

相似文献

同人小说为NLP、教育和社会科学研究提供了一个数据源的机会。然而,用这些数据回答具体的研究问题是困难的,因为同人小说比正式小说包含更多样化的写作风格。我们提出了一个用于同人小说的文本处理管道,重点是识别与角色相关的文本。管道包括字符识别和相互参照的模块,以及引用和叙述这些字符的属性。此外,管道包含一种新的字符共指方法,该方法使用引用属性中的知识来解决引用中的代词。对于每个模块,我们评估了10个注释的同人小说故事的各种方法的有效性。在角色共指和引用归属的任务上,这条管道优于为正式小说开发的工具
Fanfiction presents an opportunity as a data source for research in NLP, education, and social science. However, answering specific research questions with this data is difficult, since fanfiction contains more diverse writing styles than formal fiction. We present a text processing pipeline for fanfiction, with a focus on identifying text associated with characters. The pipeline includes modules for character identification and coreference, as well as the attribution of quotes and narration to those characters. Additionally, the pipeline contains a novel approach to character coreference that uses knowledge from quote attribution to resolve pronouns within quotes. For each module, we evaluate the effectiveness of various approaches on 10 annotated fanfiction stories. This pipeline outperforms tools developed for formal fiction on the tasks of character coreference and quote attribution