Reflections on 20,000 Victorian Newspapers: ‘Distant Reading’ The Times using The Times Digital Archive

Reflections on 20,000 Victorian Newspapers: ‘Distant Reading’ The Times using The Times Digital Archive
复制标题

对 20,000 份维多利亚报纸的反思:使用《泰晤士报数字档案》“远读”《泰晤士报》

DOI:
10.1080/13555502.2012.683151
复制
发表时间:
2012
影响因子:
0.3
通讯作者:
D. Liddle
D. Liddle
中科院分区:
人文科学4区
文献类型:
--
作者:
D. Liddle

文献摘要

被引文献

相似文献

维多利亚时代的学者和印刷文化历史学家会去报纸数据库,如19世纪大英图书馆报纸、英国报纸档案馆和泰晤士报数字档案馆,寻找文本形式的信息,但这些数据库也以数字形式给我们提供信息。例如,圣engage的Times Digital Archive的用户可以看到与我们的搜索参数匹配的文章的数量(“点击率”),我们阅读的每篇文章的字数,甚至我们下载的可移植文档格式(pdf)页面图像的文件大小。这些数字显然不是历史报纸原始文本的一部分,而是由光学字符识别(OCR)软件、搜索引擎软件或计算机操作系统计算出来的与文本内容一模一样的描述性数据——“元数据”。大多数人文学科的研究实践都忽略了元数据,但在《Style, Inc.: Reflections on 7000 Titles》一书中,佛朗哥·莫雷蒂(Franco Moretti)证明了一些关于数据的数据具有阐明文学史趋势的潜力。莫雷蒂对大约7000本英国小说的标题进行了统计和绘图,这些小说创作于18世纪中期至19世纪中期,莫雷蒂展示了英国出版业向信息密集的标题惯例的重要演变。他建议用“远读”一词来形容这种通过寻找数字抽象来揭示文本质量和模式的方法,这种方法研究了大量无法阅读的历史文本。在这篇短文中,我考虑报纸数据库产生的元数据是否有可能帮助我们“远读”英国新闻业历史的各个方面。我将只描述实验,仅略高于信封背面的水平,并且不会对报纸本身提出什么强有力的主张,而是提供初步的观察和可视化
Victorianists and print culture historians go to newspaper databases such as 19th Century British Library Newspapers, The British Newspaper Archive, and The Times Digital Archive looking for information in the form of text, but these databases also give us information in the form of numbers. Users of Cengage’s Times Digital Archive, for example, are shown the count of articles that match our search parameters (‘hits’), the count of words in each article we read, and even the file sizes of the portable document format (pdf) page images we download. 1 Such numbers are obviously not part of the original text of historical newspapers, but descriptive data at one remove from textual content–‘metadata’–calculated by optical character recognition (OCR) software, search engine software, or computer operating systems. Most research praxis in the humanities ignores metadata, but in ‘Style, Inc.: Reflections on 7,000 Titles’, Franco Moretti has demonstrated that some data about data has potential to illuminate trends in literary history. 2 Counting and graphing the words in the titles of some 7000 British novels written between the mid-eighteenth and mid-nineteenth centuries, Moretti has shown an important evolution in British publishing toward more information-dense titling conventions. He has suggested the term ‘distant reading’for this method of investigating unreadably large amounts of historical text by finding numerical abstractions that can reveal qualities and patterns within those texts. 3In this short article I consider whether any of the metadata generated by newspaper databases might have potential to help us ‘distant-read’aspects of the history of British journalism. I will describe experiments only, barely above the backof-the-envelope level, and will put forward few strong claims about newspapers themselves, instead offering preliminary observations and visualizations of what