课题基金 / 基金详情

Unlocking Digital Texts: Towards an Interoperable Text Framework

Unlocking Digital Texts: Towards an Interoperable Text Framework
解锁数字文本:迈向可互操作的文本框架
批准号:
AH/W005638/1
负责人:
Neil Jefferies
金额:
$25.63万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
数字文本项目面临的一个关键挑战是鼓励文化机构、研究人员和广大公众重新使用和利用他们的资源。尽管学术界在创建数字文本方面付出了努力,但数字文本似乎不像其他学科的数据那样具有相同的“长尾”使用模式。主要的障碍之一是文本以难以重用的格式生成和存储。许多包含嵌入底层文本的详细上下文、语义和表示标记。即使这些文本按照诸如文本编码倡议(TEI)之类的可靠标准进行编码,编码的内容和风格也从根本上受到编辑的特定研究领域、语言或文化规范的影响。重用材料通常需要包含这些原则和规范的项目特定代码,有时甚至需要复制交付这些原则和规范的基础结构。这设置了一个很高的门槛,只有最熟练、最坚定和资金充足的研究人员才能超越。我们的目标是通过定义一个可互操作的文本框架(ITF)和实现典型的测试用例来证明它的优势来纠正这种情况。我们并不是要提出一种新的编码或存储文本的格式,而是要提出一种访问和交付文本资源(无论是整个文档还是片段)的方法,这些资源既可以被人类阅读,也可以被机器友好地用于计算分析。当ITF与其他框架(如IIIF和W3C Web Annotation Data Model)结合在一起时,就有可能链接文本、图像、注释和其他在线资源,以构建可以可视化和在线导航的叙述。通过允许围绕文本创建多种新的叙述和分析,而不损害原始文本的完整性,ITF有可能将在线文本和版本转变为活跃的在线话语。合作项目都需要这样的能力,并且已经开发了具体的方法,可以为更通用和灵活的标准的开发提供信息。ITF将使研究塞缪尔·贝克特作品的贝克特数字手稿项目的研究人员能够对他的写作过程构建自己的叙述。他们可以连接、展示和分析贝克特图书馆的书籍片段,这些片段被抄写在笔记本上,贝克特随后在他的作品中互文重用。读者可以看到并比较这些多重叙述,并做出自己的推断。对于数字教学和数字版本,ITF将成为前所未有的全球和地方合作的起点。例如,早期现代数学家托马斯·哈里奥特(Thomas Harriot)丰富但杂乱无章的论文最终将受益于一个灵活的框架,该框架不假设从头到尾都是线性的,而是为不同的读者提供了多个切入点。增强的可导航性和注释将使哈里奥特的论文不仅对全球合作的研究人员来说是清晰的,而且对课堂来说也是如此,在课堂上,教师们可以在他们的文化时刻寻找现成的方法来将数学发现置于背景中。ITF还将使用户能够将计算分析工具应用于来自不同来源的异构文本集合,这通常是避免的,因为它们很难使用。研究人员可以利用现有的文本挖掘和机器学习工具,通过对信件文本和文本创作伙伴关系数字化的参考作品进行比较主题和情感分析,研究早期现代信件在线通信馆藏目录中的引用和参考模式。通过消除技术和基础设施障碍,ITF将有助于确保文本资源能够更好地实现FAIR原则的承诺[https://www.go-fair.org/fair-principles/];它们将是可查找的、可访问的、可互操作的和可重用的。
英文摘要
A key challenge faced by digital text projects is encouraging cultural institutions, researchers and the wider-public to reuse and build upon their resources. Despite the scholarly effort put into creating them, digital texts do not seem to have the same 'long tail' use patterns that data from other disciplines have. One of the chief impediments has been that texts are produced and stored in formats that are hard to reuse. Many contain detailed contextual, semantic and presentational markup embedded with the underlying text. Even when these texts are encoded according to robust standards, such as Text Encoding Initiative (TEI), the content and style of the coding is fundamentally shaped by the editors' specific fields of study, languages or cultural norms. Reusing materials often requires project-specific code that embodies those principles and norms and sometimes even replicating the infrastructure that delivers them. This sets a high bar that only the most skilled, determined and well-funded researchers are able to surmount.We aim to rectify this situation by defining an Interoperable Text Framework (ITF) and implementing exemplary test cases to demonstrate its strengths. We are not proposing a new format for encoding or storing text but rather a method for accessing and delivering textual resources (either whole documents or fragments) that are both readable by humans and also machine-friendly for computational analysis. When ITF is combined with other frameworks, such as IIIF and the W3C Web Annotation Data Model, it becomes possible to link texts, images, annotations and other online resources to construct narratives that can be visualised and navigated online.ITF has the potential to transform online texts and editions into active online discourse by allowing multiple new narratives and analyses to be created around texts, without compromising the integrity of the originals. The partner projects all require such a capability and have already developed specific approaches that can inform the development of a more general and flexible standard.ITF will enable researchers studying Samuel Beckett's works on the Beckett Digital Manuscript Project to construct their own narratives about his writing process. They can connect, display and analyse fragments from books in Beckett's library, where they were copied in notebook(s), and Beckett's subsequent intertextual reuse in his works. Readers can see and compare these multiple narratives and make their own inferences.For digital pedagogy and for digital editions, ITF will be the starting point of unprecedented global and local collaborations. The rich but disorganized papers of the early modern mathematician Thomas Harriot, for instance, will finally benefit from a flexible framework that does not assume linearity from front cover to back cover, but rather enables multiple points of entry for various readers. Enhanced navigability and annotation will make Harriot's papers legible not only for researchers collaborating worldwide but for classrooms, where teachers seek ready ways to contextualize mathematical discoveries within their cultural moments. ITF will also enable users to apply computational analysis tools to heterogeneous collections of text from diverse sources, which would have typically been avoided because they are difficult to use. A researcher could use existing text mining and machine learning tools to study patterns of citation and reference in the correspondence collections catalogues in Early Modern Letters Online by performing comparative topic and sentiment analyses of letter texts and the referenced works digitised by the Text Creation Partnership. By removing the technical and infrastructural barriers, ITF will help to ensure that textual resources will then be better able to live up to the promises of the FAIR principles [https://www.go-fair.org/fair-principles/]; they will be Findable, Accessible, Interoperable, and Reusable.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
超灵敏高分辨的Digital-CRISPR技术用于免扩增的多重核酸检测
  • 批准号:
    22104048
  • 项目类别:
    青年科学基金项目(C类)
  • 资助金额:
    30.0万元
  • 批准年份:
    2021
  • 负责人:
    陈勇
  • 依托单位:
基于Digital Twin的数控机床智能运行维护方法研究
  • 批准号:
    51875323
  • 项目类别:
    面上项目
  • 资助金额:
    60.0万元
  • 批准年份:
    2018
  • 负责人:
    胡天亮
  • 依托单位:
基于数字PCR(digital-PCR)技术的耳聋无创产前检测研究
  • 批准号:
    LQ19H040016
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2018
  • 负责人:
    严恺
  • 依托单位:
基于Digital LAMP技术的循环肿瘤细胞检测和分型新方法研究
  • 批准号:
    81702102
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    20.0万元
  • 批准年份:
    2017
  • 负责人:
    王纪东
  • 依托单位: