课题基金 / 基金详情

RI: Small: Collaborative Research: Unsupervised Transcription of Early Modern Documents

RI: Small: Collaborative Research: Unsupervised Transcription of Early Modern Documents
RI:小型:合作研究:早期现代文献的无监督转录
批准号:
1618044
负责人:
Taylor Berg-Kirkpatrick
金额:
$24.95万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-09-01 至 2021-08-31

项目摘要

项目成果

Taylor Berg-Kirkpatrick的其他基金

相似基金

相关文献

中文摘要
翻译
最近,社会科学和人文科学的研究人员在他们的工作中越来越多地使用数字技术,试图根据新的分析来回答关于人类文物的重要问题。然而,由于它们的许多方法本质上是统计的,它们需要大量的数字可读文本来操作。例如,要询问有关妇女法律权利在过去五个世纪中如何变化的统计问题,必须以数字形式提供跨越这一时期的大量公正的法庭诉讼程序样本。不幸的是,在许多时期,这些数据是不可用的,不是因为历史文件已经丢失,而是因为它们不能有效地转录。特别是印刷机发明后的400年(现代早期,约1450-1850年)是这类研究的关键黑暗时期,因为这一时期的文件很难用自动方法转录成机器可读的文本,原因有三:它们使用模糊和未知的字体,它们的文本与现代语言不同,历史上的印刷工艺不精确。该提案试图通过将转录视为一种密码破译类型来解决这些问题,并使用机器学习直接从未注释的文档图像中诱导字体和文本结构,而不依赖于注释的示例,这种方法称为无监督学习。因此,该提案不仅旨在数字化主要图书馆中现有的早期现代语料库,而且还旨在生产一种工具,研究人员可以使用它来大规模地数字化数据,并且该工具足够灵活,可以开发新的表示,例如,非标准字符集。提出的方法将文档转录问题视为语言解码问题,利用解密历史密码工作中的建模技术。关键的思想是,虽然像字体和文本结构这样的属性是特定于文档的,因此很难用一般的监督技术来处理,但这些现象实际上在单个文档中是常规的。例如,虽然系统可能不知道某个模糊的历史字体中的特定字符的形状,但该形状实际上是规则的;每次打印字符时都使用相同的模板。利用这种规律的模型将其作为一个假设,可以约束其他困难的无监督学习问题,并使其可行。本提案引入了一类基于此目标的生成模型,旨在通过捕获生成输入数据的过程的核心属性(历史打印过程),以无监督的方式学习字体并预测准确的转录。这些模型代表了早期现代文献所表现出的特定类型的印刷和排版噪声,将排版作为一个潜在变量,并在推理过程中共同考虑可能的字符分割和转录。它们的参数可以有效地估计,直接从历史文件的图像,而不附带转录。此外,通过将输入文档的损坏部分作为潜在变量处理,该建议旨在使用相同的方法自动重建损坏的文档。这里开发的无监督技术可能在自然语言处理的其他领域中使用,在这些领域中很难获得带注释的训练数据;例如,在个性化语音识别和接地语义。
英文摘要
Recently, researchers in the social sciences and humanities have made increasing use of digital technologies in their work, seeking to answer important questions about human artifacts based on new kinds of analyses. However, since many of their methods are statistical in nature, they require a large amount of digitally readable text to operate. For example, to ask statistical questions about how the legal rights of women have changed during the past five centuries, a large and unbiased sample of court proceedings spanning that time period has to be accessible in digital form. Unfortunately, for many time periods this data is not available, not because the historical documents have been lost, but because they cannot be efficiently transcribed. In particular, the 400 years just after the invention of the printing press (the early modern period, ca. 1450-1850) represents a critical dark period for such research because documents from this period are notoriously hard to transcribe into machine-readable text with automatic methods for three reasons: they use obscure and unknown fonts, their text differs from modern language, and historical printing processes were imprecise. This proposal seeks to address these issues by treating transcription as a type of code-breaking and using machine learning to induce font and text structure directly from unannotated document images without relying on annotated examples, an approach called unsupervised learning. As a result, the proposal aims not just to digitize existing early modern corpora in major libraries, but also to produce a tool that researchers can use to digitize data at scale themselves and that is sufficiently flexible to develop new representations, for example, of non-standard character sets. The proposed approach treats the problem of document transcription as a linguistic decipherment problem, leveraging modeling techniques from work on decrypting historical ciphers. The key idea is that while properties like font and text structure are document-specific and therefore difficult to treat generally with supervised techniques, these phenomena are in fact regular within individual documents. For example, while the shape of a particular character in an obscure historical font may be unknown to the system, that shape is in fact regular; every time the character is printed it uses the same template. Models that leverage this kind of regularity by incorporating it as an assumption can constrain the otherwise difficult unsupervised learning problem and make it feasible. This proposal introduces a class of generative models with this goal in mind, designed to learn fonts and predict accurate transcriptions in an unsupervised fashion by capturing the core properties of the process that generated the input data: the historical printing process. These models represent the specific types of printing and typesetting noise exhibited by early modern documents, treat typesetting as a latent variable, and jointly consider possible character segmentations and transcriptions during inference. Their parameters can be estimated efficiently, directly from images of historical documents without accompanying transcriptions. Further, by treating damaged portions of the input documents as latent variables, this proposal aims to automatically reconstruct damaged documents using the same approach. The unsupervised techniques developed here may have uses in other areas of natural language processing where annotated training data is hard to obtain; for example, in personalized speech recognition and grounded semantics.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Modeling Language Evolution via Deep Probabilistic Factorization
  • 批准号:
    2146151
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $60.0万
  • 财政年份:
    2022
  • 负责人:
    Taylor Berg-Kirkpatrick
  • 依托单位:
Collaborative Research: RI: Small: Unsupervised Islamicate Manuscript Transcription via Lacunae Reconstruction
  • 批准号:
    2200333
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2022
  • 负责人:
    Taylor Berg-Kirkpatrick
  • 依托单位:
RI: Small: Print and Probability - A Statistical Approach to Analysis of Clandestine Publication
  • 批准号:
    1936155
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.98万
  • 财政年份:
    2019
  • 负责人:
    Taylor Berg-Kirkpatrick
  • 依托单位:
RI: Small: Print and Probability - A Statistical Approach to Analysis of Clandestine Publication
  • 批准号:
    1816311
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.98万
  • 财政年份:
    2018
  • 负责人:
    Taylor Berg-Kirkpatrick
  • 依托单位:
国内基金
海外基金
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    张祥忠
  • 依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
  • 批准号:
    31972324
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    高学文
  • 依托单位: