课题基金 / 基金详情

CAREER: Robust and Secure Multi-Modal Learning for Library-Scale Text Collections

CAREER: Robust and Secure Multi-Modal Learning for Library-Scale Text Collections
职业:图书馆规模文本收藏的稳健且安全的多模式学习
批准号:
1652536
负责人:
David Mimno
金额:
$55.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-05-15 至 2024-04-30

项目摘要

项目成果

David Mimno的其他基金

相似基金

相关文献

中文摘要
翻译
社交媒体和数字化图书馆的发展使得计算文本分析成为现代学术研究的重要工具。但是,为专家用户提供标准化数据集的方法往往无法转化为现实世界的数据分析。为了使文本挖掘方法有用,需要在理论能力和实际应用之间取得平衡。真实的数据集是嘈杂和复杂的。更重要的是,由于版权的原因,包括1923年以后出版的所有书籍在内的大量数据无法直接共享。该项目将开发可应用于有限的、私有化的文件视图的工具。算法将专注于可靠性和效率,因此强大的技术可以被非专业用户在容易访问的硬件上使用,例如1000万K-12学生使用基于浏览器的低功耗chromebook,从而增加了这项工作的社会影响。主题模型和词嵌入等无监督文本挖掘方法在机器学习之外变得流行,因为它们在简单、广泛可用的表示上操作,并识别代表可识别主题、事件或概念的潜在变量。但是标准算法不能很好地扩展,需要完全访问潜在敏感的文本集合,并且不能利用图像等非文本数据。尽管最近在光谱推断方面的工作已经提高了速度,但目前的方法受到对噪声观测的敏感性的困扰。这项工作将开发一种基于矩阵和张量分解的统一的无监督文本挖掘方法。该项目将侧重于输入矩阵的数据校正方法,使简单的算法能够更好地工作,即使在存在稀疏和嘈杂的观测数据的情况下,同时也减少了模型的不确定性。该项目将通过创建非公开数据的公开视图,开发从私人和敏感文件中学习的新方法。这些将包括单个文档的嘈杂表示以及语料库级摘要矩阵,并支持强不可识别性和弱非表达性标准。最后,该项目将开发新的图像和文本建模工具,优化图像在真实语料库中的实际陪伴方式,而不是简短的人工字幕。通过对大量文本和语义相关图像进行联合建模,该项目将使用户能够搜索与上下文相关的图像,而不仅仅是视觉上相似的图像,并识别基于视觉世界而不仅仅是文本的主题。欲了解更多信息,请参阅项目网页:http://mimno.infosci.cornell.edu
英文摘要
The growth of social media and digitized libraries has made computational text analysis a vital tool for modern scholarship. But too often methods that work on standardized collections for expert users don't translate to real-world data analysis. In order to be useful, text mining methodologies need to balance theoretical power with practical application. Real data sets are noisy and complicated. More importantly, vast amounts of data cannot be shared directly due to copyright, including all published books after 1923. This project will develop tools that can be applied to limited, privatized views of documents. Algorithms will focus on reliability and efficiency, so that powerful techniques can be used by non-expert users on easily accessible hardware, such as the 10 million K-12 students using low-powered browser-based Chromebooks thereby increasing the societal impact of the work.Unsupervised text mining methods such as topic models and word embeddings have become popular outside of machine learning because they operate on simple, widely-available representations and identify latent variables that represent recognizable themes, events, or concepts. But standard algorithms do not scale well, require full access to potentially sensitive text collections, and cannot take advantage of non-textual data such as images. Although recent work in spectral inference has produced improvements in speed, current methods are plagued by sensitivity to noisy observations. This work will develop a unified approach to unsupervised text mining based on matrix and tensor factorization. The project will focus on data rectification methods for input matrices, enabling simple algorithms to work dramatically better, even in the presence of sparse and noisy observations, while also reducing model uncertainty. The project will develop new methods for learning from private and sensitive documents by creating public views of non-public data. These will include both noisy representations of individual documents as well as corpus-level summary matrices, and support both strong non-identifiability and weaker non-expressivity criteria. Finally, the project will develop new tools for modeling images and text optimized for the way images actually accompany text in real corpora, rather than short, artificial captions. By jointly modeling large volumes of text and semantically related images, the project will enable users to search for contextually related images, not just visually similar images, and identify topics that are grounded in the visual world, not just in text. For further information see the project web page: http://mimno.infosci.cornell.edu
期刊论文(13)
专著(0)
科研奖励(0)
会议论文
DOI: 10.18653/v1/d17-1308
发表时间: 2017-09
期刊:
影响因子: --
作者: [David Mimno;Laure Thompson]
通讯作者: David Mimno;Laure Thompson
DOI: 10.18653/v1/2021.emnlp-main.449
发表时间: 2021-09
期刊: ArXiv
影响因子: --
作者: [Gregory Yauney;David M. Mimno]
通讯作者: Gregory Yauney;David M. Mimno
Computational Cut-Ups: The Influence of Dada
计算剪切:达达主义的影响
DOI: --
发表时间: 2018
期刊: Journal of modern periodical studies
影响因子: 0.3
作者: [Thompson, Laure, Mimno, David]
通讯作者: Mimno, David
DOI: --
发表时间: 2019
期刊:
影响因子: --
作者: [Alexandra Schofield]
通讯作者: Alexandra Schofield
共 10 条
    Conference: Text As Data Conference 2022
    • 批准号:
      2232664
    • 项目类别:
      Standard Grant
    • 资助金额:
      $2.0万
    • 财政年份:
      2022
    • 负责人:
      David Mimno
    • 依托单位:
    国内基金
    海外基金
    供应链管理中的稳健型(Robust)策略分析和稳健型优化(Robust Optimization )方法研究
    • 批准号:
      70601028
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      7.0万元
    • 批准年份:
      2006
    • 负责人:
      王明征
    • 依托单位:
    心理紧张和应力影响下Robust语音识别方法研究
    • 批准号:
      60085001
    • 项目类别:
      专项基金项目
    • 资助金额:
      14.0万元
    • 批准年份:
      2000
    • 负责人:
      韩纪庆
    • 依托单位:
    ROBUST语音识别方法的研究
    • 批准号:
      69075008
    • 项目类别:
      面上项目
    • 资助金额:
      3.5万元
    • 批准年份:
      1990
    • 负责人:
      高雨青
    • 依托单位:
    改进型ROBUST序贯检测技术
    • 批准号:
      68671030
    • 项目类别:
      面上项目
    • 资助金额:
      2.0万元
    • 批准年份:
      1986
    • 负责人:
      刘有恒
    • 依托单位: