课题基金 / 基金详情

ITR: Machine Learning from Labeled and Unlabeled Data

ITR: Machine Learning from Labeled and Unlabeled Data
ITR:从标记和未标记数据进行机器学习
批准号:
0312814
负责人:
John Lafferty
金额:
$39.5万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2003
资助国家:
美国
项目状态:
已结题
起止时间:
2003-09-01 至 2006-08-31

项目摘要

项目成果

John Lafferty的其他基金

相似基金

相关文献

中文摘要
翻译
该项目研究了如何在机器学习中最有效地使用未标记数据和标记数据的基本问题。这项工作的目标有三个方面。首先,这项研究的目的是对这个问题有一个基本的理解,包括对未标记数据所能提供的信息进行推理的新方法。其次,本研究探索了将大量未标记数据与少量标记数据和背景知识结合使用的新算法,以获得大大超过仅使用标记数据和更传统方法的性能。研究人员使用的方法包括图算法和随机场、蒙特卡罗采样和光谱方法,这些方法与计算机科学密切相关,已在计算机视觉中得到应用,但尚未在机器学习中得到充分利用。最后,将研究目标应用,包括文本分析,图像分类和计算机安全的入侵检测,以验证所开发的理论原理,并探索算法并提出新的研究方向。这项研究的更广泛的影响将是帮助新技术利用在如此多的新领域和如此大的规模上收集的大量数据。我们对标记和未标记数据结合的可能性和基本限制的理解取得了进步,这有可能影响许多科学领域,使研究人员能够更容易地使用大量可用的数据,但这些数据不一定是为他们自己的特定需求而注释的。它也可能最终影响我们社会选择投资的未来数据收集计划。
英文摘要
This project investigates the basic question of how unlabeled data can be most effectively used together with labeled data in machine learning. The goals of this work are three-fold. First, the research aims to achieve a fundamental understanding of this problem, including new methods for reasoning about the kind of information unlabeled data can provide. Second, this research explores new algorithms for using large amounts of unlabeled data together with small amounts of labeled data and background knowledge, in order to achieve performance that greatly exceeds that available using only labeled data and more traditional methods today. The approaches used by the investigators include graph algorithms and random fields, Monte Carlo sampling and spectral methods, closely connected areas of computer science that have found application in computer vision, but that have yet to be fully exploited in machine learning. Finally, targeted applications, including text analysis, image classification, and intrusion detection for computer security, will be investigated to validate the theoretical principles that are developed, and to explore algorithms and suggest new directions for investigation.The broader impact of this research will be to help enable new technologies to use the volumes of data that are being collected in so many new domains, and on such a great scale. Advances in our understanding of the possibilities for, and fundamental limits to, combining labeled and unlabeled data has the potential to impact many scientific fields, allowing researchers to more easily use the vast quantities of data that are available but not necessarily annotated for their own specific needs. It also may ultimately influence the future data collection initiatives that our society chooses to invest in.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Generative Models for Complex Data: Inference, Sensing, and Repair
  • 批准号:
    2015397
  • 项目类别:
    Standard Grant
  • 资助金额:
    $25.0万
  • 财政年份:
    2020
  • 负责人:
    John Lafferty
  • 依托单位:
Constrained Statistical Estimation and Inference: Theory, Algorithms and Applications
  • 批准号:
    1748444
  • 项目类别:
    Standard Grant
  • 资助金额:
    $14.5万
  • 财政年份:
    2017
  • 负责人:
    John Lafferty
  • 依托单位:
Constrained Statistical Estimation and Inference: Theory, Algorithms and Applications
  • 批准号:
    1513594
  • 项目类别:
    Standard Grant
  • 资助金额:
    $32.0万
  • 财政年份:
    2015
  • 负责人:
    John Lafferty
  • 依托单位:
MSPA-MCS: Nonparametric Learning in High Dimensions
  • 批准号:
    0625879
  • 项目类别:
    Standard Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2006
  • 负责人:
    John Lafferty
  • 依托单位:
国内基金
海外基金
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位: