课题基金 / 基金详情

On Partially Supervised Text Classification

On Partially Supervised Text Classification
关于部分监督文本分类
批准号:
0307239
负责人:
Bing Liu
金额:
$23.02万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2003
资助国家:
美国
项目状态:
已结题
起止时间:
2003-05-15 至 2006-04-30

项目摘要

项目成果

Bing Liu的其他基金

相似基金

相关文献

中文摘要
翻译
文本分类是将文本文档自动分配到预定义的类别/类。传统上,分类器是使用带有预定义类标记的训练文档构建的(通常是手动的)。然后使用此分类器将新文档分类到这些类中。这个经典模型被称为监督分类,因为训练文档都有预先标记的类。尽管这个模型很重要,但在实践中,它可能不适用于某些常见情况。例如,给定一组特定类别P(正类)的文档用于训练,人们可能想要将包含P类文档和其他类型文档(负类文档)的混合(未标记)文档集M分类为来自P的文档和非来自P的文档。由于没有标记的负类文档,传统的分类技术就不适用了。这个问题被称为部分监督分类(PSC)。现有的求解PSC的方法是基于启发式的,容易出错。研究人员通常使用监督分类的评价方法来评价PSC技术,由于监督分类方法假设了标记阴性文件的可用性,因此这种评价方法是不充分的。该项目的目标是设计一种强大的、原则性的技术来解决PSC问题,实现PSC系统,设计一种方法来评估这些技术,并确定确定达到最佳精度所需的最小标记文件数量的方法,以减少人工标记工作。该项目开发的算法和系统将用于研究生课程,并放在网上供其他研究人员在他们的研究和教学中使用。这项研究的结果应该是广泛有用的,因为目标信息/文件的识别在这个信息时代是非常有价值的。
英文摘要
Text classification is the automated assignment of text documents to pre-defined categories/classes. Traditionally, a classifier is built using training documents labeled (often manually) with pre-defined classes. This classifier is then used to classify new documents into those classes. This classic model is called supervised classification because the training documents all have pre-labeled classes. Although this model is important, in practice it may not be applicable in some common situations. For example, given a set of documents of a particular class P (positive class) for training, one may want to classify a set M of mixed (unlabeled) documents that contains documents from class P along with other types of documents (negative documents) into documents from P and documents not from P. Since there are no labeled negative documents, the traditional classification techniques are inapplicable. This problem is called partially supervised classification (PSC). Existing methods for solving PSC are based on heuristics, which are prone to error. Researchers typically use the evaluation method for supervised classification to evaluate PSC techniques, which is inadequate because supervised classification approaches assume the availability of labeled negative documents. The objectives of this project are to design a robust and principled technique to solve PSC, implement a system for PSC, devise a method to evaluate such techniques, and identify methods for determining the minimum number of labeled documents needed to achieve the optimal accuracy in order to reduce manual labeling efforts. The algorithms and systems developed by this project will be used in a graduate course and put on the Web for other researchers to use in their research and teaching. The results of this research should be widely useful because the identification of targeted information/documents is of great value in this information age.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
III: Small: A Holistic Approach to Sentiment Analysis
  • 批准号:
    1910424
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.99万
  • 财政年份:
    2019
  • 负责人:
    Bing Liu
  • 依托单位:
III: Medium: Collaborative Research: Collective Opinion Fraud Detection: Identifying and Integrating Cues from Language, Behavior, and Networks
  • 批准号:
    1407927
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2014
  • 负责人:
    Bing Liu
  • 依托单位:
Collaborative Research: Using Multi-Modal Digital Footprints to Infer Public Sentiment
  • 批准号:
    1111092
  • 项目类别:
    Standard Grant
  • 资助金额:
    $44.98万
  • 财政年份:
    2011
  • 负责人:
    Bing Liu
  • 依托单位:
海外基金