课题基金 / 基金详情

On Partially Supervised Text Classification

On Partially Supervised Text Classification
关于部分监督文本分类
批准号:
0307239
负责人:
Bing Liu
金额:
$23.02万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2003
资助国家:
美国
项目状态:
已结题
起止时间:
2003-05-15 至 2006-04-30

项目摘要

项目成果

Bing Liu的其他基金

相似基金

相关文献

中文摘要
翻译
文本分类是将文本文档自动分配到预定义的类别/类别。传统上,分类器是使用用预定义类标记(通常是手动)的训练文档来构建的。然后使用该分类器将新文档分类到这些类中。这一经典模型被称为监督分类,因为训练文档都具有预先标记的类别。虽然这个模型很重要,但在实践中,它可能不适用于一些常见的情况。例如,给定用于训练的特定类P(正类)的一组文档,人们可能想要将包含来自类P的文档以及其他类型的文档(负文档)的混合(未标记)文档的集合M分类为来自P的文档和不来自P的文档。由于没有标记的负文档,传统的分类技术是不适用的。这个问题被称为部分监督分类(PSC)。现有的PSC求解方法都是基于启发式的,容易出错。研究人员通常使用监督分类的评估方法来评估PSC技术,这是不够的,因为监督分类方法假设有标记的否定文档的可用性。本项目的目标是设计一种稳健和原则性的技术来解决PSC,实现PSC的系统,设计一种评估这种技术的方法,并确定确定达到最佳准确度所需的最小标记文件数量的方法,以减少人工标记的工作量。该项目开发的算法和系统将用于研究生课程,并放到网上供其他研究人员在研究和教学中使用。这项研究的结果应该是广泛有用的,因为在这个信息时代,确定目标信息/文件具有重要价值。
英文摘要
Text classification is the automated assignment of text documents to pre-defined categories/classes. Traditionally, a classifier is built using training documents labeled (often manually) with pre-defined classes. This classifier is then used to classify new documents into those classes. This classic model is called supervised classification because the training documents all have pre-labeled classes. Although this model is important, in practice it may not be applicable in some common situations. For example, given a set of documents of a particular class P (positive class) for training, one may want to classify a set M of mixed (unlabeled) documents that contains documents from class P along with other types of documents (negative documents) into documents from P and documents not from P. Since there are no labeled negative documents, the traditional classification techniques are inapplicable. This problem is called partially supervised classification (PSC). Existing methods for solving PSC are based on heuristics, which are prone to error. Researchers typically use the evaluation method for supervised classification to evaluate PSC techniques, which is inadequate because supervised classification approaches assume the availability of labeled negative documents. The objectives of this project are to design a robust and principled technique to solve PSC, implement a system for PSC, devise a method to evaluate such techniques, and identify methods for determining the minimum number of labeled documents needed to achieve the optimal accuracy in order to reduce manual labeling efforts. The algorithms and systems developed by this project will be used in a graduate course and put on the Web for other researchers to use in their research and teaching. The results of this research should be widely useful because the identification of targeted information/documents is of great value in this information age.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
III: Small: A Holistic Approach to Sentiment Analysis
  • 批准号:
    1910424
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.99万
  • 财政年份:
    2019
  • 负责人:
    Bing Liu
  • 依托单位:
III: Medium: Collaborative Research: Collective Opinion Fraud Detection: Identifying and Integrating Cues from Language, Behavior, and Networks
  • 批准号:
    1407927
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2014
  • 负责人:
    Bing Liu
  • 依托单位:
Collaborative Research: Using Multi-Modal Digital Footprints to Infer Public Sentiment
  • 批准号:
    1111092
  • 项目类别:
    Standard Grant
  • 资助金额:
    $44.98万
  • 财政年份:
    2011
  • 负责人:
    Bing Liu
  • 依托单位:
海外基金