课题基金 / 基金详情

Development of a general classification framework under the Neyman-Pearson Paradigm, with biomedical and social applications

Development of a general classification framework under the Neyman-Pearson Paradigm, with biomedical and social applications
在内曼-皮尔逊范式下开发通用分类框架,并具有生物医学和社会应用
批准号:
1613338
负责人:
Xin Tong
金额:
$12.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-08-15 至 2019-10-31

项目摘要

项目成果

Xin Tong的其他基金

相似基金

相关文献

中文摘要
翻译
分类在各个领域有着广泛的应用,包括生物科学、医学、工程、金融和社会科学。分类的目的是基于标记的训练数据准确地预测新观测的类别标签。例如,电子邮件服务提供商需要确定传入的电子邮件是否是垃圾邮件。 在不同类型的分类问题中,二元分类是理论、方法和算法发展中最基本和最重要的类型。二进制分类中的一个重要问题是如何控制优先级类型的错误,无论是I型错误(将0类数据点错误分类为1类的机会)还是II型错误(将1类数据点错误分类为0类的机会)。Neyman-Pearson(NP)分类范式是一个旨在为控制第一类(或第二类)错误提供理论保证的理论框架。然而,如何实现NP范式与实际的分类算法仍然是一个巨大的挑战。在这项研究中,PI将通过开发新的统计理论,方法,算法和NP范式下的新评估指标来应对这一挑战。该提案的结果将具有广泛的潜在应用,例如降低疾病诊断中的假阳性率,提高社交媒体数据对社交事件的预测准确性。PI将在拟议的项目中监督不同背景的研究生和本科生,项目成果将在研究生级别的研讨会课程中教授。为协助统计及跨学科研究,研究员将把本项目所开发的方法以开放源码软件包的形式分发,并开发新的统计理论、方法、算法及应用,以控制奈曼-皮尔逊(NP)范式下的不对称分类错误。NP范式解决的情况下,用户坚持对第一类错误的具体限制,同时保持第二类错误最小化。NP范式在假设检验领域已有百年历史,但直到最近才在分类领域受到重视,其理论和方法也不完善。有了以下四个目标,PI将开发一个通用的NP分类框架,并显示它如何可以应用于生物医学和社会科学。根据目标I,PI将通过探索不同数据结构和样本大小的特征依赖性和相互作用来开发新的NP分类理论和方法。在目标II下,PI将设计一个伞形算法,以使流行的分类方法适应NP范式。在目的III下,主要研究者将根据NP分类理论和方法,构建一个NP版本的受试者工作特征(ROC)曲线:“NP-ROC”,这是一个新的评估指标。根据目标IV,PI将把目标I-III中开发的新型NP分类方法应用于大规模生物医学和社会应用。
英文摘要
Classification has broad applications in various fields, including biological sciences, medicine, engineering, finance, and social sciences. The aim of classification is to accurately predict class labels for new observations based on labeled training data. For example, an email service provider needs to decide whether an incoming email is spam. Among different types of classification problems, binary classification is the most basic and important type for theoretical, methodological and algorithmic development. An important question in binary classification is how to control a prioritized type of error, either the type I error (the chance of misclassifying a class 0 data point as class 1) or the type II error (the chance of misclassifying a class 1 data point as class 0). The Neyman-Pearson (NP) classification paradigm is a theoretic framework aiming to control the type I (or type II) error with theoretic guarantee. Yet how to implement the NP paradigm with practical classification algorithms remains a great challenge. In this research, the PIs will tackle this challenge by developing new statistical theory, methods, algorithms, and a novel evaluation metric under the NP paradigm. Results from this proposal will have broad potential applications, such as reducing false positive rates in disease diagnosis and improving prediction accuracy of social events from social media data. The PIs will supervise graduate and undergraduate students of diverse background in the proposed project, and the project outcomes will be taught in graduate-level seminar courses. To aid statistical and interdisciplinary research, the PIs will distribute methods developed in this project as open-source software packages.The PIs will develop new statistical theory, methods, algorithms and applications to control asymmetric classification errors under the Neyman-Pearson (NP) paradigm. The NP paradigm addresses cases where users insist on a specific bound on type I error while keeping type II error to a minimum. Although the NP paradigm has a century-long history in hypothesis testing, until recently it did not receive much attention in the classification area, and its theory and methodologies are as yet incomplete. With the following four aims, the PIs will develop a general NP classification framework and show how it can be applied in the biomedical and social sciences. Under Aim I, the PIs will develop new NP classification theory and methods by exploring feature dependency and interactions for different data structures and sample sizes. Under Aim II, the PIs will design an umbrella algorithm to adapt popular classification methods to the NP paradigm. Under Aim III, the PIs will construct an NP version of Receiver Operating Characteristic (ROC) curves: "NP-ROC", a new evaluation metric based on the NP classification theory and methodologies. Under Aim IV, the PIs will apply the novel NP classification methodologies developed in Aims I-III to large-scale biomedical and social applications.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Development of Classification Theory and Methods for Objective Asymmetry, Sample Size Limitation, Labeling Ambiguity, and Feature Importance
  • 批准号:
    2113500
  • 项目类别:
    Standard Grant
  • 资助金额:
    $12.0万
  • 财政年份:
    2021
  • 负责人:
    Xin Tong
  • 依托单位:
Collaborative Research: Transfer Learning for Large-Scale Inference: General Framework and Data-Driven Algorithms
  • 批准号:
    2015339
  • 项目类别:
    Standard Grant
  • 资助金额:
    $12.0万
  • 财政年份:
    2020
  • 负责人:
    Xin Tong
  • 依托单位:
Robust and Interpretable Bayesian Quantile Longitudinal Analysis in Social and Behavioral Sciences
  • 批准号:
    1951038
  • 项目类别:
    Standard Grant
  • 资助金额:
    $25.0万
  • 财政年份:
    2020
  • 负责人:
    Xin Tong
  • 依托单位:
国内基金
海外基金
Toward a general theory of intermittent aeolian and fluvial nonsuspended sediment transport
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    55万元
  • 批准年份:
    2022
  • 负责人:
    Thomas Pahtz
  • 依托单位:
一类新的连分数动力系统的研究
  • 批准号:
    11361025
  • 项目类别:
    地区科学基金项目
  • 资助金额:
    33.0万元
  • 批准年份:
    2013
  • 负责人:
    钟婷
  • 依托单位:
全身麻醉药作用于生殖系统GABAA受体对男性生殖功能的影响及机制研究
  • 批准号:
    30901390
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    20.0万元
  • 批准年份:
    2009
  • 负责人:
    王金韬
  • 依托单位:
图的一般染色数与博弈染色数
  • 批准号:
    10771035
  • 项目类别:
    面上项目
  • 资助金额:
    18.0万元
  • 批准年份:
    2007
  • 负责人:
    杨大庆
  • 依托单位: