Collaborative Research: Development of Classification Theory and Methods for Objective Asymmetry, Sample Size Limitation, Labeling Ambiguity, and Feature Importance
Collaborative Research: Development of Classification Theory and Methods for Objective Asymmetry, Sample Size Limitation, Labeling Ambiguity, and Feature Importance
批准号:
2113500
负责人:
Xin Tong
金额:
$12.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-07-01 至 2024-06-30
中文摘要
从生物医学科学到信息技术,分类是一种流行的数据分析技术。该项目将开发理论支持的统计方法和算法,以解决分类应用中的紧迫挑战。这些挑战与训练数据的不完善有关,这在疾病诊断和网络安全等高风险应用中很普遍。特别是,这个项目将专注于所谓的不对称分类问题,其中一个特定的类比其他类更重要,方法和算法将旨在控制在总体中错过最重要的类的分类误差,而不仅仅是在特定的数据集中。这一特性将使方法和算法在医学诊断方面变得强大,因为医学诊断的主要目标是在人群中诊断的准确性。此外,该项目将提供一系列项目,从理论到应用,适合培养研究生和本科生。该项目的跨学科性质预计将吸引来自不同背景的学生加入pi的努力。pi将开发一套应用驱动、理论支持的方法和算法,以解决紧迫的数据挑战,包括样本量限制、抽样偏差和模糊类标签。发展将主要在Neyman-Pearson (NP)分类范式下进行,该范式旨在将总体水平的假阴性率(p-FNR)控制在期望的水平下,同时最小化总体水平的假阳性率(p-FPR)。该项目将把NP分类集成到尖端的统计学习任务中,并使其能够解决上述现实世界的数据挑战。具体而言,该项目将包括以下四个总体目标。首先,pi将使用随机矩阵理论来解决NP分类方法中一个长期存在的问题:是否可以在没有样本分裂步骤的情况下构建NP分类器以提高数据效率。其次,由于NP范式对抽样偏差具有不变性,因此pi将开发NP分类器来解决生物医学应用中的抽样偏差问题。这些分类器可以在有偏差的样本上进行训练,但仍然可以实现p-FNR控制。第三,pi将开发一个无模型特征排序框架,以结合包括NP范式在内的多种分类范式,并反映预测目标。第四,pi将开发第一个标签噪声设置下的NP保护伞算法和第一个在多类分类中结合模糊类的信息论标准。为了传播项目成果,项目负责人将举办研究讲座,组织会议,分享开源软件包和教程,并与分类方法的实践者接触。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Classification is a popular data analytical technique in disciplines ranging from biomedical sciences to information technologies. This project will develop theory-backed statistical methods and algorithms to address pressing challenges in the application of classification. These challenges are related to imperfect aspects of training data, which are widespread in high-stake applications such as disease diagnosis and cybersecurity. In particular, this project will focus on the so-called asymmetric classification problems where a particular class is of greater importance than other classes, and the methods and algorithms will aim to control the classification error of missing the most important class in the population, not just in a particular dataset. This property will make the methods and algorithms powerful for medical diagnosis, for which the primary goal is diagnosis accuracy in the population. Moreover, this project will provide a suite of projects, ranging from theory to applications, that are suitable for training graduate and undergraduate students. The interdisciplinary nature of this project is expected to attract students from diverse background to join the PIs’ efforts.The PIs will develop a suite of application-driven, theory-backed methods and algorithms to address pressing data challenges including sample size limitations, sampling biases, and ambiguous class labels. The development will be primarily under the Neyman-Pearson (NP) classification paradigm, which was designed to control the population-level false-negative rate (p-FNR) under a desired level while minimizing the population-level false-positive rate (p-FPR). This project will integrate the NP classification into cutting-edge statistical learning tasks and enable it to address the aforementioned real-world data challenges. Specifically, this project will include the following four overarching goals. First, the PIs will use random matrix theory to address a long-standing problem in the NP classification methodology: whether NP classifiers can be constructed without a sample-splitting step to improve data efficiency. Second, because the NP paradigm has an invariance property to sampling bias, the PIs will develop NP classifiers to address the sampling bias issue in biomedical applications. These classifiers can be trained on biased samples but still achieve the p-FNR control. Third, the PIs will develop a model-free feature ranking framework to incorporate multiple classification paradigms including the NP paradigm and to reflect prediction objectives. Fourth, the PIs will develop the first NP umbrella algorithm under the label noise setting and the first information-theoretic criteria that combine ambiguous classes in multi-class classification. To disseminate the project outcomes, the PIs will give research talks, organize conference sessions, share open-source software packages with tutorials, and reach out to practitioners of classification methods.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI:
--
发表时间:
2021-09
期刊:
J. Mach. Learn. Res.
影响因子:
--
作者:
[Chihao Zhang;Y. Chen;Shihua Zhang;Jingyi Jessica Li]
通讯作者:
Chihao Zhang;Y. Chen;Shihua Zhang;Jingyi Jessica Li
Collaborative Research: Transfer Learning for Large-Scale Inference: General Framework and Data-Driven Algorithms
-
批准号:2015339
-
项目类别:Standard Grant
-
资助金额:$12.0万
-
财政年份:2020
-
负责人:Xin Tong
-
依托单位:
Robust and Interpretable Bayesian Quantile Longitudinal Analysis in Social and Behavioral Sciences
-
批准号:1951038
-
项目类别:Standard Grant
-
资助金额:$25.0万
-
财政年份:2020
-
负责人:Xin Tong
-
依托单位:
Development of a general classification framework under the Neyman-Pearson Paradigm, with biomedical and social applications
-
批准号:1613338
-
项目类别:Standard Grant
-
资助金额:$12.0万
-
财政年份:2016
-
负责人:Xin Tong
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: