课题基金 / 基金详情

Feature selection in several challenging directions

Feature selection in several challenging directions
几个具有挑战性的方向的特征选择
批准号:
2310668
负责人:
Jiashun Jin
金额:
$22.5万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-09-01 至 2026-08-31

项目摘要

项目成果

Jiashun Jin的其他基金

相似基金

相关文献

中文摘要
翻译
特征选择在许多统计问题中起着至关重要的作用,例如癌症分类以及文本和网络数据的分析。这个项目将在几个具有挑战性的、未被充分研究的方向上研究特征选择。首先,MDAStat是1971年至2015年间统计学家出版物的最新大规模数据集,为网络分析和文本分析的研究提供了丰富的资源。该项目将通过收集新数据将数据集的范围扩大到1971-2025年。其次,该项目将开发一系列相关度量,它们提供了一种更好的方法来衡量响应和预测变量之间的非线性关系。在癌症和生物医学研究中的许多应用问题中,该度量提供了更准确的特征选择结果。最后,该项目将开发从社交网络和文本文档中提取特征的新方法,并在医疗数据分析和作者归属(即识别可能古老的文本文档的正确作者)等应用程序中产生更好的结果。这项研究将产生新的思想和方法来解决现代统计研究中的许多具有挑战性的问题,并将大大增加对许多科学和工程问题的理解,如癌症和生物医学研究、网络分析、文本分析和自然语言处理。特征选择是高维数据分析的重要方法。该项目将在几个具有挑战性的方向上研究特征选择,并将在以下主题上做出贡献。首先,尽管有许多关于稀有/强信号机制的研究,但在更具挑战性的稀有/弱信号机制中,套索的性质在很大程度上仍然未知。该项目将开发新的技术,并使用它们来推导套索的汉明选择误差的敏感率,特别是对于稀有/弱信号区域。第二,特征选择中的一个具有挑战性的问题是如何度量响应和预测变量之间的非线性关系。该项目将开发一系列非线性相关度量,并使用它们来推导用于癌症分类和癌症聚集的非线性稀有/弱模型中的尖锐相变。第三,尽管社交网络的模型很多,但目前还不清楚哪种模型最适合真实的网络,部分原因是网络的适合度是一个具有挑战性的问题。该项目将开发一种新的契合度方法,并使用它来确定最适合社交网络的模型。最后,特征提取和嵌入到文本文档和网络中是一个具有挑战性的问题。该项目将开发新的特征提取和嵌入方法,并将其用于预测已发表论文的未来引文数量和作者归属。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Feature selection plays a crucial role in many statistical problems, such as cancer classification and analysis of text and network data. This project will study feature selection in several challenging, understudied directions. First, MDAStat is a recent large-scale data set on the publications of statisticians between 1971 and 2015, which provides a rich resource for research on network analysis and text analysis. The project will expand the scope of data set to 1971-2025 by collecting new data. Second, the project will develop a family of correlation metrics, which provide a better way to measure the nonlinear relationship between the response and predictive variables. The metrics provide more accurate feature selection results in many application problems in cancer and biomedical study. Last, the project will develop new approaches to extracting features from social networks and text documents and generate better results in applications such as analysis of health care data and author attribution (i.e., identifying the right authors of a possibly ancient text document). The research will generate new ideas and methods to address many challenging problems in modern statistical research, and will substantially increase the understanding of many problems in science and engineering, such as cancer and biomedical research, network analysis, text analysis, and natural language processing. Feature selection is an important approach in high dimensional data analysis. The project will study feature selection in several challenging directions and will make contributions on the following topics. First, despite many studies on the rare/strong signal regime, the property of the lasso remains largely unknown in the more challenging rare/weak signal regime. The project will develop new techniques and use them to derive sharp rates of the Hamming selection errors of the lasso, especially for the rare/weak signal regime. Second, a challenging problem in feature selection is how to measure the nonlinear relationship between the response and predictive variables. The project will develop a family of nonlinear correlation metrics and use them to derive sharp phase transitions in nonlinear rare/weak models for cancer classification and cancer clustering. Third, despite that there are more than a handful of models for social networks, it remains unclear which model fits the best with real networks, partially because network goodness-of-fit is a challenging problem. The project will develop a novel goodness-of-fit approach and use it to identify the most appropriate models for social networks. Last, feature extraction and embedding with text documents and networks is a challenging problem. The project will develop novel approaches for feature extraction and embedding and using them for predicting future citation counts of a published paper and for author attribution.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
New Tools for Analyzing Complex Network and Text Data
  • 批准号:
    2015469
  • 项目类别:
    Standard Grant
  • 资助金额:
    $25.0万
  • 财政年份:
    2020
  • 负责人:
    Jiashun Jin
  • 依托单位:
New Tools for Large-Scale Sparse Inference
  • 批准号:
    1513414
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2015
  • 负责人:
    Jiashun Jin
  • 依托单位:
Rare and Weak Signals in Big Data: How to Find Them and How to Use Them
  • 批准号:
    1208315
  • 项目类别:
    Standard Grant
  • 资助金额:
    $10.0万
  • 财政年份:
    2012
  • 负责人:
    Jiashun Jin
  • 依托单位:
CAREER: Inferences on Large-Scale Multiple Comparisons: The Temptation of the Fourier Kingdom
  • 批准号:
    0908613
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $34.22万
  • 财政年份:
    2008
  • 负责人:
    Jiashun Jin
  • 依托单位:
国内基金
海外基金
Intelligent Patent Analysis for Optimized Technology Stack Selection:Blockchain BusinessRegistry Case Demonstration
  • 批准号:
    --
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    USHARANI HAREESH GOVINDARA JAN
  • 依托单位:
高维复杂数据分析中具有可重复性的统计学习方法研究及其应用
  • 批准号:
    72071187
  • 项目类别:
    面上项目
  • 资助金额:
    48.0万元
  • 批准年份:
    2020
  • 负责人:
    郑泽敏
  • 依托单位:
基于microRNA前体性质的microRNA演化研究
最优证券设计及完善中国资本市场的路径选择
  • 批准号:
    70873012
  • 项目类别:
    面上项目
  • 资助金额:
    27.0万元
  • 批准年份:
    2008
  • 负责人:
    彭龙
  • 依托单位: