课题基金 / 基金详情

RUI: A Family of Versatile Mixture Models for Analyzing Mixed-Type Data with Asymmetry, Outliers, and Missing Values

RUI: A Family of Versatile Mixture Models for Analyzing Mixed-Type Data with Asymmetry, Outliers, and Missing Values
RUI:一系列多功能混合模型,用于分析具有不对称性、离群值和缺失值的混合类型数据
批准号:
2209974
负责人:
Cristina Tortora
金额:
$15.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-07-15 至 2025-06-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
聚类分析的目的是发现数据中的模式和同质性。它揭示了研究群体中的子群,并在许多领域有应用。例如,在心理学中,集群可以是可以从一套特定的治疗中受益的一组患者。要将聚类分析应用于数据集,数据需要具有一些特征;例如,一些技术要求数据是连续的,而且往往需要进行预处理。该项目将开发一系列新的聚类技术,这些技术适用于在没有预处理的情况下挑战数据集,例如高维、缺失值、非连续变量或具有离群值的数据集。将编制新的统计方法和软件包,供一般用户使用。本科生将直接参与研究项目,并将与研究生一起培训他们进行数据分析方面的研究。更多的学生将通过课堂项目参与进来,研究成果将丰富一些所提供课程的内容。一种广泛使用的聚类分析方法是基于模型的聚类。它假设种群是一个子种群的混合,每个子种群可以用一个密度函数来表示。虽然存在多种聚类方法和算法,但它们仍有一系列的局限性。离群点和缺失数据会影响聚类结果,较多的参数使得该技术不适用于高维数据集。此外,许多算法假定数据是连续的;并且它们不容易适应于处理离散、二进制、分类或连续和分类数据类型的混合。这是一个主要的限制,因为在许多领域,如医学、生物、营销和许多其他领域,数据都具有所有这些特征。在本项目中,将开发基于非高斯模型聚类的新的聚类技术,以绕过现有方法在聚类形状、离群点、缺失数据、维度和数据类型方面的现有限制。这些新方法将提高检测倾斜聚类的灵活性,并在处理孤立点和丢失数据时获得稳健性。降维采用隐式和显式降维技术,混合类型数据采用隐类模型进行降维。该项目将包括一项选择集群数量的指数研究,并将与基于真实和模拟数据的现有方法进行彻底比较,根据数据中的目标和挑战为用户提供使用哪种模型的指导方针。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Cluster analysis aims to discover patterns and homogeneity in the data. It reveals subgroups in a population of study, and it has applications in many fields. For example, in psychology, the clusters can be groups of patients that can benefit from a specific set of treatments. To apply cluster analysis to a data set, the data need to have some characteristics; for example, some techniques require the data to be continuous, and oftentimes they need pre-treatments. This project will develop a series of new clustering techniques that are suitable for challenging data sets without pre-treatments, such as those with high dimension, missing values, non-continuous variables, or with outliers. Novel statistical approaches and software packages will be produced and made available to general users. Undergraduate students will be directly involved in the research project, and together with graduate students, they will be trained to conduct research in data analysis. Many more students will be involved through class projects and the research outcomes will enrich the content of some of the offered courses. A widely used approach for cluster analysis is model-based clustering. It assumes that a population is a mixture of subpopulations, each of which can be represented by a density function. A variety of clustering methods and algorithms exist; however, they still have a series of limitations. Outliers and missing data can impact the clustering results, the high number of parameters makes the techniques not usable on high-dimensional data sets. Moreover, many algorithms assume continuous data; and they are not readily adaptable to handle discrete, binary, categorical, or a mixture of continuous and categorical data types. This is a major limitation because, in many fields such as medicine, biology, marketing, and many others, the data have all those characteristics. In this project, new clustering techniques based on non-Gaussian model-based clustering will be developed that will circumvent existing limitations on cluster shape, outliers, missing data, dimension, and data type of current methods. The novel methods will improve the flexibility in detecting skewed clusters and in obtaining robustness when dealing with outliers and missing data. Implicit and explicit dimension reduction techniques will be used for dimension reduction and latent class models will be adopted to deal with mixed-type data. The project will include a study on the indices to select the number of clusters and a thorough comparison with existing methods on real and simulated data will be undertaken, giving the users a guideline on which model to use based on the goal and challenges in their data.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
水稻 OVATE Family Protein 8 (OsOFP8)基因的功能研究
  • 批准号:
    31671271
  • 项目类别:
    面上项目
  • 资助金额:
    62.0万元
  • 批准年份:
    2016
  • 负责人:
    李建雄
  • 依托单位:
del Pezzo曲面的family上的E_n向量丛
  • 批准号:
    11501201
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    18.0万元
  • 批准年份:
    2015
  • 负责人:
    陈云霞
  • 依托单位:
Pim family调控白血病细胞和造血微环境之间Cross Talk在急性髓系白血病中的作用
  • 批准号:
    81100330
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2011
  • 负责人:
    吴俣
  • 依托单位: