RUI: A Family of Versatile Mixture Models for Analyzing Mixed-Type Data with Asymmetry, Outliers, and Missing Values
RUI: A Family of Versatile Mixture Models for Analyzing Mixed-Type Data with Asymmetry, Outliers, and Missing Values
批准号:
2209974
负责人:
Cristina Tortora
金额:
$15.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-07-15 至 2025-06-30
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Cluster analysis aims to discover patterns and homogeneity in the data. It reveals subgroups in a population of study, and it has applications in many fields. For example, in psychology, the clusters can be groups of patients that can benefit from a specific set of treatments. To apply cluster analysis to a data set, the data need to have some characteristics; for example, some techniques require the data to be continuous, and oftentimes they need pre-treatments. This project will develop a series of new clustering techniques that are suitable for challenging data sets without pre-treatments, such as those with high dimension, missing values, non-continuous variables, or with outliers. Novel statistical approaches and software packages will be produced and made available to general users. Undergraduate students will be directly involved in the research project, and together with graduate students, they will be trained to conduct research in data analysis. Many more students will be involved through class projects and the research outcomes will enrich the content of some of the offered courses. A widely used approach for cluster analysis is model-based clustering. It assumes that a population is a mixture of subpopulations, each of which can be represented by a density function. A variety of clustering methods and algorithms exist; however, they still have a series of limitations. Outliers and missing data can impact the clustering results, the high number of parameters makes the techniques not usable on high-dimensional data sets. Moreover, many algorithms assume continuous data; and they are not readily adaptable to handle discrete, binary, categorical, or a mixture of continuous and categorical data types. This is a major limitation because, in many fields such as medicine, biology, marketing, and many others, the data have all those characteristics. In this project, new clustering techniques based on non-Gaussian model-based clustering will be developed that will circumvent existing limitations on cluster shape, outliers, missing data, dimension, and data type of current methods. The novel methods will improve the flexibility in detecting skewed clusters and in obtaining robustness when dealing with outliers and missing data. Implicit and explicit dimension reduction techniques will be used for dimension reduction and latent class models will be adopted to deal with mixed-type data. The project will include a study on the indices to select the number of clusters and a thorough comparison with existing methods on real and simulated data will be undertaken, giving the users a guideline on which model to use based on the goal and challenges in their data.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
水稻 OVATE Family Protein 8 (OsOFP8)基因的功能研究
-
批准号:31671271
-
项目类别:面上项目
-
资助金额:62.0万元
-
批准年份:2016
-
负责人:李建雄
-
依托单位:
del Pezzo曲面的family上的E_n向量丛
-
批准号:11501201
-
项目类别:青年科学基金项目
-
资助金额:18.0万元
-
批准年份:2015
-
负责人:陈云霞
-
依托单位:
Pim family调控白血病细胞和造血微环境之间Cross Talk在急性髓系白血病中的作用
-
批准号:81100330
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2011
-
负责人:吴俣
-
依托单位: