课题基金 / 基金详情

Discrete and/or Longitudinal Data (small/big) analysis and The Behrens-Fisher problem

Discrete and/or Longitudinal Data (small/big) analysis and The Behrens-Fisher problem
离散和/或纵向数据(小/大)分析和 Behrens-Fisher 问题
批准号:
RGPIN-2018-04558
负责人:
Paul, Sudhir
金额:
$1.31万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2020
资助国家:
加拿大
项目状态:
已结题
起止时间:
2020-01-01 至 2021-12-31

项目摘要

项目成果

Paul, Sudhir的其他基金

相似基金

相关文献

中文摘要
翻译
在许多研究领域,如流行病学、生物统计学、医学和公共卫生科学、环境研究和社会科学,经常出现计数或比例形式的离散数据。这些数据经常遇到过度离散(方差大于简单模型,如二项式或泊松模型)和零膨胀(比简单模型预测的零计数更多)。 由于响应变量和/或解释变量(协变量)中存在缺失值,离散数据的回归分析可能会进一步复杂化。如果缺失不依赖于观测数据,则缺失数据称为完全随机缺失(MCAR)。如果缺失数据机制仅依赖于观察数据,则数据随机缺失(MAR)。MAR也被称为可忽略的失踪。也就是说,在这种情况下,可以忽略缺失数据机制。如果缺失数据机制取决于观察到的数据和未观察到的数据,即观察值的失败取决于本应观察到的值,则数据被称为缺失而不是随机(Mnar),在这种情况下,缺失是不可忽略的。 纵向数据(计数/二进制/连续/存活)经常出现在许多主题领域,如流行病学、生物统计学、医学和公共卫生科学、环境研究和社会科学。纵向研究的特点是在一段时间内重复观察相同的变量。通常假设受试者是独立的,而收集到的同一受试者的观察结果是相关的。 此外,在具有许多解释变量的大(大)数据集中选择模型(选择贡献最大的回归变量)过程很重要,因为在实践中解释简单模型的结果要容易得多。 在这项研究中,我们将发展离散数据回归模型中的估计程序(包括过度分散、零通货膨胀、缺失响应、协变量中的测量误差)、模型选择和参数估计的小样本偏差修正。 在许多应用领域,有时有必要比较一种方法与另一种方法(两种药物、两种教学方法、两种肥料等)的有效性。例如,在两种不同的生物条件下,我们通常对识别差异表达的基因感兴趣。通常的情况是,对于许多需要对大量基因进行筛选或排序的基因,违反了两组方差相等的假设。在这些情况下,无法进行准确的测试。在这项研究中,我计划开发近似程序,并在关于数据分布的不同假设(正态、负二项、β-二项、威布尔、伽玛等)下将它们与现有程序进行比较。
英文摘要
Discrete data in the form of counts or proportions often arise in many fields of study, such as, epidemiology, biostatistics, medical and public health sciences, environmental studies and social sciences. These data often encounter over-dispersion (variance is larger than what can be predicted by a simple model, such as, the binomial or the Poisson model) and zero-inflation (more zero counts than what can be predicted by a simple model). Regression analysis of discrete data can be further complicated by the existence of missing values in the response variable and/or in the explanatory variables (covariates). If the missingness does not depend on observed data, then the missing data are called missing completely at random (MCAR). If the missing data mechanism depends only on observed data, then the data are missing at random (MAR). The MAR is also known as ignorable missing. That is, in this case, the missing data mechanism can be ignored. If the missing data mechanism depends on both observed and unobserved data, that is, failure to observe a value depends on the value that would have been observed, then the data are called missing not at random (MNAR) in which case the missingness is nonignorable. Longitudinal data (count/binary/continuous/survival) are frequently encountered in many subject-matter areas such as epidemiology, biostatistics, medical and public health sciences, environmental studies and social sciences. Longitudinal studies are characterized by observing the same variables repeatedly over a period of time. Usually the subjects are assumed to be independent, while the collected observations of the same subject are correlated. Further, model selection (selecting regression variables that contribution most) procedures in large (big) data sets with many explanatory variables is important, as in practice interpreting results from a simple model is much easier. In this research I we will develop estimation procedures in discrete data regression models (involving over-dispersion, zero-inflation, missing responses, measurement errors in covariates), model selection, and small sample bias correction of parameter estimates in longitudinal set up or otherwise. In many applied fields sometimes it is necessary to compare effectiveness of one procedure over another (two drugs, two teaching methods, two fertilizers etc.). For example, under two biologically different conditions we are often interested in identifying differentially expressed genes. It is often the case that the assumption of equal variances of the two groups is violated for many genes where a large number of them are required to be filtered or ranked. In these cases exact tests are unavailable. In this research I plan to develop approximate procedures and compare them with existing procedures under different assumptions regarding the data distribution (normal, negative binomial, beta-binomial, Weibull, Gamma etc.).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Discrete and/or Longitudinal Data (small/big) analysis and The Behrens-Fisher problem
  • 批准号:
    RGPIN-2018-04558
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.62万
  • 财政年份:
    2022
  • 负责人:
    Paul, Sudhir
  • 依托单位:
Discrete and/or Longitudinal Data (small/big) analysis and The Behrens-Fisher problem
  • 批准号:
    RGPIN-2018-04558
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.31万
  • 财政年份:
    2021
  • 负责人:
    Paul, Sudhir
  • 依托单位:
Discrete and/or Longitudinal Data (small/big) analysis and The Behrens-Fisher problem
  • 批准号:
    RGPIN-2018-04558
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.31万
  • 财政年份:
    2019
  • 负责人:
    Paul, Sudhir
  • 依托单位:
Discrete and/or Longitudinal Data (small/big) analysis and The Behrens-Fisher problem
  • 批准号:
    RGPIN-2018-04558
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.31万
  • 财政年份:
    2018
  • 负责人:
    Paul, Sudhir
  • 依托单位:
海外基金