Longitudinal data subject to irregular observation: developing methods for variable selection, causal inference, and measurement error
Longitudinal data subject to irregular observation: developing methods for variable selection, causal inference, and measurement error
批准号:
RGPIN-2021-02733
负责人:
Pullenayegum, Eleanor
金额:
$1.75万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2021
资助国家:
加拿大
项目状态:
已结题
起止时间:
2021-01-01 至 2022-12-31
中文摘要
我们生活在一个数据丰富的社会中,许多数据集都包含同一个人的信息(称为纵向数据)。然而,收集数据的时间往往与感兴趣的结果有关。例如,在一项关于新生儿生长的研究中,生长缓慢的新生儿可能会更频繁地测量体重。取所有体重测量值的平均值会低估增长率。测量时间是信息性的:测量频率为我们提供了有关测量值的重要信息。我们必须考虑测量频率,以获得随时间变化的轨迹的准确图像。分析这类数据的方法所能做的是有限的。这项工作将侧重于:(a)处理在每个时间点测量大量事物的数据集;(B)使研究人员能够检测一个变量是否引起另一个变量的变化;(c)正确解释测量误差;(d)根据系统处于高风险状态还是低风险状态来描述系统的行为。(a)许多数据集包含大量信息。例如,关于水质的数据可能包括数千种污染物的水平;我们需要过滤出哪些污染物是重要的。我们将为具有信息测量时间的纵向数据开发这样做的方法。(b)It通常可以直接表明一个量是否会随着另一个量的增加而增加(关联),但更难表明一个量的变化是否会导致另一个量的变化(因果关系)。我们的能力,以确定因果关系的纵向数据与信息测量时间是有限的。我们将制定办法来弥补这一差距。(c)数据的计量通常会有误差。例如,高度从来没有被完美地测量过。这可能会导致偏见,除非它被解释。我们将找到方法,这样做的纵向数据与翔实的测量时间。(d)It通常有助于描述以风险状态为条件的预期结果。例如,对于患有慢性疾病的患者,如果复发和缓解,描述复发期间的健康状况,缓解期间的健康状况以及缓解时间的比例可能会更有帮助。我们将开发这样做的方法,使用纵向数据与翔实的观察时间。在加拿大,我们收集大量数据作为日常社会运作的一部分。例如,许多患者同意将其医疗记录用于研究。这些数据可以解决加拿大特有的问题,例如,确定哪些人群面临健康状况较差的风险。这就需要仔细处理信息性观察,以提供可靠的结果。我们开发的方法将使研究人员能够做到这一点,从而产生高质量的证据,作为社会决策的基础。
英文摘要
We live in a data-rich society, and many datasets include information on the same individuals repeatedly over time (known as longitudinal data). However, often the times at which data is collected are related to the outcome of interest. For example, in a study of newborn growth, newborns who grow slowly are likely to have their weight measured more often. Taking the average of all the weight measurements over time will underestimate the growth rate. The measurement times are informative: the frequency of measurements gives us important information about the value of the measurements. We must take the measurement frequency into account to get an accurate picture of the trajectory over time. Methods for analysing this type of data are limited in what they can do. This work will focus on: (a) handling datasets for which a large number things are measured at each time point; (b) enabling researchers to detect whether one variable causes a change in another; (c) correctly accounting for measurement error; (d) describing the behaviour of a system according to whether it is in a high or low risk state. (a)Many datasets contain a large amount of information. For example, data on water quality might include the levels of thousands of contaminants; we need to filter out which contaminants are important. We will develop ways of doing this for longitudinal data with informative measurement times. (b)It is often straightforward to show whether one quantity tends to increase as another increases (association), but harder to show whether a change in one quantity causes a change in another (causality). Our ability to determine causality with longitudinal data with informative measurement times is limited. We will develop approaches to address this gap. (c)Data is typically measured subject to error. For example, height is never measured perfectly. This can lead to bias unless it is accounted for. We will find ways of doing this with longitudinal data with informative measurement times. (d)It is often helpful to describe expected outcomes conditional on risk status. For example, for a patient with a chronic disease subject to relapse and remission, it may be more helpful to describe health during periods of relapse, health during periods of remission, and the proportion of time spent in remission. We will develop ways of doing this using longitudinal data with informative observation times. In Canada we collect a large amount of data as part of usual societal operations. For example, many patients provide consent for their medical records to be used for research. This data can address questions that are specific to Canada, for example, to determine which groups of people are at risk of poorer health outcomes. This requires careful handling of informative observation in order to provide reliable results. The methods we develop will equip researchers to do this, and so generate high quality evidence on which to base societal decisions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Longitudinal data subject to irregular observation: developing methods for variable selection, causal inference, and measurement error
-
批准号:RGPIN-2021-02733
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2022
-
负责人:Pullenayegum, Eleanor
-
依托单位:
Statistical Methods for Irregularly Measured Longitudinal Data
-
批准号:RGPIN-2014-03989
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$0.8万
-
财政年份:2019
-
负责人:Pullenayegum, Eleanor
-
依托单位:
Statistical Methods for Irregularly Measured Longitudinal Data
-
批准号:RGPIN-2014-03989
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$0.8万
-
财政年份:2018
-
负责人:Pullenayegum, Eleanor
-
依托单位:
Statistical Methods for Irregularly Measured Longitudinal Data
-
批准号:RGPIN-2014-03989
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$0.8万
-
财政年份:2016
-
负责人:Pullenayegum, Eleanor
-
依托单位:
Statistical Methods for Irregularly Measured Longitudinal Data
-
批准号:RGPIN-2014-03989
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$0.8万
-
财政年份:2015
-
负责人:Pullenayegum, Eleanor
-
依托单位:
Statistical Methods for Irregularly Measured Longitudinal Data
-
批准号:RGPIN-2014-03989
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$0.8万
-
财政年份:2014
-
负责人:Pullenayegum, Eleanor
-
依托单位:
Semi-parametric modelling of longitudinal data when the observation process is neither completely random nor completely deterministic
-
批准号:356042-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$0.87万
-
财政年份:2012
-
负责人:Pullenayegum, Eleanor
-
依托单位:
Semi-parametric modelling of longitudinal data when the observation process is neither completely random nor completely deterministic
-
批准号:356042-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$0.87万
-
财政年份:2011
-
负责人:Pullenayegum, Eleanor
-
依托单位:
Semi-parametric modelling of longitudinal data when the observation process is neither completely random nor completely deterministic
-
批准号:356042-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$0.87万
-
财政年份:2010
-
负责人:Pullenayegum, Eleanor
-
依托单位:
Semi-parametric modelling of longitudinal data when the observation process is neither completely random nor completely deterministic
-
批准号:356042-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$0.87万
-
财政年份:2009
-
负责人:Pullenayegum, Eleanor
-
依托单位:
Semi-parametric modelling of longitudinal data when the observation process is neither completely random nor completely deterministic
-
批准号:356042-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$0.87万
-
财政年份:2008
-
负责人:Pullenayegum, Eleanor
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
-
批准号:--
-
项目类别:外国青年学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:江洋子
-
依托单位:
复杂数据下半参数转换模型及其在老年慢性病发展中的应用研究
-
批准号:72101261
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:孙韬
-
依托单位:
Development of a Linear Stochastic Model for Wind Field Reconstruction from Limited Measurement Data
-
批准号:--
-
项目类别:--
-
资助金额:40万元
-
批准年份:2020
-
负责人:Vikrant Gupta
-
依托单位:
基于高频信息下高维波动率矩阵估计及应用
-
批准号:71901118
-
项目类别:青年科学基金项目
-
资助金额:18.0万元
-
批准年份:2019
-
负责人:穆燕
-
依托单位:
半参数空间自回归面板模型的有效估计与应用研究
-
批准号:71961011
-
项目类别:地区科学基金项目
-
资助金额:16.0万元
-
批准年份:2019
-
负责人:丁飞鹏
-
依托单位:
高频数据波动率统计推断、预测与应用
-
批准号:71971118
-
项目类别:面上项目
-
资助金额:50.0万元
-
批准年份:2019
-
负责人:孔新兵
-
依托单位:
经济管理中复杂数据和复杂行为的分析方法及其应用
-
批准号:71931004
-
项目类别:重点项目
-
资助金额:230.0万元
-
批准年份:2019
-
负责人:周勇
-
依托单位:
基于个体分析的投影式非线性非负张量分解在高维非结构化数据模式分析中的研究
-
批准号:61502059
-
项目类别:青年科学基金项目
-
资助金额:19.0万元
-
批准年份:2015
-
负责人:刘昶
-
依托单位:
基于Linked Open Data的Web服务语义互操作关键技术
-
批准号:61373035
-
项目类别:面上项目
-
资助金额:77.0万元
-
批准年份:2013
-
负责人:冯志勇
-
依托单位: