Development and application of statistical methods for addressing the heterogeneity of data collection intervals common in longitudinal datasets
Development and application of statistical methods for addressing the heterogeneity of data collection intervals common in longitudinal datasets
批准号:
1943044
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --
中文摘要
许多科学都对收集长期数据和分析变化模式很感兴趣。例子包括研究生长曲线、心理变化以及一个地区疾病患病率的变化。通常,研究人员检查这些数据模式(纵向暴露)与后来事件(结果)的关系,这需要使用描述个体纵向暴露模式的数据分析技术(例如个体的生长曲线)。有许多技术可以用来做到这一点。多水平模型(MLMs)首先描述纵向暴露的平均模式,但也提供了关于个体模式与之有多大差异的信息;这些模型的结果很容易理解,但它们不能描述非常复杂的模式。另一方面,潜在增长曲线模型(LGCM)通过创建描述它们的额外数据而不是平均轨迹来估计纵向暴露的个体模式。为此,潜在增长曲线模型以一种不寻常的方式表示时间-通过添加一个“因子加载”,将描述曲线的数据与纵向暴露的每次测量相关联。该因子加载通常设置为进行测量的时间,但可以通过模型进行估计,这使得非常复杂的曲线可以由LGCM比MLM更容易地表示。LGCM还可以扩展到增长混合模型(GMM),该模型根据纵向暴露模式的类型确定数据中的潜在亚组。然而,LGCM要求在完全相同的时间点测量所有个体的数据-称为“间隔时间”。然而,在实践中很少出现这种情况,特别是在使用观察数据时(例如,在医疗记录中记录为测量值的儿童生长曲线);因此,这些最灵活的建模技术不能广泛使用。另一种可用于描述纵向暴露的方法是功能数据分析(FDA)。这通过在较小的时间段中拟合平滑曲线来描述纵向暴露的个体模式。这些部分由“结”限定,结的数量和位置由研究人员选择。这也是一种灵活的方法,但当纵向曝光的测量之间存在较宽的空间时,可能不准确。本项目旨在通过使用本文描述的个体模式对个体测量值进行插值并创建间隔均匀性,从而检查在没有间隔均匀性的纵向暴露上执行FDA的实用性,从而允许使用潜在增长曲线建模来分析纵向暴露的模式,同时将其与以后的结果相关联。这些目标将通过使用真实的和模拟数据来实现,并且还将解决以下问题:a)如何找到测量插值的最佳点?以及B)应如何选择最佳“基函数”(即用于拟合FDA中纵向暴露分段的曲线类型)?还将结果与LGCM(假设间隔均匀性)和MLM的结果进行比较。
英文摘要
There is great interest in many sciences in gathering data over time and analysing patterns of change. Examples include the study of growth curves, psychological changes, and changes in the prevalence of diseases in an area. Often, researchers examine the relationship of these patterns of data (longitudinal exposures) to later events (outcomes), which requires the use of data analysis techniques that describe patterns of the longitudinal exposure in individuals (e.g. growth curves in individual people). There are a number of techniques that can be used to do this. Multilevel models (MLMs) start by describing the average pattern of the longitudinal exposure, but also give information on how much individual patterns differ from it; the results from these models are easy to understand but they cannot describe very complicated patterns. On the other hand, latent growth curve models (LGCMs) estimate individual patterns of a longitudinal exposure by creating extra data that describes them, rather than an average trajectory. To do this, latent growth curve models represent time in an unusual way - by adding a 'factor loading' relating the data describing the curves to each measurement of the longitudinal exposure. This factor loading is often set to the time at which the measurement was taken, but can be estimated by the model, which allows for very complex curves to be represented by LGCMs much more easily than in MLMs. LGCMs can also be extended to growth mixture models (GMMs), which identify underlying subgroups in the data based on the types of patterns of the longitudinal exposure. However, LGCMs require the data in all individuals to be measured at exactly the same time points - called 'interval homogeneity'. However, this is rarely the case in practice, especially when using observational data (e.g. children's growth curves recorded as measurements in their medical records); thus, these most flexible modelling techniques cannot be used widely. Another method that can be used to describe longitudinal exposures is functional data analysis (FDA). This describes individual patterns of the longitudinal exposure by fitting smooth curves in smaller time segments. These segments are bounded by 'knots', the number and position of which are chosen by the researcher. This is also a flexible method, but can be inaccurate when there exist wide spaces between measurements of the longitudinal exposure. This project aims to examine the utility of carrying out FDA on a longitudinal exposure without interval homogeneity by using the individual patterns this describes to interpolate individual measurements and create interval homogeneity, thereby allowing for the use of latent growth curve modelling to analyse the patterns of the longitudinal exposure while relating this to a later outcome. These aims will be addressed using real and simulated data, and the following questions will also be addressed: a) How can the optimum points for interpolation of measurements be found?; and b) How should the optimum 'basis function' be chosen (i.e. the types of curves used to fit segments of the longitudinal exposure in FDA)?. The results will also be compared to those from LGCMs (which assume interval homogeneity) and from MLMs.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
Adjustment for time-invariant and time-varying confounders in 'unexplained residuals' models for longitudinal data within a causal framework and associated challenges.
在因果关系框架和相关挑战中,调整了“无法解释的残差”模型中的时间不变和时变的混杂因素。
DOI:
10.1177/0962280218756158
发表时间:
2019-05
期刊:
Statistical methods in medical research
影响因子:
2.3
作者:
[Arnold KF, Ellison G, Gadd SC, Textor J, Tennant P, Heppenstall A, Gilthorpe MS]
通讯作者:
Gilthorpe MS
国内基金
海外基金
登录
查看更多内容
Graphon mean field games with partial observation and application to failure detection in distributed systems
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:MATHIEULOUROCHLAURIERE
-
依托单位:
均相液相生物芯片检测系统的构建及其在癌症早期诊断上的应用
-
批准号:82372089
-
项目类别:面上项目
-
资助金额:48.00万元
-
批准年份:2023
-
负责人:李万万
-
依托单位:
用于小尺寸管道高分辨成像荧光聚合物点的构建、成像机制及应用研究
-
批准号:82372015
-
项目类别:面上项目
-
资助金额:48.00万元
-
批准年份:2023
-
负责人:熊丽琴
-
依托单位:
网格中以情境为中心的应用自动化研究
-
批准号:60703054
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2007
-
负责人:黄震春
-
依托单位: