Statistical Methods for Dependent Data
Statistical Methods for Dependent Data
批准号:
0805050
负责人:
David Stoffer
金额:
$32.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-07-01 至 2014-06-30
中文摘要
本建议集中讨论与相关数据的统计分析有关的各种主题。第一个项目扩展了分析DNA序列的光谱包络概念。分析长DNA序列数据的一个常见问题是如何识别分散在整个序列中并被非编码区域分开的蛋白质编码序列。DNA序列是异质的,因此有必要扩展方法来捕捉这些序列的局部行为。为了解决局部行为的问题,我们将探讨一个通过光滑样条混合估计的局部谱包络。希望这种方法能够以快速和自动化的方式帮助强调存在于几乎任何长度的分类序列中的任何周期性特征。像人类基因组计划这样的项目已经产生了大量的数据,在这个项目中建立的方法将被证明对各种基因组计划产生的大量数据的分析是有用的。在另一个项目中,重点是对纵向数据的分析和开发一种实用的非参数程序来估计主题内相关结构。该技术用于开发数据驱动的功能主成分分析程序(FPCA)。由于纵向数据通常具有在一个主题内进行的观察是相关的属性,因此需要对这些数据进行有效的分析,以解释这种主题内的相关性。当协方差结构的参数形式未知时,使用错误指定的结构可能导致有偏差和低效的估计。该项目侧重于纵向数据的分析,这些数据可以建模为来自光滑受试者轨迹的观察结果,这些轨迹是在带有噪声的离散时间点观察到的随机过程的实现。纵向数据的高维性和复杂性使得FPCA通过捕获生成数据的随机过程的主要变化模式而成为数据简化和可视化的流行工具。科学家们通常对使用纵向数据来确定一组可能随时间变化的协变量对给定响应的影响感兴趣。函数线性模型,特别是变系数模型,为分析这些数据提供了一个框架。在许多这些数据集中,功能系数具有无法参数化建模的形状。需要对这些数据进行有效的分析,既要考虑主题内的相关性,又要考虑系数的灵活形状。由于主体内协方差的参数形式并不总是已知的,第三个项目侧重于创建一个迭代数据驱动的基于样条的过程,用于拟合变系数模型。这个建议集中于解决相关数据分析中涉及的问题。第一个项目将开发一种检测长DNA序列基因的方法。像人类基因组计划这样的项目已经产生了大量的数据,在这个项目中建立的方法将被证明对各种基因组计划产生的大量数据的分析是有用的。第二个提议的项目侧重于分析长期收集的复杂数据。这个项目也受到DNA分析的推动,特别是基因表达数据的分析。在第三个项目中,研究人员将专注于一种称为功能线性模型的技术。例如,将开发技术来研究生长因子在治疗卵巢癌时是否决定用抗血管生成疗法补充化疗时的作用。
英文摘要
This proposal concentrates on various topics relating to the statistical analysis of dependent data. The first project extends the spectral envelope concept for analyzing DNA sequences. A common problem in analyzing long DNA sequence data is in identifying protein-coding sequences that are dispersed throughout the sequence and separated by regions of noncoding. DNA sequences are heterogeneous, so it is necessary to expand the methodology to capture the local behavior of such sequences. To address the problem of local behavior, a local spectral envelope with estimation via mixtures of smoothing splines will be explored. It is the hope that this methodology will help emphasize any periodic feature that exists in a categorical sequence of virtually any length in a quick and automated fashion. Projects such as the human genome project have produced large amounts of data and the methods established in this project will prove to be useful in the analysis of the vast quantities of data being produced by various genome projects. In another project, the focus is on the analysis of longitudinal data and the development of a practical nonparametric procedure for the estimation of the within-subject correlation structure. This technique is used to develop a data driven functional principal components analysis procedure (FPCA). Because longitudinal data often possess the property that observations made within a subject are correlated, an effective analysis of these data is required to account for this within-subject correlation. When a parametric form for the covariance structure is unknown, using a misspecified structure can result in biased and inefficient estimates. This project focuses on the analysis of longitudinal data that can be modeled as observations from smooth subject trajectories that are realizations of a stochastic process observed at discrete time points with noise. The high dimensionality and complexity of longitudinal data has made FPCA a popular tool for data reduction and visualization by capturing the primary modes of variation of the stochastic process generating the data. Scientists are often interested in using longitudinal data to determine the effect that a set of possibly time-varying covariates have on a given response over time. Functional linear models, and in particular the varying-coefficient model, provide a framework for analyzing such data. In many of these data sets, the functional coefficients have shapes that cannot be modeled parametrically. An effective analysis of these data is required to both account for the within-subject correlation and to allow for the flexible shapes of the coefficients. Because a parametric form for the within-subject covariance is not always known, a third project focuses on creating an iterative data-driven spline based procedure for fitting varying-coefficient models. This proposal concentrates on solving problems involved in the analysis of dependent data. The first project will develop a method for detecting genes in a long DNA sequences. Projects such as the human genome project have produced large amounts of data and the methods established in this project will prove to be useful in the analysis of the vast quantities of data being produced by various genome projects. A second proposed project focuses on the analysis of complex data collected over time. This project is also motivated by the analysis of DNA, and in particular, the analysis of gene expression data. In a third project, the investigators will focus on a technique called functional linear models. For example, techniques will be developed for studying the effect that a growth factor should have on the decision to supplement chemotherapy with antiangiogenic therapy when treating ovarian cancer.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Nonlinear and Nonstationary Time Series
-
批准号:1506882
-
项目类别:Continuing Grant
-
资助金额:$33.74万
-
财政年份:2015
-
负责人:David Stoffer
-
依托单位:
Collaborative Research: The Analysis of Time Series Collected in Experimental Designs
-
批准号:0706723
-
项目类别:Standard Grant
-
资助金额:$4.61万
-
财政年份:2007
-
负责人:David Stoffer
-
依托单位:
Time Series Analysis and Applications
-
批准号:0405038
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2004
-
负责人:David Stoffer
-
依托单位:
Statistical Methods in the Frequency Domain
-
批准号:0102511
-
项目类别:Continuing Grant
-
资助金额:$27.0万
-
财政年份:2001
-
负责人:David Stoffer
-
依托单位:
Expanding the Spectral Envelope
-
批准号:9703720
-
项目类别:Continuing Grant
-
资助金额:$15.0万
-
财政年份:1997
-
负责人:David Stoffer
-
依托单位:
The Spectral Envelope
-
批准号:9404343
-
项目类别:Standard Grant
-
资助金额:$5.9万
-
财政年份:1994
-
负责人:David Stoffer
-
依托单位:
Mathematical Sciences: Walsh-Fourier Analysis and Categorical Time Series
-
批准号:9000522
-
项目类别:Standard Grant
-
资助金额:$3.8万
-
财政年份:1990
-
负责人:David Stoffer
-
依托单位:
国内基金
海外基金
Computational Methods for Analyzing Toponome Data
-
批准号:60601030
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2006
-
负责人:Axel Mosig
-
依托单位: