Boosting flexible functional regression models with a high number of functional historical effects

Boosting flexible functional regression models with a high number of functional historical effects
复制标题

DOI:
10.1007/s11222-016-9662-1
复制
发表时间:
2017-07-01
影响因子:
2.2
通讯作者:
Greven, Sonja
Greven, Sonja
中科院分区:
数学2区
文献类型:
--
作者:
Brockhaus, Sarah;Melcher, Michael;Greven, Sonja

文献摘要

被引文献

相似文献

我们提出了一个具有功能响应的回归模型的通用框架,其中包含功能和标量协变量的潜在大量灵活效应。特别强调历史功能效应,其中在同一时间间隔内观察功能响应和功能协变量,并且响应仅受当前网格点之前的协变量值的影响。当在公共时间间隔内观察到功能响应和协变量时,主要使用历史功能效应,因为它们解释了时间顺序。我们的公式允许灵活的集成限制,包括提前或滞后时间。可以在不规则的曲线特定网格上观察功能响应。此外,我们为历史效应引入了不同的参数化,并讨论了可识别性问题。模型是通过逐分量梯度增强算法来估计的,该算法适用于具有潜在大量协变量效应的模型,甚至比观测值还要多,并且本质上进行模型选择。通过最小化相应的损失函数,可以对条件响应分布的不同特征进行建模,包括作为特殊情况的广义回归模型和分位数回归模型。这些方法在开源 R 包 FDboost 中实现。该方法的发展受到大肠杆菌发酵生物技术数据的推动,但涵盖了更广泛的模型类别。
We propose a general framework for regression models with functional response containing a potentially large number of flexible effects of functional and scalar covariates. Special emphasis is put on historical functional effects, where functional response and functional covariate are observed over the same interval and the response is only influenced by covariate values up to the current grid point. Historical functional effects are mostly used when functional response and covariate are observed on a common time interval, as they account for chronology. Our formulation allows for flexible integration limits including, e.g., lead or lag times. The functional responses can be observed on irregular curve-specific grids. Additionally, we introduce different parameterizations for historical effects and discuss identifiability issues.The models are estimated by a component-wise gradient boosting algorithm which is suitable for models with a potentially high number of covariate effects, even more than observations, and inherently does model selection. By minimizing corresponding loss functions, different features of the conditional response distribution can be modeled, including generalized and quantile regression models as special cases. The methods are implemented in the open-source R package FDboost. The methodological developments are motivated by biotechnological data on Escherichia coli fermentations, but cover a much broader model class.