A generalized covariate-adjusted top-scoring pair algorithm with applications to diabetic kidney disease stage classification in the Chronic Renal Insufficiency Cohort (CRIC) Study.

A generalized covariate-adjusted top-scoring pair algorithm with applications to diabetic kidney disease stage classification in the Chronic Renal Insufficiency Cohort (CRIC) Study.
复制标题

DOI:
10.1186/s12859-023-05171-w
复制
发表时间:
2023-02-20
期刊:
影响因子:
3
通讯作者:
--
中科院分区:
生物学4区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

不断增长的高维生物分子数据催生了新的统计和计算模型,用于风险预测和疾病分类。然而,尽管这些方法提供了很高的分类精度,但许多方法并不能产生生物学上可解释的模型。一个例外是,最高得分对(TSP)算法推导出在疾病分类中准确和健壮的无参数、生物可解释的单对决策规则。然而,标准的TSP方法不包括可能严重影响得分最高对的特征选择的协变量。在这里,我们提出了一种协变量调整的TSP方法,它使用协变量上特征回归的残差来识别得分最高的对。我们进行了模拟和数据应用来研究我们的方法,并将其与现有的分类器、套索和随机森林进行了比较。我们的模拟发现,与临床变量高度相关的特征在标准TSP设置中被选为得分最高的对的可能性很高。然而,通过残差化,我们的协变量调整的TSP能够识别新的得分最高的配对,这些配对基本上与临床变量无关。在数据应用中,使用在慢性肾功能不全队列研究中选择进行代谢组学分析的糖尿病患者(n = 977),标准TSP算法确定(Valine-甜菜碱,二甲基精氨酸)为糖尿病肾病严重程度分类的最高评分代谢物对,而协变量调整的TSP方法将这对代谢物对(匹扎酯,辛二醇)确定为最高评分。缬氨酸甜菜碱和二甲基精氨酸分别与尿白蛋白和血肌酐有 ≥ 0.4%的绝对相关性,而尿白蛋白和血肌酐是糖尿病肾病的预测指标。因此,在没有协变量调整的情况下,得分最高的一对主要反映了疾病严重程度的已知标记物,而协变量调整后的TSP揭示了从混淆中解放出来的特征,并确定了DKD严重程度的独立预后标记物。此外,基于TSP的方法在DKD中获得了与套索和随机森林相当的分类精度,同时提供了更简约的模型。通过一个简单、易于实现的残差过程,我们扩展了基于TSP的方法以考虑协变量。我们的协变量调整的TSP方法确定了与临床协变量无关的代谢物特征,这些特征基于两个特征之间的相对顺序来区分DKD严重程度阶段,从而为未来关于早期与晚期疾病状态的顺序逆转的研究提供了洞察力。网上版载有补充材料,可在10.1186/s12859-023-05171-w查阅。
The growing amount of high dimensional biomolecular data has spawned new statistical and computational models for risk prediction and disease classification. Yet, many of these methods do not yield biologically interpretable models, despite offering high classification accuracy. An exception, the top-scoring pair (TSP) algorithm derives parameter-free, biologically interpretable single pair decision rules that are accurate and robust in disease classification. However, standard TSP methods do not accommodate covariates that could heavily influence feature selection for the top-scoring pair. Herein, we propose a covariate-adjusted TSP method, which uses residuals from a regression of features on the covariates for identifying top scoring pairs. We conduct simulations and a data application to investigate our method, and compare it to existing classifiers, LASSO and random forests. Our simulations found that features that were highly correlated with clinical variables had high likelihood of being selected as top scoring pairs in the standard TSP setting. However, through residualization, our covariate-adjusted TSP was able to identify new top scoring pairs, that were largely uncorrelated with clinical variables. In the data application, using patients with diabetes (n = 977) selected for metabolomic profiling in the Chronic Renal Insufficiency Cohort (CRIC) study, the standard TSP algorithm identified (valine-betaine, dimethyl-arg) as the top-scoring metabolite pair for classifying diabetic kidney disease (DKD) severity, whereas the covariate-adjusted TSP method identified the pair (pipazethate, octaethylene glycol) as top-scoring. Valine-betaine and dimethyl-arg had, respectively, ≥ 0.4 absolute correlation with urine albumin and serum creatinine, known prognosticators of DKD. Thus without covariate-adjustment the top-scoring pair largely reflected known markers of disease severity, whereas covariate-adjusted TSP uncovered features liberated from confounding, and identified independent prognostic markers of DKD severity. Furthermore, TSP-based methods achieved competitive classification accuracy in DKD to LASSO and random forests, while providing more parsimonious models. We extended TSP-based methods to account for covariates, via a simple, easy to implement residualizing process. Our covariate-adjusted TSP method identified metabolite features, uncorrelated from clinical covariates, that discriminate DKD severity stage based on the relative ordering between two features, and thus provide insights into future studies on the order reversals in early vs advanced disease states. The online version contains supplementary material available at 10.1186/s12859-023-05171-w.
DOI: 10.2215/cjn.04260415
发表时间: 2015-11-01
影响因子: 9.8
作者:
Denker, Matthew;Boyle, Suzanne;Feldman, Harold I.
通讯作者: Feldman, Harold I.
DOI: 10.1021/ac201267k
发表时间: 2011-09-15
影响因子: 7.4
作者:
Fuhrer, Tobias;Heer, Dominik;Zamboni, Nicola
通讯作者: Zamboni, Nicola
DOI: 10.1681/asn.2013020126
发表时间: 2013-11-01
影响因子: 13.6
作者:
Sharma, Kumar;Karl, Bethany;Naviaux, Robert K.
通讯作者: Naviaux, Robert K.
DOI: 10.1214/14-aoas738
发表时间: 2014-09-01
影响因子: 1.8
作者:
Afsari, Bahman;Braga-Neto, Ulisses M.;Geman, Donald
通讯作者: Geman, Donald
DOI: 10.1038/s41598-020-59897-1
发表时间: 2020-02-19
期刊: SCIENTIFIC REPORTS
影响因子: 4.6
作者:
Parmar, Dharmeshkumar;Bhattacharya, Nivedita;Panchagnula, Venkateswarlu
通讯作者: Panchagnula, Venkateswarlu