Intelligent Data Engineering and Automated Learning - IDEAL 2022 - 23rd International Conference, IDEAL 2022, Manchester, UK, November 24-26, 2022, Proceedings

Intelligent Data Engineering and Automated Learning - IDEAL 2022 - 23rd International Conference, IDEAL 2022, Manchester, UK, November 24-26, 2022, Proceedings
复制标题

智能数据工程和自动化学习 - IDEAL 2022 - 第 23 届国际会议,IDEAL 2022,英国曼彻斯特,2022 年 11 月 24-26 日,会议记录

DOI:
10.1007/978-3-031-21753-1_42
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Cooper J
Cooper J
中科院分区:
--
文献类型:
--
作者:
Cooper J

文献摘要

相似文献

直觉上,相似的客户应该有相似的信用风险。通常使用客户特征之间的欧几里得距离和通过逻辑回归预测信用违约来尝试捕获这种相似性。在这里,我们探索使用拓扑数据分析来描述这种相似性。特别是,持久同调算法提供了与其拓扑相关的点云的摘要。这种方法已被证明在许多应用中是有用的,但据我们所知,将拓扑数据分析应用于信用风险预测是新颖的。我们开发了一种基于客户邻居拓扑分析的管道,邻居由几何网络结构确定。我们使用来自theLending Club的三个数据集和日本信贷筛选数据集发现了一个适度的信号。leveland肿瘤学数据集用于验证该管道。结果有很高的方差,但它们表明,当将这些拓扑特征作为逻辑回归中的附加解释变量时,可以改善信用风险预测。
Intuitively, similar customers should have similar credit risk. Capturing this similarity is often attempted using Euclidean distances between customer features and predicting credit default via logistic regression. Here we explore the use of topological data analysis for describing this similarity. In particular, persistent homology algorithms provide summaries of point clouds which relate to their topology. This approach has been shown to be useful in many applications but to the best of our knowledge, applying topological data analysis to prediction of credit risk is novel. We develop a pipeline which is based on the topological analysis of neighbourhoods of customers, with the neighbourhoods determined by a geometric network construction. We find a modest signal using three data sets from theLending Club, and theJapan Credit Screeningdata set. TheCleveland oncologicaldata set is used to validate the pipeline. The results have high variance, but they indicate that including such topological features could improve credit risk prediction when used as additional explanatory variable in a logistic regression.