Interpolation of non-random missing values in financial statements' big data using CatBoost

Interpolation of non-random missing values in financial statements' big data using CatBoost
复制标题

DOI:
10.1007/s42001-022-00165-9
复制
发表时间:
2022-05-26
影响因子:
3.2
通讯作者:
Ishikawa,Atushi
Ishikawa,Atushi
中科院分区:
其他
文献类型:
--
作者:
Fujimoto,Shouji;Mizuno,Takayuki;Ishikawa,Atushi

文献摘要

相似文献

财务报表大数据具有“不完全性”和“非代表性”的特点。本文利用世界上最大的商业金融数据库ORBIS,首先发现数据缺失率因国家、金融项目类型和规模、年份而异。利用缺失数据的信息,我们从同一金融项目的前一年和/或下一年的值、其他金融项目的值以及CatBoost确定的缺失值的条件中插入非随机缺失的金融变量。由于财务价值在大范围内的分布服从齐夫定律,均值和方差发散,我们采用逆双曲函数将财务项目的价值转换为目标变量。针对目标变量缺失的两种情况,介绍了两种缺失插值模型。在验证了这些模型的准确性和稳定性之后,我们描述了插值非随机缺失值的企业规模变量的性质。在本工作的最后阶段,我们将这两个模型结合起来。从我们的观察中,我们确认齐夫定律建立的范围比插值之前更宽。
Financial statements’ big data have the characteristics of “Incompleteness” and “Nonrepresentative”. In this paper, employing the world’s largest commercial database on finance, ORBIS, we first find that the rate of missing data varies depending on the country, the type and size of financial items, and the year. Using information on missing data, we interpolate non-random missing financial variables from the previous- and/or next-year values of the same financial item, the values of other financial items, and the conditions of missing values determined by CatBoost. Because the distribution of financial values obeys Zipf’s law in the large-scale range and mean and variance diverge, we employ an inverse hyperbolic function to convert the value of a financial item as a target variable. We introduce two types of missing interpolation models according to the two types of situations involving missing objective variables. After verifying the accuracies and stabilities of these models, we describe the properties of firm-scale variables in which non-random missing values are interpolated. In the final stage of this work, we combine these two models. From our observations, we confirm that the range in which Zipf’s law is established becomes wider than before interpolation.