Transnational machine learning with screens for flagging bid‐rigging cartels

Transnational machine learning with screens for flagging bid‐rigging cartels
复制标题

DOI:
10.1111/rssa.12811
复制
发表时间:
2022-03
期刊:
Journal of the Royal Statistical Society: Series A (Statistics in Society)
影响因子:
--
通讯作者:
M. Huber;David Imhof;Rieko Ishii
M. Huber;David Imhof;Rieko Ishii
中科院分区:
其他
文献类型:
--
作者:
M. Huber;David Imhof;Rieko Ishii

文献摘要

被引文献

相似文献

我们调查了最初使用瑞士数据开发的统计筛选方法的跨国可转移性,以检测日本的投标操纵卡特尔。我们发现,将屏幕与机器学习(随机森林或由六种不同算法组成的集成方法)相结合来分类串通投标与竞争投标,在冲绳投标操纵卡特尔上训练和测试该方法时,正确分类率为88%-97%(取决于模型)。与瑞士一样,冲绳的操纵投标行为减少了差异,增加了投标分布的不对称性。当在来自一个国家的数据中训练模型以测试它们在来自另一个国家的数据中的性能时,所有机器学习器对真正串通投标和竞争性投标的正确预测之间的不平衡增加,并且当使用随机森林作为机器学习器时,分类率大幅下降,因为竞争性日本投标的一些筛选与串通性瑞士投标的筛选类似。贬低屏幕减少了由于国家之间的制度差异而导致的这种扭曲,使得基于一个国家的训练和另一个国家的测试的正确分类率达到85%,当使用集成方法作为机器学习者时达到90%,这通常优于随机森林。
We investigate the transnational transferability of statistical screening methods originally developed using Swiss data for detecting bid‐rigging cartels in Japan. We find that combining screens with machine learning (either a random forest or an ensemble method consisting of six different algorithms) to classify collusive versus competitive tenders entails (depending on the model) correct classification rates of 88%–97% when training and testing the method on the Okinawa bid‐rigging cartel. As in Switzerland, bid rigging in Okinawa reduced the variance and increased the asymmetry in the distribution of bids. When training the models in data from one country to test their performance in the data from the other country, imbalance increases between the correct prediction of truly collusive and competitive tenders for all machine learners and classification rates go down substantially when using the random forest as machine learner, due to some screens for competitive Japanese tenders being similar to those for collusive Swiss tenders. Demeaning the screens reduces such distortions due to institutional differences across countries such that correct classification rates based on training in one and testing in the other country amount to 85% and to 90% when using the ensemble method as machine learner, which generally outperforms the random forest.