Comparative text analytics via topic modeling in banking

Comparative text analytics via topic modeling in banking
复制标题

DOI:
10.1109/ssci.2017.8280945
复制
发表时间:
2017-11
期刊:
2017 IEEE Symposium Series on Computational Intelligence (SSCI)
影响因子:
--
通讯作者:
Yu Chen;Rhaad M. Rabbani;Aparna Gupta;Mohammed J. Zaki
Yu Chen;Rhaad M. Rabbani;Aparna Gupta;Mohammed J. Zaki
中科院分区:
其他
文献类型:
--
作者:
Yu Chen;Rhaad M. Rabbani;Aparna Gupta;Mohammed J. Zaki

文献摘要

相似文献

在本文中,我们比较和评估多个主题建模方法及其有效性,在分析了大量的美国上市银行提交给美国证券交易委员会。更具体地说,我们将四种主要的主题建模方法应用于2005-2016年578家银行控股公司的8-K和10-K文件的语料库。这些方法包括主成分分析,非负矩阵分解,潜在的Dirichlet分配和KATE,一种新的k-竞争文本文档的自动编码器。分别对于8-K和10-K,通过比较它们在两个分类任务上的性能来评估这些方法的有用性和有效性:(i)预测每个文件对应的部分,我们将8-K或10-K文件中的每个部分视为单独的文件,以及(ii)检测银行倒闭年份的文本,我们使用了2008年金融危机的银行倒闭数据。此外,我们定性地比较了不同方法发现的主题。我们得出的结论是,主题建模可以成为财务决策和风险管理的有效工具。
In this paper, we compare and evaluate multiple topic modeling approaches and their effectiveness in analyzing a large set of SEC filings by US public banks. More specifically, we apply four major topic modeling methods to a corpus of 8-K and 10-K filings, from the years 2005–2016, of 578 bank holding companies. These methods include Principal Component Analysis, Non-negative Matrix Factorization, Latent Dirichlet Allocation and KATE, a novel k-competitive autoencoder for text documents. Separately for 8-K and 10-K, the usefulness and effectiveness of these methods is evaluated by comparing their performances on two classification tasks: (i) predicting which section each document corresponds to, where we consider each section within an 8-K or 10-K filing as an individual document, and (ii) detecting text from a bank's year of failure, a task for which we use bank failure data from the 2008 financial crisis. In addition, we qualitatively compare the topics discovered by the different methods. We conclude that topic modeling can be an effective tool in financial decision making and risk management.