Analysis of year-over-year changes in Risk Factors Disclosure in 10-K filings

Analysis of year-over-year changes in Risk Factors Disclosure in 10-K filings
复制标题

DOI:
10.1145/3220547.3220555
复制
发表时间:
2018-06
期刊:
Proceedings of the Fourth International Workshop on Data Science for Macro-Modeling with Financial and Economic Datasets
影响因子:
--
通讯作者:
Vipula Rawte;Aparna Gupta;Mohammed J. Zaki
Vipula Rawte;Aparna Gupta;Mohammed J. Zaki
中科院分区:
其他
文献类型:
--
作者:
Vipula Rawte;Aparna Gupta;Mohammed J. Zaki

文献摘要

被引文献

相似文献

向SEC提交的10-K表格中的风险因素披露(第1A项)是重要的部分之一,因为它包含了公司的年度风险更新,从而帮助投资者决定是否投资于一家公司。为了做出更好的投资选择,仔细阅读本节至关重要。鉴于每年提交的此类表格数量庞大,人类要理解和分析它们以做出明智的决定是非常麻烦的。我们讨论了银行失败分类的任务,使用文本分析的项目1A的各种银行的10-K表格,即,来预测一家银行是否会倒闭我们还分析了其他定量银行绩效指标,如杠杆率和资产回报率(罗阿),并看到如何以及基于文本的方法可以预测这些风险指标。特别是,为了创建我们的文本语料库,我们专注于1A部分的变化,只保留那些连续两年(同一家银行)相似度低于30%和40%的句子。我们实施深度学习和其他监督学习技术,如卷积神经网络(CNN),支持向量机(SVM)和线性回归。我们还将单词情感极性沿着它们的计数结合起来作为我们的加权特征向量。
Risk Factor Disclosures -- Item 1A -- in 10-K forms filed with SEC is one of the important sections since it contains a company's yearly risk updates, and thus helps investors decide whether to invest in a company or not. It is crucial to read this section carefully in order to make better investment choices. Given the large number of such forms filed on a yearly basis, it is very cumbersome for humans to understand and analyze them to make informed decisions. We discuss the task of bank failure classification using textual analysis on item 1A for various banks' 10-K forms, i.e., to predict whether a bank will fail or not. We also analyze other quantitative bank performance indicators like leverage and Return On Assets (ROA), and see how well text-based methods can predict those risk indicators. In particular, to create our textual corpora, we focus on the changes in the 1A sections, retaining only those sentences that have under 30% and 40% similarity over two consecutive years (for the same bank). We implement deep learning and other supervised learning techniques like Convolutional Neural Networks (CNN), Support Vector Machines (SVM) and Linear Regression. We also combine the word sentiment polarities along with their count as our weighted feature vector.