"Sentometrics": Econometrics of textual sentiment with applications in economics and finance
"Sentometrics": Econometrics of textual sentiment with applications in economics and finance
批准号:
RGPIN-2022-03767
负责人:
Ardia, David
金额:
$1.97万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31
中文摘要
在计量经济学建模中,将情绪作为参数或变量使用的传统由来已久。从历史上看,使用问卷和代理来量化情绪变量一直是占主导地位的。近年来,由于通信媒体的数字化以及自然语言处理和机器学习技术的进步,对文本数据中蕴含的情感进行分析已经成为一种流行。确定媒体是否为经济和金融分析提供了潜在有价值的信息,这是“情报计量学”研究议程的目标之一。Sentometrics将计量经济学和机器学习技术结合起来,研究大量定性文本情感数据到定量情感变量的转换,以及它们在分析情感与其他变量之间的关系中的应用。这个项目的长期目标是设计新的计量经济学工具,以利用新闻媒体文章中关于情绪动态和来源的更多信息。这一研究项目的成果将填补这方面的几个研究空白。首先,它将开发模型,以确定新闻媒体是否有助于预测市场机制中经济和金融变量的变化。目前还缺乏这样的模型和分析。其次,它将调查是否可以以改善金融风险预测为具体目标来设计词典。虽然词典的优点是不是黑匣子,但目前的做法是依赖人类注释的词典,这些词典相当通用,可能包含偏见。第三,该建议将构建算法来改进联合情感-话题模型的估计。这些模型由于其计算成本和收敛性能较差,目前还没有得到广泛应用。我们希望设计算法使它们变得实用,并扩大它们的使用范围。最后,在项目每个阶段开发的开放源码将促进在社区中部署方法,以扩展工具或开发新的应用程序。改进的主题提取和特定领域的词典构建在其他领域也有意义,如营销和政策监测。在这方面,拟议项目的公开性非常重要,因此将加强直接知识转让。
英文摘要
There is a long-standing tradition of using sentiment as either a parameter or a variable in econometric modeling. Historically, the use of questionnaires and proxies to quantify sentiment variables has been predominant. In recent years, it has become popular to analyze the sentiment embedded in textual data due to the digitalization of communication media and progress in natural language processing and machine learning techniques. Determining whether media are carriers of potentially valuable information for economic and financial analysis is among the objectives of the "Sentometrics" research agenda. Sentometrics bridges econometrics and machine learning techniques to investigate the transformation of large volumes of qualitative textual sentiment data into quantitative sentiment variables and their subsequent application in analyzing the relationship between sentiment and other variables. The long-term objective of this project is to design new econometric tools to exploit more information on the dynamics and sources of sentiment in news media articles. The outputs of this research project will fill several research gaps in that direction. First, it will develop models to determine whether news media help anticipate changes in market regimes of economic and financial variables. Such models and analyses are currently lacking. Second, it will investigate if lexicons can be designed with the specific goal of improving financial risk forecasts. While lexicons have the advantage of not being black boxes, the current practice is to rely on human-annotated dictionaries, which are rather generic and may contain biases. Third, the proposal will construct algorithms to improve the estimation of joint sentiment-topic models. These models are not yet widely used due to their computational costs and poor convergence. We expect to design algorithms to render them practical and widen the scope of their usage. Finally, the open-source code developed at each phase of the project will foster the deployment of the methodologies in the community to extend the tools or develop new applications. Improved topic extraction and domain-specific lexicon construction are relevant in other fields such as marketing and policy monitoring. In this context, the openness of the proposed project is very relevant and will therefore enhance direct knowledge transfer.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金