课题基金 / 基金详情

Effective Computational Optimization in Data Mining and Financial Applications

Effective Computational Optimization in Data Mining and Financial Applications
数据挖掘和金融应用中的有效计算优化
批准号:
RGPIN-2014-03978
负责人:
Li, Yuying
金额:
$1.89万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2017
资助国家:
加拿大
项目状态:
已结题
起止时间:
2017-01-01 至 2018-12-31

项目摘要

项目成果

Li, Yuying的其他基金

相似基金

相关文献

中文摘要
翻译
在过去五年里,有两个词主导了全球新闻的讨论:市场不稳定和大数据。这项建议的主要重点是开发计算效率高和有效的方法,以最佳地利用现有数据,以改善医疗保健、商业和金融,包括检测和预防金融市场的系统性风险。在2008年金融市场崩溃后,许多人提出了一个问题:为什么没有对问题的迹象发出警告?如果这些迹象是存在的,为什么没有被检测到?我们如何才能更好地发现和预防这种未来的市场灭亡?《多德·弗兰克华尔街改革和消费者保护法》(Dodd Frank Wall Street改革and Consumer Protection Act)产生了一项社会使命,即从金融公司的事件中识别金融稳定面临的风险。尽管没有系统风险衡量的明确定义,但Biasias等人(2012)最近建议采用一个包含多种视角和流程的稳健框架,以动态调整系统风险衡量以适应金融市场结构的变化。Khandani等人(2010)将机器学习技术应用于银行交易和客户的信用局数据,以预测消费者信用风险。特别是,有人认为,预计违约的比例是消费贷款中系统性风险指标的一个信号。Hardle等人(2007)使用支持向量机的得分来估计金融公司的违约概率。随着数据的积累,高效和有效的数据挖掘方法有望为金融、商业和一般生活中面临的挑战性问题提供解决方案。《2011年麦肯锡大数据报告》估计,数据挖掘可能为美国医疗保健带来3000亿美元的年价值,为欧洲公共行政部门带来2500亿美元的年价值,以及6000亿美元的潜在消费者盈余。虽然在信息收集方面取得了重大进展,但要将这些估计变为现实,我们还需要在数据分析方面取得相应的进展。最近由遗产提供者网络(HPN)赞助的全球激励竞争表明了解决具有挑战性的数据分析问题的紧迫性。为了更早地识别高危人群并确保他们得到及时治疗,比赛的目标是创建使用患者数据预测住院人数的算法。大赛历时两年,奖金300万美元,吸引了来自世界各地不同学科的近2000名参赛者。与我的博士生Aditya Tayal*和同事Thomas Coleman一起,我们研究并开发了几种计算优化算法,最终使我们在比赛中获得了第四名的排名。数据挖掘方法有三个组成部分:最小化训练误差,最大化稳定性,以及平衡这两个目标的机制。数据挖掘中剩余的具有挑战性的优化问题通常是非凸的和大规模的。例如,在许多真实的数据分析问题中,可用的标签非常有限。我们如何使用部分标签信息来学习预测模型?我们如何从一组与特定预测任务相关的可用数据中最佳地选择特征?许多实际的数据挖掘问题都有罕见的类和多数类。我们如何开发计算效率高的非线性方法来解决这些不平衡问题?本文提出的研究的主要目标是解决数据挖掘中这些具有挑战性但相关的优化问题,并将其应用于医疗、金融、商业等行业。
英文摘要
In the past five years, two phrases dominate discussion in the global news: market instability and big data. The main focus of this proposal is to develop computationally efficient and effective methods to optimally utilize available data, in order to improve health care, business and finance, including detection and prevention of systematic risk in financial markets.After the 2008 financial market collapse, many have asked the question: why had there been no warnings of the signs of troubles? If these signs exist, why were they not detected? How can we achieve better detection and prevention of such a future market demise?The Dodd Frank Wall Street Reform and Consumer Protection Act has resulted in a societal mandate to identify risks to financial stability from the events of financial firms. Although there is no clear definition of systemic risk measure, Biasias et al (2012) recently propose that a robust framework incorporating a diverse collection of perspectives and processes be adopted to dynamically adapt systemic risk measures to changes in financial market structures. Khandani et al (2010) have applied machine learning techniques to bank transactions and credit-bureau data of customers in order to predict consumer credit risk. In particular, it is suggested that the proportion of predicted delinquencies is a signal to systemic risk indicator in consumer lending. Hardle et al (2007) use scores from support vector machines to estimate default probabilities of financial firms.With increasing accumulation of data, efficient and effective data mining methods stand to potentially offer solutions to challenging problems faced in finance, business, and our lives in general. The “2011 McKinsey Report on Big Data” estimates that data mining could potentially bring $300 billion annual value to US health care, 250 billion annual value to the European public administration sector, and a $600 billion potential consumer surplus. While there have been major advances in information gathering, to turn these estimates into realities, we need a commensurate advance in data analytics. The urgency of solving challenging data analysis problems is illustrated by the recent Heritage Provider Network (HPN) sponsored global incentivized competition. In an effort to identify at-risk individuals earlier and ensure they receive prompt treatment, the objective of the competition was to create algorithms that use patient data to predict hospitalizations. The competition ran for two years with a grand prize of $3 million, attracting nearly 2000 participants from various disciplines around the world. Together with my PhD student Aditya Tayal* and colleague Thomas Coleman, we investigated and developed several computational optimization algorithms that ultimately led us to securing a fourth place ranking in the competition.A data mining method has three components: minimizing training error, maximizing stability and a mechanism to balance the trade off between the two objectives. The remaining challenging optimization problems in data mining are typically nonconvex and large scale. For example, in many real data analysis problems, only very limited labels are available. How do we learn a predictive model, using partial label information? How do we optimally select features from a collection of available data that are relevant for a particular prediction task? Many practical data mining problems have a rare class and a majority class. How do we develop computationally efficiently nonlinear methods for these unbalanced problems? The main goal of the research proposed here is to solve these challenging but relevant optimization problems in data mining and apply them to health care, finance, business, and other industries.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Methodology of Learning Optimal Decisions from Market Data in Financial Technology
  • 批准号:
    RGPIN-2020-04331
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.99万
  • 财政年份:
    2022
  • 负责人:
    Li, Yuying
  • 依托单位:
Methodology of Learning Optimal Decisions from Market Data in Financial Technology
  • 批准号:
    RGPIN-2020-04331
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.99万
  • 财政年份:
    2021
  • 负责人:
    Li, Yuying
  • 依托单位:
A data driven approach for optimal stochastic control in finance
  • 批准号:
    530985-2018
  • 项目类别:
    Collaborative Research and Development Grants
  • 资助金额:
    $2.8万
  • 财政年份:
    2020
  • 负责人:
    Li, Yuying
  • 依托单位:
Methodology of Learning Optimal Decisions from Market Data in Financial Technology
  • 批准号:
    RGPIN-2020-04331
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.99万
  • 财政年份:
    2020
  • 负责人:
    Li, Yuying
  • 依托单位:
国内基金
海外基金
Computational Methods for Analyzing Toponome Data