Designing Highly Recoverable Cloud Based Software Applications
Designing Highly Recoverable Cloud Based Software Applications
批准号:
RGPIN-2014-04611
负责人:
Khomh, Foutse
金额:
$1.68万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2015
资助国家:
加拿大
项目状态:
已结题
起止时间:
2015-01-01 至 2016-12-31
中文摘要
云计算是一种越来越流行的模式,它允许个人和企业通过互联网配置和部署软件应用程序。客户可以租用这些“云”应用程序(也称为云应用程序)提供的服务,根据需要增加或减少容量,只为自己使用的内容付费。云应用程序通常运行在Google App Engine、Windows Azure或OpenStack等云平台上。从金融、零售、教育和通信,到制造业、公用事业和交通运输,如今几乎每个行业都在使用云应用。Forrester Research预测,到2016年,云应用的销售额将翻两番以上(从今年的212亿美元增加到928亿美元),占整个软件市场的16%左右。然而,云应用程序的可靠性仍然是提供商和用户面临的一个主要问题。云应用程序的故障通常会导致巨大的经济损失,因为现在的核心业务活动都依赖于它们。2012年12月24日的情况就是这样,当时亚马逊网络服务故障导致Netflix云服务中断19小时。如今,对高度可靠的云应用的需求达到了前所未有的高水平。然而,业界仍没有明确的方法来开发高度可靠的云应用程序。开发人员通常将可靠性问题委托给运行这些应用程序的云平台。一条经验法则是跨多个可用区(AZ)复制服务,Netflix的策略总结道:“在多个AZ中部署,无需额外实例-目标自动扩展30%-60%,直到您有50%的净空空间来应对负载高峰。失去一个AZ会导致90%的利用率。”然而,悉尼研究人员进行的压力测试显示,由于服务过载、硬件故障、软件错误和运营商错误,亚马逊、谷歌和微软等大公司提供的基础设施和平台服务经常出现性能和可用性问题。据发现,这些服务的反应时间因一天中的时间不同而相差20倍。因此,云应用程序如果要高度可靠,就应该对故障具有健壮性。该研究计划的长期目标是开发技术和工具,以提高云应用程序的可恢复性。通过减少云应用的恢复时间,我们将能够提高它们的可靠性,并减少服务停机期间的损失。我将通过开发和应用一种新颖而创新的方法来实现这一目标,该方法由一个框架支持,将故障恢复机制纳入云应用程序的架构中。云计算的一大资产是不断增加的应用程序编程接口(API),允许开发人员将多个第三方服务集成到云应用程序中。云API平台的示例包括Apache CloudStack、Amazon Web Services、Eucalyptus、Simple Cloud和OpenStack。这些服务级别的API提供了大量冗余服务,这些服务可以整合到云应用程序中以实现容错。使用心跳或看门狗等模式,云应用程序可以监控它所依赖的特定服务,并在该服务出现故障时将请求重定向到备份服务,并触发恢复机制以维护高可用性。我将提出集成和监控服务的架构模式,以及在云应用架构中集成故障恢复机制的框架。
英文摘要
Cloud computing is an increasingly popular paradigm that allows individuals and enterprises to provision and deploy software applications over the Internet. Customers can lease services provided by these ‘cloud’ applications (a.k.a cloud apps), ramping up or down the capacity as they need and paying only for what they use. Cloud apps typically run on cloud platforms such as Google App Engine, Windows Azure, or OpenStack. Cloud apps are used in about every industry today; from financial, retail, education, and communications, to manufacturing, utilities and transportation. Forrester Research predicts that cloud apps sales will more than quadruple by 2016 (from $21.2 billion this year to $92.8 billion) to account for around 16% of the total software market. However, cloud apps dependability is still a major issue for both providers and users. Failures of cloud apps generally result in big economic losses as core business activities now rely on them. This was the case in December 24, 2012 when a failure of Amazon web services caused an outage of Netflix cloud services for 19 hours. The demand for highly dependable cloud apps has reached unprecedentedly high levels today. Yet, there is still no clear methodology in the industry for developing highly dependable cloud apps. Developers usually delegate dependability issues to the cloud platforms running the apps. A rule of thumb is to replicate services across multiple availability zones (AZ) as summarized by Netflix’s strategy: "Deploy in multiple AZ with no extra instances – target autoscale 30-60% until you have 50% headroom for load spikes. Lose an AZ leads to 90% utilization". Yet, stress tests conducted by Sydney-based researchers have revealed that infrastructure and platform services offered by big players like Amazon, Google, and Microsoft suffer from regular performance and availability issues due to service overload, hardware failures, software errors, and operator errors. The response times of these services was found to vary by a factor of twenty depending on the time of day. Therefore, cloud apps should be robust to failures if they are to be highly dependable. The long-term goal of this research program is to develop techniques and tools to improve the recoverability of cloud apps. By reducing the recovery time of cloud apps, we will be able to improve their dependability and reduce the amount of money lost during service-downtime. I will achieve this goal by developing and applying a novel and innovative methodology supported by a framework to incorporate fault recovery mechanisms in the architecture of cloud apps. One big asset of cloud computing is the constantly increasing number of Application Programming Interfaces (API) that allow developers to integrate multiple third party services into cloud apps. Examples of cloud API platforms include Apache CloudStack, Amazon Web Services, Eucalyptus, Simple Cloud, and OpenStack. These service-level APIs provide a lot of redundant services that can be incorporate in cloud apps to implement fault-tolerance. Using patterns like Heartbeat or Watchdog, a cloud app can monitor a specific service on which it depends and, in case of failure of this service, redirect requests to a backup service and trigger a recovery mechanism to maintain high availability. I will propose architectural patterns to integrate and monitor services, as well as a framework to integrate fault recovery mechanisms in the architecture of cloud apps.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Improving the Quality Assurance of Machine-Learning Software Applications
-
批准号:RGPIN-2019-06956
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.99万
-
财政年份:2022
-
负责人:Khomh, Foutse
-
依托单位:
Improving the Quality Assurance of Machine-Learning Software Applications
-
批准号:RGPIN-2019-06956
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.99万
-
财政年份:2021
-
负责人:Khomh, Foutse
-
依托单位:
A Comprehensive Framework for the Automatic Evaluation of the Quality of ML-based Software Systems
-
批准号:561420-2020
-
项目类别:Alliance Grants
-
资助金额:$4.74万
-
财政年份:2021
-
负责人:Khomh, Foutse
-
依托单位:
Improving the Quality Assurance of Machine-Learning Software Applications
-
批准号:RGPIN-2019-06956
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.99万
-
财政年份:2020
-
负责人:Khomh, Foutse
-
依托单位:
Improving the Quality Assurance of Machine-Learning Software Applications
-
批准号:RGPAS-2019-00083
-
项目类别:Discovery Grants Program - Accelerator Supplements
-
资助金额:$5.83万
-
财政年份:2020
-
负责人:Khomh, Foutse
-
依托单位:
Improving the Quality Assurance of Machine-Learning Software Applications
-
批准号:RGPAS-2019-00083
-
项目类别:Discovery Grants Program - Accelerator Supplements
-
资助金额:$2.91万
-
财政年份:2019
-
负责人:Khomh, Foutse
-
依托单位:
Improving the Quality Assurance of Machine-Learning Software Applications
-
批准号:RGPIN-2019-06956
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.99万
-
财政年份:2019
-
负责人:Khomh, Foutse
-
依托单位:
Designing Highly Recoverable Cloud Based Software Applications
-
批准号:RGPIN-2014-04611
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.68万
-
财政年份:2018
-
负责人:Khomh, Foutse
-
依托单位:
Applying Machine Learning Techniques to Automatically Process and Match Candidates Applications to Job Descriptions
-
批准号:521807-2017
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2017
-
负责人:Khomh, Foutse
-
依托单位:
Designing Highly Recoverable Cloud Based Software Applications
-
批准号:RGPIN-2014-04611
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.68万
-
财政年份:2017
-
负责人:Khomh, Foutse
-
依托单位:
Designing Highly Recoverable Cloud Based Software Applications
-
批准号:RGPIN-2014-04611
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.68万
-
财政年份:2016
-
负责人:Khomh, Foutse
-
依托单位:
Designing Highly Recoverable Cloud Based Software Applications
-
批准号:RGPIN-2014-04611
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.68万
-
财政年份:2014
-
负责人:Khomh, Foutse
-
依托单位:
海外基金