课题基金 / 基金详情

CAREER: Towards a Big Data Application Server Stack

CAREER: Towards a Big Data Application Server Stack
职业:迈向大数据应用服务器堆栈
批准号:
1351047
负责人:
Tyson Condie
金额:
$46.47万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-02-01 至 2019-01-31

项目摘要

项目成果

Tyson Condie的其他基金

相似基金

相关文献

中文摘要
翻译
b谷歌的MapReduce启发了很多大数据分析工作,并作为Apache Hadoop等开源系统的模板。MapReduce编程模型具有广泛的适用性,但广泛的采用暴露了一些局限性,例如缺乏对迭代(这在机器学习算法中很常见)、流处理、图形分析、实时和交互式查询的支持。除了编程框架之外,底层实现还提供了一个如何向外扩展大规模分布式计算的模板:将它们分解为可以通过划分底层数据并行执行的小任务,并保存中间状态以减轻部分故障的影响(在大型集群上运行时必须对部分故障进行规划)。因此,挑战在于构建其他编程框架(例如SQL和机器学习)的实现,这些框架共享MapReduce相同的横向扩展和容错运行时特性,而不强加其局限性。Apache Hadoop YARN、谷歌Omega和Berkeley Mesos等资源管理器通过将资源分配从高级编程模型和语言的细节中分离出来,在这个方向上迈出了第一步。资源管理器在相同的底层机器集群上复用多个作业,从而提高利用率并培养全新的软件堆栈。当在容器中执行的任务(单个机器的资源(CPU/GPU、内存、磁盘)的一部分)完成时,容器将返回给资源管理器,供其他作业使用。与高级堆栈不同,容器是一个空白进程,设计用于承载任意计算。该项目规定了进一步的可重用软件层,以捕获诸如我应该为一项工作投入多少资源之类的问题?什么是冗余代码路径,我可以在可重用库中提供它们吗?什么是正确的语言和运行时抽象?在诸如MapReduce和相关SQL实现、ML工具包、存储系统和消息传递系统等系统的上下文中,在下一代资源管理器上探索这些问题,是我们工作的主要焦点。目标是在单一运行时层上统一一套大规模数据处理任务,构建在现代资源管理器(云操作系统)上。我们的结果将在专门的系统中提取出共性,并在单个底层运行时系统中提供它们,从而缩短上市时间。为下一个可供使用的大数据工具包,这反过来又会增加这些工具在更广泛社区的可用性。通过下一代资源管理器大规模实施和部署应用程序所获得的经验,可以帮助为未来云计算平台开发中的关键设计选择提供信息,从而影响广泛的科学、工程、国家安全、医疗保健和商业应用程序。该项目为研究生和本科生提供了更多基于研究的高级培训机会,包括目前在计算机科学、数据库、机器学习和云计算领域代表性不足的小组成员。
英文摘要
Google's MapReduce inspired much of the Big Data Analytics work and has served as a template for open source systems like Apache Hadoop. The MapReduce programming model has wide applicability, but widespread adoption has exposed some limitations, such as the lack of support for iteration (which is common in machine learning algorithms), stream processing, graph analytics, real-time and interactive queries. Beyond the programming framework, the underlying implementation offers a template for how to scale-out massively distributed computations: break them up into small tasks that can be carried out in parallel by partitioning the underlying data, and save intermediate state to mitigate the impact of partial failures (which must be planned for when running on large clusters). The challenge then, is to build implementations of other programming frameworks (e.g., SQL and machine learning) that share the same scale-out and fault-tolerance runtime characteristics of MapReduce without imposing its limitations. Resource managers such as Apache Hadoop YARN, Google Omega and Berkeley Mesos take a first step in this direction by separating resource allocation from the details of higher-level programming models and languages. Resource managers multiplex several jobs on the same underlying machine cluster, thereby increasing utilization and fostering clean-slate software stacks. When the task executing in a container a slice of a single machine's resources (CPU/GPU, memory, disk) is finished, the container is returned to the resource manager, where it is made available to other jobs. Unlike in higher-level stacks, a container is a blank-slate process, designed to host arbitrary computations. This project prescribes further reusable software layers that capture issues like how many resources should I dedicate to a job?; what are the redundant code-pathways and can I provide them in a reusable library?; what are the right language and runtime abstractions? Exploring these questions in the context of systems like MapReduce and related SQL implementations, ML toolkits, storage systems, and messaging systems, on next generation resource managers, is the primary focus of our work.The goal is to unify a suite of large-scale data processing tasks on a single runtime layer, built on modern resource managers (the cloud operating systems). Our results will factor out commonalities in specialized systems and provide them in a single underlying runtime system, shortening the time to ?market? for the next ready-to-use Big Data toolkit, which in turn would increase the availability of such tools to the broader community. Experience gained by implementing and deploying applications at scale, over next generation resource managers, could help inform critical design choices in the development of future cloud computing platforms, and hence impact a broad range of scientific, engineering, national security, healthcare and business applications. The project offers enhanced opportunities for research-based advanced training of graduate and undergraduate students, including members of groups that are currently under-represented in computer science, in databases, machine learning, and cloud computing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Student Travel Fellowships for ACM Symposium on Cloud Computing 2017
III: Medium: Collaborative Research: Scaling Machine Learning to Massive Datasets---A Logic Based Approach
  • 批准号:
    1302698
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $66.7万
  • 财政年份:
    2013
  • 负责人:
    Tyson Condie
  • 依托单位:
海外基金