课题基金 / 基金详情

III: Medium: SimSQL: A Database System Supporting Implementation and Execution of Distributed Machine Learning Codes

III: Medium: SimSQL: A Database System Supporting Implementation and Execution of Distributed Machine Learning Codes
III:媒介:SimSQL:支持分布式机器学习代码实现和执行的数据库系统
批准号:
1409543
负责人:
Christopher Jermaine
金额:
$120.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-01 至 2020-03-31

项目摘要

项目成果

Christopher Jermaine的其他基金

相似基金

相关文献

中文摘要
翻译
统计机器学习(ML)是一种用于分析超大型数据集的常用框架。在统计机器学习中,目标是学习一个统计模型,可以用来理解数据,发现模式或进行预测。 因此,许多新的软件系统已经被设计为支持在大型数据集上的并行/分布式ML计算机代码的简单实现和快速执行。 几乎所有这些系统都是“非关系”的,因为它们使用的数据和编程模型与今天的关系数据库管理系统非常不同。 尽管如此,关系或面向数据库的数据处理方法的吸引力仍然存在。例如,在数据库上运行的代码是声明性的,因此程序员只需要关心他或她想要什么,而不需要关心如何获得它。这使得编写代码并让它们在分布式环境中运行变得更容易,从而实现代码与数据库数据处理算法、存储、硬件和索引之间的强烈分离,甚至与数据库模式之间的分离。 此外,世界上的大部分结构化数据都位于关系数据库中,并且提取任何超过大型数据集的一小部分子样本以供外部使用通常都是不可行的。能够使用数据库引擎在数据库中执行ML推理代码,将大大提高统计ML的适用性。 该项目将进行必要的基础研究,使ML-in-the-database成为一项成熟的技术。该项目开发的所有想法都将在SimSQL的上下文中进行原型化,评估和分发,SimSQL是一个并行的关系数据库系统,具有执行“随机分析”的能力。这意味着SimSQL具有特殊的功能,允许用户定义具有模拟数据的特殊数据库表--这些数据实际上并不存储在数据库中,而是通过调用统计分布产生的。 由于SimSQL中的模拟数据表可以具有这种递归依赖关系,因此很容易使用SimSQL在“大数据”上运行随机ML推理算法(如MCMC)。研究任务包括通过利用大规模迭代ML计算提供的优化机会来提高SimSQL的性能水平。它们还包括扩展可以在SimSQL的SQL方言中轻松指定的ML推理算法的类型,使SimSQL适用于各种随机推理算法,如MCMC(马尔可夫链蒙特卡罗)和蒙特卡洛EM(期望最大化)。此外,该项目还将研究将R和类似BUG的ML算法规范自动编译为SimSQL SQL。该项目开发的所有软件都将在Apache许可证下开源。 代码可以从http://cmj4.web.rice.edu/SimSQL/SimSQL.html下载(更多信息可以在www.example.com找到)
英文摘要
Statistical machine learning (ML) is a commonly-applied framework for analyzing very large data sets. In statistical ML, the goal is to learn a statistical model that can be used to understand the data, find patterns, or make predictions. Thus, many new software systems have been designed to support easy implementation and fast execution of parallel/distributed ML computer codes over large data sets. Almost all of those systems are "non-relational" in the sense that they utilize data and programming models that are very different from today's relational database management systems. Still, the attractiveness of the relational or database-oriented approach to data processing persists. For example, codes running on top of a database are declarative, so a programmer need only be concerned with what he or she wants, and not how to obtain it. This makes it easier to write codes and get them to run in a distributed environment, enabling a strong separation between the code and the database data processing algorithms, storage, hardware, and indexing, and even from the database schema. Further, much of the world's structured data sits in relational databases, and extracting anything more than a small subsample of a large data set for external use is typically a non-starter. Being able to execute a ML inference code within the database, using the database engine, would greatly increase applicability of statistical ML. This project will perform the fundamental research necessary to make ML-in-the-database a mature technology. All of the ideas developed by the project will be prototyped, evaluated, and distributed within the context of SimSQL, which is a parallel, relational database system, augmented with the ability to perform "stochastic analytics". This means that SimSQL has special facilities that allow a user to define special database tables that have simulated data---these are data that are not actually stored in the database, but are produced by calls to statistical distributions. Since tables of simulated data in SimSQL can have such recursive dependencies, it is easy to use SimSQL to run stochastic ML inference algorithms (such as MCMC) over "Big Data". Research tasks include increasing the level of performance of SimSQL by exploiting the optimization opportunities presented by large-scale, iterative, ML computations. They also include expanding the types of ML inference algorithms that can easily be specified in SimSQL's SQL dialect, making SimSQL applicable for various stochastic inference algorithms such as MCMC (Markov Chain Monte Carlo) and Monte Carlo EM (Expectation Maximization). Further, the project will investigate automatically compiling R and BUGS-like ML algorithm specifications into SimSQL SQL. All of the software developed by the project will be available open source under the Apache license. The code can be downloaded from (and more information can be found at) http://cmj4.web.rice.edu/SimSQL/SimSQL.html
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: SHF: Medium: Semantics-Aware Neural Models of Code
  • 批准号:
    2212557
  • 项目类别:
    Standard Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2022
  • 负责人:
    Christopher Jermaine
  • 依托单位:
Collaborative Research: CISE-MSI: RPEP: III: celtSTEM Research Collaborative: Catapulting MSI Faculty and Students into Computational Research.
  • 批准号:
    2131294
  • 项目类别:
    Standard Grant
  • 资助金额:
    $48.85万
  • 财政年份:
    2021
  • 负责人:
    Christopher Jermaine
  • 依托单位:
III: Small: Applying Relational Database Design Principles to Machine Learning System Design
  • 批准号:
    2008240
  • 项目类别:
    Standard Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2020
  • 负责人:
    Christopher Jermaine
  • 依托单位:
MLWiNS: Wireless On-the-Edge Training of Deep Networks Using Independent Subnets
  • 批准号:
    2003137
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2020
  • 负责人:
    Christopher Jermaine
  • 依托单位:
海外基金