课题基金 / 基金详情

III: Medium: Linear Algebra Operators in Databases to Support Analytic and Machine-Learning Workloads

III: Medium: Linear Algebra Operators in Databases to Support Analytic and Machine-Learning Workloads
III:中:数据库中的线性代数运算符支持分析和机器学习工作负载
批准号:
2312991
负责人:
Kenneth Ross
金额:
$101.63万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-07-01 至 2027-06-30

项目摘要

项目成果

Kenneth Ross的其他基金

相似基金

相关文献

中文摘要
翻译
机器学习工具在现代信息系统中无处不在。这些工具使用的数据输入通常来自关系数据库。这些工具生成的数据输出通常存储在数据库中,可用于后续的数据分析。但是,通常,学习过程本身是在数据库系统之外执行的。这个项目研究了在数据库本身内执行更多机器学习工作的机会,避免了昂贵的(通常是冗余的)数据导出和导入。哥伦比亚大学的团队将与relationship - ai和微软的研究人员合作,设计并构建两个相互作用的开源系统,分别名为MARQUE和ZORK。这些系统将使数据库驻留信息的数据分析更加高效和有效。提高效率将带来更快、更经济的机器学习,在DBMS中执行ML将简化操作复杂性,并受益于DBMS的特性,如可扩展性、访问控制和数据管理。最终,这项工作将扩大机器学习技术在广泛的数据密集型学科中的应用。MARQUE将是一个数据库管理系统,支持机器学习原语,如查询处理引擎上下文中的线性代数操作。该系统将使用数据库内的嵌入式机器学习模型有效地编译SQL查询,将最先进的查询处理技术与高度设计的线性代数算法相结合。MARQUE将允许机器学习管道本身的组件被制定为数据库内操作,避免不必要的数据复制。传统的SQL分析查询可以使用像矩阵乘法这样的运算符的扩展来重新表述,可以优化为使用高效的执行计划,包括针对这些运算符的专门算法。为了进一步支持数据库内机器学习,项目研究人员将构建ZORK,这是一个利用MARQUE提供的基础设施支持大规模机器学习的系统。ZORK将通过处理数据的因式表示来扩展到非常大的数据集,而不是显式地实现大型连接。该项目将为涉及传统关系运算符和广义线性代数运算符的查询开发新的和创新的查询处理技术。紧密集成将有助于操作符内部和操作符之间的查询优化。使用该系统,将开发一系列完全在数据库管理系统内运行的机器学习技术,避免数据导出并简化数据隐私管理等问题。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Machine-learning tools have become ubiquitous in modern information systems. The data inputs used by these tools often originate from relational databases. The data outputs generated by those tools are often stored in databases, where they can be used for subsequent data analysis. Typically, however, the learning process itself is performed outside the database system. This project investigates the opportunity for performing more of the machine learning work within the database itself, avoiding expensive (and often redundant) data export and import. In partnership with researchers from Relational-AI and Microsoft, the Columbia University team will design and build two interacting open-source systems named MARQUE and ZORK. These systems will make data analysis more efficient and effective for database-resident information. Improved efficiency will lead to faster, more cost-effective machine learning, and executing ML within the DBMS will simplify operational complexity and benefit from DBMS features such as scalability, access control, and data management. Ultimately, this work will broaden the adoption of machine learning technologies in a wide range of data-intensive disciplines.MARQUE will be a database management system that supports machine learning primitives such as linear algebra operations within the context of a query processing engine. The system will efficiently compile SQL queries using embedded machine learning models within the database, combining state-of-the-art query processing techniques with highly engineered linear algebra algorithms. MARQUE will allow components of the machine-learning pipeline itself to be formulated as in-database operations, avoiding unnecessary data copying. Conventional SQL analytic queries that can be reformulated using extensions of operators like matrix multiplication can be optimized to use efficient execution plans involving specialized algorithms for such operators. To further support in-database machine learning, the project investigators will build ZORK, a system to support machine learning at scale that will make use of the infrastructure provided by MARQUE. ZORK will scale to very large datasets by processing factorized representations of the data rather than explicitly materializing large joins. This project will develop new and innovative query processing techniques for queries involving both conventional relational operators and generalized linear algebra operators. Tight integration will facilitate query optimization within and between operators. Using this system, a range of machine learning techniques will be developed that operate entirely within the database management system, avoiding data export and simplifying concerns such as data privacy administration.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
III: Small: Bringing database query optimization to data intensive applications
  • 批准号:
    2008295
  • 项目类别:
    Standard Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2020
  • 负责人:
    Kenneth Ross
  • 依托单位:
Evolutionary Genomics of a Supergene Implicated in Social Evolution
III: Small: Database Algorithms for Modern CPU Memory Hierarchies
  • 批准号:
    1422488
  • 项目类别:
    Standard Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2014
  • 负责人:
    Kenneth Ross
  • 依托单位:
III: Small: Database Processing on GPUs
  • 批准号:
    1218222
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2012
  • 负责人:
    Kenneth Ross
  • 依托单位:
海外基金