课题基金 / 基金详情

SHF: Medium: Collaborative Research: From Volume to Velocity: Big Data Analytics in Near-Realtime

SHF: Medium: Collaborative Research: From Volume to Velocity: Big Data Analytics in Near-Realtime
SHF:媒介:协作研究:从数量到速度:近实时的大数据分析
批准号:
1564207
负责人:
Tiark Rompf
金额:
$33.28万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-08-01 至 2021-07-31

项目摘要

项目成果

Tiark Rompf的其他基金

相似基金

相关文献

中文摘要
翻译
大多数现有的数据分析技术和系统都只关注大数据的体积方面,即体积、速度和种类。相比之下,有明确的迹象表明,在不久的将来,速度部分将成为主要的要求,最重要的是因为全球各地移动的设备的扩散。最新的数据通常包含最有价值的信息,并且用户已经习惯了通过复杂的机器学习(ML)技术进行深入分析和处理的数据,以实现他们的“永远在线”体验。例如,在大多数移动的交互中,一个或潜在的许多用户的物理位置起作用,但系统需要处理实际位置,而不是十分钟前的位置。在金融、情报和其他领域存在许多类似的用例。在所有这些领域中,对新鲜数据和高度处理数据的需求都处于根本性的紧张状态,因为高质量的分析在计算上是昂贵的,并且通常是大批量完成的。这个项目的智力价值是研究新思想的组合来应对这一挑战,跨越机器学习算法,专用硬件加速器,特定领域的语言和编译器技术。该项目更广泛的意义和重要性是为新型高速大数据分析铺平道路,这有可能彻底改变人们与世界互动的方式。该项目研究了新的增量ML原语和新算法,可以在速度和精度之间进行权衡,但保留可证明的保证。新的DSL(领域特定语言)使应用程序开发人员可以使用这些算法和技术,新的编译技术将DSL程序映射到专门的加速器。特别是,该项目展示了如何通过这些新的编译技术,机器学习算法可以特别受益于FPGA的硬件加速。最后,该项目研究了端到端数据路径优化的新编译技术,包括将外部格式的传入数据转换为DSL数据结构,以及在网络接口和FPGA加速器之间传输数据。将这些新的想法和技术结合在一起,该项目将产生一个集成的全栈解决方案(跨越算法,语言,编译器和架构),以实现大数据分析的高速问题。
英文摘要
Most existing techniques and systems for data analytics focus exclusively on the volume side of the common definition of Big Data as volume, velocity and variety. In contrast, there are clear indications that the velocity component will become the dominant requirement in the near future, most significantly because of the proliferation of mobile devices across the planet. This is compounded by the fact that the freshest data often contains the most valuable information and that users have grown accustomed to data that is deeply analyzed and processed by sophisticated machine learning (ML) techniques, to enable their "always on" experience. In most mobile interactions, for example, the physical locations of one or potentially many users play a role, but the system needs to process the actual locations, not the ones from ten minutes ago. Many similar use cases exist in finance, intelligence and other domains. In all of them, the desires for fresh and for highly processed data are in a fundamental tension, as high quality analysis is computationally expensive and often done in large batches. The intellectual merits of this project are to investigate a combination of new ideas to address this challenge, spanning machine learning algorithms, specialized hardware accelerators, domain-specific languages, and compiler technology. The project's broader significance and importance are to pave the way for new kinds of high-velocity big-data analytics, which have the potential to revolutionize the way that people interact with the world.The project investigates new incremental ML primitives and new algorithms that can trade off speed with precision, but retain provable guarantees. Novel DSLs (domain-specific languages) make such algorithms and techniques available to application developers, and new compilation techniques map DSL programs to specialized accelerators. In particular, the project shows how through these novel compilation techniques, machine learning algorithms can especially benefit from hardware acceleration with FPGAs. Finally, the project investigates new compilation techniques for end-to-end data path optimizations, including conversion of incoming data from external formats into DSL data structures, and transferring data between network interfaces and FPGA accelerators. Tying these new ideas and techniques together, this project will result in an integrated full-stack solution (spanning algorithms, languages, compilers, and architecture) to the problem of achieving high velocity in big data analytics.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
FMitF: Track I: Symbolic Reasoning with Graph Networks
  • 批准号:
    1918483
  • 项目类别:
    Standard Grant
  • 资助金额:
    $75.0万
  • 财政年份:
    2019
  • 负责人:
    Tiark Rompf
  • 依托单位:
CAREER: Generative Programming and DSLs for Safe Performance Critical Systems
  • 批准号:
    1553471
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $51.72万
  • 财政年份:
    2016
  • 负责人:
    Tiark Rompf
  • 依托单位:
海外基金