Ember: No-Code Context Enrichment via Similarity-Based Keyless Joins

Ember: No-Code Context Enrichment via Similarity-Based Keyless Joins
复制标题

DOI:
10.14778/3494124.3494149
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
S. Suri;Ihab F. Ilyas;Christopher R'e;Theodoros Rekatsinas
S. Suri;Ihab F. Ilyas;Christopher R'e;Theodoros Rekatsinas
中科院分区:
其他
文献类型:
--
作者:
S. Suri;Ihab F. Ilyas;Christopher R'e;Theodoros Rekatsinas

文献摘要

相似文献

结构化数据或遵守预定义架构的数据可能会受到上下文碎片化的影响:描述单个实体的信息可能分散在为特定业务需求量身定做的多个数据集或表中,没有显式的链接键。在结构化数据源上的机器学习(ML)管道中,使用无键联接来丰富上下文或重建碎片化的上下文是一个隐式或显式的步骤。这个过程是乏味的、特定于领域的,并且在现在流行的无代码ML系统中缺乏支持,这些系统允许用户仅使用输入数据和高级配置文件来创建ML管道。作为回应,我们提出了Ember,这是一个抽象和自动化无键连接以泛化上下文丰富的系统。我们的主要见解是,Ember可以通过构建一个填充了特定于任务的嵌入的索引来启用通用的无键连接操作符。EMBER通过利用基于Transformer的表示学习技术来学习这些嵌入。我们在开发Ember时描述了我们的架构原则和运算符,并经验证明,Ember允许用户为包括搜索、推荐和问题回答在内的五个领域开发无代码上下文丰富管道,并且可以超过替代方案高达39%的召回率,只需更改一行配置即可。
Structured data, or data that adheres to a pre-defined schema, can suffer from fragmented context: information describing a single entity can be scattered across multiple datasets or tables tailored for specific business needs, with no explicit linking keys. Context enrichment, or rebuilding fragmented context, using keyless joins is an implicit or explicit step in machine learning (ML) pipelines over structured data sources. This process is tedious, domain-specific, and lacks support in now-prevalent no-code ML systems that let users create ML pipelines using just input data and high-level configuration files. In response, we propose Ember, a system that abstracts and automates keyless joins to generalize context enrichment. Our key insight is that Ember can enable a general keyless join operator by constructing an index populated with task-specific embeddings. Ember learns these embeddings by leveraging Transformer-based representation learning techniques. We describe our architectural principles and operators when developing Ember, and empirically demonstrate that Ember allows users to develop no-code context enrichment pipelines for five domains, including search, recommendation and question answering, and can exceed alternatives by up to 39% recall, with as little as a single line configuration change.