An Extensible Multimodal Multi-task Object Dataset with Materials

An Extensible Multimodal Multi-task Object Dataset with Materials
复制标题

DOI:
10.48550/arxiv.2305.14352
复制
发表时间:
2023-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Trevor Scott Standley;Ruohan Gao;Dawn Chen;Jiajun Wu;S. Savarese
Trevor Scott Standley;Ruohan Gao;Dawn Chen;Jiajun Wu;S. Savarese
中科院分区:
其他
文献类型:
--
作者:
Trevor Scott Standley;Ruohan Gao;Dawn Chen;Jiajun Wu;S. Savarese

文献摘要

相似文献

我们提出了EMMa,一个可扩展的,多模式的亚马逊产品列表数据集,包含丰富的材料注释。它包含超过280万个对象,每个对象都有图像,列表文本,质量,价格,产品评级以及在亚马逊产品类别分类中的位置。我们还设计了182种物理材料的综合分类法(例如,塑料$\rightarrow$热塑性$\rightarrow$丙烯酸)。对象使用此分类中的一个或多个材质进行注释。由于每个对象都有许多属性,我们开发了一个智能标签框架,可以快速为所有对象添加新的二进制标签,只需很少的手动标签工作,使数据集可扩展。我们数据集中的每个对象属性都可以包含在模型输入或输出中,从而在任务配置中产生组合可能性。例如,我们可以训练一个模型来从列表文本中预测对象类别,或者从产品列表图像中预测质量和价格。EMMa为计算机视觉和NLP的多任务学习提供了一个新的基准,并允许从业者有效地大规模添加新任务和对象属性。
We present EMMa, an Extensible, Multimodal dataset of Amazon product listings that contains rich Material annotations. It contains more than 2.8 million objects, each with image(s), listing text, mass, price, product ratings, and position in Amazon's product-category taxonomy. We also design a comprehensive taxonomy of 182 physical materials (e.g., Plastic $\rightarrow$ Thermoplastic $\rightarrow$ Acrylic). Objects are annotated with one or more materials from this taxonomy. With the numerous attributes available for each object, we develop a Smart Labeling framework to quickly add new binary labels to all objects with very little manual labeling effort, making the dataset extensible. Each object attribute in our dataset can be included in either the model inputs or outputs, leading to combinatorial possibilities in task configurations. For example, we can train a model to predict the object category from the listing text, or the mass and price from the product listing image. EMMa offers a new benchmark for multi-task learning in computer vision and NLP, and allows practitioners to efficiently add new tasks and object attributes at scale.