Interactive Curation of Datasets for Training and Refining Generative Models

Interactive Curation of Datasets for Training and Refining Generative Models
复制标题

DOI:
10.1111/cgf.13844
复制
发表时间:
2019-10
影响因子:
2.5
通讯作者:
Wenjie Ye;Yue Dong;P. Peers
Wenjie Ye;Yue Dong;P. Peers
中科院分区:
计算机科学4区
文献类型:
--
作者:
Wenjie Ye;Yue Dong;P. Peers

文献摘要

相似文献

我们提出了一种基于交互式学习的新颖方法,用于使用用户定义的标准来整理数据集,以训练和优化生成对抗网络。我们采用一种新颖的批量模式主动学习策略,逐步选择小批量候选样本,要求用户指出这些样本是否符合(可能是主观的)选择标准。在每一批之后,对模拟用户意图的分类器进行优化,随后用于选择下一批候选样本。在选择过程结束后,使用有限但自适应选择的训练数据训练的最终分类器,用于筛选大量输入样本,以提取足够大的子集来训练或优化符合用户选择标准的生成模型。我们系统的一个关键区别特征是,我们不假定用户总是能够对每个候选样本做出明确的二元决策(即“符合”或“不符合”选择标准),并且我们允许用户将一个样本标记为“未决定”。我们依靠一种非二元的委员会查询策略来区分用户的不确定性和训练后的分类器的不确定性,并开发了一种新颖的分歧距离度量来鼓励多样化的候选集。此外,还采用了一些优化策略来实现交互式体验。我们在与训练或优化生成模型相关的几个应用中展示了我们的交互式整理系统:训练一个符合用户定义标准的生成对抗网络,调整现有生成模型的输出分布,以及从生成模型中去除不需要的样本。
We present a novel interactive learning‐based method for curating datasets using user‐defined criteria for training and refining Generative Adversarial Networks. We employ a novel batch‐mode active learning strategy to progressively select small batches of candidate exemplars for which the user is asked to indicate whether they match the, possibly subjective, selection criteria. After each batch, a classifier that models the user's intent is refined and subsequently used to select the next batch of candidates. After the selection process ends, the final classifier, trained with limited but adaptively selected training data, is used to sift through the large collection of input exemplars to extract a sufficiently large subset for training or refining the generative model that matches the user's selection criteria. A key distinguishing feature of our system is that we do not assume that the user can always make a firm binary decision (i.e., “meets” or “does not meet” the selection criteria) for each candidate exemplar, and we allow the user to label an exemplar as “undecided”. We rely on a non‐binary query‐by‐committee strategy to distinguish between the user's uncertainty and the trained classifier's uncertainty, and develop a novel disagreement distance metric to encourage a diverse candidate set. In addition, a number of optimization strategies are employed to achieve an interactive experience. We demonstrate our interactive curation system on several applications related to training or refining generative models: training a Generative Adversarial Network that meets a user‐defined criteria, adjusting the output distribution of an existing generative model, and removing unwanted samples from a generative model.