Guided Task Planning Under Complex Constraints

Guided Task Planning Under Complex Constraints
复制标题

DOI:
10.1109/icde53745.2022.00067
复制
发表时间:
2022-05
期刊:
2022 IEEE 38th International Conference on Data Engineering (ICDE)
影响因子:
--
通讯作者:
Sepideh Nikookar;Paras Sakharkar;Baljinder Smagh;S. Amer-Yahia;Senjuti Basu Roy
Sepideh Nikookar;Paras Sakharkar;Baljinder Smagh;S. Amer-Yahia;Senjuti Basu Roy
中科院分区:
其他
文献类型:
--
作者:
Sepideh Nikookar;Paras Sakharkar;Baljinder Smagh;S. Amer-Yahia;Senjuti Basu Roy

文献摘要

相似文献

创建计划,即组成一系列项目来完成一项任务,如果手动完成,本质上是复杂的。这不仅需要找到一系列相关的项目,还需要了解用户需求并将其合并为约束。例如,在课程规划中,项目是核心和选修课程,学位要求将它们复杂的依赖关系作为约束条件。在旅行计划中,项目是兴趣点(POI),约束代表时间和金钱预算,这两个用户指定的要求。最重要的是,计划必须符合项目的理想交错,以实现一个目标,如提高学生的技能,以实现教育计划的更广泛的学习目标,或者在旅行场景中,改善整体用户体验。我们研究任务规划问题(TPP),目标是生成一系列在满足复杂约束的同时优化多个目标的项目。TPP被建模为一个受约束的马尔可夫决策过程,我们采用加权强化学习来学习满足项目、用户需求和满意度之间的复杂依赖关系的策略。我们提出了一个TPP的计算框架RL-Planner。RL-Planner需要来自领域专家(课程的学术顾问或旅行的旅行社)的最少输入,但生成满足所有限制的个性化计划。我们在大学项目和旅行社的数据集上进行了广泛的实验。我们将我们的解决方案与人类专家起草的计划和完全自动化的方法进行比较。我们的实验证实了现有的自动化解决方案不适合解决TPP问题,并且我们的计划与昂贵的手工制作的计划具有很高的可比性。
Creating a plan, i.e., composing a sequence of items to achieve a task is inherently complex if done manually. This requires not only finding a sequence of relevant items but also understanding user requirements and incorporating them as constraints. For instance, in course planning, items are core and elective courses, and degree requirements capture their complex dependencies as constraints. In trip planning, items are points of interest (POIs) and constraints represent time and monetary budget, two user-specified requirements. Most importantly, a plan must comply with the ideal interleaving of items to achieve a goal such as enhancing students' skills towards the broader learning goal of an education program, or in the travel scenario, improving the overall user experience. We study the Task Planning Problem (TPP) with the goal of generating a sequence of items that optimizes multiple objectives while satisfying complex constraints. TPP is modeled as a Constrained Markov Decision Process, and we adapt weighted Reinforcement Learning to learn a policy that satisfies complex dependencies between items, user requirements, and satisfaction. We present a computational framework RL-Planner for TPP. RL-Planner requires minimal input from domain experts (academic advisors for courses, or travel agents for trips), yet produces personalized plans satisfying all constraints. We run extensive experiments on datasets from university programs and from travel agencies. We compare our solutions with plans drafted by human experts and with fully automated approaches. Our experiments corroborate that existing automated solutions are not suitable to solve TPP and that our plans are highly comparable to expensive handcrafted ones.