课题基金 / 基金详情

III: Medium: RUI: Collaborative Research: Counterfactual Learning and Evaluation for Interactive Information Systems

III: Medium: RUI: Collaborative Research: Counterfactual Learning and Evaluation for Interactive Information Systems
III:媒介:RUI:协作研究:交互式信息系统的反事实学习和评估
批准号:
1901330
负责人:
Douglas Turnbull
金额:
$22.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-08-15 至 2024-07-31

项目摘要

项目成果

Douglas Turnbull的其他基金

相似基金

相关文献

中文摘要
翻译
许多信息系统通过以下交互循环与用户互动:系统接收上下文作为输入(例如查询,用户简介),响应上下文相关的操作(例如排名,推荐,广告),然后接收一些关于操作质量的显式或隐式反馈(例如星级评分,跟随搜索结果,点击广告)。虽然这种交互循环的日志数据无处不在且丰富,但它并不符合监督式学习的标准模式,因为反馈是有偏见的和不完整的——系统通过自己的行动来决定从哪里获得反馈,甚至对于选择的行动,它通常也不会观察到所有的反馈(例如,在排名中缺少相关结果的点击)。该项目将解决如何将这些记录的数据用于评估和学习新系统的问题。重用现有日志数据的潜在好处是显而易见的。对于评估,使用历史日志数据使工程师能够快速评估许多离线新系统(例如,新的排名功能,推荐策略),而不会延迟数周,也不会对在线A/B测试所隐含的用户体验产生潜在的负面影响。对于学习,它同样支持离线重用现有数据,而不是通过在线学习算法缓慢地收集新数据。这可以大大加快机器学习的开发周期,因为模型选择、特征选择和最终的质量控制可以在任何学习策略部署给用户之前离线进行。重复使用现有的日志数据对于小规模的信息系统(例如学术搜索)特别重要,因为它往往是唯一一种随时可以获得足够数量的潜在训练数据。该项目的智力价值在于开发有原则的机器学习方法,使信息系统能够可靠地从它们产生的部分和有偏见的反馈日志中学习。本研究的理论基础在于与反事实推理和因果推理之间的深层联系,利用日志和以行动为处理手段、现行制度为分配机制的对照实验之间的类比。这项研究建立在反事实估计器的最新进展之上,回答了这样一个问题:如果一个新系统被使用,而不是记录数据的系统,它将如何表现。该项目将开发新的反事实估计器,专门为信息系统(例如排名)中通常遇到的行动空间设计,新的倾向模型,以及结合两者的新的反事实政策学习算法。最后,为了验证研究在现实世界中的有效性,该项目将构建本地化系统,该系统提供本地音乐活动推荐和个性化播放列表。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Many information systems engage with their users through the following loop of interactions: the system receives a context as input (e.g. query, user profile), responds with a context-dependent action (e.g. ranking, recommendation, ad), and then receives some explicit or implicit feedback on the quality of the action (e.g. star rating, following a search result, clicking on an ad). While ubiquitous and plentiful, log data from this interaction loop does not fit the standard mold of supervised learning, since the feedback is both biased and partial -- the system determines through its actions where it gets feedback, and even for the chosen actions it typically doesn't observe all feedback (e.g. missing clicks on relevant results in ranking). This project will address the question of how this logged data can nevertheless be used for evaluating and learning new systems. The potential upsides of reusing the existing log data are evident. For evaluation, the use of historic log data enables engineers to rapidly evaluate many new systems offline (e.g. new ranking functions, recommendation policies), without the weeks of delay and the potential negative impact on user experience implied by online A/B testing. For learning, it similarly enables offline reuse of existing data instead of slowly collecting new data through an online learning algorithm. This can greatly speed up the machine-learning development cycle, since model selection, feature selection, and eventual quality control can happen offline before any learned policy gets deployed to the users. Reusing existing log data is particularly important for small-scale information systems (e.g. scholarly search), where it is often the only type of potential training data that is readily available in sufficient quantity.The intellectual merit of the project will lie in the development of principled machine learning methods that enable information systems to reliably learn from logs of the partial and biased feedback they produce. The theoretical basis for the research lies in deep connections to counterfactual and causal inference, exploiting the analogy between logs and controlled experiments with actions as treatments and the current system as the assignment mechanism. The research builds upon recent advances in counterfactual estimators, answering the question of how a new system would have performed, if it had been used instead of the system that logged the data. The project will develop new counterfactual estimators specifically designed for the action spaces typically encountered in information systems (e.g. rankings), new propensity models, and new counterfactual policy learning algorithms that incorporate both. Finally, to validate the real-world effectiveness of the research, the project will build the Localify system, which provides local music-event recommendations and personalized playlists.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
Exploring Acoustic Similarity and Preference for Novel Music Recommendation
探索小说音乐推荐的声学相似性和偏好
DOI: --
发表时间: 2020
期刊: International Symposium on Music Information Retrieval
影响因子: --
作者: [Cheng, Derek, Joachims, Thorsten, Turnbull, Douglas]
通讯作者: Turnbull, Douglas
Towards Quantifying the Strength of Music Scenes Using Live Event Data
使用现场活动数据量化音乐场景的强度
DOI: --
发表时间: 2022
期刊: International Society for Music Information Retrieval Conference
影响因子: --
作者: [Zhou, Michael, Mcgraw, Andrew, Turnbull, Douglas R.]
通讯作者: Turnbull, Douglas R.
Collaborative Research: III: Medium: Designing AI Systems with Steerable Long-Term Dynamics
  • 批准号:
    2312866
  • 项目类别:
    Standard Grant
  • 资助金额:
    $22.0万
  • 财政年份:
    2023
  • 负责人:
    Douglas Turnbull
  • 依托单位:
RI: Small: Collaborative Research: RUI: Batch Learning from Logged Bandit Feedback
  • 批准号:
    1615679
  • 项目类别:
    Standard Grant
  • 资助金额:
    $10.0万
  • 财政年份:
    2016
  • 负责人:
    Douglas Turnbull
  • 依托单位:
III: Small: Collaborative Research: RUI: Learning to Model Sequences
  • 批准号:
    1217485
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $18.6万
  • 财政年份:
    2012
  • 负责人:
    Douglas Turnbull
  • 依托单位:
NSF East Asia Summer Institutes for US Graduate Students
  • 批准号:
    0610260
  • 项目类别:
    Fellowship
  • 资助金额:
    $0.0万
  • 财政年份:
    2006
  • 负责人:
    Douglas Turnbull
  • 依托单位:
海外基金