课题基金 / 基金详情

CRII: III: Real-World Machine Learning: Adaptation Methods for Addressing Temporal, Geographic, and Demographic Confounds in User-Generated Content

CRII: III: Real-World Machine Learning: Adaptation Methods for Addressing Temporal, Geographic, and Demographic Confounds in User-Generated Content
CRII:III:现实世界的机器学习:解决用户生成内容中的时间、地理和人口统计混乱的适应方法
批准号:
1657338
负责人:
Michael Paul
金额:
$17.41万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-09-01 至 2020-08-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
利用网络上的用户生成内容(如社交媒体信息)来得出关于世界的结论的研究正在迅速增长。使用机器学习和自然语言处理方法,可以根据人们在网上公开分享的想法和行为来估计公众舆论、消费者情绪和人口健康状况。例如,如果有人写他们发烧了,我们可能会推断他们得了流感;如果我们像这样汇总所有信息,我们就可以在人口水平上追踪流感的流行和传播。然而,将机器学习应用于用户生成内容的一个挑战是,内容的特征高度依赖于用户的身份、时间和地点。在线讨论发展迅速;在某一年构建的系统可能在下一年就不能很好地工作,为一个用户社区构建的系统可能不适用于另一个用户社区。拟议的项目旨在创建对内容和内容创作者的时间、地理和人口统计数据变化具有鲁棒性的机器学习方法。与机器学习中的领域适应技术相关,PI提出了学习跨这些不同内容属性进行泛化的方法。总体目标是创建健壮的、开放源代码的工具,以便其他研究人员可以轻松地采用。该项目的一个特别成果将是改进先前在基于社交媒体的疾病监测工作中使用的机器学习分类器。PI的健康分析系统的输出将被整合到HealthTweets.org,这是一个可公开访问的网站,为其他研究人员和卫生官员分享每日疾病流行的估计。该项目将为训练分类器创建层次贝叶斯模型,该模型可以适应不同的内容属性。感兴趣的具体属性包括时间、地理和作者的人口统计群体,但所提出的模型不依赖于具体属性,可以广泛应用于其他机器学习设置。作为起点,将构建一个预测模型(分类或回归),该模型可以一次适用于一个属性。然后,PI将为模型创建新的扩展,以适应多个属性(如时间和位置)的连接。这些扩展与PI之前构建结构化主题模型的工作有关,该模型可以学习内容的不同特征之间的关系。最后,除了创建预测模型外,PI还将构建可用于推断缺失属性的内容模型(例如,如果用户的位置未知),这些模型可以与预测模型相结合,共同执行推理和分类。在各种数据集上的新设置中的分类性能,以及对不同参数的影响和敏感性的探索将进行测试。具体的交付成果包括改进了Twitter上检测流感感染的分类器,并将分类器集成到HealthTweets.org网站中。
英文摘要
There is a rapidly growing body of research that uses user-generated content from the web, e.g., social media messages, to draw conclusions about the world. Using machine learning and natural language processing methods, it is possible to estimate public opinion, consumer sentiment, and population health based on what people are publicly sharing about their thoughts and actions online. For example, if someone writes that they have a fever, we might infer that they have the flu; if we aggregate all messages like this, we can track the prevalence and spread of the flu at a population level. However, a challenge with applying machine learning to user-generated content is that the characteristics of the content are highly dependent on the Who, When, and Where of the users. Online discussions evolve rapidly; a system built in one year might not work well in the next, and a system built for one community of users might not work for another. The proposed project seeks to create machine learning methods that are robust to variations in time, geography, and demographics of content and content creators. Related to domain adaptation techniques in machine learning, the PI proposes methods that learn to generalize across these various content attributes. The general goal is to create robust, open source tools that can be easily adopted by other researchers. One particular outcome of the project will be to improve the machine learning classifiers used in prior work on social media-based disease surveillance. The output of the PI's health analysis systems will be integrated into HealthTweets.org, a publicly accessible website that shares daily estimates of disease prevalence for other researchers and health officials. The project will create hierarchical Bayesian models for training classifiers that can be adapted across different content attributes. The specific attributes of interest include time, geography, and demographic group of the author, but the proposed models do not depend on the specific attributes, and can be broadly applied to other machine learning settings. As a starting point, a predictive model (classification or regression) will be constructed that can be adapted across one attribute at a time. The PI will then create novel extensions to the model that can adapt across conjunctions of multiple attributes, such as time AND location. These extensions are related to the PI's prior work on building structured topic models that learn relationships between different features of content. Finally, in addition to creating predictive models, the PI will also build models of content that can be used to infer missing attributes (e.g., the location of a user if it is unknown), which can be combined with the predictive models to jointly perform inference and classification. Classification performance in new settings on a variety of datasets and exploration of the effects of, and sensitivity to, different parameters will be tested. Specific deliverables include the improvement a classifier for detecting influenza infection on Twitter, and integrating the classifier into the website, HealthTweets.org.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
DOI: 10.18653/v1/s19-1015
发表时间: 2019-06
期刊:
影响因子: --
作者: [Xiaolei Huang;Michael J. Paul]
通讯作者: Xiaolei Huang;Michael J. Paul
DOI: 10.18653/v1/p19-1403
发表时间: 2019-07
期刊: EngRN: Computer-Aided Engineering (Topic)
影响因子: --
作者: [Xiaolei Huang;Michael J. Paul]
通讯作者: Xiaolei Huang;Michael J. Paul
Examining Temporality in Document Classification
检查文档分类中的临时性
DOI: --
发表时间: 2018
期刊: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers
影响因子: --
作者: [Huang, Xiaolei, Paul, Michael J.]
通讯作者: Paul, Michael J.
国内基金
海外基金
基于人工智能与多组学的III期结核性脓胸CT“低密度线”形成机制及手术时机预测模型研究
基于MOF–CRISPR微流控平台的雄黄As(III)/As(V)价态识别与炮制耦合机制研究
  • 批准号:
    JCZRLH202600780
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
  • 依托单位:
白术内酯III靶向IRF4-CD36轴通过调控脂质代谢重编程提升结直肠癌奥沙利铂敏感性的机制研究
  • 批准号:
    2026JJ82690
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    张卓
  • 依托单位:
基于废水零排放的FeS-As(III)置换法从污酸中清洁脱砷处理技术研究
  • 批准号:
    2026JJ30130
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    张二军
  • 依托单位: