课题基金 / 基金详情

CRII: III: Real-World Machine Learning: Adaptation Methods for Addressing Temporal, Geographic, and Demographic Confounds in User-Generated Content

CRII: III: Real-World Machine Learning: Adaptation Methods for Addressing Temporal, Geographic, and Demographic Confounds in User-Generated Content
CRII:III:现实世界的机器学习:解决用户生成内容中的时间、地理和人口统计混乱的适应方法
批准号:
1657338
负责人:
Michael Paul
金额:
$17.41万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-09-01 至 2020-08-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
有一个快速增长的研究机构,使用用户生成的内容从网络上,例如,社交媒体上的信息,来得出关于这个世界的结论。 使用机器学习和自然语言处理方法,可以根据人们在网上公开分享的想法和行动来估计公众舆论,消费者情绪和人口健康。 例如,如果有人写他们发烧了,我们可能会推断他们得了流感;如果我们像这样汇总所有消息,我们就可以在人群水平上跟踪流感的流行和传播。 然而,将机器学习应用于用户生成内容的挑战在于,内容的特征高度依赖于用户的身份、时间和地点。 在线讨论发展迅速;一年内建立的系统可能在下一年就不能很好地工作,为一个用户群体建立的系统可能不适用于另一个用户群体。 拟议的项目旨在创建机器学习方法,这些方法对内容和内容创建者的时间,地理和人口统计数据的变化具有鲁棒性。 与机器学习中的领域自适应技术相关,PI提出了学习在这些不同的内容属性之间进行泛化的方法。 总体目标是创建强大的开源工具,可以很容易地被其他研究人员采用。 该项目的一个特别成果将是改进先前基于社交媒体的疾病监测工作中使用的机器学习分类器。 PI的健康分析系统的输出将被整合到HealthTweets.org,这是一个可公开访问的网站,为其他研究人员和卫生官员分享疾病流行率的每日估计。该项目将创建分层贝叶斯模型,用于训练分类器,这些分类器可以适应不同的内容属性。 感兴趣的特定属性包括作者的时间,地理和人口统计组,但所提出的模型不依赖于特定属性,并且可以广泛应用于其他机器学习设置。 作为起点,将构建一个预测模型(分类或回归),该模型可以一次在一个属性上进行调整。 然后PI将为模型创建新的扩展,这些扩展可以适应多个属性的连接,例如时间和位置。 这些扩展与PI先前构建结构化主题模型的工作有关,该模型可以学习内容的不同特征之间的关系。 最后,除了创建预测模型之外,PI还将构建可用于推断缺失属性的内容模型(例如,用户的位置(如果其是未知的),其可以与预测模型组合以联合执行推断和分类。将测试在新设置下对各种数据集的分类性能,以及对不同参数的影响和敏感性的探索。 具体的交付成果包括改进一个用于在Twitter上检测流感感染的分类器,并将该分类器集成到网站HealthTweets.org中。
英文摘要
There is a rapidly growing body of research that uses user-generated content from the web, e.g., social media messages, to draw conclusions about the world. Using machine learning and natural language processing methods, it is possible to estimate public opinion, consumer sentiment, and population health based on what people are publicly sharing about their thoughts and actions online. For example, if someone writes that they have a fever, we might infer that they have the flu; if we aggregate all messages like this, we can track the prevalence and spread of the flu at a population level. However, a challenge with applying machine learning to user-generated content is that the characteristics of the content are highly dependent on the Who, When, and Where of the users. Online discussions evolve rapidly; a system built in one year might not work well in the next, and a system built for one community of users might not work for another. The proposed project seeks to create machine learning methods that are robust to variations in time, geography, and demographics of content and content creators. Related to domain adaptation techniques in machine learning, the PI proposes methods that learn to generalize across these various content attributes. The general goal is to create robust, open source tools that can be easily adopted by other researchers. One particular outcome of the project will be to improve the machine learning classifiers used in prior work on social media-based disease surveillance. The output of the PI's health analysis systems will be integrated into HealthTweets.org, a publicly accessible website that shares daily estimates of disease prevalence for other researchers and health officials. The project will create hierarchical Bayesian models for training classifiers that can be adapted across different content attributes. The specific attributes of interest include time, geography, and demographic group of the author, but the proposed models do not depend on the specific attributes, and can be broadly applied to other machine learning settings. As a starting point, a predictive model (classification or regression) will be constructed that can be adapted across one attribute at a time. The PI will then create novel extensions to the model that can adapt across conjunctions of multiple attributes, such as time AND location. These extensions are related to the PI's prior work on building structured topic models that learn relationships between different features of content. Finally, in addition to creating predictive models, the PI will also build models of content that can be used to infer missing attributes (e.g., the location of a user if it is unknown), which can be combined with the predictive models to jointly perform inference and classification. Classification performance in new settings on a variety of datasets and exploration of the effects of, and sensitivity to, different parameters will be tested. Specific deliverables include the improvement a classifier for detecting influenza infection on Twitter, and integrating the classifier into the website, HealthTweets.org.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
DOI: 10.18653/v1/s19-1015
发表时间: 2019-06
期刊:
影响因子: --
作者: [Xiaolei Huang;Michael J. Paul]
通讯作者: Xiaolei Huang;Michael J. Paul
Examining Temporality in Document Classification
检查文档分类中的临时性
DOI: --
发表时间: 2018
期刊: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers
影响因子: --
作者: [Huang, Xiaolei, Paul, Michael J.]
通讯作者: Paul, Michael J.
DOI: 10.18653/v1/p19-1403
发表时间: 2019-07
期刊: EngRN: Computer-Aided Engineering (Topic)
影响因子: --
作者: [Xiaolei Huang;Michael J. Paul]
通讯作者: Xiaolei Huang;Michael J. Paul
国内基金
海外基金
基于人工智能与多组学的III期结核性脓胸CT“低密度线”形成机制及手术时机预测模型研究
基于MOF–CRISPR微流控平台的雄黄As(III)/As(V)价态识别与炮制耦合机制研究
  • 批准号:
    JCZRLH202600780
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
  • 依托单位:
白术内酯III靶向IRF4-CD36轴通过调控脂质代谢重编程提升结直肠癌奥沙利铂敏感性的机制研究
  • 批准号:
    2026JJ82690
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    张卓
  • 依托单位:
基于废水零排放的FeS-As(III)置换法从污酸中清洁脱砷处理技术研究
  • 批准号:
    2026JJ30130
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    张二军
  • 依托单位: