课题基金 / 基金详情

Aspect Classification of Social Documents

Aspect Classification of Social Documents
社会文献方面分类
批准号:
488936-2015
负责人:
Sadat, Fatiha
金额:
$1.82万
依托单位国家:
加拿大
项目类别:
Engage Grants Program
财政年份:
2015
资助国家:
加拿大
项目状态:
已结题
起止时间:
2015-01-01 至 2016-12-31

项目摘要

项目成果

Sadat, Fatiha的其他基金

相似基金

相关文献

中文摘要
翻译
目前的研究项目旨在创建工具,用于分析、建模和将社会文件分类为一般类别。这项研究涉及两种语言:法语和英语。Twitter是一种流行的短信服务,可通过网页、桌面和移动软件使用。在当前的研究中,Twitter帖子将被视为社会文档的语料库。通过单语言社会文献分类和跨语言社会文献分类说明了项目的主要步骤。
英文摘要
The current research project aims at creating tools for analyzing, modelling and classifying social documents into general categories. Two languages are concerned in this research: French and English. Twitter is a popular short messaging service available through Web page, desktop and mobile software. In the current research, Twitter posts will be considered as a corpus of social documents. The main steps of the project are explained through the monolingual social documents classification and the Cross-Language social documents classification. In this research project, various well-known classification models in the machine learning field will be tested and compared to each others such as Naïve Bayes, Random Forest, Support Vector Machine, etc. Moreover, Tweet Natural Language Processing (NLP) tools such as part-of-speech taggers and a dependency parser will be used in order to extract several features based on NLP knowledge of the tweets for the classifiers. Our objective is to identify the optimal combination of features that yields good prediction results, while avoiding overfitting. The Cross-lingual text classification is a major challenge in NLP, since often training data is available in only one language (target language), but not available for the language of the document we want to classify (source language). Classifying French tweets will be more complex and challenging as we do not have affordable Tweet NLP tools for this language in order to apply the same classification method. One can proceed in two ways: First, a monolingual classification method as explained earlier, will be applied on the French social documents, with considering fewer NLP features. The second solution is a Cross-lingual Text Classification Using topic-dependent word probabilities. Having social documents in the two languages, one can adopt a naïve approach by considering the combination of the multiple independent monolingual (and cross-language) text classifiers.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Coping with Zero-Shot Translation and its Explainability
  • 批准号:
    RGPIN-2019-07242
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.48万
  • 财政年份:
    2022
  • 负责人:
    Sadat, Fatiha
  • 依托单位:
Coping with Zero-Shot Translation and its Explainability
  • 批准号:
    RGPIN-2019-07242
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.48万
  • 财政年份:
    2021
  • 负责人:
    Sadat, Fatiha
  • 依托单位:
Coping with Zero-Shot Translation and its Explainability
  • 批准号:
    RGPIN-2019-07242
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.48万
  • 财政年份:
    2020
  • 负责人:
    Sadat, Fatiha
  • 依托单位:
Coping with Zero-Shot Translation and its Explainability
  • 批准号:
    RGPIN-2019-07242
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.48万
  • 财政年份:
    2019
  • 负责人:
    Sadat, Fatiha
  • 依托单位:
海外基金