课题基金 / 基金详情

Harmonizing String and Unification-based Methodology with Machine Learning for Text Mining and Processing

Harmonizing String and Unification-based Methodology with Machine Learning for Text Mining and Processing
将基于字符串和统一的方法与用于文本挖掘和处理的机器学习相协调
批准号:
RGPIN-2019-05683
负责人:
Keselj, Vlado
金额:
$2.04万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2020
资助国家:
加拿大
项目状态:
已结题
起止时间:
2020-01-01 至 2021-12-31

项目摘要

项目成果

Keselj, Vlado的其他基金

相似基金

相关文献

中文摘要
翻译
该研究计划旨在从三个层面推进自然语言处理的最新技术,以满足更好的信息处理的需求。 在最低级别,字符和单词 n-gram 级别处理,我们的目标是通过使用可变长度 n-gram 配置文件来改进基于 n-gram 的文本挖掘,通过 n-gram 配置文件和相应欧拉图的可视化基于 n-gram 的视觉文本分析,在更深的模型级别将当前 CNG 距离度量与其他度量(例如,Jaccard、Dice)进行比较,使用 Google N-grams 数据改进标准语言 n-gram 配置文件,并适应规范化 google Distance 以实现离线距离。 在中间处理级别(基于 RegEx),我们将推进用于定向情感分析和噪声文本解析的正则表达式模式的开发,检查生成基于 RegEx 的模式的方法,生成 来自 Google N-grams 数据的模式,并扩展 Starfish 系统以进行文本嵌入处理。 在第三层,即统一层,我们的目标是:将子图同构技术从生物医学科学领域的分析转移到社交媒体的信息收集、维基百科数据的概念语义关系生成以及流文本数据的基于语义的可视化,例如电子邮件流的可视化。 我们的方法基于这三个级别的语言处理的先前工作:(1)通用 N 元语法分析(CNG),其中使用字符 N 元语法配置文件对文本数据进行建模; (2)基于正则表达式的文本数据处理,基于应用正则表达式重写模式,并将数据与相似模式进行匹配,以及(3)在统一级别,我们使用统一或子图同构应用信息提取和匹配,并且结构数据本身是通过使用基于随机统一的语法进行解析来生成的。 该方法的新颖性和预期意义基于改进方法以提供可视化文本分析,即可视化和与用户更密切的交互,并更好地 适应和开发来自互联网数据和社交媒体扩展的新型文本数据和新颖应用的方法。 这些方法的重要性得到了行业合作伙伴对社交媒体数据汇总分析领域的强烈兴趣的支持。
英文摘要
This research proposal aims at advancing the state of the art in the natural language processing at three levels in order to meet demands for better information processing. At the lowest level, the character and word n-gram level processing, our objectives are to improve n-gram based text mining through the use of variable-length n-gram profiles, n-gram based visual text analytics through visualization of n-gram profiles and corresponding Eulerian graphs, comparison of current CNG distance measure with other measures (e.g., Jaccard, Dice) at a deeper model level, use of Google N-grams data in improving the standard language n-gram profiles, and adaptation of Normalized google Distance to achieve an off-line distance. At the middle level of processing (RegEx based), we will advance development of regular expression patterns for directed sentiment analysis and parsing of noisy text, examining the ways to generate RegEx-based patterns, generating patterns from Google N-grams data, and extending the Starfish system for text-embedded processing. At the third level, the unification level, our bojectives are: to transfer sub-graph isomorphism technique from analysis in biomedical scientific domain to information gathering from social media, concept semantic relationship generation from Wikipedia data, and semantic-based visualization of stream textual data, such as visualization of e-mail streams. Our Approach is based on the previous work ot these three levels of language processing: (1) Common N-Gram analysis (CNG), where the text data is modelled using character n-gram profiles; (2) Regular Expression based processing of textual data, based on applying RegEx rewriting patterns, and matching the data with similar patterns, and (3) at the Unification level, we apply information extraction and matching using unification or sub-graph isomorphism, and the structural data itself is generated by parsing using the stochastic unification-based grammars. Novelty and Expected Significance of the approach is based on improving methodology to provide for visual text analysis, i.e., visualization and closer interaction with the user, and for better adaptation and development of methodology for new kind of textual data and novel applications coming from the expansion of Internet data and social media. The significance of the approaches is supported by strong interest coming from industrial partners in the area of summarized analysis of social media data.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Harmonizing String and Unification-based Methodology with Machine Learning for Text Mining and Processing
  • 批准号:
    RGPIN-2019-05683
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2022
  • 负责人:
    Keselj, Vlado
  • 依托单位:
Harmonizing String and Unification-based Methodology with Machine Learning for Text Mining and Processing
  • 批准号:
    RGPIN-2019-05683
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2021
  • 负责人:
    Keselj, Vlado
  • 依托单位:
Harmonizing String and Unification-based Methodology with Machine Learning for Text Mining and Processing
  • 批准号:
    RGPIN-2019-05683
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2019
  • 负责人:
    Keselj, Vlado
  • 依托单位:
String-based and Unification-based Methodology for Text mining and Processing
  • 批准号:
    262059-2013
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.46万
  • 财政年份:
    2018
  • 负责人:
    Keselj, Vlado
  • 依托单位:
国内基金
海外基金
带应力string方法及其在材料计算中的应用
  • 批准号:
    11001244
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    17.0万元
  • 批准年份:
    2010
  • 负责人:
    靳聪明
  • 依托单位: