课题基金 / 基金详情

Kolmogorov complexity and its applications

Kolmogorov complexity and its applications
柯尔莫哥洛夫复杂度及其应用
批准号:
RGPIN-2016-03687
负责人:
Li, Ming
金额:
$4.59万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2019
资助国家:
加拿大
项目状态:
已结题
起止时间:
2019-01-01 至 2020-12-31

项目摘要

项目成果

Li, Ming的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
I am interested in developing a compelling theory of big data. Such a theory will depend on Kolmogorov complexity and information distance. Kolmogorov complexity is defined on one object. Information distance [C. Bennett, P. Gacs, M. Li, P. Vitanyi, W. Zurek, Information distance, IEEE Tran-IT, 44:4(1998)] is defined on two objects. This concept can be generalized to many objects. Using such a theory it is possible to optimally approximate the intuitive concept of "semantic distance" or closeness of two piece of data, in general. The key to this theory is to compress the data. Many ways of compressing data will be studied, including error encoding, clustering, and especially deep neural networks. Deep neural networks can be considered as ways of compressing data, especially big data. The following short-term goals are in tune with the above main theme of this research: ***1) Deep learning in natural language processing (NLP). My group has trained a Convolutional Neural Network (CNN) to map natural language questions to a database structured query with a limited number of relations. This work will continue. My group also has trained a Recurrent Neural Network (RNN) for conversation or chatting. This work will be extended to context sensitive chatting. This work will have two implications with the long term goal: a) Neural network will be studied as one way to approximate semantic distance; and b) Only with big data from the internet, this approach is practically useful.***2) Bioinformatics. A CNN has also been trained for protein identification as well as for peak-picking in mass spectrometry protein quantitation. These studies and methodologies will be extended to protein quantitation. This work again depends on huge amount of training data I have obtained from industry. ***These deep learning approaches will not be studied in isolation. They will be studied together with my theory of approximating semantic distance by information distance, experimenting with the efficiency of using deep neural networks as compression methods to deal with big data when there are no clear rules of compressing. I will also spend 8 months full time to revise his research book with Paul Vitanyi "An introduction to Kolmogorov complexity and its applications", that will include these new results. ***Several other short-term topics in bioinformatics will be studied. One is an antibody sequencing algorithm. I plan to design a new algorithm using linear programming to solve a bioinformatics industrial problem of antibody sequencing. Another problem is to apply the ideas in bioinformatics to other fields: optimal spaced seeds were invented by my group to do homology search. This has been considered one of the most influential innovations in bioinformatics during the last 15 years. I have the idea of using the optimal spaced seed idea to develop an observation theory to detect the trends in time series. Initial experiments were performed successfully.**
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Bioinformatics
  • 批准号:
    CRC-2015-00208
  • 项目类别:
    Canada Research Chairs
  • 资助金额:
    $14.57万
  • 财政年份:
    2022
  • 负责人:
    Li, Ming
  • 依托单位:
Kolmogorov complexity and algorithms for immunopeptidomics
  • 批准号:
    RGPIN-2022-02942
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $4.01万
  • 财政年份:
    2022
  • 负责人:
    Li, Ming
  • 依托单位:
Bioinformatics
  • 批准号:
    CRC-2015-00208
  • 项目类别:
    Canada Research Chairs
  • 资助金额:
    $14.57万
  • 财政年份:
    2021
  • 负责人:
    Li, Ming
  • 依托单位:
Kolmogorov complexity and its applications
  • 批准号:
    RGPIN-2016-03687
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $4.59万
  • 财政年份:
    2021
  • 负责人:
    Li, Ming
  • 依托单位:
海外基金