Hidden Components in Modern Applications
Hidden Components in Modern Applications
批准号:
1925845
负责人:
Zheng Ke
金额:
$14.18万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-07-01 至 2020-06-30
中文摘要
在大数据时代,研究人员经常会遇到规模庞大、结构复杂的数据集,感兴趣的信息通常包含在隐藏在大量噪声中的“组件”中。例子包括大型社会网络中的社区、文本文档中的主题以及全基因组关联研究(GWAS)中的混杂因素。提取这些隐藏的组件是一个有趣但具有挑战性的问题。该项目将解决这些挑战,并将应用于许多科学领域,包括社交网络、文本挖掘、基因组学和遗传学。该项目将包括(a)收集大型社会网络数据,(b)开发新的模型、方法和理论,用于提取网络分析、文本挖掘和全基因组关联研究中的隐藏成分,以及(c)利用学术研究数据(如合著者和引文关系)研究知识发现。这项研究将对语言学、社会科学、癌症研究和知识发现产生影响。本项目旨在发展统计模型、方法和理论,以推断和利用复杂数据,特别是矩阵数据中的隐藏成分。项目目标包括:(1)开发简单快速的网络混合隶属度估计和主题模型估计方法。这些方法基于主成分分析(PCA)的非平凡修改,易于实现,可以处理非常大的数据。(2)为GWAS中罕见和微弱效应的检测和估计提供了新的方法和理论。在复杂的相关结构和检测大协方差矩阵的弱尖峰存在的基因的最优排序的问题将被考虑。(3)科研人员社会网络结构研究。PI和她的合作者将从代表性统计期刊上发表的文章中收集元信息,以了解社会网络结构和统计社区的其他特征。(4)发展用于统计分析的新随机矩阵理论(RMT)。
英文摘要
In the era of Big Data, researchers often encounter datasets that are large in size and complex in structure where the information of interest is usually contained in "components" hidden in the enormous amount of noise. Examples include communities in large social networks, topics in text documents, and confounding factors in genome-wide association studies (GWAS). Extracting these hidden components is an interesting but challenging problem. This project will address these challenges and include applications to many scientific areas including social networks, text mining, genomics, and genetics. The project will include (a) collection of large social networks data, (b) development of new models, methods, and theory for extracting hidden components in network analysis, text mining, and genome-wide association studies, and (c) a study of knowledge discovery using academic research data such as co-authorship and citation relationships. The research will have an impact in linguistics, social sciences, cancer research, and knowledge discovery. This project aims to develop statistical models, methods, and theory for inferring and utilizing hidden components in complex data, especially matrix data. The goals of the project include: (1) Development of simple and fast methods for network mixed membership estimation and topic model estimation. These methods, based on nontrivial modifications of Principal Component Analysis (PCA), are easy to implement and can handle very large data. (2) New methods and theory for detecting and estimating rare and weak effects in GWAS. Problems related to optimal ranking of genes in the presence of complex correlation structures and detection of weak spikes in large covariance matrices will be considered. (3) A study of social network structures of scientific researchers. The PI and her collaborators will collect meta-information from published articles in representative statistics journals to understand social network structures and other features of the statistics community. (4) Development of new random matrix theory (RMT) for statistical analysis.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Improvements on SCORE, Especially for Weak Signals
SCORE 的改进,特别是对于弱信号
DOI:
10.1007/s13171-020-00240-1
发表时间:
2021
期刊:
Sankhya A
影响因子:
--
作者:
[Jin, Jiashun, Ke, Zheng Tracy, Luo, Shengming]
通讯作者:
Luo, Shengming
DOI:
--
发表时间:
2018-11
期刊:
ArXiv
影响因子:
--
作者:
[Yaqi Duan;Z. Ke;Mengdi Wang]
通讯作者:
Yaqi Duan;Z. Ke;Mengdi Wang
Network global testing by counting graphlets
通过计算 graphlet 进行网络全局测试
DOI:
--
发表时间:
2018
期刊:
Proceedings of Machine Learning Research
影响因子:
--
作者:
[Jin, Jiashun, Ke, Zheng Tracy, Luo, Shengming]
通讯作者:
Luo, Shengming
DOI:
10.1214/16-aos1522
发表时间:
2017-10-01
期刊:
ANNALS OF STATISTICS
影响因子:
4.5
作者:
[Jin, Jiashun, Ke, Zheng Tracy, Wang, Wanjie]
通讯作者:
Wang, Wanjie
DOI:
10.1111/ctr.13195
发表时间:
2018-05
期刊:
Clinical transplantation
影响因子:
2.1
作者:
[Peng RB, Lee H, Ke ZT, Saunders MR]
通讯作者:
Saunders MR
共 8 条
CAREER: Learning Probabilistic Factor Models
-
批准号:1943902
-
项目类别:Continuing Grant
-
资助金额:$40.0万
-
财政年份:2020
-
负责人:Zheng Ke
-
依托单位:
Hidden Components in Modern Applications
-
批准号:1712958
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2017
-
负责人:Zheng Ke
-
依托单位:
海外基金