课题基金 / 基金详情

CAREER: Distilling information structure from big and dirty data: Efficient learning of clusters and graphs in modern datasets

CAREER: Distilling information structure from big and dirty data: Efficient learning of clusters and graphs in modern datasets
职业:从大数据和脏数据中提取信息结构:现代数据集中集群和图的高效学习
批准号:
1252412
负责人:
Aarti Singh
金额:
$50.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2013
资助国家:
美国
项目状态:
已结题
起止时间:
2013-03-01 至 2018-02-28

项目摘要

项目成果

Aarti Singh的其他基金

相似基金

相关文献

中文摘要
翻译
这个职业项目旨在推进从现代应用领域中出现的大量肮脏数据集中提取聚类和图形的理论和方法的最新水平。集群和图表提供了数据中包含的信息结构的有意义的表示,例如在神经科学和医疗保健领域,集群具有相似表型和基因的患者有助于确定药物设计的目标群体,集群由高分辨率数字表面成像(DSI)大脑扫描生成的纤维轨迹有助于识别重要的神经路径,而图表结构可以反映大脑区域之间的连接。这项工作的结果将显著提高利用这种现代数据集的能力,通过新的方法从大规模、高维、欠采样、损坏并且通常仅在压缩或流表示中可用的数据中学习簇和图。具体地说,该项目将开发用于学习簇和图的计算高效和原则性的方法,这些方法可以(I)执行无监督特征选择以丢弃高维中不相关的特征,(Ii)利用基于将资源集中在最具信息量的变量和特征的智能自适应查询的反馈,(Iii)使用适应于测量和计算效率的信息结构的压缩测量设计,以及(Iv)能够处理有噪声的流数据。算法将伴随着错误聚类率和图形恢复误差的精确表征的形式的性能保证。此外,该项目将调查这些问题中测量数量、计算复杂性和稳健性之间的权衡。开发的方法和理论将通过与神经科学和医疗保健领域的从业者合作,通过模拟以及它们在神经科学和医疗保健领域的真实数据集的适用性进行评估。这项研究的结果可能会潜在地改变许多应用领域,这些领域涉及基于大而脏的数据集对相似变量进行分组并学习它们之间的复杂交互作用。特别是,神经科学和医疗保健应用可能会对社会产生非常直接和重大的影响。准确绘制神经路径图将有助于早期诊断和治疗大脑病理,并有助于了解大脑功能。根据对相关基因特征或指标的少量测量,对患者进行分组并发现疾病传播途径,可以帮助预防和治愈疾病,并将医疗成本降至最低。研究活动将与教育努力紧密结合,旨在培养一支多样化的劳动力队伍,更好地配备跨学科工具,以应对现代数据集的挑战。该教育计划包括开发两门跨学科课程,以及加强卡内基梅隆大学(CMU)的统计与机器学习联合博士项目。外展活动包括促进本科生研究,通过OURCS(计算机科学本科生研究机会)、Andrew?S Leap(地区高中和中学生暑期充实计划)和针对卡内基梅隆大学高中和K-8教师的CS4HS计划,扩大女性和未被充分代表的群体在STEM领域的参与。该项目的成果(包括出版物、数据集和软件)将在http://www.cs.cmu.edu/~aarti/research_projects/.网站上在线发布
英文摘要
This CAREER project aims to advance the state-of-the-art in theory and methods for extracting clusters and graphs from big and dirty datasets arising in modern application domains. Clusters and graphs provide a meaningful representation of the structure of information contained in data, e.g. in neuroscience and health care domains, clustering patients with similar phenotypes and genotypes helps identify target groups for drug design, clustering fiber tracks generated by high-resolution Digital Surface Imaging (DSI) scans of brains help identify significant neural pathways, and graph structures can reflect connectivity between brain regions. The results of this work will significantly enhance the ability to exploit such modern datasets through new methods for learning clusters and graphs from data that is large-scale, high-dimensional, under-sampled, corrupted, and often only available in a compressed or streaming representation. Specifically, this project will develop computationally efficient and principled methods for learning clusters and graphs that can (i) perform unsupervised feature selection to discard irrelevant features in high dimensions, (ii) leverage feedback based on intelligent adaptive queries that focus resources on most informative variables and features, (iii) use compressive measurement design that adapts to the information structure for measurement and computation efficiency, and (iv) be able to handle noisy streaming data. The algorithms will be accompanied with performance guarantees in the form of a precise characterization of the mis-clustering rate and graph recovery error. Additionally, the project will investigate the tradeoffs between number of measurements, computational complexity and robustness in these problems. The methods and theory developed will be evaluated through simulations as well as their applicability to real datasets in neuroscience and healthcare domain, in collaboration with practitioners from these fields. The results of this research could potentially transform many application domains that involve grouping similar variables and learning complex interactions between them, based on big and dirty datasets. In particular, the neuroscience and healthcare applications are likely have very direct and significant implications for society. Accurately mapping neural pathways will help diagnose and treat brain pathologies at an early stage, and help understand brain functioning. Clustering patients and discovering disease spreading pathways based on few measurements of relevant genetic features or indicators could help prevent and cure diseases, and also minimize healthcare costs. The research activities will be tightly integrated with education efforts that aim to develop a diverse workforce that is better equipped with cross-disciplinary tools to address the challenges of modern datasets. The education plan includes development of two inter-disciplinary courses, and enhancement of the joint Statistics & Machine Learning PhD program at Carnegie Mellon University (CMU). Outreach activities include promoting undergraduate research, broadening participation of women and underrepresented groups in STEM fields through OurCS (Opportunities for Undergraduate Research in Computer Science), Andrew?s Leap (a summer enrichment program for area high school and middle school students) and CS4HS program aimed at High School and K-8 teachers at Carnegie Mellon University. The results of this project (including publications, data sets, and software) will be disseminated online at http://www.cs.cmu.edu/~aarti/research_projects/.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
AI Institute for Societal Decision Making (AI-SDM)
  • 批准号:
    2229881
  • 项目类别:
    Cooperative Agreement
  • 资助金额:
    $1987.97万
  • 财政年份:
    2023
  • 负责人:
    Aarti Singh
  • 依托单位:
Collaborative Research: New Perspectives on Deep Learning: Bridging Approximation, Statistical, and Algorithmic Theories
  • 批准号:
    2134133
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.0万
  • 财政年份:
    2021
  • 负责人:
    Aarti Singh
  • 依托单位:
QuBBD: Collaborative Research: Personalized Predictive Neuromarkers for Stress-Related Health Risks
  • 批准号:
    1557572
  • 项目类别:
    Standard Grant
  • 资助金额:
    $9.12万
  • 财政年份:
    2015
  • 负责人:
    Aarti Singh
  • 依托单位:
15th IMS New Researchers Conference
  • 批准号:
    1301845
  • 项目类别:
    Standard Grant
  • 资助金额:
    $2.5万
  • 财政年份:
    2013
  • 负责人:
    Aarti Singh
  • 依托单位:
海外基金