课题基金 / 基金详情

BIGDATA: IA: Exploring Analysis of Environment and Health Through Multiple Alternative Clustering

BIGDATA: IA: Exploring Analysis of Environment and Health Through Multiple Alternative Clustering
BIGDATA:IA:通过多重替代聚类探索环境与健康分析
批准号:
1546428
负责人:
Jennifer Dy
金额:
$86.06万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-01-01 至 2020-12-31

项目摘要

项目成果

Jennifer Dy的其他基金

相似基金

相关文献

中文摘要
翻译
虽然由于大规模和多源数据收集已经变得普遍,许多学科已经变得越来越具有探索性,但我们发现,为了回答特定于项目的问题而仔细收集和研究的大量数据被忽视和未被研究,即使它们掌握着未来问题的关键答案。该项目通过开发新的数据分析、替代集群、数据可视化和加速解决方案来应对这一挑战,从而能够探索和识别隐藏在不同数据集中的联系,从而产生新的发现和知识。特别是,数据分析算法将应用于一个大型数据集,该数据集来自正在进行的国家环境健康科学研究所(NIEHS)项目,该项目正在评估波多黎各水基污染物对早产率的影响。探索性分析将侧重于发现可能更广泛地影响健康和环境的未知潜在环境因素和过程。因此,这项研究促进了数据科学、环境科学和健康方面的进步。请注意,该项目将针对服务不足人群中的妇女--S的健康。此外,该项目通过研究生研究支持、开发跨学科教程和创建一个新的本科生班级来支持教育,该班级解决了机器学习方法和并行计算之间的交叉。环境健康数据由多个具有不同时间和空间分辨率的不同来源组成:识别井水和自来水中目标化合物和非目标化合物的质谱仪读数,详细说明个人护理和家庭用品使用情况的参与者调查,以及分析胎盘、血液和尿液样本。这些复杂的数据在以下方面对传统的聚类算法提出了挑战。第一个挑战是为每种类型的数据源定义适当的相似性度量。第二步涉及如何集成来自这些多个来源的信息以进行集群。第三个挑战是,在探索性分析中,找到的解决方案可能不是分析师正在寻找的。有了这些知识,人们如何才能找到替代的解决方案呢?在现实世界的应用程序中,数据通常可以用许多不同的方式解释。然而,现有的多源融合方法只能找到单一的解。这项研究将开发新的可供选择的聚类方法,探索多种不同的信息源。该项目将提供可视化和可扩展的解决方案,以便高效地筛选堆积如山的数据。此外,该项目将在Spark内生成并行库,并展示这些新方法在环境健康应用中的力量。
英文摘要
While many disciplines have become increasingly exploratory given that large-scale and multi-source data collection has become prevalent, we find that volumes of data that were carefully collected and studied to answer project-specific questions, are neglected and unstudied, even if they hold key answers to tomorrow's questions. This project addresses this challenge by developing novel data analysis, alternative clustering, data visualization and acceleration solutions to enable exploration and identification of connections hidden in diverse data sets, leading to new discoveries and knowledge. In particular, the data analysis algorithms will be applied to a large dataset taken from an ongoing National Institute of Environmental Health Sciences (NIEHS) project that is assessing the impact of water-borne pollutants on premature birth rates in Puerto Rico. The exploratory analysis will focus on discovering the unknown underlying environmental factors and processes that may more broadly impact health and the environment. As such, this study promotes progress in data science, environmental science and health. Note that this project will address women?s health in an under-served population. In addition, this project supports education through graduate research support, development of inter-disciplinary tutorials, and creation of a new undergraduate class that addresses the intersection between machine learning approaches and parallel computing. The environmental health data comprises of multiple heterogeneous sources with varying temporal and spatial resolutions: mass spectrometer readings to identify targeted and non-targeted compounds in well and tap water, participant surveys detailing personal care and household products use in the home, and analyzed placental, blood and urine samples. Such complex data challenges traditional clustering algorithms in the following ways. The first challenge is in defining the appropriate similarity measure for each type of data source. The second step involves how to integrate information from these multiple sources for clustering. The third challenge is that in exploratory analysis, the solution found may not be what the analyst is looking for. How can one discover alternative solutions given this knowledge? In real world applications, data can often be interpreted in many different ways. However, existing multi-source fusion methods can only find a single solution. This study will develop new alternative clustering approaches, exploring multiple heterogeneous information sources. The project will deliver both visual and scalable solutions to enable sifting through the mountains of data efficiently. Moreover, the project will produce parallel libraries within Spark, and demonstrate the power of these new methods to this environmental health application.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CyberSEES: Type 2: SEA-MASCOT: Spatio-temporal Extremes and Associations : Marine Adaptation and Survivorship under Changes in extreme Ocean Temperatures
  • 批准号:
    1442728
  • 项目类别:
    Standard Grant
  • 资助金额:
    $119.96万
  • 财政年份:
    2014
  • 负责人:
    Jennifer Dy
  • 依托单位:
SIAM International Conference on Data Mining Student Travel Awards 2012
  • 批准号:
    1227843
  • 项目类别:
    Standard Grant
  • 资助金额:
    $2.6万
  • 财政年份:
    2012
  • 负责人:
    Jennifer Dy
  • 依托单位:
III: Small: Exploring Data in Multiple Clustering Views
  • 批准号:
    0915910
  • 项目类别:
    Standard Grant
  • 资助金额:
    $47.01万
  • 财政年份:
    2009
  • 负责人:
    Jennifer Dy
  • 依托单位:
CAREER: A Foundation for Unsupervised Learning of High-Dimensional Data
  • 批准号:
    0347532
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $0.0万
  • 财政年份:
    2004
  • 负责人:
    Jennifer Dy
  • 依托单位:
国内基金
海外基金
多任务深度学习融合多模态数据术前精准预测IA期非小细胞肺癌亚肺叶切除术复发风险
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    李琦
  • 依托单位:
Ia型超新星多波段实测特性及其机理研究
  • 批准号:
    JCZRYB202500270
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
  • 依托单位:
Ia型超新星及相关特殊天体研究
  • 批准号:
    12333008
  • 项目类别:
    重点项目
  • 资助金额:
    239.00万元
  • 批准年份:
    2023
  • 负责人:
    孟祥存
  • 依托单位:
南方根结线虫Mi-UNP与Bt-Cry1Ia36互作研究及其功能分析
  • 批准号:
    2023JJ30355
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2023
  • 负责人:
    成飞雪
  • 依托单位: