课题基金 / 基金详情

RII Track-4:NSF: Federated Analytics Systems with Fine-grained Knowledge Comprehension: Achieving Accuracy with Privacy

RII Track-4:NSF: Federated Analytics Systems with Fine-grained Knowledge Comprehension: Achieving Accuracy with Privacy
RII Track-4:NSF:具有细粒度知识理解的联合分析系统:通过隐私实现准确性
批准号:
2327480
负责人:
Hao Wang
金额:
$30.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-02-01 至 2026-01-31

项目摘要

项目成果

Hao Wang的其他基金

相似基金

相关文献

中文摘要
翻译
在医疗保健、广告、金融和公共交通等各个部门迅速采用大数据分析技术的同时,人们对数据隐私的认识和关注也在不断提高。数据监管领域的最新发展促使人们向保护隐私的数据分析转变,从而产生了联邦分析(FA),这是协作数据科学的主要范例,无需数据收集。这种数据分析范式的核心原则允许打破从有限的集中数据中获得分析的限制,包括隐私问题和操作成本。然而,FA系统的分布式特性和非数据共享的实施对数据分析的准确性和效率提出了严峻的挑战。首先,参与客户端的数据分布偏斜导致了严重的偏差和不一致。其次,在整个分析过程中使用的隐私保护技术导致数据效用和分析效率较差。该项目创新了精确、高效和可信的FA系统,具有细粒度的知识理解,以优化FA过程的整个生命周期。该项目将为与IBM t.j. Watson研究中心的研究人员在开发隐私保护FA解决方案方面的长期合作奠定坚实的基础。最后,该项目将培训路易斯安那州急需的保护隐私的数据科学人才。这项研究基础设施改进轨道4 EPSCoR研究研究员(RII轨道4)提案将为路易斯安那州立大学的助理教授提供奖学金,并为研究生提供培训。本项目旨在研究细粒度知识理解,以整体优化整个FA生命周期,包括数据偏度估计,参与者选择和隐私保护分析。探索细粒度知识理解将提供新的见解,以更好地理解和解释分布式非数据共享场景中的数据偏度和效用、隐私保护数据表示和分析结果。它将催化新的FA算法和系统设计,以优化FA性能和安全性。我们将具体的研究活动分解为两个协同目标:(1)通过数据偏度估计和自适应改进查询和客户端选择,提高具有细粒度数据偏度感知的FA准确性;(2)通过分离公共和个人特征表示,优化具有细粒度隐私保护的FA实用程序。拟议的FA系统将在LSU A&M和IBM的Maximo应用程序套件和OpenShift数据科学平台的公共数据集的现实大规模测试平台上进行广泛评估。所有数据集、基准测试和源代码都将在GitHub上发布,以产生更广泛的影响。通过利用细粒度的知识理解来提高FA的效率和隐私性,所提出的解决方案将推动FA功能的极限,并在现实场景中推动FA应用程序的发展。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
In parallel with the rapid adoption of big data analytics techniques by various sectors, such as healthcare, advertising, finance, and public transportation, there has been growing awareness and concern about data privacy. Recent developments in the data regulation landscape have prompted a seismic shift towards privacy-preserving data analytics, leading to Federated Analytics (FA), the leading paradigm for collaborative data science without data collection. The core principles of this data analysis paradigm allow for breaking the limitations of deriving analytics from limited centralized data in terms of privacy concerns and operational costs. However, FA systems' distributed nature and non-data-sharing enforcement raise critical challenges in the accuracy and efficiency of data analytics. First, skewed data distribution across participating clients leads to severe bias and inconsistency. Second, privacy-preserving techniques applied to the entire FA process leave poor data utility and analysis efficiency. This project innovates accurate, efficient, and credible FA systems with fine-grained knowledge comprehension to optimize the entire lifecycle of the FA process. This project will establish a solid foundation for long-term collaboration with researchers at IBM T. J. Watson Research Center toward developing privacy-preserving FA solutions. Lastly, the project will train a privacy-preserving data science workforce urgently needed in Louisiana.This Research Infrastructure Improvement Track-4 EPSCoR Research Fellows (RII Track-4) proposal would provide a fellowship to an Assistant Professor and training for a graduate student at Louisiana State University. This project aims to investigate fine-grained knowledge comprehension to optimize the entire FA lifecycle holistically, including data skewness estimation, participant selection, and privacy-preserved analysis. Exploring fine-grained knowledge comprehension will offer new insights to better understand and explain data skewness and utility, privacy-preserved data representations, and analytics results in a distributed non-data-sharing scenario. It will catalyze new FA algorithms and system designs toward optimizing FA performance and security. We decouple our specific research activities into two synergistic aims: (1) Improving FA accuracy with fine-grained data skew awareness by data skewness estimation and adaptive refinement of query and client selection; and (2) Optimizing FA utility with fine-grained privacy preservation by separating common and personal feature representations. The proposed FA systems will be extensively evaluated on realistic large-scale testbeds with public datasets at LSU A&M and IBM's Maximo Application Suite and OpenShift Data Science platform. All datasets, benchmarks, and source code will be released on GitHub for a broader impact. By harnessing fine-grained knowledge comprehension for escalating FA efficiency and privacy, the proposed solutions will push the envelope of FA's capabilities and spur the landscape of FA applications in real-world scenarios.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: OAC: Core: Harvesting Idle Resources Safely and Timely for Large-scale AI Applications in High-Performance Computing Systems
  • 批准号:
    2403398
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2024
  • 负责人:
    Hao Wang
  • 依托单位:
Collaborative Research: SaTC: CORE: Small: Critical Learning Periods Augmented Robust Federated Learning
  • 批准号:
    2315612
  • 项目类别:
    Standard Grant
  • 资助金额:
    $20.0万
  • 财政年份:
    2023
  • 负责人:
    Hao Wang
  • 依托单位:
CRII: OAC: High-Efficiency Serverless Computing Systems for Deep Learning: A Hybrid CPU/GPU Architecture
  • 批准号:
    2153502
  • 项目类别:
    Standard Grant
  • 资助金额:
    $17.49万
  • 财政年份:
    2022
  • 负责人:
    Hao Wang
  • 依托单位:
RI: Small: Enabling Interpretable AI via Bayesian Deep Learning
  • 批准号:
    2127918
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $49.99万
  • 财政年份:
    2021
  • 负责人:
    Hao Wang
  • 依托单位:
海外基金