RII Track-4:NSF: Federated Analytics Systems with Fine-grained Knowledge Comprehension: Achieving Accuracy with Privacy
RII Track-4:NSF: Federated Analytics Systems with Fine-grained Knowledge Comprehension: Achieving Accuracy with Privacy
批准号:
2327480
负责人:
Hao Wang
金额:
$30.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-02-01 至 2026-01-31
中文摘要
在医疗、广告、金融和公共交通等各个行业迅速采用大数据分析技术的同时,人们对数据隐私的意识和担忧也越来越多。数据监管领域的最新发展推动了向隐私保护数据分析的巨大转变,导致了联合分析(FA),这是无需数据收集的协作数据科学的领先范例。这种数据分析范式的核心原则允许打破从有限的集中数据中获取分析在隐私问题和运营成本方面的限制。然而,FA系统的分布式性质和非数据共享强制执行在数据分析的准确性和效率方面提出了严峻的挑战。首先,跨参与客户的数据分布不对称会导致严重的偏见和不一致。其次,隐私保护技术应用于整个FA过程,导致数据利用率和分析效率较低。该项目用细粒度的知识理解创新了准确、高效和可信的FA系统,以优化FA过程的整个生命周期。该项目将为与IBM T.J.沃森研究中心的研究人员开发隐私保护FA解决方案的长期合作奠定坚实的基础。最后,该项目将培训路易斯安那州迫切需要的隐私保护数据科学劳动力。这项研究基础设施改善轨道4 EPSCoR研究人员(RII轨道4)提案将为路易斯安那州立大学的一名助理教授提供奖学金,并为一名研究生提供培训。该项目旨在研究细粒度的知识理解,以整体优化整个FA生命周期,包括数据偏斜度估计、参与者选择和隐私保护分析。探索细粒度的知识理解将提供新的见解,以更好地理解和解释数据偏斜和效用、隐私保护的数据表示以及分布式非数据共享场景中的分析结果。它将促进新的FA算法和系统设计,以优化FA的性能和安全性。我们将具体的研究活动分解为两个协同目标:(1)通过数据倾斜估计和查询和客户选择的自适应精化来提高FA的精确度和细粒度数据倾斜感知;(2)通过分离公共和个人特征表示来优化FA效用和细粒度隐私保护。建议的FA系统将在现实的大规模试验台上进行广泛评估,这些试验台使用LSU A&Amp;M和IBM的Maximo应用套件和OpenShift数据科学平台的公共数据集。所有数据集、基准测试和源代码都将在GitHub上发布,以产生更广泛的影响。通过利用细粒度的知识理解来提升FA的效率和隐私,建议的解决方案将推动FA的能力极限,并促进FA在现实世界场景中的应用。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
In parallel with the rapid adoption of big data analytics techniques by various sectors, such as healthcare, advertising, finance, and public transportation, there has been growing awareness and concern about data privacy. Recent developments in the data regulation landscape have prompted a seismic shift towards privacy-preserving data analytics, leading to Federated Analytics (FA), the leading paradigm for collaborative data science without data collection. The core principles of this data analysis paradigm allow for breaking the limitations of deriving analytics from limited centralized data in terms of privacy concerns and operational costs. However, FA systems' distributed nature and non-data-sharing enforcement raise critical challenges in the accuracy and efficiency of data analytics. First, skewed data distribution across participating clients leads to severe bias and inconsistency. Second, privacy-preserving techniques applied to the entire FA process leave poor data utility and analysis efficiency. This project innovates accurate, efficient, and credible FA systems with fine-grained knowledge comprehension to optimize the entire lifecycle of the FA process. This project will establish a solid foundation for long-term collaboration with researchers at IBM T. J. Watson Research Center toward developing privacy-preserving FA solutions. Lastly, the project will train a privacy-preserving data science workforce urgently needed in Louisiana.This Research Infrastructure Improvement Track-4 EPSCoR Research Fellows (RII Track-4) proposal would provide a fellowship to an Assistant Professor and training for a graduate student at Louisiana State University. This project aims to investigate fine-grained knowledge comprehension to optimize the entire FA lifecycle holistically, including data skewness estimation, participant selection, and privacy-preserved analysis. Exploring fine-grained knowledge comprehension will offer new insights to better understand and explain data skewness and utility, privacy-preserved data representations, and analytics results in a distributed non-data-sharing scenario. It will catalyze new FA algorithms and system designs toward optimizing FA performance and security. We decouple our specific research activities into two synergistic aims: (1) Improving FA accuracy with fine-grained data skew awareness by data skewness estimation and adaptive refinement of query and client selection; and (2) Optimizing FA utility with fine-grained privacy preservation by separating common and personal feature representations. The proposed FA systems will be extensively evaluated on realistic large-scale testbeds with public datasets at LSU A&M and IBM's Maximo Application Suite and OpenShift Data Science platform. All datasets, benchmarks, and source code will be released on GitHub for a broader impact. By harnessing fine-grained knowledge comprehension for escalating FA efficiency and privacy, the proposed solutions will push the envelope of FA's capabilities and spur the landscape of FA applications in real-world scenarios.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: OAC: Core: Harvesting Idle Resources Safely and Timely for Large-scale AI Applications in High-Performance Computing Systems
-
批准号:2403398
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2024
-
负责人:Hao Wang
-
依托单位:
Collaborative Research: SaTC: CORE: Small: Critical Learning Periods Augmented Robust Federated Learning
-
批准号:2315612
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2023
-
负责人:Hao Wang
-
依托单位:
CRII: OAC: High-Efficiency Serverless Computing Systems for Deep Learning: A Hybrid CPU/GPU Architecture
-
批准号:2153502
-
项目类别:Standard Grant
-
资助金额:$17.49万
-
财政年份:2022
-
负责人:Hao Wang
-
依托单位:
RI: Small: Enabling Interpretable AI via Bayesian Deep Learning
-
批准号:2127918
-
项目类别:Continuing Grant
-
资助金额:$49.99万
-
财政年份:2021
-
负责人:Hao Wang
-
依托单位:
US-China planning visit: Development of High Performance and Multifunctional Infrastructure Material
-
批准号:1338297
-
项目类别:Standard Grant
-
资助金额:$1.28万
-
财政年份:2013
-
负责人:Hao Wang
-
依托单位:
SBIR Phase II: SAFE: Behavior-based Malware Detection and Prevention
-
批准号:0750299
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2008
-
负责人:Hao Wang
-
依托单位:
SBIR Phase I: SpiderWeb - Self-Healing Networks for Spyware Detection
-
批准号:0638170
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2007
-
负责人:Hao Wang
-
依托单位:
Constructibility and Large Cardinal Numbers
-
批准号:7902941
-
项目类别:Standard Grant
-
资助金额:$1.56万
-
财政年份:1979
-
负责人:Hao Wang
-
依托单位:
海外基金