课题基金 / 基金详情

Distributed and Scalable Privacy-Preserving Data Mining Techniques for Big Data

Distributed and Scalable Privacy-Preserving Data Mining Techniques for Big Data
分布式、可扩展的大数据隐私保护数据挖掘技术
批准号:
RGPIN-2014-04520
负责人:
Samet, Saeed
金额:
$1.09万
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2015
资助国家:
加拿大
项目状态:
已结题
起止时间:
2015-01-01 至 2016-12-31

项目摘要

项目成果

Samet, Saeed的其他基金

相似基金

相关文献

中文摘要
翻译
在健康、商业、政府和社交网络等各个领域的大数据时代,保护隐私的重要性是不可能忽视的。随着数据存储能力的巨大增长,计算能力的增强,数据分析的进步,以及连接到互联网和专用网络的设备和传感器数量的大幅增加,隐私和安全风险也随之增加。用于小数据集的典型隐私保护技术,如去身份识别、访问控制、安全计算和数据加密,不能简单地用于大数据。因此,在大数据的有益利用和个人隐私之间建立平衡至关重要。处理大数据研究领域的隐私问题,就像大数据的其他方面,如数据收集、存储、分析和结果发布一样,已经成为一项挑战。 本研究计划提出和开发新的隐私保护技术,并将现有的可扩展和增量的数据挖掘和统计分析方法中的技术进行扩展,以便它们可以实际应用于大数据,同时将应用这些技术对提取的知识的整体性能、准确性和实用性的负面影响降至最低。这项研究计划的发现和成果将应用于2型糖尿病的基因组-环境-关联,以揭示与这种高度常见的疾病相关的基因-环境相互作用,作为概念的证明,并使用真实数据进行测试。 这项研究的长期目标是为大数据上的数据挖掘算法开发可扩展的隐私保护方法和协议。这项研究将侧重于在模拟和真实数据(2型糖尿病)上保护隐私协议的实用方法和技术。这些发现将有助于在卫生、商业和政府领域类似复杂的应用中进行比较。为了利用2型糖尿病数据集,有必要保护个人隐私,同时允许有意义的数据挖掘和计算操作。 结果将扩展安全协议集,以涵盖统计分析和数据挖掘方法,建议的技术将适用于使用大数据的其他卫生、商业和政府领域。因此,短期目标的重点将是提出、设计和实施有效的隐私保护工具,使用适用于模拟和2型糖尿病数据的新的和现有的隐私保护技术。这项研究的具体短期目标是(1)指出大数据的哪些步骤(从数据收集到结果传播)需要隐私保护,以及在哪些地方使用了当前可用的隐私保护技术;(2)确定目前在遗传数据上使用的统计和数据挖掘技术;(3)为这些应用开发隐私保护的数据挖掘和计算程序和算法;以及(4)在模拟大数据和2型糖尿病数据集上测试这些隐私保护算法。
英文摘要
It is impossible to ignore the importance of preserving privacy especially in the era of Big Data in various fields, such as health, business, government and social networks. With the immense growth in the ability to store data, the increased computing power, advances in data analytics, and very large increases in the number of devices and sensors connected to the internet and dedicated networks, there has been an increase in privacy and security risks. Typical privacy-preserving techniques used with small data sets, such as de-identification, access control, secure computation and data encryption cannot be simply used with Big Data. Therefore, it is crucial to create a balance between beneficial uses of Big Data and individual privacy. Dealing with privacy issues in the research area of Big Data, like other aspects of Big Data such as data collection, storage, analysis, and result dissemination, has become a challenge. This research program plans to propose and develop new privacy-preserving techniques and extend the existing ones in data mining and statistical analysis methods that are scalable and incremental, such that they can be practically applied on Big Data, while minimizing the negative effects of applying these techniques on the overall performance, the accuracy and utility of the extracted knowledge. The findings and outputs of this research program will be applied on genome-environment-associations in type 2 diabetes to uncover gene-environment interactions associated with this highly common disease as proof of concept and test using real data. The long-term objective of this research is to develop scalable privacy-preserving methods and protocols for data mining algorithms on Big Data. The research will focus on practical methods and techniques for privacy-preserving protocols on both simulated and real data (type 2 diabetes). The findings will be useful for comparison in similarly complex applications in health, business and government. In order to utilize the type 2 diabetes dataset it will be necessary to preserve the individual’s privacy while allowing meaningful data mining and computational operations. The results will extend the set of secure protocols to cover statistical analysis and data mining methods, and the proposed techniques will be applicable in other areas of health, business, and government where Big Data are used. Therefore, the focus of the short-term objectives will be to propose, design and implement efficient privacy-preserving tools, using new and existing privacy-preserving techniques applied to simulated and type 2 diabetes data. The specific short-term objectives of this research are to (1) indicate which steps (from data gathering to dissemination of results) of Big Data require privacy protection and where currently available privacy-preserving techniques are used; (2) identify the statistical and data-mining techniques that are currently used on genetic data; (3) develop privacy-protected data-mining and computational procedures and algorithms for these applications and (4) test these privacy-protected algorithms on simulated Big Data and type 2 diabetes datasets.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Distributed and Scalable Privacy-Preserving Data Mining Techniques for Big Data
  • 批准号:
    RGPIN-2014-04520
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.09万
  • 财政年份:
    2019
  • 负责人:
    Samet, Saeed
  • 依托单位:
Distributed and Scalable Privacy-Preserving Data Mining Techniques for Big Data
  • 批准号:
    RGPIN-2014-04520
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.09万
  • 财政年份:
    2018
  • 负责人:
    Samet, Saeed
  • 依托单位:
Distributed and Scalable Privacy-Preserving Data Mining Techniques for Big Data
  • 批准号:
    RGPIN-2014-04520
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $0.33万
  • 财政年份:
    2017
  • 负责人:
    Samet, Saeed
  • 依托单位:
Distributed and Scalable Privacy-Preserving Data Mining Techniques for Big Data
  • 批准号:
    RGPIN-2014-04520
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $0.76万
  • 财政年份:
    2017
  • 负责人:
    Samet, Saeed
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis